<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2014.00110</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research Article</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Pan-genome analyses identify lineage- and niche-specific markers of evolution and adaptation in <italic>Epsilonproteobacteria</italic></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhang</surname> <given-names>Ying</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<xref ref-type="author-notes" rid="fn003"><sup>&#x02020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Sievert</surname> <given-names>Stefan M.</given-names></name>
</contrib>
</contrib-group>
<aff><institution>Biology Department, Woods Hole Oceanographic Institution</institution> <country>Woods Hole, MA, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Martin G. Klotz, University of North Carolina at Charlotte, USA</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Barbara J. Campbell, Clemson University, USA; Amrita Pati, DOE Joint Genome Institute, USA</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Ying Zhang, Department of Cell and Molecular Biology, College of the Environment and Life Sciences, University of Rhode Island, 487 CBLS, 120 Flagg Road, Kingston, RI 02881, USA e-mail: <email>yingzhang&#x00040;mail.uri.edu</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Evolutionary and Genomic Microbiology, a section of the journal Frontiers in Microbiology.</p></fn>
<fn fn-type="present-address" id="fn003"><p>&#x02020;Present address: Ying Zhang, Department of Cell and Molecular Biology, College of the Environment and Life Sciences, University of Rhode Island, Kingston, USA</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>03</month>
<year>2014</year>
</pub-date>
<pub-date pub-type="collection">
<year>2014</year>
</pub-date>
<volume>5</volume>
<elocation-id>110</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>11</month>
<year>2013</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>03</month>
<year>2014</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2014 Zhang and Sievert.</copyright-statement>
<copyright-year>2014</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/3.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract><p>The rapidly increasing availability of complete bacterial genomes has created new opportunities for reconstructing bacterial evolution, but it has also highlighted the difficulty to fully understand the genomic and functional variations occurring among different lineages. Using the class <italic>Epsilonproteobacteria</italic> as a case study, we investigated the composition, flexibility, and function of its pan-genomes. Models were constructed to extrapolate the expansion of pan-genomes at three different taxonomic levels. The results show that, for <italic>Epsilonproteobacteria</italic> the seemingly large genome variations among strains of the same species are less noticeable when compared with groups at higher taxonomic ranks, indicating that genome stability is imposed by the potential existence of taxonomic boundaries. The analyses of pan-genomes has also defined a set of universally conserved core genes, based on which a phylogenetic tree was constructed to confirm that thermophilic species from deep-sea hydrothermal vents represent the most ancient lineages of <italic>Epsilonproteobacteria</italic>. Moreover, by comparing the flexible genome of a chemoautotrophic deep-sea vent species to (1) genomes of species belonging to the same genus, but inhabiting different environments, and (2) genomes of other vent species, but belonging to different genera, we were able to delineate the relative importance of lineage-specific versus niche-specific genes. This result not only emphasizes the overall importance of phylogenetic proximity in shaping the variable part of the genome, but also highlights the adaptive functions of niche-specific genes. Overall, by modeling the expansion of pan-genomes and analyzing core and flexible genes, this study provides snapshots on how the complex processes of gene acquisition, conservation, and removal affect the evolution of different species, and contribute to the metabolic diversity and versatility of <italic>Epsilonproteobacteria</italic>.</p></abstract>
<kwd-group>
<kwd>pan-genome</kwd>
<kwd>core genes</kwd>
<kwd>flexible genes</kwd>
<kwd><italic>Epsilonproteobacteria</italic></kwd>
<kwd><italic>Sulfurimonas</italic></kwd>
<kwd><italic>Helicobacter</italic></kwd>
<kwd><italic>Campylobacter</italic></kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="63"/>
<page-count count="13"/>
<word-count count="8708"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="introduction" id="s1">
<title>Introduction</title>
<p>The evolution of bacterial genomes is characterized by a massive amount of insertions, deletions, and rearrangements, permitting the differentiation and adaptation of evolutionarily related lineages into vastly diverse environments (Romero and Palacios, <xref ref-type="bibr" rid="B47">1997</xref>; Cohan, <xref ref-type="bibr" rid="B10">2001</xref>; Mira et al., <xref ref-type="bibr" rid="B32">2002</xref>; Reams and Neidle, <xref ref-type="bibr" rid="B45">2003</xref>). Over the past two decades, the realization of extended pan-genomes in bacterial species has stimulated many discussions about the meaning of species boundaries (Medini et al., <xref ref-type="bibr" rid="B31">2005</xref>). The fact that strains within a single species can have widely different repertoires of genes has highlighted the complexity in identifying bacterial species. Despite the complete sequencing of over 2000 bacterial genomes and the in-depth study of a number of model organisms, a coherent model is still lacking for genome evolution both within and among bacterial species.</p>
<p>In practice, bacteria have been classified using a polyphasic approach combining information from multiple molecular, morphological, and physiological analyses (Rossell&#x000F3;-Mora and Amann, <xref ref-type="bibr" rid="B48">2001</xref>), with molecular methods playing an increasingly important role. The universally conserved 16S ribosomal RNA (rRNA) gene has been widely used to assess the phylogenetic diversity and, by inference, even the functional diversity of microbes in environmental samples. However, inferring function based on the 16S rRNA gene is only possible in selected cases where all members of a 16S-defined clade share the same physiology, e.g., in case of cyanobacteria or certain groups of sulfate-reducing <italic>Deltaproteobacteria</italic>. Further, various studies have demonstrated substantial genomic variations among strains that differ only slightly in 16S rRNA sequences (Coleman et al., <xref ref-type="bibr" rid="B11">2006</xref>; Rasko et al., <xref ref-type="bibr" rid="B42">2008</xref>; Tettelin et al., <xref ref-type="bibr" rid="B57">2008</xref>). Therefore, a genome-scale understanding of how genes evolve and what determines the acquisition and deletion of genes is essential for mapping the complete genetic variations of bacteria, while at the same time assisting in the classification and functional identification of bacterial species.</p>
<p>Traditionally, the term pan-genome has been used to describe the full repertoire of genes found in different strains of a single species (Hanage et al., <xref ref-type="bibr" rid="B19">2005</xref>; Konstantinidis et al., <xref ref-type="bibr" rid="B23">2006</xref>; Read and Ussery, <xref ref-type="bibr" rid="B44">2006</xref>; Lef&#x000E9;bure et al., <xref ref-type="bibr" rid="B25">2010</xref>; Lukjancenko et al., <xref ref-type="bibr" rid="B29">2010</xref>), but more recently this concept has been extended to represent the total genes in any pre-defined group of bacteria or archaea (Polz et al., <xref ref-type="bibr" rid="B40">2013</xref>). Here, we compared the pan-genomes at different taxonomic ranks within the class <italic>Epsilonproteobacteria</italic>. By examining the genomic variations within same species as well as among different species, we investigated how the processes of gene conservation and transfer affect the evolution of pan-genomes and contribute to the functional adaptation of individual species.</p>
<p>The class <italic>Epsilonproteobacteria</italic> is metabolically diverse and contains organisms with different life styles, including both free-living and host-associated species (Campbell et al., <xref ref-type="bibr" rid="B7">2006</xref>). To date, most studies have focused on species that are associated with the human digestive system, such as members of the genera <italic>Helicobacter</italic> and <italic>Campylobacter</italic>, where they exist either asymptomatically or cause diseases like peptic ulcers or gastric cancer (Engberg et al., <xref ref-type="bibr" rid="B15">2000</xref>). The environmental relevance of <italic>Epsilonproteobacteria</italic> had not been recognized until the late 90&#x02032; s (Moyer et al., <xref ref-type="bibr" rid="B34">1995</xref>; Polz and Cavanaugh, <xref ref-type="bibr" rid="B39">1995</xref>; Longnecker and Reysenbach, <xref ref-type="bibr" rid="B26">2001</xref>). As more and more free-living species were identified and isolated, it became clear that <italic>Epsilonproteobacteria</italic> play important roles in the biogeochemical cycling of nitrogen, sulfur, and carbon in various marine and terrestrial environments (Campbell et al., <xref ref-type="bibr" rid="B7">2006</xref>). At deep-sea hydrothermal vents, chemoautotrophic <italic>Epsilonproteobacteria</italic> serve as important primary producers by utilizing the abundantly available geochemical energy sources to assimilate inorganic carbon through a process known as chemosynthesis (Sievert and Vetriani, <xref ref-type="bibr" rid="B52">2012</xref>, and referenes therein). <italic>Epsilonproteobacteria</italic> also play important roles in coastal and open ocean environments characterized by reducing conditions, such as sulfidic sediments, euxinic water columns, and oxygen minimum zones (e.g., Sievert et al., <xref ref-type="bibr" rid="B51">2008</xref>; Labrenz et al., <xref ref-type="bibr" rid="B24">2013</xref>).</p>
<p>The present study aimed at understanding the function and evolution of the pan-genomes of <italic>Epsilonproteobacteria</italic> at three different levels of taxonomy: species, genus, and class. To this end, we compared all published full genomes of <italic>Epsilonproteobacteria</italic> at the time of our study (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>). These genomes represented a wide-ranging set of isolates and provided an excellent opportunity for studying the connections between genome diversity and phenotypic diversity. Specifically, we set out to answer three fundamental questions. First, do the pan-genomes of different taxonomic ranks within <italic>Epsilonproteobacteria</italic> show different rates of expansion? Second, how is the evolution of the class reflected in the core genes of its pan-genome? Finally, how are the adaptive features reflected in flexible genes of a species?</p>
</sec>
<sec sec-type="results" id="s2">
<title>Results</title>
<sec>
<title>Modeling the expansion of <italic>Epsilonproteobacterial</italic> pan-genomes</title>
<p>The expansion of a pan-genome can be examined by plotting the number of genomes considered against the total number of genes observed. The associations on the plot can then be mathematically evaluated through fitting regression models (Tettelin et al., <xref ref-type="bibr" rid="B57">2008</xref>), according to which it can be classified as &#x0201C;open&#x0201D; or &#x0201C;closed&#x0201D; (see Materials and Methods for details). While the size of an open pan-genome would increase unboundedly with the inclusion of new genomes, the size of a closed pan-genome would reach a plateau after a certain number of sample genomes were included. Previously, such analyses have only been performed at the species level (Tettelin et al., <xref ref-type="bibr" rid="B57">2008</xref>). In order to compare the pan-genomes of different taxonomic ranks, we constructed regression models for five groups of <italic>Epsilonproteobacteria</italic> that correspond to three different levels of taxonomy (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>): two groups representing strains of same species (intra-species), two representing species of same genera (intra-genus), and one representing different species of the entire class (intra-class or inter-species). It is worth mentioning that groups of higher taxonomic rank encompassed those of lower rank. For example, the inter-species group (Class) included all the different species in the intra-genus groups (Genus), and the intra-genus groups included a representative strain from each of the intra-species groups (Species). Hence, our analyses compared the extent of pan-genome expansion at all three taxonomic levels.</p>
<p>We used a two-step process to model the expansion of the pan-genomes. First, complete or random permutations were carried out with step-wise additions of new genomes, and a median was taken on the size of pan-genomes after each step. Second, the median counts was extrapolated using two different models: power-law regression (Tettelin et al., <xref ref-type="bibr" rid="B57">2008</xref>) and exponential regression (Tettelin et al., <xref ref-type="bibr" rid="B56">2005</xref>). The resulting extrapolations were normalized by the median genome sizes in their respective sets to assist the visualization and comparison of the fitted curves (Figure <xref ref-type="fig" rid="F1">1</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Modeling epsilonproteobacterial pan-genomes under different taxonomic levels.</bold> The curves show the extrapolated models. The data points present the median size of pan-genomes over complete or random permutations (Materials and Methods), normalized by the median sizes of individual genomes in respective genome sets, and the error bars represented the 25th and 75th percentiles (Materials and Methods). <bold>(A)</bold> Power law models. The curves for <italic>E. coli</italic> (Rasko et al., <xref ref-type="bibr" rid="B42">2008</xref>), <italic>Sar11</italic> (Grote et al., <xref ref-type="bibr" rid="B17">2012</xref>), and <italic>S. agalactiae</italic> (Tettelin et al., <xref ref-type="bibr" rid="B56">2005</xref>) were plotted following published data. <bold>(B)</bold> Comparison of power law models (solid curves) with exponential models (dotted curves). <bold>(C,D)</bold> Evaluating the stability of power-law versus exponential models by varying the number of genomes (data points) included in the extrapolation.</p></caption>
<graphic xlink:href="fmicb-05-00110-g0001.tif"/>
</fig>
<p>The power-law regression showed that all five pan-genomes in our dataset are open, with an average &#x003B3; parameter of 0.29, 0.53, and 0.66, respectively, for intra-species, intra-genus, and inter-species sets (Materials and Methods). The extrapolated curves of intra-species (red), intra-genus (blue), and inter-species (green) pan-genomes followed distinct slopes: while the intra-species curves were the shallowest, the inter-species curves were the steepest. In contrast, the curves within the same category (the two intra-species or the two intra-genus groups) presented similar slopes (Figure <xref ref-type="fig" rid="F1">1A</xref> and Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>). While both intra-species curves (<italic>H. pylori</italic> and <italic>C. jejuni</italic>) were slightly shallower than the pan-genome curves of <italic>E. coli</italic> (Rasko et al., <xref ref-type="bibr" rid="B42">2008</xref>) and SAR11 (Grote et al., <xref ref-type="bibr" rid="B17">2012</xref>) and steeper than that of <italic>Streptococcus agalactiae</italic> (Tettelin et al., <xref ref-type="bibr" rid="B56">2005</xref>), the two intra-genus curves (<italic>Helicobacter</italic> and <italic>Campylobacter</italic>) were much steeper than all evaluated intra-species curves (Figure <xref ref-type="fig" rid="F1">1A</xref>).</p>
<p>The power-law model fit the available data with high <italic>R</italic><sup>2</sup> values of more than 0.98 for all extrapolations, while the exponential model had on average lower <italic>R</italic><sup>2</sup> values, especially when the number of considered genomes is relatively small (Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>). To further evaluate this, we fitted both models to a varying number of data points to monitor how the number of available genomes may affect the accuracy of pan-genome modeling (Figures <xref ref-type="fig" rid="F1">1C,D</xref>). For example, when using a range of 6&#x02013;10 considered genomes in the modeling of <italic>H. pylori</italic> (Figure <xref ref-type="fig" rid="F1">1C</xref>), the exponential curves (dotted lines) are more spread out than the power-law curves (solid lines). Similarly, modeling of the entire <italic>Epsilonproteobacteria</italic> also showed that the power-law model is more stable than the exponential regression model and less influenced by the availability of genomic data (Figure <xref ref-type="fig" rid="F1">1D</xref>). Besides this, results in Figures <xref ref-type="fig" rid="F1">1C</xref>,<xref ref-type="fig" rid="F1">D</xref> also showed that the number of available genomes are far from saturating the power-law model. In other words, new genes are still been discovered with the addition of each genome. Therefore, the diversity of epsilonproteobacterial pan-genomes is yet to be fully explored with additional genomic sequences.</p>
</sec>
<sec>
<title>Classifying the core, flexible, and singleton genes in the pan-genome</title>
<p>The pan-genome of all the examined <italic>Epsilonproteobacteria</italic> contains 16,349 clusters of non-redundant protein coding genes (Materials and Methods). Among them, 289 clusters (1.7%) represented Conserved Single Copy Genes (CSCGs) that appear only once in every examined genome. These CSCG clusters contain 11,271 genes, which account for 16% of the more than 70,000 genes in all analyzed genomes. At the other end of the spectrum are 10,944 clusters (67%) that occur only in a single genome, accounting for about 15% of the total genes. We classified the total genes in the pan-genomes into three sets based on their occurrence: (1) the &#x0201C;core&#x0201D; genes, which are universally conserved in all considered genomes, include both CSCGs and conserved genes of multiple copies per genome, (2) the &#x0201C;singleton&#x0201D; genes are specific to single genomes, and (3) the &#x0201C;flexible&#x0201D; genes are found in more than one, but not all genomes.</p>
<p>The functional distribution of core, flexible, and singleton genes was examined for the intra-genus groups of <italic>Helicobacter</italic>, <italic>Campylobacter</italic>, and <italic>Sulfurimonas</italic>, as well as the inter-species group. Figure <xref ref-type="fig" rid="F2">2</xref> shows the classifications based on the Clusters of Orthologous Groups (COG) database (Tatusov et al., <xref ref-type="bibr" rid="B55">2003</xref>). The genes that cannot be classified into any existing COG clusters were grouped into the &#x0201C;unmapped&#x0201D; category. For all groups, the majority of the singleton genes (56&#x02013;79%) and a large fraction of the flexible genes (38&#x02013;48%) were unmapped or poorly characterized, while only a small fraction of the core genes (13&#x02013;27%) had no clear functional assignment. The core, which represents an indispensable part of all genes in <italic>Epsilonproteobacteria</italic>, contained mainly housekeeping genes that encode the central machinery of a cell, such as translation, protein modification and turnover, replication and repair, cell wall biogenesis, as well as co-enzyme metabolism. In contrast, the flexible and singleton genes that could be mapped to a functional category mainly encoded functions in signal transduction, inorganic ion transport, and energy production.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Functional distribution of the Core, Flexible, and Singleton genes in the pan-genomes of the class <italic>Epsilonproteobacteria</italic> and the genera of <italic>Campylobacter</italic>, <italic>Helicobacter</italic>, and <italic>Sulfurimonas</italic>.</bold> The X-axis was labeled with the one-letter codes that correspond to different functional categories in the COG database (Tatusov et al., <xref ref-type="bibr" rid="B55">2003</xref>), which was further grouped based on the broader categories: I&#x02014;Cellular processes and signaling, II&#x02014;Information storage and processing, III&#x02014;Metabolism, and IV&#x02014;Poorly characterized or unmapped. The right-most bars were unlabeled and indicate the &#x0201C;unmapped&#x0201D; category. The Y-axis indicates the fraction of genes in a COG family out of the total genes in each of the core, flexible and singleton groups.</p></caption>
<graphic xlink:href="fmicb-05-00110-g0002.tif"/>
</fig>
<p>While both considered to be non-core (Medini et al., <xref ref-type="bibr" rid="B31">2005</xref>), the flexible and singleton genes presented slightly different functional distributions. A larger fraction of the flexible genes (66&#x02013;78%) were mapped to existing COG families than the singleton genes (33&#x02013;55%) despite both being poorly characterized. Additionally, a slightly larger fraction of the flexible genes (10&#x02013;14%) encoded functions in energy production and inorganic ion transport than the core (5&#x02013;12%) and singleton (3&#x02013;7%) genes. Therefore, these processes were more likely shared among subgroups of <italic>Epsilonproteobacteria</italic> than being common to all or unique to a single species. In the following two sections, we further investigate the composition and evolution of pan-genomes by analyzing the core and flexible genes. First, phylogeny of the core genes was reconstructed and used to infer the evolutionary history of <italic>Epsilonproteobacteria</italic>. Second, an analysis was carried out to identify lineage-specific and niche-specific signals within the flexible genome of selected species.</p>
</sec>
<sec>
<title>Phylogenomic reconstruction based on core genes</title>
<p>The core genome of <italic>Epsilonproteobacteria</italic> contains 289 CSCGs and 40 conserved genes with multiple copies per genome. Specifically, the CSCGs accounted for 15% of all the genes in an average epsilonproteobacterial genome. The fact that these genes are universally present in a single copy makes them useful markers for inferring the phylogenetic relationships within the class (Figure <xref ref-type="fig" rid="F3">3A</xref>), as well as evaluating the phylogenetic position of the <italic>Epsilonproteobacteria</italic> as a whole (Figure <xref ref-type="fig" rid="F4">4</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Phylogeny of <italic>Epsilonproteobacteria</italic> based on (A) concatenated protein sequences of a subset of the CSCGs (Materials and Methods), and (B) 16S rRNA gene sequences.</bold> The species labels are color-coded to reflect the taxonomic information and environmental conditions of the isolates. Specifically, the free-living species are classified into two groups, Vent (red) versus Non-Vent (blue), depending on whether or not a species was isolated from deep-sea hydrothermal vents. Whereas the host-associated species are classified into three groups, <italic>Helicobacter</italic> (green), <italic>Campylobacter</italic> (dark green), and <italic>Wolinella</italic> (cyan), in accordance with their taxonomic classifications. The phylogenomic tree in <bold>(A)</bold> is marked with colored vertical lines at the right-hand side to highlight taxonomic groups at the species (red lines) and genera (blue lines) levels. The dotted line that crosses the internal branches of <bold>(A)</bold> indicates the proposed cutoff for family-level taxonomic classification of <italic>Epsilonproteobacteria</italic>, and the roots of the five proposed families are identified with arrows.</p></caption>
<graphic xlink:href="fmicb-05-00110-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Phylogenomic reconstruction of a global bacterial tree based on the concatenated protein sequences of 37 markers.</bold> The leaves are collapsed when possible to present taxonomic groups. Bootstrapping values of more than 50% are shown. The <italic>Deltaproteobacteria</italic> is highlighted with bold font. The <italic>Epsilonproteobacteria</italic> clade is expanded to show individual species and genera, and the root of <italic>Epsilonproteobacteria</italic> is indicated with a black arrow.</p></caption>
<graphic xlink:href="fmicb-05-00110-g0004.tif"/>
</fig>
<p>The phylogenomic analysis (Figure <xref ref-type="fig" rid="F3">3A</xref>) is overall consistent with a phylogenetic tree based on 16S rRNA sequences (Figure <xref ref-type="fig" rid="F3">3B</xref>), in which species of the same taxonomic groups were tightly clustered. According to both trees, the deepest lineages within the <italic>Epsilonproteobacteria</italic> are represented by <italic>Nautilia profundicola</italic> (Smith et al., <xref ref-type="bibr" rid="B53">2008</xref>) and <italic>Caminibacter mediatlanticus</italic> (Voordeckers et al., <xref ref-type="bibr" rid="B59">2005</xref>), both of which are moderate thermophiles and obligate anaerobes isolated from deep-sea hydrothermal vents. The other species from either vent or non-vent environments emerged later in evolution, which paralleled the emergence of host-associated species. Despite these similarities, the CSCG-based tree (Figure <xref ref-type="fig" rid="F3">3A</xref>) was overall better resolved compared to the 16S rRNA gene tree (Figure <xref ref-type="fig" rid="F3">3B</xref>), and it indicated slightly different branching patterns of certain clades, such as the precise phylogenetic positions of the genera <italic>Arcobacter</italic> and <italic>Helicobacter</italic>.</p>
<p>Besides assisting in the evolutionary reconstruction within the <italic>Epsilonproteobacteria</italic>, the CSCGs can also help in evaluating the phylogenetic affiliation of this class as a whole (Figure <xref ref-type="fig" rid="F4">4</xref>). We used a set of 37 phylogenomic markers that are universally present in single copies in a set of fully sequenced genomes to build the bacterial tree. These markers included 31 universal genes published in a previous study (Wu and Eisen, <xref ref-type="bibr" rid="B63">2008</xref>), as well as six additional genes that were identified through a global search of epsilonproteobacterial CSCGs using Hidden Markov Models (HMMs). The concatenated protein tree revealed a close affiliation of the <italic>Epsilonproteobacteria</italic> with the <italic>Aquificae</italic>, as well as provided evidence that the <italic>Epsilonproteobacteria</italic> represent a distinct clade that is separated from <italic>Proteobacteria</italic>. The global analyses also positioned <italic>Deltaproteobacteria</italic> into a distal branch. With the exception of <italic>Hippea maritima</italic> (Anderson et al., <xref ref-type="bibr" rid="B4">2011</xref>), all other examined deltaproteobacterial genomes grouped with the phyla <italic>Thermodesulfobacteria</italic> (Anderson et al., <xref ref-type="bibr" rid="B3">2012</xref>; Elkins et al., <xref ref-type="bibr" rid="B14">2013</xref>), <italic>Acidobacteria</italic> (Ward et al., <xref ref-type="bibr" rid="B60">2009</xref>; Challacombe et al., <xref ref-type="bibr" rid="B9">2011</xref>; Rawat et al., <xref ref-type="bibr" rid="B43">2012</xref>), and <italic>Nitrospirae</italic> (L&#x000FC;cker et al., <xref ref-type="bibr" rid="B28">2010</xref>; Fujimura et al., <xref ref-type="bibr" rid="B16">2012</xref>).</p>
</sec>
<sec>
<title>Niche-specific genes at the deep-sea hydrothermal vents</title>
<p>While a subset of the core genes supported the phylogenomic reconstruction of <italic>Epsilonproteobacteria</italic>, they only covered a small fraction of an average genome. The remaining genes (85%) were either shared among a subset (flexible genes), or they only occurred in one genome (singleton genes) of currently sequenced <italic>Epsilonproteobacteria</italic>. Therefore, analyses of these genes can provide a more complete picture of evolution of this class. Since the singleton genes were largely unmapped to COG families (Figure <xref ref-type="fig" rid="F2">2</xref>) and because the definition of singleton genes is subject to change, i.e., what appears to be a singleton today may be reclassified into a multi-gene cluster when new genome data becomes available in the future, we decided to focus our study on the flexible genes.</p>
<p>We hypothesized that the flexible genome carries two types of information: (1) lineage specific signals that trace the evolutionary heritage of a single lineage, and (2) niche specific signals that represent the adaptation of different species to similar environments. We tested this hypothesis by comparing the genome of a free-living species from a particular environment to genomes of species within the same genus, but inhabiting different environments, as well as to genomes of other species inhabiting a similar environment, but belonging to different genera. We chose <italic>Sulfurimonas autotrophica</italic> for our studies, which was the only free-living species that matched our criteria. We compared the genome of <italic>S. autotrophica</italic> with two subsets (Figure <xref ref-type="fig" rid="F5">5</xref>), one containing its phylogenetic neighbors of the same genus that were isolated from different environments (<italic>S. gotlandica</italic>, and <italic>S. denitrificans</italic>), and the other containing organisms with similar habitats, but belonging to different genera (<italic>Sulfurovum</italic> sp. NBC37-1, and <italic>Nitratiruptor</italic> sp. SB155-2) (Figure <xref ref-type="fig" rid="F3">3</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>Lineage-specific and niche-specific genes in <italic>Sulfurimonas autotrophica</italic>. (A)</bold> Mapping of the <italic>Sulfurimonas</italic>-specific and vent-specific genes. The genomic sequence of <italic>S. autotrophica</italic> is represented by the left semicircle. The other four genomes, colored according to their groups (blue for <italic>Sulfurimonas</italic> and red for vent), were shown on the right, with each genome scaled to be 25% of the semicircle. The links in the center connect the genes that are uniquely shared among <italic>Sulfurimonas</italic> (blue) or vent species (red). <bold>(B)</bold> Cartoon diagram of a vent-specific operon with a putative function of a Formate hydrogenlyase complex based on annotations in the SEED database (Overbeek et al., <xref ref-type="bibr" rid="B36">2005</xref>). The homologs were indicated by arrows with same colors and labeled with gene names.</p></caption>
<graphic xlink:href="fmicb-05-00110-g0005.tif"/>
</fig>
<p>Our analyses successfully identified genes that were either uniquely shared among the <italic>Sulfurimonas</italic> species, i.e., genus specific, or among the vent isolates, i.e., habitat specific (Figure <xref ref-type="fig" rid="F5">5</xref>). This confirmed our hypothesis that both lineage-specific and niche-specific evolutionary events were recorded in the flexible genome. Overall, there were almost three times as many genes specifically inherited within the genus of <italic>Sulfurimonas</italic> (195 unique genes, Table <xref ref-type="supplementary-material" rid="SM1">S3</xref>) than shared among the vent isolates (67 unique genes, Table <xref ref-type="supplementary-material" rid="SM1">S4</xref>). Further, the <italic>Sulfurimonas</italic>-specific genes were distributed evenly across the genome of <italic>S. autotrophica</italic>, while the vent-specific genes clustered tightly at certain locations of the genome. The vent-specific gene clusters encode three major functions, including glycogen production and utilization, inorganic ion and efflux transport, and energy production and conversion. We specifically focused on two vent-specific protein complexes related to energy production and conversion (Hyf-like and PYDH complex) as examples to further elucidate the evolution of vent-specific genes in the flexible genome.</p>
<p>The vent-specific <italic>hyf</italic>-like operon (<italic>hyfBCEFGI</italic>) encodes a putative hydrogenase-4 complex (Figure <xref ref-type="fig" rid="F5">5B</xref>). This operon is conserved among all the vent species, while the genomic neighborhood of the operon is completely unrelated between different species (Figure <xref ref-type="fig" rid="F5">5B</xref>). Besides the <italic>hyf</italic>-like operon, all examined genomes of vent-inhabiting <italic>Epsilonproteobacteria</italic> also encode a homolog of formate dehydrogenase H (Fdh-H). This combination resembles the formate hydrogenlyase (FHL-2) complex of <italic>E. coli</italic>, which oxidizes formic acid to carbon dioxide and molecular hydrogen (Andrews et al., <xref ref-type="bibr" rid="B5">1997</xref>). The pyruvate dehydrogenase complex (PYDH) is a multi-enzyme complex composed of three different enzymes: a decarboxylase (E1p), a dihydrolipoamide acyltransferase (E2p), and a dihydrolipoamide dehydrogenase (LPD), among which the E1p carries out the initial step of an enzymatic reaction that converts pyruvate to acetyl-CoA, NADH and CO<sub>2</sub> (Neveling et al., <xref ref-type="bibr" rid="B35">1998</xref>). The E1p can be present in two forms, one containing multiple copies of a single subunit (type-I) and the other containing multiple copies of two subunits, alpha and beta (type-II). These two forms are not evolutionarily related, and either or both forms may be present in the same organism (Schreiner et al., <xref ref-type="bibr" rid="B50">2005</xref>). Our analyses confirmed the presence of PYDH in seven epsilonproteobacterial genomes: two encode the type-I form and belong to the genus <italic>Arcobacter</italic>, while the other five encode the type-II form and are from species inhabiting deep-sea hydrothermal vents (Table <xref ref-type="supplementary-material" rid="SM1">S5</xref>). We constructed phylogenetic trees of the E1p enzymes in order to further investigate the evolution of the two forms of PYDH in <italic>Epsilonproteobacteria</italic> (Figure <xref ref-type="fig" rid="F6">6</xref>). According to the phylogenetic trees, the type-I form, which existed solely in the genus <italic>Arcobacter</italic>, was closely related to the PYDHs encoded in <italic>Gammaproteobacteria</italic> (Figure <xref ref-type="fig" rid="F6">6A</xref>), whereas the type-II form, which existed solely in the vent species, grouped with the enzymes of <italic>Sulfurihydrogenibium azorense</italic>, which belongs to the <italic>Aquificales</italic>, and of <italic>Geobacter</italic> species (Figure <xref ref-type="fig" rid="F6">6B</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Protein phylogeny of the type-I (A) and type-II (B) E1p enzymes of the pyruvate dehydrogenase complex.</bold> The leaves of the tree were labeled with the genome names, and the internal nodes were labeled with bootstrapping values, in which only values of at least 80 were shown. <bold>(A)</bold> The type-I form of <italic>Epsilonproteobacteria</italic> was found within the genus <italic>Arcobacter</italic> (blue). <bold>(B)</bold> Alpha-subunit of the type-II form was used to construct the tree. The type-II form of <italic>Epsilonproteobacteria</italic> was only found in species isolated from deep-sea hydrothermal vents (red).</p></caption>
<graphic xlink:href="fmicb-05-00110-g0006.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s3">
<title>Discussion</title>
<p>The recognition of bacterial pan-genomes has resulted in discussions regarding the bacterial species concept (Hanage et al., <xref ref-type="bibr" rid="B19">2005</xref>; Konstantinidis et al., <xref ref-type="bibr" rid="B23">2006</xref>; Read and Ussery, <xref ref-type="bibr" rid="B44">2006</xref>; Lef&#x000E9;bure et al., <xref ref-type="bibr" rid="B25">2010</xref>; Lukjancenko et al., <xref ref-type="bibr" rid="B29">2010</xref>). A fundamental problem in microbiology today remains how to define bacterial species according to their genomic information, or in other words, how to accurately reconstruct the acquisition, removal, or conservation of genes during divergence and speciation. Here, we approached this problem using a case study of the class <italic>Epsilonproteobacteria</italic> by systematically comparing all available genomes from this class. We quantitatively modeled the expansion of pan-genomes at the intra-species, intra-genus, and inter-species levels using current taxonomic classifications. Our model not only indicates a steady rate of conservation and divergence for epsilonproteobacterial genomes of the same taxonomic rank, but also verified that strains of a same species have much lower levels of genomic variation than the different species (Figure <xref ref-type="fig" rid="F1">1</xref>). Moreover, comparing two different approaches of modeling bacterial pan-genomes, our study suggested the power-law regression provides a better fit and is more stable than the approach based on the exponential regression, especially when the number of available genomes is low (Figures <xref ref-type="fig" rid="F1">1C,D</xref>).</p>
<p>While the modeling of bacterial pan-genomes can be useful for evaluating the overall genomic diversity, the analyses of individual or subgroups of genes can provide a detailed picture of genomic and functional evolution. Typically, a pan-genome is composed of a core and a non-core fraction, and the phylogeny of CSCGs in the core can provide important insights into the evolution of various lineages. We performed evolutionary reconstruction using the concatenated protein sequences encoded by a subset of CSCGs and compared it with a reconstruction based on the 16S rRNA genes (Figure <xref ref-type="fig" rid="F3">3</xref>). In general, both approaches agreed with each other, further corroborating that <italic>Epsilonproteobacteria</italic> evolved from thermophilic and anaerobic species, and subsequently diversified into mesophilic and microaerobic species that colonized non-vent environments, as well as became host-associated mutualists, commensals or pathogens (Campbell et al., <xref ref-type="bibr" rid="B7">2006</xref>). Compared to the current taxonomic classification that features two main families, i.e., <italic>Campylobacteraceae</italic> and <italic>Helicobacteraceae</italic>, the CSCG-bsed phylogenomic reconstruction suggests a new scheme of family-level taxonomy, in which five different families are identified among the analyzed genomes (marked with arrows in the internal nodes of Figure <xref ref-type="fig" rid="F3">3A</xref>). The identification of CSCGs also assisted in the construction of a global bacterial tree by introducing new protein-coding genes to a previous set of globally conserved proteins (Wu and Eisen, <xref ref-type="bibr" rid="B63">2008</xref>). The global tree indicates that the <italic>Epsilon</italic>- and <italic>Deltaproteobacteria</italic> might not belong to the phylum <italic>Proteobacteria</italic>, but form two distinct lineages within the Bacteria (Figure <xref ref-type="fig" rid="F4">4</xref>), providing further evidence for the need to reclassify these two proteobacterial classes. The exact location of the <italic>Epsilon</italic>- and <italic>Deltaproteobacteria</italic> taxa, however, is still uncertain. While the grouping of <italic>Epsilonproteobacteria</italic> with <italic>Aquificae</italic> and the grouping of <italic>Deltaproteobacteria</italic> with <italic>Acidobacteria</italic> is in line with some studies (Wu et al., <xref ref-type="bibr" rid="B62">2009</xref>; L&#x000FC;cker et al., <xref ref-type="bibr" rid="B27">2013</xref>), others have suggested distinct branching patterns for these taxa (Rinke et al., <xref ref-type="bibr" rid="B46">2013</xref>).</p>
<p>The non-core fraction of pan-genomes can be further divided into flexible genes that are shared among a subset of genomes, and singleton genes that are unique to individual genomes. Consistent with previous studies (Mira et al., <xref ref-type="bibr" rid="B33">2010</xref>; Grote et al., <xref ref-type="bibr" rid="B17">2012</xref>), our analysis of functional profiles showed that while the majority of core genes can be classified into known functional categories, the non-core genes are largely unknown (Figure <xref ref-type="fig" rid="F2">2</xref>). Moreover, the functional distribution of flexible genes suggest that the processes of signal transduction, inorganic ion transport, and energy production can be important in driving genomic variations in <italic>Epsilonproteobacteria</italic>. This is different from observations made on SAR11, where the processes of amino acid and carbohydrate transport dominates the flexible genes (Grote et al., <xref ref-type="bibr" rid="B17">2012</xref>). In analyzing the functional profiles of flexible genes, we did not differentiate between genes that are present in the majority of the genomes from those that are present in only a few genomes, but the presence or absence of a gene in a particular genome could potentially provide useful information. This motivated us to perform a detailed study on a free-living species, <italic>S. autotrophica</italic>, to examine the lineage-specific versus niche-specific signals that were shared among its phylogenetic versus environmental neighbors.</p>
<p>The comparison of <italic>S. autotrophica</italic> with organisms that belong to the same genus but live in different environments, or with organisms that share the same environment but belong to different genera, has provided important insights into how the evolution of bacterial genomes is driven by either phylogenetic relatedness or environmental similarities (Figure <xref ref-type="fig" rid="F5">5</xref>). The results confirmed the presence of both lineage-specific and niche-specific signals in the flexible genome. Moreover, the data revealed multiple niche-specific gene clusters that could benefit the adaptation of <italic>S. autotrophica</italic> to the fluctuating and metal-rich environment of deep-sea hydrothermal vents. Detailed analyses of two niche-specific gene clusters that encode Hyf-like hydrogenase and pyruvate dehydrogenase (PYDH) complexes provided insights into the acquisition of new functions by <italic>Epsilonproteobacteria</italic>.</p>
<p>The co-occurrence of a <italic>hyf</italic>-like operon and a homolog of fdh-H in the genomes of vent species suggest the potential existence of a vent-specific formate hydrogenlyase complex (FHL-2). The genomic neighborhood of the <italic>hyf</italic>-like operon was completely unrelated (Figure <xref ref-type="fig" rid="F5">5B</xref>), suggesting that this operon was transferred independently into the vent species, potentially in response to the presence of formate as a substrate and as a way of coping with the changing environment of deep-sea hydrothermal vents. The protein phylogeny based on the E1p subunit of the PYDH complex presented distinct evolutionary paths for the type-I and type-II forms, which evolved independently in the genus <italic>Arcobacter</italic> and the vent species (Figure <xref ref-type="fig" rid="F6">6</xref>). The observed phylogenetic distribution and branching patterns of the two forms of PYDH may be interpreted based on two different evolutionary scenarios. In the first scenario, the type-II form was the more ancestral type that was initially present in the deepest branching lineages of <italic>Epsilonproteobacteria</italic>, i.e., <italic>Nautilia</italic> and <italic>Nitratiruptor</italic>, but was subsequently preserved only in the vent-inhabiting species and was lost in the other species, while the type-I form was acquired by the <italic>Arcobacter</italic> lineage independently via lateral gene transfer. In the second scenario, both types were absent in the common ancestor of <italic>Epsilonproteobacteria</italic> and were acquired by the different lineages due to adaptations that were specific to the vents (type-II form) or to the genus <italic>Arcobacter</italic> (type-I form). Based on the available data, the first scenario requires gene losses in multiple lineages of <italic>Epsilonproteobacteria</italic>. However, it is in line with the observation that the protein phylogeny of the type-II form is congruent with the species phylogeny based on both 16S rRNA and the concatenated core genes. The second scenario appears to be equally or slightly more parsimonious, but it fails to account for the observed phylogeny. At this point, it might not be possible to fully differentiate between the two scenarios due to the limited availability of genomes of free-living epsilonproteobacterial species, underlying the need for sequencing the genomes of other free-living vent and non-vent species to obtain a more comprehensive understanding of the evolution of PYDH and potentially other flexible genes of <italic>Epsilonproteobacteria</italic>.</p>
<p>Overall, the present study provided insights into the composition, flexibility, and function of epsilonproteobacterial pan-genomes and linked these factors with lineage-specific versus niche-specific gene evolution. The present study has benefited greatly from the advances in genomic sequencing and the application of such technologies to increase the breath and depth of genomic samples (Wu et al., <xref ref-type="bibr" rid="B62">2009</xref>). Many of the genomes we analyzed were simply not available a few years ago. The availability of genomic data, however, is still biased toward a small number of well-studied lineages. Today we are still far from obtaining a comprehensive picture on the evolution and adaptation of bacterial genomes. This is to some extent due to the lack of cultivated strains in the under-represented branches. The development of environmental sequencing, such as metagenomics (Schmeisser et al., <xref ref-type="bibr" rid="B49">2007</xref>; Johnson and Slatkin, <xref ref-type="bibr" rid="B21">2009</xref>; Caro-Quintero and Konstantinidis, <xref ref-type="bibr" rid="B8">2011</xref>) and single-cell genomics (Woyke et al., <xref ref-type="bibr" rid="B61">2009</xref>; Marshall et al., <xref ref-type="bibr" rid="B30">2012</xref>; Rinke et al., <xref ref-type="bibr" rid="B46">2013</xref>), has created new frontiers in solving this problem. With more genomes of organisms from a wider range of environments going to be sequenced, we will be able to more accurately quantify the acquisition, deletion, and maintenance of genes during evolution of various bacterial genomes, hence providing improved models to simulate the process of bacterial speciation.</p>
</sec>
<sec sec-type="materials and methods" id="s4">
<title>Materials and methods</title>
<sec>
<title>Genome sequences of <italic>Epsilonproteobacteria</italic></title>
<p>A total of 39 published genomes were collected for this study (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>), the majority of which came from human or animal associated species of clinical relevance, with 16 genomes representing different strains of two widely studied pathogens, <italic>Helicobacter pylori</italic> and <italic>Campylobacter jejuni</italic>. Despite this bias toward clinical species, nine complete genomes in our dataset represented free-living species from multiple genera that were isolated from a variety of environments, including deep-sea hydrothermal vents, coastal marine sediments, oil fields, and salt marshes. Additionally, we also included near-complete draft genomes from two free-living species, <italic>Caminibacter mediatlanticus</italic> TB-2 and <italic>Sulfurimonas gotlandica</italic> GD1, in order to increase the number of genomes from non-pathogenic species.</p>
</sec>
<sec>
<title>Comparative genomics</title>
<p>In order to identify non-redundant gene clusters, we performed pairwise comparison on the collected genomes to identify bi-directional best hits of the encoded proteins. The determination of bi-directional best hits was based on the BLAST software (Altschul et al., <xref ref-type="bibr" rid="B1">1990</xref>, <xref ref-type="bibr" rid="B2">1997</xref>) using three criteria: (1) the <italic>e</italic>-value should be better than 0.001; (2) the alignment should cover at least 70% of the sequences; and (3) the sequence similarity should be better than 30%. Here the sequence similarity cutoff may seem low, but we used it to accommodate the orthologs in distantly related species. We showed that for the majority of the multi-gene clusters we identified, at least 90% of genes in these clusters encode proteins of the same Pfam family assignment, COG family assignment and functional annotation based on exact word matches (Figure <xref ref-type="fig" rid="F7">7</xref>). This is significant considering the many different existing ways to name identical functions. Figure <xref ref-type="fig" rid="F7">7</xref> only presented a lower estimation of the accuracy of our approach in obtaining coherent orthologous clusters.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Histogram of functional consistency within the non-redundant gene clusters.</bold> The consistence score in the x-axis indicates the fraction of genes in a cluster that were assigned to the same function, with 1 being the most consistent and 0 being the least consistent. The y-axis indicates the fraction of multi-gene clusters in the pan-genome that carry a consistence score of a certain range.</p></caption>
<graphic xlink:href="fmicb-05-00110-g0007.tif"/>
</fig>
</sec>
<sec>
<title>Random permutation of the pan-genomes</title>
<p>We performed permutations on five genome sets: one inter-species, two intra-genus, as well as two intra-species sets (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>). At the inter-species level, we examined all of the 25 distinct species in our genome set of <italic>Epsilonproteobacteria</italic>; at the intra-genus level, we looked at the different species within the genera <italic>Helicobacter</italic> (6 species) and <italic>Campylobacter</italic> (6 species), disregarding strain variations over the same species by selecting one representative strain each for <italic>H. pylori</italic> and <italic>C. jejuni</italic>; at the intra-species level, we compared the different strains of <italic>H. pylori</italic> (10 strains) and <italic>C. jejuni</italic> (6 strains). All possible permutations were explored for the intra-species and intra-genus datasets. However, the inter-species dataset contains too many genomes to be fully permuted. Alternatively, we created 25,000 random permutations by randomly selecting genomes one after another over the whole dataset until all genomes has been visited.</p>
<p>Through each permutation, a data vector is produced to record the increased number of unique genes in the pan-genome with the step-wise addition of new genomes. Then, medians and the 25th and 75th percentile values were calculated over all permutations for each data point of the vectors. Finally, these counts were normalized by the median size of genomes, respectively (Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>), so that pan-genomes of different datasets can be compared with one another (Figure <xref ref-type="fig" rid="F1">1</xref>).</p>
</sec>
<sec>
<title>Power law regression model</title>
<p>The power law regression was performed following the approach described in Tettelin et al. (<xref ref-type="bibr" rid="B57">2008</xref>). The regression function <italic>n</italic> &#x0003D; &#x003C3;<italic>N</italic><sup>&#x003B3;</sup> was used to model the median sizes of the pan-genomes generated from all permutations, where <italic>n</italic> is the total number of non-orthologous genes in the pan-genome, <italic>N</italic> is the number of genomes considered, and &#x003C3; and &#x003B3; are free parameters (Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>). When 0 &#x0003C; &#x003B3; &#x0003C; 1, the pan-genome is considered open because it is an unbounded function over the number of genomes. When &#x003B3; &#x0003C; 0, the pan-genome is considered closed since it approaches a constant as more genomes are considered.</p>
</sec>
<sec>
<title>Exponential regression model</title>
<p>Instead of fitting the total number of genes in pan-genomes, the exponential regression model fits the number of new genes per added genome, which were implemented with an exponential decay function <italic>F</italic><sub><italic>s</italic></sub>(<italic>N</italic>) &#x0003D; <italic>K</italic><sub><italic>s</italic></sub> exp[&#x02212;<italic>N</italic>/<italic>T</italic><sub><italic>s</italic></sub>] &#x0002B; <italic>tg</italic>(&#x003B8;), where <italic>F</italic><sub><italic>s</italic></sub>(<italic>N</italic>) is the number of new genes with the addition of each new genome, <italic>tg</italic>(&#x003B8;) is the same number when <italic>N</italic> approaches infinity, and <italic>K</italic><sub><italic>s</italic></sub>, <italic>T</italic><sub><italic>s</italic></sub>, and <italic>tg</italic>(&#x003B8;) are free parameters (Tettelin et al., <xref ref-type="bibr" rid="B56">2005</xref>). Based on the estimated parameters in the exponential decay function, the pan-genomes were modeled with the formula <inline-formula><mml:math id="M1"><mml:mrow><mml:mi>P</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo> <mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>[</mml:mo> <mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>j</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow> <mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>t</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x003B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow> <mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mrow></mml:math></inline-formula>, where <italic>D</italic> is the median number of genes per sequenced genome at each dataset (Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>).</p>
</sec>
<sec>
<title>Phylogenomic reconstruction with concatenated core proteins</title>
<p>We adapted the protocol by Wu et al. for phylogenomic reconstructions (Wu and Eisen, <xref ref-type="bibr" rid="B63">2008</xref>). In a first step, the individual clusters of CSCG-encoded proteins were aligned using MUSCLE (Edgar, <xref ref-type="bibr" rid="B13">2004</xref>), and HMMs were built for each cluster using hmmbuild from the HMMER package (Eddy, <xref ref-type="bibr" rid="B12">2011</xref>). Then, the models were used as queries to search against other genomes and the resulting alignments were trimmed adapting scripts from AMPHORA (Wu and Eisen, <xref ref-type="bibr" rid="B63">2008</xref>). In a next step the trimmed alignments were concatenated with one another into a master alignment, which was further refined using Gblocks (Talavera and Castresana, <xref ref-type="bibr" rid="B54">2007</xref>) to remove the less conserved columns. Finally, the refined master alignment was used as the input for PhyML (Guindon et al., <xref ref-type="bibr" rid="B18">2010</xref>) for phylogenetic reconstruction.</p>
<p>The CSCG tree of <italic>Epsilonproteobacteria</italic> (Figure <xref ref-type="fig" rid="F3">3A</xref>) included six additional draft or complete genomes that were published after our initial steps of data collection. These included <italic>Sulfurospirillum barnesii</italic> SES-3, Uncultured <italic>Sulfuricurvum</italic> sp. RIFRC-1, <italic>Arcobacter butzleri</italic> ED-1 (Toh et al., <xref ref-type="bibr" rid="B58">2011</xref>), <italic>Arcobacter</italic> sp. L (Toh et al., <xref ref-type="bibr" rid="B58">2011</xref>), <italic>Sulfurovum</italic> sp. AR (Park et al., <xref ref-type="bibr" rid="B37">2012</xref>), as well as the single-cell genomes of <italic>Thiovulum</italic> sp. ES (Marshall et al., <xref ref-type="bibr" rid="B30">2012</xref>). To accommodate the incompleteness of draft genomes, we selected a subset of the CSCG-encoded proteins that occurred once in every draft genomes, and used only these as markers for tree construction. As a result, 194 of the CSCG-encoded proteins were used in the above procedure to construct the local phylogeny for <italic>Epsilonproteobacteria</italic>.</p>
<p>The global bacterial phylogeny was constructed with 37 globally conserved single copy markers (Figure <xref ref-type="fig" rid="F4">4</xref>). In addition to the 31 applied in the AMPHORA package (Wu and Eisen, <xref ref-type="bibr" rid="B63">2008</xref>), we identified six additional phylogenetic markers using the HMM of core proteins: DNA gyrase subunit B (gyrB), Tryptophanyl-tRNA synthetase (TrpRS), SSU ribosomal protein S12p (S23e), LSU ribosomal protein L17p, SSU ribosomal protein S4p (S9e), and SSU ribosomal protein S15p (S13e). Among these new marker genes, GyrB (Kasai et al., <xref ref-type="bibr" rid="B22">2000</xref>; Holmes et al., <xref ref-type="bibr" rid="B20">2004</xref>; Peeters and Willems, <xref ref-type="bibr" rid="B38">2011</xref>) and TrpRS (Rajendran et al., <xref ref-type="bibr" rid="B41">2008</xref>) have been used in previous studies to determine the phylogeny of selected taxonomic groups, and the rest are ribosomal proteins.</p>
<p>The global bacterial tree in Figure <xref ref-type="fig" rid="F4">4</xref> was rooted using mid-point rooting. The 16S and CSCG trees in Figure <xref ref-type="fig" rid="F3">3</xref> were rooted based on the relative positions of different epsilonproteobacterial species at the global bacterial tree and using all other bacteria as an outgroup. As indicated with a black arrow in Figure <xref ref-type="fig" rid="F4">4</xref>, the root of <italic>Epsilonproteobacteria</italic> is located between <italic>Nautiliales</italic> and the other examined lineages.</p>
</sec>
<sec>
<title>Protein phylogeny of PYDH</title>
<p>We collected representative sequences of the type-I and type-II forms of PYDH based on annotations in the UniProt database (Bairoch et al., <xref ref-type="bibr" rid="B6">2005</xref>). The protein phylogeny was reconstructed with PhyML (Guindon et al., <xref ref-type="bibr" rid="B18">2010</xref>) using default settings. The tree of type-I form PYDH (Figure <xref ref-type="fig" rid="F6">6A</xref>) is rooted using mid-point rooting, and the tree of type-II form PYDH (Figure <xref ref-type="fig" rid="F6">6B</xref>) is rooted with the <italic>Eukaryotes</italic> as an outgroup.</p>
</sec>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</sec>
</body>
<back>
<ack>
<p>This study was supported by a WHOI postdoctoral scholarship to Ying Zhang and National Science Foundation grant OCE-1136727 to Stefan M. Sievert, and in part by the National Science Foundation EPSCoR Cooperative Agreement &#x00023;EPS-1004057 to Rhode Island. We thank Jesse McNichol and Fran&#x000E7;ois Thomas for suggestions on the manuscript. We thank Lingsheng Dong for technical supports in setting up the phylogenetic reconstructions.</p>
</ack>
<sec sec-type="supplementary material" id="s5">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://www.frontiersin.org/journal/10.3389/fmicb.2014.00110/abstract">http://www.frontiersin.org/journal/10.3389/fmicb.2014.00110/abstract</ext-link></p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Gish</surname> <given-names>W.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <name><surname>Myers</surname> <given-names>E. W.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Basic local alignment search tool</article-title>. <source>J. Mol. Biol</source>. <volume>215</volume>, <fpage>403</fpage>&#x02013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id><pub-id pub-id-type="pmid">2231712</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Madden</surname> <given-names>T. L.</given-names></name> <name><surname>Sch&#x000E4;ffer</surname> <given-names>A. A.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>1997</year>). <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>. <source>Nucleic Acids Res</source>. <volume>25</volume>, <fpage>3389</fpage>&#x02013;<lpage>3402</lpage> <pub-id pub-id-type="pmid">9254694</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>I.</given-names></name> <name><surname>Saunders</surname> <given-names>E.</given-names></name> <name><surname>Lapidus</surname> <given-names>A.</given-names></name> <name><surname>Nolan</surname> <given-names>M.</given-names></name> <name><surname>Lucas</surname> <given-names>S.</given-names></name> <name><surname>Tice</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Complete genome sequence of the thermophilic sulfate-reducing ocean bacterium <italic>Thermodesulfatator indicus</italic> type strain (CIR29812(T))</article-title>. <source>Stand. Genomic Sci</source>. <volume>6</volume>, <fpage>155</fpage>&#x02013;<lpage>164</lpage>. <pub-id pub-id-type="doi">10.4056/sigs.2665915</pub-id><pub-id pub-id-type="pmid">22768359</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>I.</given-names></name> <name><surname>Sikorski</surname> <given-names>J.</given-names></name> <name><surname>Zeytun</surname> <given-names>A.</given-names></name> <name><surname>Nolan</surname> <given-names>M.</given-names></name> <name><surname>Lapidus</surname> <given-names>A.</given-names></name> <name><surname>Lucas</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Complete genome sequence of <italic>Nitratifractor salsuginis</italic> type strain (E9I37-1)</article-title>. <source>Stand. Genomic Sci</source>. <volume>4</volume>, <fpage>322</fpage>&#x02013;<lpage>330</lpage>. <pub-id pub-id-type="doi">10.4056/sigs.1844518</pub-id><pub-id pub-id-type="pmid">21886859</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andrews</surname> <given-names>S.</given-names></name> <name><surname>Berks</surname> <given-names>B.</given-names></name> <name><surname>McClay</surname> <given-names>J.</given-names></name> <name><surname>Ambler</surname> <given-names>A.</given-names></name> <name><surname>Quail</surname> <given-names>M.</given-names></name> <name><surname>Golby</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>1997</year>). <article-title>A 12-cistron <italic>Escherichia coli</italic> operon (hyf) encoding a putative proton-translocating formate hydrogenlyase system</article-title>. <source>Microbiology</source> <volume>143</volume>, <fpage>3633</fpage>&#x02013;<lpage>3647</lpage>. <pub-id pub-id-type="doi">10.1099/00221287-143-11-3633</pub-id><pub-id pub-id-type="pmid">9387241</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bairoch</surname> <given-names>A.</given-names></name> <name><surname>Apweiler</surname> <given-names>R.</given-names></name> <name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Barker</surname> <given-names>W.</given-names></name> <name><surname>Boeckmann</surname> <given-names>B.</given-names></name> <name><surname>Ferro</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>The Universal Protein Resource (UniProt)</article-title>. <source>Nucleic Acids Res</source>. <volume>33</volume>, <fpage>D154</fpage>&#x02013;<lpage>D159</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gki070</pub-id><pub-id pub-id-type="pmid">15608167</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>B. J.</given-names></name> <name><surname>Engel</surname> <given-names>A. S.</given-names></name> <name><surname>Porter</surname> <given-names>M. L.</given-names></name> <name><surname>Takai</surname> <given-names>K.</given-names></name></person-group> (<year>2006</year>). <article-title>The versatile epsilon-proteobacteria: key players in sulphidic habitats</article-title>. <source>Nat. Rev. Microbiol</source>. <volume>4</volume>, <fpage>458</fpage>&#x02013;<lpage>468</lpage>. <pub-id pub-id-type="doi">10.1038/nrmicro1414</pub-id><pub-id pub-id-type="pmid">16652138</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caro-Quintero</surname> <given-names>A.</given-names></name> <name><surname>Konstantinidis</surname> <given-names>K. T.</given-names></name></person-group> (<year>2011</year>). <article-title>Bacterial species may exist, metagenomics reveal</article-title>. <source>Environ. Microbiol</source>. <volume>14</volume>, <fpage>347</fpage>&#x02013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1111/j.1462-2920.2011.02668.x</pub-id><pub-id pub-id-type="pmid">22151572</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Challacombe</surname> <given-names>J.</given-names></name> <name><surname>Eichorst</surname> <given-names>S.</given-names></name> <name><surname>Hauser</surname> <given-names>L.</given-names></name> <name><surname>Land</surname> <given-names>M.</given-names></name> <name><surname>Xie</surname> <given-names>G.</given-names></name> <name><surname>Kuske</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>Biological consequences of ancient gene acquisition and duplication in the large genome of <italic>Candidatus Solibacter usitatus</italic> Ellin6076</article-title>. <source>PLoS ONE</source> <volume>6</volume>: <fpage>e24882</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0024882</pub-id><pub-id pub-id-type="pmid">21949776</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohan</surname> <given-names>F. M.</given-names></name></person-group> (<year>2001</year>). <article-title>Bacterial species and speciation</article-title>. <source>Syst. Biol</source>. <volume>50</volume>, <fpage>513</fpage>&#x02013;<lpage>524</lpage>. <pub-id pub-id-type="doi">10.1080/10635150118398</pub-id><pub-id pub-id-type="pmid">12116650</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coleman</surname> <given-names>M. L.</given-names></name> <name><surname>Sullivan</surname> <given-names>M. B.</given-names></name> <name><surname>Martiny</surname> <given-names>A. C.</given-names></name> <name><surname>Steglich</surname> <given-names>C.</given-names></name> <name><surname>Barry</surname> <given-names>K.</given-names></name> <name><surname>Delong</surname> <given-names>E. F.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>Genomic islands and the ecology and evolution of Prochlorococcus</article-title>. <source>Science</source> <volume>311</volume>, <fpage>1768</fpage>&#x02013;<lpage>1770</lpage>. <pub-id pub-id-type="doi">10.1126/science.1122050</pub-id><pub-id pub-id-type="pmid">16556843</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eddy</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Accelerated profile HMM searches</article-title>. <source>PLoS Comput. Biol</source>. <volume>7</volume>:<fpage>e1002195</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002195</pub-id><pub-id pub-id-type="pmid">22039361</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edgar</surname> <given-names>R. C.</given-names></name></person-group> (<year>2004</year>). <article-title>MUSCLE: multiple sequence alignment with high accuracy and high throughput</article-title>. <source>Nucleic Acids Res</source>. <volume>32</volume>, <fpage>1792</fpage>&#x02013;<lpage>1797</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh340</pub-id><pub-id pub-id-type="pmid">15034147</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elkins</surname> <given-names>J.</given-names></name> <name><surname>Scott</surname> <given-names>H.-B.</given-names></name> <name><surname>Lucas</surname> <given-names>S.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Lapidus</surname> <given-names>A.</given-names></name> <name><surname>Cheng</surname> <given-names>J.-F.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Complete Genome Sequence of the Hyperthermophilic Sulfate-Reducing Bacterium <italic>Thermodesulfobacterium geofontis</italic> OPF15T</article-title>. <source>Genome Announc</source>. <volume>1</volume>, <fpage>e0016213</fpage>. <pub-id pub-id-type="doi">10.1128/genomeA.00162-13</pub-id><pub-id pub-id-type="pmid">23580711</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Engberg</surname> <given-names>J.</given-names></name> <name><surname>On</surname> <given-names>S.</given-names></name> <name><surname>Harrington</surname> <given-names>C.</given-names></name> <name><surname>Gerner-Smidt</surname> <given-names>P.</given-names></name></person-group> (<year>2000</year>). <article-title>Prevalence of Campylobacter, Arcobacter, Helicobacter, and Sutterella spp. in human fecal samples as estimated by a reevaluation of isolation methods for Campylobacters</article-title>. <source>J. Clin. Microbiol</source>. <volume>38</volume>, <fpage>286</fpage>&#x02013;<lpage>291</lpage>. <pub-id pub-id-type="doi">10.1016/j.jinf.2006.10.047</pub-id><pub-id pub-id-type="pmid">10618103</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fujimura</surname> <given-names>R.</given-names></name> <name><surname>Sato</surname> <given-names>Y.</given-names></name> <name><surname>Nishizawa</surname> <given-names>T.</given-names></name> <name><surname>Oshima</surname> <given-names>K.</given-names></name> <name><surname>Kim</surname> <given-names>S.-W.</given-names></name> <name><surname>Hattori</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Complete genome sequence of <italic>Leptospirillum ferrooxidans</italic> strain C2-3, isolated from a fresh volcanic ash deposit on the island of Miyake, Japan</article-title>. <source>J. Bacteriol</source>. <volume>194</volume>, <fpage>4122</fpage>&#x02013;<lpage>4123</lpage>. <pub-id pub-id-type="doi">10.1128/JB.00696-12</pub-id><pub-id pub-id-type="pmid">22815442</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grote</surname> <given-names>J.</given-names></name> <name><surname>Thrash</surname> <given-names>J. C.</given-names></name> <name><surname>Huggett</surname> <given-names>M. J.</given-names></name> <name><surname>Landry</surname> <given-names>Z. C.</given-names></name> <name><surname>Carini</surname> <given-names>P.</given-names></name> <name><surname>Giovannoni</surname> <given-names>S. J.</given-names></name></person-group> (<year>2012</year>). <article-title>Streamlining and core genome conservation among highly divergent members of the SAR11</article-title> <source>Clade</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1128/mBio.00252-12</pub-id><pub-id pub-id-type="pmid">22991429</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guindon</surname> <given-names>S.</given-names></name> <name><surname>Dufayard</surname> <given-names>J.-F.</given-names></name> <name><surname>Lefort</surname> <given-names>V.</given-names></name> <name><surname>Anisimova</surname> <given-names>M.</given-names></name> <name><surname>Hordijk</surname> <given-names>W.</given-names></name> <name><surname>Gascuel</surname> <given-names>O.</given-names></name></person-group> (<year>2010</year>). <article-title>New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0</article-title>. <source>Syst. Biol</source>. <volume>59</volume>, <fpage>307</fpage>&#x02013;<lpage>321</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/syq010</pub-id><pub-id pub-id-type="pmid">20525638</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanage</surname> <given-names>W. P.</given-names></name> <name><surname>Fraser</surname> <given-names>C.</given-names></name> <name><surname>Spratt</surname> <given-names>B. G.</given-names></name></person-group> (<year>2005</year>). <article-title>Fuzzy species among recombinogenic bacteria</article-title>. <source>BMC Biol</source>. <volume>3</volume>:<fpage>6</fpage>. <pub-id pub-id-type="doi">10.1186/1741-7007-3-6</pub-id><pub-id pub-id-type="pmid">15752428</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holmes</surname> <given-names>D.</given-names></name> <name><surname>Nevin</surname> <given-names>K.</given-names></name> <name><surname>Lovley</surname> <given-names>D.</given-names></name></person-group> (<year>2004</year>). <article-title>Comparison of 16S rRNA, nifD, recA, gyrB, rpoB and fusA genes within the family Geobacteraceae fam. nov</article-title>. <source>Int. J. Syst. Evol. Microbiol</source>. <volume>54</volume>, <fpage>1591</fpage>&#x02013;<lpage>1599</lpage>. <pub-id pub-id-type="doi">10.1099/ijs.0.02958-0</pub-id><pub-id pub-id-type="pmid">15388715</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>P.</given-names></name> <name><surname>Slatkin</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Inference of microbial recombination rates from metagenomic data</article-title>. <source>PLoS Genet</source>. <volume>5</volume>:<fpage>e1000674</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1000674</pub-id><pub-id pub-id-type="pmid">19798447</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kasai</surname> <given-names>H.</given-names></name> <name><surname>Ezaki</surname> <given-names>T.</given-names></name> <name><surname>Harayama</surname> <given-names>S.</given-names></name></person-group> (<year>2000</year>). <article-title>Differentiation of phylogenetically related slowly growing mycobacteria by their gyrB sequences</article-title>. <source>J. Clin. Microbiol</source>. <volume>38</volume>, <fpage>301</fpage>&#x02013;<lpage>308</lpage>. <pub-id pub-id-type="pmid">10618105</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Konstantinidis</surname> <given-names>K. T.</given-names></name> <name><surname>Ramette</surname> <given-names>A.</given-names></name> <name><surname>Tiedje</surname> <given-names>J. M.</given-names></name></person-group> (<year>2006</year>). <article-title>The bacterial species definition in the genomic era</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci</source>. <volume>361</volume>, <fpage>1929</fpage>&#x02013;<lpage>1940</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2006.1920</pub-id><pub-id pub-id-type="pmid">17062412</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Labrenz</surname> <given-names>M.</given-names></name> <name><surname>Grote</surname> <given-names>J.</given-names></name> <name><surname>Mammitzsch</surname> <given-names>K.</given-names></name> <name><surname>Boschker</surname> <given-names>H. T.</given-names></name> <name><surname>Laue</surname> <given-names>M.</given-names></name> <name><surname>Jost</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title><italic>Sulfurimonas gotlandica</italic> sp. nov., a chemoautotrophic and psychrotolerant epsilonproteobacterium isolated from a pelagic Baltic Sea redoxcline, and an emended description of the genus Sulfurimonas</article-title>. <source>Int. J. Syst. Evol. Microbiol</source>. <volume>63</volume>, <fpage>4141</fpage>&#x02013;<lpage>4148</lpage>. <pub-id pub-id-type="doi">10.1099/ijs.0.048827-0</pub-id><pub-id pub-id-type="pmid">23749282</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lef&#x000E9;bure</surname> <given-names>T.</given-names></name> <name><surname>Bitar</surname> <given-names>P. D. P.</given-names></name> <name><surname>Suzuki</surname> <given-names>H.</given-names></name> <name><surname>Stanhope</surname> <given-names>M. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Evolutionary dynamics of complete Campylobacter pan-genomes and the bacterial species concept</article-title>. <source>Genome Biol. Evol</source>. <volume>2</volume>, <fpage>646</fpage>&#x02013;<lpage>655</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evq048</pub-id><pub-id pub-id-type="pmid">20688752</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Longnecker</surname> <given-names>K.</given-names></name> <name><surname>Reysenbach</surname> <given-names>A.-L.</given-names></name></person-group> (<year>2001</year>). <article-title>Expansion of the geographic distribution of a novel lineage of epsilon-Proteobacteria to a hydrothermal vent site on the Southern East Pacific Rise</article-title>. <source>FEMS Microbiol. Ecol</source>. <volume>35</volume>, <fpage>287</fpage>&#x02013;<lpage>293</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6941.2001.tb00814.x</pub-id><pub-id pub-id-type="pmid">11311439</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>L&#x000FC;cker</surname> <given-names>S.</given-names></name> <name><surname>Nowka</surname> <given-names>B.</given-names></name> <name><surname>Rattei</surname> <given-names>T.</given-names></name> <name><surname>Spieck</surname> <given-names>E.</given-names></name> <name><surname>Daims</surname> <given-names>H.</given-names></name></person-group> (<year>2013</year>). <article-title>The genome of nitrospina gracilis illuminates the metabolism and evolution of the major marine nitrite oxidizer</article-title>. <source>Front. Microbiol</source>. <volume>4</volume>:<issue>27</issue>. <pub-id pub-id-type="doi">10.3389/fmicb.2013.00027</pub-id><pub-id pub-id-type="pmid">23439773</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>L&#x000FC;cker</surname> <given-names>S.</given-names></name> <name><surname>Wagner</surname> <given-names>M.</given-names></name> <name><surname>Maixner</surname> <given-names>F.</given-names></name> <name><surname>Pelletier</surname> <given-names>E.</given-names></name> <name><surname>Koch</surname> <given-names>H.</given-names></name> <name><surname>Vacherie</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>A Nitrospira metagenome illuminates the physiology and evolution of globally important nitrite-oxidizing bacteria</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>107</volume>, <fpage>13479</fpage>&#x02013;<lpage>13484</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1003860107</pub-id><pub-id pub-id-type="pmid">20624973</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lukjancenko</surname> <given-names>O.</given-names></name> <name><surname>Wassenaar</surname> <given-names>T. M.</given-names></name> <name><surname>Ussery</surname> <given-names>D. W.</given-names></name></person-group> (<year>2010</year>). <article-title>Comparison of 61 sequenced <italic>Escherichia coli</italic> genomes</article-title>. <source>Microb. Ecol</source>. <volume>60</volume>, <fpage>708</fpage>&#x02013;<lpage>720</lpage>. <pub-id pub-id-type="doi">10.1007/s00248-010-9717-3</pub-id><pub-id pub-id-type="pmid">20623278</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marshall</surname> <given-names>I. P. G.</given-names></name> <name><surname>Blainey</surname> <given-names>P. C.</given-names></name> <name><surname>Spormann</surname> <given-names>A. M.</given-names></name> <name><surname>Quake</surname> <given-names>S. R.</given-names></name></person-group> (<year>2012</year>). <article-title>A single-cell genome for Thiovulum sp</article-title>. <source>Appl. Environ. Microbiol</source>. <volume>78</volume>, <fpage>8555</fpage>&#x02013;<lpage>8563</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.02314-12</pub-id><pub-id pub-id-type="pmid">23023751</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Medini</surname> <given-names>D.</given-names></name> <name><surname>Donati</surname> <given-names>C.</given-names></name> <name><surname>Tettelin</surname> <given-names>H.</given-names></name> <name><surname>Masignani</surname> <given-names>V.</given-names></name> <name><surname>Rappuoli</surname> <given-names>R.</given-names></name></person-group> (<year>2005</year>). <article-title>The microbial pan-genome</article-title>. <source>Curr. Opin. Genet. Dev</source>. <volume>15</volume>, <fpage>589</fpage>&#x02013;<lpage>594</lpage>. <pub-id pub-id-type="doi">10.1016/j.gde.2005.09.006</pub-id><pub-id pub-id-type="pmid">16185861</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mira</surname> <given-names>A.</given-names></name> <name><surname>Klasson</surname> <given-names>L.</given-names></name> <name><surname>Andersson</surname> <given-names>S. G. E.</given-names></name></person-group> (<year>2002</year>). <article-title>Microbial genome evolution: sources of variability</article-title>. <source>Curr. Opin. Microbiol</source>. <volume>5</volume>, <fpage>506</fpage>&#x02013;<lpage>512</lpage>. <pub-id pub-id-type="doi">10.1016/S1369-5274(02)00358-2</pub-id><pub-id pub-id-type="pmid">12354559</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mira</surname> <given-names>A.</given-names></name> <name><surname>Mart&#x000ED;n-cuadrado</surname> <given-names>A. B.</given-names></name> <name><surname>Auria</surname> <given-names>G. D.</given-names></name> <name><surname>Rodr&#x000ED;guez-valera</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>The bacterial pan-genome?: a new paradigm in microbiology</article-title>. <source>Int. Microbiol</source>. <volume>13</volume>, <fpage>45</fpage>&#x02013;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.2436/20.1501.01.110</pub-id><pub-id pub-id-type="pmid">20890839</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moyer</surname> <given-names>C.</given-names></name> <name><surname>Dobbs</surname> <given-names>F.</given-names></name> <name><surname>Karl</surname> <given-names>D.</given-names></name></person-group> (<year>1995</year>). <article-title>Phylogenetic diversity of the bacterial community from a microbial mat at an active, hydrothermal vent system, Loihi Seamount, Hawaii</article-title>. <source>Appl. Environ. Microbiol</source>. <volume>61</volume>, <fpage>1555</fpage>&#x02013;<lpage>1562</lpage>. <pub-id pub-id-type="pmid">7538279</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neveling</surname> <given-names>U.</given-names></name> <name><surname>S</surname> <given-names>B.-M.</given-names></name> <name><surname>Sahm</surname> <given-names>H.</given-names></name></person-group> (<year>1998</year>). <article-title>Gene and subunit organization of bacterial pyruvate dehydrogenase complexes</article-title>. <source>Biochim. Biophys. Acta</source> <volume>1385</volume>, <fpage>367</fpage>&#x02013;<lpage>372</lpage>. <pub-id pub-id-type="doi">10.1016/S0167-4838(98)00080-6</pub-id><pub-id pub-id-type="pmid">9655937</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Overbeek</surname> <given-names>R.</given-names></name> <name><surname>Begley</surname> <given-names>T.</given-names></name> <name><surname>Butler</surname> <given-names>R. M.</given-names></name> <name><surname>Choudhuri</surname> <given-names>J. V.</given-names></name> <name><surname>Chuang</surname> <given-names>H.-Y.</given-names></name> <name><surname>Cohoon</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>The subsystems approach to genome annotation and its use in the project to annotate 1000 genomes</article-title>. <source>Nucleic Acids Res</source>. <volume>33</volume>, <fpage>5691</fpage>&#x02013;<lpage>5702</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gki866</pub-id><pub-id pub-id-type="pmid">16214803</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>S.-J.</given-names></name> <name><surname>Ghai</surname> <given-names>R.</given-names></name> <name><surname>Mart&#x000ED;n-Cuadrado</surname> <given-names>A.-B.</given-names></name> <name><surname>Rodr&#x000ED;guez-Valera</surname> <given-names>F.</given-names></name> <name><surname>Jung</surname> <given-names>M.-Y.</given-names></name> <name><surname>Kim</surname> <given-names>J.-G.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Draft genome sequence of the sulfur-oxidizing bacterium &#x0201C;<italic>Candidatus Sulfurovum sediminum</italic>&#x0201D; AR, which belongs to the Epsilonproteobacteria</article-title>. <source>J. Bacteriol</source>. <volume>194</volume>, <fpage>4128</fpage>&#x02013;<lpage>4129</lpage>. <pub-id pub-id-type="doi">10.1128/JB.00741-12</pub-id><pub-id pub-id-type="pmid">22815446</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peeters</surname> <given-names>K.</given-names></name> <name><surname>Willems</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>The gyrB gene is a useful phylogenetic marker for exploring the diversity of Flavobacterium strains isolated from terrestrial and aquatic habitats in Antarctica</article-title>. <source>FEMS Microbiol. Lett</source>. <volume>321</volume>, <fpage>130</fpage>&#x02013;<lpage>140</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6968.2011.02326.x</pub-id><pub-id pub-id-type="pmid">21645050</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Polz</surname> <given-names>M.</given-names></name> <name><surname>Cavanaugh</surname> <given-names>C.</given-names></name></person-group> (<year>1995</year>). <article-title>Dominance of one bacterial phylotype at a Mid-Atlantic Ridge hydrothermal vent site</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>92</volume>, <fpage>7232</fpage>&#x02013;<lpage>7236</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.92.16.7232</pub-id><pub-id pub-id-type="pmid">7543678</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Polz</surname> <given-names>M. F.</given-names></name> <name><surname>Alm</surname> <given-names>E. J.</given-names></name> <name><surname>Hanage</surname> <given-names>W. P.</given-names></name></person-group> (<year>2013</year>). <article-title>Horizontal gene transfer and the evolution of bacterial and archaeal population structure</article-title>. <source>Trends Genet</source>. <volume>29</volume>, <fpage>170</fpage>&#x02013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1016/j.tig.2012.12.006</pub-id><pub-id pub-id-type="pmid">23332119</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rajendran</surname> <given-names>N.</given-names></name> <name><surname>Rajnarayanan</surname> <given-names>R.</given-names></name> <name><surname>Demuth</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Molecular phylogenetic analysis of tryptophanyl-tRNA synthetase of <italic>Actinobacillus actinomycetemcomitans</italic></article-title>. <source>Z. Naturforsch. C</source> <volume>63</volume>, <fpage>418</fpage>&#x02013;<lpage>428</lpage>. <pub-id pub-id-type="pmid">18669030</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rasko</surname> <given-names>D. A.</given-names></name> <name><surname>Rosovitz</surname> <given-names>M. J.</given-names></name> <name><surname>Myers</surname> <given-names>G. S.</given-names></name> <name><surname>Mongodin</surname> <given-names>E. F.</given-names></name> <name><surname>Fricke</surname> <given-names>W. F.</given-names></name> <name><surname>Gajer</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>The pangenome structure of <italic>Escherichia coli</italic>: comparative genomic analysis of <italic>E. coli</italic> commensal and pathogenic isolates</article-title>. <source>J. Bacteriol</source>. <volume>190</volume>, <fpage>6881</fpage>&#x02013;<lpage>6893</lpage>. <pub-id pub-id-type="doi">10.1128/JB.00619-08</pub-id><pub-id pub-id-type="pmid">18676672</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rawat</surname> <given-names>S.</given-names></name> <name><surname>M&#x000E4;nnist&#x000F6;</surname> <given-names>M.</given-names></name> <name><surname>Bromberg</surname> <given-names>Y.</given-names></name> <name><surname>H&#x000E4;ggblom</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Comparative genomic and physiological analysis provides insights into the role of Acidobacteria in organic carbon utilization in Arctic tundra soils</article-title>. <source>FEMS Microbiol. Ecol</source>. <volume>82</volume>, <fpage>341</fpage>&#x02013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6941.2012.01381.x</pub-id><pub-id pub-id-type="pmid">22486608</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Read</surname> <given-names>T. D.</given-names></name> <name><surname>Ussery</surname> <given-names>D. W.</given-names></name></person-group> (<year>2006</year>). <article-title>Opening the pan-genomics box</article-title>. <source>Curr. Opin. Microbiol</source>. <volume>9</volume>, <fpage>496</fpage>&#x02013;<lpage>498</lpage>. <pub-id pub-id-type="doi">10.1016/j.mib.2006.08.010</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reams</surname> <given-names>A. B.</given-names></name> <name><surname>Neidle</surname> <given-names>E. L.</given-names></name></person-group> (<year>2003</year>). <article-title>Genome plasticity in Acinetobacter: new degradative capabilities acquired by the spontaneous amplification of large chromosomal segments</article-title>. <source>Mol. Microbiol</source>. <volume>47</volume>, <fpage>1291</fpage>&#x02013;<lpage>1304</lpage>. <pub-id pub-id-type="doi">10.1046/j.1365-2958.2003.03342.x</pub-id><pub-id pub-id-type="pmid">12603735</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rinke</surname> <given-names>C.</given-names></name> <name><surname>Schwientek</surname> <given-names>P.</given-names></name> <name><surname>Sczyrba</surname> <given-names>A.</given-names></name> <name><surname>Ivanova</surname> <given-names>N.</given-names></name> <name><surname>Anderson</surname> <given-names>I.</given-names></name> <name><surname>Cheng</surname> <given-names>J.-F.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Insights into the phylogeny and coding potential of microbial dark matter</article-title>. <source>Nature</source> <volume>499</volume>, <fpage>431</fpage>&#x02013;<lpage>437</lpage>. <pub-id pub-id-type="doi">10.1038/nature12352</pub-id><pub-id pub-id-type="pmid">23851394</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romero</surname> <given-names>D.</given-names></name> <name><surname>Palacios</surname> <given-names>R.</given-names></name></person-group> (<year>1997</year>). <article-title>Gene amplification and genomic plasticity in prokaryotes</article-title>. <source>Annu. Rev. Genet</source>. <volume>31</volume>, <fpage>91</fpage>&#x02013;<lpage>9111</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.genet.31.1.91</pub-id><pub-id pub-id-type="pmid">9442891</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rossell&#x000F3;-Mora</surname> <given-names>R.</given-names></name> <name><surname>Amann</surname> <given-names>R.</given-names></name></person-group> (<year>2001</year>). <article-title>The species concept for prokaryotes</article-title>. <source>FEMS Microbiol. Rev</source>. <volume>25</volume>, <fpage>39</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6976.2001.tb00571.x</pub-id><pub-id pub-id-type="pmid">11152940</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmeisser</surname> <given-names>C.</given-names></name> <name><surname>Steele</surname> <given-names>H.</given-names></name> <name><surname>Streit</surname> <given-names>W. R.</given-names></name></person-group> (<year>2007</year>). <article-title>Metagenomics, biotechnology with non-culturable microbes</article-title>. <source>Appl. Microbiol. Biotechnol</source>. <volume>75</volume>, <fpage>955</fpage>&#x02013;<lpage>962</lpage>. <pub-id pub-id-type="doi">10.1007/s00253-007-0945-5</pub-id><pub-id pub-id-type="pmid">17396253</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schreiner</surname> <given-names>M.</given-names></name> <name><surname>Fiur</surname> <given-names>D.</given-names></name> <name><surname>Hol&#x000E1;tko</surname> <given-names>J.</given-names></name> <name><surname>P&#x000E1;tek</surname> <given-names>M.</given-names></name> <name><surname>Eikmanns</surname> <given-names>B.</given-names></name></person-group> (<year>2005</year>). <article-title>E1 enzyme of the pyruvate dehydrogenase complex in <italic>Corynebacterium glutamicum</italic>: molecular analysis of the gene and phylogenetic aspects</article-title>. <source>J. Bacteriol</source>. <volume>187</volume>, <fpage>6005</fpage>&#x02013;<lpage>6018</lpage>. <pub-id pub-id-type="doi">10.1128/JB.187.17.6005-6018.2005</pub-id><pub-id pub-id-type="pmid">16109942</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sievert</surname> <given-names>S. M.</given-names></name> <name><surname>Scott</surname> <given-names>K. M.</given-names></name> <name><surname>Klotz</surname> <given-names>M. G.</given-names></name> <name><surname>Chain</surname> <given-names>P. S. G.</given-names></name> <name><surname>Hauser</surname> <given-names>L. J.</given-names></name> <name><surname>Hemp</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Genome of the epsilonproteobacterial chemolithoautotroph <italic>Sulfurimonas denitrificans</italic></article-title>. <source>Appl. Environ. Microbiol</source>. <volume>74</volume>, <fpage>1145</fpage>&#x02013;<lpage>1156</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.01844-07</pub-id><pub-id pub-id-type="pmid">18065616</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sievert</surname> <given-names>S. M.</given-names></name> <name><surname>Vetriani</surname> <given-names>C.</given-names></name></person-group> (<year>2012</year>). <article-title>Chemoautotrophy at deep-sea vents: Past, present, and future</article-title>. <source>Oceanography</source> <volume>25</volume>, <fpage>218</fpage>&#x02013;<lpage>233</lpage>. <pub-id pub-id-type="doi">10.5670/oceanog.2012.21</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>J. L.</given-names></name> <name><surname>Campbell</surname> <given-names>B. J.</given-names></name> <name><surname>Hanson</surname> <given-names>T. E.</given-names></name> <name><surname>Zhang</surname> <given-names>C. L.</given-names></name> <name><surname>Cary</surname> <given-names>S. C.</given-names></name></person-group> (<year>2008</year>). <article-title><italic>Nautilia profundicola</italic> sp. nov., a thermophilic, sulfur-reducing epsilonproteobacterium from deep-sea hydrothermal vents</article-title>. <source>Int. J. Syst. Evol. Microbiol</source>. <volume>58</volume>, <fpage>1598</fpage>&#x02013;<lpage>1602</lpage>. <pub-id pub-id-type="doi">10.1099/ijs.0.65435-0</pub-id><pub-id pub-id-type="pmid">18599701</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Talavera</surname> <given-names>G.</given-names></name> <name><surname>Castresana</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments</article-title>. <source>Syst. Biol</source>. <volume>56</volume>, <fpage>564</fpage>&#x02013;<lpage>577</lpage>. <pub-id pub-id-type="doi">10.1080/10635150701472164</pub-id><pub-id pub-id-type="pmid">17654362</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tatusov</surname> <given-names>R.</given-names></name> <name><surname>Fedorova</surname> <given-names>N.</given-names></name> <name><surname>Jackson</surname> <given-names>J.</given-names></name> <name><surname>Jacobs</surname> <given-names>A.</given-names></name> <name><surname>Kiryutin</surname> <given-names>B.</given-names></name> <name><surname>Koonin</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>The COG database: an updated version includes eukaryotes</article-title>. <source>BMC Bioinform</source>. <volume>4</volume>:<fpage>41</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-4-41</pub-id><pub-id pub-id-type="pmid">12969510</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tettelin</surname> <given-names>H.</given-names></name> <name><surname>Masignani</surname> <given-names>V.</given-names></name> <name><surname>Cieslewicz</surname> <given-names>M. J.</given-names></name> <name><surname>Donati</surname> <given-names>C.</given-names></name> <name><surname>Medini</surname> <given-names>D.</given-names></name> <name><surname>Ward</surname> <given-names>N. L.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Genome analysis of multiple pathogenic isolates of <italic>Streptococcus agalactiae</italic>: implications for the microbial &#x0201C;pan-genome.&#x0201D;</article-title> <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>102</volume>, <fpage>13950</fpage>&#x02013;<lpage>13955</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0506758102</pub-id><pub-id pub-id-type="pmid">16172379</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tettelin</surname> <given-names>H.</given-names></name> <name><surname>Riley</surname> <given-names>D.</given-names></name> <name><surname>Cattuto</surname> <given-names>C.</given-names></name> <name><surname>Medini</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Comparative genomics: the bacterial pan-genome</article-title>. <source>Curr. Opin. Microbiol</source>. <volume>11</volume>, <fpage>472</fpage>&#x02013;<lpage>477</lpage>. <pub-id pub-id-type="doi">10.1016/j.mib.2008.09.006</pub-id><pub-id pub-id-type="pmid">19086349</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toh</surname> <given-names>H.</given-names></name> <name><surname>Sharma</surname> <given-names>V. K.</given-names></name> <name><surname>Oshima</surname> <given-names>K.</given-names></name> <name><surname>Kondo</surname> <given-names>S.</given-names></name> <name><surname>Hattori</surname> <given-names>M.</given-names></name> <name><surname>Ward</surname> <given-names>F. B.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Complete genome sequences of <italic>Arcobacter butzleri</italic> ED-1 and Arcobacter sp. strain L, both isolated from a microbial fuel cell</article-title>. <source>J. Bacteriol</source>. <volume>193</volume>, <fpage>6411</fpage>&#x02013;<lpage>6412</lpage>. <pub-id pub-id-type="doi">10.1128/JB.06084-11</pub-id><pub-id pub-id-type="pmid">22038970</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Voordeckers</surname> <given-names>J. W.</given-names></name> <name><surname>Starovoytov</surname> <given-names>V.</given-names></name> <name><surname>Vetriani</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). <article-title><italic>Caminibacter mediatlanticus</italic> sp. nov., a thermophilic, chemolithoautotrophic, nitrate-ammonifying bacterium isolated from a deep-sea hydrothermal vent on the Mid-Atlantic Ridge</article-title>. <source>Int. J. Syst. Evol. Microbiol</source>. <volume>55</volume>, <fpage>773</fpage>&#x02013;<lpage>779</lpage>. <pub-id pub-id-type="doi">10.1099/ijs.0.63430-0</pub-id><pub-id pub-id-type="pmid">15774661</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ward</surname> <given-names>N.</given-names></name> <name><surname>Challacombe</surname> <given-names>J.</given-names></name> <name><surname>Janssen</surname> <given-names>P.</given-names></name> <name><surname>Henrissat</surname> <given-names>B.</given-names></name> <name><surname>Coutinho</surname> <given-names>P.</given-names></name> <name><surname>Wu</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Three genomes from the phylum Acidobacteria provide insight into the lifestyles of these microorganisms in soils</article-title>. <source>Appl. Environ. Microbiol</source>. <volume>75</volume>, <fpage>2046</fpage>&#x02013;<lpage>2056</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.02294-08</pub-id><pub-id pub-id-type="pmid">19201974</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woyke</surname> <given-names>T.</given-names></name> <name><surname>Xie</surname> <given-names>G.</given-names></name> <name><surname>Copeland</surname> <given-names>A.</given-names></name> <name><surname>Gonz&#x000E1;lez</surname> <given-names>J. M.</given-names></name> <name><surname>Han</surname> <given-names>C.</given-names></name> <name><surname>Kiss</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Assembling the marine metagenome, one cell at a time</article-title>. <source>PLoS ONE</source> <volume>4</volume>:<fpage>e5299</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0005299</pub-id><pub-id pub-id-type="pmid">19390573</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>D.</given-names></name> <name><surname>Hugenholtz</surname> <given-names>P.</given-names></name> <name><surname>Mavromatis</surname> <given-names>K.</given-names></name> <name><surname>Pukall</surname> <given-names>R.</given-names></name> <name><surname>Dalin</surname> <given-names>E.</given-names></name> <name><surname>Ivanova</surname> <given-names>N. N.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>A phylogeny-driven genomic encyclopaedia of Bacteria and Archaea</article-title>. <source>Nature</source> <volume>462</volume>, <fpage>1056</fpage>&#x02013;<lpage>1060</lpage>. <pub-id pub-id-type="doi">10.1038/nature08656</pub-id><pub-id pub-id-type="pmid">20033048</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>M.</given-names></name> <name><surname>Eisen</surname> <given-names>J. A.</given-names></name></person-group> (<year>2008</year>). <article-title>A simple, fast, and accurate method of phylogenomic inference</article-title>. <source>Genome Biol</source>. <volume>9</volume>, <fpage>R151</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2008-9-10-r151</pub-id><pub-id pub-id-type="pmid">18851752</pub-id></citation>
</ref>
</ref-list>
</back>
</article>