<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Cell. Infect. Microbiol.</journal-id>
<journal-title>Frontiers in Cellular and Infection Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Cell. Infect. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">2235-2988</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcimb.2017.00028</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title><italic>In silico</italic> Comparison of 19 <italic>Porphyromonas gingivalis</italic> Strains in Genomics, Phylogenetics, Phylogenomics and Functional Genomics</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Chen</surname> <given-names>Tsute</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/378818/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Siddiqui</surname> <given-names>Huma</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/363810/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Olsen</surname> <given-names>Ingar</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/363500/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Microbiology, The Forsyth Institute</institution> <country>Cambridge, MA, USA</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Oral Biology, University of Oslo</institution> <country>Oslo, Norway</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jan Potempa, University of Louisville, USA</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Gena D. Tribble, University of Texas Health Science Center at Houston, USA; Deborah R. Yoder-Himes, University of Louisville, USA</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Ingar Olsen <email>ingar.olsen&#x00040;odont.uio.no</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>02</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>7</volume>
<elocation-id>28</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>10</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>01</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Chen, Siddiqui and Olsen.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Chen, Siddiqui and Olsen</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Currently, genome sequences of a total of 19 <italic>Porphyromonas gingivalis</italic> strains are available, including eight completed genomes (strains W83, ATCC 33277, TDC60, HG66, A7436, AJW4, 381, and A7A1-28) and 11 high-coverage draft sequences (JCVI SC001, F0185, F0566, F0568, F0569, F0570, SJD2, W4087, W50, Ando, and MP4-504) that are assembled into fewer than 300 contigs. The objective was to compare these genomes at both nucleotide and protein sequence levels in order to understand their phylogenetic and functional relatedness. Four copies of <italic>16S rRNA</italic> gene sequences were identified in each of the eight complete genomes and one in the other 11 unfinished genomes. These 43 <italic>16S rRNA</italic> sequences represent only 24 unique sequences and the derived phylogenetic tree suggests a possible evolutionary history for these strains. Phylogenomic comparison based on shared proteins and whole genome nucleotide sequences consistently showed two groups with closely related members: one consisted of ATCC 33277, 381, and HG66, another of W83, W50, and A7436. At least 1,037 core/shared proteins were identified in the 19 <italic>P. gingivalis</italic> genomes based on the most stringent detecting parameters. Comparative functional genomics based on genome-wide comparisons between NCBI and RAST annotations, as well as additional approaches, revealed functions that are unique or missing in individual <italic>P. gingivalis</italic> strains, or species-specific in all <italic>P. gingivalis</italic> strains, when compared to a neighboring species <italic>P. asaccharolytica</italic>. All the comparative results of this study are available online for download at <ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425/">ftp://www.homd.org/publication_data/20160425/</ext-link>.</p></abstract>
<kwd-group>
<kwd>comparative genomics</kwd>
<kwd>phylogenetics</kwd>
<kwd>phylogenomics</kwd>
<kwd>Porphyromonas gingivalis</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="9"/>
<equation-count count="0"/>
<ref-count count="56"/>
<page-count count="24"/>
<word-count count="17240"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>The Gram-negative anaerobic rod-shaped bacterium <italic>Porphyromonas gingivalis</italic> is one of the most important pathogens in chronic adult periodontitis (Socransky et al., <xref ref-type="bibr" rid="B48">1998</xref>; Darveau et al., <xref ref-type="bibr" rid="B8">2012</xref>; Hajishengallis et al., <xref ref-type="bibr" rid="B16">2012</xref>). Colonization with <italic>P. gingivalis</italic> is also associated with some systemic diseases, including cardiovascular diseases, rheumatoid arthritis, and Alzheimer&#x00027;s disease (Demmer and Desvarieux, <xref ref-type="bibr" rid="B10">2006</xref>; Lundberg et al., <xref ref-type="bibr" rid="B29">2010</xref>; Olsen and Singhrao, <xref ref-type="bibr" rid="B38">2015</xref>). It has become increasingly clear that strains of <italic>P. gingivalis</italic> differ in their pathogenicity and their ability to invade tissues and cells varies as much as three orders of magnitude (Dorn et al., <xref ref-type="bibr" rid="B12">2000</xref>; Lundberg et al., <xref ref-type="bibr" rid="B29">2010</xref>; Dolgilevich et al., <xref ref-type="bibr" rid="B11">2011</xref>; Olsen and Progulske-Fox, <xref ref-type="bibr" rid="B37">2015</xref>). Thus, W83 is considered a virulent strain while ATCC 33277 is less virulent. The AJW4 strain had the lowest invasion ability of 27 strains tested (Dolgilevich et al., <xref ref-type="bibr" rid="B11">2011</xref>).</p>
<p>A comparative genomics study focusing on differences that affect virulence in a mouse model identified over 150 divergent genes (Chen et al., <xref ref-type="bibr" rid="B7">2004</xref>). Dolgilevich et al. (<xref ref-type="bibr" rid="B11">2011</xref>) suggested deficiency in multiple genes as a basis for the <italic>P. gingivalis</italic> non-invasive phenotype. Actually, more than 100 genes were missing from the genome of a non-invading strain. The interstrain genomic polymorphisms and the individual host response have been suggested to be the key to disease initiation and progression (Dolgilevich et al., <xref ref-type="bibr" rid="B11">2011</xref>). Genomic arrangement may also play a key role in the difference in virulence. For example, Naito et al. (<xref ref-type="bibr" rid="B35">2008</xref>) found that although the genome size and GC content were almost the same in strain ATCC 33277 and W83 there were extensive rearrangements between the two strains. <italic>P. gingivalis</italic> has been suggested to harbor many genetic mobile elements such as insertion sequence (IS), miniature inverted-repeat transposable element (MITE) and conjugative transposons CTns (Duncan, <xref ref-type="bibr" rid="B13">2003</xref>; Naito et al., <xref ref-type="bibr" rid="B35">2008</xref>; Tribble et al., <xref ref-type="bibr" rid="B52">2013</xref>; Klein et al., <xref ref-type="bibr" rid="B25">2015</xref>).</p>
<p>Together they are responsible for the fluidic genomic structure of this species (Naito et al., <xref ref-type="bibr" rid="B35">2008</xref>; Tribble et al., <xref ref-type="bibr" rid="B52">2013</xref>). The structural changes of the <italic>P. gingivalis</italic> genomes caused by these elements might have generated many strain-specific protein-coding sequences (CDs) and may have resulted in differences in various phenotypes including important virulence factors (Naito et al., <xref ref-type="bibr" rid="B35">2008</xref>).</p>
<p>To date, a total of 19 <italic>P. gingivalis</italic> genome sequences have been published including eight completed (strains W83, ATCC 33277, TDC60, HG66, A7436, AJW4, 381, and A7A1-28); and 11 high-coverage draft sequences (JCVI SC001, F0185, F0566, F0568, F0569, F0570, SJD2, W4087, W50, Ando, and MP4-504) that are assembled into fewer than 300 contigs. These strains were isolated from various sources including the well-studied laboratory cultures with different degree of virulence, clinical samples from patients with different disease states, as well as an environmental strain isolated from a hospital bathroom sink drain. Together these sequences provide a great opportunity for a comparative genomics study and the results will provide valuable information to better understand the disease mechanism of this important periodontal pathogen. The aim of this study was to conduct <italic>in-silico</italic> genomics comparison for theses genomes using various approaches in the areas of phylogenetics, phylogenomics, and functional genomics. Results that we found most important and interesting are presented in this paper whereas complete results derived from this study are also made available for download online for further investigation.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and methods</title>
<sec>
<title>Sequence sources</title>
<p>Genomic sequences used in this study were downloaded from the NCBI FTP site (<ext-link ext-link-type="uri" xlink:href="ftp://ftp.ncbi.nlm.nih.gov/genomes/all">ftp://ftp.ncbi.nlm.nih.gov/genomes/all</ext-link>). The versions that were downloaded are also available online at <ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425/">ftp://www.homd.org/publication_data/20160425</ext-link>. A summary of all the meta information for each genome is available in the Excel file PG_Genome_Summary.xlsx in the above FTP folder. This file lists all the detail information that are provided by NCBI, such as methods for sequencing, assembling and annotation, as well as various IDs for the same genome including GenBank Accession, GenBank Assembly Accession, Refseq Accession, Refseq Assembly Accession. Table <xref ref-type="table" rid="T1">1</xref> lists the basic information and sources of the sequence data of the 19 <italic>P. gingivalis</italic> genomes analyzed in this report.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Summary of all the <italic>P. gingivalis</italic> genome sequences compared in this report<xref ref-type="table-fn" rid="TN1"><sup>a</sup></xref></bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Strain</bold></th>
<th valign="top" align="center"><bold>Sequence release date<xref ref-type="table-fn" rid="TN2"><sup>b</sup></xref></bold></th>
<th valign="top" align="center"><bold>Genome size (bps)</bold></th>
<th valign="top" align="center"><bold>Contigs</bold></th>
<th valign="top" align="left"><bold>GenBank accession</bold></th>
<th valign="top" align="left"><bold>Bioproject</bold></th>
<th valign="top" align="left"><bold>Biosample<xref ref-type="table-fn" rid="TN3"><sup>c</sup></xref></bold></th>
<th valign="top" align="left"><bold>Submitter</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">W83</td>
<td valign="top" align="center">2003-09-02</td>
<td valign="top" align="center">2,343,476</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AE015924">AE015924</ext-link></td>
<td valign="top" align="left">PRJNA48</td>
<td valign="top" align="left">SAMN02603720</td>
<td valign="top" align="left"><italic>Porphyromonas gingivalis</italic> Genome Project</td>
</tr>
<tr>
<td valign="top" align="left">ATCC_33277</td>
<td valign="top" align="center">2008-05-20</td>
<td valign="top" align="center">2,354,886</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AP009380">AP009380</ext-link></td>
<td valign="top" align="left">PRJDA19051</td>
<td/>
<td valign="top" align="left">Kitasato Univ.</td>
</tr>
<tr>
<td valign="top" align="left">TDC60</td>
<td valign="top" align="center">2011-05-23</td>
<td valign="top" align="center">2,339,898</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AP012203">AP012203</ext-link></td>
<td valign="top" align="left">PRJDA66755</td>
<td/>
<td valign="top" align="left">Tokyo Medical and Dental Univ.</td>
</tr>
<tr>
<td valign="top" align="left">W50</td>
<td valign="top" align="center">2012-06-25</td>
<td valign="top" align="center">2,242,062</td>
<td valign="top" align="center">104</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AJZS01000000">AJZS01000000</ext-link></td>
<td valign="top" align="left">PRJNA78905</td>
<td valign="top" align="left">SAMN00792205</td>
<td valign="top" align="left">J. Craig Venter Institute</td>
</tr>
<tr>
<td valign="top" align="left">JCVI_SC001</td>
<td valign="top" align="center">2013-04-24</td>
<td valign="top" align="center">2,426,396</td>
<td valign="top" align="center">1,284</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CM001843">CM001843</ext-link><xref ref-type="table-fn" rid="TN4"><sup>d</sup></xref> <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="APMB01000000">APMB01000000</ext-link></td>
<td valign="top" align="left">PRJNA167667</td>
<td valign="top" align="left">SAMN02436407</td>
<td valign="top" align="left">J. Craig Venter Institute</td>
</tr>
<tr>
<td valign="top" align="left">F0568</td>
<td valign="top" align="center">2013-09-16</td>
<td valign="top" align="center">2,334,744</td>
<td valign="top" align="center">154</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AWUU01000000">AWUU01000000</ext-link></td>
<td valign="top" align="left">PRJNA173937</td>
<td valign="top" align="left">SAMN02436723</td>
<td valign="top" align="left">Washington Univ.</td>
</tr>
<tr>
<td valign="top" align="left">F0569</td>
<td valign="top" align="center">2013-09-16</td>
<td valign="top" align="center">2,249,227</td>
<td valign="top" align="center">111</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AWUV01000000">AWUV01000000</ext-link></td>
<td valign="top" align="left">PRJNA173938</td>
<td valign="top" align="left">SAMN02436724</td>
<td valign="top" align="left">Washington Univ.</td>
</tr>
<tr>
<td valign="top" align="left">F0570</td>
<td valign="top" align="center">2013-09-16</td>
<td valign="top" align="center">2,282,791</td>
<td valign="top" align="center">117</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AWUW01000000">AWUW01000000</ext-link></td>
<td valign="top" align="left">PRJNA173939</td>
<td valign="top" align="left">SAMN02436747</td>
<td valign="top" align="left">Washington Univ.</td>
</tr>
<tr>
<td valign="top" align="left">F0185</td>
<td valign="top" align="center">2013-09-16</td>
<td valign="top" align="center">2,246,368</td>
<td valign="top" align="center">113</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AWVC01000000">AWVC01000000</ext-link></td>
<td valign="top" align="left">PRJNA198891</td>
<td valign="top" align="left">SAMN02436815</td>
<td valign="top" align="left">Washington Univ.</td>
</tr>
<tr>
<td valign="top" align="left">F0566</td>
<td valign="top" align="center">2013-09-16</td>
<td valign="top" align="center">2,306,092</td>
<td valign="top" align="center">192</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AWVD01000000">AWVD01000000</ext-link></td>
<td valign="top" align="left">PRJNA198892</td>
<td valign="top" align="left">SAMN02436881</td>
<td valign="top" align="left">Washington Univ.</td>
</tr>
<tr>
<td valign="top" align="left">W4087</td>
<td valign="top" align="center">2013-09-16</td>
<td valign="top" align="center">2,216,597</td>
<td valign="top" align="center">114</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AWVE01000000">AWVE01000000</ext-link></td>
<td valign="top" align="left">PRJNA198893</td>
<td valign="top" align="left">SAMN02436749</td>
<td valign="top" align="left">Washington Univ.</td>
</tr>
<tr>
<td valign="top" align="left">SJD2</td>
<td valign="top" align="center">2013-12-04</td>
<td valign="top" align="center">2,329,548</td>
<td valign="top" align="center">117</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="ASYL01000000">ASYL01000000</ext-link></td>
<td valign="top" align="left">PRJNA205615</td>
<td valign="top" align="left">SAMN02470968</td>
<td valign="top" align="left">Shanghai Jiao Tong Univ. School of Medicine</td>
</tr>
<tr>
<td valign="top" align="left">HG66</td>
<td valign="top" align="center">2014-08-14</td>
<td valign="top" align="center">2,441,780</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP007756">CP007756</ext-link></td>
<td valign="top" align="left">PRJNA245225</td>
<td valign="top" align="left">SAMN02732406</td>
<td valign="top" align="left">Univ. of Louisville</td>
</tr>
<tr>
<td valign="top" align="left">A7436</td>
<td valign="top" align="center">2015-08-11</td>
<td valign="top" align="center">2,367,029</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP011995">CP011995</ext-link></td>
<td valign="top" align="left">PRJNA276132</td>
<td valign="top" align="left">SAMN03366764</td>
<td valign="top" align="left">Univ. of Florida</td>
</tr>
<tr>
<td valign="top" align="left">AJW4</td>
<td valign="top" align="center">2015-08-26</td>
<td valign="top" align="center">2,372,492</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP011996">CP011996</ext-link></td>
<td valign="top" align="left">PRJNA276132</td>
<td valign="top" align="left">SAMN03372093</td>
<td valign="top" align="left">Univ. of Florida</td>
</tr>
<tr>
<td valign="top" align="left">Ando</td>
<td valign="top" align="center">2015-09-17</td>
<td valign="top" align="center">2,229,994</td>
<td valign="top" align="center">112</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="BCBV01000000">BCBV01000000</ext-link></td>
<td valign="top" align="left">PRJDB4201</td>
<td valign="top" align="left">SAMD00040429</td>
<td valign="top" align="left">Lab. of Plant Genomics and Genetics, Dept. of Plant Genome Research, Kazusa DNA Research Institute</td>
</tr>
<tr>
<td valign="top" align="left">381</td>
<td valign="top" align="center">2015-10-14</td>
<td valign="top" align="center">2,378,872</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP012889">CP012889</ext-link></td>
<td valign="top" align="left">PRJNA276132</td>
<td valign="top" align="left">SAMN03656156</td>
<td valign="top" align="left">Univ. of Florida</td>
</tr>
<tr>
<td valign="top" align="left">A7A1-28</td>
<td valign="top" align="center">2015-11-17</td>
<td valign="top" align="center">2,249,024</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP013131">CP013131</ext-link></td>
<td valign="top" align="left">PRJNA276132</td>
<td valign="top" align="left">SAMN03653671</td>
<td valign="top" align="left">Univ. of Florida</td>
</tr>
<tr>
<td valign="top" align="left">MP4-504</td>
<td valign="top" align="center">2016-02-09</td>
<td valign="top" align="center">2,373,453</td>
<td valign="top" align="center">92</td>
<td valign="top" align="left"><ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="LOEL01000000">LOEL01000000</ext-link></td>
<td valign="top" align="left">PRJNA305025</td>
<td valign="top" align="left">SAMN04309157</td>
<td valign="top" align="left">Univ. of Washington</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN1"><label>a</label><p><italic>For a more detailed list of this table please follow this web link: <ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425/">ftp://www.homd.org/publication_data/20160425/</ext-link></italic>.</p></fn>
<fn id="TN2"><label>b</label><p><italic>Genomes of this table are sorted by the original sequence release date</italic>.</p></fn>
<fn id="TN3"><label>c</label><p><italic>Unassembled raw sequence reads from which the assembly that was done can be traced back by the Biosample ID, if available</italic>.</p></fn>
<fn id="TN4"><label>d</label><p><italic>This Genbank number shows the sequence as &#x0201C;circular,&#x0201D; however it is a single pseudo-contig with many Ns filling the gaps. Thus, it should not be considered as a complete genome</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Strain information</title>
<sec>
<title>W83, ATCC 33277, and W50</title>
<p>These most-studied laboratory cultures were among the first <italic>P. gingivalis</italic> strains sequenced. Strain W83 was isolated in the 1950s by H. Werner (Bonn, Germany) from an undocumented human oral infection and was brought to The Pasteur Institute by Madeleine Sebald during the 1960s. It was subsequently obtained by Christian Mouton (Quebec, Canada) during the late 1970s. W83 was reported to be also known as strain HG66 (Nelson et al., <xref ref-type="bibr" rid="B36">2003</xref>), however it has been demonstrated that the two are very different strains based on data shown in this report. Strain W50 was originally isolated from a clinical specimen by H. Werner and first studied for known virulence (Marsh et al., <xref ref-type="bibr" rid="B32">1994</xref>). W50 is also known as ATCC 53978 based on the description of the BioSample ID SAMN00792205 (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/?term=SAMN00792205">http://www.ncbi.nlm.nih.gov/biosample/?term=SAMN00792205</ext-link>). The strain ATCC 33277 used for genomic sequencing was directly obtained from the American Type Culture Collection (ATCC) and was described as &#x0201C;has been kept for more than 20 years&#x0201D; by the authors (Naito et al., <xref ref-type="bibr" rid="B35">2008</xref>).</p>
</sec>
<sec>
<title>TDC60</title>
<p>This strain was isolated from a severe periodontal lesion at Tokyo Dental College in Japan. Strain TDC60 exhibited higher pathogenicity in causing abscesses in mice than strains W83 and ATCC 33277 and other strains tested in the college (Watanabe et al., <xref ref-type="bibr" rid="B55">2011</xref>).</p>
</sec>
<sec>
<title>JCVI SC001</title>
<p>This strain was not isolated from the human oral cavity; instead the genomic sequence was derived from single cells found in the biofilm of a hospital bathroom sink drain. The sequence was the first report of a human pathogen sequenced based on a single-cell genomic sequencing approach by capturing DNA from a complex environmental sample outside of the human host (McLean et al., <xref ref-type="bibr" rid="B33">2013</xref>). An automated platform was used to generate genomic DNA by the multiple displacement amplification (MDA) technique from hundreds of single cells in parallel. Thus, the bacterial culture or DNA source of the genomic sequence obtained through MDA cannot be made available (Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/SAMN02436407">http://www.ncbi.nlm.nih.gov/biosample/SAMN02436407</ext-link>, also see reference (McLean et al., <xref ref-type="bibr" rid="B33">2013</xref>).</p>
</sec>
<sec>
<title>Strains sequenced by HMP</title>
<p>A total of six strains (F0185, F0566, F0568, F0569, F0570, and W4087) were sequenced by The Genome Institute of Washington University collaborated with the Data Analysis and Coordination Center (DACC) of the Human Microbiome Project (HMP) and the Human Oral Microbiome Database and were funded by a consortium of institutes including the National Human Genome Research Institute (NHGRI)/National Institutes of Health (NIH), and the National Institute of Dental and Craniofacial Research (NIDCR). Strain F0568 and F0569 were isolated in the 1980s in the USA from the subgingival plaque biofilm of black, non-Hispanic male subjects (53 and 39 years old respectively) diagnosed with moderate periodontitis. F0570 was isolated in the 1980s in the USA from a 39 years old non-Hispanic white male diagnosed with moderate periodontitis. Strain F0185, F0566, and W4087 were reported to be isolated from the oral cavity/mouth of human subjects. Information source: GenBank records in Table <xref ref-type="table" rid="T1">1</xref>.</p>
</sec>
<sec>
<title>SJD2</title>
<p>This strain was isolated from subgingival plaque of a patient in China with chronic periodontitis. It was shown to have high virulent properties comparable with those of the strain W83 in a mouse abscess model (Liu et al., <xref ref-type="bibr" rid="B28">2014</xref>). It was reported to have a higher number of SJD2-specific genes which suggests that strains isolated from a periodontal pocket of Chinese patients with chronic periodontitis may have distinct genes (Liu et al., <xref ref-type="bibr" rid="B28">2014</xref>).</p>
</sec>
<sec>
<title>HG66</title>
<p>HG66 (also known as DSM 28984) was isolated in Roland R. Arnold&#x00027;s laboratory at the Emory School of Dentistry, Atlanta, GA in the 1960s and was maintained in Jan Potempa&#x00027;s laboratory since 1989. This strain was of interest because it does not retain gingipains on the cell surface, instead releases the majority of proteases in a soluble form. In fact HG66 secretes all carboxy terminal domain-bearing proteins as soluble substances. Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/SAMN02732406">http://www.ncbi.nlm.nih.gov/biosample/SAMN02732406</ext-link> and Siddiqui et al. (<xref ref-type="bibr" rid="B46">2014</xref>).</p>
</sec>
<sec>
<title>A7436</title>
<p>This strain was isolated from the subgingival plaque of the tooth abscess of a refractory periodontitis patient by V.R. Dowell, Jr., at the Centers for Disease Control and Prevention in Atlanta, GA, in the mid-1980s. Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/SAMN03366764">http://www.ncbi.nlm.nih.gov/biosample/SAMN03366764</ext-link> and Chastain-Gross et al. (<xref ref-type="bibr" rid="B5">2015</xref>)</p>
</sec>
<sec>
<title>AJW4</title>
<p>This strain was isolated from the subgingival plaque of the tooth abscess of a periodontitis patient by R.J. Genco and colleagues in 1988 at SUNY-Buffalo, and described by A. Progulske-Fox and colleagues as a minimally invasive strain during <italic>in vitro</italic> cell culture studies. Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/SAMN03372093">http://www.ncbi.nlm.nih.gov/biosample/SAMN03372093</ext-link>.</p>
</sec>
<sec>
<title>Ando</title>
<p>This strain was isolated from the gingival sulcus of a human oral cavity in Japan in 1985. The genome of this strain was sequenced because it was reported to express a 53-kDa-type Mfa1 fimbrium, a major fimbrilin variant of Mfa1 previously known in many <italic>P. gingivalis</italic> strains. Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/?term=SAMD00040429">http://www.ncbi.nlm.nih.gov/biosample/?term=SAMD00040429</ext-link> and Nagano et al. (<xref ref-type="bibr" rid="B34">2015</xref>), Goto et al. (<xref ref-type="bibr" rid="B15">2015</xref>).</p>
</sec>
<sec>
<title>381</title>
<p>Strain 381 was isolated from the subgingival plaque of the tooth abscess of a localized chronic periodontitis patient by S. Socransky, A. Tanner, A. Crawford and colleagues at the Forsyth Dental Center (currently The Forsyth Institute), in the early 1970s. Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/SAMN03656156">http://www.ncbi.nlm.nih.gov/biosample/SAMN03656156</ext-link> and Chastain-Gross et al. (<xref ref-type="bibr" rid="B6">2017</xref>).</p>
</sec>
<sec>
<title>A7A1-28</title>
<p>A strain isolated from subgingival plaque of the tooth abscess of a periodontitis patient, with non-insulin dependent diabetes mellitus, by M.E. Neiders and colleagues in the mid-1987 at SUNY-Buffalo, and was described as a virulent strain with atypical fimbriae and capsule phenotypes. Information source: <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/biosample/SAMN03653671">http://www.ncbi.nlm.nih.gov/biosample/SAMN03653671</ext-link>.</p>
</sec>
<sec>
<title>MP4-504</title>
<p>This strain is a low-passage (fewer than five passages) clinical isolate sampled from the periodontal pocket (8 mm probing depth) of a chronic periodontitis patient at the University of Washington Graduate Periodontics Clinic in 1991. The important characteristics of this strain include stable adherence to oral streptococci, enhanced invasion of gingival epithelial cells (GECs), strong inhibition of IL-8 production by GECs, and the ability to transfer DNA by conjugation at high efficiencies (To et al., <xref ref-type="bibr" rid="B51">2016</xref>).</p>
</sec>
</sec>
<sec>
<title>Data analysis</title>
<sec>
<title>16S rRNA phylogeny</title>
<p>For the 16S rRNA gene phylogeny, <italic>16S rRNA</italic> gene sequences were extracted from the genomes of the 19 <italic>P. gingivalis</italic> strains based on NCBI&#x00027;s annotation (the <sup>&#x0002A;</sup>genomic.gff file in each of the downloaded genome folder). Sequences were pre-aligned with MAFFT v6.935b (2012/08/21) (Katoh and Standley, <xref ref-type="bibr" rid="B23">2013</xref>) and leading and trailing sequences not present in all sequences were trimmed. The trimmed and aligned sequences, with an alignment length of 1,425 bases and representing 20 unique sequences, were subjected to QuickTree V 1.1 (Howe et al., <xref ref-type="bibr" rid="B20">2002</xref>) using the &#x0201C;-kimura&#x0201D; option to calculate the substitution rate. A copy of the <italic>16S rRNA</italic> gene sequence from <italic>Porphyromonas asaccharolytica</italic> (PaDSM20707) was used as the out-group during the phylogenetic tree construction.</p>
</sec>
<sec>
<title>Core and unique proteins</title>
<p>To study the phylogenetic relationship based on more genes/proteins, protein sequences annotated by NCBI were used. Together with the outgroup species PaDSM20707, a total of 41,625 proteins were annotated by NCBI, including 39,926 from the 19 <italic>P. gingivalis</italic> genomes and 1,699 from PaDSM20707. Of the 39,926 <italic>P. gingivalis</italic> proteins, 37,667 are &#x02265; 50 amino acids in length and were searched for homologous clusters using the &#x0201C;blastclust&#x0201D; software V.2.2.25 (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html">http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html</ext-link>). Various sequence identity cutoffs ranging from 10 to 95% and two minimal alignment length cutoffs 50 and 90% were used for identifying the protein clusters. Proteins in each set of the identified clusters were aligned with MAFFT and poorly aligned regions were filtered by Gblocks 0.91b (Talavera and Castresana, <xref ref-type="bibr" rid="B49">2007</xref>). Trees were constructed with FastTree 2.1.9 (Price et al., <xref ref-type="bibr" rid="B40">2010</xref>) using the JTT protein mutation model (Jones et al., <xref ref-type="bibr" rid="B21">1992</xref>) and CAT&#x0002B;&#x02013;gemma options to account for the different rates of evolution at different sites. The reliability of tree splits were reported as &#x0201C;local support values&#x0201D; based on the Shimodaira-Hasegawa test (Shimodaira and Hasegawa, <xref ref-type="bibr" rid="B45">2001</xref>). For comparison, all 41,625 proteins were also subject to the PhyloPhlAn software (Segata et al., <xref ref-type="bibr" rid="B43">2013</xref>) version 0.99 (8 May 2013).</p>
<p>To identify proteins that are unique for each genome, all the 39,926 <italic>P. gingivalis</italic> proteins were searched against each other using BLASTP 2.2.25 with default parameters (Altschul et al., <xref ref-type="bibr" rid="B1">1997</xref>). Those that did not match any other protein with expected <italic>e</italic> value &#x02264; 10 were considered unique among the 19 genomes.</p>
</sec>
<sec>
<title>Whole genome nucleotide comparisons</title>
<p>Pairwise whole genome nucleotide to nucleotide sequence alignment were plotted using NUCmer (NUCleotide MUMmer) version 3.1 (Delcher et al., <xref ref-type="bibr" rid="B9">2002</xref>). To compare the whole genome DNA similarity by the oligonucleotide frequency, all possible 20-mer sequences present in the 20 genomes, including that of <italic>P. asaccharolytica</italic> strain DSM 20707 used as an out-group, were categorized and the number of genomes in which a 20-mer was present was recorded. Any given oligonucleotide can have a maximum of 20 (i.e., present in all 20 genomes) and a minimum of 1 (unique, found in only a single genome). To plot the oligonucleotide frequencies, an overall frequency for every 500 bases across the entire genome was calculated by recording the total number of genomes that all the possible 20-mer in the 500 bases can be found in (maximal 20, minimal 1). Each of the 500 bases windows was colored based on the genome frequency. Another plot was created similarly except that the non-coding regions were masked with light blue color to highlight the oligonucleotide frequencies for the areas that correspond to both forward (upper) and reverse-complement (lower) protein coding sequences.</p>
</sec>
<sec>
<title>Comparative functional genomics</title>
<p>Three functional annotation systems were used and compared in this study for all the 20 genomes&#x02013; (1) the NCBI prokaryotic genome annotation pipeline (Tatusova et al., <xref ref-type="bibr" rid="B50">2016</xref>), (2) the SEED and RAST (Rapid Annotation using Subsystem Technology) (Overbeek et al., <xref ref-type="bibr" rid="B39">2014</xref>), and (3) the KOALA (KEGG Orthology And Links Annotation) (Kanehisa et al., <xref ref-type="bibr" rid="B22">2016</xref>). The NCBI annotation results were downloaded from the NCBI FTP site described in the Sequence Sources above. The genomic DNA sequences were sent to the SEED server (Aziz et al., <xref ref-type="bibr" rid="B3">2012</xref>) using the Linux command-line and network-based SEED API downloaded from the SEED server web site (<ext-link ext-link-type="uri" xlink:href="http://blog.theseed.org/servers/installation/distribution-of-the-seed-server-packages.html">http://blog.theseed.org/servers/installation/distribution-of-the-seed-server-packages.html</ext-link>). The NCBI annotated proteins were sent to the BLastKoala website (<ext-link ext-link-type="uri" xlink:href="http://www.kegg.jp/blastkoala">http://www.kegg.jp/blastkoala</ext-link>) to identify the KEGG Orthologs. The results of both NCBI and RAST annotations were compared by several text based keyword searches. To identify more proteins in a particular functional category that were somehow annotated in certain genomes but not in others, protein sequences that were annotated in the same category from all 20 genomes were collected and used as the query to search for more proteins of the same functional category. NCBI BLASTP was used for this purpose and proteins with &#x02265; 95% sequence identity to and &#x02265; 95% coverage of the query sequences were identified as highly similar proteins. The number of proteins related to the IS5 transposase family was identified by the BlastKOALA program (Kanehisa et al., <xref ref-type="bibr" rid="B22">2016</xref>) with the matching to the KEGG Orthology (KO) number K07481. Additional functional comparison results were also made available as several files in Excel format.</p>
</sec>
</sec>
<sec>
<title>Data and results availability</title>
<p>To facilitate further comparison and future studies, all the data and results generated in this study, including the original downloaded sequences, annotations, the comparative results presented in this paper, as well as additional complete results that were not mentioned or discussed, are available for download from this FTP data repository site: <ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425">ftp://www.homd.org/publication_data/20160425</ext-link>.</p>
</sec>
</sec>
<sec id="s3">
<title>Results and discussion</title>
<sec>
<title>Summary of genome annotations</title>
<p>The first <italic>P. gingivalis</italic> genome released was that of the strain W83 in 2003 and the latest one was released in February 2016. Of the 19 genomes, eight were assembled into a single contig and were considered complete and finished genomes; the remaining were released as various numbers of sequence contigs assembled from whole genome shotgun (WGS) sequence reads. The sequence of JCVI SC001 appears to have a 1-contig circular sequence under the Genbank Accession number <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CM001843">CM001843</ext-link>, however it is a pseudo-contig generated by ordering the 284 unassembled contigs (accession number <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="APMB01000000">APMB01000000</ext-link>) based on the homologous matches to the genome of TDC60 (McLean et al., <xref ref-type="bibr" rid="B33">2013</xref>) and joining the ordered contigs with 282 100-N spacer sequences (total N length is 28,200 bps). Thus, it is not considered a complete or finished genome. Examining the sequences for the presence of Ns reveals the &#x0201C;completeness&#x0201D; of the genomes. Table <xref ref-type="table" rid="T2">2</xref> shows the reported length, non-N length, total number of Ns and the distribution of the N fragments in the genomic sequences. Overall strain A7A1-28 is the smallest of the completed <italic>P. gingivalis</italic> genomes with a size of 2,249,024 bps. HG66 has the largest size of all the sequenced <italic>P. gingivalis</italic> genomes at 2,441,680 bps after removing the 100 Ns placed at the end of the sequence. The placement of the 100 Ns at the end of the sequence was due to the unsuccessful attempt to circularize the sequence with the minimus2 software used by the PacBio sequencer at default settings (personal communication). For this reason the HG66 genome should not be considered complete. Almost all the unfinished draft genomes consist of various numbers of Ns ranging from 698 Ns in SDJ2 to 7,200 Ns in F0569 (Table <xref ref-type="table" rid="T2">2</xref>). It is likely that some of these published contigs were assembled based on a reference genome and the Ns had been filled in the gaps. Hence the true order of genes identified by the annotation process may not be correct.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><bold>Effective (non-Ns) sizes of the genomes</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Strain</bold></th>
<th valign="top" align="center"><bold>Contigs</bold></th>
<th valign="top" align="center"><bold>Size(bps)</bold></th>
<th valign="top" align="center"><bold>Non-N size(bps)<xref ref-type="table-fn" rid="TN5"><sup>a</sup></xref></bold></th>
<th valign="top" align="center"><bold>Ns (bps)</bold></th>
<th valign="top" align="center"><bold>N fragment size range (fragment count)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">HG66</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,441,780</td>
<td valign="top" align="center">2,441,680</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">100 (1)</td>
</tr>
<tr>
<td valign="top" align="left">JCVI_SC001</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,426,396</td>
<td valign="top" align="center">2,398,196</td>
<td valign="top" align="center">28,200</td>
<td valign="top" align="center">100 (282)</td>
</tr>
<tr>
<td valign="top" align="left">381</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,378,872</td>
<td valign="top" align="center">2,378,872</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">MP4-504</td>
<td valign="top" align="center">92</td>
<td valign="top" align="center">2,373,453</td>
<td valign="top" align="center">2,373,453</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">AJW4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,372,492</td>
<td valign="top" align="center">2,372,492</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">A7436</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,367,029</td>
<td valign="top" align="center">2,367,029</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">ATCC_33277</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,354,886</td>
<td valign="top" align="center">2,354,886</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">W83</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,343,476</td>
<td valign="top" align="center">2,343,476</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">TDC60</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,339,898</td>
<td valign="top" align="center">2,339,897</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1 (1)</td>
</tr>
<tr>
<td valign="top" align="left">SJD2</td>
<td valign="top" align="center">117</td>
<td valign="top" align="center">2,329,548</td>
<td valign="top" align="center">2,328,850</td>
<td valign="top" align="center">698</td>
<td valign="top" align="center">4&#x02013;256 (23)</td>
</tr>
<tr>
<td valign="top" align="left">F0568</td>
<td valign="top" align="center">154</td>
<td valign="top" align="center">2,334,744</td>
<td valign="top" align="center">2,328,244</td>
<td valign="top" align="center">6,500</td>
<td valign="top" align="center">100 (65)</td>
</tr>
<tr>
<td valign="top" align="left">F0566</td>
<td valign="top" align="center">192</td>
<td valign="top" align="center">2,306,092</td>
<td valign="top" align="center">2,300,992</td>
<td valign="top" align="center">5,100</td>
<td valign="top" align="center">100 (51)</td>
</tr>
<tr>
<td valign="top" align="left">F0570</td>
<td valign="top" align="center">117</td>
<td valign="top" align="center">2,282,791</td>
<td valign="top" align="center">2,278,391</td>
<td valign="top" align="center">4,400</td>
<td valign="top" align="center">100 (44)</td>
</tr>
<tr>
<td valign="top" align="left">A7A1-28</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2,249,024</td>
<td valign="top" align="center">2,249,024</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">None</td>
</tr>
<tr>
<td valign="top" align="left">W50</td>
<td valign="top" align="center">104</td>
<td valign="top" align="center">2,242,062</td>
<td valign="top" align="center">2,242,060</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">1 (2)</td>
</tr>
<tr>
<td valign="top" align="left">F0569</td>
<td valign="top" align="center">111</td>
<td valign="top" align="center">2,249,227</td>
<td valign="top" align="center">2,242,027</td>
<td valign="top" align="center">7,200</td>
<td valign="top" align="center">100 (72)</td>
</tr>
<tr>
<td valign="top" align="left">F0185</td>
<td valign="top" align="center">113</td>
<td valign="top" align="center">2,246,368</td>
<td valign="top" align="center">2,240,268</td>
<td valign="top" align="center">6,100</td>
<td valign="top" align="center">100 (61)</td>
</tr>
<tr>
<td valign="top" align="left">Ando</td>
<td valign="top" align="center">112</td>
<td valign="top" align="center">2,229,994</td>
<td valign="top" align="center">2,227,972</td>
<td valign="top" align="center">2,022</td>
<td valign="top" align="center">10&#x02013;100 (61)</td>
</tr>
<tr>
<td valign="top" align="left">W4087</td>
<td valign="top" align="center">114</td>
<td valign="top" align="center">2,216,597</td>
<td valign="top" align="center">2,212,597</td>
<td valign="top" align="center">4,000</td>
<td valign="top" align="center">100 (40)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN5"><label>a</label><p><italic>Genomes are ordered based on the non-N size</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Table <xref ref-type="table" rid="T3">3</xref> gives a numeric summary of the genome annotation results by the NCBI Prokaryotic Genome Annotation Pipeline (released 2013, <ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/genome/annotation_prok/">http://www.ncbi.nlm.nih.gov/genome/annotation_prok/</ext-link>). The NCBI pipeline is capable of identifying more than just the protein-coding genes, rRNAs and tRNAs, including several interesting types of genes such as binding sites, repeat sequences, pseudo-genes, and several types of non-coding RNAs (ncRNAs). However, since the NCBI pipeline is quite new, more features are still being added and since some of the annotations of these <italic>P. gingivalis</italic> genomes were done prior to 2013, the annotation results may not be comprehensive until the annotation is updated again based on the latest NCBI pipeline.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p><bold>Summary of the NCBI annotation<xref ref-type="table-fn" rid="TN6"><sup>a</sup></xref><sup>,</sup><xref ref-type="table-fn" rid="TN7"><sup>b</sup></xref></bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Strain</bold></th>
<th valign="top" align="center"><bold>Protein coding</bold></th>
<th valign="top" align="center"><bold>tRNA</bold></th>
<th valign="top" align="center"><bold>rRNA</bold></th>
<th valign="top" align="center"><bold>tmRNA<xref ref-type="table-fn" rid="TN9"><sup>d</sup></xref></bold></th>
<th valign="top" align="center"><bold>Repeat region</bold></th>
<th valign="top" align="center"><bold>Binding site</bold></th>
<th valign="top" align="center"><bold>Pseudo-gene</bold></th>
<th valign="top" align="center" colspan="4" style="border-bottom: thin solid #000000;"><bold>ncRNA<xref ref-type="table-fn" rid="TN8"><sup>c</sup></xref></bold></th>
<th valign="top" align="center"><bold>Other</bold></th>
<th valign="top" align="center"><bold>Annotation Release Date<xref ref-type="table-fn" rid="TN10"><sup>e</sup></xref></bold></th>
</tr>
<tr>
<th/>
<th/>
<th/>
<th/>
<th/>
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>Antisense-RNA</bold></th>
<th valign="top" align="center"><bold>RNase-P-RNA</bold></th>
<th valign="top" align="center"><bold>Auto-catalytically spliced intron</bold></th>
<th valign="top" align="center"><bold>Other ncRNA</bold></th>
<th/>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">5W83</td>
<td valign="top" align="center">1,909</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">41</td>
<td valign="top" align="center">2014-01-31</td>
</tr>
<tr>
<td valign="top" align="left">ATCC_33277</td>
<td valign="top" align="center">2,090</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">210</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2011-11-26</td>
</tr>
<tr>
<td valign="top" align="left">TDC60</td>
<td valign="top" align="center">2,220</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">380</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">34</td>
<td valign="top" align="center">2011-08-17</td>
</tr>
<tr>
<td valign="top" align="left">W50</td>
<td valign="top" align="center">2,016</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2012-06-25</td>
</tr>
<tr>
<td valign="top" align="left">JCVI_SC001</td>
<td valign="top" align="center">2,354</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-04-23</td>
</tr>
<tr>
<td valign="top" align="left">F0568</td>
<td valign="top" align="center">2,410</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-09-16</td>
</tr>
<tr>
<td valign="top" align="left">F0569</td>
<td valign="top" align="center">2,297</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-09-16</td>
</tr>
<tr>
<td valign="top" align="left">F0570</td>
<td valign="top" align="center">2,315</td>
<td valign="top" align="center">44</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-09-16</td>
</tr>
<tr>
<td valign="top" align="left">F0185</td>
<td valign="top" align="center">2,233</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-09-16</td>
</tr>
<tr>
<td valign="top" align="left">F0566</td>
<td valign="top" align="center">2,392</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-09-16</td>
</tr>
<tr>
<td valign="top" align="left">W4087</td>
<td valign="top" align="center">2,202</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-09-16</td>
</tr>
<tr>
<td valign="top" align="left">SJD2</td>
<td valign="top" align="center">2,012</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">62</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2013-12-04</td>
</tr>
<tr>
<td valign="top" align="left">HG66</td>
<td valign="top" align="center">1,958</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">38</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2014-10-22</td>
</tr>
<tr>
<td valign="top" align="left">A7436</td>
<td valign="top" align="center">2,004</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2015-08-11</td>
</tr>
<tr>
<td valign="top" align="left">AJW4</td>
<td valign="top" align="center">2,002</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2015-08-26</td>
</tr>
<tr>
<td valign="top" align="left">Ando</td>
<td valign="top" align="center">1,770</td>
<td valign="top" align="center">47</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2015-11-27</td>
</tr>
<tr>
<td valign="top" align="left">381</td>
<td valign="top" align="center">1,968</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">9</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2015-10-14</td>
</tr>
<tr>
<td valign="top" align="left">A7A1-28</td>
<td valign="top" align="center">1,841</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">37</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2015-11-17</td>
</tr>
<tr>
<td valign="top" align="left">MP4-504</td>
<td valign="top" align="center">1,889</td>
<td valign="top" align="center">47</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">99</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2016-02-09</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN6"><label>a</label><p><italic>Data analyzed based on the gff files of each genome generated by the NCBI annotation pipeline</italic>.</p></fn>
<fn id="TN7"><label>b</label><p><italic>Detail information provided by NCBI can also be downloaded from <ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425/1_Sequence_Sources/">ftp://www.homd.org/publication_data/20160425/1_Sequence_Sources/</ext-link></italic>.</p></fn>
<fn id="TN8"><label>c</label><p><italic>Mon-coding RNA</italic>.</p></fn>
<fn id="TN9"><label>d</label><p><italic>Trans-messenger RNA: a bacterial RNA molecule with dual tRNA-like and mRNA-like properties</italic>.</p></fn>
<fn id="TN10"><label>e</label><p><italic>NCBI annotation release dates were based on the dates reported in the protein.gbff file in the above FTP link</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>In addition to the NCBI annotations, RAST (Rapid Annotations using Subsystems Technology) is also a popular pipeline for annotating microbial genomes (Aziz et al., <xref ref-type="bibr" rid="B2">2008</xref>). All the 19 <italic>P. gingivalis</italic> genomes, as well as the chosen outgroup <italic>P. asaccharolytica</italic> DSM20707 were subjected to the RAST pipeline and the results were compared with those done by the NCBI pipeline. As shown in Table <xref ref-type="table" rid="T4">4</xref>, both the RAST and NCBI pipelines identified almost the same number of rRNA and tRNA genes. However, the numbers of protein-coding genes varied quite significantly between the two pipelines. Although most of the genes were commonly identified, up to hundreds of protein-coding sequences can be missed by either system. Moreover, 86% (6,422 of 7,382 for all the 19 genomes) of these uniquely identified genes code for hypothetical proteins and 80% are shorter than 100 amino acids in length (in fact, only 94 have lengths &#x02265; 500 amino acids), thus the impact due to the annotation discrepancy may not be as significant especially when drawing conclusions in genome-wide systematic analysis or metabolic pathway capability.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p><bold>Comparison of NCBI and RAST genome annotations<xref ref-type="table-fn" rid="TN11"><sup>a</sup></xref></bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Strain</bold></th>
<th valign="top" align="center"><bold>Total NCBI</bold></th>
<th valign="top" align="center"><bold>Total RAST</bold></th>
<th valign="top" align="center"><bold>Common/Unique<xref ref-type="table-fn" rid="TN12"><sup>b</sup></xref></bold></th>
<th valign="top" align="center"><bold>5S rRNA</bold></th>
<th valign="top" align="center"><bold>16S rRNA</bold></th>
<th valign="top" align="center"><bold>23S rRNA</bold></th>
<th valign="top" align="center"><bold>tRNA</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">W83</td>
<td valign="top" align="center">1,909</td>
<td valign="top" align="center">2,163</td>
<td valign="top" align="center">1784/80/334</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">ATCC_33277</td>
<td valign="top" align="center">2,090</td>
<td valign="top" align="center">2,092</td>
<td valign="top" align="center">1911/154/144</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">TDC60</td>
<td valign="top" align="center">2,220</td>
<td valign="top" align="center">2,090</td>
<td valign="top" align="center">1880/286/167</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">W50</td>
<td valign="top" align="center">2,016</td>
<td valign="top" align="center">2,036</td>
<td valign="top" align="center">1887/102/123</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">48</td>
</tr>
<tr>
<td valign="top" align="left">JCVI_ SC001</td>
<td valign="top" align="center">2,354</td>
<td valign="top" align="center">2,136</td>
<td valign="top" align="center">2030/276/78</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">45/42</td>
</tr>
<tr>
<td valign="top" align="left">F0568</td>
<td valign="top" align="center">2,417</td>
<td valign="top" align="center">2,096</td>
<td valign="top" align="center">1939/403/111</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">46</td>
</tr>
<tr>
<td valign="top" align="left">F0569</td>
<td valign="top" align="center">2,297</td>
<td valign="top" align="center">1,982</td>
<td valign="top" align="center">1845/377/92</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">46</td>
</tr>
<tr>
<td valign="top" align="left">F0570</td>
<td valign="top" align="center">2,316</td>
<td valign="top" align="center">2,063</td>
<td valign="top" align="center">1912/338/107</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">44</td>
</tr>
<tr>
<td valign="top" align="left">F0185</td>
<td valign="top" align="center">2,236</td>
<td valign="top" align="center">2,005</td>
<td valign="top" align="center">1862/319/107</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">45</td>
</tr>
<tr>
<td valign="top" align="left">F0566</td>
<td valign="top" align="center">2,395</td>
<td valign="top" align="center">2,044</td>
<td valign="top" align="center">1885/428/112</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">45</td>
</tr>
<tr>
<td valign="top" align="left">W4087</td>
<td valign="top" align="center">2,204</td>
<td valign="top" align="center">1,973</td>
<td valign="top" align="center">1850/303/92</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">45</td>
</tr>
<tr>
<td valign="top" align="left">SJD2</td>
<td valign="top" align="center">2,020</td>
<td valign="top" align="center">2,166</td>
<td valign="top" align="center">1845/136/271</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">48/47</td>
</tr>
<tr>
<td valign="top" align="left">HG66</td>
<td valign="top" align="center">1,958</td>
<td valign="top" align="center">2,215</td>
<td valign="top" align="center">1881/58/298</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">A7436</td>
<td valign="top" align="center">2,004</td>
<td valign="top" align="center">2,173</td>
<td valign="top" align="center">1898/84/239</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">AJW4</td>
<td valign="top" align="center">2,002</td>
<td valign="top" align="center">2,139</td>
<td valign="top" align="center">1884/104/226</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">Ando</td>
<td valign="top" align="center">1,788</td>
<td valign="top" align="center">1,989</td>
<td valign="top" align="center">1674/76/275</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">47</td>
</tr>
<tr>
<td valign="top" align="left">381</td>
<td valign="top" align="center">1,968</td>
<td valign="top" align="center">2,108</td>
<td valign="top" align="center">1853/91/221</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">A7A1-28</td>
<td valign="top" align="center">1,841</td>
<td valign="top" align="center">2,039</td>
<td valign="top" align="center">1736/89/269</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">53</td>
</tr>
<tr>
<td valign="top" align="left">MP4-504</td>
<td valign="top" align="center">1,891</td>
<td valign="top" align="center">2,181</td>
<td valign="top" align="center">1806/68/347</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">47</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN11"><label>a</label><p><italic>Only protein-coding, rRNA and tRNA genes were compared since these are the only types of genes annotated by RAST</italic>.</p></fn>
<fn id="TN12"><label>b</label><p><italic>The three numbers shown (X/Y/Z) are X, common genes, genes with &#x02265; 80% overlapped based on the annotated start and end postion; Y, RAST unique genes, gene annotated by RAST without overlap of any NCBI gene; Z, NCBI unique genes, genes annotated by NCBI without overlap to any RAST gene. There are genes that are partially overlapping to each other with &#x0003C; 80% of the length not included.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>A list of the 960 (7,382&#x02013;6,422) non-hypothetical proteins is provided at the link (<ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425/2_Summary_of_Genome_Annotations/Non-overlap_Non-hypothetical_protein_identified_by_NCBI_or_RAST.fasta">ftp://www.homd.org/publication_data/20160425/2_Summary_of_Genome_Annotations/Non-overlap_Non-hypothetical_protein_identified_by_NCBI_or_RAST.fasta</ext-link>).</p>
</sec>
<sec>
<title>16S rRNA phylogeny</title>
<p>The 16S rRNA sequences have been used to infer the evolutionary relatedness of the prokaryotes due to its slow rate of evolution (Woese et al., <xref ref-type="bibr" rid="B56">1990</xref>). However, multiple <italic>rRNA</italic> genes including 16S rRNAs are common in prokaryotic genomes (Klappenbach et al., <xref ref-type="bibr" rid="B24">2000</xref>) and the genomic copy number of 16S rRNA varies greatly among species from 1 to 15 (Vetrovsky and Baldrian, <xref ref-type="bibr" rid="B54">2013</xref>). The number of rRNA genes was reported to correlate with the rate at which phylogenetically diverse bacteria respond to resource availability (Klappenbach et al., <xref ref-type="bibr" rid="B24">2000</xref>). As shown in Table <xref ref-type="table" rid="T4">4</xref>, all of the eight genomes which had been assembled to a single contig contain four copies of 5S, 16S, and 23S rRNA genes respectively, thus it is reasonable to believe that all P. gingivalis genomes have four copies of the rRNA operons. The lower number of rRNA genes in the unfinished genomes is likely due to the incompleteness of the sequences and is also likely due to the fact that genomes sequenced by short reads sequencing platforms such as those of the Illumina sequencers cannot be easily assembled across the repeated regions such as the highly conserved rRNA operons.</p>
<p>The 16S rRNA sequences of all the 19 genomes annotated by NCBI were extracted and aligned for the construction of a phylogenetic tree. Based on the annotation, there are a total of 24 unique 16S rRNA gene sequences identified from the 19 genomes (Table <xref ref-type="table" rid="T5">5</xref>, first column), plus the sequence of a close species P. asaccharolytica strain DSM 20707 (Accession Number <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP002689">CP002689</ext-link>), making it a total of 25 unique sequences in the study. However, many of the sequence differences are due to different annotated lengths. After aligning all the 24 (25 if including the outgroup sequence) unique sequences and trimming off the leading and trailing sequences not present in all copies (trimmed aligned length &#x0003D; 1,425 bps), the aligned portion of several sequences are identical and the number of unique sequences was reduced to 20 (second column of Table <xref ref-type="table" rid="T5">5</xref>). Strains 381, A7A1-28, ATCC 33277, and W83 all have four copies of identical sequences in their genomes (Table <xref ref-type="table" rid="T5">5</xref>, last column). The aligned regions of all the 16S rRNA gene sequences from ATCC 33277, 381, and HG66 have identical sequences, indicating the close evolutional distance of these strains. Strain A7436 shared three of its four copies of 16S rRNA sequences identically with those of W83. Together with the single copy from W50, they formed an identical group of sequences. W50 has been known to be a close strain of W83, thus the identical sequences between these two are not surprising. The explanation of identical copies of the 16S rRNA sequence in the genome is apparently due to the gene duplication event and the fact that several strains shared identical duplicated sequences suggested that the duplication event occurred after the speciation. Strains A7436, AJW4, and HG66 had three strain-specific identical sequences with the 4th copy different from the other three. Overall, all the P. gingivalis 16S rRNA gene sequences were extremely similar and often have only a single number of nucleotide mismatches between any two strains (if not identical). Altogether only 16 loci on the gene had nucleotide variations (some could have variations in more than one strain), with the exception of one copy in TDC60, which had a series of A &#x02192; C or G &#x02192; C transversions between position 50 and 130 and two single nucleotide insertions at position 174 and 233. It is thus the most divergent sequence of among all 19 P. gingivalis genomes. These aligned and trimmed sequences, including the outgroup sequence from P. asaccharolytica strain DSM 20707, were used to construct a phylogenetic tree based on Kimura&#x00027;s nucleotide substitution model and the result is shown in Figure <xref ref-type="fig" rid="F1">1</xref>. The phylogenetic tree depicts a likely evolutionary path for these different P. gingivalis strains. The strains 381, ATCC 33277, and HG66 appeared to be closer to the potential common ancestor, based on the tree topology inferred with a close species as the outgroup sequence. The other strains gradually diversified into deeper branching nodes with two of the sequences from strains F0566 and TDC60 as the most deeply branched and mutated from the common ancestor, which was inferred by using the sequence of a neighboring species (PaDSM20707) as an outgroup.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p><bold>Unique <italic>16S rRNA</italic> gene sequences in <italic>P. gingivalis</italic> genomes</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Original sequence</bold></th>
<th valign="top" align="center"><bold>Trimmed sequence<xref ref-type="table-fn" rid="TN13"><sup>a</sup></xref></bold></th>
<th valign="top" align="center"><bold>Copy number</bold></th>
<th valign="top" align="center"><bold>Original length (bps)</bold></th>
<th valign="top" align="center"><bold>Strains (copy number)<xref ref-type="table-fn" rid="TN14"><sup>b</sup></xref></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Unique Seq 1</td>
<td valign="top" align="center">Unique Trimmed Seq 1</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1,422</td>
<td valign="top" align="center">381 (4)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 2</td>
<td/>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1,475</td>
<td valign="top" align="center">ATCC33277 (4)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 3</td>
<td/>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1,538</td>
<td valign="top" align="center">HG66 (3)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 4</td>
<td valign="top" align="center">Unique Trimmed Seq 2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,538</td>
<td valign="top" align="center">HG66</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 5</td>
<td valign="top" align="center">Unique Trimmed Seq 3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1,422</td>
<td valign="top" align="center">A7436 (3)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 6</td>
<td/>
<td valign="top" align="center">5</td>
<td valign="top" align="center">1,475</td>
<td valign="top" align="center">W50; W83 (4)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 7</td>
<td valign="top" align="center">Unique Trimmed Seq 4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,422</td>
<td valign="top" align="center">A7436</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 8</td>
<td valign="top" align="center">Unique Trimmed Seq 5</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1,422</td>
<td valign="top" align="center">A7A1-28 (4)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 9</td>
<td valign="top" align="center">Unique Trimmed Seq 6</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1,422</td>
<td valign="top" align="center">AJW4 (3)</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 10</td>
<td valign="top" align="center">Unique Trimmed Seq 7</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,422</td>
<td valign="top" align="center">AJW4</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 11</td>
<td valign="top" align="center">Unique Trimmed Seq 8</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,521</td>
<td valign="top" align="center">TDC60</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 12</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">TDC60</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 13</td>
<td valign="top" align="center">Unique Trimmed Seq 9</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,522</td>
<td valign="top" align="center">TDC60</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 14</td>
<td valign="top" align="center">Unique Trimmed Seq 10</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">TDC60</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 15</td>
<td valign="top" align="center">Unique Trimmed Seq 11</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,475</td>
<td valign="top" align="center">JCVI SC001</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 16</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,538</td>
<td valign="top" align="center">SJD2</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 17</td>
<td valign="top" align="center">Unique Trimmed Seq 12</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,475</td>
<td valign="top" align="center">Ando</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 18</td>
<td valign="top" align="center">Unique Trimmed Seq 13</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">W4087</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 19</td>
<td valign="top" align="center">Unique Trimmed Seq 14</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">F0569</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 20</td>
<td valign="top" align="center">Unique Trimmed Seq 15</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">F0568</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 21</td>
<td valign="top" align="center">Unique Trimmed Seq 16</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">F0185</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 22</td>
<td valign="top" align="center">Unique Trimmed Seq 17</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">F0566</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 23</td>
<td valign="top" align="center">Unique Trimmed Seq 18</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,542</td>
<td valign="top" align="center">MP4-504</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 24</td>
<td valign="top" align="center">Unique Trimmed Seq 19</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1,520</td>
<td valign="top" align="center">F0570</td>
</tr>
<tr>
<td valign="top" align="left">Unique Seq 25<xref ref-type="table-fn" rid="TN15"><sup>c</sup></xref></td>
<td valign="top" align="center">Unique Trimmed Seq 20</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">1,517</td>
<td valign="top" align="center">PaDSM20707 (2)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN13"><label>a</label><p><italic>Sequences were pre-aligned with the software MAFFT v6.935b (2012/08/21) (Katoh and Standley, <xref ref-type="bibr" rid="B23">2013</xref>) with default setting; after trimming the leading and trailing sequences not present for all genomes, the trimmed aligned sequence length is 1,425 bps in length</italic>.</p></fn>
<fn id="TN14"><label>b</label><p><italic>If multiple copies of identical sequences are present, the copy number is indicated in the parenthesis</italic>.</p></fn>
<fn id="TN15"><label>c</label><p><italic>Sequence of P. asaccharolytica strain DSM 20707 (from Genbank ID: <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CP002689">CP002689</ext-link>) was included as outgroup</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Phylogenetic tree of <italic>P. gingivalis 16S rRNA</italic> gene sequences</bold>. A total of 24 unique <italic>16S rRNA</italic> gene sequences were extracted from the genomes of 19 <italic>P. gingivalis</italic> strains annotated by NCBI. Sequences were pre-aligned with MAFFT v6.935b (2012/08/21) (Katoh and Standley, <xref ref-type="bibr" rid="B23">2013</xref>) and leading and trailing sequences not present in all sequences were trimmed. The trimmed aligned sequences represent 20 unique sequences and were subject to QuickTree V 1.1 (Howe et al., <xref ref-type="bibr" rid="B20">2002</xref>) using the &#x0201C;-kimura&#x0201D; option to calculate the substitution rate. Sequence of <italic>P. asaccharolytica</italic> strain DSM 20707 (PaDSM20707) was used as out-group. The branch length of the out-group was truncated to fit the tree in the figure and the substitution rate is indicated with the blue number. The red numbers next to the branching point are the bootstrap values based on 100 iterations. Sequences of different strains were separated by semicolons and the number of sequences were indicated in the parentheses in the format of (x&#x02013;y/z), where x and y are the start and end IDs and z the total number in the strain.</p></caption>
<graphic xlink:href="fcimb-07-00028-g0001.tif"/>
</fig>
</sec>
<sec>
<title>Core and unique proteins</title>
<p>The phylogenetic relationship inferred based on the <italic>16S rRNA</italic> gene sequences reported above can only represent the evolution of this particular gene, hence a gene tree. A more comprehensive way of studying the evolutionary relatedness of different genomes is to use as much genomic information as possible in the analysis (i.e., phylogenomics). A popular approach is to use the core proteins for the construction of a tree that may be closer to a true species tree, if such a tree exists, or if there is no true species tree, may reflect more on the relatedness of these strains at the genomic level. The concept of &#x0201C;core&#x0201D; proteins is ideally defined as proteins that are present and required by all the genomes in study, however the identification of such a group of proteins, namely orthologs, is not straightforward and the results vary depending on the criteria used. It is challenging, if not impossible, to identify all the orthologous proteins among a group of genomes. In general, genomes of closer species or strains share more orthologs; however any percent protein sequence identity chosen as the cutoff to test whether a group of homologous proteins are truly orthologs (or paralogs) can always include some false positive and negative orthologs. Nevertheless, one can still hypothesize that a more reliable evolutionary relationship of a group of genomes can be obtained if the use of higher or lower percent identity constrains does not affect the overall tree topology.</p>
<p>To test this hypothesis, the &#x0201C;core&#x0201D; proteins were first identified among all <italic>P. gingivalis</italic> genomes under different cutoffs. Based on the NCBI annotation, a total of 39,926 protein sequences were identified. However, for some unknown reasons, some of the annotated protein lengths were as short as one or two amino acids. For example, proteins with Genbank IDs <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="GAP82138.1">GAP82138.1</ext-link>, <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="GAP81676.1">GAP81676.1</ext-link>, and <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="GAP81848.1">GAP81848.1</ext-link> in strain Ando were identified with only 1, 2, and 2 amino acids in length respectively. These clearly were annotation errors caused by the computational bugs in the annotation pipeline. In this analysis, only protein sequences with a minimal length of 50 amino acids were used (a total of 37,667 proteins) for identifying core proteins. They were subject to the &#x0201C;blastclust&#x0201D; program (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html">http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html</ext-link>) to identify clusters of proteins that share a certain degree of sequence homology and with specified alignment length coverage. In this analysis, if a protein is present (i.e., meets the % identity and alignment cutoffs specified) in all 19 genomes (or 20 genomes if PaDSM20707 was used as the outgroup in some results) it is considered as a core/shared protein, and if a protein is only present in a single genome it is considered as a strain-specific unique protein.</p>
<p>Figure <xref ref-type="fig" rid="F2">2</xref> shows the potential numbers of both core and unique proteins in the 19 genomes analyzed with &#x0201C;blastclust&#x0201D; by varying two parameters: Sequence percent identity cutoffs (from 95 to 10%) and percent alignment length (90 and 50%). Figure <xref ref-type="fig" rid="F2">2A</xref> shows that regardless of sequence identity cutoffs; the number of core proteins stays relatively constant around 1,000 with 90% as the alignment length cutoff. The number of core protein groups increased gradually from 1,037 at 95% identity, maximized at 1,045 at 60%, then decreased to 910 at 10%. The reason for the increase from 95 to 60% was due to more core protein groups clustered together at lower % identity. The decrease after 60% identity was due to the fact that different protein groups identified with identity &#x02265; 60% began to merge into fewer groups. The 1,037 shared proteins were detected under most stringent conditions thus it is reasonable to state that at least 1,037 core proteins were detected based on 19 strains. This number is expectedly smaller than the 1,476 detected in the core genome based on eight <italic>P. gingivalis</italic> strains (Brunner et al., <xref ref-type="bibr" rid="B4">2010</xref>). It should also be noted that the core genes/proteins are not the same as the &#x0201C;essential&#x0201D; genes, for which only ca. 400 were experimentally detected previously (Klein et al., <xref ref-type="bibr" rid="B26">2012</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Core and unique genes in <italic>P</italic>. <italic>gingivalis</italic> surveyed by sequence identity and alignment length</bold>. Of the 39,926 NCBI annotated <italic>P. gingivalis</italic> proteins, 37,667 are &#x02265; 50 amino acids in length and were searched for homologous clusters using the &#x0201C;blastclust&#x0201D; software V.2.2.25 (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html">http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html</ext-link>). Various sequence identity cutoffs ranging from 10 to 95% and two minimal alignment length cutoffs 50 and 90% were used as the program parameters to identify the protein clusters in the three categories <bold>(A)</bold> clusters containing proteins from all 19 genomes; <bold>(B)</bold> clusters containing proteins from 2 to 18 genomes; and <bold>(C)</bold> clusters with protein from only 1 genome.</p></caption>
<graphic xlink:href="fcimb-07-00028-g0002.tif"/>
</fig>
<p>For the purpose of identifying a core/shared set of proteins for constructing a phylogenomic tree, the 1,045 core proteins identified at 60% sequence identity and 90% alignment length cutoffs were used for sequence alignment and tree building. This set of sequences is available for download in the data repository FTP site mentioned in the Material and Methods. In addition, as expected when the percent alignment length was decreased from 90 to 50%, more proteins were identified as core proteins, e.g., from 1,289 at 95% identity cutoff to 1,301 at 60%, due to the fact that more proteins share the same percent identity over shorter sequence length.</p>
<p>Figure <xref ref-type="fig" rid="F2">2B</xref> shows the number of protein groups that are shared by 2 to 18 genomes (partially shared proteins). The number decreases with the lower % identity because similar protein groups that were identified as separate groups merged into a single (but larger) group due to the more relaxed (lower) % identity (e.g., from 1,927 at 95% to 1,651 at 10%). However, contrary to the core proteins above, which require proteins present in all 19 genomes, when the percent alignment decreased, fewer partially-shared proteins were identified. This is as expected because when the percent alignment cutoff was lowered, a protein group which consists of only members from for example 18 genomes, at higher cutoff, now may find a member in the 19th genome thus disqualifying it as the 18-genome partially-shared group.</p>
<p>Figure <xref ref-type="fig" rid="F2">2C</xref> shows the number of strain-specific proteins that are present in only one genome with a single copy. Similar to the partially-shared groups, as the % identity decreases the number of unique proteins becomes smaller because more proteins from different genomes were lumped together as a homologous group under a lower % identity, resulting in the loss of the &#x0201C;uniqueness&#x0201D;. In addition, for example, at 60% sequence identity and 90% alignment cutoffs, there were 2,289 proteins identified as present in a single genome, but the number was reduced to 1,044 at 50% alignment cutoff&#x02013;1,245 proteins lost their uniqueness due to the presence of more &#x0201C;similar&#x0201D; proteins found in other genomes.</p>
<p>For the unique proteins identified, it would be interesting to observe their distribution in the 19 genomes and the result may help understand which genomes possess more or fewer unique proteins. Figure <xref ref-type="fig" rid="F3">3</xref> shows the distribution of the 1,044 unique proteins identified with the 50% alignment cutoff (Figure <xref ref-type="fig" rid="F3">3A</xref>) and 2,289 with 90% cutoff (Figure <xref ref-type="fig" rid="F3">3B</xref>). Regardless of the sequence identity and percent alignment cutoff, the results show that some strains possess significantly more unique proteins than others. Four strains, F0566, F0568, F0569, and JCVI SC001 have a significantly higher number of unique proteins under all identification conditions - as high as 96&#x02013;249 unique proteins (for percent identity 10&#x02013;95% at 50% alignment length) in the case of F0566 (Figure <xref ref-type="fig" rid="F3">3A</xref>). On the other hand, strains 381, ATCC 33277, A7436, and W83 are the four strains with the lowest number of unique proteins, only 10&#x02013;15 unique proteins (10&#x02013;95% identity; 50% alignment) in the case of W83 (Figure <xref ref-type="fig" rid="F3">3A</xref>). Interestingly, strain W50, the closest strain to W83, encodes more unique proteins (30&#x02013;46) even though it is an unfinished draft genome. Apparently the incompletes of the draft genomes are not the cause for the difference in the number of unique proteins (JCVI SC001 and all the F strains are draft genomes). This further suggests that the gaps in the draft genomes may contain only repeated sequences that either do not encode for proteins, or encode repeated proteins that do not contribute much to the genome&#x00027;s uniqueness. However, we cannot rule out the possibility that these regions could still have some unique proteins which would drive up the number already observed.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Unique proteins in 19 <italic>P. gingivalis</italic> strains</bold>. Of the 39,926 NCBI annotated <italic>P. gingivalis</italic> proteins, 37,667 are &#x02265; 50 amino acids in length and were searched for homologous clusters using the &#x0201C;blastclust&#x0201D; software V.2.2.25 (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html">http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html</ext-link>). Unique proteins of each of the 19 <italic>P. gingivalis</italic> genomes were identified as proteins found in only one genome without any similar counterpart in any other. The total number of clusters that contain only unique proteins for each genome were plotted. Various sequence identity cutoffs ranging from 10 to 95% (dots with varying grayscale color intensity) and two minimal alignment length cutoffs 50% <bold>(A)</bold> and 90% <bold>(B)</bold> were used as the program parameters.</p></caption>
<graphic xlink:href="fcimb-07-00028-g0003.tif"/>
</fig>
<p>Another noteworthy observation is that there is a consistent gap between the data points 95% and 90% identity when searching for unique proteins in all strains under both alignment conditions (Figure <xref ref-type="fig" rid="F3">3</xref>). This suggests that the more proteins identified as unique at 95% became &#x0201C;similar&#x0201D; at 90%. Hence 90% sequence identity may be an ideal cutoff for differentiating homologs and unique proteins, at least at the strain level.</p>
<p>Table <xref ref-type="table" rid="T6">6</xref> lists the percentage of proteins that were annotated in NCBI as the hypothetical proteins (functionally unknown) and the percentage of the unique proteins that were identified with 80% as the sequence identity and 50% alignment length. The total percent hypothetical proteins range from 26% (W50) to as high as 46% (F0566 and F0568), whereas the majority of the unique proteins are hypothetical, from 68% (TDC60) to 100% (W83). Thus, until more functions of the hypothetical proteins are understood, it will still be challenging to understand what each genome&#x00027;s overall &#x0201C;specialty&#x0201D; functions conferred by the unique genes. To give a glimpse of what each genome&#x00027;s most unique functions are, based on the currently available information, Table <xref ref-type="table" rid="T7">7</xref> lists the functional annotations of the non-hypothetical unique proteins for each of the 19 genomes (with default BLASTP parameter, i.e., expected e value &#x02264; 10) (Altschul et al., <xref ref-type="bibr" rid="B1">1997</xref>). All of them are among the proteins identified above under the most stringent parameters in terms of uniqueness&#x02013;50% sequence identity and 50% alignment length. Strain JCVI SC001, an environment isolate from a hospital sink drain, has the most diverse functions encoded by these unique proteins. The unique toxin-antitoxin system detected in strain F0569 (a clinical isolate from subgingival plaque biofilm) is also of interest. The toxin-antitoxin system genes when carried on a plasmid is often referred to as the post-segregational killing (PSK) system (Gerdes, <xref ref-type="bibr" rid="B14">2000</xref>) while when carried in the chromosome such a system is involved in stress response and &#x0201C;programmed cell death&#x0201D; (Hayes, <xref ref-type="bibr" rid="B18">2003</xref>). Although indigenous plasmids have never been detected in <italic>P. gingivalis</italic> this does not rule out finding one in the future. Thus, which type of function that this toxin-antitoxin system is involved in can only be speculated at this point of time. Whether these annotations translate to unique functions of the genome, require further investigation to ensure there are no other non-homologous proteins that play similar functions. All of the unique proteins identified are available by strain in the FTP data repository.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p><bold>Percent hypothetical proteins 19 <italic>P. gingivalis</italic> genomes</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Strain<xref ref-type="table-fn" rid="TN16"><sup>a</sup></xref></bold></th>
<th valign="top" align="center"><bold>Total</bold></th>
<th valign="top" align="center"><bold>% Total hypothetical(%)</bold></th>
<th valign="top" align="center"><bold>Total unique<xref ref-type="table-fn" rid="TN17"><sup>b</sup></xref> (80% identity)</bold></th>
<th valign="top" align="center"><bold>% Unique hypothetical (80% identity)(%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">HG66</td>
<td valign="top" align="center">1,958</td>
<td valign="top" align="center">28</td>
<td valign="top" align="center">53</td>
<td valign="top" align="center">81</td>
</tr>
<tr>
<td valign="top" align="left">381</td>
<td valign="top" align="center">1,968</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">85</td>
</tr>
<tr>
<td valign="top" align="left">ATCC_33277</td>
<td valign="top" align="center">2,090</td>
<td valign="top" align="center">42</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">79</td>
</tr>
<tr>
<td valign="top" align="left">A7A1-28</td>
<td valign="top" align="center">1,841</td>
<td valign="top" align="center">28</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">78</td>
</tr>
<tr>
<td valign="top" align="left">MP4-504</td>
<td valign="top" align="center">1,891</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">34</td>
<td valign="top" align="center">85</td>
</tr>
<tr>
<td valign="top" align="left">Ando</td>
<td valign="top" align="center">1,788</td>
<td valign="top" align="center">29</td>
<td valign="top" align="center">61</td>
<td valign="top" align="center">70</td>
</tr>
<tr>
<td valign="top" align="left">F0568</td>
<td valign="top" align="center">2,417</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">114</td>
<td valign="top" align="center">88</td>
</tr>
<tr>
<td valign="top" align="left">F0569</td>
<td valign="top" align="center">2,297</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">125</td>
<td valign="top" align="center">86</td>
</tr>
<tr>
<td valign="top" align="left">W4087</td>
<td valign="top" align="center">2,204</td>
<td valign="top" align="center">43</td>
<td valign="top" align="center">94</td>
<td valign="top" align="center">78</td>
</tr>
<tr>
<td valign="top" align="left">F0185</td>
<td valign="top" align="center">2,236</td>
<td valign="top" align="center">43</td>
<td valign="top" align="center">72</td>
<td valign="top" align="center">88</td>
</tr>
<tr>
<td valign="top" align="left">F0570</td>
<td valign="top" align="center">2,316</td>
<td valign="top" align="center">44</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">90</td>
</tr>
<tr>
<td valign="top" align="left">JCVI_ SC001</td>
<td valign="top" align="center">2,354</td>
<td valign="top" align="center">30</td>
<td valign="top" align="center">172</td>
<td valign="top" align="center">72</td>
</tr>
<tr>
<td valign="top" align="left">SJD2</td>
<td valign="top" align="center">2,020</td>
<td valign="top" align="center">35</td>
<td valign="top" align="center">79</td>
<td valign="top" align="center">82</td>
</tr>
<tr>
<td valign="top" align="left">AJW4</td>
<td valign="top" align="center">2,002</td>
<td valign="top" align="center">29</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">76</td>
</tr>
<tr>
<td valign="top" align="left">A7436</td>
<td valign="top" align="center">2,004</td>
<td valign="top" align="center">28</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">80</td>
</tr>
<tr>
<td valign="top" align="left">W50</td>
<td valign="top" align="center">2,016</td>
<td valign="top" align="center">26</td>
<td valign="top" align="center">34</td>
<td valign="top" align="center">91</td>
</tr>
<tr>
<td valign="top" align="left">W83</td>
<td valign="top" align="center">1,909</td>
<td valign="top" align="center">35</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left">F0566</td>
<td valign="top" align="center">2,395</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">161</td>
<td valign="top" align="center">86</td>
</tr>
<tr>
<td valign="top" align="left">TDC60</td>
<td valign="top" align="center">2,220</td>
<td valign="top" align="center">41</td>
<td valign="top" align="center">78</td>
<td valign="top" align="center">68</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN16"><label>a</label><p><italic>The strains were ordered somewhat according to the 16S rRNA phylogenetic tree shown in Figure <xref ref-type="fig" rid="F1">1</xref></italic>.</p></fn>
<fn id="TN17"><label>b</label><p><italic>The unique proteins were identified by &#x0201C;blastclust&#x0201D; program with parameters 80% as the sequence identity and 50% alignment length</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p><bold>Non-hypothetical unique<xref ref-type="table-fn" rid="TN18"><sup>a</sup></xref> proteins in 19 <italic>P. gingivalis</italic> genomes</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Strain</bold></th>
<th valign="top" align="left"><bold>Annotation</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">HG66</td>
<td valign="top" align="left">Glyoxalase</td>
</tr>
<tr>
<td valign="top" align="left">A7A1-28</td>
<td valign="top" align="left">Beta-galactosidase; putative hydrolase or acyltransferase of alpha/beta superfamily</td>
</tr>
<tr>
<td valign="top" align="left">Ando</td>
<td valign="top" align="left">DNA polymerase III subunits gamma and tau, partial external scaffolding protein D replication-associated protein A major spike protein G</td>
</tr>
<tr>
<td valign="top" align="left">F0568</td>
<td valign="top" align="left">DGQHR domain protein</td>
</tr>
<tr>
<td valign="top" align="left">F0569</td>
<td valign="top" align="left">Toxin-antitoxin system, toxin component, Fic domain protein</td>
</tr>
<tr>
<td valign="top" align="left">W4087</td>
<td valign="top" align="left">CAAX amino terminal protease family protein phage portal protein, SPP1 family phage uncharacterized protein</td>
</tr>
<tr>
<td valign="top" align="left">F0185</td>
<td valign="top" align="left">Peptidase S24-like protein</td>
</tr>
<tr>
<td valign="top" align="left">JCVI_SC001</td>
<td valign="top" align="left">Thioesterase family protein, partial starch-binding protein, SusD-like domain protein, partial spermine/spermidine synthase, partial phage portal protein, lambda family, partial head to-tail joining protein W serine carboxypeptidase domain protein, partial NYN domain protein imidazoleglycerol-phosphate dehydratase domain protein, partial carbohydrate kinase, PfkB domain protein PF13785 domain protein, partial DNA-binding helix-turn-helix protein</td>
</tr>
<tr>
<td valign="top" align="left">SJD2</td>
<td valign="top" align="left">Transposase ISPsy14</td>
</tr>
<tr>
<td valign="top" align="left">AJW4</td>
<td valign="top" align="left">Geranylgeranyl pyrophosphate synthase T5orf172 domain-containing protein</td>
</tr>
<tr>
<td valign="top" align="left">A7436</td>
<td valign="top" align="left">Transposase</td>
</tr>
<tr>
<td valign="top" align="left">W50</td>
<td valign="top" align="left">Transposase, mutator-like family protein</td>
</tr>
<tr>
<td valign="top" align="left">TDC60</td>
<td valign="top" align="left">Terminase</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN18"><label>a</label><p><italic>These proteins were searched against all the proteins in the 19 genomes and matched none but itself at the default BLASTP 2.2.25 parameter (i.e., with expected e value &#x02264; 10) (Altschul et al., <xref ref-type="bibr" rid="B1">1997</xref>)</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Phylogenomics by homologous proteins</title>
<p>Once a group of putative core proteins is identified, they can be concatenated and aligned together and used for compiling a phylogenomic tree to infer a possible evolutionary relationship at a level closer to the species than just any single gene. In this analysis, the 1,045 proteins shared by all 19 genomes at 60% sequence identity and 90% alignment cutoffs (Figure <xref ref-type="fig" rid="F1">1A</xref>), were first aligned individually with the &#x0201C;mafft&#x0201D; software (Katoh and Standley, <xref ref-type="bibr" rid="B23">2013</xref>). Each of the 1,045 protein sets contained exactly 19 aligned sequences, one from each of the 19 genomes. The aligned proteins were concatenated in the same protein order. This generated a set of 19 mega protein sequences with each consisting of 1,045 concatenated aligned sequences. The poorly aligned sequence regions, including leading and trailing unaligned portions of the sequences, as well as low-confidence parts of the alignment, such as positions that contain many gaps, were removed with the &#x0201C;Gblock&#x0201D; tool (V 0.91) (Talavera and Castresana, <xref ref-type="bibr" rid="B49">2007</xref>). After the Gblock screening, a final set of 19 aligned protein sequences, each with a length of 395,174 amino acids were used for constructing an unrooted tree. However, among the 395,174 aligned amino acids, only 17,389 positions had at least two different amino acids across proteins of all 19 <italic>P. gingivalis</italic> genomes, the remaining 377,785 were all the same amino acids across all genomes. Thus, only those 17,389 informative or effective positions contributed to the pairwise distances calculated among all genomes. Figure <xref ref-type="fig" rid="F4">4A</xref> is the result of the unrooted tree compiled based on the 1,045 shared proteins processed as described above. The overall topology is quite different from that of the <italic>16S rRNA</italic> tree (Figure <xref ref-type="fig" rid="F1">1</xref>) with the exception of two very closely related groups of strains, one consists of strains 381, ATCC 33277, and HG66 and another A7436, W50, and W83. This is not surprising because both groups have members with identical <italic>16S rRNA</italic> sequences hence their shared protein sequences are closer to each other in the group than other genomes.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold><italic>P. gingivalis</italic> phylogenomic trees based on core proteins identified at various percent sequence identities</bold>. Of the 39,926 NCBI annotated <italic>P. gingivalis</italic> proteins, 37,667 are &#x02265; 50 amino acids in length and were searched for homologous clusters using the &#x0201C;blastclust&#x0201D; software V.2.2.25 (<ext-link ext-link-type="uri" xlink:href="http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html">http://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html</ext-link>). <bold>(A)</bold> unrooted tree based on the 1,045 shared proteins identified by &#x0201C;blastclust&#x0201D; with 60% as the sequence identity and 90% as the alignment length cutoffs; the alignment generated a total of 17,389 effective (non-identical) protein sequence positions across all 19 genomes and the tree was constructed based on these positions; <bold>(B)</bold> rooted tree based on 436 proteins (out of 1,045) that are also found in <italic>P. asaccharolytica</italic> strain DSM 20707 (PaDSM20707) with &#x02265; 50% sequence identity and &#x02265; 90% alignment length; the alignment generated 4,771 effective protein sequence positions; <bold>(C)</bold> rooted tree based on 36 proteins shared among 20 genomes with &#x02265; 80% sequence identity and &#x02265; 90% alignment length. Proteins were aligned with MAFFT v6.935b (2012/08/21) (Katoh and Standley, <xref ref-type="bibr" rid="B23">2013</xref>) and poorly aligned regions were filtered by Gblocks 0.91b (Talavera and Castresana, <xref ref-type="bibr" rid="B49">2007</xref>). Trees were constructed with FastTree 2.1.9 (Price et al., <xref ref-type="bibr" rid="B40">2010</xref>) using the JTT protein mutation model (Jones et al., <xref ref-type="bibr" rid="B21">1992</xref>) and CAT&#x0002B;&#x02013;gemma options to account for the different rates of evolution at different sites. The reliability of tree splits were reported as &#x0201C;local support values&#x0201D; based on Shimodaira-Hasegawa test (Shimodaira and Hasegawa, <xref ref-type="bibr" rid="B45">2001</xref>) and are printed in blue on the split. The branch length (substitution rate) of the outgroup PaDSM20707 was truncated and the length were printed in black <bold>(B,C)</bold>; <bold>(D)</bold> Rooted tree constructed using PhyloPhlAn (Segata et al., <xref ref-type="bibr" rid="B43">2013</xref>) by directly subjecting all NCBI annotated proteins of the 20 genomes to the software, resulting in 840 effective protein positions from 225 aligned proteins.</p></caption>
<graphic xlink:href="fcimb-07-00028-g0004.tif"/>
</fig>
<p>To test whether including proteins from the outgroup species will result in a tree more similar to that of <italic>16S rRNA</italic>, i.e., a tree that is rooted at a potential common ancestor for these strains, ortholog candidates were first identified from the genome of <italic>P. asaccharolytica</italic> DSM 20707, of which the <italic>16S rRNA</italic> sequence was also used for the <italic>16S rRNA</italic> tree. At 90% alignment length cutoff, the number of homologous proteins in <italic>P. asaccharolytica</italic> decreases as the percent sequence identity cutoff increases. The numbers of protein homologous to any of the 1,045 core proteins used for the unrooted tree above are 436, 271, 146, 36, 7, 1, and 0 respectively for percent identity cutoffs 50, 60, 70, 80, 85, 90, and 95%. Figures <xref ref-type="fig" rid="F4">4B,C</xref> are the two rooted phylogenetic trees constructed based on the 436 (50% identity) and 36 (80% identity) proteins shared between <italic>P. asaccharolytica</italic> DSM 20707 and all 19 <italic>P. gingivalis</italic> strains. After Gblocks screening, the length of the aligned sequences were 12,646 (80% identity) and 177,272 (50% identity) amino acids respectively and the number of effective amino acids positions are 154 and 4,771 respectively. In general, the branch lengths increased with more effective amino acids positions which resulted in greater distances. Again, the only consistent close clusters were the two grouped with identical <italic>16S rRNA</italic>, i.e., the group of 381, ATCC 32277, and HG66, and of A7436, W50, and W83.</p>
<p>Figure <xref ref-type="fig" rid="F4">4D</xref> is the rooted tree constructed using the software PhyloPhlAn (Segata et al., <xref ref-type="bibr" rid="B43">2013</xref>) version 0.99 (8 May 2013). All 41,625 proteins annotated for the 20 genomes were subject to PhyloPhlAn with the default parameters that excluded proteins shorter than 30 amino acids in length. PhyloPhlAn finds among the input protein matches to a pre-set of the 400 most conserved proteins for extracting the phylogenetic signals. A total of 264 query proteins were matched to the 400 preset core but only 225 were present in all 20 genomes. These proteins were then aligned individually and subsampled based on a sophisticated procedure provided by PhyloPhlAn, which emphasizes regions both universally conserved and phylogenetically discriminating. The final aligned, subsampled, and concatenated sequences had a length of 3,082 aligned amino acids with 840 effective positions. The PhyloPhlAn tree is shown in Figure <xref ref-type="fig" rid="F4">4D</xref>. Similar to the two rooted trees (Figures <xref ref-type="fig" rid="F4">4B,C</xref>) and the 16S rRNA tree (Figure <xref ref-type="fig" rid="F1">1</xref>) the PhyloPhlAn tree also placed the three strains ATCC 33277, 381, and HG66 closest (but much closer) to the root and the remaining strains in a more linearly nested topology.</p>
<p>In summary, the only consensus based on interpretation of the three rooted protein trees and the <italic>16S rRNA</italic> tree is that the group ATCC 33277, 381, and HG66 is less evolved and closest to the common ancestor of this species (inferred based on the distance to the root). Strains W83, W50, and A7436 consistently formed a close group regardless of how the trees were built, but their exact phylogenetic position is inconclusive based on these analyses. Strains F0185, F0568, F0570, W4087, and MP4-504 are also found to be in the proximity of each other, although not as close as the two groups mentioned above. In general more effective/informative aligned amino acid positions resulted in longer branches and pairwise distances. To this end, the unrooted tree (Figure <xref ref-type="fig" rid="F4">4A</xref>) has the best resolution to reveal the similarity/differences among these strains, in the most genome-wide manner. Until a group of true orthologous proteins are identified (together with the outgroup) a true phylogenetic tree that infers the evolutionary path for this species will not be accessible.</p>
</sec>
</sec>
<sec id="s4">
<title>Comparisons based on whole-genome nucleotide sequences</title>
<sec>
<title>NUCmer nucleotide plots</title>
<p>Moving up the scale for comparison, one possible way is the whole genome nucleotide alignment with a commonly used software NUCleotide MUMmer (NUCmer), which identified nucleotide MUMs&#x02013;minimal unique matches between two genomic sequences (Delcher et al., <xref ref-type="bibr" rid="B9">2002</xref>). Figure <xref ref-type="fig" rid="F5">5</xref> shows some of the pairwise alignment results of the 19 <italic>P. gingivalis</italic> genomes. Figure <xref ref-type="fig" rid="F5">5A</xref> is the nucleotide alignment between strains 381 and ATCC 33277 and the almost perfect diagonal high similarity (red) match line indicates highly similar sequences, with only two visible exceptions &#x02013; one inversion and one insertion (to 381)/deletion (to ATCC 33277). Interestingly the inverted sequence almost matches the inserted sequence; apparently the inverted sequence was duplicated in the 381 genome and inserted somewhere else in the genome, where the ATCC 33277 genome shows no counterpart. The high DNA sequence similarity between 381 and ATCC 33277 is also supported by the identical <italic>16S rRNA</italic> gene sequence and copy numbers (Figure <xref ref-type="fig" rid="F1">1</xref>) as well as the protein-based phylogenetic relationships (Figure <xref ref-type="fig" rid="F4">4</xref>), even though their genomes are not far from identical. The phenomenon that a fairly large chunk of genomic sequence was duplicated and inserted elsewhere in the genome is only observed in strain 381, as evidenced by the NUCmer self-alignment of its genome (data available from the FTP site), but very similar to the alignment between 381 and ATCC 33277). No duplication event was observed in the self-alignment of the other 18 genomes.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>DNA-DNA sequence alignment between <italic>P. gingivalis</italic> genomes</bold>. Genomic sequence alignment between several pairs of <italic>P. gingivalis</italic> strains were plotted using NUCmer (NUCleotide MUMmer) version 3.1 (Delcher et al., <xref ref-type="bibr" rid="B9">2002</xref>). The sequence percent identities of detected homologous fragments were plotted in gradient colors based on the percentage. The axes are the nucleotide coordination in the genomes. The orders of the contigs in the unfinished genomes were rearranged based on the reference genome (genome on X- axis). <bold>(A)</bold> strain 381 vs. ATCC 33277; <bold>(B)</bold> HG66 vs. ATCC 33277; <bold>(C)</bold> strain 381 vs. HG66; <bold>(D)</bold> W50 vs. W83; <bold>(E)</bold> A7436 vs. W83; <bold>(F)</bold> AJW4 vs. A7436; <bold>(G)</bold> TDC60 vs. JCVI SC001; and <bold>(H)</bold> TDC60 vs. JCVI SC001 showing only the region with percent identity &#x02265; 99%.</p></caption>
<graphic xlink:href="fcimb-07-00028-g0005.tif"/>
</fig>
<p>Strain HG66 is the genome that is closest to 381 and ATCC 33277 based on <italic>16S rRNA</italic> genes and protein sequences, on the other hand it shows the disconnected high similarity match lines, which indicates more large-scale genomic arrangement between the two close strains &#x02013; between 381 and HG66 (Figure <xref ref-type="fig" rid="F5">5B</xref>) and between ATCC 33277 and HG66 (Figure <xref ref-type="fig" rid="F5">5C</xref>).</p>
<p>The second closest groups of strains are A7436, W50, and W83 and their nucleotide sequences are also highly similar based on the NUMMER plots (Figures <xref ref-type="fig" rid="F5">5D,E</xref>). However, the contigs of the unfinished draft genome of W50 were rearranged by NUCmer in the order based on the similarity to the W83 sequence. Whether there is a large scale genomic rearrangement between W83 and W50 cannot be known until the genome of W50 is completed. Strain A7436, a finished genome, shows only one inversion of the genome when compared to that of W83 (Figure <xref ref-type="fig" rid="F5">5E</xref>). The fact that A7436 is not as close to W50 and W83 as the distance between HG66 and 381 (or ATCC 33277) based on <italic>16S rRNA</italic> and protein phylogeny (Figures <xref ref-type="fig" rid="F1">1</xref>, <xref ref-type="fig" rid="F4">4</xref>), suggests that the genomes of the group of HG66, 381, and ATCC 33277 have higher genomic sequence rearrangement activity than the A7436-W50-W83 group. The next genome which is closest to the A7436-W50-W83 group is strain AJW4, with several visible (larger fragments) of insertions/deletions and inversions when compared to A7436 (arrows heads in Figure <xref ref-type="fig" rid="F5">5F</xref>). This relationship is also consistent with the <italic>16S rRNA</italic> gene tree (Figure <xref ref-type="fig" rid="F1">1</xref>).</p>
<p>Another interesting observation is the alignment between JCVI SC001 and TDC60. These two strains are not among the closest groups based on the <italic>16S rRNA</italic> and protein sequences (Figures <xref ref-type="fig" rid="F1">1</xref>, <xref ref-type="fig" rid="F4">4</xref>). The NUCmer plot between these two genomes appears to be a straight diagonal red line (Figure <xref ref-type="fig" rid="F5">5G</xref>), similar to that between 381 and ATCC 33277. However, since the genomic sequence of JCVI SC001 was not really completed and closed to a circular chromosomal format, the 284 <italic>de novo</italic> assembled contigs were mapped to the genome of TDC60 and the gaps were filled with Ns to form a single pseudo-contig (Genbank Accession <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="CM001843">CM001843</ext-link>) (McLean et al., <xref ref-type="bibr" rid="B33">2013</xref>). Thus, the contig order in the published single contig genomic sequence of JCVI SC001 may not be correct and the sequence similarity between JCVI SC001 and TDC60 may not be as &#x0201C;straight&#x0201D; as indicated in the NUCmer plot. In fact when the plot was filtered to show only the region with percent identity &#x02265; 99%, the red line became fragmented with large gaps (Figure <xref ref-type="fig" rid="F5">5H</xref>), indicating that a large portion of the genomic sequences between these two strains are under 99% similarity.</p>
<p>The complete pair-wise NUCmer plots of the 19 <italic>P. gingivalis</italic> genomes can be viewed in a specifically designed interactive webpage at <ext-link ext-link-type="uri" xlink:href="http://bioinformatics.forsyth.org/publication/20160425">http://bioinformatics.forsyth.org/publication/20160425</ext-link>. The web page provides interactive tools to choose any two <italic>P. gingivalis</italic> genomes for the NUCmer results, as well as the possibility of viewing the alignment at various percent sequence identity cutoffs.</p>
</sec>
<sec>
<title>Oligonucleotide frequency</title>
<p>The NUCmer plots above are limited to viewing comparisons only between two given genomes. To view and compare nucleotide difference/similarity for all genomes on the same plot, the overall oligonucleotide composition and frequency can be measured along the entire genome and the results can be plotted out and visually compared to each other. This analysis started by collecting all the possible 20-mer sequences in all 20 genomes and then count for each 20-mer how many genomes have each particular sequence. The number of genomes (genomic frequency) for each 20-mer thus ranges from 1 (unique) to 20 (universal). The frequencies can be calculated and plotted along the entire genome by taking every 20-mer from the beginning to the end of the genome. Figure <xref ref-type="fig" rid="F6">6A</xref> depicts the results of the 20-mer oligonucleotide frequencies among all 20 genomes (including the out-group <italic>P. asaccharolytica</italic> DSM 20707). If a region of a genome is shared by all other 20 genomes, it is colored black; and if a region is unique to the genome itself, it is colored bright yellow. In other words, a black region means that all the possible 20-mer sequences appeared in all tested genomes, whereas the brightest yellow regions have unique 20-mer sequences that are only found in one genome. For easy comparison, the order or the genomes shown in Figures <xref ref-type="fig" rid="F6">6A,B</xref> were arranged according to that of the <italic>16S rRNA</italic> tree (Figure <xref ref-type="fig" rid="F1">1</xref>) with a dendrogram reflecting similar tree topology. As expected, the two closest strains ATCC 33277 and 381 share almost identical 20-mer frequency patterns, with the exception of a small insertion at nucleotide position ca. 1,400,000, which is also detected by the NUCmer plot in Figure <xref ref-type="fig" rid="F5">5A</xref>. The genome of strain 381 is ca. 24 Kbps longer than that of ATCC 33277 due to this insert and the length difference is illustrated in Figure <xref ref-type="fig" rid="F6">6A</xref> because the length of the bars were based on actual genome sizes. Another interesting example observation is that even though strain JCVI SC001 is closest to SDJ2 due to their identical <italic>16S rRNA</italic> sequences (Figure <xref ref-type="fig" rid="F1">1</xref>), their oligonucleotide frequency patterns are quite different, with each showing unique regions (brighter colors) at different places. This can be due to two possibilities: (1) the artificial order of the unfinished sequence contigs (in the plots, contig order was the same as that in the downloaded sequences); and (2) the bona fide differences in sequence. This is also true for the other three genomes A7436, W83, and W50, which share identical <italic>16S rRNA</italic> sequences but exhibit distinct frequency patterns.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Genomic DNA similarity of 19 <italic>P. gingivalis</italic> genomes compared by oligonucleotide frequency</bold>. All possible 20-mer sequences present in all genomes, including that of <italic>P. asaccharolytica</italic> strain DSM 20707 (PaDSM2070) used as an out-group, were categorized and the number of genomes in which a 20-mer is present, was recorded. <bold>(A)</bold> was generated by first calculating the average number of genomes for all the 20 mers present in every 500-nucleotide windows across the entire genome and then color each window based on the genome frequency (minimum 1 in yellow and maximum 20 in black). <bold>(B)</bold> was similar to <bold>(A)</bold> but the non-coding regions were masked with light blue color to highlight the oligonucleotide frequencies for the areas that correspond to both forward (upper) and reverse-complement (lower) protein coding sequences. The order of the unfinished genomic contigs was arranged in the same order as appeared in the sequences downloaded from NCBI. The genomes in the plot were ordered based on the 16S rRNA phylogenetic tree (Figure <xref ref-type="fig" rid="F1">1</xref>) with a dendrogram derived from the same tree to show the relatedness.</p></caption>
<graphic xlink:href="fcimb-07-00028-g0006.tif"/>
</fig>
<p>When the frequencies were plotted only in the open reading frames (ORFs), most of the ORFs appeared dark in color and are shared by the majority of the <italic>P. gingivalis</italic> genomes (Figure <xref ref-type="fig" rid="F6">6B</xref>). Those ORFs with colors close to yellow should account for the differences of the number of unique proteins previously identified (Figure <xref ref-type="fig" rid="F3">3</xref>). These differences in nucleotide sequences are essentially reflected in the genes and then translated to proteins, and ultimately accounting for differences in biological functions.</p>
<p>Finally the out-group <italic>P. asaccharolytica</italic> DSM 20707 shows mostly yellow colors in the plot, which is as expected and means that <italic>P. asaccharolytica</italic> does not share much of the 20-mer oligonucleotide sequences with <italic>P. gingivalis</italic>. Interestingly, by lowering the oligomer size to 14 bases (14-mer), most of the DSM 20707 genome appears black, meaning that these two different species share most of the 14-mer sequences (data not shown). When the plot was generated with 15-mer sequences, the <italic>P. asaccharolytica</italic> DSM 20707 genome started to show patches of yellow area (data not shown), indicating that some unique 15-mer sequences are present between these two species. With 15-mer all <italic>P. gingivalis</italic> genomes are black in the plot (data not shown), meaning 15-mer is too short and have not enough resolution power to differentiate unique regions among <italic>P. gingivalis</italic> genomes. Hence whether the oligomer frequency analysis can detect unique/shared regions in a group of genomes, depends on the size of the oligomer. The choice of 20-mer was able to identify unique regions among the strains of <italic>P. gingivalis</italic>, as shown in Figures <xref ref-type="fig" rid="F6">6A,B</xref>, yet is too sensitive for a different species.</p>
</sec>
</sec>
<sec id="s5">
<title>Comparative functional genomics</title>
<p>The comparative genomics is less meaningful without association with biological functions. Most functional genomic annotations rely on either DNA or protein sequence homology to other sequences with known biological functions. The most popular genome annotation pipeline is probably the NCBI Prokaryotic Genome Annotation Pipeline (Tatusova et al., <xref ref-type="bibr" rid="B50">2016</xref>), which is the current default annotation pipeline when a microbial genome sequence is deposited to and published in NCBI. Several other microbial genome annotation pipelines have also been published and commonly used, including the RAST system&#x02014;Rapid Annotation of microbial genomes using Subsystems Technology (Overbeek et al., <xref ref-type="bibr" rid="B39">2014</xref>); the BASys&#x02014;Bacterial Annotation System, a web server for automated bacterial genome annotation (Van Domselaar et al., <xref ref-type="bibr" rid="B53">2005</xref>). Above the gene level, there have also been tools and databases available for constructing and comparing metabolic pathways of microbial genomes. Examples in this category are IMG&#x02014;the integrated microbial genomes comparative analysis system (Markowitz et al., <xref ref-type="bibr" rid="B31">2014</xref>) and BlastKOALA&#x02014;a KEGG tool for functional characterization of genome sequences (Kanehisa et al., <xref ref-type="bibr" rid="B22">2016</xref>). Both systems provide annotation information beyond individual gene and protein level, such as, in the case of IMG, conserved protein domain and groups COGs and families (Pfam), as well as the enzymes and metabolic pathways inferred by BlastKOALA.</p>
<p>In this report we compared the <italic>P. gingivalis</italic> genomes at the functional level based on three systems: The NCBI annotation, the RAST annotation, and the BlastKOALA inferred metabolic pathways. The results of these analyses are too voluminous to be presented in text however the complete results are provided in a central online site for download (<ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425">ftp://www.homd.org/publication_data/20160425</ext-link>). Here we summarized all the comparisons into a single table (Table <xref ref-type="table" rid="T8">8</xref>). Initially, functional comparisons were done based on simple text search&#x02014;by counting the number of genes with functional annotations containing several categories of keywords listed in the table (in Italic font).</p>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p><bold>Comparative functional genomics of <italic>P. gingivalis</italic> genomes<xref ref-type="table-fn" rid="TN19"><sup>a</sup></xref><sup>,</sup><xref ref-type="table-fn" rid="TN20"><sup>b</sup></xref><sup>,</sup><xref ref-type="table-fn" rid="TN21"><sup>c</sup></xref></bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Annotation Source</bold></th>
<th valign="top" align="center"><bold>ATCC33277</bold></th>
<th valign="top" align="center"><bold>HG66</bold></th>
<th valign="top" align="center"><bold>381</bold></th>
<th valign="top" align="center"><bold>W83</bold></th>
<th valign="top" align="center"><bold>W50</bold></th>
<th valign="top" align="center"><bold>A7436</bold></th>
<th valign="top" align="center"><bold>AJW4</bold></th>
<th valign="top" align="center"><bold>F0570</bold></th>
<th valign="top" align="center"><bold>JCVISC001</bold></th>
<th valign="top" align="center"><bold>SJD2</bold></th>
<th valign="top" align="center"><bold>F0568</bold></th>
<th valign="top" align="center"><bold>F0569</bold></th>
<th valign="top" align="center"><bold>Ando</bold></th>
<th valign="top" align="center"><bold>F0185</bold></th>
<th valign="top" align="center"><bold>W4087</bold></th>
<th valign="top" align="center"><bold>MP4-504</bold></th>
<th valign="top" align="center"><bold>A7A1-28</bold></th>
<th valign="top" align="center"><bold>F0566</bold></th>
<th valign="top" align="center"><bold>TDC60</bold></th>
<th valign="top" align="center"><bold>PaDSM20707</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>GINGIPAIN:</bold> <italic><bold>rgp, kgp, gingipain</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BLAST<xref ref-type="table-fn" rid="TN22"><sup>d</sup></xref></td>
<td valign="top" align="center"><bold>7</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>7</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>4</bold></td>
<td valign="top" align="center"><bold>2</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>4</bold></td>
<td valign="top" align="center"><bold>4</bold></td>
<td valign="top" align="center"><bold>4</bold></td>
<td valign="top" align="center"><bold>4</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>0</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>ATTACHMENT:</bold> <italic><bold>adhesin, fim, pili, pilus, fimbriae, fimbrilin</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BLAST</td>
<td valign="top" align="center"><bold>16</bold></td>
<td valign="top" align="center"><bold>17</bold></td>
<td valign="top" align="center"><bold>18</bold></td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>11</bold></td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>17</bold></td>
<td valign="top" align="center"><bold>12</bold></td>
<td valign="top" align="center"><bold>16</bold></td>
<td valign="top" align="center"><bold>11</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>12</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>11</bold></td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>0</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>HEME:</bold> <italic><bold>heme, haga, hagb, hagc, hemaglu, hemoglo</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">BLAST</td>
<td valign="top" align="center"><bold>10</bold></td>
<td valign="top" align="center"><bold>11</bold></td>
<td valign="top" align="center"><bold>10</bold></td>
<td valign="top" align="center"><bold>8</bold></td>
<td valign="top" align="center"><bold>8</bold></td>
<td valign="top" align="center"><bold>10</bold></td>
<td valign="top" align="center"><bold>10</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>7</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>7</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>8</bold></td>
<td valign="top" align="center"><bold>8</bold></td>
<td valign="top" align="center"><bold>8</bold></td>
<td valign="top" align="center"><bold>7</bold></td>
<td valign="top" align="center"><bold>10</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>GENE MOBILITY:</bold> <italic><bold>transposon, ISPg, transposase, conjugation, insertion element</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">118</td>
<td valign="top" align="center">68</td>
<td valign="top" align="center">94</td>
<td valign="top" align="center">73</td>
<td valign="top" align="center">20</td>
<td valign="top" align="center">98</td>
<td valign="top" align="center">73</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">30</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">35</td>
<td valign="top" align="center">23</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">26</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">26</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">65</td>
<td valign="top" align="center">32</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">56</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">35</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">69</td>
<td valign="top" align="center">42</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">38</td>
<td valign="top" align="center">28</td>
<td valign="top" align="center">40</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">87</td>
<td valign="top" align="center">46</td>
<td valign="top" align="center">61</td>
<td valign="top" align="center">51</td>
<td valign="top" align="center">24</td>
</tr>
<tr>
<td valign="top" align="left">BLAST</td>
<td valign="top" align="center"><bold>131</bold></td>
<td valign="top" align="center"><bold>133</bold></td>
<td valign="top" align="center"><bold>139</bold></td>
<td valign="top" align="center"><bold>138</bold></td>
<td valign="top" align="center"><bold>46</bold></td>
<td valign="top" align="center"><bold>149</bold></td>
<td valign="top" align="center"><bold>132</bold></td>
<td valign="top" align="center"><bold>46</bold></td>
<td valign="top" align="center"><bold>71</bold></td>
<td valign="top" align="center"><bold>54</bold></td>
<td valign="top" align="center"><bold>66</bold></td>
<td valign="top" align="center"><bold>45</bold></td>
<td valign="top" align="center"><bold>40</bold></td>
<td valign="top" align="center"><bold>43</bold></td>
<td valign="top" align="center"><bold>34</bold></td>
<td valign="top" align="center"><bold>110</bold></td>
<td valign="top" align="center"><bold>68</bold></td>
<td valign="top" align="center"><bold>72</bold></td>
<td valign="top" align="center"><bold>89</bold></td>
<td valign="top" align="center"><bold>37</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>TRANSPOSASE, IS5 FAMILY; K07481<xref ref-type="table-fn" rid="TN23"><sup>e</sup></xref></bold></td>
</tr>
<tr>
<td valign="top" align="left">KEGG Orthology</td>
<td valign="top" align="center">47</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>CAPSULE:</bold> <italic><bold>capsul</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BLAST</td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>CRISPR:</bold> <italic><bold>crispr</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">7</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">7</td>
</tr>
<tr>
<td valign="top" align="left">BLAST</td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>0</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>5</bold></td>
<td valign="top" align="center"><bold>6</bold></td>
<td valign="top" align="center"><bold>14</bold></td>
<td valign="top" align="center"><bold>11</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>7</bold></td>
</tr>
<tr>
<td valign="top" align="left">CRISPR arrays<xref ref-type="table-fn" rid="TN24"><sup>f</sup></xref></td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left" colspan="21" style="background-color:#bbbdc0"><bold>PHAGE:</bold> <italic><bold>phage</bold></italic></td>
</tr>
<tr>
<td valign="top" align="left">NCBI</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">9</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">RAST</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">BLAST</td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>15</bold></td>
<td valign="top" align="center"><bold>20</bold></td>
<td valign="top" align="center"><bold>18</bold></td>
<td valign="top" align="center"><bold>19</bold></td>
<td valign="top" align="center"><bold>22</bold></td>
<td valign="top" align="center"><bold>19</bold></td>
<td valign="top" align="center"><bold>25</bold></td>
<td valign="top" align="center"><bold>19</bold></td>
<td valign="top" align="center"><bold>18</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>17</bold></td>
<td valign="top" align="center"><bold>18</bold></td>
<td valign="top" align="center"><bold>25</bold></td>
<td valign="top" align="center"><bold>17</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>13</bold></td>
<td valign="top" align="center"><bold>12</bold></td>
<td valign="top" align="center"><bold>3</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN19"><label>a</label><p><italic>Results were compiled based on the NCBI or RAST genome annotations. Total number of proteins containing any of the keywords shown in each category were recorded for each genome and for NCBI and RAST annotations separately. The detail results are provided in the Supplemental Files available from the FTP site: <ext-link ext-link-type="uri" xlink:href="ftp://bioinformatics.forsyth.org/publication_data/20160425/">ftp://bioinformatics.forsyth.org/publication_data/20160425/</ext-link></italic></p></fn>
<fn id="TN20"><label>b</label><p><italic>The keyword search was performed in a case-insensitive manner and allowed matching of the partial word</italic>.</p></fn>
<fn id="TN21"><label>c</label><p><italic>The order of genomes was based on that similar to the 16S rRNA phylogenetic tree</italic>.</p></fn>
<fn id="TN22"><label>d</label><p><italic>BLAST: all the proteins identified by NCBI and RAST were collected and the sequences searched against all the proteins of all 20 genomes using BPLSTP. The numbers (in bold) indicated for each genome are the number of proteins with &#x02265; 95% sequence identity and &#x02265; 95% coverage of the query sequences. The numbers were calculated separatly for NCBI and RAST annotated proteins, and the larger number of the two are shown in this table</italic>.</p></fn>
<fn id="TN23"><label>e</label><p><italic>The number of proteins related to the IS5 transposase family was identified by the BlastKOALA program (Kanehisa et al., <xref ref-type="bibr" rid="B22">2016</xref>) with the matching to the KEGG Orthology (KO) number K0748</italic>.</p></fn>
<fn id="TN24"><label>f</label><p><italic>The number of CRISPR arrays detected by the online software CRISPRfinger (<ext-link ext-link-type="uri" xlink:href="http://crispr.i2bc.paris-saclay.fr/Server/">http://crispr.i2bc.paris-saclay.fr/Server/</ext-link>); only the number of &#x0201C;confirmed&#x0201D; candidates were reported thus excluding those &#x0201C;questionable&#x0201D; ones, which only have two DR and one spacer sequences</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Interestingly marked differences were observed in the NCBI and RAST annotations either by total protein count or by keyword searches. For example, as shown in Table <xref ref-type="table" rid="T4">4</xref>, the difference of the total number of protein encoding genes annotated by the two systems can be as large as 351 (for strain F0566). The difference in functional annotation is also quite noticeable (Table <xref ref-type="table" rid="T8">8</xref>). For example, only five of the 19 <italic>P. gingivalis</italic> genomes have proteins annotated as &#x0201C;gingipain&#x0201D; by NCBI, whereas three other different genomes were annotated by RAST to have a single gingipain gene.</p>
<p>To remedy the differences and apparent incompleteness of the two annotation systems, a more effective way to detect most, if not all, proteins of the same function, is to perform sequence similarity searches using protein sequences that had been identified. For example, the 19 proteins that were annotated as &#x0201C;gingipain&#x0201D; (16 by NCBI, three by RAST, Table <xref ref-type="table" rid="T8">8</xref>) were grouped together and used as the baits to search against all proteins of all 20 genomes identified by both systems. A total of 84 proteins highly similar to the 19 gingipain proteins were detected this way among all 19 <italic>P. gingivalis</italic> genomes ranging from two to seven gingipains per genome (none was detected in <italic>P. asaccharolytica</italic>). These searches were conservative by setting a high percent sequence identity and coverage, and so the numbers can be under-estimated. This approach was done repeatedly for each of the seven functional categories that were deemed of high interest by authors. All the proteins identified by either NCBI or RAST in each category were collectively searched against all protein sequences in all 20 and the number of proteins with &#x02265; 95% sequence identity and &#x02265; 95% alignment coverage to the query sequences were recorded. The results of the BLAST searches were listed in the third row of each category in Table <xref ref-type="table" rid="T8">8</xref>. Unsurprisingly, the number of genes identified in all categories is higher than those provided by either annotation system, and often higher than both systems combined. The fact that stringent BLAST search identified more proteins of the same function indicates that the currently microbial genome annotation pipelines are quite incomprehensive and are in need of improvement.</p>
<p>For gingipains, using the 16 NCBI identified proteins and three RAST ones (Table <xref ref-type="table" rid="T8">8</xref>), the BLAST search of these sequences matched with many more proteins that are highly similar to gingipains in all 19 <italic>P. gingivalis</italic> genomes. Examining the annotation for those proteins highly matched with annotated gingipains, most of them were simply annotated as &#x0201C;hypothetical&#x0201D; or &#x0201C;functionally unknown&#x0201D; proteins, while some were annotated as &#x0201C;peptidase&#x0201D;.</p>
<p>Another notable observation is the high prevalence of the transposase proteins encoded in this species, as high as 149 copies in strain A7436. The lower number of transposases detected in those unfinished genomes is most likely due to the in-between-contig sequence gaps that may contain highly repeated sequences such as the transposases and the IS elements. The completed genome with lowest number of mobility related genes is strain A7A1-28 where only 68 were detected in the genome.</p>
<p>Capsular polysaccharide (CPS) has long been recognized as an important virulence factor for <italic>P. gingivalis</italic> (Singh et al., <xref ref-type="bibr" rid="B47">2011</xref>) and encapsulated strains are known to be more virulent than the non-encapsulated ones (Laine and van Winkelhoff, <xref ref-type="bibr" rid="B27">1998</xref>). When all the annotated capsule related proteins were BLASTP searched against all genomes, the total number of capsule related proteins ranged consistently between five and six copies (Table <xref ref-type="table" rid="T8">8</xref>). For example, W83 is known as an encapsulated strain and ATCC 33277 is non-encapsulated. However, both strains encode six copies of capsulated related genes. Of these, four were annotated as &#x0201C;CPS/capsule biosynthesis proteins&#x0201D; by both NCBI and RAST. Interestingly, NCBI only identified three of these four CPS biosynthesis protein. The 5th one was annotated as &#x0201C;CPS transport protein&#x0201D; in W83 by NCBI (Genbank ID <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AAQ65636.1">AAQ65636.1</ext-link>) but was annotated as &#x0201C;conserved hypothetic protein&#x0201D; in ATCC 33277 by NCBI (BAG34043.1) or &#x0201C;tyrosine-protein kinase Wzc&#x0201D; in both W83 and ATCC 33277 by RAST. The 6th capsule related gene was annotated as &#x0201C;sugar isomerase&#x0201D; in ATCC 33277 (BAG34552.1) or &#x0201C;SIS domain protein&#x0201D; in W83 (AAQ65335.1) by NCBI. This same gene was annotated as &#x0201C;arabinose 5-phosphate isomerase&#x0201D; in both W83 and ATCC 3327 by RAST and &#x0201C;sugar phosphate isomerase involved in capsule formation&#x0201D; in several other strains by NCBI. Taken together, this serves as an example of how inconsistent both annotations are, for genes involved in a single biological function. By BLAST searching using proteins annotated as capsule related genes annotated across all 19 <italic>P. gingivalis</italic> genomes, we were able to detect consistently between five to six copies of genes involved in encapsulation for this species. The fact the all <italic>P. gingivalis</italic> genomes contain a similar number of capsule related genes yet some are encapsulated and others are not, indicates that these genes may subject to different gene expression controls. It is thus likely that some non-encapsulated strains may become encapsulated under certain specific <italic>in vivo</italic> conditions.</p>
<p>In a very different functional aspect, there is a high prevalence of the bacterial phage related proteins, such as phage integrase/site-specific recombinase, phage tail component proteins, and phage-related lysozyme. The number of phage related proteins detected in the 19 <italic>P. gingivalis</italic> genomes ranged from 12 to 25. Functional bacteriophage have so far never been detected in this species (Sandmeier et al., <xref ref-type="bibr" rid="B42">1993</xref>) yet contrarily many proteins related to phage reproduction were detected in all the 19 <italic>P. gingivalis</italic> strains. One most plausible explanation is the prevalence of the CRISPR/Cas systems in this species (discussed below); another is also the presence of the abortive phage infection proteins found in several strains (ATCC 33277, HG66, W83, AJW4, SJD2, and MP4-504, data not shown).</p>
<p>As mentioned above, another very interesting category of enzymes reported in Table <xref ref-type="table" rid="T8">8</xref> is the prevalence of proteins associated with the CRISPR (clustered regularly interspaced short palindromic repeats) elements. CRISPR, together with the Cas (CRISPR associated) proteins, have been dubbed as the adaptive immune system for Bacteria and Archaea to ward off invading foreign DNA (Horvath and Barrangou, <xref ref-type="bibr" rid="B19">2010</xref>). However, although CRISPR arrays were detected in all genomes (including outgroup <italic>P. asaccharolytica</italic>, 4th row in the CRISPR category of Table <xref ref-type="table" rid="T8">8</xref>), that is not the case for the Cas proteins. Cas was not detected in the genome of strain JCVI SC001, and only one copy detected in strain AJW4 (3rd row in Table <xref ref-type="table" rid="T8">8</xref> CRISPR category). Strain F0569 has the highest number or CRISPR arrays detected using the online software CRISPRfinger (<ext-link ext-link-type="uri" xlink:href="http://crispr.i2bc.paris-saclay.fr/Server">http://crispr.i2bc.paris-saclay.fr/Server</ext-link>) but this strain does not have the highest number of Cas proteins. Of all the CRISPR arrays detected, the length of the direct repeat (DR) element ranged from 23 to 47 bps and the number of the DRs in the array ranged from 5 to 121 copies (data not shown but available from the online FTP site). Both ATCC 33277 and strain 381 had three copies of nearly identical CRISPR arrays and both had one copy of the arrays with 121 DR sequences (and 120 spacer sequences). The high DR copy number may be an indication for the CRISPR activity in the past. On the other hand, JSVI SC001 had three copies of CRISPR arrays detected with DR of 31, 26, and 45 bps and repeat number 5, 7, and 6 respectively. Whether or not this strain possesses a type of Cas protein that is very different from those in other strains remains to be investigated. If this strain lacks any functional Cas protein, it is likely to be susceptible to bacteriophage infection or the activation of the possible presence of prophages as evidenced by the detection of 25 copies of phage related proteins (Table <xref ref-type="table" rid="T8">8</xref>).</p>
<p>On the other side of the scale, at the metabolic pathway level, the KEGG pathways and KEGG Orthology identified by BlastKOALA are BLAST-based, i.e., all the proteins sequences regardless of their annotations, were BLAST-searched against the online protein database used by BlastKOALA (<ext-link ext-link-type="uri" xlink:href="http://www.kegg.jp/blastkoala/">http://www.kegg.jp/blastkoala/</ext-link>). Hence the comprehensiveness of the KEGG pathways and the KO terms inferred by BlastKOALA depend on the completeness of the proteins in the database. At any rate, several metabolic pathways identified to be unique to the species <italic>P. gingivalis</italic> are: Glycosphingolipid biosynthesis&#x02013;globoseries; sphingolipid metabolism; lysosome; glycosphingolipid biosynthesis&#x02014;ganglio series; and glycosaminoglycan degradation. These species specific pathways were suggested based on the fact that they were detected in all of the 19 <italic>P. gingivalis</italic> genomes but not in <italic>P. asaccharolytica</italic> DSM 20707 (detailed data available in the FTP site). When compared to <italic>P. asaccharolytica</italic> DSM 20707, BlastKOALA determined that <italic>P. gingivalis</italic> lacks proteins involved in the following pathways: C5-branched dibasic acid metabolism; AMPK signaling pathway; amoebiasis; thyroid hormone synthesis; apoptosis; and arachidonic acid metabolism. Note that the specific-specific observations made above were only based on using a different species of the same genus. Inclusion of more outgroups would certainly reduce the number of species-specific functions or pathways.</p>
</sec>
<sec id="s6">
<title>Comparison with other species</title>
<p>The focus of this paper is the within-species comparison for the different strains of <italic>P. gingivalis</italic> (PG). It is however also very interesting and important to compare to other related species. Due to the great amount of data generated already from the within-species comparisons, a full multi-species comparison study will be reported in a separate publication. Here we present only a brief summary of the comparison of <italic>P. gingivalis</italic> genomes with several selected species (Table <xref ref-type="table" rid="T9">9</xref>).</p>
<table-wrap position="float" id="T9">
<label>Table 9</label>
<caption><p><bold>Genome statistics of selected species<xref ref-type="table-fn" rid="TN25"><sup>a</sup></xref></bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Category</bold></th>
<th valign="top" align="center"><bold>PG</bold></th>
<th valign="top" align="center"><bold>TF</bold></th>
<th valign="top" align="center"><bold>TD</bold></th>
<th valign="top" align="center"><bold>AA</bold></th>
<th valign="top" align="center"><bold>BU</bold></th>
<th valign="top" align="center"><bold>BF</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Number of genomes</td>
<td valign="top" align="center">19</td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">17</td>
<td valign="top" align="center">38</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">103</td>
</tr>
<tr>
<td valign="top" align="left">Number of genomes with single contig</td>
<td valign="top" align="center">9</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">Number of contigs (genomes with &#x0003E; 1 contig)</td>
<td valign="top" align="center">92&#x02013;192</td>
<td valign="top" align="center">71&#x02013;141</td>
<td valign="top" align="center">2&#x02013;12</td>
<td valign="top" align="center">4&#x02013;787</td>
<td valign="top" align="center">8&#x02013;343</td>
<td valign="top" align="center">2&#x02013;2,566</td>
</tr>
<tr>
<td valign="top" align="left">Genome sizes</td>
<td valign="top" align="center">2,216&#x02013;2,441 kb</td>
<td valign="top" align="center">3,233&#x02013;3,405 kb</td>
<td valign="top" align="center">2,742&#x02013;2,990 kb</td>
<td valign="top" align="center">1,860&#x02013;2,382 kb</td>
<td valign="top" align="center">4,270&#x02013;5,047 kb</td>
<td valign="top" align="center">4,457&#x02013;8,029 kb</td>
</tr>
<tr>
<td valign="top" align="left">Mean genome size</td>
<td valign="top" align="center">2,320 kb</td>
<td valign="top" align="center">3,312 kb</td>
<td valign="top" align="center">2,838 kb</td>
<td valign="top" align="center">2,155 kb</td>
<td valign="top" align="center">4,771 kb</td>
<td valign="top" align="center">5,416 kb</td>
</tr>
<tr>
<td valign="top" align="left">Number of ORFs<xref ref-type="table-fn" rid="TN26"><sup>b</sup></xref></td>
<td valign="top" align="center">1,788&#x02013;2,417</td>
<td valign="top" align="center">2,492&#x02013;3,001</td>
<td valign="top" align="center">2,520&#x02013;2,793</td>
<td valign="top" align="center">1,829&#x02013;2,364</td>
<td valign="top" align="center">3,254&#x02013;4,663</td>
<td valign="top" align="center">3,593&#x02013;8,060</td>
</tr>
<tr>
<td valign="top" align="left">Mean ORF number</td>
<td valign="top" align="center">2,101</td>
<td valign="top" align="center">2,740</td>
<td valign="top" align="center">2,634</td>
<td valign="top" align="center">2,046</td>
<td valign="top" align="center">4,045</td>
<td valign="top" align="center">4,760</td>
</tr>
<tr>
<td valign="top" align="left">Number of core proteins<xref ref-type="table-fn" rid="TN27"><sup>c</sup></xref></td>
<td valign="top" align="center">1,037</td>
<td valign="top" align="center">1,560</td>
<td valign="top" align="center">1,129</td>
<td valign="top" align="center">424</td>
<td valign="top" align="center">1,191</td>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left">Number of unique proteins<xref ref-type="table-fn" rid="TN28"><sup>d</sup></xref></td>
<td valign="top" align="center">1,044</td>
<td valign="top" align="center">801</td>
<td valign="top" align="center">692</td>
<td valign="top" align="center">1,233</td>
<td valign="top" align="center">4,040</td>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left">Non-hypothetical proteins<xref ref-type="table-fn" rid="TN29"><sup>e</sup></xref> (percentage)</td>
<td valign="top" align="center">1,206&#x02013;1,637 (54.16&#x02013;74.26%)</td>
<td valign="top" align="center">1,573&#x02013;1,800 (58.77&#x02013;64.65%)</td>
<td valign="top" align="center">685&#x02013;1,618 (25.17&#x02013;58.47%)</td>
<td valign="top" align="center">1,127&#x02013;2,035 (48.43&#x02013;89.75%)</td>
<td valign="top" align="center">1,038&#x02013;3,156 (24.84&#x02013;79.48%)</td>
<td valign="top" align="center">677&#x02013;5,732 (13.99&#x02013;71.91%)</td>
</tr>
<tr>
<td valign="top" align="left">Non-hypothetical proteins: Mean (Percentage)</td>
<td valign="top" align="center">1,346 (64.60%)</td>
<td valign="top" align="center">1,702 (62.18%)</td>
<td valign="top" align="center">801 (30.38%)</td>
<td valign="top" align="center">1,656 (81.19%)</td>
<td valign="top" align="center">2,450 (60.71%)</td>
<td valign="top" align="center">2,993 (62.63%)</td>
</tr>
<tr>
<td valign="top" align="left">Hypothetical proteins<xref ref-type="table-fn" rid="TN29"><sup>e</sup></xref> (Percentage)</td>
<td valign="top" align="center">515&#x02013;1,108 (25.74&#x02013;45.84%)</td>
<td valign="top" align="center">919&#x02013;1,201 (35.35&#x02013;41.23%)</td>
<td valign="top" align="center">1,149&#x02013;2,090 (41.53&#x02013;74.83%)</td>
<td valign="top" align="center">211&#x02013;1,200 (10.25&#x02013;51.57%)</td>
<td valign="top" align="center">784&#x02013;3,149 (20.52&#x02013;75.16%)</td>
<td valign="top" align="center">1,193&#x02013;4,162 (28.09&#x02013;86.01%)</td>
</tr>
<tr>
<td valign="top" align="left">Hypothetical proteins: Mean (Percentage)</td>
<td valign="top" align="center">755 (35.40%)</td>
<td valign="top" align="center">1,038 (37.82%)</td>
<td valign="top" align="center">1,832 (69.62%)</td>
<td valign="top" align="center">390 (18.81%)</td>
<td valign="top" align="center">1,594 (39.29%)</td>
<td valign="top" align="center">1,767 (37.37%)</td>
</tr>
<tr>
<td valign="top" align="left" colspan="7" style="background-color:#bbbdc0"><bold>UNIQUE RASE SUBSYSTEM<xref ref-type="table-fn" rid="TN30"><sup>f</sup></xref></bold></td>
</tr>
<tr>
<td valign="top" align="left">Level 1</td>
<td/>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left">Level 2</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left">Level 3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">37</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left">Level 4</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">28</td>
<td valign="top" align="center">141</td>
<td valign="top" align="center">397</td>
<td valign="top" align="center">84</td>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left" colspan="7" style="background-color:#bbbdc0"><bold>MISSING RAST SUBSYSTEMS<xref ref-type="table-fn" rid="TN30"><sup>f</sup></xref></bold></td>
</tr>
<tr>
<td valign="top" align="left">Level 3</td>
<td/>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td valign="top" align="center">NA</td>
</tr>
<tr>
<td valign="top" align="left">Level 4</td>
<td/>
<td valign="top" align="center">11</td>
<td/>
<td/>
<td/>
<td valign="top" align="center">NA</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN25"><label>a</label><p><italic>PG, Porphyromonas gingivalis; TF, Tannerella forsythia; TD, Treponema denticola; AA, Aggregatibacter actinomycetemcomitans; BU, Bacteroides uniformis; BF, B. fragilis Genomes are downloaded from the NCBI FTP site; only those genomes with annotation were analyzed</italic>.</p></fn>
<fn id="TN26"><label>b</label><p><italic>Number of NCBI predicted ORFs based on the number of proteins found in the &#x0201C;.faa&#x0201D; file for each genome</italic>.</p></fn>
<fn id="TN27"><label>c</label><p><italic>Number of core proteins are those present in all genomes of the species, identified with the &#x0201C;blastclust&#x0201D; software (<ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html">https://www.ncbi.nlm.nih.gov/Web/Newsltr/Spring04/blastlab.html</ext-link>) using 95% sequence identity and 90% length coverage as the parameters</italic>.</p></fn>
<fn id="TN28"><label>d</label><p><italic>Number of unique proteins are those present in a single genome of the species, identified with the &#x0201C;blastclust&#x0201D; software using 60% sequence identity and 50% length coverage as the parameters</italic>.</p></fn>
<fn id="TN29"><label>e</label><p><italic>Non-hypothetic proteins predicted by NCBI, are proteins with annotation that does not contain the key words &#x0201C;hypothetic&#x0201D; and &#x0201C;uncharacterized&#x0201D;; hypothetical ones are those with annotation that contains either words</italic>.</p></fn>
<fn id="TN30"><label>f</label><p><italic>The 4 levels of subsystems were defined by the RAST (Aziz et al., <xref ref-type="bibr" rid="B2">2008</xref>); the unique systems for a species (regardless of levels) are those that can be found in &#x02265; 85% of all the genomes in the species, and &#x0003C; 15% in all other species; the missing subsystems are those found in &#x0003C; 15% of the target species but present in &#x02265;85% of all other species. The blank cells are those with none found; NA: not analyzed. The 85/15% cutoff was chosen based on the number of genomes in TF so that it allowed detection of a subsystem present or missing in only 1 out of 7 genomes</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p><italic>P. gingivalis</italic> originally belonged to the genus <italic>Bacteroides</italic> but was reclassified to the new genus <italic>Porphyromonas</italic> due to its marked biochemical and chemical differences (Shah and Collins, <xref ref-type="bibr" rid="B44">1988</xref>). Based on the 16S rRNA phylogeny the closest bacterial genus found in human oral cavity is <italic>Tannerella</italic>, which was also re-classified from <italic>Bacteroides</italic> (Sakamoto et al., <xref ref-type="bibr" rid="B41">2002</xref>). <italic>Bacteroides</italic> is the next closest genus to <italic>Porphyromonas</italic>. Naturally it is highly interesting and important to compare genomes of these close genera. So far genomics of seven strains of the species <italic>T. forsythia</italic> (formerly <italic>T. forsythensis</italic>) (TF) have been sequenced. As to <italic>Bacteroides</italic>, the most sequenced species is <italic>B. fragilis</italic> (BF), an important human gut bacterium, with a total of 114 strains/genomes sequenced to-date (only 103 were annotated). The runner-up is <italic>B. uniformis</italic> (BU), another human gut bacterial species with 21 genomic sequences available (16 were annotated). However, both of these two gut <italic>Bacteroides</italic> species are not considered human oral species. The only oral <italic>Bacteroides</italic> species that has been sequenced is <italic>B. pyogenes</italic> (BP), with five sequenced genomes that actually belong to four different strains (strain DSM 20611 &#x0003D; JCM 6294 was sequenced twice separately).</p>
<p>Table <xref ref-type="table" rid="T9">9</xref> also includes summary statistics for two other important human oral pathogenic species&#x02014;<italic>Treponema denticola</italic> (TD) and <italic>Aggregatibacter actinomycetemcomitans</italic> (AA). It has long been recognized that bacterial species exist in complexes in subgingival plaque. One complex was found by Socransky et al. (<xref ref-type="bibr" rid="B48">1998</xref>) to consist of the tightly related group <italic>P. gingivalis, T. denticola</italic>, and <italic>T. forsythia</italic>. This complex related strikingly to clinical measures such as pocket depth and bleeding on probing in chronic periodontitis. <italic>A. actinomycetemcomitans</italic> has for decades been associated with aggressive forms of periodontitis in adolescents (Haubek and Johansson, <xref ref-type="bibr" rid="B17">2014</xref>).</p>
<p>Of the six species summarized in Table <xref ref-type="table" rid="T9">9</xref>, <italic>A. actinomycetemcomitans</italic> has the smallest mean genome size (2,155 kb) and <italic>B. fragilis</italic> the largest (5,416 kb). The average number of protein genes encoded by these genomes are as expected proportional to the genome sizes, from 2,046 to 4,760 ORFs for AA and BF respectively. Using the most stringent criteria to define the core and unique proteins used for <italic>P. gingivalis</italic> in this study, <italic>T. denticola</italic> and <italic>B. uniformis</italic> both have similar numbers of core proteins when compared to PG. <italic>T. forsythia</italic> appeared to have higher number of core proteins (1,560). This is most likely due to the lower number of genomes analyzed) and that if more genomes are sequenced fewer core proteins would be identified. AA appeared to have the lowest number of core proteins detected but it also has twice the number of genomes analyzed compared to PG, TD, and BU (38 vs. 19, 17, and 16). We were not able to analyze the 103 BF genomes using the &#x0201C;blastclust&#x0201D; software due to the extreme slow speed with larger number of genomes. It would otherwise be interesting to examine the number of core proteins detected using the same criteria (95% sequence identity and 90% length coverage) for this large group of genomes. Similarly the number of unique proteins detected (with 60% sequence identity and 60% length coverage) is also, if not completely, determined by the size and number of genomes used for the analysis. The larger the sizes of the genomes, or the fewer the number of the genomes used for analysis, the more unique proteins were detected.</p>
<p>For functional comparison, Table <xref ref-type="table" rid="T9">9</xref> summarizes the total number of non-hypothetical and hypothetical proteins annotated by NCBI based on the absence or presence of either keyword &#x0201C;hypothetical&#x0201D; or &#x0201C;uncharacterized&#x0201D; associated with the protein annotation. The average percentage of non-hypothetical proteins differs greatly between species from 30.38% in TD to as high as 81.19% in AA (69.62% and 18.81% respectively for hypothetical proteins). A preliminary functional comparison based on RAST annotation for these genomes revealed several unique higher level subsystems. For example, the unique &#x0201C;Level 2&#x0201D; RAST subsystem identified in TD is &#x0201C;Motility and Chemotaxis&#x0201D; which is also expected because TD is a highly motile spirochaete species. The only Level 1 subsystem determined by this analysis is &#x0201C;Secondary Metabolism&#x02014;Lanthionine biosynthesis&#x0201D; which is associated with the LanB and LanC proteins identified in TF. Lanthionine has been found in bacterial cell walls and is also a component of a group of genes encoding peptide antibiotics called lantibiotics. They are a type of bacteriocins commonly found and made by different genera of actinomycetes (Maffioli et al., <xref ref-type="bibr" rid="B30">2015</xref>). The true roles of the lanB and lanC genes identified in TF may be worth investigating. For PG, there are four Level 3 subsystems detected that are unique and not present in the other 5 species: (1) Amino Acids and Derivatives&#x02014;Lysine, threonine, methionine, and cysteine&#x02014;Lysine fermentation; (2) Amino Acids and Derivatives&#x02014;Proline, 4-hydroxyproline uptake and utilization; (3) Stress Response&#x02014;Dimethylarginine metabolism; and (4) Virulence, Disease and Defense&#x02014;Resistance to antibiotics Vancomycin. The vancomycin resistance is conferred by a gene encoding vancomycin B-type resistance protein VanW, found exactly 1 copy in all of the 19 PG genomes, but not in all other genomes of other species.</p>
<p>Above is only a very brief description of what were identified as unique or missing functions among this selected group of species based on a very preliminary analysis. More data from which Table <xref ref-type="table" rid="T9">9</xref> was derived can be found in the FTP site dedicated to this publication <ext-link ext-link-type="uri" xlink:href="ftp://ftp.homd.org/publication_data/20160425/8_Comparison_to_other_species/">ftp://ftp.homd.org/publication_data/20160425/8_Comparison_to_other_species/</ext-link>. A more comprehensive comparative genomics study for these, and more interesting species and genomes is under investigation and will be reported in a separate publication in the future.</p>
</sec>
<sec id="s7">
<title>Concluding remarks</title>
<p>In this report 19 genomes of the species <italic>P. gingivalis</italic> as well as the outgroup species <italic>P. asaccharolytica</italic> were compared at several different levels of information ranging from nucleotide to genes to proteins and metabolic functions. Based on the single gene 16S rRNA phylogeny and multi-gene pholygenomic approach using core/shared protein sequences, several plausible evolutionary paths were suggested. Although there is no single evolutionary path concluded by these analyses, two closely related groups were consistently observed throughout the analyses. The first group consists of strains ATCC 33277, 381, and HG66 and the second of W83, W50, and A7436. The group of ATCC 33277, 381, and HG66 is also closer to the possible common ancestor inferred based on the use of an outgroup species <italic>P. asaccharolytica</italic>. We also detected at least 1,037 core/shared proteins for this species based on 95% sequence similarity and 90% alignment length. However, the number of core proteins increases with the lowering of the two detecting parameters. Functional and metabolic pathways were also compared and suggested several important functions of pathways that are unique to this species, to each strain, or missing in any particular strain. <italic>P. gingivalis</italic> has many genes encoding proteins related to or involved in gingipains, attachment (e.g., adhesins and fimbrins), capsules, and phages. These proteins were either missing or present in very few copies in the neighbor species <italic>P. asaccharolytica</italic>. Particularly intriguing observations were prevalence of many proteins related in phage productions and the equal prevalence of the CRISPR system in this species, with the exception of one strain lacking the Cas proteins.</p>
<p>Despite the large amount of comparative results generated in this study, there are still many different ways and software tools for analyzing and comparing a group of genomes. The complete results presented in this report, together with several other results that were only mentioned briefly here, are made available for download online at <ext-link ext-link-type="uri" xlink:href="ftp://www.homd.org/publication_data/20160425/">ftp://www.homd.org/publication_data/20160425/</ext-link>. We hope these data are useful to the research community and more hypotheses can be formulated based on the current or future analyses in order to gain deeper understanding on this important periodontal pathogen.</p>
</sec>
<sec id="s8">
<title>Author contributions</title>
<p>TC: Data acquisition, data analysis, data interpretation, writing of the manuscript, final approval of the version to be published; HS: Data acquisition, data analysis, data interpretation, writing; IO: Initiating the study, writing of the manuscript, revising the manuscript, final approval of the version to be published.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>This work was supported by The Forsyth Institute Bioinformatics Core and the European Commission (FP7-HEALTH-306029 &#x0201C;TRIGGER&#x0201D;).</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The reviewer DY and handling Editor declared their shared affiliation and the handling Editor states that the process nevertheless met the standards of a fair and objective review.</p></sec>
</sec>
</body>
<back>
<ack><p>TC acknowledges the Human Oral Microbiome Database (HOMD) for use of the computational resource and hosting the online data. IO and HS acknowledge funding through the European Commission.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Madden</surname> <given-names>T. L.</given-names></name> <name><surname>Sch&#x000E4;ffer</surname> <given-names>A. A.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>1997</year>). <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>. <source>Nucleic Acids Res.</source> <volume>25</volume>, <fpage>3389</fpage>&#x02013;<lpage>3402</lpage>. <pub-id pub-id-type="doi">10.1093/nar/25.17.3389</pub-id><pub-id pub-id-type="pmid">9254694</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aziz</surname> <given-names>R. K.</given-names></name> <name><surname>Bartels</surname> <given-names>D.</given-names></name> <name><surname>Best</surname> <given-names>A. A.</given-names></name> <name><surname>DeJongh</surname> <given-names>M.</given-names></name> <name><surname>Disz</surname> <given-names>T.</given-names></name> <name><surname>Edwards</surname> <given-names>R. A.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>The RAST Server: rapid annotations using subsystems technology</article-title>. <source>BMC Genomics</source> <volume>9</volume>:<fpage>75</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-9-75</pub-id><pub-id pub-id-type="pmid">18261238</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aziz</surname> <given-names>R. K.</given-names></name> <name><surname>Devoid</surname> <given-names>S.</given-names></name> <name><surname>Disz</surname> <given-names>T.</given-names></name> <name><surname>Edwards</surname> <given-names>R. A.</given-names></name> <name><surname>Henry</surname> <given-names>C. S.</given-names></name> <name><surname>Olsen</surname> <given-names>G. J.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>SEED servers: high-performance access to the SEED genomes, annotations, and metabolic models</article-title>. <source>PLoS ONE</source> <volume>7</volume>:<fpage>e48053</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0048053</pub-id><pub-id pub-id-type="pmid">23110173</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brunner</surname> <given-names>J.</given-names></name> <name><surname>Wittink</surname> <given-names>F. R.</given-names></name> <name><surname>Jonker</surname> <given-names>M. J.</given-names></name> <name><surname>de Jong</surname> <given-names>M.</given-names></name> <name><surname>Breit</surname> <given-names>T. M.</given-names></name> <name><surname>Laine</surname> <given-names>M. L.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>The core genome of the anaerobic oral pathogenic bacterium <italic>Porphyromonas gingivalis</italic></article-title>. <source>BMC Microbiol.</source> <volume>10</volume>:<fpage>252</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2180-10-252</pub-id><pub-id pub-id-type="pmid">20920246</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chastain-Gross</surname> <given-names>R. P.</given-names></name> <name><surname>Xie</surname> <given-names>G.</given-names></name> <name><surname>B&#x000E9;langer</surname> <given-names>M.</given-names></name> <name><surname>Kumar</surname> <given-names>D.</given-names></name> <name><surname>Whitlock</surname> <given-names>J. A.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Genome sequence of <italic>Porphyromonas gingivalis</italic> strain A7436</article-title>. <source>Genome Announc</source> <volume>3</volume>:<fpage>e00927</fpage>-<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1128/genomeA.00927-15</pub-id><pub-id pub-id-type="pmid">26404590</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chastain-Gross</surname> <given-names>R. P.</given-names></name> <name><surname>Xie</surname> <given-names>G.</given-names></name> <name><surname>B&#x000E9;langer</surname> <given-names>M.</given-names></name> <name><surname>Kumar</surname> <given-names>D.</given-names></name> <name><surname>Whitlock</surname> <given-names>J. A.</given-names></name> <name><surname>Liu</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Genome sequence of <italic>Porphyromonas gingivalis</italic> strain 381</article-title>. <source>Genome Announc</source>. <volume>5</volume>:<fpage>e01467</fpage>-<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1128/genomeA.01467-16</pub-id><pub-id pub-id-type="pmid">28082501</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>T.</given-names></name> <name><surname>Hosogi</surname> <given-names>Y.</given-names></name> <name><surname>Nishikawa</surname> <given-names>K.</given-names></name> <name><surname>Abbey</surname> <given-names>K.</given-names></name> <name><surname>Fleischmann</surname> <given-names>R. D.</given-names></name> <name><surname>Walling</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>Comparative whole-genome analysis of virulent and avirulent strains of <italic>Porphyromonas gingivalis</italic></article-title>. <source>J. Bacteriol.</source> <volume>186</volume>, <fpage>5473</fpage>&#x02013;<lpage>5479</lpage>. <pub-id pub-id-type="doi">10.1128/JB.186.16.5473-5479.2004</pub-id><pub-id pub-id-type="pmid">15292149</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Darveau</surname> <given-names>R. P.</given-names></name> <name><surname>Hajishengallis</surname> <given-names>G.</given-names></name> <name><surname>Curtis</surname> <given-names>M. A.</given-names></name></person-group> (<year>2012</year>). <article-title><italic>Porphyromonas gingivalis</italic> as a potential community activist for disease</article-title>. <source>J. Dent. Res.</source> <volume>91</volume>, <fpage>816</fpage>&#x02013;<lpage>820</lpage>. <pub-id pub-id-type="doi">10.1177/0022034512453589</pub-id><pub-id pub-id-type="pmid">22772362</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Delcher</surname> <given-names>A. L.</given-names></name> <name><surname>Phillippy</surname> <given-names>A.</given-names></name> <name><surname>Carlton</surname> <given-names>J.</given-names></name> <name><surname>Salzberg</surname> <given-names>S. L.</given-names></name></person-group> (<year>2002</year>). <article-title>Fast algorithms for large-scale genome alignment and comparison</article-title>. <source>Nucleic Acids Res.</source> <volume>30</volume>, <fpage>2478</fpage>&#x02013;<lpage>2483</lpage>. <pub-id pub-id-type="doi">10.1093/nar/30.11.2478</pub-id><pub-id pub-id-type="pmid">12034836</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Demmer</surname> <given-names>R. T.</given-names></name> <name><surname>Desvarieux</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>Periodontal infections and cardiovascular disease: the heart of the matter</article-title>. <source>J. Am. Dent. Assoc</source>. <volume>137</volume>, <fpage>14S</fpage>&#x02013;<lpage>20S</lpage>. <pub-id pub-id-type="doi">10.14219/jada.archive.2006.0402</pub-id><pub-id pub-id-type="pmid">17012731</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dolgilevich</surname> <given-names>S.</given-names></name> <name><surname>Rafferty</surname> <given-names>B.</given-names></name> <name><surname>Luchinskaya</surname> <given-names>D.</given-names></name> <name><surname>Kozarov</surname> <given-names>E.</given-names></name></person-group> (<year>2011</year>). <article-title>Genome comparison of invasive and rare non-invasive strains reveals <italic>Porphyromonas gingivalis</italic> genetic polymorphisms</article-title>. <source>J. Oral Microbiol</source>. <volume>3</volume>:<fpage>5764</fpage>. <pub-id pub-id-type="doi">10.3402/jom.v3i05764</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dorn</surname> <given-names>B. R.</given-names></name> <name><surname>Burks</surname> <given-names>J. N.</given-names></name> <name><surname>Seifert</surname> <given-names>K. N.</given-names></name> <name><surname>Progulske-Fox</surname> <given-names>A.</given-names></name></person-group> (<year>2000</year>). <article-title>Invasion of endothelial and epithelial cells by strains of <italic>Porphyromonas gingivalis</italic></article-title>. <source>FEMS Microbiol. Lett.</source> <volume>187</volume>, <fpage>139</fpage>&#x02013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6968.2000.tb09150.x</pub-id><pub-id pub-id-type="pmid">10856647</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duncan</surname> <given-names>M. J.</given-names></name></person-group> (<year>2003</year>). <article-title>Genomics of oral bacteria</article-title>. <source>Crit. Rev. Oral Biol. Med.</source> <volume>14</volume>, <fpage>175</fpage>&#x02013;<lpage>187</lpage>. <pub-id pub-id-type="doi">10.1177/154411130301400303</pub-id><pub-id pub-id-type="pmid">12799321</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gerdes</surname> <given-names>K.</given-names></name></person-group> (<year>2000</year>). <article-title>Toxin-antitoxin modules may regulate synthesis of macromolecules during nutritional stress</article-title>. <source>J. Bacteriol.</source> <volume>182</volume>, <fpage>561</fpage>&#x02013;<lpage>572</lpage>. <pub-id pub-id-type="doi">10.1128/JB.182.3.561-572.2000</pub-id><pub-id pub-id-type="pmid">10633087</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goto</surname> <given-names>T.</given-names></name> <name><surname>Nagano</surname> <given-names>K.</given-names></name> <name><surname>Hirakawa</surname> <given-names>H.</given-names></name> <name><surname>Tanaka</surname> <given-names>K.</given-names></name> <name><surname>Yoshimura</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>Draft genome sequence of <italic>Porphyromonas gingivalis</italic> strain Ando expressing a 53-kilodalton-type fimbrilin variant of Mfa1 fimbriae</article-title>. <source>Genome Announc.</source> <volume>3</volume>:<fpage>e01292</fpage>-<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1128/genomeA.01292-15</pub-id><pub-id pub-id-type="pmid">26543123</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hajishengallis</surname> <given-names>G.</given-names></name> <name><surname>Darveau</surname> <given-names>R. P.</given-names></name> <name><surname>Curtis</surname> <given-names>M. A.</given-names></name></person-group> (<year>2012</year>). <article-title>The keystone pathogen hypothesis</article-title>. <source>Nat. Rev. Microbiol.</source> <volume>10</volume>, <fpage>717</fpage>&#x02013;<lpage>725</lpage>. <pub-id pub-id-type="doi">10.1038/nrmicro2873</pub-id><pub-id pub-id-type="pmid">22941505</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haubek</surname> <given-names>D.</given-names></name> <name><surname>Johansson</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Pathogenicity of the highly leukotoxic JP2 clone of <italic>Aggregatibacter actinomycetemcomitans</italic> and its geographic dissemination and role in aggressive periodontitis</article-title>. <source>J. Oral Microbiol.</source> <volume>14</volume>:<fpage>6</fpage>. <pub-id pub-id-type="doi">10.3402/jom.v6.23980</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hayes</surname> <given-names>F.</given-names></name></person-group> (<year>2003</year>). <article-title>Toxins-antitoxins: plasmid maintenance, programmed cell death, and cell cycle arrest</article-title>. <source>Science</source> <volume>301</volume>, <fpage>1496</fpage>&#x02013;<lpage>1499</lpage>. <pub-id pub-id-type="doi">10.1126/science.1088157</pub-id><pub-id pub-id-type="pmid">12970556</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Horvath</surname> <given-names>P.</given-names></name> <name><surname>Barrangou</surname> <given-names>R.</given-names></name></person-group> (<year>2010</year>). <article-title>CRISPR/Cas, the immune system of bacteria and archaea</article-title>. <source>Science</source> <volume>327</volume>, <fpage>167</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1126/science.1179555</pub-id><pub-id pub-id-type="pmid">20056882</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Howe</surname> <given-names>K.</given-names></name> <name><surname>Bateman</surname> <given-names>A.</given-names></name> <name><surname>Durbin</surname> <given-names>R.</given-names></name></person-group> (<year>2002</year>). <article-title>QuickTree: building huge Neighbour- Joining trees of protein sequences</article-title>. <source>Bioinformatics</source> <volume>18</volume>, <fpage>1564</fpage>&#x02013;<lpage>1567</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/18.11.1546</pub-id><pub-id pub-id-type="pmid">12424131</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jones</surname> <given-names>D. T.</given-names></name> <name><surname>Taylor</surname> <given-names>W. R.</given-names></name> <name><surname>Thornton</surname> <given-names>J. M.</given-names></name></person-group> (<year>1992</year>). <article-title>The rapid generation of mutation data matrices from protein sequences</article-title>. <source>Comput. Appl. Biosci.</source> <volume>8</volume>, <fpage>275</fpage>&#x02013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/8.3.275</pub-id><pub-id pub-id-type="pmid">1633570</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kanehisa</surname> <given-names>M.</given-names></name> <name><surname>Sato</surname> <given-names>Y.</given-names></name> <name><surname>Morishima</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>BlastKOALA and GhostKOALA: KEGG tools for functional characterization of genome and metagenome sequences</article-title>. <source>J. Mol. Biol.</source> <volume>428</volume>, <fpage>726</fpage>&#x02013;<lpage>731</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmb.2015.11.006</pub-id><pub-id pub-id-type="pmid">26585406</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katoh</surname> <given-names>K.</given-names></name> <name><surname>Standley</surname> <given-names>D. M.</given-names></name></person-group> (<year>2013</year>). <article-title>MAFFT multiple sequence alignment software version 7: improvements in performance and usability</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>772</fpage>&#x02013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mst010</pub-id><pub-id pub-id-type="pmid">23329690</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klappenbach</surname> <given-names>J. A.</given-names></name> <name><surname>Dunbar</surname> <given-names>J. M.</given-names></name> <name><surname>Schmidt</surname> <given-names>T. M.</given-names></name></person-group> (<year>2000</year>). <article-title>rRNA operon copy number reflects ecological strategies of bacteria</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>66</volume>, <fpage>1328</fpage>&#x02013;<lpage>1333</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.66.4.1328-1333.2000</pub-id><pub-id pub-id-type="pmid">10742207</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klein</surname> <given-names>B. A.</given-names></name> <name><surname>Chen</surname> <given-names>T.</given-names></name> <name><surname>Scott</surname> <given-names>J. C.</given-names></name> <name><surname>Koenigsberg</surname> <given-names>A. L.</given-names></name> <name><surname>Duncan</surname> <given-names>M. J.</given-names></name> <name><surname>Hu</surname> <given-names>L. T.</given-names></name></person-group> (<year>2015</year>). <article-title>Identification and characterization of a minisatellite contained within a novel miniature inverted-repeat transposable element (MITE) of <italic>Porphyromonas gingivalis</italic></article-title>. <source>Mob. DNA</source> <volume>6</volume>:<fpage>18</fpage>. <pub-id pub-id-type="doi">10.1186/s13100-015-0049-1</pub-id><pub-id pub-id-type="pmid">26448788</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klein</surname> <given-names>B. A.</given-names></name> <name><surname>Tenorio</surname> <given-names>E. L.</given-names></name> <name><surname>Lazinski</surname> <given-names>D. W.</given-names></name> <name><surname>Camilli</surname> <given-names>A.</given-names></name> <name><surname>Duncan</surname> <given-names>M. J.</given-names></name> <name><surname>Hu</surname> <given-names>L. T.</given-names></name></person-group> (<year>2012</year>). <article-title>Identification of essential genes of the periodontal pathogen <italic>Porphyromonas gingivalis</italic></article-title>. <source>BMC Genomics</source> <volume>13</volume>:<fpage>578</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-13-578</pub-id><pub-id pub-id-type="pmid">23114059</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laine</surname> <given-names>M. L.</given-names></name> <name><surname>van Winkelhoff</surname> <given-names>A. J.</given-names></name></person-group> (<year>1998</year>). <article-title>Virulence of six capsular serotypes of <italic>Porphyromonas gingivalis</italic> in a mouse model</article-title>. <source>Oral Microbiol. Immunol.</source> <volume>13</volume>, <fpage>322</fpage>&#x02013;<lpage>325</lpage>. <pub-id pub-id-type="doi">10.1111/j.1399-302X.1998.tb00714.x</pub-id><pub-id pub-id-type="pmid">9807125</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>D.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Naito</surname> <given-names>M.</given-names></name> <name><surname>Yumoto</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Miyake</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Draft genome sequence of <italic>Porphyromonas gingivalis</italic> strain SJD2, isolated from the periodontal pocket of a patient with periodontitis in China</article-title>. <source>Genome Announc.</source> <volume>2</volume>:<fpage>e01091</fpage>-<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1128/genomeA.01091-13</pub-id><pub-id pub-id-type="pmid">24385574</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lundberg</surname> <given-names>K.</given-names></name> <name><surname>Wegner</surname> <given-names>N.</given-names></name> <name><surname>Yucel-Lindberg</surname> <given-names>T.</given-names></name> <name><surname>Venables</surname> <given-names>P. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Periodontitis in RA- citrullinated enolase connection</article-title>. <source>Nat. Rev. Rheumatol.</source> <volume>6</volume>, <fpage>727</fpage>&#x02013;<lpage>730</lpage>. <pub-id pub-id-type="doi">10.1038/nrrheum.2010.139</pub-id><pub-id pub-id-type="pmid">20820197</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maffioli</surname> <given-names>S. I.</given-names></name> <name><surname>Monciardini</surname> <given-names>P.</given-names></name> <name><surname>Catacchio</surname> <given-names>B.</given-names></name> <name><surname>Mazzetti</surname> <given-names>C.</given-names></name> <name><surname>M&#x000FC;nch</surname> <given-names>D.</given-names></name> <name><surname>Brunati</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Family of class I lantibiotics from actinomycetes and improvement of their antibacterial activities</article-title>. <source>ACS Chem. Biol.</source> <volume>10</volume>, <fpage>1034</fpage>&#x02013;<lpage>1042</lpage>. <pub-id pub-id-type="doi">10.1021/cb500878h</pub-id><pub-id pub-id-type="pmid">25574687</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markowitz</surname> <given-names>V. M.</given-names></name> <name><surname>Chen</surname> <given-names>I. M.</given-names></name> <name><surname>Palaniappan</surname> <given-names>K.</given-names></name> <name><surname>Chu</surname> <given-names>K.</given-names></name> <name><surname>Szeto</surname> <given-names>E.</given-names></name> <name><surname>Pillay</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>IMG 4 version of the integrated microbial genomes comparative analysis system</article-title>. <source>Nucleic Acids Res</source>. <volume>42</volume>(Database issue), <fpage>D560</fpage>&#x02013;<lpage>D567</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkt963</pub-id><pub-id pub-id-type="pmid">24165883</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>P. D.</given-names></name> <name><surname>McDermid</surname> <given-names>A. S.</given-names></name> <name><surname>McKee</surname> <given-names>A. S.</given-names></name> <name><surname>Baskerville</surname> <given-names>A.</given-names></name></person-group> (<year>1994</year>). <article-title>The effect of growth rate and haemin on the virulence and proteolytic activity of <italic>Porphyromonas gingivalis</italic> W50</article-title>. <source>Microbiology</source> <volume>140</volume>, <fpage>861</fpage>&#x02013;<lpage>865</lpage>. <pub-id pub-id-type="doi">10.1099/00221287-140-4-861</pub-id><pub-id pub-id-type="pmid">8012602</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McLean</surname> <given-names>J. S.</given-names></name> <name><surname>Lombardo</surname> <given-names>M. J.</given-names></name> <name><surname>Ziegler</surname> <given-names>M. G.</given-names></name> <name><surname>Novotny</surname> <given-names>M.</given-names></name> <name><surname>Yee-Greenbaum</surname> <given-names>J.</given-names></name> <name><surname>Badger</surname> <given-names>J. H.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Genome of the pathogen <italic>Porphyromonas gingivalis</italic> recovered from a biofilm in a hospital sink using a high-throughput single-cell genomics platform</article-title>. <source>Genome Res.</source> <volume>23</volume>, <fpage>867</fpage>&#x02013;<lpage>877</lpage>. <pub-id pub-id-type="doi">10.1101/gr.150433.112</pub-id><pub-id pub-id-type="pmid">23564253</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nagano</surname> <given-names>K.</given-names></name> <name><surname>Hasegawa</surname> <given-names>Y.</given-names></name> <name><surname>Yoshida</surname> <given-names>Y.</given-names></name> <name><surname>Yoshimura</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>A major fimbrilin variant of Mfa1fimbriae in <italic>Porphyromonas gingivalis</italic></article-title>. <source>J. Dent. Res.</source> <volume>94</volume>, <fpage>1143</fpage>&#x02013;<lpage>1148</lpage>. <pub-id pub-id-type="doi">10.1177/0022034515588275</pub-id><pub-id pub-id-type="pmid">26001707</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naito</surname> <given-names>M.</given-names></name> <name><surname>Hirakawa</surname> <given-names>H.</given-names></name> <name><surname>Yamashita</surname> <given-names>A.</given-names></name> <name><surname>Ohara</surname> <given-names>N.</given-names></name> <name><surname>Shoji</surname> <given-names>M.</given-names></name> <name><surname>Yukitake</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Determination of the genome sequence of <italic>Porphyromonas gingivalis</italic> strain ATCC 33277 and genomic comparison with strain W83 revealed extensive genome rearrangements in <italic>P</italic></article-title>. <source>gingivalis. DNA Res.</source> <volume>15</volume>, <fpage>215</fpage>&#x02013;<lpage>225</lpage>. <pub-id pub-id-type="doi">10.1093/dnares/dsn013</pub-id><pub-id pub-id-type="pmid">18524787</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nelson</surname> <given-names>K. E.</given-names></name> <name><surname>Fleischmann</surname> <given-names>R. D.</given-names></name> <name><surname>DeBoy</surname> <given-names>R. T.</given-names></name> <name><surname>Paulsen</surname> <given-names>I. T.</given-names></name> <name><surname>Fouts</surname> <given-names>D. E.</given-names></name> <name><surname>Eisen</surname> <given-names>J. A.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Complete genome sequence of the oral pathogenic Bacterium <italic>Porphyromonas gingivalis</italic> strain W83</article-title>. <source>J. Bacteriol.</source> <volume>185</volume>, <fpage>5591</fpage>&#x02013;<lpage>5601</lpage>. <pub-id pub-id-type="doi">10.1128/JB.185.18.5591-5601.2003</pub-id><pub-id pub-id-type="pmid">12949112</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olsen</surname> <given-names>I.</given-names></name> <name><surname>Progulske-Fox</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Invasion of <italic>Porphyromonas gingivalis</italic> into vascular cells and tissue</article-title>. <source>J. Oral Microbiol.</source> <volume>7</volume>:<fpage>28788</fpage>. <pub-id pub-id-type="doi">10.3402/jom.v7.28788</pub-id><pub-id pub-id-type="pmid">26329158</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olsen</surname> <given-names>I.</given-names></name> <name><surname>Singhrao</surname> <given-names>S. K.</given-names></name></person-group> (<year>2015</year>). <article-title>Can oral infection be a risk factor for Alzheimer&#x00027;s disease?</article-title> <source>J. Oral Microbiol.</source> <volume>7</volume>:<fpage>29143</fpage>. <pub-id pub-id-type="doi">10.3402/jom.v7.29143</pub-id><pub-id pub-id-type="pmid">26385886</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Overbeek</surname> <given-names>R.</given-names></name> <name><surname>Olson</surname> <given-names>R.</given-names></name> <name><surname>Pusch</surname> <given-names>G. D.</given-names></name> <name><surname>Olsen</surname> <given-names>G. J.</given-names></name> <name><surname>Davis</surname> <given-names>J. J.</given-names></name> <name><surname>Disz</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>The SEED and the rapid annotation of microbial genomes using subsystems technology (RAST)</article-title>. <source>Nucleic Acids Res</source>. <volume>42</volume>(Database issue), <fpage>D206</fpage>&#x02013;<lpage>D214</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkt122</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Price</surname> <given-names>M. N.</given-names></name> <name><surname>Dehal</surname> <given-names>P. S.</given-names></name> <name><surname>Arkin</surname> <given-names>A. P.</given-names></name></person-group> (<year>2010</year>). <article-title>FastTree 2-approximately maximum-likelyhood trees for large alignments</article-title>. <source>PLoS ONE</source> <volume>5</volume>:<fpage>e9490</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0009490</pub-id><pub-id pub-id-type="pmid">20224823</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakamoto</surname> <given-names>M.</given-names></name> <name><surname>Suzuki</surname> <given-names>M.</given-names></name> <name><surname>Umeda</surname> <given-names>M.</given-names></name> <name><surname>Ishikawa</surname> <given-names>I.</given-names></name> <name><surname>Benno</surname> <given-names>Y.</given-names></name></person-group> (<year>2002</year>). <article-title>Reclassification of Bacteroides forsythus (Tanner et al. 1986) as Tannerella forsythensis corrig., gen. nov., comb. nov</article-title>. <source>Int. J. Syst. Evol. Microbiol</source>. <volume>52</volume>, <fpage>841</fpage>&#x02013;<lpage>849</lpage>. <pub-id pub-id-type="doi">10.1099/00207713-52-3-841</pub-id><pub-id pub-id-type="pmid">12054248</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sandmeier</surname> <given-names>H.</given-names></name> <name><surname>Bar</surname> <given-names>K.</given-names></name> <name><surname>Meyer</surname> <given-names>J.</given-names></name></person-group> (<year>1993</year>). <article-title>Search for bacteriophages of black-pigmented gram-negative anaerobes from dental plaque</article-title>. <source>FEMS Immunol. Med. Microbiol.</source> <volume>6</volume>, <fpage>193</fpage>&#x02013;<lpage>194</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-695X.1993.tb00324.x</pub-id><pub-id pub-id-type="pmid">8390892</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Segata</surname> <given-names>N.</given-names></name> <name><surname>Bornigen</surname> <given-names>D.</given-names></name> <name><surname>Morgan</surname> <given-names>X. C.</given-names></name> <name><surname>Huttenhower</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>PhyloPhlAn is a new method for improved phylogenetic and taxonomic placement of microbes</article-title>. <source>Nat. Commun.</source> <volume>4</volume>:<fpage>2304</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms3304</pub-id><pub-id pub-id-type="pmid">23942190</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shah</surname> <given-names>H. N.</given-names></name> <name><surname>Collins</surname> <given-names>M. D.</given-names></name></person-group> (<year>1988</year>). <article-title>Proposal for reclassification of <italic>Bacteroides asaccharolyticus, Bacteroides gingivalis</italic>, and <italic>Bacteroides endodontalis</italic> in a new genus, <italic>Porphyromonas</italic></article-title>. <source>Int. J. Syst. Bacteriol.</source> <volume>38</volume>, <fpage>128</fpage>&#x02013;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.1099/00207713-38-1-128</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shimodaira</surname> <given-names>H.</given-names></name> <name><surname>Hasegawa</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>CONSEL: for assessing the confidence of phylogenetic tree selection</article-title>. <source>Bioinformatics</source> <volume>17</volume>, <fpage>1246</fpage>&#x02013;<lpage>1247</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/17.12.1246</pub-id><pub-id pub-id-type="pmid">11751242</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siddiqui</surname> <given-names>H.</given-names></name> <name><surname>Yoder-Himes</surname> <given-names>D. R.</given-names></name> <name><surname>Mizgalska</surname> <given-names>D.</given-names></name> <name><surname>Nguyen</surname> <given-names>K. A.</given-names></name> <name><surname>Potempa</surname> <given-names>J.</given-names></name> <name><surname>Olsen</surname> <given-names>I.</given-names></name></person-group> (<year>2014</year>). <article-title>Genome sequence of <italic>Porphyromonas gingivalis</italic> strain HG66 (DSM 28984)</article-title>. <source>Genome Announc.</source> <volume>2</volume>:<fpage>e00947</fpage>-<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1128/genomeA.00947-14</pub-id><pub-id pub-id-type="pmid">25291768</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>A.</given-names></name> <name><surname>Wyant</surname> <given-names>T.</given-names></name> <name><surname>Anaya-Bergman</surname> <given-names>C.</given-names></name> <name><surname>Aduse-Opoku</surname> <given-names>J.</given-names></name> <name><surname>Brunner</surname> <given-names>J.</given-names></name> <name><surname>Laine</surname> <given-names>M. L.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>The capsule of <italic>Porphyromonas gingivalis</italic> leads to a reduction in the host inflammatory response, evasion of phagocytosis, and increase in virulence</article-title>. <source>Infect. Immun.</source> <volume>79</volume>, <fpage>4533</fpage>&#x02013;<lpage>4542</lpage>. <pub-id pub-id-type="doi">10.1128/IAI.05016-11</pub-id><pub-id pub-id-type="pmid">21911459</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Socransky</surname> <given-names>S. S.</given-names></name> <name><surname>Haffajee</surname> <given-names>A. D.</given-names></name> <name><surname>Cugini</surname> <given-names>M. A.</given-names></name> <name><surname>Smith</surname> <given-names>C.</given-names></name> <name><surname>Kent</surname> <given-names>R. L.</given-names> <suffix>Jr.</suffix></name></person-group> (<year>1998</year>). <article-title>Microbial complexes in subgingival plaque</article-title>. <source>J. Clin. Periodontol.</source> <volume>2</volume>, <fpage>134</fpage>&#x02013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1111/j.1600-051X.1998.tb02419.x</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Talavera</surname> <given-names>G.</given-names></name> <name><surname>Castresana</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments</article-title>. <source>Syst. Biol.</source> <volume>56</volume>, <fpage>564</fpage>&#x02013;<lpage>577</lpage>. <pub-id pub-id-type="doi">10.1080/10635150701472164</pub-id><pub-id pub-id-type="pmid">17654362</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tatusova</surname> <given-names>T.</given-names></name> <name><surname>DiCuccio</surname> <given-names>M.</given-names></name> <name><surname>Badretdin</surname> <given-names>A.</given-names></name> <name><surname>Chetvernin</surname> <given-names>V.</given-names></name> <name><surname>Nawrocki</surname> <given-names>E. P.</given-names></name> <name><surname>Zaslavsky</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>NCBI prokaryotic genome annotation pipeline</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>6614</fpage>&#x02013;<lpage>6624</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw569</pub-id><pub-id pub-id-type="pmid">27342282</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>To</surname> <given-names>T. T.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Watling</surname> <given-names>M.</given-names></name> <name><surname>Bumgarner</surname> <given-names>R. E.</given-names></name> <name><surname>Darveau</surname> <given-names>R. P.</given-names></name> <name><surname>McLean</surname> <given-names>J. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Draft genome sequence of low-passage clinical isolate <italic>Porphyromonas gingivalis</italic> MP4-504</article-title>. <source>Genome Announc.</source> <volume>4</volume>:<fpage>e00256</fpage>-<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1128/genomeA.00256-16</pub-id><pub-id pub-id-type="pmid">27056232</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tribble</surname> <given-names>G. D.</given-names></name> <name><surname>Kerr</surname> <given-names>J. E.</given-names></name> <name><surname>Wang</surname> <given-names>B. Y.</given-names></name></person-group> (<year>2013</year>). <article-title>Genetic diversity in the oral pathogen <italic>Porphyromonas gingivalis</italic>: molecular mechanisms and biological consequences</article-title>. <source>Future Microbiol.</source> <volume>8</volume>, <fpage>607</fpage>&#x02013;<lpage>620</lpage>. <pub-id pub-id-type="doi">10.2217/fmb.13.30</pub-id><pub-id pub-id-type="pmid">23642116</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Domselaar</surname> <given-names>G. H.</given-names></name> <name><surname>Stothard</surname> <given-names>P.</given-names></name> <name><surname>Shrivastava</surname> <given-names>S.</given-names></name> <name><surname>Cruz</surname> <given-names>J. A.</given-names></name> <name><surname>Guo</surname> <given-names>A.</given-names></name> <name><surname>Dong</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>BASys: a web server for automated bacterial genome annotation</article-title>. <source>Nucleic Acids Res</source>. <volume>33</volume>(Web Server issue), <fpage>W455</fpage>&#x02013;<lpage>W459</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gki593</pub-id><pub-id pub-id-type="pmid">15980511</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vetrovsky</surname> <given-names>T.</given-names></name> <name><surname>Baldrian</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>The variability of the 16S rRNA gene in bacterial genomes and its consequences for bacterial community analyses</article-title>. <source>PLoS ONE</source> <volume>8</volume>:<fpage>e57923</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0057923</pub-id><pub-id pub-id-type="pmid">23460914</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watanabe</surname> <given-names>T.</given-names></name> <name><surname>Maruyama</surname> <given-names>F.</given-names></name> <name><surname>Nozawa</surname> <given-names>T.</given-names></name> <name><surname>Aoki</surname> <given-names>A.</given-names></name> <name><surname>Okano</surname> <given-names>S.</given-names></name> <name><surname>Shibata</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Complete genome sequence of the bacterium <italic>Porphyromonas gingivalis</italic> TDC60, which causes periodontal disease</article-title>. <source>J. Bacteriol.</source> <volume>193</volume>, <fpage>4259</fpage>&#x02013;<lpage>4260</lpage>. <pub-id pub-id-type="doi">10.1128/JB.05269-11</pub-id><pub-id pub-id-type="pmid">21705612</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woese</surname> <given-names>C. R.</given-names></name> <name><surname>Kandler</surname> <given-names>O.</given-names></name> <name><surname>Wheelis</surname> <given-names>M. L.</given-names></name></person-group> (<year>1990</year>). <article-title>Towards a natural system of organisms: proposal for the domains Archaea, Bacteria, and Eucarya</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>87</volume>, <fpage>4576</fpage>&#x02013;<lpage>4579</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.87.12.4576</pub-id><pub-id pub-id-type="pmid">2112744</pub-id></citation>
</ref>
</ref-list>
</back>
</article>