<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2017.01345</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Pan-genome Analyses of the Species <italic>Salmonella enterica</italic>, and Identification of Genomic Markers Predictive for Species, Subspecies, and Serovar</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Laing</surname> <given-names>Chad R.</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/417418/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Whiteside</surname> <given-names>Matthew D.</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Gannon</surname> <given-names>Victor P. J.</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/336905/overview"/>
</contrib>
</contrib-group>
<aff><institution>National Microbiology Laboratory, Public Health Agency of Canada</institution> <country>Lethbridge, AB, Canada</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Sandra Torriani, University of Verona, Italy</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Jinshui Zheng, Huazhong Agricultural University, China; Dapeng Wang, Shanghai Jiao Tong University, China</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>Chad R. Laing, <email>chadr.laing@canada.ca</email></italic></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Food Microbiology, a section of the journal Frontiers in Microbiology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>31</day>
<month>07</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>1345</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>02</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>07</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2017 Laing, Whiteside and Gannon.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Laing, Whiteside and Gannon</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Food safety is a global concern, with upward of 2.2 million deaths due to enteric disease every year. Current whole-genome sequencing platforms allow routine sequencing of enteric pathogens for surveillance, and during outbreaks; however, a remaining challenge is the identification of genomic markers that are predictive of strain groups that pose the most significant health threats to humans, or that can persist in specific environments. We have previously developed the software program Panseq, which identifies the pan-genome among a group of sequences, and the SuperPhy platform, which utilizes this pan-genome information to identify biomarkers that are predictive of groups of bacterial strains. In this study, we examined the pan-genome of 4893 genomes of <italic>Salmonella enterica</italic>, an enteric pathogen responsible for the loss of more disability adjusted life years than any other enteric pathogen. We identified a pan-genome of 25.3 Mbp, a strict core of 1.5 Mbp present in all genomes, and a conserved core of 3.2 Mbp found in at least 96% of these genomes. We also identified 404 genomic regions of 1000 bp that were specific to the species <italic>S. enterica</italic>. These species-specific regions were found to encode mostly hypothetical proteins, effectors, and other proteins related to virulence. For each of the six <italic>S. enterica</italic> subspecies, markers unique to each were identified. No serovar had pan-genome regions that were present in all of its genomes and absent in all other serovars; however, each serovar did have genomic regions that were universally present among all constituent members, and statistically predictive of the serovar. The phylogeny based on SNPs within the conserved core genome was found to be highly concordant to that produced by a phylogeny using the presence/absence of 1000 bp regions of the entire pan-genome. Future studies could use these predictive regions as components of a vaccine to prevent salmonellosis, as well as in simple and rapid diagnostic tests for both <italic>in silico</italic> and wet-lab applications, with uses ranging from food safety to public health. Lastly, the tools and methods described in this study could be applied as a pan-genomics framework to other population genomic studies seeking to identify markers for other bacterial species and their sub-groups.</p>
</abstract>
<kwd-group>
<kwd>genomics</kwd>
<kwd>pan-genome</kwd>
<kwd><italic>Salmonella</italic></kwd>
<kwd>predictive markers</kwd>
<kwd>food safety</kwd>
</kwd-group>
<contract-sponsor id="cn001">Public Health Agency of Canada<named-content content-type="fundref-id">10.13039/100011094</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="6"/>
<equation-count count="0"/>
<ref-count count="69"/>
<page-count count="15"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>The global burden of bacterial enteric disease, much of it foodborne, results in an estimated 2.2 million deaths per year, and an annual loss of 112,000 disability adjusted life years in the United States alone (<xref ref-type="bibr" rid="B5">Bergholz et al., 2014</xref>; <xref ref-type="bibr" rid="B53">Scallan et al., 2015</xref>). Nationwide molecular diagnostic networks, such as PulseNet in North America, were designed to enable the rapid identification of outbreaks by genetic fingerprinting the etiological agents of disease, and keeping nationwide databases of genetic fingerprints of specific pathogens associated with human disease. Since its inception, PulseNet has relied on pulsed-field gel electrophoresis (PFGE) for fingerprinting of bacterial pathogens to identify the specific sources of outbreaks and prevent further infections. Using this approach, it has been estimated that PulseNet prevents 277,000 illnesses from bacterial pathogens annually in the United States, reducing the costs associated with medical care and loss of productivity due to worker illness (<xref ref-type="bibr" rid="B54">Scharff et al., 2016</xref>).</p>
<p>Despite the usefulness of PulseNet, the PFGE technique itself is often unable to distinguish between related and unrelated strains, due to its reliance on rare-cutting restriction enzyme sites within the genome (<xref ref-type="bibr" rid="B2">Allard et al., 2012</xref>). Additionally, the interpretation of the banding patterns among labs requires extensive training and standardization to enable meaningful comparisons. Lastly, the banding patterns provide no information on the actual content of the genomes they represent, so important information regarding human virulence, such as the presence or absence of known toxins, is not available.</p>
<p>Lastly, while the presence of known virulence factors has been correlated with severe human disease in a number of bacterial species, it has also been shown that some lineages or clades within these same species, while possessing specific virulence factors, are rarely associated with human disease (<xref ref-type="bibr" rid="B32">Lupolova et al., 2016</xref>; <xref ref-type="bibr" rid="B63">Waryah et al., 2016</xref>). Thus, multiple virulence factors, and regulatory genes that influence the expression of key virulence factors, or otherwise modulate the virulence of these strains, need to be taken into consideration when attempting to predict the strains of a bacterial species that are potential human health threats (<xref ref-type="bibr" rid="B41">Opijnen et al., 2012</xref>).</p>
<p>Recently, whole-genome sequencing (WGS) has displaced PFGE as the <italic>de facto</italic> standard for the complete characterization of bacterial pathogens, in both ongoing surveillance and outbreak investigations (<xref ref-type="bibr" rid="B12">Deng et al., 2016</xref>; <xref ref-type="bibr" rid="B14">Franz et al., 2016</xref>). WGS allows clear definition between outbreak-related strains and those from unrelated sources, and it has the ability to identify routes of transmission, and attribute bacterial contaminants to specific sources (<xref ref-type="bibr" rid="B11">den Bakker et al., 2014</xref>). It is currently being utilized in reference laboratories worldwide. Examples of its application include the sequencing of all <italic>Listeria monocytogenes</italic> isolated in the United States, all <italic>Salmonella</italic> isolated by the Food and Drug Administration in the USA, and by Public Health England as part of routine surveillance (<xref ref-type="bibr" rid="B3">Ashton et al., 2016</xref>), and a large-scale survey of <italic>Staphylococcus aureus</italic> in continental Europe. In the latter study, the applicability of WGS for the identification of the emergence and spread of clinically relevant <italic>Staphylococcus aureus</italic> was demonstrated (<xref ref-type="bibr" rid="B1">Aanensen et al., 2016</xref>).</p>
<p>It has also recently been shown that antimicrobial resistance (<xref ref-type="bibr" rid="B62">Tyson et al., 2015</xref>; <xref ref-type="bibr" rid="B36">McDermott et al., 2016</xref>; <xref ref-type="bibr" rid="B69">Zhao et al., 2016</xref>), serovar (<xref ref-type="bibr" rid="B31">Levine et al., 2016</xref>; <xref ref-type="bibr" rid="B66">Yoshida et al., 2016b</xref>), and the results of other traditional sub-typing schemes such as multi-locus sequence typing (<xref ref-type="bibr" rid="B57">Sheppard et al., 2012</xref>) can be accurately predicted <italic>in silico</italic> through the analysis of bacterial genome sequences. However, identifying bacterial isolates that are most likely to cause disease in humans, based on the genome sequence alone, is a more complex task. In addition, markers that can identify bacteria likely to exhibit particular phenotypes, such as the ability to survive in a particular niche, or the ability to tolerate harsh environments such as those found in food processing plants are also required.</p>
<p>We have previously developed the software platform Panseq, for the analyses of thousands of genomes in a pan-genome context, where both the presence/absence of the accessory genome and SNPs within the shared core-genome are computed (<xref ref-type="bibr" rid="B29">Laing et al., 2010</xref>). Additionally, we recently released a platform for the predictive genomics of <italic>Escherichia coli</italic>, called SuperPhy, in which markers statistically biased within groups of bacteria, based on any metadata category, can be identified (<xref ref-type="bibr" rid="B64">Whiteside et al., 2016</xref>).</p>
<p>In this study we use our previously created software to examine the pan-genome of <italic>Salmonella enterica</italic>, a pathogen that causes an estimated 93.8 million cases of enteric illness worldwide each year (<xref ref-type="bibr" rid="B33">Majowicz et al., 2010</xref>; <xref ref-type="bibr" rid="B17">Gal-Mor et al., 2014</xref>). The species <italic>S. enterica</italic> is divided into six subspecies: <italic>enterica</italic>, <italic>salamae</italic>, <italic>arizonae</italic>, <italic>diarizonae</italic>, <italic>houtenae</italic>, and <italic>indica</italic>. Over 99% of human disease caused by <italic>S. enterica</italic> is done so by subspecies <italic>enterica</italic>, with the World Health Organization estimating that <italic>S. enterica</italic> infections from contaminated food alone constitute a loss of 6.43 million disability adjusted life years worldwide, more than any other enteric pathogen (<xref ref-type="bibr" rid="B25">Kirk et al., 2015</xref>). Within this bacterial subspecies, are human-adapted strains responsible for typhoid fever, as well as a large number of animal-derived non-typhoidal strains responsible for foodborne illness. In this study, we have identified species- and subspecies-specific markers, as well as markers predictive of serovar for subspecies enterica. While this study focused on <italic>S. enterica</italic>, the tools and approach are broadly applicable to any species or collection of genomes.</p>
</sec>
<sec id="s1" sec-type="materials|methods">
<title>Materials and Methods</title>
<p>All commands and parameters used to analyze the data and generate the Figures are available as Supplementary File <xref ref-type="supplementary-material" rid="SM1">1</xref>. The scripts used for analyses are available at <ext-link ext-link-type="uri" xlink:href="https://github.com/superphy/gamechanger">https://github.com/superphy/gamechanger</ext-link>. The following is a summary of the methods used.</p>
<sec><title>Data Collection</title>
<p>All <italic>S. enterica</italic> genomes were downloaded from GenBank in nucleotide fasta format. A full listing of the initial 4939 genomes, including GenBank identifier, subspecies, serovar, the number of species-specific core regions present, the number of contigs, and whether the genome passed the quality filtering steps are listed in Supplementary File <xref ref-type="supplementary-material" rid="SM2">2</xref>.</p>
</sec>
<sec><title>Serovar Identification</title>
<p>Most of the <italic>S. enterica</italic> genomes in GenBank had serovar provided as part of their metadata; however, 321 were missing this designation. The SISTR web-server, as well as the SISTR commandline app were used to predict the serovar for these strains (<xref ref-type="bibr" rid="B66">Yoshida et al., 2016b</xref>).</p>
</sec>
<sec><title>Pan-genome Analyses</title>
<p>Panseq (commit:1d0ab9d37e8e358d266e1d0aa80e9b27f28a1def) was used to identify the pan-genome of the 4939 strains in this study (<xref ref-type="bibr" rid="B29">Laing et al., 2010</xref>). Genomes were initially fragmented into 1000 bp segments, and subsequently clustered using cd-hit v.4.6 to remove potential duplicates/paralogs from the analyses using a 90% sequence identity threshold (<xref ref-type="bibr" rid="B15">Fu et al., 2012</xref>). Initially Panseq was used to determine the distribution of the pan-genome among the genomes at a 90% sequence identity threshold, from which a &#x201C;conserved core&#x201D; was identified. Within the conserved core, Panseq was then used to identify single-nucleotide polymorphisms.</p>
</sec>
<sec><title>Identification of <italic>S. enterica</italic> Species-Specific Regions</title>
<p>To identify regions that were likely to represent the species as a whole, we initially examined the 211 closed <italic>S. enterica</italic> genomes in GenBank (Supplementary File <xref ref-type="supplementary-material" rid="SM2">2</xref>), and identified 3832 regions of 1000 bp that were found in 90% (190) of the 211 closed genomes using Panseq, at a 90% sequence identity threshold. These regions were then screened against the online GenBank nr database using megablast as a first-pass filter with default parameters, searching across bacteria (taxid:2), and excluding all <italic>Salmonella</italic> (taxid:590) hits that had greater than 80% identity across 80% of the query length from the results. The remaining 1482 genomic regions were subsequently screened against the online GenBank nr database of all bacteria (taxid:2), using the blastn algorithm, to identify matches that were missed using the less-specific megablast algorithm, with word size 11, an e-value cutoff of 0.001, and excluding all <italic>Salmonella</italic> (taxid:590). These results were filtered in the same manner, leaving 405 potentially species-specific regions. Lastly, these regions were compared against <italic>Salmonella bongori</italic> genomes in GenBank; one <italic>S. bongori</italic> hit was identified, which left 404 genomic regions present in <italic>S. enterica</italic> but no other bacterial genomic sequences within the GenBank nr database.</p>
<p>The putative function of these regions was determined by screening them across the GenBank nr database using blastx with &#x201C;max hits:10,&#x201D; &#x201C;taxid limit:1236 (gammaproteobacteria),&#x201D; and an &#x201C;e-value threshold: 0.001.&#x201D; The best matching hit above a 90% sequence identity threshold was used for the putative functional assignment.</p>
</sec>
<sec><title>Identification of Subspecies- and Serovar-Specific Regions</title>
<p>The Fisher&#x2019;s Exact test, using the Bonferroni correction for multiple testing was applied as in the SuperPhy platform (<xref ref-type="bibr" rid="B64">Whiteside et al., 2016</xref>), implemented here as the standalone program feht<sup><xref ref-type="fn" rid="fn01">1</xref></sup>. The input for the program was Supplementary File <xref ref-type="supplementary-material" rid="SM2">2</xref>, which contained metadata for all the strains, as well as the binary_table.txt output file from the Panseq analyses, which denotes the presence/absence of each 1000 bp pan-genome region among all the strains.</p>
</sec>
<sec><title><italic>S. enterica</italic> Phylogenetic Analyses</title>
<p>The phylogeny based on SNPs within the core genome was generated using RAxML v8.2.9, with the snp.phylip output file from Panseq (<xref ref-type="bibr" rid="B58">Stamatakis, 2014</xref>). The phylogeny based on the presence/absence of the pan-genome was also generated using RAxML v8.2.9, with the binary.phylip output file from Panseq.</p>
</sec>
<sec><title>Generation of Figures and Tables</title>
<p>The R-statistical language v3.3.2 was used to generate the summary Figures and Tables (<xref ref-type="bibr" rid="B48">R Core Team, 2016</xref>). The R-scripts and all others used for the analyses can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/superphy/gamechanger/tree/master/src">https://github.com/superphy/gamechanger/tree/master/src</ext-link>. The ggtree package for R was used in the generation of the phylogenetic tree images (<xref ref-type="bibr" rid="B68">Yu et al., 2016</xref>).</p>
</sec>
</sec>
<sec><title>Results</title>
<sec><title><italic>S. enterica</italic> Pan-genome</title>
<p>We initially determined the size and distribution of the <italic>S. enterica</italic> pan-genome as genome fragments of 1000 bp in size, across the 4939 genome sequences of this study, which are summarized by subspecies in <bold>Table <xref ref-type="table" rid="T1">1</xref></bold>, and within subspecies <italic>enterica</italic> by serovar in <bold>Table <xref ref-type="table" rid="T2">2</xref></bold>. As can be seen in <bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>, the pan-genome comprised of 4939 <italic>S. enterica</italic> genomes was found to be 25.3 Mbp in size, with 70% of the pan-genome present in fewer than 100 strains. Conversely, the core genome was found to be 1.5 Mbp in size, with all but 200 genomes (96%) containing 3.2 Mbp of shared genomic core. Only 17% of the pan-genome was found in greater than 100 genomes, but fewer than 4739 genomes.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>The frequency of the subspecies observed within the study set of 4936 <italic>Salmonella enterica</italic> genomes, prior to any quality filtering.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Subspecies</th>
<th valign="top" align="center">No.</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>enterica</italic></td>
<td valign="top" align="center">4913</td>
</tr>
<tr>
<td valign="top" align="left"><italic>arizonae</italic></td>
<td valign="top" align="center">7</td>
</tr>
<tr>
<td valign="top" align="left"><italic>diarizonae</italic></td>
<td valign="top" align="center">7</td>
</tr>
<tr>
<td valign="top" align="left"><italic>houtenae</italic></td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left"><italic>salamae</italic></td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left"><italic>indica</italic></td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>The serovars with more than 20 representatives in the current study set of 4936 <italic>Salmonella enterica</italic> genomes, and their frequency, prior to any quality filtering.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Serovar</th>
<th valign="top" align="center">No.</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Typhi</td>
<td valign="top" align="center">1977</td>
</tr>
<tr>
<td valign="top" align="left">Typhimurium</td>
<td valign="top" align="center">758</td>
</tr>
<tr>
<td valign="top" align="left">Enteritidis</td>
<td valign="top" align="center">413</td>
</tr>
<tr>
<td valign="top" align="left">Heidelberg</td>
<td valign="top" align="center">201</td>
</tr>
<tr>
<td valign="top" align="left">Paratyphi</td>
<td valign="top" align="center">158</td>
</tr>
<tr>
<td valign="top" align="left">Kentucky</td>
<td valign="top" align="center">155</td>
</tr>
<tr>
<td valign="top" align="left">Agona</td>
<td valign="top" align="center">136</td>
</tr>
<tr>
<td valign="top" align="left">Weltevreden</td>
<td valign="top" align="center">120</td>
</tr>
<tr>
<td valign="top" align="left">Bareilly</td>
<td valign="top" align="center">106</td>
</tr>
<tr>
<td valign="top" align="left">Newport</td>
<td valign="top" align="center">82</td>
</tr>
<tr>
<td valign="top" align="left">Tennessee</td>
<td valign="top" align="center">77</td>
</tr>
<tr>
<td valign="top" align="left">Montevideo</td>
<td valign="top" align="center">69</td>
</tr>
<tr>
<td valign="top" align="left">Saintpaul</td>
<td valign="top" align="center">48</td>
</tr>
<tr>
<td valign="top" align="left">Infantis</td>
<td valign="top" align="center">39</td>
</tr>
<tr>
<td valign="top" align="left">Senftenberg</td>
<td valign="top" align="center">35</td>
</tr>
<tr>
<td valign="top" align="left">Bovismorbificans</td>
<td valign="top" align="center">34</td>
</tr>
<tr>
<td valign="top" align="left">Hadar</td>
<td valign="top" align="center">33</td>
</tr>
<tr>
<td valign="top" align="left">Muenchen</td>
<td valign="top" align="center">30</td>
</tr>
<tr>
<td valign="top" align="left">Anatum</td>
<td valign="top" align="center">27</td>
</tr>
<tr>
<td valign="top" align="left">Schwarzengrund</td>
<td valign="top" align="center">27</td>
</tr>
<tr>
<td valign="top" align="left">Dublin</td>
<td valign="top" align="center">24</td>
</tr>
<tr>
<td valign="top" align="left">Cerro</td>
<td valign="top" align="center">21</td>
</tr>
<tr>
<td valign="top" align="left"></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<attrib><italic>The list of all serovars and their frequency within the current study is available as Supplementary File <xref ref-type="supplementary-material" rid="SM2">2</xref>.</italic></attrib>
</table-wrap-foot>
</table-wrap>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The distribution of the <italic>Salmonella enterica</italic> pan-genome, as 1000 bp fragments, among 4939 whole-genome sequences (WGSs).</p></caption>
<graphic xlink:href="fmicb-08-01345-g001.tif"/>
</fig>
</sec>
<sec><title><italic>S. enterica</italic> Species-Specific Regions</title>
<p>To identify regions of <italic>S. enterica</italic> that were likely to be shared among most genomes of the species, we examined all 211 closed genomes of <italic>S. enterica</italic> in GenBank, looking for genomic regions that were present in at least 190 (90%) of these genomes. We identified 3832 regions of 1000 bp that were present in at least 90% of the closed genomes. These regions were subsequently screened against the GenBank nr database, and any present in non-<italic>Salmonella</italic> genomes were removed, leaving 404 putative <italic>S. enterica</italic> species-specific regions (Supplementary File <xref ref-type="supplementary-material" rid="SM3">3</xref>).</p>
<p><bold>Figure <xref ref-type="fig" rid="F2">2</xref></bold> shows the carriage of these 404 regions among the 4939 genomes of this study. All but 105 genomes contained at least 330 of these putative <italic>S. enterica</italic> specific regions. A stark difference in carriage of these species-specific markers was observed, with 4742 genomes containing at least 350 species-specific markers, while only 2674 genomes contained 360 or more species-specific markers.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>The carriage of the 404 <italic>S. enterica</italic> species-specific regions among each of the 4939 genomes of this study. Each dot represents a single <italic>S. enterica</italic> genome, which are arranged in order from those that contain the fewest species-specific regions to those that contain the most.</p></caption>
<graphic xlink:href="fmicb-08-01345-g002.tif"/>
</fig>
</sec>
<sec><title>Quality Filtering for Subsequent Analyses</title>
<p>To ensure the quality of the genomes in use for subsequent analyses, we plotted carriage of the 404 species-specific regions versus the number of contigs that each sequenced genome was comprised of (<bold>Figure <xref ref-type="fig" rid="F3">3</xref></bold>). As can be seen, the two genomes marked in yellow contained only one, and the same, species-specific region each, despite being comprised of relatively few contigs. Subsequent searches against the GenBank nr database identified these two genomes as <italic>Citrobacter</italic> spp. contamination, mislabeled as <italic>S. enterica</italic> (GCA_001570325 and GCA_001570345). The &#x201C;<italic>Salmonella enterica</italic> species-specific region&#x201D; found in both of the contaminant <italic>Citrobacter</italic> genomes, did not match any other <italic>Citrobacter</italic> spp. in GenBank above the thresholds used for determining presence/absence in this study. However, due to the presence of this region in what have been identified as <italic>Citrobacter</italic> genomes, the region was removed from subsequent analyses.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>The carriage of the 404 <italic>S. enterica</italic> species-specific regions, versus the number of contigs for each of the 4936 genomes. Colors indicate the subspecies within <italic>S. enterica</italic> as follows: red: arizonae, lime: diarizonae, teal: enterica, blue: houtenae, lavender: indica, magenta: salamae and yellow: sample with <italic>Citrobacter</italic> contamination.</p></caption>
<graphic xlink:href="fmicb-08-01345-g003.tif"/>
</fig>
<p>The majority of genomes (4913) were from subspecies <italic>enterica</italic>, with genomes from the five other <italic>S. enterica</italic> subspecies present in drastically fewer numbers (<bold>Table <xref ref-type="table" rid="T1">1</xref></bold>). All closed genomes from subspecies <italic>enterica</italic> contained greater than 250 species-specific regions, which was more than the genomes from any other subspecies, with the exception of <italic>enterica</italic> genomes that were of poor quality and comprised of many 1000s of contigs (<bold>Figure <xref ref-type="fig" rid="F3">3</xref></bold>). Genomes from subspecies <italic>houtenae</italic> and <italic>arizonae</italic> contained fewer than 100 species-specific regions, while genomes from <italic>diarizonae</italic>, <italic>indica</italic>, and <italic>salamae</italic> contained between 100 and 200 species-specific regions. All regions were screened against <italic>S. bongori</italic> to ensure specificity to <italic>S. enterica</italic>; one region was found to also be present in genomes from <italic>S. bongori</italic> and was removed from further analyses.</p>
<p>Within subspecies <italic>enterica</italic>, a negative linear relationship was observed among the number of species-specific regions contained within a genome, and the number of contigs the genome was comprised of, with the worst-case genome (GCA_000495155) being comprised of 6945 contigs, but containing only 13 species-specific regions. Other genomes such as <italic>S. enterica</italic> Bovismorbificans strain GCA_001114865 contained both few contigs (140) as well as fewer species-specific regions (209) than other <italic>enterica</italic> genomes. Additional searches discovered sequencing gaps within the genome totaling over 464 Kbp. A final outlier genome harbored nearly 5000 contigs, but also contained 403 of the species-specific regions. It was determined by searching the GenBank database, that this sequence (GCA_000765055) was actually a combination of multiple genomes in a single file.</p>
<p>Given the above information, all genomes from the five subspecies other than <italic>enterica</italic> were included in subsequent analyses, while the thresholds for inclusion of <italic>enterica</italic> genomes were set at a maximum of 1000 contigs, and a minimum of 250 species-specific regions. Following this quality filtering, 43 genomes were removed, leaving 4870 <italic>S. enterica</italic> subspecies <italic>enterica</italic> genomes for the following analyses.</p>
</sec>
<sec><title>Phylogeny of <italic>S. enterica</italic> Using the Conserved Core Genome</title>
<p>Based on the distribution of the pan-genome presented in <bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>, the &#x201C;conserved core&#x201D; of <italic>S. enterica</italic> was set at being present in more that 4500 genomes, to fully capture the conserved genomic regions within the species. A phylogeny based on the SNPs among these shared regions was created, and is shown along with the distribution of the <italic>S. enterica</italic> species-specific regions in <bold>Figure <xref ref-type="fig" rid="F4">4</xref></bold>. As can be seen, the majority of the genomes are subspecies <italic>enterica</italic>, and the other five subspecies are relatively more distant in the order of <italic>indica</italic>, <italic>salamae</italic>, <italic>houtenae</italic>, <italic>diarizonae</italic>, and <italic>arizonae</italic>. However, the order of subspecies in declining number of species-specific regions is: <italic>enterica</italic>, <italic>diarizonae</italic>, <italic>salamae</italic>, <italic>indica</italic>, <italic>houtenae</italic>, and <italic>arizonae</italic>, which is shown in <bold>Figure <xref ref-type="fig" rid="F3">3</xref></bold>.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>The phylogeny of the 4893 <italic>S. enterica</italic> genomes post quality-filtering, and limiting the number of genomes from each serovar to five. The name of each serovar is presented as text, and the six subspecies are shown as colored circles as follows: teal: arizonae, blue: diarizonae, dark orange: enterica, peach: houtenae, dark green: indica, light orange: salamae.</p></caption>
<graphic xlink:href="fmicb-08-01345-g004.tif"/>
</fig>
<p>The serovar distribution within subspecies <italic>enterica</italic> was shown to be largely concordant with phylogeny, as demonstrated in <bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>, where the 10 most abundant serovars in the current study are highlighted. However, not all serovars clustered as monophyletic groups, as can be seen with serovar Bareilly; nor were all clades found to be comprised of single serovars, demonstrated by the clade containing genomes of serovars Bareilly and Agona.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>The phylogeny of the 4893 <italic>S. enterica</italic> genomes post quality-filtering based on SNPs found within the conserved core genome. The 10 most abundant serovars of subspecies enterica in the current study (Agona, Bareilly, Enteritidis, Heidelberg, Kentucky, Newport, Paratyphi, Typhi, Typhimurium, Weltevreden) are labeled on the tree. The matrix to the right of the phylogeny represents the 404 species-specific regions, with blue being the absence of a region, and green being the presence of a region, for each of the genomes of the study.</p></caption>
<graphic xlink:href="fmicb-08-01345-g005.tif"/>
</fig>
<p>The large clades within the phylogenetic tree also demonstrate clade-specific patterns of presence/absence for the 404 species-specific markers. Among the most abundant serovars, Typhimurium, Heidelberg, Newport, and Enteritidis were found to contain the most species-specific markers, and grouped together near the center of the tree. Likewise, serovars Agona, Welevreden, and Kentucky contained fewer species-specific regions, and group together near the bottom of the tree, closer to the non-<italic>enterica</italic> sub-species genomes.</p>
<p><bold>Table <xref ref-type="table" rid="T3">3</xref></bold> considers all serovars with at least 10 members in the dataset, and the average number of species-specific markers per serovar. As can be seen, the serovars with the largest average number of species-specific regions were: Enteritidis (401.7), Anatum (401.5), Muenchen (400.5), Hadar (400.3), and Typhimurium (400.1); conversely, the serovars with the fewest average number of species-specific regions were: Derby (360.7), Montevideo (360.1), Typhi (358.1), Bovismorbificans (355.3), and Cerro (342.0).</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>The average number of species-specific genomic regions found among serovars of subspecies <italic>enterica</italic>, that contained at least 10 representative genomes, within the 4870 quality filtered subspecies <italic>enterica</italic> genomes of this study.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Serovar</th>
<th valign="top" align="center">Average no. species-specific regions</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Enteritidis</td>
<td valign="top" align="center">401.7</td>
</tr>
<tr>
<td valign="top" align="left">Anatum</td>
<td valign="top" align="center">401.5</td>
</tr>
<tr>
<td valign="top" align="left">Muenchen</td>
<td valign="top" align="center">400.5</td>
</tr>
<tr>
<td valign="top" align="left">Hadar</td>
<td valign="top" align="center">400.3</td>
</tr>
<tr>
<td valign="top" align="left">Typhimurium</td>
<td valign="top" align="center">400.1</td>
</tr>
<tr>
<td valign="top" align="left">Newport</td>
<td valign="top" align="center">399.8</td>
</tr>
<tr>
<td valign="top" align="left">Thompson</td>
<td valign="top" align="center">399.7</td>
</tr>
<tr>
<td valign="top" align="left">Saintpaul</td>
<td valign="top" align="center">399.6</td>
</tr>
<tr>
<td valign="top" align="left">Heidelberg</td>
<td valign="top" align="center">397.4</td>
</tr>
<tr>
<td valign="top" align="left">Dublin</td>
<td valign="top" align="center">395.2</td>
</tr>
<tr>
<td valign="top" align="left">Infantis</td>
<td valign="top" align="center">394.9</td>
</tr>
<tr>
<td valign="top" align="left">Braenderup</td>
<td valign="top" align="center">392.8</td>
</tr>
<tr>
<td valign="top" align="left">Weltevreden</td>
<td valign="top" align="center">390.0</td>
</tr>
<tr>
<td valign="top" align="left">Bareilly</td>
<td valign="top" align="center">388.5</td>
</tr>
<tr>
<td valign="top" align="left">Kentucky</td>
<td valign="top" align="center">380.3</td>
</tr>
<tr>
<td valign="top" align="left">Plymouth/Zega</td>
<td valign="top" align="center">377.9</td>
</tr>
<tr>
<td valign="top" align="left">Senftenberg</td>
<td valign="top" align="center">376.5</td>
</tr>
<tr>
<td valign="top" align="left">Mbandaka</td>
<td valign="top" align="center">374.5</td>
</tr>
<tr>
<td valign="top" align="left">Lubbock</td>
<td valign="top" align="center">374.1</td>
</tr>
<tr>
<td valign="top" align="left">Reading</td>
<td valign="top" align="center">370.4</td>
</tr>
<tr>
<td valign="top" align="left">Agona</td>
<td valign="top" align="center">369.5</td>
</tr>
<tr>
<td valign="top" align="left">Tennessee</td>
<td valign="top" align="center">368.3</td>
</tr>
<tr>
<td valign="top" align="left">Schwarzengrund</td>
<td valign="top" align="center">362.3</td>
</tr>
<tr>
<td valign="top" align="left">Paratyphi</td>
<td valign="top" align="center">361.5</td>
</tr>
<tr>
<td valign="top" align="left">Derby</td>
<td valign="top" align="center">360.7</td>
</tr>
<tr>
<td valign="top" align="left">Montevideo</td>
<td valign="top" align="center">360.1</td>
</tr>
<tr>
<td valign="top" align="left">Typhi</td>
<td valign="top" align="center">358.1</td>
</tr>
<tr>
<td valign="top" align="left">Bovismorbificans</td>
<td valign="top" align="center">355.3</td>
</tr>
<tr>
<td valign="top" align="left">Cerro</td>
<td valign="top" align="center">342.0</td>
</tr>
<tr>
<td valign="top" align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec><title>Phylogeny of <italic>S. enterica</italic> Using the Pan-genome</title>
<p>A phylogeny based on the presence/absence of the pan-genome among the 4893 <italic>S. enterica</italic> genomes was created, and is shown along with the distribution of the <italic>S. enterica</italic> species-specific regions in <bold>Figure <xref ref-type="fig" rid="F6">6</xref></bold>. As can be seen this phylogeny based on the presence/absence of the entire 25.3 Mbp pan-genome is highly concordant with the phylogeny based on the SNPs found in the conserved core of the same strains (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). In both trees the serovars cluster together and in the same relation to each other, for example serovars Typhi and Paratyphi strains form a discrete monophyletic clade. However, the branch lengths in the pan-genome tree are larger than those in the conserved SNP tree, due to the larger variation among the presence/absence of the pan-genome than to sequence variation among shared core regions.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>The phylogeny of the 4893 <italic>S. enterica</italic> genomes post quality-filtering based on the presence/absence of the entire pan-genome as 1000 bp fragments. The 10 most abundant serovars of subspecies enterica in the current study (Agona, Bareilly, Enteritidis, Heidelberg, Kentucky, Newport, Paratyphi, Typhi, Typhimurium, Weltevreden) are labeled on the tree. The matrix to the right of the phylogeny represents the 404 species-specific regions, with blue being the absence of a region, and green being the presence of a region, for each of the genomes of the study.</p></caption>
<graphic xlink:href="fmicb-08-01345-g006.tif"/>
</fig>
</sec>
<sec><title>Identification of a Minimum Set of Species-Specific Genomic Markers</title>
<p>Within the 404 species-specific markers, none were specific for any of the subspecies. That is, a marker was always present in genomes from at least two subspecies.</p>
<p>We next determined that the presence of a minimum set of two genomic regions was required to unambiguously identify genomes of <italic>S. enterica</italic>, within the 4893 genomes of the current study. A combination of two genomic regions were all that was required, and two such markers that were also present in the most <italic>S. enterica</italic> genomes were found at the following locations within the Typhimurium reference genome LT2: (1336001.. 1337000) and (2467001.. 2468000) (Supplementary File <xref ref-type="supplementary-material" rid="SM3">3</xref>). All members of <italic>S. enterica</italic> examined contained at least one of these markers, but many other combinations within the 404 species-specific markers are also possible.</p>
</sec>
<sec><title>Putative Functional Identification of the <italic>S. enterica</italic> Species-Specific Regions</title>
<p>The putative function of the 404 quality-filtered <italic>S. enterica</italic> species-specific regions were determined from the GenBank nr database. The annotation of each of the 404 regions is available as Supplementary File <xref ref-type="supplementary-material" rid="SM1">1</xref>. <bold>Table <xref ref-type="table" rid="T4">4</xref></bold> summarizes the frequency of functional annotation categories, after annotating each region with the single best match. As can be seen, hypothetical proteins accounted for the majority (64) of the 404 annotations, with secreted effector and membrane proteins being the next most frequent category among the species-specific regions. Other membrane, transport, and secretion proteins were observed. The species-specific regions also included proteins involved in core metabolic functions, protein and DNA synthesis, and response to stress.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>The putative function of the <italic>S. enterica</italic> species-specific regions for functions that were identified more than once, utilizing the best hit for each region.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Putative protein function</th>
<th valign="top" align="center">Frequency</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Hypothetical</td>
<td valign="top" align="center">64</td>
</tr>
<tr>
<td valign="top" align="left">Secreted effector</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="left">Membrane</td>
<td valign="top" align="center">7</td>
</tr>
<tr>
<td valign="top" align="left">Secretion system apparatus</td>
<td valign="top" align="center">5</td>
</tr>
<tr>
<td valign="top" align="left">Uncharacterized</td>
<td valign="top" align="center">5</td>
</tr>
<tr>
<td valign="top" align="left">Fimbrial</td>
<td valign="top" align="center">5</td>
</tr>
<tr>
<td valign="top" align="left">Pathogenicity island 2 effector</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">Fimbrial assembly</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">Outer membrane usher</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">mfs transporter</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Oxidoreductase</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Histidine kinase</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Putative inner membrane</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Putative cytoplasmic</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">lysr family transcriptional regulator</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Transcriptional regulator</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Permease</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Outer membrane</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Type III secretion</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Phosphoglycerate transport</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">arac family transcriptional regulator</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Conserved hypothetical</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Methyl-accepting chemotaxis</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Hybrid sensor histidine kinase/response regulator</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Glycosyl transferase, partial</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Phenylacetaldehyde dehydrogenase</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Pathogenicity island 1 effector</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left"><italic>n</italic>-Acetylneuraminic acid mutarotase, partial</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Type III secretion system</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Transcriptional regulator, partial</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Cytoplasmic</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Fimbrial chaperone</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Putative sialic acid transporter</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left"></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<attrib><italic>The complete list of all putative functions is available as Supplementary File <xref ref-type="supplementary-material" rid="SM3">3</xref>.</italic></attrib>
</table-wrap-foot>
</table-wrap>
</sec>
<sec><title>Identification of Subspecies-Specific Markers from the Pan-genome</title>
<p>Having identified species-specific markers, we employed the same techniques, utilizing the presence/absence of all pan-genome markers, just as was carried out in identifying the 404 species-specific ones, to identify subspecies-specific markers. The number of markers that were completely unique to a subspecies is given in <bold>Table <xref ref-type="table" rid="T5">5</xref></bold>. Subspecies <italic>arizonae</italic> contained the most unique markers, at 207, and <italic>enterica</italic> contained the least, at 9.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>The number of subspecies-specific pan-genome markers that were universally present or absent among members of the subspecies, and not absent or present among genomes from any other subspecies.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Subspecies</th>
<th valign="top" align="center">No. markers</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>arizonae</italic></td>
<td valign="top" align="center">207</td>
</tr>
<tr>
<td valign="top" align="left"><italic>diarizonae</italic></td>
<td valign="top" align="center">93</td>
</tr>
<tr>
<td valign="top" align="left"><italic>enterica</italic></td>
<td valign="top" align="center">9</td>
</tr>
<tr>
<td valign="top" align="left"><italic>houtenae</italic></td>
<td valign="top" align="center">134</td>
</tr>
<tr>
<td valign="top" align="left"><italic>indica</italic></td>
<td valign="top" align="center">192</td>
</tr>
<tr>
<td valign="top" align="left"><italic>salamae</italic></td>
<td valign="top" align="center">135</td>
</tr>
<tr>
<td valign="top" align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec><title>Identification of Universal Serovar Markers within <italic>Subspecies enterica</italic> from the Pan-genome</title>
<p>Subspecies <italic>enterica</italic> genomes were the vast majority of those available, so we attempted to identify serovar-specific markers for the top 10 serovars, in the same manner that we identified subspecies-specific markers. We found that there were no genomic markers that uniquely defined any of the serovars based on their presence or absence; however, there were a number of genomic regions that were universally present or absent among serovars, as well as statistically over- or under- represented with respect to all other serovar genomes from this study; they are shown in <bold>Table <xref ref-type="table" rid="T6">6</xref></bold>.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>The number of pan-genome regions that were universally present and absent, as well as statistically over- or under-represented in comparison to all other genomes, within the 10 most abundant serovars within the 4870 subspecies <italic>enterica</italic> genomes of this study.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Serovar</th>
<th valign="top" align="center">No. universally present</th>
<th valign="top" align="center">No. universally absent</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Typhi</td>
<td valign="top" align="center">288</td>
<td valign="top" align="center">2720</td>
</tr>
<tr>
<td valign="top" align="left">Typhimurium</td>
<td valign="top" align="center">41</td>
<td valign="top" align="center">698</td>
</tr>
<tr>
<td valign="top" align="left">Enteritidis</td>
<td valign="top" align="center">18</td>
<td valign="top" align="center">440</td>
</tr>
<tr>
<td valign="top" align="left">Heidelberg</td>
<td valign="top" align="center">121</td>
<td valign="top" align="center">840</td>
</tr>
<tr>
<td valign="top" align="left">Paratyphi</td>
<td valign="top" align="center">65</td>
<td valign="top" align="center">202</td>
</tr>
<tr>
<td valign="top" align="left">Kentucky</td>
<td valign="top" align="center">177</td>
<td valign="top" align="center">331</td>
</tr>
<tr>
<td valign="top" align="left">Agona</td>
<td valign="top" align="center">161</td>
<td valign="top" align="center">638</td>
</tr>
<tr>
<td valign="top" align="left">Weltevreden</td>
<td valign="top" align="center">426</td>
<td valign="top" align="center">608</td>
</tr>
<tr>
<td valign="top" align="left">Bareilly</td>
<td valign="top" align="center">87</td>
<td valign="top" align="center">436</td>
</tr>
<tr>
<td valign="top" align="left">Newport</td>
<td valign="top" align="center">226</td>
<td valign="top" align="center">360</td>
</tr>
<tr>
<td valign="top" align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To further assess the validity of these markers, a dataset comprised of 3948 genomes from EnteroBase<sup><xref ref-type="fn" rid="fn02">2</xref></sup>, was selected to have an identical number of strains belonging to each of nine serovars in our GenBank dataset. The EnteroBase dataset was used to test the predictive markers we identified from the GenBank dataset in the first part of the study. The results of this comparison are shown in <bold>Figure <xref ref-type="fig" rid="F7">7</xref></bold>. As can be seen, the markers were well-conserved among the EnteroBase dataset, with eight of the nine serovars having a subset of the predictive markers present among all of the test genomes; serovar Typhimurium had a marker subset that was present in all but one of the test genomes.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>The number of predictive markers from the GenBank dataset found within the EnteroBase dataset for nine serovars of <italic>S. enterica</italic>, which encompassed a test set of 3948 genomes. The number of genomes for each serovar was the same between the GenBank and EnteroBase datasets, as shown in <bold>Table <xref ref-type="table" rid="T6">6</xref></bold>. The size of the circles is proportional to the number of predictive markers from the GenBank dataset found in the EnteroBase dataset. The number of genomes for each serovar is given in the horizontal axis label. Using serovar Agona as an example, there were 136 genomes in both the GenBank and EnteroBase datasets, and 129 of the 161 predictive markers from the GenBank dataset were found in all of the genomes from the EnteroBase dataset, whereas 21 of the GenBank predictive markers were found in all but one (135) of the EnteroBase genomes examined.</p></caption>
<graphic xlink:href="fmicb-08-01345-g007.tif"/>
</fig>
</sec>
</sec>
<sec><title>Discussion</title>
<sec><title><italic>S. enterica</italic> Pan-genome</title>
<p>Previous examinations of the <italic>S. enterica</italic> pan-genome were based on relatively small datasets of 45 and 73 genomes (<xref ref-type="bibr" rid="B24">Jacobsen et al., 2011</xref>; <xref ref-type="bibr" rid="B30">Leekitcharoenphon et al., 2012</xref>). While others have analyzed 1000s of <italic>S. enterica</italic> genomes, the analyses were not conducted to examine the population structure. For example, in demonstrating the software program Roary, 1000 <italic>S.</italic> Typhi genomes were used to test the program (<xref ref-type="bibr" rid="B43">Page et al., 2015</xref>). Likewise, the GenomeTrackR project utilized 32 <italic>S. enterica</italic> genomes to identify a <italic>S. enterica</italic> core, which was subsequently used as the basis for genetic distance estimates for nearly 20,000 genomes (<xref ref-type="bibr" rid="B45">Pettengill et al., 2016</xref>).</p>
<p>Previous estimates placed the core-genome size of <italic>S. enterica</italic> at &#x223C;2800 gene families, and the pan-genome at &#x223C;10,000 gene families (<xref ref-type="bibr" rid="B24">Jacobsen et al., 2011</xref>). The current study identified a strict core of 1.5 Mbp, and a conserved core of 3.2 Mbp shared among 96% of the genomes, which given an average gene size of 1000 bp is &#x223C;1500 and &#x223C;3200 genes respectively, with a much larger pan-genome of &#x223C;25,300 genes. Previous analyses found <italic>S. enterica</italic> to have a closed pan-genome (<xref ref-type="bibr" rid="B24">Jacobsen et al., 2011</xref>), and thus the rate of discovery for new genomic regions would decrease for each new genome of the species sequenced (<xref ref-type="bibr" rid="B60">Tettelin et al., 2005</xref>).</p>
<p>In line with <italic>S. enterica</italic> having a closed pan-genome, when we compared it to <italic>E. coli</italic>, a related bacterial species with an open pan-genome (<xref ref-type="bibr" rid="B60">Tettelin et al., 2005</xref>), we found that the <italic>E. coli</italic> pan-genome was larger (37.4 Mbp), despite the fact that the <italic>E. coli</italic> study used less than half the number of strains in the current <italic>Salmonella enterica</italic> study. Additionally, more of the pan-genome of <italic>S. enterica</italic> was distributed among more genomes than in <italic>E. coli</italic> (<xref ref-type="bibr" rid="B64">Whiteside et al., 2016</xref>). Specifically, in <italic>S. enterica</italic> 70% of the pan-genome was found to belong to 100 or fewer of the genomes examined, while in <italic>E. coli</italic> 80% of the pan-genome was found in 100 or fewer genomes.</p>
<p>It should be noted that erroneously labeled, and poor quality assemblies, can greatly affect the size, analyses, and composition of the pan-genome. Software tools to evaluate assembly quality have been created to help researchers identify bad data. These include QUAST (<xref ref-type="bibr" rid="B20">Gurevich et al., 2013</xref>), which summarizes the assembly statistics including average contig size and number of contigs; as well as CGAL (<xref ref-type="bibr" rid="B49">Rahman and Pachter, 2013</xref>), which uses a likelihood approach to infer assembly quality rather than summary statistics. As demonstrated in the current study, having a known set of species-specific genome regions can facilitate rapid quality assessment and filtering of genome assemblies. Others have proposed whole-genome MLST for this purpose as well (<xref ref-type="bibr" rid="B4">Babenko et al., 2016</xref>; <xref ref-type="bibr" rid="B66">Yoshida et al., 2016b</xref>), but the benefit of a pan-genome analysis is that it is schema free, requiring no agreed upon reference set or central repository of alleles.</p>
</sec>
<sec><title><italic>S. enterica</italic> Species-Specific Regions</title>
<p>Previous studies have identified gene targets that are useful in the identification of <italic>Salmonella</italic>. These include the <italic>fimA</italic> gene (<xref ref-type="bibr" rid="B10">Cohen et al., 1996</xref>), <italic>hilA</italic> (<xref ref-type="bibr" rid="B19">Guo et al., 2000</xref>), <italic>invA</italic> (<xref ref-type="bibr" rid="B34">Malorny et al., 2003</xref>), <italic>ttr</italic> (<xref ref-type="bibr" rid="B35">Malorny et al., 2004</xref>), and <italic>ssaN</italic> (<xref ref-type="bibr" rid="B9">Chen et al., 2010</xref>). Other markers, and combinations thereof have been developed for use in RT-PCR (<xref ref-type="bibr" rid="B47">Postollec et al., 2011</xref>), and other detection platforms such as loop-mediated isothermal amplification (<xref ref-type="bibr" rid="B27">Kokkinos et al., 2014</xref>). Additionally, the identification of serovar based on allelic variation in somatic and flagellar genes has previously been conducted, with at least four laboratory methods currently available [the <italic>Salmonella</italic> genoserotyping assay (<xref ref-type="bibr" rid="B67">Yoshida et al., 2014</xref>), and the commerical assays: <italic>Salmonella</italic> Serogenotyping Assay, Check&#x0026;Trace <italic>Salmonella</italic>, and xMAP <italic>Salmonella</italic> serotyping assay], capable of identifying over 100 of the most common <italic>Salmonella enterica</italic> serovars in some cases (<xref ref-type="bibr" rid="B65">Yoshida et al., 2016a</xref>). The recently released software, the <italic>Salmonella in silico</italic> typing resource (SISTR), is capable of providing <italic>Salmonella</italic> serovar prediction from WGSs for 90% (2,190) of all serovars (<xref ref-type="bibr" rid="B66">Yoshida et al., 2016b</xref>).</p>
<p>Despite the utility of the previously mentioned methods, previous marker-discover studies have used at most 100s of <italic>Salmonella</italic> strains, while the current study examines nearly 5000. Further, the current study analyzes the entire pan-genome for predictive markers, and identified over 400 that were specific to the species, as well as others being predictive for both subspecies and serovar.</p>
<p>The host intestinal environment consists of a multitude of bacterial species competing for scarce nutritional sources such as carbohydrates, direct antagonistic competition with other bacterial cells, and competition for access to the host intestine, where stable attachment and colonization of the local environment are possible (<xref ref-type="bibr" rid="B52">Sana et al., 2016</xref>). The normal intestinal microflora offer protection to the host against enteric pathogens such as <italic>S. enterica</italic>, but disruption of the intestinal environment by virulence factors and effector proteins secreted by the pathogen itself, or external factors including antibiotics, have been shown to alter the composition of the microbiota, and allow pathogens such as <italic>S. enterica</italic> to proliferate (<xref ref-type="bibr" rid="B40">Ng et al., 2013</xref>).</p>
<p>Nutritional competition exists for free metabolic compounds, such as carbohydrates that are readily available, as well as others that are sequestered in forms such as the intestinal mucus, which is composed of sialic sugar acids (<xref ref-type="bibr" rid="B37">McDonald et al., 2016</xref>). In the gut, these sugar acids exists as a conjugate in the alpha form, which to be useful for bacteria such as <italic>Salmonella</italic>, need to be converted to the beta form by a mutarotase enzyme (<xref ref-type="bibr" rid="B56">Severi et al., 2008</xref>). In this study, we identified n-acetylneuraminic acid mutarotase genes as species-specific genomic regions, along with sialic acid transporter genes. It is possible the presence of these systems allow <italic>S. enterica</italic> to more efficiently compete with the host microbiota by efficiently utilizing scarce metabolic sources.</p>
<p>It was also previously found that sialic acid on the surface of host colon cells increased colonization by <italic>S.</italic> Typhi, and disialylation of these cells reduced the adherence of the <italic>Salmonella</italic> strains by 41% (<xref ref-type="bibr" rid="B51">Sakarya et al., 2010</xref>). This was also demonstrated in <italic>S.</italic> Typhimurium, where following antibiotic treatment, the presence of free sialic acid increased, and the ability to utilize it was correlated with higher levels of bacterial colonization of the host gut (<xref ref-type="bibr" rid="B40">Ng et al., 2013</xref>).</p>
<p>Enzymes that utilize sialic acids have previously been shown to be present in 452 bacterial species, including other pathogens such as <italic>Vibrio cholerae</italic>, but the genomic regions found in the current study were sufficiently unique at the nucleotide level to be determinative for <italic>S. enterica</italic> (<xref ref-type="bibr" rid="B37">McDonald et al., 2016</xref>).</p>
<p>In addition to species-specific regions used to gain a metabolic advantage, a number of secretion system and effector proteins were identified as diagnostic of <italic>S. enterica</italic>. These included components of the Type VI secretion system (T6SS), which is a contact-dependent, syringe-like secretion system that allows <italic>S. enterica</italic> to directly kill other competing bacteria that it comes into physical contact with (<xref ref-type="bibr" rid="B7">Brunet et al., 2015</xref>), and is encoded on the <italic>Salmonella</italic> Pathogenicity Island 6 (<xref ref-type="bibr" rid="B52">Sana et al., 2016</xref>). It has been demonstrated that silencing the T6SS via H-NS repression (histone-like nucleoid structuring), reduces inter-bacterial killing of <italic>S. enterica</italic> (<xref ref-type="bibr" rid="B7">Brunet et al., 2015</xref>). It was also previously shown that commensal bacteria are killed by <italic>S. enterica</italic> in a T6SS-dependent manner, that the T6SS was required for <italic>Salmonella</italic> to establish infection in the host gut, and that increased concentrations of bile salts resulted in a concomitant increase in T6SS anti-bacterial activity (<xref ref-type="bibr" rid="B52">Sana et al., 2016</xref>). The T6SS itself has been shown to have been independently acquired from four separate lineages within five of the six <italic>S. enterica</italic> subspecies (<xref ref-type="bibr" rid="B13">Desai et al., 2013</xref>).</p>
<p>Like the T6SS, the type III secretion system (T3SS) found within <italic>S. enterica</italic> is a syringe like apparatus that injects effector proteins into host cells (<xref ref-type="bibr" rid="B28">Kubori et al., 2000</xref>). There are two T3SS found within <italic>S. enterica</italic>: the first is encoded on the <italic>Salmonella</italic> Pathogenicity Island 1 (SPI1) and is required for invasion of host cells; the second is encoded on <italic>Salmonella</italic> Pathogenicity Island 2 (SPI2), and is required for survival and proliferation within the host macrophage cells (<xref ref-type="bibr" rid="B22">Hensel et al., 1998</xref>; <xref ref-type="bibr" rid="B6">Bijlsma and Groisman, 2005</xref>). The innate host immune system utilizes the inflammatory response to help reduce the proliferation of bacterial pathogens (<xref ref-type="bibr" rid="B59">Sun et al., 2016</xref>). <italic>S. enterica</italic> has developed a means of regulating host inflammation via the SPI1 T3SS, whereby secreted effector proteins target the NF-&#x03BA;B signaling pathway, reduce inflammation and host tissue damage, and allow increased <italic>S. enterica</italic> propagation within the host. <italic>S. enterica</italic> also relies on free long-chain fatty acids within the host to regulate T3SS expression, and provide a cue to the bacteria to up-regulate genes necessary for host intestinal colonization (<xref ref-type="bibr" rid="B18">Golubeva et al., 2016</xref>).</p>
<p>The current study identified many secretion system and effector proteins as being species-specific, as well as proteins for attachment to the host, such as fimbriae. These proteins allow <italic>S. enterica</italic> to compete within the intestinal environment, and take up residence within the host, where it can proliferate.</p>
<p>Effector proteins and other virulence factors aid in the colonization of the host, and are frequently horizontally acquired and are present on mobile elements such as integrated bacteriophages (<xref ref-type="bibr" rid="B39">Moreno Switt et al., 2013</xref>). Previous work identified clusters of phages that carried virulence factors such as adhesins and antimicrobial resistance determinants within <italic>S. enterica</italic> (<xref ref-type="bibr" rid="B39">Moreno Switt et al., 2013</xref>).</p>
<p>Additionally, many of the genes associated with bacteriophage in <italic>S. enterica</italic> have been found to be of the putative and hypothetical class (<xref ref-type="bibr" rid="B44">Penad&#x00E9;s et al., 2015</xref>). The current study identified a large accessory gene pool that contained many hypothetical and putative genes, which were also the most abundant category of species-specific genomic regions. The proteins of putative and unknown function may aid in colonizing warm-blooded animals, or specific animal or environmental niches. Previous studies identified genotype/phenotype correlations of <italic>S.</italic> Typhimurium that had particular gene complements associated with specific food sources (<xref ref-type="bibr" rid="B21">Hayden et al., 2016</xref>). The same study also postulated that specific phage repertoires may give phylogenetically distant strains a similar accessory gene content, and therefore similar niche specificity. Previously, 285 gene families were identified as being recruited into <italic>S. enterica</italic>, where most of these genes had unknown function, but were postulated to be important for its survival and infection of its host (<xref ref-type="bibr" rid="B13">Desai et al., 2013</xref>). It is therefore not surprising to find that the most abundant species-specific category of genomic regions are those of unknown or putative function; they likely represent genes enhancing the ability of <italic>S. enterica</italic> to propagate within warm-blooded animals, but they have not yet been fully characterized. The other genomic regions diagnostic of <italic>S. enterica</italic> include means for disseminating these fitness genes within the population, competing for resources in the host, and attaching and proliferating. The <italic>S. enterica</italic> species-specific regions likely give a good overview of the factors responsible for making it such an effective pathogen and intestinal inhabitant.</p>
</sec>
<sec><title>Specific Regions for Subspecies and Serovar</title>
<p>The current study recapitulates the phylogenetic relationship of the six <italic>S. enterica</italic> subspecies that has been previously described by others (<xref ref-type="bibr" rid="B13">Desai et al., 2013</xref>). However, the number of species-specific regions found within each subspecies does not follow the same pattern. For example, <italic>diarizonae</italic> is more distantly related to <italic>enterica</italic> than subspecies <italic>indica</italic>, but contains more species-specific regions, and the branch lengths on the tree are shorter. This indicates that although the <italic>diarizonae</italic> strains diverged longer ago than the houtenae strains, they have accumulated less genomic change. Both subspecies <italic>diarizonae</italic> and <italic>houtenae</italic> strains are associated with reptile-acquired salmonellosis (<xref ref-type="bibr" rid="B55">Schroter et al., 2004</xref>; <xref ref-type="bibr" rid="B23">Horvath et al., 2016</xref>), but the differences in genomic change may reflect the specific reptile niches that each inhabit.</p>
<p>Genomic regions specific to each subspecies were identified, the presence of which were unambiguously indicative of each subspecies. The most abundant subspecies in the current analyses, <italic>enterica</italic>, had the fewest specific markers present (9), while the most distantly related subspecies <italic>arizonae</italic>, had the most specific markers (207). These results indicate that just as core genome size decreases with the number of genomes examined, so too do the number of markers &#x201C;core&#x201D; to each subspecies. As more genomes in subspecies <italic>arizonae</italic> and closely related subspecies are examined, we would expect fewer genomic regions to remain specific for the subspecies. This has important implications for designing a set of markers indicative for subspecies, indicating that a group of redundant markers should be used, and that a sampling of the diversity within a subspecies is first required to identify genomic regions that are truly core.</p>
<p>This was also observed within serovar for subspecies <italic>enterica</italic> strains. The original study examining the pan-genome of <italic>S. enterica</italic> used a set of 45 genomes and was able to identify unique gene families for each serovar examined, with Enteritidis having the fewest (29), and Typhi having the most (349) (<xref ref-type="bibr" rid="B24">Jacobsen et al., 2011</xref>). The results of the current study showed no unique genomic regions for any of the serovars with a sample set of 4893 quality filtered genomes. Although genomic regions universally present for each serovar were observed, and followed the same pattern with Enteritidis having the fewest (18), and Typhi having the most (288), these regions were also observed among genomes of other serovars, even though they were statistically over-represented for the serovar in question. The presence of these predictive markers in nearly all of the genomes within the EnteroBase test dataset indicates that the markers are robust, indicative of serovar, and could be combined to determine the likelihood of a genome being of a particular serovar.</p>
<p>When examining the average number of the 404 species-specific regions found among the <italic>enterica</italic> serovars, it was interesting to observe that Enteritidis, which had the fewest number of universal genomic regions, had the highest average number of species-specific regions; likewise Typhi, which had the most universally shared genomic regions, had one of the lowest averages of species-specific regions present. These results indicate that Enteritidis is the serovar that is closest to being the &#x201C;core&#x201D; example of a <italic>S. enterica</italic> genome, while Typhi is the serovar that is the most divergent. <italic>S.</italic> Enteritidis is the most common cause of enteric <italic>Salmonella</italic> infection, causing upward of one quarter of all infections, and is prevalent in chickens as well as their eggs (<xref ref-type="bibr" rid="B8">Chai et al., 2012</xref>). Conversely, <italic>S.</italic> Typhi is a human adapted serovar, responsible for Typhoid fever, and observed to have undergone genome degradation, rearrangement, and acquisition through horizontal gene-transfer, as it has evolved within its human host (<xref ref-type="bibr" rid="B50">Sabbagh et al., 2010</xref>; <xref ref-type="bibr" rid="B26">Klemm et al., 2016</xref>). It thus appears that genomic change enabling adaptation to a host creates a genomic pool that distinguishes a group from others of the same species. At the same time, genetically similar serovars that maintain a broad host range do not undergo as much selection for genomic change are much harder to distinguish as separate groups, but much easier to identify as members of the subspecies.</p>
</sec>
<sec><title>Core and Pan-genome Comparison</title>
<p>Most phylogentic studies focus on variation within homologs in the core genome to infer evolutionary relationships (<xref ref-type="bibr" rid="B61">Treangen et al., 2014</xref>), as paralogs and horizontally transferred elements confound the evolutionary signal found in genes obtained through vertical descent over time (<xref ref-type="bibr" rid="B16">Gabald&#x00F3;n and Koonin, 2013</xref>). While this approach is undoubtedly useful for long-term evolutionary analyses, when attempting to identify phenotypic linkages between phylogenetic clades, the accessory genome needs to be taken into account, as non-ubiquitous genomic regions allow different groups within the species to occupy and thrive in specific niches (<xref ref-type="bibr" rid="B46">Polz et al., 2013</xref>). Additionally, it has recently been shown that regulatory switching to non-homologous regulatory regions acquired via horizontal gene transfer happens in many bacteria (<xref ref-type="bibr" rid="B42">Oren et al., 2014</xref>). It was further shown that regulatory regions can move without the genes they regulate moving, and that at least 16% of the differences in expression observed within an <italic>E. coli</italic> population were explained by this regulatory switching.</p>
<p>It is therefore prudent to examine both the accessory genome, and not just genes, but non-coding DNA as well, as both have been shown to influence gene expression, and niche specificity. Recent studies have shown that the concordance between a phylogeny based on core genome SNPs and the presence/absence of pan-genome regions is high. For example, in a study examining <italic>E. coli</italic> lineage ST131, the core and accessory genomes showed high concordance, and the combined analyses of both allowed the analyses of the evolution of the <italic>E. coli</italic> lineage at a resolution not possible if only a restricted portion of the genome had been considered (<xref ref-type="bibr" rid="B38">McNally et al., 2016</xref>). The current study shows the same concordant relationship within <italic>S. enterica</italic> between the core and accessory genome, indicating that the accessory genome is not just randomly acquired genomic material, but that selection within specific niches establishes a complement of genes and regulatory elements that enable the survival of the <italic>S. enterica</italic> strains present. It also suggests that to understand why particular clades are more virulent, or possess a particular phenotype, a pan-genomic approach should be used in comparative analyses.</p>
</sec>
</sec>
<sec><title>Conclusion</title>
<p>We examined a quality filtered set of 4893 genomes, the largest pan-genomic study of the <italic>S. enterica</italic> species to date. We identified a pan-genome of 25.3 Mbp, a strict core of 1.5 Mbp present in all genomes, and a conserved core of 3.2 Mbp found in at least 96% of the genomes in this study. In addition we identified 404 species-specific regions, within which a minimum set of two was required to unambiguously identify a genome as being part of the species <italic>S. enterica</italic>. These species-specific regions were found to have functions related to the propagation in and colonization of the host, including the utilization of sialic acid in intestinal mucus, secretion systems for attachment to the host, and the killing of other host microbiota. Within subspecies <italic>enterica</italic>, the species-specific regions were found most frequently in serovar Enteritidis. Each of the six subspecies was found to have genomic regions specific to it; however, the number of subspecies-specific regions appeared to be correlated with the level of sampling of the diversity within the subspecies. No serovar had pan-genome regions that were present in all of its representative genomes and absent in all other serovar genomes; however, each serovar did have genomic regions that were universally present among all constituent members, and statistically predictive of the serovar. <italic>S.</italic> Typhi, which is host-adapted to humans, was found to have the most universal markers predictive of its serovar. The phylogeny based on SNPs within the conserved core genome was found to be highly concordant to that produced by a phylogeny using the presence/absence of the entire pan-genome, and both agreed with phylogenies previously reported for <italic>S. enterica</italic>. Together, the core and accessory genome offered a more complete picture of the diversity within the genomes than either alone. The genomic regions identified in this study that are predictive of the species <italic>S. enterica</italic>, its six subspecies, and the serovar groups within subspecies <italic>enterica</italic>, could be developed into simple and rapid diagnostic tests, with uses ranging from food safety to public health. Additionally, the tools and methods described in this study could be generally applicable as a pan-genomics framework for future population studies, or those looking for genotype/phenotype linkages.</p>
</sec>
<sec><title>Author Contributions</title>
<p>CL: designed the experiments, analyzed the data, and wrote the manuscript. MW: designed the experiments, and wrote the manuscript. VG: designed the experiments, and wrote the manuscript.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>Thanks to Peter Kruczkiewicz of the Public Health Agency of Canada for providing the curated metadata for the EnteroBase genomes used in this study.</p>
</ack>
<sec sec-type="supplementary material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fmicb.2017.01345/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fmicb.2017.01345/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.CSV" id="SM1" mimetype="text/csv" xmlns:xlink="http://www.w3.org/1999/xlink">
</supplementary-material>
<supplementary-material xlink:href="Data_Sheet_2.csv" id="SM2" mimetype="text/csv" xmlns:xlink="http://www.w3.org/1999/xlink">
</supplementary-material>
<supplementary-material xlink:href="Data_Sheet_3.CSV" id="SM3" mimetype="text/csv" xmlns:xlink="http://www.w3.org/1999/xlink">
</supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aanensen</surname> <given-names>D. M.</given-names></name> <name><surname>Feil</surname> <given-names>E. J.</given-names></name> <name><surname>Holden</surname> <given-names>M. T. G.</given-names></name> <name><surname>Dordel</surname> <given-names>J.</given-names></name> <name><surname>Yeats</surname> <given-names>C. A.</given-names></name> <name><surname>Fedosejev</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Whole-genome sequencing for routine pathogen surveillance in public health: a population snapshot of invasive <italic>Staphylococcus aureus</italic> in Europe.</article-title> <source><italic>mBio</italic></source> <volume>7</volume>:<issue>e00444</issue>&#x2013;<issue>16</issue>. <pub-id pub-id-type="doi">10.1128/mBio.00444-16</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Allard</surname> <given-names>M. W.</given-names></name> <name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Strain</surname> <given-names>E.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Keys</surname> <given-names>C. E.</given-names></name> <name><surname>Son</surname> <given-names>I.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>High resolution clustering of <italic>Salmonella enterica</italic> serovar Montevideo strains using a next-generation sequencing approach.</article-title> <source><italic>BMC Genomics</italic></source> <volume>13</volume>:<issue>32</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-13-32</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashton</surname> <given-names>P. M.</given-names></name> <name><surname>Nair</surname> <given-names>S.</given-names></name> <name><surname>Peters</surname> <given-names>T. M.</given-names></name> <name><surname>Bale</surname> <given-names>J. A.</given-names></name> <name><surname>Powell</surname> <given-names>D. G.</given-names></name> <name><surname>Painset</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Identification of <italic>Salmonella</italic> for public health surveillance using whole genome sequencing.</article-title> <source><italic>PeerJ</italic></source> <volume>4</volume>:<issue>e1752</issue>. <pub-id pub-id-type="doi">10.7717/peerj.1752</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Babenko</surname> <given-names>D.</given-names></name> <name><surname>Azizov</surname> <given-names>I.</given-names></name> <name><surname>Toleman</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>wgMLST as a standardized tool for assessing the quality of genome assembly data.</article-title> <source><italic>Int. J. Infect. Dis.</italic></source> <volume>45</volume>:<issue>329</issue>. <pub-id pub-id-type="doi">10.1016/j.ijid.2016.02.714</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergholz</surname> <given-names>T. M.</given-names></name> <name><surname>Moreno Switt</surname> <given-names>A. I.</given-names></name> <name><surname>Wiedmann</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Omics approaches in food safety: Fulfilling the promise?</article-title> <source><italic>Trends Microbiol.</italic></source> <volume>22</volume> <fpage>275</fpage>&#x2013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.1016/j.tim.2014.01.006</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bijlsma</surname> <given-names>J. J.</given-names></name> <name><surname>Groisman</surname> <given-names>E. A.</given-names></name></person-group> (<year>2005</year>). <article-title>The PhoP/PhoQ system controls the intramacrophage type three secretion system of <italic>Salmonella enterica</italic>.</article-title> <source><italic>Mol. Microbiol.</italic></source> <volume>57</volume> <fpage>85</fpage>&#x2013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2958.2005.04668.x</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brunet</surname> <given-names>Y. R.</given-names></name> <name><surname>Khodr</surname> <given-names>A.</given-names></name> <name><surname>Logger</surname> <given-names>L.</given-names></name> <name><surname>Aussel</surname> <given-names>L.</given-names></name> <name><surname>Mignot</surname> <given-names>T.</given-names></name> <name><surname>Rimsky</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>H-NS silencing of the <italic>Salmonella</italic> pathogenicity island 6-encoded type VI secretion system limits <italic>Salmonella enterica</italic> serovar typhimurium interbacterial killing.</article-title> <source><italic>Infect. Immun.</italic></source> <volume>83</volume> <fpage>2738</fpage>&#x2013;<lpage>2750</lpage>. <pub-id pub-id-type="doi">10.1128/IAI.00198-15</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chai</surname> <given-names>S. J.</given-names></name> <name><surname>White</surname> <given-names>P. L.</given-names></name> <name><surname>Lathrop</surname> <given-names>S. L.</given-names></name> <name><surname>Solghan</surname> <given-names>S. M.</given-names></name> <name><surname>Medus</surname> <given-names>C.</given-names></name> <name><surname>McGlinchey</surname> <given-names>B. M.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title><italic>Salmonella enterica</italic> serotype enteritidis: increasing incidence of domestically acquired infections.</article-title> <source><italic>Clin. Infect. Dis.</italic></source> <volume>54(Suppl. 5)</volume>, <fpage>S488</fpage>&#x2013;<lpage>S497</lpage>. <pub-id pub-id-type="doi">10.1093/cid/cis231</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Paoli</surname> <given-names>G. C.</given-names></name> <name><surname>Shi</surname> <given-names>C.</given-names></name> <name><surname>Tu</surname> <given-names>S. I.</given-names></name> <name><surname>Shi</surname> <given-names>X.</given-names></name></person-group> (<year>2010</year>). <article-title>A real-time PCR method for the detection of <italic>Salmonella enterica</italic> from food using a target sequence identified by comparative genomic analysis.</article-title> <source><italic>Int. J. Food Microbiol.</italic></source> <volume>137</volume> <fpage>168</fpage>&#x2013;<lpage>174</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijfoodmicro.2009.12.004</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>H.</given-names></name> <name><surname>Mechanda</surname> <given-names>S.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name></person-group> (<year>1996</year>). <article-title>PCR amplification of the fimA gene sequence of <italic>Salmonella</italic> typhimurium, a specific method for detection of <italic>Salmonella</italic> spp.</article-title> <source><italic>Appl. Environ. Microbiol.</italic></source> <volume>62</volume> <fpage>4303</fpage>&#x2013;<lpage>4308</lpage>.</citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>den Bakker</surname> <given-names>H. C.</given-names></name> <name><surname>Allard</surname> <given-names>M. W.</given-names></name> <name><surname>Bopp</surname> <given-names>D.</given-names></name> <name><surname>Brown</surname> <given-names>E. W.</given-names></name> <name><surname>Fontana</surname> <given-names>J.</given-names></name> <name><surname>Iqbal</surname> <given-names>Z.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Rapid whole-genome sequencing for surveillance of <italic>Salmonella enterica</italic> serovar enteritidis.</article-title> <source><italic>Emerg. Infect. Dis.</italic></source> <volume>20</volume> <fpage>1306</fpage>&#x2013;<lpage>1314</lpage>. <pub-id pub-id-type="doi">10.3201/eid2008.131399</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>X.</given-names></name> <name><surname>den Bakker</surname> <given-names>H. C.</given-names></name> <name><surname>Hendriksen</surname> <given-names>R. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Genomic epidemiology: whole-genome-sequencing-powered surveillance and outbreak investigation of foodborne bacterial pathogens.</article-title> <source><italic>Annu. Rev. Food Sci. Technol.</italic></source> <volume>7</volume> <fpage>353</fpage>&#x2013;<lpage>374</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-food-041715-033259</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Desai</surname> <given-names>P. T.</given-names></name> <name><surname>Porwollik</surname> <given-names>S.</given-names></name> <name><surname>Long</surname> <given-names>F.</given-names></name> <name><surname>Cheng</surname> <given-names>P.</given-names></name> <name><surname>Wollam</surname> <given-names>A.</given-names></name> <name><surname>Bhonagiri-Palsikar</surname> <given-names>V.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Evolutionary genomics of <italic>Salmonella enterica</italic> subspecies.</article-title> <source><italic>mBio</italic></source> <volume>4</volume>:<issue>e00579</issue>&#x2013;<issue>12</issue>. <pub-id pub-id-type="doi">10.1128/mBio.00579-12</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Franz</surname> <given-names>E.</given-names></name> <name><surname>Gras</surname> <given-names>L. M.</given-names></name> <name><surname>Dallman</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Significance of whole genome sequencing for surveillance, source attribution and microbial risk assessment of foodborne pathogens.</article-title> <source><italic>Curr. Opin. Food Sci.</italic></source> <volume>8</volume> <fpage>74</fpage>&#x2013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1016/j.cofs.2016.04.004</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>L.</given-names></name> <name><surname>Niu</surname> <given-names>B.</given-names></name> <name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2012</year>). <article-title>CD-HIT: accelerated for clustering the next-generation sequencing data.</article-title> <source><italic>Bioinformatics</italic></source> <volume>28</volume> <fpage>3150</fpage>&#x2013;<lpage>3152</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts565</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gabald&#x00F3;n</surname> <given-names>T.</given-names></name> <name><surname>Koonin</surname> <given-names>E. V.</given-names></name></person-group> (<year>2013</year>). <article-title>Functional and evolutionary implications of gene orthology.</article-title> <source><italic>Nat. Rev. Genet.</italic></source> <volume>14</volume> <fpage>360</fpage>&#x2013;<lpage>366</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3456</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gal-Mor</surname> <given-names>O.</given-names></name> <name><surname>Boyle</surname> <given-names>E. C.</given-names></name> <name><surname>Grassl</surname> <given-names>G. A.</given-names></name></person-group> (<year>2014</year>). <article-title>Same species, different diseases: how and why typhoidal and non-typhoidal <italic>Salmonella enterica</italic> serovars differ.</article-title> <source><italic>Front. Microbiol.</italic></source> <volume>5</volume>:<issue>391</issue>. <pub-id pub-id-type="doi">10.3389/fmicb.2014.00391</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Golubeva</surname> <given-names>Y. A.</given-names></name> <name><surname>Ellermeier</surname> <given-names>J. R.</given-names></name> <name><surname>Chubiz</surname> <given-names>J. E. C.</given-names></name> <name><surname>Slauch</surname> <given-names>J. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Intestinal long-chain fatty acids act as a direct signal to modulate expression of the <italic>Salmonella</italic> pathogenicity island 1 type III secretion system.</article-title> <source><italic>mBio</italic></source> <volume>7</volume>:<issue>e02170</issue>&#x2013;<issue>15</issue>. <pub-id pub-id-type="doi">10.1128/mBio.02170-15</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Beuchat</surname> <given-names>L. R.</given-names></name> <name><surname>Robert</surname> <given-names>E.</given-names></name></person-group> (<year>2000</year>). <article-title>PCR detection of <italic>Salmonella enterica</italic> serotype montevideo in and on raw tomatoes using primers derived from <italic>hilA</italic>.</article-title> <source><italic>Appl. Environ. Microbiol.</italic></source> <volume>66</volume> <fpage>5248</fpage>&#x2013;<lpage>5252</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.66.12.5248-5252.2000.Updated</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gurevich</surname> <given-names>A.</given-names></name> <name><surname>Saveliev</surname> <given-names>V.</given-names></name> <name><surname>Vyahhi</surname> <given-names>N.</given-names></name> <name><surname>Tesler</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>QUAST: quality assessment tool for genome assemblies.</article-title> <source><italic>Bioinformatics</italic></source> <volume>29</volume> <fpage>1072</fpage>&#x2013;<lpage>1075</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt086</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hayden</surname> <given-names>H. S.</given-names></name> <name><surname>Matamouros</surname> <given-names>S.</given-names></name> <name><surname>Hager</surname> <given-names>K. R.</given-names></name> <name><surname>Brittnacher</surname> <given-names>M. J.</given-names></name> <name><surname>Rohmer</surname> <given-names>L.</given-names></name> <name><surname>Radey</surname> <given-names>M. C.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Genomic analysis of <italic>Salmonella enterica</italic> serovar Typhimurium characterizes strain diversity for recent U.S. salmonellosis cases and identifies mutations linked to loss of fitness under nitrosative and oxidative stress.</article-title> <source><italic>mBio</italic></source> <volume>7</volume>:<issue>e00154</issue>&#x2013;<issue>16</issue>. <pub-id pub-id-type="doi">10.1128/mBio.00154-16</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hensel</surname> <given-names>M.</given-names></name> <name><surname>Shea</surname> <given-names>J. E.</given-names></name> <name><surname>Waterman</surname> <given-names>S. R.</given-names></name> <name><surname>Mundy</surname> <given-names>R.</given-names></name> <name><surname>Nikolaus</surname> <given-names>T.</given-names></name> <name><surname>Banks</surname> <given-names>G.</given-names></name><etal/></person-group> (<year>1998</year>). <article-title>Genes encoding putative effector proteins of the type III secretion system of <italic>Salmonella</italic> pathogenicity island 2 are required for bacterial virulence and proliferation in macrophages.</article-title> <source><italic>Mol. Microbiol.</italic></source> <volume>30</volume> <fpage>163</fpage>&#x2013;<lpage>174</lpage>. <pub-id pub-id-type="doi">10.1046/j.1365-2958.1998.01047.x</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Horvath</surname> <given-names>L.</given-names></name> <name><surname>Kraft</surname> <given-names>M.</given-names></name> <name><surname>Fostiropoulos</surname> <given-names>K.</given-names></name> <name><surname>Falkowski</surname> <given-names>A.</given-names></name> <name><surname>Tarr</surname> <given-names>P. E.</given-names></name></person-group> (<year>2016</year>). <article-title><italic>Salmonella enterica</italic> subspecies <italic>diarizonae</italic> maxillary sinusitis in a snake handler: first report.</article-title> <source><italic>Open Forum Infect. Dis.</italic></source> <volume>3</volume>:<issue>ofw066</issue>. <pub-id pub-id-type="doi">10.1093/ofid/ofw066</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jacobsen</surname> <given-names>A.</given-names></name> <name><surname>Hendriksen</surname> <given-names>R. S.</given-names></name> <name><surname>Aaresturp</surname> <given-names>F. M.</given-names></name> <name><surname>Ussery</surname> <given-names>D. W.</given-names></name> <name><surname>Friis</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>The <italic>Salmonella enterica</italic> Pan-genome.</article-title> <source><italic>Microb. Ecol.</italic></source> <volume>62</volume> <fpage>487</fpage>&#x2013;<lpage>504</lpage>. <pub-id pub-id-type="doi">10.1007/s00248-011-9880-1</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kirk</surname> <given-names>M. D.</given-names></name> <name><surname>Pires</surname> <given-names>S. M.</given-names></name> <name><surname>Black</surname> <given-names>R. E.</given-names></name> <name><surname>Caipo</surname> <given-names>M.</given-names></name> <name><surname>Crump</surname> <given-names>J. A.</given-names></name> <name><surname>Devleesschauwer</surname> <given-names>B.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>World Health Organization estimates of the global and regional disease burden of 22 foodborne bacterial, protozoal, and viral diseases, 2010: a data synthesis.</article-title> <source><italic>PLoS Med.</italic></source> <volume>12</volume>:<issue>e1001921</issue>. <pub-id pub-id-type="doi">10.1371/journal.pmed.1001921</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klemm</surname> <given-names>E. J.</given-names></name> <name><surname>Gkrania-Klotsas</surname> <given-names>E.</given-names></name> <name><surname>Hadfield</surname> <given-names>J.</given-names></name> <name><surname>Forbester</surname> <given-names>J. L.</given-names></name> <name><surname>Harris</surname> <given-names>S. R.</given-names></name> <name><surname>Hale</surname> <given-names>C.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Emergence of host-adapted <italic>Salmonella</italic> Enteritidis through rapid evolution in an immunocompromised host.</article-title> <source><italic>Nat. Microbiol.</italic></source> <volume>1</volume>:<issue>15023</issue>. <pub-id pub-id-type="doi">10.1038/NMICROBIOL.2015.23</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kokkinos</surname> <given-names>P. A.</given-names></name> <name><surname>Ziros</surname> <given-names>P. G.</given-names></name> <name><surname>Bellou</surname> <given-names>M.</given-names></name> <name><surname>Vantarakis</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Loop-Mediated Isothermal Amplification (LAMP) for the detection of <italic>Salmonella</italic> in food.</article-title> <source><italic>Food Anal. Methods</italic></source> <volume>7</volume> <fpage>512</fpage>&#x2013;<lpage>526</lpage>. <pub-id pub-id-type="doi">10.1007/s12161-013-9748-8</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kubori</surname> <given-names>T.</given-names></name> <name><surname>Sukhan</surname> <given-names>A.</given-names></name> <name><surname>Aizawa</surname> <given-names>S. I.</given-names></name> <name><surname>Gal&#x00E1;n</surname> <given-names>J. E.</given-names></name></person-group> (<year>2000</year>). <article-title>Molecular characterization and assembly of the needle complex of the <italic>Salmonella</italic> typhimurium type III protein secretion system.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>97</volume> <fpage>10225</fpage>&#x2013;<lpage>10230</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.170128997</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laing</surname> <given-names>C.</given-names></name> <name><surname>Buchanan</surname> <given-names>C.</given-names></name> <name><surname>Taboada</surname> <given-names>E. N.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Kropinski</surname> <given-names>A.</given-names></name> <name><surname>Villegas</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>Pan-genome sequence analysis using Panseq: an online tool for the rapid analysis of core and accessory genomic regions.</article-title> <source><italic>BMC Bioinformatics</italic></source> <volume>11</volume>:<issue>461</issue>. <pub-id pub-id-type="doi">10.1186/1471-2105-11-461</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leekitcharoenphon</surname> <given-names>P.</given-names></name> <name><surname>Lukjancenko</surname> <given-names>O.</given-names></name> <name><surname>Friis</surname> <given-names>C.</given-names></name> <name><surname>Aarestrup</surname> <given-names>F. M.</given-names></name> <name><surname>Ussery</surname> <given-names>D. W.</given-names></name></person-group> (<year>2012</year>). <article-title>Genomic variation in <italic>Salmonella enterica</italic> core genes for epidemiological typing.</article-title> <source><italic>BMC Genomics</italic></source> <volume>13</volume>:<issue>88</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-13-88</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levine</surname> <given-names>M. M.</given-names></name> <name><surname>Stinear</surname> <given-names>T.</given-names></name> <name><surname>Holt</surname> <given-names>K. E.</given-names></name> <name><surname>Robins-Browne</surname> <given-names>R. M.</given-names></name> <name><surname>Ingle</surname> <given-names>D. J.</given-names></name> <name><surname>Kuzevski</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title><italic>In silico</italic> serotyping of <italic>E. coli</italic> from short read data identifies limited novel O-loci but extensive diversity of O:H serotype combinations within and between pathogenic lineages.</article-title> <source><italic>Microb. Genomics</italic></source> <volume>2</volume>:<issue>e000064</issue>. <pub-id pub-id-type="doi">10.1099/mgen.0.000064</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lupolova</surname> <given-names>N.</given-names></name> <name><surname>Dallman</surname> <given-names>T. J.</given-names></name> <name><surname>Matthews</surname> <given-names>L.</given-names></name> <name><surname>Bono</surname> <given-names>J. L.</given-names></name> <name><surname>Gally</surname> <given-names>D. L.</given-names></name></person-group> (<year>2016</year>). <article-title>Support vector machine applied to predict the zoonotic potential of <italic>E. coli</italic> O157 cattle isolates.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>113</volume> <fpage>11312</fpage>&#x2013;<lpage>11317</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1606567113</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Majowicz</surname> <given-names>S. E.</given-names></name> <name><surname>Musto</surname> <given-names>J.</given-names></name> <name><surname>Scallan</surname> <given-names>E.</given-names></name> <name><surname>Angulo</surname> <given-names>F. J.</given-names></name> <name><surname>Kirk</surname> <given-names>M.</given-names></name> <name><surname>O&#x2019;Brien</surname> <given-names>S. J.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>The global burden of nontyphoidal <italic>Salmonella</italic> gastroenteritis.</article-title> <source><italic>Clin. Infect. Dis.</italic></source> <volume>50</volume> <fpage>882</fpage>&#x2013;<lpage>889</lpage>. <pub-id pub-id-type="doi">10.1086/650733</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malorny</surname> <given-names>B.</given-names></name> <name><surname>Hoorfar</surname> <given-names>J.</given-names></name> <name><surname>Bunge</surname> <given-names>C.</given-names></name> <name><surname>Helmuth</surname> <given-names>R.</given-names></name></person-group> (<year>2003</year>). <article-title>Multicenter validation of the analytical accuracy of <italic>Salmonella</italic> PCR: towards an international standard.</article-title> <source><italic>Appl. Environ. Microbiol.</italic></source> <volume>69</volume> <fpage>290</fpage>&#x2013;<lpage>296</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.69.1.290</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malorny</surname> <given-names>B.</given-names></name> <name><surname>Paccassoni</surname> <given-names>E.</given-names></name> <name><surname>Fach</surname> <given-names>P.</given-names></name> <name><surname>Martin</surname> <given-names>A.</given-names></name> <name><surname>Helmuth</surname> <given-names>R.</given-names></name> <name><surname>Bunge</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <article-title>Diagnostic real-time PCR for detection of <italic>Salmonella</italic> in food.</article-title> <source><italic>Appl. Environ. Microbiol.</italic></source> <volume>70</volume> <fpage>7046</fpage>&#x2013;<lpage>7052</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.70.12.7046</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McDermott</surname> <given-names>P. F.</given-names></name> <name><surname>Tyson</surname> <given-names>G. H.</given-names></name> <name><surname>Kabera</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Folster</surname> <given-names>J. P.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>The use of whole genome sequencing for detecting antimicrobial resistance in nontyphoidal <italic>Salmonella</italic>.</article-title> <source><italic>Antimicrob. Agents Chemother.</italic></source> <volume>60</volume> <fpage>5515</fpage>&#x2013;<lpage>5520</lpage>. <pub-id pub-id-type="doi">10.1128/AAC.01030-16</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McDonald</surname> <given-names>N. D.</given-names></name> <name><surname>Lubin</surname> <given-names>J. B.</given-names></name> <name><surname>Chowdhury</surname> <given-names>N.</given-names></name> <name><surname>Boyd</surname> <given-names>E. F.</given-names></name></person-group> (<year>2016</year>). <article-title>Host-derived sialic acids are an important nutrient source required for optimal bacterial fitness in vivo.</article-title> <source><italic>mBio</italic></source> <volume>7</volume>:<issue>e02237</issue>&#x2013;<issue>15</issue>. <pub-id pub-id-type="doi">10.1128/mBio.02237-15</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McNally</surname> <given-names>A.</given-names></name> <name><surname>Oren</surname> <given-names>Y.</given-names></name> <name><surname>Kelly</surname> <given-names>D.</given-names></name> <name><surname>Pascoe</surname> <given-names>B.</given-names></name> <name><surname>Dunn</surname> <given-names>S.</given-names></name> <name><surname>Sreecharan</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Combined analysis of variation in core, accessory and regulatory genome regions provides a super-resolution view into the evolution of bacterial populations.</article-title> <source><italic>PLoS Genet.</italic></source> <volume>12</volume>:<issue>e1006280</issue>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1006280</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moreno Switt</surname> <given-names>A. I.</given-names></name> <name><surname>Orsi</surname> <given-names>R. H.</given-names></name> <name><surname>den Bakker</surname> <given-names>H. C.</given-names></name> <name><surname>Vongkamjan</surname> <given-names>K.</given-names></name> <name><surname>Altier</surname> <given-names>C.</given-names></name> <name><surname>Wiedmann</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Genomic characterization provides new insight into <italic>Salmonella</italic> phage diversity.</article-title> <source><italic>BMC Genomics</italic></source> <volume>14</volume>:<issue>481</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-14-481</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ng</surname> <given-names>K. M.</given-names></name> <name><surname>Ferreyra</surname> <given-names>J. A.</given-names></name> <name><surname>Higginbottom</surname> <given-names>S. K.</given-names></name> <name><surname>Lynch</surname> <given-names>J. B.</given-names></name> <name><surname>Kashyap</surname> <given-names>P. C.</given-names></name> <name><surname>Gopinath</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Microbiota-liberated host sugars facilitate post-antibiotic expansion of enteric pathogens.</article-title> <source><italic>Nature</italic></source> <volume>502</volume> <fpage>96</fpage>&#x2013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1038/nature12503</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Opijnen</surname> <given-names>T. V.</given-names></name> <name><surname>Camilli</surname> <given-names>A.</given-names></name> <name><surname>Opijnen</surname> <given-names>T. V.</given-names></name> <name><surname>Camilli</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>A fine scale phenotype-genotype virulence map of a bacterial pathogen.</article-title> <source><italic>Genome Res.</italic></source> <volume>22</volume> <fpage>2541</fpage>&#x2013;<lpage>2551</lpage>. <pub-id pub-id-type="doi">10.1101/gr.137430.112</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oren</surname> <given-names>Y.</given-names></name> <name><surname>Smith</surname> <given-names>M. B.</given-names></name> <name><surname>Johns</surname> <given-names>N. I.</given-names></name> <name><surname>Kaplan Zeevi</surname> <given-names>M.</given-names></name> <name><surname>Biran</surname> <given-names>D.</given-names></name> <name><surname>Ron</surname> <given-names>E. Z.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Transfer of noncoding DNA drives regulatory rewiring in bacteria.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>111</volume> <fpage>16112</fpage>&#x2013;<lpage>16117</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1413272111</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Page</surname> <given-names>A. J.</given-names></name> <name><surname>Cummins</surname> <given-names>C. A.</given-names></name> <name><surname>Hunt</surname> <given-names>M.</given-names></name> <name><surname>Wong</surname> <given-names>V. K.</given-names></name> <name><surname>Reuter</surname> <given-names>S.</given-names></name> <name><surname>Holden</surname> <given-names>M. T. G.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Roary: rapid large-scale prokaryote pan genome analysis.</article-title> <source><italic>Bioinformatics</italic></source> <volume>31</volume> <fpage>3691</fpage>&#x2013;<lpage>3693</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv421</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Penad&#x00E9;s</surname> <given-names>J. R.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Quiles-Puchalt</surname> <given-names>N.</given-names></name> <name><surname>Carpena</surname> <given-names>N.</given-names></name> <name><surname>Novick</surname> <given-names>R. P.</given-names></name></person-group> (<year>2015</year>). <article-title>Bacteriophage-mediated spread of bacterial virulence genes.</article-title> <source><italic>Curr. Opin. Microbiol.</italic></source> <volume>23</volume> <fpage>171</fpage>&#x2013;<lpage>178</lpage>. <pub-id pub-id-type="doi">10.1016/j.mib.2014.11.019</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pettengill</surname> <given-names>J. B.</given-names></name> <name><surname>Pightling</surname> <given-names>A. W.</given-names></name> <name><surname>Baugher</surname> <given-names>J. D.</given-names></name> <name><surname>Rand</surname> <given-names>H.</given-names></name> <name><surname>Strain</surname> <given-names>E.</given-names></name></person-group> (<year>2016</year>). <article-title>Real-time pathogen detection in the era of whole-genome sequencing and big data: comparison of k-mer and site-based methods for inferring the genetic distances among tens of thousands of <italic>Salmonella</italic> samples.</article-title> <source><italic>PLoS ONE</italic></source> <volume>11</volume>:<issue>e0166162</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0166162</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Polz</surname> <given-names>M. F.</given-names></name> <name><surname>Alm</surname> <given-names>E. J.</given-names></name> <name><surname>Hanage</surname> <given-names>W. P.</given-names></name></person-group> (<year>2013</year>). <article-title>Horizontal gene transfer and the evolution of bacterial and archaeal population structure.</article-title> <source><italic>Trends Genet.</italic></source> <volume>29</volume> <fpage>170</fpage>&#x2013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1016/j.tig.2012.12.006</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Postollec</surname> <given-names>F.</given-names></name> <name><surname>Falentin</surname> <given-names>H.</given-names></name> <name><surname>Pavan</surname> <given-names>S.</given-names></name> <name><surname>Combrisson</surname> <given-names>J.</given-names></name> <name><surname>Sohier</surname> <given-names>D.</given-names></name></person-group> (<year>2011</year>). <article-title>Recent advances in quantitative PCR (qPCR) applications in food microbiology.</article-title> <source><italic>Food Microbiol.</italic></source> <volume>28</volume> <fpage>848</fpage>&#x2013;<lpage>861</lpage>. <pub-id pub-id-type="doi">10.1016/j.fm.2011.02.008</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><collab>R Core Team</collab> (<year>2016</year>). <source><italic>R: A Language and Environment for Statistical Computing</italic>.</source> <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Foundation for Statistical Computing</publisher-name>.</citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rahman</surname> <given-names>A.</given-names></name> <name><surname>Pachter</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>CGAL: computing genome assembly likelihoods.</article-title> <source><italic>Genome Biol.</italic></source> <volume>14</volume>:<issue>R8</issue>. <pub-id pub-id-type="doi">10.1186/gb-2013-14-1-r8</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sabbagh</surname> <given-names>S. C.</given-names></name> <name><surname>Forest</surname> <given-names>C. G.</given-names></name> <name><surname>Lepage</surname> <given-names>C.</given-names></name> <name><surname>Leclerc</surname> <given-names>J. M.</given-names></name> <name><surname>Daigle</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>So similar, yet so different: uncovering distinctive features in the genomes of <italic>Salmonella enterica</italic> serovars Typhimurium and Typhi.</article-title> <source><italic>FEMS Microbiol. Lett.</italic></source> <volume>305</volume> <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6968.2010.01904.x</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sakarya</surname> <given-names>S.</given-names></name> <name><surname>G&#x00F6;kt&#x00FC;rk</surname> <given-names>C.</given-names></name> <name><surname>&#x00D6;zt&#x00FC;rk</surname> <given-names>T.</given-names></name> <name><surname>Ertugrul</surname> <given-names>M. B.</given-names></name></person-group> (<year>2010</year>). <article-title>Sialic acid is required for nonspecific adherence of <italic>Salmonella enterica</italic> ssp. <italic>enterica</italic> serovar Typhi on Caco-2 cells.</article-title> <source><italic>FEMS Immunol. Med. Microbiol.</italic></source> <volume>58</volume> <fpage>330</fpage>&#x2013;<lpage>335</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-695X.2010.00650.x</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sana</surname> <given-names>T. G.</given-names></name> <name><surname>Flaugnatti</surname> <given-names>N.</given-names></name> <name><surname>Lugo</surname> <given-names>K. A.</given-names></name> <name><surname>Lam</surname> <given-names>L. H.</given-names></name> <name><surname>Jacobson</surname> <given-names>A.</given-names></name> <name><surname>Baylot</surname> <given-names>V.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title><italic>Salmonella</italic> Typhimurium utilizes a T6SS-mediated antibacterial weapon to establish in the host gut.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>113</volume> <fpage>E5044</fpage>&#x2013;<lpage>E5051</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1608858113</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scallan</surname> <given-names>E.</given-names></name> <name><surname>Hoekstra</surname> <given-names>R. M.</given-names></name> <name><surname>Mahon</surname> <given-names>B. E.</given-names></name> <name><surname>Jones</surname> <given-names>T. F.</given-names></name> <name><surname>Griffin</surname> <given-names>P. M.</given-names></name></person-group> (<year>2015</year>). <article-title>An assessment of the human health impact of seven leading foodborne pathogens in the United States using disability adjusted life years.</article-title> <source><italic>Epidemiol. Infect.</italic></source> <volume>143</volume> <fpage>2795</fpage>&#x2013;<lpage>2804</lpage>. <pub-id pub-id-type="doi">10.1017/S0950268814003185</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scharff</surname> <given-names>R. L.</given-names></name> <name><surname>Besser</surname> <given-names>J.</given-names></name> <name><surname>Sharp</surname> <given-names>D. J.</given-names></name> <name><surname>Jones</surname> <given-names>T. F.</given-names></name> <name><surname>Peter</surname> <given-names>G. S.</given-names></name> <name><surname>Hedberg</surname> <given-names>C. W.</given-names></name></person-group> (<year>2016</year>). <article-title>An economic evaluation of PulseNet: a network for foodborne disease surveillance.</article-title> <source><italic>Am. J. Prev. Med.</italic></source> <volume>50</volume> <fpage>S66</fpage>&#x2013;<lpage>S73</lpage>. <pub-id pub-id-type="doi">10.1016/j.amepre.2015.09.018</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schroter</surname> <given-names>M.</given-names></name> <name><surname>Roggentin</surname> <given-names>P.</given-names></name> <name><surname>Hofmann</surname> <given-names>J.</given-names></name> <name><surname>Speicher</surname> <given-names>A.</given-names></name> <name><surname>Laufs</surname> <given-names>R.</given-names></name> <name><surname>Mack</surname> <given-names>D.</given-names></name></person-group> (<year>2004</year>). <article-title>Pet snakes as a reservoir for <italic>Salmonella enterica</italic> subsp. <italic>diarizonae</italic> (serogroup IIIb): a prospective study.</article-title> <source><italic>Appl. Environ. Microbiol.</italic></source> <volume>70</volume> <fpage>613</fpage>&#x2013;<lpage>615</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.70.1.613-615.2004</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Severi</surname> <given-names>E.</given-names></name> <name><surname>M&#x00FC;ller</surname> <given-names>A.</given-names></name> <name><surname>Potts</surname> <given-names>J. R.</given-names></name> <name><surname>Leech</surname> <given-names>A.</given-names></name> <name><surname>Williamson</surname> <given-names>D.</given-names></name> <name><surname>Wilson</surname> <given-names>K. S.</given-names></name><etal/></person-group> (<year>2008</year>). <article-title>Sialic acid mutarotation is catalyzed by the <italic>Escherichia coli</italic> &#x03B2;-propeller protein YjhT.</article-title> <source><italic>J. Biol. Chem.</italic></source> <volume>283</volume> <fpage>4841</fpage>&#x2013;<lpage>4849</lpage>. <pub-id pub-id-type="doi">10.1074/jbc.M707822200</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sheppard</surname> <given-names>S. K.</given-names></name> <name><surname>Jolley</surname> <given-names>K. A.</given-names></name> <name><surname>Maiden</surname> <given-names>M. C. J.</given-names></name></person-group> (<year>2012</year>). <article-title>A gene-by-gene approach to bacterial population genomics: whole genome MLST of <italic>Campylobacter</italic>.</article-title> <source><italic>Genes</italic></source> <volume>3</volume> <fpage>261</fpage>&#x2013;<lpage>277</lpage>. <pub-id pub-id-type="doi">10.3390/genes3020261</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stamatakis</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies.</article-title> <source><italic>Bioinformatics</italic></source> <volume>30</volume> <fpage>1312</fpage>&#x2013;<lpage>1313</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu033</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>H.</given-names></name> <name><surname>Kamanova</surname> <given-names>J.</given-names></name> <name><surname>Lara-Tejero</surname> <given-names>M.</given-names></name> <name><surname>Gal&#x00E1;n</surname> <given-names>J. E.</given-names></name></person-group> (<year>2016</year>). <article-title>A Family of <italic>Salmonella</italic> type III secretion effector proteins selectively targets the NF-&#x03BA;B signaling pathway to preserve host homeostasis.</article-title> <source><italic>PLoS Pathog.</italic></source> <volume>12</volume>:<issue>e1005484</issue>. <pub-id pub-id-type="doi">10.1371/journal.ppat.1005484</pub-id></citation></ref>
<ref id="B60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tettelin</surname> <given-names>H.</given-names></name> <name><surname>Masignani</surname> <given-names>V.</given-names></name> <name><surname>Cieslewicz</surname> <given-names>M. J.</given-names></name> <name><surname>Donati</surname> <given-names>C.</given-names></name> <name><surname>Medini</surname> <given-names>D.</given-names></name> <name><surname>Ward</surname> <given-names>N. L.</given-names></name><etal/></person-group> (<year>2005</year>). <article-title>Genome analysis of multiple pathogenic isolates of <italic>Streptococcus agalactiae</italic>: implications for the microbial &#x201C;pan-genome&#x201D;.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>102</volume> <fpage>13950</fpage>&#x2013;<lpage>13955</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0506758102</pub-id></citation></ref>
<ref id="B61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Treangen</surname> <given-names>T. J.</given-names></name> <name><surname>Ondov</surname> <given-names>B. D.</given-names></name> <name><surname>Koren</surname> <given-names>S.</given-names></name> <name><surname>Phillippy</surname> <given-names>A. M.</given-names></name></person-group> (<year>2014</year>). <article-title>The Harvest suite for rapid core-genome alignment and visualization of thousands of intraspecific microbial genomes.</article-title> <source><italic>Genome Biol.</italic></source> <volume>15</volume>:<issue>524</issue>. <pub-id pub-id-type="doi">10.1186/s13059-014-0524-x</pub-id></citation></ref>
<ref id="B62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tyson</surname> <given-names>G. H.</given-names></name> <name><surname>McDermott</surname> <given-names>P. F.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Tadesse</surname> <given-names>D. A.</given-names></name> <name><surname>Mukherjee</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>WGS accurately predicts antimicrobial resistance in <italic>Escherichia coli</italic>.</article-title> <source><italic>J. Antimicrob. Chemother.</italic></source> <volume>70</volume> <fpage>2763</fpage>&#x2013;<lpage>2769</lpage>. <pub-id pub-id-type="doi">10.1093/jac/dkv186</pub-id></citation></ref>
<ref id="B63"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Waryah</surname> <given-names>C. B.</given-names></name> <name><surname>Gogoi-Tiwari</surname> <given-names>J.</given-names></name> <name><surname>Wells</surname> <given-names>K.</given-names></name> <name><surname>Eto</surname> <given-names>K. Y.</given-names></name> <name><surname>Masoumi</surname> <given-names>E.</given-names></name> <name><surname>Costantino</surname> <given-names>P.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Diversity of virulence factors associated with West Australian methicillin-sensitive <italic>Staphylococcus aureus</italic> isolates of human origin.</article-title> <source><italic>BioMed Res. Int.</italic></source> <volume>2016</volume>:<issue>8651918</issue>. <pub-id pub-id-type="doi">10.1155/2016/8651918</pub-id></citation></ref>
<ref id="B64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Whiteside</surname> <given-names>M. D.</given-names></name> <name><surname>Laing</surname> <given-names>C. R.</given-names></name> <name><surname>Manji</surname> <given-names>A.</given-names></name> <name><surname>Kruczkiewicz</surname> <given-names>P.</given-names></name> <name><surname>Taboada</surname> <given-names>E. N.</given-names></name> <name><surname>Gannon</surname> <given-names>V. P. J.</given-names></name></person-group> (<year>2016</year>). <article-title>SuperPhy: predictive genomics for the bacterial pathogen <italic>Escherichia coli</italic>.</article-title> <source><italic>BMC Microbiol.</italic></source> <volume>16</volume>:<issue>65</issue>. <pub-id pub-id-type="doi">10.1186/s12866-016-0680-0</pub-id></citation></ref>
<ref id="B65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoshida</surname> <given-names>C.</given-names></name> <name><surname>Gurnik</surname> <given-names>S.</given-names></name> <name><surname>Ahmad</surname> <given-names>A.</given-names></name> <name><surname>Blimkie</surname> <given-names>T.</given-names></name> <name><surname>Murphy</surname> <given-names>S. A.</given-names></name> <name><surname>Kropinski</surname> <given-names>A. M.</given-names></name><etal/></person-group> (<year>2016a</year>). <article-title>Evaluation of molecular methods for identification of <italic>Salmonella</italic> serovars.</article-title> <source><italic>J. Clin. Microbiol.</italic></source> <volume>54</volume> <fpage>1992</fpage>&#x2013;<lpage>1998</lpage>. <pub-id pub-id-type="doi">10.1128/JCM.00262-16</pub-id></citation></ref>
<ref id="B66"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoshida</surname> <given-names>C.</given-names></name> <name><surname>Kruczkiewicz</surname> <given-names>P.</given-names></name> <name><surname>Laing</surname> <given-names>C. R.</given-names></name> <name><surname>Lingohr</surname> <given-names>E. J.</given-names></name> <name><surname>Gannon</surname> <given-names>V. P. J.</given-names></name> <name><surname>Nash</surname> <given-names>J. H. E.</given-names></name><etal/></person-group> (<year>2016b</year>). <article-title>The <italic>Salmonella in silico</italic> typing resource (SISTR): an open web-accessible tool for rapidly typing and subtyping draft <italic>Salmonella</italic> genome assemblies.</article-title> <source><italic>PLoS ONE</italic></source> <volume>11</volume>:<issue>e0147101</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0147101</pub-id></citation></ref>
<ref id="B67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoshida</surname> <given-names>C.</given-names></name> <name><surname>Lingohr</surname> <given-names>E. J.</given-names></name> <name><surname>Trognitz</surname> <given-names>F.</given-names></name> <name><surname>MacLaren</surname> <given-names>N.</given-names></name> <name><surname>Rosano</surname> <given-names>A.</given-names></name> <name><surname>Murphy</surname> <given-names>S. A.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Multi-laboratory evaluation of the rapid genoserotyping array (SGSA) for the identification of <italic>Salmonella</italic> serovars.</article-title> <source><italic>Diagn. Microbiol. Infect. Dis.</italic></source> <volume>80</volume> <fpage>185</fpage>&#x2013;<lpage>190</lpage>. <pub-id pub-id-type="doi">10.1016/j.diagmicrobio.2014.08.006</pub-id></citation></ref>
<ref id="B68"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>G.</given-names></name> <name><surname>Smith</surname> <given-names>D. K.</given-names></name> <name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Guan</surname> <given-names>Y.</given-names></name> <name><surname>Lam</surname> <given-names>T. T.-Y.</given-names></name></person-group> (<year>2016</year>). <article-title><sc>GGTREE</sc>: an <sc>R</sc> package for visualization and annotation of phylogenetic trees with their covariates and other associated data.</article-title> <source><italic>Methods Ecol. Evol.</italic></source> <volume>8</volume> <fpage>28</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1111/2041-210X.12628</pub-id></citation></ref>
<ref id="B69"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>S.</given-names></name> <name><surname>Tyson</surname> <given-names>G. H.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Mukherjee</surname> <given-names>S.</given-names></name> <name><surname>Young</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Whole-genome sequencing analysis accurately predicts antimicrobial resistance phenotypes in <italic>Campylobacter</italic> spp.</article-title> <source><italic>Appl. Environ. Microbiol.</italic></source> <volume>82</volume> <fpage>459</fpage>&#x2013;<lpage>466</lpage>. <pub-id pub-id-type="doi">10.1128/AEM.02873-15</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn01"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="https://github.com/chadlaing/feht">https://github.com/chadlaing/feht</ext-link></p></fn>
<fn id="fn02"><label>2</label><p><ext-link ext-link-type="uri" xlink:href="https://enterobase.warwick.ac.uk/species/index/senterica">https://enterobase.warwick.ac.uk/species/index/senterica</ext-link></p></fn>
</fn-group>
</back>
</article>