<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2017.00021</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Robust Inference of Genetic Exchange Communities from Microbial Genomes Using TF-IDF</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Cong</surname> <given-names>Yingnan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Chan</surname> <given-names>Yao-ban</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/395555/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Phillips</surname> <given-names>Charles A.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Langston</surname> <given-names>Michael A.</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Ragan</surname> <given-names>Mark A.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/264928/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute for Molecular Bioscience and ARC Centre of Excellence in Bioinformatics, University of Queensland, St Lucia</institution> <country>QLD, Australia</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Mathematics and Statistics, University of Melbourne, Parkville</institution> <country>VIC, Australia</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Electrical Engineering and Computer Science, University of Tennessee, Knoxville</institution> <country>TN, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Johann Peter Gogarten, University of Connecticut, USA</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Luis Delaye, CINVESTAV, Mexico; Jonathan Badger, National Cancer Institute, USA; Gregory Fournier, Massachusetts Institute of Technology, USA</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>Mark A. Ragan, <email>m.ragan@uq.edu.au</email></italic></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Evolutionary and Genomic Microbiology, a section of the journal Frontiers in Microbiology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>01</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>21</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>10</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>01</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2017 Cong, Chan, Phillips, Langston and Ragan.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Cong, Chan, Phillips, Langston and Ragan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Bacteria and archaea can exchange genetic material across lineages through processes of lateral genetic transfer (LGT). Collectively, these exchange relationships can be modeled as a network and analyzed using concepts from graph theory. In particular, densely connected regions within an LGT network have been defined as genetic exchange communities (GECs). However, it has been problematic to construct networks in which edges solely represent LGT. Here we apply term frequency-inverse document frequency (TF-IDF), an alignment-free method originating from document analysis, to infer regions of lateral origin in bacterial genomes. We examine four empirical datasets of different size (number of genomes) and phyletic breadth, varying a key parameter (word length <italic>k</italic>) within bounds established in previous work. We map the inferred lateral regions to genes in recipient genomes, and construct networks in which the nodes are groups of genomes, and the edges natively represent LGT. We then extract maximum and maximal cliques (i.e., GECs) from these graphs, and identify nodes that belong to GECs across a wide range of <italic>k</italic>. Most surviving lateral transfer has happened within these GECs. Using Gene Ontology enrichment tests we demonstrate that biological processes associated with metabolism, regulation and transport are often over-represented among the genes affected by LGT within these communities. These enrichments are largely robust to change of <italic>k</italic>.</p>
</abstract>
<kwd-group>
<kwd>TF-IDF</kwd>
<kwd>lateral genetic transfer</kwd>
<kwd>horizontal genetic transfer</kwd>
<kwd>microbial genomes</kwd>
<kwd>genetic exchange community</kwd>
<kwd>lateral genetic transfer network</kwd>
<kwd>clique analysis</kwd>
</kwd-group>
<contract-num rid="cn001">220020272</contract-num>
<contract-sponsor id="cn001">James S. McDonnell Foundation<named-content content-type="fundref-id">10.13039/100000913</named-content></contract-sponsor>
<counts>
<fig-count count="5"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="59"/>
<page-count count="11"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>Bacteria and archaea (BA) comprise much of the planet&#x2019;s biodiversity. Although individually inconspicuous, communities of these organisms are responsible for key biological and geochemical processes including nitrogen fixation, aerobic and anaerobic digestion of biomass, and oxidative dissolution of minerals. Bacteria also cause a range of diseases in plants, animals, and humans. Since 1996, genome-sequencing technologies have been applied initially to study bacterial pathogenesis, and more recently to understand environmental processes and explore biodiversity. Genome sequences are publicly available for more than 30,000 BA, and large international projects are underway to sequence many thousands more.</p>
<p>Arguably, the two most-notable discoveries from the first two decades of microbial genomics have been the extent of strain-to-strain variation in gene content (<xref ref-type="bibr" rid="B57">Tettelin et al., 2005</xref>; <xref ref-type="bibr" rid="B51">Segerman, 2012</xref>; <xref ref-type="bibr" rid="B21">Croucher et al., 2014</xref>), and the prevalence of lateral genetic transfer (LGT). It has long been known that bacteria can take up genetic material from their surroundings, incorporate it into their main genome (or maintain it on extrachromosomal elements) and transmit it to subsequent generations. More than 35 years ago, unexpected patterns of gene presence among bacterial taxa and anomalous topologies of phylogenetic trees inferred for bacterial proteins were attributed, somewhat controversially, to LGT (<xref ref-type="bibr" rid="B2">Ambler et al., 1979a</xref>,<xref ref-type="bibr" rid="B3">b</xref>; <xref ref-type="bibr" rid="B24">Dickerson, 1980</xref>; <xref ref-type="bibr" rid="B58">Woese et al., 1980</xref>). In the last 10&#x2013;15 years, large-scale analysis has revealed the surprising extent of LGT among BA, with many estimates indicating that 10&#x2013;40% of genes may have a relatively recent lateral origin; for details see the review by <xref ref-type="bibr" rid="B50">Ragan and Beiko (2009)</xref>. Thus while all organisms transmit genetic information vertically from parent to offspring, BA simultaneously operate an orthogonal genetics that links important components of their genomes with viruses, phage, plasmids and free environmental DNA in a vast web (<xref ref-type="bibr" rid="B25">Doolittle, 1999</xref>; <xref ref-type="bibr" rid="B13">Bryant and Moulton, 2004</xref>; <xref ref-type="bibr" rid="B8">Beiko et al., 2005</xref>; <xref ref-type="bibr" rid="B41">Kunin et al., 2005</xref>; <xref ref-type="bibr" rid="B22">Dagan et al., 2008</xref>; <xref ref-type="bibr" rid="B23">Dagan and Martin, 2009</xref>; <xref ref-type="bibr" rid="B36">Halary et al., 2010</xref>; <xref ref-type="bibr" rid="B46">Puigb&#x00F2; et al., 2010</xref>; <xref ref-type="bibr" rid="B6">Bapteste et al., 2013</xref>; <xref ref-type="bibr" rid="B40">Koonin, 2015</xref>).</p>
<p>We and others (<xref ref-type="bibr" rid="B8">Beiko et al., 2005</xref>; <xref ref-type="bibr" rid="B22">Dagan et al., 2008</xref>; <xref ref-type="bibr" rid="B23">Dagan and Martin, 2009</xref>; <xref ref-type="bibr" rid="B45">Popa et al., 2011</xref>) have sought to model this web of genetic relationships as a graph in which vertices represent observed entities that carry DNA (genomes, and in some applications also plasmids and phage), and edges represent the inferred transmission of genetic material between them. However, resolving the lateral signal turns out to be unexpectedly tricky. Two genomes that have descended only recently from a common ancestor are unlikely to differ greatly in genome sequence or gene content, and if they are accorded individual vertices, the similarity between them will arise almost entirely from vertical signal. To the extent that our graph is intended to help us understand patterns of LGT, it makes sense to combine such genomes into a single vertex (node). As genomes diversify through time, it becomes increasingly desirable to represent them as separate vertices, because doing so potentially increases the resolution at which LGT can be studied; but pairwise edges represent a mixture of vertical and lateral signal. Moreover, older LGT (more-basal in the tree of vertical signal) becomes established in lineages and begins to be allocated among present-day genomes in hierarchical patterns that reinforce local vertical signal (<xref ref-type="bibr" rid="B31">Gogarten et al., 2002</xref>; <xref ref-type="bibr" rid="B32">Gogarten and Townsend, 2005</xref>). Thus by flattening the temporal (historical) dimension into the plane of the (present-day) graph, we hide sequence diversity in the vertices and admix vertical and lateral signal in the edges. Although an optimal balance (or multiple locally optimal balances across the tree) can be sought, these issues remain.</p>
<p>Until now, the nature of the edges has received the most attention. LGT detection methods can be classified into two general types: surrogate and phylogenetic (<xref ref-type="bibr" rid="B48">Ragan, 2001a</xref>,<xref ref-type="bibr" rid="B49">b</xref>; <xref ref-type="bibr" rid="B32">Gogarten and Townsend, 2005</xref>). The former include methods based on all-versus-all sequence comparison (<xref ref-type="bibr" rid="B5">Bansal et al., 1998</xref>; <xref ref-type="bibr" rid="B42">Lima-Mendez et al., 2008</xref>; <xref ref-type="bibr" rid="B29">Fondi and Fani, 2010</xref>; <xref ref-type="bibr" rid="B36">Halary et al., 2010</xref>) or reciprocal best matches (<xref ref-type="bibr" rid="B56">Tatusov et al., 1997</xref>; <xref ref-type="bibr" rid="B12">Bork et al., 1998</xref>) of genes or proteins. Some additional filter must then be applied to distinguish matches that are unexpectedly strong after correction for shared vertical relationship (and perhaps other factors, e.g., functional constraints), and therefore candidates for LGT. This filter might involve a more-stringent match threshold (<xref ref-type="bibr" rid="B36">Halary et al., 2010</xref>) and/or subtracting edges present in a trusted reference tree (<xref ref-type="bibr" rid="B22">Dagan et al., 2008</xref>; <xref ref-type="bibr" rid="B23">Dagan and Martin, 2009</xref>). A converse strategy was employed by <xref ref-type="bibr" rid="B18">Clarke et al. (2002)</xref>, who were interested only in the vertical component. Alternatively in the phylogenetic approach, a test tree (inferred for a putatively orthologous gene or protein family) is compared with a reference (genome or organismal) tree, and instances of topological incongruence that meet a statistical support criterion are considered <italic>prima facie</italic> cases of LGT (<xref ref-type="bibr" rid="B33">Goldman et al., 2000</xref>; <xref ref-type="bibr" rid="B8">Beiko et al., 2005</xref>; <xref ref-type="bibr" rid="B59">Zhaxybayeva et al., 2006</xref>; <xref ref-type="bibr" rid="B9">Beiko and Ragan, 2008</xref>). Even so, reconstructing the pathway of inferred LGT as shortest edit paths is computationally hard and may not yield a unique solution, or any solution at all (<xref ref-type="bibr" rid="B7">Beiko and Hamilton, 2006</xref>). <xref ref-type="bibr" rid="B45">Popa et al. (2011)</xref> employ a hybrid approach in which only genes assessed as having regions of anomalous G+C content are input into phylogenetic discordance analysis.</p>
<p>Several objections have been raised to these approaches, both individually and collectively. We have repeatedly argued that as genes are not the actual units of LGT, gene families should not be the primary units of analysis (<xref ref-type="bibr" rid="B14">Chan et al., 2009a</xref>,<xref ref-type="bibr" rid="B16">b</xref>). <xref ref-type="bibr" rid="B27">Doolittle and Bapteste (2007)</xref> and <xref ref-type="bibr" rid="B26">Doolittle (2009)</xref> have argued that by using a reference tree external to the analysis, we impose a higher standard of evidence on rejecting the reference topology (and thereby inferring LGT) than on accepting (or failing to reject) it, thereby according the vertical paradigm a methodologically unfair and theoretically unjustified advantage; for a conflicting opinion see <xref ref-type="bibr" rid="B43">O&#x2019;Malley and Koonin (2011)</xref>. A way is needed to infer LGT directly, positively and fairly in large genome-scale datasets.</p>
<p>Recently we (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>,<xref ref-type="bibr" rid="B20">b</xref>) introduced term frequency-inverse document frequency (TF-IDF) as an accurate, scalable approach to infer LGT among microbial genomes. Using TF-IDF, edges represent only lateral signal and can be inferred directly from whole genomes without first parsing them into individual genes. These edges are directional: transfers are inferred from a group of donor genomes to a single recipient genome. No comparison with an external topology is required, although inference quality may be improved if the group structure reflects phylogeny (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>).</p>
<p>Direct access to edges that represent only the lateral component of genetic relationships greatly simplifies the interpretation of such graphs: they are natively LGT networks. <xref ref-type="bibr" rid="B54">Skippington and Ragan (2011)</xref> defined a genetic exchange community (GEC) as a densely connected region of an LGT network. Recognizing the limitations of then-existing methods and data, these authors operationally defined a GEC as &#x201C;a set of entities, each of which has over time both donated genetic material to, and received genetic material from, every other entity in that GEC, via a path of lateral transfer.&#x201D; These GECs do not exist <italic>a priori</italic> in nature, but rather are &#x201C;actively fashioned (and continually refashioned) by the complex ongoing interplay among habitats, donors, vectors, recipients, mechanisms, sequences, population structures and selection&#x201D; (<xref ref-type="bibr" rid="B54">Skippington and Ragan, 2011</xref>). Biological problems that could be modeled as involving dense edge sets in LGT graphs include the number, size, geospatial extent, taxonomic or habitat diversity of GECs in the microbial biosphere, and the role of vectors in mediating the exchange of pathogenicity, virulence or resistance factors among pathogens, primary hosts and secondary hosts (<xref ref-type="bibr" rid="B36">Halary et al., 2010</xref>; <xref ref-type="bibr" rid="B45">Popa et al., 2011</xref>; <xref ref-type="bibr" rid="B54">Skippington and Ragan, 2011</xref>, <xref ref-type="bibr" rid="B55">2012</xref>).</p>
<p><xref ref-type="bibr" rid="B54">Skippington and Ragan (2011)</xref> further proposed that dense regions in LGT graphs might be described using concepts from graph theory, including cliques (complete subgraphs, i.e., groups of nodes that are all connected directly to each other), paracliques (cliques missing a few edges: <xref ref-type="bibr" rid="B17">Chesler and Langston, 2007</xref>; <xref ref-type="bibr" rid="B35">Hagan et al., 2016</xref>), other forms of near-cliques, or looser structures such as transitively closed sets, cycles, paths or walks. They were not, however, in a position to recommend one of these notions over the others. Our previous results make it clear that edges, hence dense edge sets in LGT graphs and their biological interpretations, can be sensitive to the choice of TF-IDF parameters. Notably, precision and recall can be sensitive to the size of <italic>k</italic> (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>), and edges to the structure and delineation of groups (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>). It may be that different values of <italic>k</italic> are more sensitive to different assumptions or biological processes; because of this, we are interested in inferences of GECs that are robust to change of <italic>k</italic>. In the present work, three empirical genome-scale datasets we studied in detail earlier (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>) provide a solid foundation for addressing these issues. We add a fourth dataset to control further for balance across taxa, while removing a few poorly represented and/or anomalous taxa; and present the alignment-free LGT network analytical workflow end-to-end, including extraction of maximum and maximal cliques.</p>
<p>Specifically, here we examine (a) whether and how <italic>k</italic> affects cliques in LGT networks; (b) whether <italic>core nodes</italic>, stable to variation of <italic>k</italic> within biologically reasonable bounds, exist in different cliques; and (c) whether and how our biological process (functional) interpretation is consequently affected. More broadly, we believe that the approach pioneered here will provide a framework for understanding the extent and biological significance of LGT in complex environments.</p>
</sec>
<sec id="s1" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec><title>Datasets and Groups</title>
<p>Here we analyze four datasets, three of which we introduced earlier (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>): 20 <italic>Escherichia coli</italic> and seven <italic>Shigella</italic> genomes (ECS dataset), 110 enteric bacterial genomes (EB) and 143 genomes from BA. To these we now add a dataset of 144 bacterial genomes (BAC) purpose-built for this analysis. When this latter dataset was constructed, 24 bacterial orders in 12 classes were represented by at least one genus from which at least six genomes had been sequenced to high quality. Within each of these orders we selected one genus at random, and if that genus was represented by more than six genomes we chose six of them at random, thereby constituting BAC with 144 genomes in 24 genera. In this way we attempt to achieve as broad and balanced selection of genomes across Bacteria as possible, given the available data, a synthetic classification (NCBI) and the underlying biology.</p>
<p>As noted above, TF-IDF infers transfers from identified groups of donor genomes into a single recipient genome. It is therefore necessary to delineate groups prior to analysis. Here we recognize groups within the ECS dataset according to multi-locus sequence type (MLST; <xref ref-type="bibr" rid="B34">Gordon et al., 2008</xref>); within EB by genus, sometimes combining <italic>Escherichia</italic> and <italic>Shigella</italic> genomes into a single group; within BA by phylum, or alternatively by class; and within BAC by order. Other approaches to grouping are possible, some of which we explored earlier (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>).</p>
</sec>
<sec><title>Inference of Lateral Segments Using TF-IDF</title>
<p>The TF-IDF method <italic>per se</italic> proceeds in four steps, as follows: for each dataset we (A) extract unique <italic>k</italic>-mers and construct a <italic>k</italic>-mer dictionary; and (B) build a relationship matrix <italic>R</italic> in which rows represent individual genomes, columns represent the identified groups of genomes, and elements count the number of identical <italic>k</italic>-mers present in a genome <italic>and</italic> in each group other than its own. These counts are then normalized, and the mean element value computed over <italic>R</italic>. Unless indicated otherwise, this mean is used as the threshold for recognizing that a genome in the dataset may contain <italic>k</italic>-mers donated by a group in the dataset. (C) Within each genome, we then construct segments from neighboring <italic>k</italic>-mers that are present in the same donor group. We further merge these segments if they are separated by less than a gap threshold <italic>G</italic>, yielding <italic>potential lateral segments</italic>. (D) If the average frequency of <italic>k</italic>-mers in a potential lateral segment is less than that of all <italic>k</italic>-mers in the target genome&#x2019;s own group, then we consider it an <italic>inferred lateral segment</italic>. Step B implements the IDF component, and step D the TF component (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>,<xref ref-type="bibr" rid="B20">b</xref>). Pseudocode is available in the Supplementary Material to <xref ref-type="bibr" rid="B19">Cong et al. (2016a)</xref>, and the TF-IDF source code at <ext-link ext-link-type="uri" xlink:href="https://github.com/congyingnan/TF-IDF.git">https://github.com/congyingnan/TF-IDF.git</ext-link>.</p>
<p>Because the TF component requires a potential lateral segment to be infrequent in genomes of its own group, TF-IDF is expected to identify recent LGT events, i.e., those affecting one or a very few genomes in a target group. By contrast, <italic>k</italic>-mers descendant from transfers more ancient than the common ancestor of a target group would tend to occur widely within that group, and thus fail to be inferred as lateral.</p>
</sec>
<sec><title>Mapping Inferred Lateral Regions to Genes</title>
<p>We consider a gene to be lateral if it contains, or is overlapped by, at least one inferred lateral segment such that two distinct length thresholds are met. The inferred lateral segment must itself contain at least a specified minimum number of <italic>k</italic>-mers (including <italic>k</italic>-mers in any intervening gaps up to <italic>G</italic> = 2<italic>k</italic>); this minimum number is 10 for the BA dataset, 100 for EB, 500 for ECS (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>) and 10 for BAC. These values approximate the average length of all LGT detections in each dataset, thereby controlling in part for differences in sequence diversity among the datasets. In addition, the overlap must extend for at least a specified minimum number of <italic>k</italic>-mers (again including <italic>k</italic>-mers in gaps up to <italic>G</italic> = 2<italic>k</italic>); this minimum number is 10 for BA, 100 for ECS and EB (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>) and 10 for BAC.</p>
</sec>
<sec><title>End-to-End Workflow: Overview</title>
<p>As introduced above, GECs might variously be described as <italic>paths, transitively closed sets, paracliques</italic> or <italic>cliques</italic> (<xref ref-type="bibr" rid="B54">Skippington and Ragan, 2011</xref>). The first two structures fail to capture the density of connectivity, and many such structures of nearly equivalent size or value can often be found in relatively highly connected graphs such as the LGT networks we derive above. Paracliques differ from the corresponding cliques by relaxing the strict requirement that all edges be present, and in this way might better ameliorate the effects of incomplete or imperfect data. In the absence of theory or established practice, paraclique parameters would have to be optimized for each dataset, requiring intense computation. Constrained by these considerations, here we adopt the strictest yet clearest definition of GEC, as <italic>a set of vertices that share (donate or receive) genetic material from all other nodes within this set</italic>. That is, there must be at least one direct path between each node and every other node. Using this definition, GECs correspond to cliques in the LGT network.</p>
<p>The discovery and analysis of such cliques proceeds in four main steps:</p>
<list list-type="simple" prefix-word="simple">
<list-item><label>(a)</label><p>construct LGT networks based on the results of TF-IDF;</p></list-item>
<list-item><label>(b)</label><p>consolidate these networks by collapsing recipient genomes to recipient groups;</p></list-item>
<list-item><label>(c)</label><p>extract maximum and maximal cliques from the LGT network; and</p></list-item>
<list-item><label>(d)</label><p>perform enrichment tests on biological processes underlying the cliques.</p></list-item>
</list>
</sec>
<sec><title>Construction of LGT Networks</title>
<p>From our previous work (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>,<xref ref-type="bibr" rid="B20">b</xref>) we know that <italic>k</italic> can strongly affect the detection of LGT, hence potentially the topologies of LGT networks. For that reason, we explore different values of <italic>k</italic> to test the stability of clique topology. In step (<italic>a</italic>) we explore values of <italic>k</italic> from 20 to 40. Depending on the data, false positives can predominate at <italic>k</italic> &#x2264; 20, while at <italic>k</italic> &#x2265; 40 shared <italic>k</italic>-mers become too rare, resulting in diminished performance. For consistency with earlier studies on the ECS, EB and BA datasets (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>), gap size <italic>G</italic> was fixed at 2<italic>k</italic>. The step size is 10 for the ECS and EB datasets, while for BA (where LGT signal is much weaker) and BAC (not studied heretofore) we set step size as 5 for improved resolution against <italic>k</italic>.</p>
</sec>
<sec><title>Consolidation of the LGT Network Graph</title>
<p>Our TF-IDF procedure infers LGT from a donor group to a recipient sequence, so at this point the vertices in our inferred networks are of two types: individual genomes when they are recipients of LGT, and groups of genomes when they are donors. Of course, members of a group may individually be (and often are) recipients of LGT from outside that group. Edges are directional, so we depict them using an arrow from donor to recipient (<bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p><bold>Merger of LGT relationships from group-to-sequence into group-to-group</bold>.</p></caption>
<graphic xlink:href="fmicb-08-00021-g001.tif"/>
</fig>
<p>We aim to delineate GECs, introduced above as sets of nodes that have both donated genetic material to, and received genetic material from, each other. However, it is unclear how to extract these relationships when nodes are of two types, individual sequences and groups. For this reason we take only groups as nodes in our network analysis. In step (<italic>b</italic>) we subsume each genome into its respective own-group identity, and merge all directed edges into those sequences into a single directed edge from the donor group to the recipient group (<bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>). The integer weight on each edge gives the total number of genes inferred in this way to have been affected by LGT from donor to recipient groups.</p>
</sec>
<sec><title>Extraction of Maximum and Maximal Cliques</title>
<p>Extracting cliques [step (<italic>c</italic>)] is known to be NP-hard (<xref ref-type="bibr" rid="B39">Karp, 1972</xref>), although a parameterized complexity approach (<xref ref-type="bibr" rid="B1">Abu-Khzam et al., 2006</xref>) is possible whenever the clique size can be bounded (<xref ref-type="bibr" rid="B28">Downey et al., 1999</xref>). For this we used the Graph Algorithms Pipeline for Pathway Analysis (GrAPPA) software suite (Langston Lab, the University of Tennessee<sup><xref ref-type="fn" rid="fn01">1</xref></sup>). GrAPPA integrates multiple graph-theoretical tools for biological data analysis, including those designed to find cliques and paracliques. It implements tools to extract patterns efficiently from graphs, but deals only with undirected and unweighted networks. For this reason, we reformulated our directed networks as undirected graphs (i.e., we disregarded the arrowhead), deleted all weight annotations (number of inferred LGT genes) on each edge, and merged edges between pairs of nodes that are both donors and recipients. Such reformulation does not make full use of the LGT information (e.g., directionality) provided by TF-IDF, but nonetheless it preserves connectivity information sufficient for discovery of GECs as currently defined.</p>
<p>Here we use GrAPPA to find, for each dataset, one <italic>maximum clique</italic> (a clique with the greatest number of vertices) and all <italic>maximal cliques</italic> (cliques that are not included in a larger clique). We report only those maximal cliques with at least three vertices.</p>
</sec>
<sec><title>Gene Ontology Enrichment Tests</title>
<p>In step (<italic>d</italic>) we determine the biological processes that are enriched in the cliques previously extracted. As we show below (Results), the cliques inferred for the ECS and EB datasets encompass almost all the respective lateral genes, so the biological process enrichments are essentially the same as described previously (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>). Most LGT inferred for BAC involves the EB genera. Thus, here we report biological process enrichment only for cliques inferred, at different values of <italic>k</italic>, for the BA dataset. These genes are extracted from the GenBank genome record using GI numbers and coordinates, and collected as a test set. All genes in the dataset form the respective reference set. The enrichment statistic is a Fisher&#x2019;s exact test, for which we set false discovery rate FDR = 0.05 as the threshold for selecting over- and under-represented Gene Ontology (<xref ref-type="bibr" rid="B4">Ashburner et al., 2000</xref>; <xref ref-type="bibr" rid="B30">Gene Ontology Consortium, 2004</xref>) terms.</p>
</sec>
</sec>
<sec><title>Results</title>
<p>Detailed results including LGT networks, maximum and maximal cliques, gene lists, NCBI accession numbers for all genome sequences, and group composition are available as Supplementary Material (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">1&#x2013;19</xref>, and Supplementary Tables <xref ref-type="supplementary-material" rid="SM1">1&#x2013;34</xref>). Very large or detailed Supplementary Figures are also available for download in high resolution at <ext-link ext-link-type="uri" xlink:href="http://bioinformatics.org.au/tools-data/">http://bioinformatics.org.au/tools-data/</ext-link> under the category &#x201C;Other.&#x201D;</p>
<sec><title>ECS Dataset</title>
<p>We divide the ECS dataset (20 <italic>Escherichia coli</italic> and seven <italic>Shigella</italic> genomes) into six groups according to MLST (<xref ref-type="bibr" rid="B34">Gordon et al., 2008</xref>). In an earlier analysis of this dataset (<xref ref-type="bibr" rid="B55">Skippington and Ragan, 2012</xref>), lateral events identified by topological incongruence between trees inferred from individual putative orthogroups and an MRP (<xref ref-type="bibr" rid="B47">Ragan, 1992</xref>) reference supertree were shown to be biased more by phylogeny than by environment or lifestyle; concern was also expressed that defining GECs as cliques or paracliques might be too rigorous a standard.</p>
<p>Here, we use our TF-IDF method to infer LGT networks (<bold>Figure <xref ref-type="fig" rid="F2">2</xref></bold>). For all <italic>k</italic> examined here, all six phyletic groups belong to a single clique, so the whole dataset forms one large GEC. Indeed, at <italic>k</italic> = 30 or 40, topologies of the two networks are identical (as before: <xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>). There is a clear trend overall of more detections on each edge as <italic>k</italic> increases, but with some exceptions: at <italic>k</italic> = 20 we find three edges not seen at <italic>k</italic> = 30, from group D to B2 (257 transfers), from B2 to S (443) and from B2 to E (3574), while transfers from D to B1 decrease from 4659 at <italic>k</italic> = 20 to 3200 at <italic>k</italic> = 30. For all other edges, more genes are affected by LGT at <italic>k</italic> = 30 than at <italic>k</italic> = 20. Likewise, when <italic>k</italic> is increased from 30 to 40, three edges show fewer detections (D to B1, 3200 to 1842; B1 to B2, 3658 to 3563; E to B2, 3804 to 3363) but all others have more.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p><bold>LGT networks for ECS at (A)</bold> <italic>k</italic> = 20 and <bold>(B)</bold> <italic>k</italic> = 30. At <italic>k</italic> = 40 connectivity is the same as in <bold>(B)</bold>, although values on the edges are usually larger.</p></caption>
<graphic xlink:href="fmicb-08-00021-g002.tif"/>
</fig>
<p>Although clique topology is stable for 20 &#x2264;<italic>k</italic> &#x2264; 40, the total number of lateral genes underlying each clique increases with <italic>k</italic> (<bold>Figure <xref ref-type="fig" rid="F3">3</xref></bold>). This increase might appear to contradict our earlier finding that when <italic>k</italic> increases, the total number of detections and detection length should remain the same or decrease (at <italic>G</italic> = 2<italic>k</italic>). However, when <italic>k</italic> is small, more short segments tend to be detected as lateral (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">1</xref>). For example, at <italic>k</italic> = 20, 26% of lateral segments are &#x2265;500 <italic>k</italic>-mers in length, our threshold for selecting the segments for mapping to genes. This proportion increases to 31% at <italic>k</italic> = 40. Thus, we infer more lateral segments of &#x2265;500 <italic>k</italic>-mers at <italic>k</italic> = 40, which leads to more genes being inferred as affected by LGT.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p><bold>Total number of lateral segments of length &#x2265;500 <italic>k</italic>-mers (i.e., not only those mapping to genes), and total number of lateral genes, within the (maximum) clique inferred for the ECS dataset, as a function of <italic>k</italic>.</bold> For numerical values see Supplementary Table <xref ref-type="supplementary-material" rid="SM1">32</xref>.</p></caption>
<graphic xlink:href="fmicb-08-00021-g003.tif"/>
</fig>
</sec>
<sec><title>EB Dataset</title>
<p>The enteric bacteria dataset contains 110 genome sequences from five genera: <italic>Escherichia, Shigella, Salmonella, Klebsiella</italic> and <italic>Yersinia</italic>. Because the delineation of groups affects the detection results (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>), we infer LGT networks and extract cliques from five variants of this dataset: all genera present (referred to as EB-1); all genera except <italic>Shigella</italic> (EB-2) or alternatively, all except <italic>Escherichia</italic> (EB-3); with <italic>Escherichia</italic> and <italic>Shigella</italic> combined into a single group (EB-4); and with both <italic>Escherichia</italic> and <italic>Shigella</italic> removed (EB-5). These are five of the six variants we examined earlier (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>).</p>
<p>If we keep all 110 sequences and group them by genus (EB-1 dataset), the LGT network topologies change as <italic>k</italic> steps from 20 to 30 to 40. At <italic>k</italic> = 20, <italic>Escherichia, Shigella, Klebsiella</italic> constitute a single clique. At <italic>k</italic> = 30, we find two cliques, one consisting of <italic>Escherichia</italic> and <italic>Shigella</italic>, the other of <italic>Escherichia</italic> and <italic>Klebsiella</italic>. At <italic>k</italic> = 40 only one clique is found, consisting of <italic>Escherichia</italic> and <italic>Shigella</italic>. We infer many more LGT events between <italic>Escherichia</italic> and <italic>Shigella</italic> than between any other pair of genera. As <italic>Escherichia</italic> and <italic>Shigella</italic> are present in the clique across the examined range of <italic>k</italic>, we can say that they are the <italic>core nodes</italic> of this GEC.</p>
<p>Because genomes from <italic>Escherichia</italic> and <italic>Shigella</italic> share many more identical <italic>k</italic>-mers than do other groups, the lateral signal between these genera can drown out weaker lateral signal from or between other genera. This happens because the IDF values (elements of the <italic>R</italic> matrix) for these genomes are much higher than for the others (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>). This pushes up the IDF threshold, with the consequence that few lateral events are detected involving the other genera. To explore this effect, we also analyzed variant datasets which are modified so that <italic>Escherichia</italic> and <italic>Shigella</italic> do not both appear in the dataset as separate genera.</p>
<p>We first removed the <italic>Shigella</italic> genomes from the dataset while retaining those from <italic>Escherichia</italic>, thereby eliminating the effect of <italic>Shigella</italic> (EB-2 dataset). We now infer additional lateral events in both directions between all pairs of <italic>Escherichia, Salmonella</italic>, and <italic>Klebsiella</italic>. Thus we find a GEC composed of <italic>Escherichia, Klebsiella</italic> and <italic>Salmonella</italic> that remains stable with respect to <italic>k</italic>. We find similar results when we instead retain <italic>Shigella</italic> sequences while removing those of <italic>Escherichia</italic> (EB-3 dataset); the GEC here is <italic>Shigella, Klebsiella</italic> and <italic>Salmonella.</italic> We do infer lateral events between <italic>Klebsiella</italic> and <italic>Yersinia</italic> in EB-3, but these are not sufficient for <italic>Yersinia</italic> to join the GEC. In EB-4 we combine <italic>Escherichia</italic> and <italic>Shigella</italic> into a single group (ES); more lateral events were inferred from <italic>Salmonella</italic> to ES, but the GEC membership remains ES, <italic>Salmonella</italic> and <italic>Klebsiella</italic>. Lastly, to eliminate completely the effects of <italic>Escherichia</italic> and <italic>Shigella</italic> on LGT inference, we use only <italic>Klebsiella, Salmonella</italic> and <italic>Yersinia</italic> as input (EB-5). At <italic>k</italic> = 20 the sole clique contains all three genera, but at <italic>k</italic> = 30 or 40 the previous clique is split into two, one containing <italic>Klebsiella</italic> and <italic>Salmonella</italic> and the other <italic>Klebsiella</italic> and <italic>Yersinia</italic>. We thus conclude that <italic>Escherichia, Shigella, Klebsiella</italic> and <italic>Salmonella</italic> are all members of a larger GEC. Details are provided in <bold>Table <xref ref-type="table" rid="T1">1</xref></bold>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Lateral genes and cliques inferred for variants of the EB dataset at <italic>k</italic> = 20, 30, or 40.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Dataset</th>
<th valign="top" align="left"><italic>k</italic> size</th>
<th valign="top" align="left">Nodes in clique</th>
<th valign="top" align="center">Number of lateral genes in cliques</th>
<th valign="top" align="center">Number of lateral genes in network</th>
<th valign="top" align="center">Proportion (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">EB-1</td>
<td valign="top" align="left">20</td>
<td valign="top" align="left"><italic>Escherichia, Shigella, Klebsiella</italic></td>
<td valign="top" align="center">29527</td>
<td valign="top" align="center">29527</td>
<td valign="top" align="center">100%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>Escherichia, Shigella</italic></td>
<td valign="top" align="center">29258</td>
<td valign="top" align="center">29264</td>
<td valign="top" align="center">99.9%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>Escherichia, Klebsiella</italic></td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">29264</td>
<td valign="top" align="center">0.1%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">40</td>
<td valign="top" align="left"><italic>Escherichia, Shigella</italic></td>
<td valign="top" align="center">16968</td>
<td valign="top" align="center">16968</td>
<td valign="top" align="center">100%</td>
</tr>
<tr>
<td valign="top" align="left">EB-2</td>
<td valign="top" align="left">20</td>
<td valign="top" align="left"><italic>Escherichia, Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">23964</td>
<td valign="top" align="center">23970</td>
<td valign="top" align="center">99.9%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>Escherichia, Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">10840</td>
<td valign="top" align="center">10840</td>
<td valign="top" align="center">100%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">40</td>
<td valign="top" align="left"><italic>Escherichia, Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">7420</td>
<td valign="top" align="center">7426</td>
<td valign="top" align="center">99.9%</td>
</tr>
<tr>
<td valign="top" align="left">EB-3</td>
<td valign="top" align="left">20</td>
<td valign="top" align="left"><italic>Klebsiella, Salmonella, Shigella</italic></td>
<td valign="top" align="center">15290</td>
<td valign="top" align="center">15290</td>
<td valign="top" align="center">100%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>Klebsiella, Salmonella, Shigella</italic></td>
<td valign="top" align="center">6473</td>
<td valign="top" align="center">6501</td>
<td valign="top" align="center">99.5%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">40</td>
<td valign="top" align="left"><italic>Klebsiella, Salmonella, Shigella</italic></td>
<td valign="top" align="center">3869</td>
<td valign="top" align="center">3909</td>
<td valign="top" align="center">98.9%</td>
</tr>
<tr>
<td valign="top" align="left">EB-4</td>
<td valign="top" align="left">20</td>
<td valign="top" align="left"><italic>ES, Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">24806</td>
<td valign="top" align="center">24811</td>
<td valign="top" align="center">99.9%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>ES, Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">10762</td>
<td valign="top" align="center">10762</td>
<td valign="top" align="center">100%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">40</td>
<td valign="top" align="left"><italic>ES, Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">7951</td>
<td valign="top" align="center">7952</td>
<td valign="top" align="center">99.9%</td>
</tr>
<tr>
<td valign="top" align="left">EB-5</td>
<td valign="top" align="left">20</td>
<td valign="top" align="left"><italic>Klebsiella, Salmonella, Yersinia</italic></td>
<td valign="top" align="center">6721</td>
<td valign="top" align="center">6721</td>
<td valign="top" align="center">100%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>Klebsiella, Yersinia</italic></td>
<td valign="top" align="center">123</td>
<td valign="top" align="center">2586</td>
<td valign="top" align="center">4.8%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">30</td>
<td valign="top" align="left"><italic>Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">2463</td>
<td valign="top" align="center">2586</td>
<td valign="top" align="center">95.2%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">40</td>
<td valign="top" align="left"><italic>Klebsiella, Yersinia</italic></td>
<td valign="top" align="center">140</td>
<td valign="top" align="center">1559</td>
<td valign="top" align="center">9%</td>
</tr>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left">40</td>
<td valign="top" align="left"><italic>Klebsiella, Salmonella</italic></td>
<td valign="top" align="center">1419</td>
<td valign="top" align="center">1559</td>
<td valign="top" align="center">91%</td></tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec><title>BA Dataset</title>
<p>The BA dataset, 143 genome sequences across BA, has been studied in our group using classical alignment-based and other computational methods for more than a decade (<xref ref-type="bibr" rid="B8">Beiko et al., 2005</xref>; <xref ref-type="bibr" rid="B14">Chan et al., 2009a</xref>,<xref ref-type="bibr" rid="B16">b</xref>). Like many empirical datasets it is unbalanced, with many more genomes representing some taxa (e.g., Proteobacteria, Firmicutes) than others. We group these genomes into fifteen phyla or, alternatively, into 31 classes. With more nodes than in the two previous datasets, there is potential for inferred LGT networks to be more complex. On the other hand, these genomes are more dissimilar to each other (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>), so fewer <italic>k</italic>-mers are shared and fewer instances of LGT are inferred.</p>
<p>When groups are delineated by phylum, the number of total LGT detections decreases significantly as <italic>k</italic> increases (<bold>Figure <xref ref-type="fig" rid="F4">4</xref></bold>), and this causes edges in the LGT network to vanish and the cliques to shrink. At the smallest value of <italic>k</italic> = 20 six maximal cliques are found, each with five phyla. Five of these contain the High G+C Firmicutes, Proteobacteria and Low G+C Firmicutes, which together represent 14797 lateral genes, 95.5% of the total inferred over the entire network. Thus these phyla form the core of the inter-phylum GEC. We also observe a smaller GEC of Nanoarchaeota, Euryarchaeota and Crenarchaeota; although based on only 10 lateral genes, it is notable for showing potential GECs among Archaea. In addition, the <italic>Thermus/Deinococcus</italic> phylum contributes 244 lateral events, 1.5% of the total; as our dataset contains only one strain in this phylum, this particular genome appears to be more LGT-active than many other bacterial genomes.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p><bold>Number of inferred lateral genes in the BA dataset at 20 &#x2264; <italic>k</italic> &#x2264; 40, analyzed at the level of phylum or of class.</bold> For numerical values see Supplementary Table <xref ref-type="supplementary-material" rid="SM1">33</xref>.</p></caption>
<graphic xlink:href="fmicb-08-00021-g004.tif"/>
</fig>
<p>The number of detections drops sharply at <italic>k</italic> > 20; recall that our earlier simulations (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>) indicate potential false positives at <italic>k</italic> &#x2264; 20, presumably due to identical <italic>k</italic>-mers shared between sequences and groups simply by coincidence. As <italic>k</italic> increases and LGT detections decrease in number, some edges in the LGT network vanish, but the core nodes &#x2013; the High G+C Firmicutes, Low G+C Firmicutes and Proteobacteria &#x2013; remain as members of the maximal cliques. <italic>Thermus</italic>/<italic>Deinococcus</italic> also remains active in sharing LGT with Proteobacteria for all investigated <italic>k</italic>.</p>
<p>When these genomes are alternatively grouped by class, the LGT networks are more complex. Again we see a sharp drop in detections for <italic>k</italic> &#x2265; 20. At <italic>k</italic> = 20, all but one of the 31 classes are involved in LGT (30696 genes), and we observe 23 maximal cliques (&#x2265;3 nodes) in the LGT network; however, five classes form core members of the GEC, with each being present in 17 maximal cliques (&#x2265;5 nodes) and in the maximum clique. These classes are the Actinomycetales (5377 genes with lateral origin), <italic>Bacillus</italic>/<italic>Clostridium</italic> (2277) and the &#x03B1;- (5944), &#x03B2;- (7322) and &#x03B3;-Proteobacteria (8596). Together they contain 77.7% of all genes that contain regions of inferred lateral origin.</p>
<p>Since the sequences within BA are relatively dissimilar from each other, many fewer <italic>k</italic>-mers are shared between sequences than in the ECS and EB datasets. Thus the LGT detections are very sensitive to <italic>k</italic> (<bold>Figure <xref ref-type="fig" rid="F4">4</xref></bold>). At <italic>k</italic> = 30 the &#x03B3;- and &#x03B2;-Proteobacteria, Actinomycetales and <italic>Bacillus</italic>/<italic>Clostridium</italic> are hubs and play key roles in most cliques; at <italic>k</italic> = 40 fewer genes are inferred as lateral, and only the former two classes remain as the core.</p>
<p><italic>Deinococcus</italic> is inferred to exchange genetic material with &#x03B2;- and &#x03B3;-Proteobacteria at 20 &#x2264;<italic>k</italic> &#x2264; 40. Lateral events are also inferred between <italic>Deinococcus</italic> and Actinomycetales, and between <italic>Deinococcus</italic> and Chlorococcales, at <italic>k</italic> &#x003C; 40.</p>
</sec>
<sec><title>BAC Dataset</title>
<p>With the BAC dataset we again explore a broad phyletic range (24 orders representing 12 classes across Bacteria); but unlike the situation with BA (above), with BAC we maintain numerical balance (six genomes per order) and a comparable degree of local sequence diversity (each set of six genomes represents a single genus) to the extent possible, given the underlying biology and the availability of high-quality genome sequences. As simulations (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>) indicate a high likelihood of false positive detections at <italic>k</italic> = 20, here we vary <italic>k</italic> from 25 to 40 in steps of 5. As above, the total number of genes inferred to be affected by LGT events decreases with increased <italic>k</italic> (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>).</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p><bold>Number of inferred lateral genes in the BAC dataset at 25 &#x2264; k &#x2264; 40.</bold> For numerical values see Supplementary Table <xref ref-type="supplementary-material" rid="SM1">34</xref>.</p></caption>
<graphic xlink:href="fmicb-08-00021-g005.tif"/>
</fig>
<p>At the smallest value of <italic>k</italic> = 25 we infer 81 edges in the LGT network, connecting 17 nodes corresponding to orders of Proteobacteria (12), Low G+C Firmicutes (four) and High G+C Firmicutes (one). The largest clique inferred contains seven orders (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">16</xref> and Supplementary Table <xref ref-type="supplementary-material" rid="SM1">23</xref>). At <italic>k</italic> = 30 only four of these orders remain (Neisseriales, Enterobacteriales, Pasteurellales and Lactobacillales) in the maximum clique, and this clique persists through to <italic>k</italic> = 40 (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">17&#x2013;19</xref> and Supplementary Tables <xref ref-type="supplementary-material" rid="SM1">25, 27</xref>, and <xref ref-type="supplementary-material" rid="SM1">29</xref>). Thus these four orders form the core nodes of the GEC for the BAC dataset. This is in complete agreement with our results from the BA dataset.</p>
<p>At <italic>k</italic> = 25, we also infer Enterobacteriales (represented here by six <italic>E. coli</italic> genomes) to have donated via LGT to 125 genes in other orders and to have accepted LGT from other orders into 111 genes, together 34% of all affected genes across this dataset. These results support the developing themes of LGT being more successful among more-closely related genomes, with enteric bacteria and Firmicutes particularly active.</p>
<p>Other maximal cliques containing more than three orders are also found in the BAC dataset (Supplementary Tables <xref ref-type="supplementary-material" rid="SM1">24, 26, 28</xref>, and <xref ref-type="supplementary-material" rid="SM1">30</xref>) at different <italic>k</italic>. Some are subsets of the maximum clique, and reflect the fractions of the maximum clique in specific parts of the LGT network, while others are independent of the maximum clique and indicate other regions of dense connection among bacterial orders.</p>
</sec>
<sec><title>Enrichment of Biological Processes within Cliques</title>
<p>In addition to clique membership and topology, we are also interested in the biological processes enriched among the genes affected by inferred lateral events, as these may point to physiological, ecological and other processes that help to construct and maintain bacterial communities in nature. Analyses of the LGT networks inferred for the ECS and EB datasets reveal that more than 90% of the genes affected by LGT are represented in the corresponding cliques. In the ECS dataset, all vertices are in the maximum clique. In such cases there is no need to carry out enrichment tests: biological processes contributing to clique formation will be indistinguishable from those of the whole LGT network to which the cliques belong, i.e., the total LGT edge sets (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>). For the BAC dataset (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">31</xref>), genes annotated for involvement in metabolic processes (e.g., small-molecule and amino acid biosynthesis) are about twice as numerous as those in the next most-numerous category, ribosomal proteins: see the Supplementary Material for <xref ref-type="bibr" rid="B20">Cong et al. (2016b)</xref>, particularly Section 4.2.</p>
<p>For the BA dataset, however, clique topologies change significantly with <italic>k</italic>. Few LGT events are detected at <italic>k</italic> > 30, particularly when sequences are grouped by phylum (<bold>Figure <xref ref-type="fig" rid="F4">4</xref></bold>). For optimal comparison, we carried out enrichment tests on lateral genes of maximum cliques in each network at <italic>k</italic> = 20, 25 or 30, with genomes grouped either by phylum or by class. These tests identify biological processes related to metabolism, transport and regulation as over-represented when sequences are grouped by phylum. The term <italic>translational elongation</italic> (GO:0006414) ranks in first position at <italic>k</italic> = 20, and seventh at <italic>k</italic> = 25, among over-represented terms. The most significantly under-represented biological processes relate to transposition and to RNA modification at <italic>k</italic> = 20, and to RNA processing and biosynthetic processes at <italic>k</italic> = 25. The only term under-represented at <italic>k</italic> = 30 describes the modification of macromolecules.</p>
<p>When the genomes are grouped instead by class, the main categories of GO terms significantly over-represented remain those describing metabolism, transport and regulation. Those most under-represent relate to transposition, RNA metabolisms and regulation at <italic>k</italic> = 20 and 25; at <italic>k</italic> = 30, processes of protein modification are under-represented.</p>
<p>In general, the patterns of over-representation are similar between analyses at phylum and class levels. Interestingly, <italic>translation elongation</italic> is significantly over-represented at phylum level, but much less so at class level. <italic>Transposition</italic> (GO:0032196) is significantly under-represented in most cases.</p>
</sec>
</sec>
<sec><title>Discussion</title>
<p>Here we inferred LGT networks for four datasets of different phyletic breadth, hence evolutionary depth. For the ECS dataset, the entire LGT network is captured within a single clique encompassing all nodes, consistent with previous research (<xref ref-type="bibr" rid="B55">Skippington and Ragan, 2012</xref>). Interplay with the IDF threshold is seen clearly with the EB dataset and its variants. For the full EB dataset (EB-1), the LGT signal between <italic>Escherichia</italic> and <italic>Shigella</italic> is much stronger than that of any other pairwise comparison and dominates the lateral signal, with the result that the only community that can be found is <italic>Escherichia</italic> and <italic>Shigella</italic>. If we remove <italic>Escherichia</italic> or (alternatively) <italic>Shigella</italic>, or combine them into a single group, we detect LGT events from (and/or to) <italic>Klebsiella</italic> and <italic>Salmonella</italic>. This reveals a larger clique containing either <italic>Escherichia</italic> or <italic>Shigella</italic>, plus <italic>Klebsiella</italic> and <italic>Salmonella</italic> (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">2&#x2013;5</xref>). By contrast, <italic>Yersinia</italic> is relatively silent to LGT, and contributes little to the community.</p>
<p>Particularly in the BA dataset, we see that different parts of the LGT network are differentially sensitive to change of <italic>k</italic>. When <italic>k</italic> is small (here <italic>k</italic> = 20), many <italic>k</italic>-mers are shared by chance, resulting in many false positive inferences (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>,<xref ref-type="bibr" rid="B20">b</xref>). Edges supported by large numbers of lateral events (e.g., those with high weights) tend to persist, whereas those representing smaller numbers of events may disappear as <italic>k</italic> is incremented. Even so, when the sequences are grouped by phylum, the High-G+C Firmicutes, Low-G+C Firmicutes and Proteobacteria are found in all cliques inferred across the investigated range of parameter values (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">6&#x2013;10</xref>, Supplementary Tables <xref ref-type="supplementary-material" rid="SM1">2&#x2013;11</xref>). For this reason we identify them as core nodes of the GEC for the BA phyla. Although it does not contribute many LGT events, <italic>Thermus</italic>/<italic>Deinococcus</italic> is also a member of most communities.</p>
<p>When the BA dataset is grouped into 31 classes, many more clique structures are found. The &#x03B1;-, &#x03B2;- and &#x03B3;-Proteobacteria, Actinomycetales and <italic>Bacillus</italic>/<italic>Clostridium</italic> are always present in at least one clique (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">11&#x2013;15</xref>, Supplementary Tables <xref ref-type="supplementary-material" rid="SM1">12&#x2013;21</xref>), i.e., are core nodes. This agrees with an earlier conclusion, based on classical alignment-based phylogenomic methods, that these groups are connected by major highways of LGT (<xref ref-type="bibr" rid="B8">Beiko et al., 2005</xref>). By contrast, the &#x1D700;-Proteobacteria appear relatively silent to LGT, with fewer inferred events per genome (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">22</xref>). In the class-level LGT network, the sole <italic>Deinococcus</italic> genome is also involved in many (maximum and maximal) cliques, linked through a lateral edge with subdivisions from Proteobacteria. Stronger connectivity might be expected if more sequences from Deinococci and its immediate relatives were represented in this dataset.</p>
<p>Although many fewer instances of LGT are inferred involving archaea, we nonetheless recognize one clique among them. The low frequency of inferred LGT events may arise because these genomes are relatively diverse in gene content and phylogenetically distant from each other, and/or because in reality these genomes have exchanged little genetic material, for example because they live in specialized environments (<xref ref-type="bibr" rid="B8">Beiko et al., 2005</xref>; <xref ref-type="bibr" rid="B45">Popa et al., 2011</xref>). In the former case TF-IDF should find instances of LGT but the pairwise values may not pass the IDF threshold, whereas in the latter case there would be little true-positive LGT to be found and lowering the IDF threshold would lead only to false-positive inferences. Comparing the results of TF-IDF with those of classical alignment-based methods may help distinguish between these alternative explanations.</p>
<p>Enrichment tests on the BA data reveal that a wide range of biological processes are over-represented in the LGT events that underpin the cliques identified. As expected (<xref ref-type="bibr" rid="B37">Jain et al., 1999</xref>, <xref ref-type="bibr" rid="B38">2003</xref>), metabolic processes, gene regulation, and trans-membrane and intracellular transport are broadly represented. For example, at <italic>k</italic> = 25 with genomes grouped by class, 39 of the 50 most over-represented processes describe metabolism. Terms associated with transposition or antibiotic resistance are not seen: these genes are usually transferred within-phylum or within-class (or indeed more narrowly) and often occur on plasmids, which are not represented in the genome data files we used. As expected, few terms describing processes of transcription, translation or DNA replication (<xref ref-type="bibr" rid="B37">Jain et al., 1999</xref>, <xref ref-type="bibr" rid="B38">2003</xref>) are overrepresented.</p>
<p>Fewer biological process terms are under-represented among the LGT events that underpin the BA cliques, although <italic>transposition</italic> (GO:0032196) is very significantly under-represented. A similar result was also found for the ECS dataset (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">4</xref>). From previous research (<xref ref-type="bibr" rid="B20">Cong et al., 2016b</xref>) we know that genes annotated with this term are widespread in the ECS genomes, making it difficult for genes annotated with this term to pass the TF threshold for detection. In the BA dataset, genomes of <italic>E. coli</italic> and <italic>Shigella</italic> are a major source of genes associated with transposition; as these are members of the same group (&#x03B3;-Proteobacteria), they are not detected by TF-IDF. In the EB dataset, when <italic>Escherichia</italic> and <italic>Shigella</italic> are not treated as separate groups, <italic>transposition</italic> is not significantly under-represented (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">5</xref>). Thus TF-IDF is not blind to such mobile biological processes, but the way groups are delimited can limit their discovery.</p>
<p>This work represents the first systematic exploration of the sensitivity of densely connected structures (maximum and maximal cliques) in LGT graphs to choice of parameter values in an alignment-free framework. Our workflow is the first to implement alignment-free and other highly scalable methods end-to-end, from whole genome sequences to delineation GECs and functional analysis of the genes affected by LGT. Our results confirm the promise of this approach, notably the robustness of clique structure and membership at sufficiently large <italic>k</italic>, here <italic>k</italic> &#x2265; 25. Nonetheless, important challenges remain.</p>
<p>Computational simulations and empirical studies demonstrate that approaches based on <italic>k</italic>-mer count can support the scalable inference of phylogenies (<xref ref-type="bibr" rid="B15">Chan et al., 2014</xref>; <xref ref-type="bibr" rid="B10">Bernard et al., 2016a</xref>,<xref ref-type="bibr" rid="B11">b</xref>) and identify regions of lateral transfer within a dataset (<xref ref-type="bibr" rid="B19">Cong et al., 2016a</xref>,<xref ref-type="bibr" rid="B20">b</xref>). Parameters including <italic>k</italic> can be adjusted to minimize the effects of sequence divergence and genome rearrangement. However, word-count methods are less robust to sequence loss or truncation (<xref ref-type="bibr" rid="B15">Chan et al., 2014</xref>). As is the case with classical phylogenetics, other scenarios likely to erode the performance of word-count methods include compositional bias and/or rate variation within genomes or across lineages, including convergent processes in distantly related sequences. Methods will need to be developed such that alignment-free approaches, including TF-IDF, can mitigate or avoid these situations.</p>
<p>Graph-theoretical research has primarily concentrated on difficult combinatorial problems posed on finite, simple graphs. Graph analytical software packages such as GrAPPA (Langston Lab, the University of Tennessee<sup><xref ref-type="fn" rid="fn02">2</xref></sup>), therefore, are designed mainly for undirected, unweighted graphs. This has required us to ignore both directionality (by merger of incoming and outgoing edges) and weights. Such simplifications represent a classic pre-processing step for a directed network (<xref ref-type="bibr" rid="B53">Seidman and Foster, 1978</xref>). While other strategies have been introduced to find cliques in directed networks, all involve weakening the edges, and none can guarantee a better interpretation of properties of the original directed network (<xref ref-type="bibr" rid="B52">Seidman, 1980</xref>; <xref ref-type="bibr" rid="B44">Palla et al., 2007</xref>). Comparing these approaches across various application domains remains an open problem. Despite this limitation, some features of the role played by LGT in the evolution of microbes are still accessible. A good example is the frequent exchange inferred among <italic>Escherichia</italic> and <italic>Shigella</italic> contrasted with the relative isolation of <italic>Yersinia</italic>.</p>
<p>It is noteworthy that we have defined GECs as cliques, because the clique is a rigorous graph-theoretical structure that maps particularly well onto numerous biological concepts, in the present case the sharing of genetic information via LGT. While this makes sense in the quest for biological fidelity, in mapping GECs onto LGT graphs <xref ref-type="bibr" rid="B54">Skippington and Ragan (2011)</xref> expressed concern that missing data might make clique too rigorous a definition. We share this reservation, and observe that noise-resilient options such as paraclique may fare better. We need only a criterion, e.g., paraclique&#x2019;s <italic>glom</italic> term (<xref ref-type="bibr" rid="B35">Hagan et al., 2016</xref>), by which to estimate the number or proportion of &#x201C;missing&#x201D; edges. An exploration of such criteria may be the subject of future work.</p>
</sec>
<sec><title>Author Contributions</title>
<p>All authors designed the experiments. YC implemented and carried out the computational analyses. CP and ML provided software. All authors analyzed the data, and wrote and edited the paper.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> YC acknowledges the China Scholarship Council and The University of Queensland for stipend and tuition fee support. The project was funded in part by grant 220020272 from the James S. McDonnell Foundation to MR. The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.</p></fn>
</fn-group>
<ack>
<p>This research utilized resources of QCIF (Queensland Cyber Infrastructure Foundation), which is supported by the Queensland and Australian Governments. We thank Mr Brett Dunsmore and staff of Information Technology Services, Institute for Molecular Bioscience for additional computational support.</p>
</ack>
<sec sec-type="supplementary material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fmicb.2017.00021/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fmicb.2017.00021/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abu-Khzam</surname> <given-names>F. N.</given-names></name> <name><surname>Langston</surname> <given-names>M. A.</given-names></name> <name><surname>Shanbhag</surname> <given-names>P.</given-names></name> <name><surname>Symons</surname> <given-names>C. T.</given-names></name></person-group> (<year>2006</year>). <article-title>Scalable parallel algorithms for FPT problems.</article-title> <source><italic>Algorithmica</italic></source> <volume>45</volume> <fpage>269</fpage>&#x2013;<lpage>284</lpage>. <pub-id pub-id-type="doi">10.1007/s00453-006-1214-1</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ambler</surname> <given-names>R. P.</given-names></name> <name><surname>Daniel</surname> <given-names>M.</given-names></name> <name><surname>Hermoso</surname> <given-names>J.</given-names></name> <name><surname>Meyer</surname> <given-names>T. E.</given-names></name> <name><surname>Bartsch</surname> <given-names>R. G.</given-names></name> <name><surname>Kamen</surname> <given-names>M. D.</given-names></name></person-group> (<year>1979a</year>). <article-title>Cytochrome c2 sequence variation among the recognised species of purple nonsulphur photosynthetic bacteria.</article-title> <source><italic>Nature</italic></source> <volume>278</volume> <fpage>659</fpage>&#x2013;<lpage>660</lpage>. <pub-id pub-id-type="doi">10.1038/278661a0</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ambler</surname> <given-names>R. P.</given-names></name> <name><surname>Meyer</surname> <given-names>T. E.</given-names></name> <name><surname>Kamen</surname> <given-names>M. D.</given-names></name></person-group> (<year>1979b</year>). <article-title>Anomalies in amino acid sequences of small cytochromes c and cytochromes c&#x2019; from two species of purple photosynthetic bacteria.</article-title> <source><italic>Nature</italic></source> <volume>278</volume> <fpage>661</fpage>&#x2013;<lpage>662</lpage>. <pub-id pub-id-type="doi">10.1038/278661a0</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashburner</surname> <given-names>M.</given-names></name> <name><surname>Ball</surname> <given-names>C. A.</given-names></name> <name><surname>Blake</surname> <given-names>J. A.</given-names></name> <name><surname>Botstein</surname> <given-names>D.</given-names></name> <name><surname>Butler</surname> <given-names>H.</given-names></name> <name><surname>Cherry</surname> <given-names>J. M.</given-names></name><etal/></person-group> (<year>2000</year>). <article-title>Gene ontology: tool for the unification of biology.</article-title> <source><italic>Nat. Genet.</italic></source> <volume>25</volume> <fpage>25</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1038/75556</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bansal</surname> <given-names>A. K.</given-names></name> <name><surname>Bork</surname> <given-names>P.</given-names></name> <name><surname>Stuckey</surname> <given-names>P. J.</given-names></name></person-group> (<year>1998</year>). <article-title>Automated pair-wise comparisons of microbial genomes.</article-title> <source><italic>Math. Model. Sci. Comput.</italic></source> <volume>9</volume> <fpage>1</fpage>&#x2013;<lpage>23</lpage>.</citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bapteste</surname> <given-names>E.</given-names></name> <name><surname>van Iersel</surname> <given-names>L.</given-names></name> <name><surname>Janke</surname> <given-names>A.</given-names></name> <name><surname>Kelchner</surname> <given-names>S.</given-names></name> <name><surname>Kelk</surname> <given-names>S.</given-names></name> <name><surname>McInerney</surname> <given-names>J. O.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Networks: expanding evolutionary thinking.</article-title> <source><italic>Trends Genet.</italic></source> <volume>29</volume> <fpage>439</fpage>&#x2013;<lpage>441</lpage>. <pub-id pub-id-type="doi">10.1016/j.tig.2013.05.007</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beiko</surname> <given-names>R. G.</given-names></name> <name><surname>Hamilton</surname> <given-names>N.</given-names></name></person-group> (<year>2006</year>). <article-title>Phylogenetic identification of lateral genetic transfer events.</article-title> <source><italic>BMC Evol. Biol.</italic></source> <volume>6</volume>:<issue>15</issue>. <pub-id pub-id-type="doi">10.1186/1471-2148-6-15</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beiko</surname> <given-names>R. G.</given-names></name> <name><surname>Harlow</surname> <given-names>T. J.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2005</year>). <article-title>Highways of gene sharing in prokaryotes.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>102</volume> <fpage>14332</fpage>&#x2013;<lpage>14337</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0504068102</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beiko</surname> <given-names>R. G.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2008</year>). <article-title>Detecting lateral genetic transfer : a phylogenetic approach.</article-title> <source><italic>Methods Mol. Biol.</italic></source> <volume>452</volume> <fpage>457</fpage>&#x2013;<lpage>469</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-60327-159-2_21</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernard</surname> <given-names>G.</given-names></name> <name><surname>Chan</surname> <given-names>C. X.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2016a</year>). <article-title>Alignment-free microbial phylogenomics under scenarios of sequence divergence, genome rearrangement and lateral generic transfer.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>6</volume> <issue>28970</issue>. <pub-id pub-id-type="doi">10.1038/srep28970</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernard</surname> <given-names>G.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name> <name><surname>Chan</surname> <given-names>C. X.</given-names></name></person-group> (<year>2016b</year>). <article-title>Recapitulating phylogenies using k-mers: from trees to networks.</article-title> <source><italic>F1000 Research</italic></source> <volume>5</volume> <issue>2789</issue>. <pub-id pub-id-type="doi">10.12688/f1000research.10225.2</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bork</surname> <given-names>P.</given-names></name> <name><surname>Dandekar</surname> <given-names>T.</given-names></name> <name><surname>Diaz-Lazcoz</surname> <given-names>Y.</given-names></name> <name><surname>Eisenhaber</surname> <given-names>F.</given-names></name> <name><surname>Huynen</surname> <given-names>M.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name></person-group> (<year>1998</year>). <article-title>Predicting function: from genes to genomes and back.</article-title> <source><italic>J. Mol. Biol.</italic></source> <volume>283</volume> <fpage>707</fpage>&#x2013;<lpage>725</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.1998.2144</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bryant</surname> <given-names>D.</given-names></name> <name><surname>Moulton</surname> <given-names>V.</given-names></name></person-group> (<year>2004</year>). <article-title>Neighbor-net: an agglomerative method for the construction of phylogenetic networks.</article-title> <source><italic>Mol. Biol. Evol.</italic></source> <volume>21</volume> <fpage>255</fpage>&#x2013;<lpage>265</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msh018</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>C. X.</given-names></name> <name><surname>Beiko</surname> <given-names>R. G.</given-names></name> <name><surname>Darling</surname> <given-names>A. E.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2009a</year>). <article-title>Lateral transfer of genes and gene fragments in prokaryotes.</article-title> <source><italic>Genome Biol. Evol.</italic></source> <volume>1</volume> <fpage>429</fpage>&#x2013;<lpage>438</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evp044</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>C. X.</given-names></name> <name><surname>Bernard</surname> <given-names>G.</given-names></name> <name><surname>Poirion</surname> <given-names>O.</given-names></name> <name><surname>Hogan</surname> <given-names>J. M.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2014</year>). <article-title>Inferring phylogenies of evolving sequences without multiple sequence alignment.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>4</volume> <issue>6504</issue>. <pub-id pub-id-type="doi">10.1038/srep06504</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>C. X.</given-names></name> <name><surname>Darling</surname> <given-names>A. E.</given-names></name> <name><surname>Beiko</surname> <given-names>R. G.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2009b</year>). <article-title>Are protein domains modules of lateral genetic transfer?</article-title> <source><italic>PLoS ONE</italic></source> <volume>4</volume>:<issue>e4524</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0004524</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chesler</surname> <given-names>E. J.</given-names></name> <name><surname>Langston</surname> <given-names>M. A.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x201C;Combinatorial genetic regulatory network analysis tools for high throughput transcriptomic data,&#x201D; in</article-title> <source><italic>Systems Biology and Regulatory Genomics, Lecture Notes in Computer Science Series 4023</italic></source> <role>eds</role> <person-group person-group-type="editor"><name><surname>Eskin</surname> <given-names>E.</given-names></name> <name><surname>Ideker</surname> <given-names>T.</given-names></name> <name><surname>Raphael</surname> <given-names>B.</given-names></name> <name><surname>Workman</surname> <given-names>C.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>150</fpage>&#x2013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-540-48540-7_13</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clarke</surname> <given-names>G. D.</given-names></name> <name><surname>Beiko</surname> <given-names>R. G.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name> <name><surname>Charlebois</surname> <given-names>R. L.</given-names></name></person-group> (<year>2002</year>). <article-title>Inferring genome trees by using a filter to eliminate phylogenetically discordant sequences and a distance matrix based on mean normalized BLASTP scores.</article-title> <source><italic>J. Bacteriol.</italic></source> <volume>184</volume> <fpage>2072</fpage>&#x2013;<lpage>2080</lpage>. <pub-id pub-id-type="doi">10.1128/JB.184.8.2072-2080.2002</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cong</surname> <given-names>Y.</given-names></name> <name><surname>Chan</surname> <given-names>Y.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2016a</year>). <article-title>A novel alignment-free method for detection of lateral genetic transfer based on TF-IDF.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>6</volume> <issue>30308</issue>. <pub-id pub-id-type="doi">10.1038/srep30308</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cong</surname> <given-names>Y.</given-names></name> <name><surname>Chan</surname> <given-names>Y.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2016b</year>). <article-title>Exploring lateral genetic transfer among microbial genomes using TF-IDF.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>6</volume> <issue>29319</issue>. <pub-id pub-id-type="doi">10.1038/srep29319</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Croucher</surname> <given-names>N. J.</given-names></name> <name><surname>Coupland</surname> <given-names>P. G.</given-names></name> <name><surname>Stevenson</surname> <given-names>A. E.</given-names></name> <name><surname>Callendrello</surname> <given-names>A.</given-names></name> <name><surname>Bentley</surname> <given-names>S. D.</given-names></name> <name><surname>Hanage</surname> <given-names>W. P.</given-names></name></person-group> (<year>2014</year>). <article-title>Diversification of bacterial genome content through distinct mechanisms over different timescales.</article-title> <source><italic>Nat. Commun.</italic></source> <volume>5</volume> <issue>5471</issue>. <pub-id pub-id-type="doi">10.1038/ncomms6471</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dagan</surname> <given-names>T.</given-names></name> <name><surname>Artzy-Randrup</surname> <given-names>Y.</given-names></name> <name><surname>Martin</surname> <given-names>W.</given-names></name></person-group> (<year>2008</year>). <article-title>Modular networks and cumulative impact of lateral transfer in prokaryote genome evolution.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>105</volume> <fpage>10039</fpage>&#x2013;<lpage>10044</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0800679105</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dagan</surname> <given-names>T.</given-names></name> <name><surname>Martin</surname> <given-names>W.</given-names></name></person-group> (<year>2009</year>). <article-title>Getting a better picture of microbial evolution en route to a network of genomes.</article-title> <source><italic>Philos. Trans. R. Soc. Lond. B Biol. Sci.</italic></source> <volume>364</volume> <fpage>2187</fpage>&#x2013;<lpage>2196</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2009.0040</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dickerson</surname> <given-names>R. E.</given-names></name></person-group> (<year>1980</year>). <article-title>Evolution and gene transfer in purple photosynthetic bacteria.</article-title> <source><italic>Nature</italic></source> <volume>283</volume> <fpage>210</fpage>&#x2013;<lpage>212</lpage>. <pub-id pub-id-type="doi">10.1038/283210a0</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doolittle</surname> <given-names>W. F.</given-names></name></person-group> (<year>1999</year>). <article-title>Phylogenetic classification and the universal tree.</article-title> <source><italic>Science</italic></source> <volume>284</volume> <fpage>2124</fpage>&#x2013;<lpage>2129</lpage>. <pub-id pub-id-type="doi">10.1126/science.284.5423.2124</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doolittle</surname> <given-names>W. F.</given-names></name></person-group> (<year>2009</year>). <article-title>The practice of classification and the theory of evolution, and what the demise of Charles Darwin&#x2019;s tree of life hypothesis means for both of them.</article-title> <source><italic>Philos. Trans. R. Soc. Lond. B Biol. Sci.</italic></source> <volume>364</volume> <fpage>2221</fpage>&#x2013;<lpage>2228</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2009.0032</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doolittle</surname> <given-names>W. F.</given-names></name> <name><surname>Bapteste</surname> <given-names>E.</given-names></name></person-group> (<year>2007</year>). <article-title>Pattern pluralism and the Tree of Life hypothesis.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>104</volume> <fpage>2043</fpage>&#x2013;<lpage>2049</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0610699104</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Downey</surname> <given-names>R. G.</given-names></name> <name><surname>Fellows</surname> <given-names>M. R.</given-names></name> <name><surname>Stege</surname> <given-names>U.</given-names></name></person-group> (<year>1999</year>). <article-title>Parameterized complexity: a framework for systematically confronting computational intractability.</article-title> <source><italic>DIMACS</italic></source> <volume>49</volume> <fpage>49</fpage>&#x2013;<lpage>99</lpage>.</citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fondi</surname> <given-names>M.</given-names></name> <name><surname>Fani</surname> <given-names>R.</given-names></name></person-group> (<year>2010</year>). <article-title>The horizontal flow of the plasmid resistome: clues from inter-generic similarity networks.</article-title> <source><italic>Environ. Microbiol.</italic></source> <volume>12</volume> <fpage>3228</fpage>&#x2013;<lpage>3242</lpage>. <pub-id pub-id-type="doi">10.1111/j.1462-2920.2010.02295.x</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><collab>Gene Ontology Consortium</collab> (<year>2004</year>). <article-title>The Gene Ontology (GO) database and informatics resource.</article-title> <source><italic>Nucl. Acids Res.</italic></source> <volume>32</volume> <fpage>D258</fpage>&#x2013;<lpage>D261</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh036</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gogarten</surname> <given-names>J. P.</given-names></name> <name><surname>Doolittle</surname> <given-names>W. F.</given-names></name> <name><surname>Lawrence</surname> <given-names>J. G.</given-names></name></person-group> (<year>2002</year>). <article-title>Prokaryotic evolution in light of gene transfer.</article-title> <source><italic>Mol. Biol. Evol.</italic></source> <volume>19</volume> <fpage>2226</fpage>&#x2013;<lpage>2238</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a004046</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gogarten</surname> <given-names>J. P.</given-names></name> <name><surname>Townsend</surname> <given-names>J. P.</given-names></name></person-group> (<year>2005</year>). <article-title>Horizontal gene transfer, genome innovation and evolution.</article-title> <source><italic>Nat. Rev. Microbiol.</italic></source> <volume>3</volume> <fpage>679</fpage>&#x2013;<lpage>687</lpage>. <pub-id pub-id-type="doi">10.1038/nrmicro1204</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldman</surname> <given-names>N.</given-names></name> <name><surname>Anderson</surname> <given-names>J. P.</given-names></name> <name><surname>Rodrigo</surname> <given-names>A. G.</given-names></name></person-group> (<year>2000</year>). <article-title>Likelihood-based tests of topologies in phylogenetics.</article-title> <source><italic>Syst. Biol.</italic></source> <volume>49</volume> <fpage>652</fpage>&#x2013;<lpage>670</lpage>. <pub-id pub-id-type="doi">10.1080/106351500750049752</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gordon</surname> <given-names>D. M.</given-names></name> <name><surname>Clermont</surname> <given-names>O.</given-names></name> <name><surname>Tolley</surname> <given-names>H.</given-names></name> <name><surname>Denamur</surname> <given-names>E.</given-names></name></person-group> (<year>2008</year>). <article-title>Assigning <italic>Escherichia coli</italic> strains to phylogenetic groups: multi-locus sequence typing versus the PCR triplex method.</article-title> <source><italic>Environ. Microbiol.</italic></source> <volume>10</volume> <fpage>2484</fpage>&#x2013;<lpage>2496</lpage>. <pub-id pub-id-type="doi">10.1111/j.1462-2920.2008.01669.x</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hagan</surname> <given-names>R. D.</given-names></name> <name><surname>Langston</surname> <given-names>M. A.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Lower bounds on paraclique density.</article-title> <source><italic>Discr. Appl. Math.</italic></source> <volume>204</volume> <fpage>208</fpage>&#x2013;<lpage>212</lpage>. <pub-id pub-id-type="doi">10.1016/j.dam.2015.11.010</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Halary</surname> <given-names>S.</given-names></name> <name><surname>Leigh</surname> <given-names>J. W.</given-names></name> <name><surname>Cheaib</surname> <given-names>B.</given-names></name> <name><surname>Lopez</surname> <given-names>P.</given-names></name> <name><surname>Bapteste</surname> <given-names>E.</given-names></name></person-group> (<year>2010</year>). <article-title>Network analyses structure genetic diversity in independent genetic worlds.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>107</volume> <fpage>127</fpage>&#x2013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0908978107</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jain</surname> <given-names>R.</given-names></name> <name><surname>Rivera</surname> <given-names>M. C.</given-names></name> <name><surname>Lake</surname> <given-names>J. A.</given-names></name></person-group> (<year>1999</year>). <article-title>Horizontal gene transfer among genomes: the complexity hypothesis.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>96</volume> <fpage>3801</fpage>&#x2013;<lpage>3806</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.96.7.3801</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jain</surname> <given-names>R.</given-names></name> <name><surname>Rivera</surname> <given-names>M. C.</given-names></name> <name><surname>Moore</surname> <given-names>J. E.</given-names></name> <name><surname>Lake</surname> <given-names>J. A.</given-names></name></person-group> (<year>2003</year>). <article-title>Horizontal gene transfer accelerates genome innovation and evolution.</article-title> <source><italic>Mol. Biol. Evol.</italic></source> <volume>20</volume> <fpage>1598</fpage>&#x2013;<lpage>1602</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msg154</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Karp</surname> <given-names>R. M.</given-names></name></person-group> (<year>1972</year>). <article-title>&#x201C;Reducibility among combinatorial problems,&#x201D; in</article-title> <source><italic>Complexity of Computer Computations</italic></source> <role>eds</role> <person-group person-group-type="editor"><name><surname>Miller</surname> <given-names>R. E.</given-names></name> <name><surname>Thatcher</surname> <given-names>J. W.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Plenum</publisher-name>) <fpage>85</fpage>&#x2013;<lpage>103</lpage>.</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koonin</surname> <given-names>E. V.</given-names></name></person-group> (<year>2015</year>). <article-title>The turbulent network dynamics of microbial evolution and the statistical Tree of Life.</article-title> <source><italic>J. Mol. Evol.</italic></source> <volume>80</volume> <fpage>244</fpage>&#x2013;<lpage>250</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-015-9679-7</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kunin</surname> <given-names>V.</given-names></name> <name><surname>Goldovsky</surname> <given-names>L.</given-names></name> <name><surname>Darzentas</surname> <given-names>N.</given-names></name> <name><surname>Ouzounis</surname> <given-names>C. A.</given-names></name></person-group> (<year>2005</year>). <article-title>The net of life: reconstructing the microbial phylogenetic network.</article-title> <source><italic>Genome Res.</italic></source> <volume>15</volume> <fpage>954</fpage>&#x2013;<lpage>959</lpage>. <pub-id pub-id-type="doi">10.1101/gr.3666505</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lima-Mendez</surname> <given-names>G.</given-names></name> <name><surname>Van Helden</surname> <given-names>J.</given-names></name> <name><surname>Toussaint</surname> <given-names>A.</given-names></name> <name><surname>Leplae</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Reticulate representation of evolutionary and functional relationships between phage genomes.</article-title> <source><italic>Mol. Biol. Evol.</italic></source> <volume>25</volume> <fpage>762</fpage>&#x2013;<lpage>777</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msn023</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x2019;Malley</surname> <given-names>M. A.</given-names></name> <name><surname>Koonin</surname> <given-names>E. V.</given-names></name></person-group> (<year>2011</year>). <article-title>How stands the Tree of Life a century and a half after The Origin?</article-title> <source><italic>Biol. Direct</italic></source> <volume>6</volume> <issue>32</issue>. <pub-id pub-id-type="doi">10.1186/1745-6150-6-32</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Palla</surname> <given-names>G.</given-names></name> <name><surname>Farkas</surname> <given-names>I. J.</given-names></name> <name><surname>Pollner</surname> <given-names>P.</given-names></name> <name><surname>Derenyi</surname> <given-names>I.</given-names></name> <name><surname>Vicsek</surname> <given-names>T.</given-names></name></person-group> (<year>2007</year>). <article-title>Directed network modules.</article-title> <source><italic>New J. Phys.</italic></source> <volume>9</volume> <issue>186</issue>. <pub-id pub-id-type="doi">10.1088/1367-2630/9/6/186</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Popa</surname> <given-names>O.</given-names></name> <name><surname>Hazkani-Covo</surname> <given-names>E.</given-names></name> <name><surname>Landan</surname> <given-names>G.</given-names></name> <name><surname>Martin</surname> <given-names>W.</given-names></name> <name><surname>Dagan</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>Directed networks reveal genomic barriers and DNA repair bypasses to lateral gene transfer among prokaryotes.</article-title> <source><italic>Genome Res.</italic></source> <volume>21</volume> <fpage>599</fpage>&#x2013;<lpage>609</lpage>. <pub-id pub-id-type="doi">10.1101/gr.115592.110</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Puigb&#x00F2;</surname> <given-names>P.</given-names></name> <name><surname>Wolf</surname> <given-names>Y. I.</given-names></name> <name><surname>Koonin</surname> <given-names>E. V.</given-names></name></person-group> (<year>2010</year>). <article-title>The tree and net components of prokaryote evolution.</article-title> <source><italic>Genome Biol. Evol.</italic></source> <volume>2</volume> <fpage>745</fpage>&#x2013;<lpage>756</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evq062</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>1992</year>). <article-title>Phylogenetic inference based on matrix representation of trees.</article-title> <source><italic>Mol. Phylogenet. Evol.</italic></source> <volume>1</volume> <fpage>53</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1016/1055-7903(92)90035-F</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2001a</year>). <article-title>Detection of lateral gene transfer among microbial genomes.</article-title> <source><italic>Curr. Opin. Genet. Dev.</italic></source> <volume>11</volume> <fpage>620</fpage>&#x2013;<lpage>626</lpage>. <pub-id pub-id-type="doi">10.1016/S0959-437X(00)00244-6</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2001b</year>). <article-title>On surrogate methods for detecting lateral gene transfer.</article-title> <source><italic>FEMS Microbiol. Lett.</italic></source> <volume>201</volume> <fpage>187</fpage>&#x2013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6968.2001.tb10755.x</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ragan</surname> <given-names>M. A.</given-names></name> <name><surname>Beiko</surname> <given-names>R. G.</given-names></name></person-group> (<year>2009</year>). <article-title>Lateral genetic transfer: open issues.</article-title> <source><italic>Philos. Trans. R. Soc. Lond. B Biol. Sci.</italic></source> <volume>364</volume> <fpage>2241</fpage>&#x2013;<lpage>2251</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2009.0031</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Segerman</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <article-title>The genetic integrity of bacterial species: the core genome and the accessory genome, two different stories.</article-title> <source><italic>Front. Cell. Infect. Microbiol.</italic></source> <volume>2</volume>:<issue>116</issue>. <pub-id pub-id-type="doi">10.3389/fcimb.2012.00116</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seidman</surname> <given-names>S. B.</given-names></name></person-group> (<year>1980</year>). <article-title>Clique-like structures in directed networks.</article-title> <source><italic>J. Soc. Biol. Struct.</italic></source> <volume>3</volume> <fpage>43</fpage>&#x2013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1016/0140-1750(80)90019-6</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seidman</surname> <given-names>S. B.</given-names></name> <name><surname>Foster</surname> <given-names>B. L.</given-names></name></person-group> (<year>1978</year>). <article-title>A graph-theoretic generalization of the clique concept.</article-title> <source><italic>J. Math. Sociol.</italic></source> <volume>6</volume> <fpage>139</fpage>&#x2013;<lpage>154</lpage>. <pub-id pub-id-type="doi">10.1080/0022250X.1978.9989883</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skippington</surname> <given-names>E.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Lateral genetic transfer and the construction of genetic exchange communities.</article-title> <source><italic>FEMS Microbiol. Rev.</italic></source> <volume>35</volume> <fpage>707</fpage>&#x2013;<lpage>735</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6976.2010.00261.x</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skippington</surname> <given-names>E.</given-names></name> <name><surname>Ragan</surname> <given-names>M. A.</given-names></name></person-group> (<year>2012</year>). <article-title>Phylogeny rather than ecology or lifestyle biases the construction of <italic>Escherichia coli</italic>-<italic>Shigella</italic> genetic exchange communities.</article-title> <source><italic>Open Biol.</italic></source> <volume>2</volume> <issue>120112</issue>. <pub-id pub-id-type="doi">10.1098/rsob.120112</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tatusov</surname> <given-names>R. L.</given-names></name> <name><surname>Koonin</surname> <given-names>E. V.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1997</year>). <article-title>A genomic perspective on protein families.</article-title> <source><italic>Science</italic></source> <volume>278</volume> <fpage>631</fpage>&#x2013;<lpage>637</lpage>. <pub-id pub-id-type="doi">10.1126/science.278.5338.631</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tettelin</surname> <given-names>H.</given-names></name> <name><surname>Masignani</surname> <given-names>V.</given-names></name> <name><surname>Cieslewicz</surname> <given-names>M. J.</given-names></name> <name><surname>Donati</surname> <given-names>C.</given-names></name> <name><surname>Medini</surname> <given-names>D.</given-names></name> <name><surname>Ward</surname> <given-names>N. L.</given-names></name><etal/></person-group> (<year>2005</year>). <article-title>Genome analysis of multiple pathogenic isolates of <italic>Streptococcus agalactiae</italic>: implications for the microbial &#x201C;pan-genome&#x201D;.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>102</volume> <fpage>13950</fpage>&#x2013;<lpage>13955</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0506758102</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woese</surname> <given-names>C. R.</given-names></name> <name><surname>Gibson</surname> <given-names>J.</given-names></name> <name><surname>Fox</surname> <given-names>G. E.</given-names></name></person-group> (<year>1980</year>). <article-title>Do genealogical patterns in purple photosynthetic bacteria reflect interspecific gene transfer?</article-title> <source><italic>Nature</italic></source> <volume>283</volume> <fpage>212</fpage>&#x2013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1038/283212a0</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhaxybayeva</surname> <given-names>O.</given-names></name> <name><surname>Gogarten</surname> <given-names>J. P.</given-names></name> <name><surname>Charlebois</surname> <given-names>R. L.</given-names></name> <name><surname>Doolittle</surname> <given-names>W. F.</given-names></name> <name><surname>Papke</surname> <given-names>R. T.</given-names></name></person-group> (<year>2006</year>). <article-title>Phylogenetic analyses of cyanobacterial genomes: quantification of horizontal gene transfer events.</article-title> <source><italic>Genome Res.</italic></source> <volume>16</volume> <fpage>1099</fpage>&#x2013;<lpage>1108</lpage>. <pub-id pub-id-type="doi">10.1101/gr.5322306</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn01"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="https://grappa.eecs.utk.edu/">https://grappa.eecs.utk.edu/</ext-link></p></fn>
<fn id="fn02"><label>2</label><p><ext-link ext-link-type="uri" xlink:href="https://grappa.eecs.utk.edu/">https://grappa.eecs.utk.edu/</ext-link></p></fn>
</fn-group>
</back>
</article>