<?xml version="1.0" encoding="UTF-8" standalone="no"?> 
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2017.01178</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Phylogenetic Tracings of Proteome Size Support the Gradual Accretion of Protein Structural Domains and the Early Origin of Viruses from Primordial Cells</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Nasir</surname> <given-names>Arshan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/66060/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kim</surname> <given-names>Kyung Mo</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>Gustavo</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/54948/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Biosciences, COMSATS Institute of Information Technology</institution> <country>Islamabad, Pakistan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Evolutionary Bioinformatics Laboratory, Department of Crop Sciences, University of Illinois at Urbana-Champaign</institution> <country>Urbana, IL, United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Division of Polar Life Sciences, Korea Polar Research Institute</institution> <country>Incheon, South Korea</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Ricardo Flores, Instituto de Biolog&#x000ED;a Molecular y Celular de Plantas (CSIC), Spain</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Hendrik Huthoff, King&#x00027;s College London, United Kingdom; Carmen Hernandez, Instituto de Biolog&#x000ED;a Molecular y Celular de Plantas (CSIC), Spain</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Gustavo Caetano-Anoll&#x000E9;s <email>gca&#x00040;illinois.edu</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Virology, a section of the journal Frontiers in Microbiology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>06</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>1178</elocation-id>
<history>
<date date-type="received">
<day>27</day>
<month>03</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>06</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Nasir, Kim and Caetano-Anoll&#x000E9;s.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Nasir, Kim and Caetano-Anoll&#x000E9;s</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Untangling the origin and evolution of viruses remains a challenging proposition. We recently studied the global distribution of protein domain structures in thousands of completely sequenced viral and cellular proteomes with comparative genomics, phylogenomics, and multidimensional scaling methods. A tree of life describing the evolution of proteomes revealed viruses emerging from the base of the tree as a fourth supergroup of life. A tree of domains indicated an early origin of modern viral lineages from ancient cells that co-existed with the cellular ancestors. However, it was recently argued that the rooting of our trees and the basal placement of viruses was artifactually induced by small genome (proteome) size. Here we show that these claims arise from misunderstanding and misinterpretations of cladistic methodology. Trees are reconstructed unrooted, and thus, their topologies cannot be distorted <italic>a posteriori</italic> by the rooting methodology. Tracing proteome size in trees and multidimensional views of evolutionary relationships as well as tests of leaf stability and exclusion/inclusion of taxa demonstrated that the smallest proteomes were neither attracted toward the root nor caused any topological distortions of the trees. Simulations confirmed that taxa clustering patterns were independent of proteome size and were determined by the presence of known evolutionary relatives in data matrices, highlighting the need for broader taxon sampling in phylogeny reconstruction. Instead, phylogenetic tracings of proteome size revealed a slowdown in innovation of the structural domain vocabulary and four regimes of allometric scaling that reflected a Heaps law. These regimes explained increasing economies of scale in the evolutionary growth and accretion of kernel proteome repertoires of viruses and cellular organisms that resemble growth of human languages with limited vocabulary sizes. Results reconcile dynamic and static views of domain frequency distributions that are consistent with the axiom of spatiotemporal continuity that is tenet of evolutionary thinking.</p></abstract>
<kwd-group>
<kwd>phylogenomics</kwd>
<kwd>tree of life</kwd>
<kwd>origin of viruses</kwd>
<kwd>protein structure</kwd>
<kwd>Heaps law</kwd>
<kwd>proteome growth</kwd>
</kwd-group>
<contract-num rid="cn001">OISE-1132791</contract-num>
<contract-num rid="cn002">ILLU-802-909</contract-num>
<contract-num rid="cn002">ILLU-483-625</contract-num>
<contract-num rid="cn003">PJT200620</contract-num>
<contract-num rid="cn004">21-519/SRGP/R&#x00026;D/HEC/2014</contract-num>
<contract-sponsor id="cn001">National Science Foundation<named-content content-type="fundref-id">10.13039/100000001</named-content></contract-sponsor>
<contract-sponsor id="cn002">National Institute of Food and Agriculture<named-content content-type="fundref-id">10.13039/100005825</named-content></contract-sponsor>
<contract-sponsor id="cn003">Ministry of Oceans and Fisheries<named-content content-type="fundref-id">10.13039/501100003566</named-content></contract-sponsor>
<contract-sponsor id="cn004">Higher Education Commission, Pakistan<named-content content-type="fundref-id">10.13039/501100004681</named-content></contract-sponsor>
<counts>
<fig-count count="9"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="105"/>
<page-count count="18"/>
<word-count count="14543"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Untangling the origin and evolution of viruses is one of the most challenging questions in evolutionary biology. Two major competing scenarios have been proposed: (i) viruses are very ancient and evolved (or co-existed) prior to the origin of modern cells, and (ii) viruses evolved recently from genetic material in host cells that &#x0201C;escaped&#x0201D; cellular control and became infectious (reviewed in Claverie, <xref ref-type="bibr" rid="B18">2006</xref>; Forterre, <xref ref-type="bibr" rid="B33">2006</xref>, <xref ref-type="bibr" rid="B34">2016</xref>; Koonin et al., <xref ref-type="bibr" rid="B58">2006</xref>; Bandea, <xref ref-type="bibr" rid="B6">2009</xref>; Holmes, <xref ref-type="bibr" rid="B46">2011a</xref>; Abergel et al., <xref ref-type="bibr" rid="B2">2015</xref>; Nasir et al., <xref ref-type="bibr" rid="B81">2015</xref>). The &#x0201C;virus-early&#x0201D; vs. &#x0201C;virus-late&#x0201D; debate is central to answering some of the toughest questions in biological research such as how and when did life originate on Earth, how to define and treat viruses (are they alive?), did viruses evolve once or multiple times in evolution, and how viruses and cells interact with each other in their bid for survival. Naturally, the topic has remained contentious (Raoult and Forterre, <xref ref-type="bibr" rid="B86">2008</xref>; Claverie and Ogata, <xref ref-type="bibr" rid="B21">2009</xref>; Koonin et al., <xref ref-type="bibr" rid="B59">2009</xref>; Moreira and Lopez-Garcia, <xref ref-type="bibr" rid="B73">2009</xref>; Claverie and Abergel, <xref ref-type="bibr" rid="B19">2013</xref>, <xref ref-type="bibr" rid="B20">2016</xref>).</p>
<p>The deep evolutionary exploration of viral origins however is often impossible with traditional phylogenetic and sequence-recognition methods (e.g., BLAST) due to the relatively higher mutation rates of viral genes that can lead to mutational saturation of genomic sequences (Krupovic and Bamford, <xref ref-type="bibr" rid="B60">2011</xref>; Abrescia et al., <xref ref-type="bibr" rid="B3">2012</xref>). This is well known among structural biologists who have shown that viral lineages infecting distantly related hosts sometimes exhibit strong morphological and three-dimensional (3D) similarities in capsid and coat protein structural components of virions, even in the presence of negligible sequence similarities (Benson et al., <xref ref-type="bibr" rid="B9">2004</xref>; Abrescia et al., <xref ref-type="bibr" rid="B3">2012</xref>). We therefore embarked on a large-scale data-driven study of the origins and evolution of viruses (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>) taking full advantage of the conservation of protein structure over long evolutionary timespans (Chothia and Lesk, <xref ref-type="bibr" rid="B17">1986</xref>; Illerg&#x000E5;rd et al., <xref ref-type="bibr" rid="B49">2009</xref>; Caetano-Anoll&#x000E9;s and Nasir, <xref ref-type="bibr" rid="B15">2012</xref>; Lundin et al., <xref ref-type="bibr" rid="B69">2012</xref>). We studied the evolution of protein fold superfamilies (FSFs), as defined by the Structural Classification of Proteins (SCOP) database, which include protein domains harboring common structural cores and biochemical functions indicative of a common origin (Andreeva et al., <xref ref-type="bibr" rid="B4">2008</xref>; Fox et al., <xref ref-type="bibr" rid="B36">2014</xref>). FSF domains are not subject to the effects of non-orthologous replacement and lineage sorting by sequence polymorphisms (Philippe and Laurent, <xref ref-type="bibr" rid="B83">1998</xref>; Kim and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B54">2012</xref>). In addition, only a small proportion of FSF domains (i.e., between 0.4 and 4% in Gough, <xref ref-type="bibr" rid="B39">2005</xref>) might have experienced convergent evolution including horizontal gene transfer (HGT). FSF domains are thus evolutionarily highly conserved and represent reliable markers to explore deep evolutionary relationships (Nasir et al., <xref ref-type="bibr" rid="B76">2012a</xref>).</p>
<p>Our large-scale analysis utilized a combination of comparative genomics, phylogenomics, and multidimensional scaling methods to study the evolution of a <italic>total</italic> of 1,995 FSF domain structures in &#x0007E;11 million proteins from 5,080 proteomes sampled from 1,420 cellular organisms and 3,460 viruses from the seven known viral replicon types (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). The most parsimonious interpretation of our data strongly supported the virus-early scenario of viral evolution, indicating that viral lineages originated multiple times in evolution (i.e., in a polyphyletic manner) from ancient cells (either by primordial reduction or escape; Forterre and Krupovic, <xref ref-type="bibr" rid="B35">2012</xref>; Nasir et al., <xref ref-type="bibr" rid="B77">2012b</xref>; Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>) that predated and/or co-existed with the early ancestors of superkingdoms Archaea (A), Bacteria (B), and Eukarya (E). However, the study disfavored the possibility of viral origins prior to the &#x0201C;first cell&#x0201D; (i.e., the virus-first scenario, Koonin et al., <xref ref-type="bibr" rid="B58">2006</xref>) because viruses by definition must reproduce in an intracellular environment and because the early co-existence of viral and cellular ancestors was supported by several lines of evidence, including:</p>
<list list-type="roman-lower">
<list-item><p>A cohort of 442 <italic>universal</italic> (i.e., ABEV) FSFs out of <italic>total</italic> 1,995 (22%) that was enriched in ancient proteins associated with cell membranes and appeared first as a group in a timeline of FSFs derived from a phylogenomic tree of domains (ToD). The ABEV domains suggested an early cell-like existence in the history of modern viruses.</p></list-item>
<list-item><p>A core of 68 FSFs common to viruses infecting Archaea (i.e., archaeoviruses), Bacteria (bacterioviruses), and Eukarya (eukaryoviruses) (hereafter the V<sub><italic>abe</italic></sub> group, Table <xref ref-type="supplementary-material" rid="SM2">S1</xref>) indicating that these viral lineages existed prior to the diversification of cellular life.</p></list-item>
<list-item><p>The abundance of virus-specific proteins lacking any homologs in cellular proteomes (&#x0003E;75% putative viral ORFans) that endowed unique identity to the viral supergroup (V).</p></list-item>
<list-item><p>The reconstruction of phylogenomic trees (and networks) that placed viruses at the base of a rooted tree of life (ToL).</p></list-item>
<list-item><p>An evolutionary principal coordinate (evoPCO) analysis projecting a &#x0201C;four-domain&#x0201D; view of cellular and viral proteomes rooted in evolutionary and geological time (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>).</p></list-item>
</list>
<p>We also ruled out the virus-late scenario because it implies little or no genetic overlap among archaeoviruses, bacterioviruses, and eukaryoviruses, an assumption shown to be false by structural studies (Bamford, <xref ref-type="bibr" rid="B5">2003</xref>; Benson et al., <xref ref-type="bibr" rid="B9">2004</xref>; Abrescia et al., <xref ref-type="bibr" rid="B3">2012</xref>) and the existence of the V<sub><italic>abe</italic></sub> group of FSF domains (Table <xref ref-type="supplementary-material" rid="SM2">S1</xref>; Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>).</p>
<p>Recently, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) criticized our phylogenomic methods and the virus-early scenario claiming that the basal position of viruses in our ToLs was due to a so-called &#x0201C;small genome attraction&#x0201D; (SGA) artifact attracting viruses (and other organisms) encoding small-sized proteomes toward the base of the rooted ToLs. Two of the authors are proponents of an origin of life in Eukarya and previously reconstructed a very complex most recent universal common ancestor of life encoding &#x0007E;75% of the total protein folds known today (Harish et al., <xref ref-type="bibr" rid="B41">2013</xref>). Their proposal, which goes counter to modern evolutionary thinking, relies on an evolutionary model that penalizes protein domain gains three times over losses (3:1), violates the &#x0201C;triangle inequality&#x0201D; property of phylogenetic distances needed for valid phylogenetic optimization, and produces an &#x0201C;upside down&#x0201D; phylogeny that attracts organisms with large genomes such as plants and animals to the base of their ToL (see Kim et al., <xref ref-type="bibr" rid="B55">2014</xref> for a discussion of these shortcomings). Here, we objectively address the criticism of a proteome size-induced basal placement of viral and prokaryotic proteomes. We show that Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) confused key concepts of our phylogenomic methodology (summarized in Table <xref ref-type="table" rid="T1">1</xref>), including our rooting methodology and character polarization scheme, the meaning of &#x0201C;genome size,&#x0201D; and downplayed &#x0201C;rules of thumb&#x0201D; for taxa selection in genome content and composition-based phylogenies. Importantly, they missed the crucial fact that our phylogenomic trees are reconstructed unrooted, and thus, their topologies cannot be distorted <italic>a posteriori</italic> by the rooting methodology, as claimed by Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>). Here we make explicit that the basal placement of viral and prokaryotic proteomes in our trees represents the <italic>modus operandi</italic> of long-term evolutionary processes of gene gains and losses that result in the gradual accretion of structural domains and the collective growth of proteomes over evolutionary time. While both gains and losses frequently participate in proteome evolution, their systematic phylogenetic tracing on a ToL indicated that gains significantly outnumbered losses (80,904 gains vs. 47,848 losses in Nasir et al., <xref ref-type="bibr" rid="B78">2014b</xref>), especially in prokaryotic proteomes. Because there are several ways to gain proteins (e.g., HGT, <italic>de novo</italic> gene creation, and neo/sub-functionalization following gene duplication) relative to losing them (e.g., gene loss as a one-time irreversible event), numerically gains override losses resulting in gradual accretion of domains and proteome growth (Nasir et al., <xref ref-type="bibr" rid="B78">2014b</xref>). This complex interplay extends to the viral supergroup and results in universal scaling patterns, which are discovered by phylogenomic reconstructions but cannot be predicted by the effects of ill-defined proxies of &#x0201C;genome size.&#x0201D;</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Fact-checking the narrative of Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Fiction (Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>)</bold></th>
<th valign="top" align="left"><bold>Fact</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">&#x0201C;<italic>A re-examination</italic> of Nasir and Caetano-Anoll&#x000E9;s&#x00027; <italic>phylogenomic approach suggests that small genomes systematically distort their phylogenetic reconstructions</italic>.&#x0201D;</td>
<td valign="top" align="left">In their re-examination, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) reconstructed trees (their Figures 1, 2) without paying attention to the rooting, character polarization, and taxa sampling details of our phylogenomic methodology. To exacerbate, they added extreme examples of cellular endosymbionts that complicate the definition of valid taxa in phylogenetic reconstructions.</td>
</tr>
<tr>
<td valign="top" align="left">We &#x0201C;<italic>use a hypothetical (ancestor) pseudo-outgroup,&#x0201D; &#x0201C;a hypothetical ancestor,&#x0201D;</italic> or &#x0201C;<italic>&#x02026;an artificial &#x02018;all-zero&#x02019; taxon</italic>&#x02026;<italic>an &#x02018;all-absent&#x02019; hypothetical ancestor&#x0201D;</italic> to root the ToL, or &#x0201C;<italic>an ancestor that is assumed to be an empty set of protein domains&#x0201D;</italic> as outgroup to &#x0201C;<italic>create specific phylogenetic artifacts</italic>.&#x0201D;</td>
<td valign="top" align="left">Outgroups indicate sister taxa external to the ingroup (the taxon set being studied), which are defined <italic>a priori</italic> as being of more ancestral nature. Unless taxa describe either resurrected or <italic>in vitro</italic> evolved molecules or microbes in long-term evolution experiments (e.g., artificial phylogenies, Hillis et al., <xref ref-type="bibr" rid="B45">1992</xref>), outgroups are never ancestors. They are typically extant taxa, which are <italic>a priori</italic> assumed to form one of two separate convex groups together with the ingroup. No outgroup taxon (presumably extant, hypothetical or artificial) was ever used or defined in our study or used as an ancestor (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Furthermore, we do not combine outgroups and ancestors, an approach known to be invalid (Bryant, <xref ref-type="bibr" rid="B12">1997</xref>).</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Including the hypothetical ancestor during tree estimation amounts to a priori character polarization</italic>.&#x0201D;</td>
<td valign="top" align="left">We polarize character transformations <italic>a posteriori</italic>, empirically and most parsimoniously, and complying with Weston&#x00027;s generality criterion (Weston, <xref ref-type="bibr" rid="B100">1988</xref>, <xref ref-type="bibr" rid="B101">1994</xref>).</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Unrooted trees describe relatedness of taxa based on graded compositional similarities of characters</italic>.&#x0201D;</td>
<td valign="top" align="left">The search of tree space using maximum parsimony as an optimality criterion is defined by homology relationships manifesting in tree branches not graded compositional similarities.</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Accordingly, we can expect the &#x02018;all-zero&#x02019; ancestor to cluster among genomes (proteomes) in which the smallest number of superfamilies is present. The latter are the proteomes described by the largest number of &#x0201C;0s&#x0201D; in the data matrix.&#x0201D;</italic></td>
<td valign="top" align="left">During phylogenetic searches, we first optimize character change in unrooted trees using the Wagner algorithm (Farris, <xref ref-type="bibr" rid="B27">1970</xref>). The topology of rooted trees cannot be predicted from patterns in character state vectors of ingroup or outgroup taxa and thus cannot be affected by genome size.</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Including viruses in the analyses draws the root toward the smaller viral proteomes.&#x0201D;</italic></td>
<td valign="top" align="left">A simple node distance (<italic>nd</italic>) vs. genome size plot dispels their putative SGA artifact for viruses (Figure <xref ref-type="fig" rid="F4">4</xref>). Contrary to their claim, including viruses decreases overall tree instability (Figure <xref ref-type="fig" rid="F8">8</xref>, Table <xref ref-type="table" rid="T2">2</xref>).</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Half of the sampled proteomes were analyzed (Figures 1, 2) for computational simplicity.&#x0201D;</italic></td>
<td valign="top" align="left">They included only 16 eukaryal (not 17 as they claim), 17 archaeal, 17 bacterial, and 5-9 viral proteomes, which only represent &#x0007E;16% of our taxa and likely missed representation of key phyla/groups in their trees (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Trees are not comparable.</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>The exclusion of highly reduced &#x02018;parasitic&#x02019; proteomes appears to be inconsistent with the inclusion of viruses.&#x0201D;</italic></td>
<td valign="top" align="left">Our exclusion and inclusion of taxa followed clear rationale. Exclusion of organisms engaged in obligate cellular endosymbiosis ensured integrity of definition of taxa. Inclusion of representatives of all viral groups portrayed the entire viral supergroup, which is unified by its parasitic lifestyle.</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Small proteome size is not an irreconcilable feature of genome-tree reconstructions.&#x0201D;</italic></td>
<td valign="top" align="left">The article referred by the authors (Harish et al., <xref ref-type="bibr" rid="B41">2013</xref>) has resulted in the reconstruction of a very complex most recent common ancestor of cells encoding almost 75% of existing protein folds. Two of the authors are proponents of an origin of Eukarya (and highly complex organisms) at the base of the ToL, which goes against modern evolutionary thinking. Their phylogenomic method uses polarized characters with arbitrary transformation costs, which violate the &#x0201C;triangle inequality&#x0201D; of phylogenetic distances and are engineered to attract large genomes to the base of their trees. Their use of unrealistic evolutionary assumptions does have irreconcilable consequences for the correct reconstruction of trees (Kim et al., <xref ref-type="bibr" rid="B55">2014</xref>).</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>49 of 68 core-SFs are unique to dsDNA viruses and 32 of these are found in Mimivirus genes. The latter are known to be acquired by cell-to-virus HGT, either from the host amoeba or from bacteria that parasitize the host amoeba.&#x0201D;</italic></td>
<td valign="top" align="left">All 49 core-FSFs (i.e., V<sub><italic>abe</italic></sub> FSFs common to archaeoviruses, bacterioviruses, and eukaryoviruses) are found in mimiviruses (Table <xref ref-type="supplementary-material" rid="SM2">S1</xref>). The majority of core-FSFs are indeed commonly detected in dsDNA viruses as hitherto no RNA viruses are known to infect Archaea and are rare in Bacteria (Nasir et al., <xref ref-type="bibr" rid="B75">2014a</xref>; Koonin et al., <xref ref-type="bibr" rid="B57">2015</xref>). They further stated that core-FSFs were acquired by viruses from their cellular hosts, specifically belonging to Acanthamoeba. However, core-FSFs are by definition not restricted to dsDNA viruses of Eukarya but are widespread among archaeoviruses and bacterioviruses. The argument about possible horizontal acquisition of core FSFs from amoeba or bacterial hosts is highly speculative and goes against recent bioinformatics explorations revealing an abundance of virus-specific genes lacking cellular homologs (Daubin et al., <xref ref-type="bibr" rid="B24">2003</xref>; Cortez et al., <xref ref-type="bibr" rid="B23">2009</xref>). Furthermore, the authors do not provide any evidence to support their statements. Core-FSFs do not cross the superkingdom barrier to infect eukaryotic hosts (e.g., a total of 10,427 instances of core-FSFs were detected in bacterioviruses compared to 5,823 in eukaryoviruses, Table <xref ref-type="supplementary-material" rid="SM2">S1</xref>). Virus transfers between superkingdoms have never been observed either in nature or the laboratory (Forterre, <xref ref-type="bibr" rid="B34">2016</xref>).</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>Likewise, their supporting data and analyses seem to be biased by limited sampling and highly skewed superfamily distributions. Indeed, the data presented here undermine the inferred relative antiquity of viruses in the ToL.&#x0201D;</italic></td>
<td valign="top" align="left">To compare, our genomic dataset included 5,080 proteomes of 3,460 viruses and 1,620 cells in comparison to their inclusion of only 9 viruses and 51 cells (their Figures 1, 2). Clearly, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) performed limited sampling and explored highly skewed FSF distributions.</td>
</tr>
<tr>
<td valign="top" align="left">&#x0201C;<italic>The instability of rooting with an all-zero ancestor becomes clear when the smallest proteome in a given taxon sampling varies in the rooting experiments.&#x0201D;</italic></td>
<td valign="top" align="left">Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) misunderstood the rooting methodology, confused stability of rooting with leaf stability, and did not report tree metrics of any kind to test the validity of their trees. They wrongly labeled one of the two most basal bacteria (taxid: 262724) an as archaeon (their Figure 1B). They selected taxa with larger genomes than those we sampled (their Figure 2D). Thus, genome size cannot be the culprit of the alleged tree distortions since our trees harbor smaller genomes and are stable. Instead and unsurprisingly, their choice of adding rogue taxa destabilized their phylogenies.</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>A somehow similar table can be found in a eLetter exchange (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>)</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s2">
<title>Results and discussion</title>
<sec>
<title>A brief overview of structural phylogenomics methodology</title>
<p>There are several pre-processing steps involved in the reconstruction of rooted phylogenies to ensure maximum protection from biological and technical artifacts. First, <italic>taxa</italic> are sampled broadly while ensuring participation from each major group of organisms (and viruses) since increased taxon sampling is known to decrease phylogenetic error (Heath et al., <xref ref-type="bibr" rid="B43">2008</xref>). Taxa are distinguished by the &#x0201C;profile&#x0201D; distribution of molecular characters, which in this case represent abundance (i.e., <italic>reuse</italic>) of FSF domains in sampled taxa. Data matrices are then processed to remove group-specific FSFs (e.g., the large number of eukaryote-specific immunoglobulin FSFs lacking counterparts in prokaryotic and viral proteomes) and FSFs with zero abundance. These filtering steps reduce the data matrix to comprise only of <italic>universal</italic> (i.e., ABEV) FSFs to increase resolution in the deep branches of the ToL. Data matrices are then transformed and normalized to an alpha-numeric scale indicating 24 (or 32 or 64) possible character states (e.g., 0&#x02013;9 and A&#x02013;N) representing FSF abundances in sampled taxa. These matrices are imported into the PAUP<sup>&#x0002A;</sup> software for phylogeny reconstruction (Swofford, <xref ref-type="bibr" rid="B91">2002</xref>). During searches of tree space and <italic>prior to rooting</italic>, we optimize character changes in <italic>unrooted</italic> trees allowing for both increases and decreases in FSF abundance (e.g., see gains vs. loss tracings in Nasir et al., <xref ref-type="bibr" rid="B78">2014b</xref>). The resulting most parsimonious unrooted trees that are retained are then rooted using the Lundberg approach (Lundberg, <xref ref-type="bibr" rid="B68">1972</xref>; i.e., <italic>a posteriori</italic>), which still preserves the optimized topology. Thus, <italic>tree topology is established prior to rooting and theoretically cannot be distorted by genome size</italic> (see empirical data discussed below), which is a property of taxa (i.e., proteomes) and not individual characters (i.e., FSFs) changing on trees. In other words, our tree building methodology precludes the systematic SGA artifacts proposed by Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) because decreasing proteome size decreases the number of contributed phylogenetic characters, not how character states change during phylogenetic reconstruction.</p>
</sec>
<sec>
<title>Rooting trees of life (ToLs): outgroup vs. generality criterion</title>
<p>Contrary to the claims of Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>), our rooting approach does not involve any outgroup taxon presumably extant, hypothetical, artificial, or treated as an ancestor (see Table <xref ref-type="table" rid="T1">1</xref>). Therefore, the indirectly rooted ToLs they build using their &#x0201C;<italic>hypothetical &#x02018;all-zero&#x02019; ancestor&#x0201D;</italic> do not mimic or undermine our methods (Figures 1, 2 in Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>). Their tree searches were also conducted differently and with the undesirable property of being dependent on the location of the root. In contrast, we minimize Farris&#x00027; <italic>f</italic>-values, a measure of the goodness-of-fit of the matrix of path length distances to the matrix of original distances, which describes total pairwise homoplasy and is independent on the location of the root (Farris, <xref ref-type="bibr" rid="B28">1972</xref>).</p>
<p>To clarify, the rooting method we applied is grounded in early and well-established cladistic formalizations (Farris, <xref ref-type="bibr" rid="B27">1970</xref>; Lundberg, <xref ref-type="bibr" rid="B68">1972</xref>) and is <italic>direct</italic> because it polarizes character transformations with information solely present in ingroup taxa, distinguishing ancestral from derived character states (Figure <xref ref-type="fig" rid="F1">1</xref>). Character polarization is only applied empirically and <italic>a posteriori</italic> to root the trees: <italic>(a)</italic> considering character spread in nested branches while accounting unproblematically for homoplasy, <italic>(b)</italic> searching for the most parsimonious solutions out of the two possible polarization schemes of the ordered characters while treating homologies as taxic hypotheses, and <italic>(c)</italic> allowing both gradual and punctuated build-up of evolutionary emergence of protein structures, including gain and loss, that complies with the principle of spatiotemporal continuity, Leibniz&#x00027;s <italic>lex continui</italic> (Leibniz, <xref ref-type="bibr" rid="B64">1687</xref>). Trees are rooted using Weston&#x00027;s generality criterion (Weston, <xref ref-type="bibr" rid="B100">1988</xref>, <xref ref-type="bibr" rid="B101">1994</xref>), which states that as long as ancestral characters are preponderantly retained in descendants, ancestral character states will always be more generic than their derivatives given their nested hierarchical distribution in rooted phylogenies (Figure <xref ref-type="fig" rid="F1">1</xref>). Biologically, protein domain structures spread in evolution when genes duplicate and diversify, genomes rearrange, and genetic information is exchanged. This is a process of accumulation and retention of iterative homologies, such as serial homologs in morphology and paralogous genes in genomes (Weston, <xref ref-type="bibr" rid="B101">1994</xref>), which is global, universal and largely unaffected by proteome size. This same process is widely used to generate rooted phylogenies from paralogous gene sequences. The Lundberg method (Lundberg, <xref ref-type="bibr" rid="B68">1972</xref>), which does not attach outgroup taxa to the ingroup as Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) claim, simply enables rooting by the generality criterion (Bryant, <xref ref-type="bibr" rid="B12">1997</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Comparing the indirect outgroup comparison method of rooting trees and the direct generality criterion. Rooting involves orienting an unrooted tree and pulling down a branch that will hold the ancestor of all taxa examined. In outgroup comparison, sister (outgroup) taxa external to the study group (ingroup taxa) are identified <italic>a priori</italic> of being of ancestral origin and the branch that is closest to the ingroup pulled down. This creates a new outgroup node for rooting the phylogeny. The outgroup node adds a character state vector that includes character state <italic>o</italic>, which is diagnostic of the outgroup and is assumed to be ancestral and absent in the ingroup. Once the outgroup is made ancestral, the tree is rooted and character state <italic>i</italic> is shared and derived, making it a synapomorphy. In Weston&#x00027;s generality criterion (Weston, <xref ref-type="bibr" rid="B100">1988</xref>, <xref ref-type="bibr" rid="B101">1994</xref>), the character state distributions in the phylogeny are used to polarize character transformations. Character state <italic>z</italic> is less distributed than <italic>y</italic> within the ingroup (it is present only in a minority subset of taxa) and is considered shared and derived. The figure was modified from Bryant (<xref ref-type="bibr" rid="B13">2001</xref>).</p></caption>
<graphic xlink:href="fmicb-08-01178-g0001.tif"/>
</fig>
<p>Weston&#x00027;s rule was repeatedly validated by inverse polarization (Felsenstein, <xref ref-type="bibr" rid="B30">1983</xref>) of our ordered (Wagner) characters, which always produced suboptimal trees (e.g., Figures 3, 4 in Kim et al., <xref ref-type="bibr" rid="B55">2014</xref>). In contrast, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) did not take into account that rooting is not a neutral procedure. While the length of the most parsimonious trees is unaffected by the position of the root, making <italic>a priori</italic> polarization unnecessary (Farris, <xref ref-type="bibr" rid="B27">1970</xref>), rooting impacts the homology statements of the undirected networks (Lundberg, <xref ref-type="bibr" rid="B68">1972</xref>). &#x0201C;<italic>The length of a tree is unaffected by the position of the root but is certainly not unaffected by the inclusion of a root&#x0201D;</italic> (Brower and de Pinna, <xref ref-type="bibr" rid="B11">2012</xref>). Importantly, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) did not report tree metrics, making their tree reconstructions open to speculative interpretations. Wheeler (<xref ref-type="bibr" rid="B102">2012</xref>) made it clear: &#x0201C;<italic>For trees to participate in hypothesis testing, we must be able to evaluate them and determine their relative quality. In order to do this, we require a comparable index of merit.&#x0201D;</italic> Generally this comes in the form of a cost or some other objective function based on data and tree. &#x0201C;<italic>Without such a cost, trees are mere pictures&#x02014;&#x02018;tree-shaped-objects&#x02019; of no use to science&#x0201D;</italic> (Wheeler, <xref ref-type="bibr" rid="B102">2012</xref>).</p>
</sec>
<sec>
<title>Limitations of taxon sampling and use of Ill-defined genome size proxies</title>
<p>Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) claimed that &#x0201C;genome size&#x0201D; defined by the total number of distinct FSFs encoded by each genome (i.e., FSF occurrence that we here term FSF <italic>use</italic>) was the determinant of taxa positions in their rooted 60-taxon ToLs (representing subsets of our 368-taxon trees in Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). They argued that organisms encoding small-sized genomes clustered together leading to topological distortions and caused mixing of taxa from different superkingdoms. It is important to first note differences between the two experimental designs before we address the existence of the alleged SGA artifact:</p>
<list list-type="roman-lower">
<list-item><p><italic>Taxon sampling:</italic> Our 368-taxon ToL described evolutionary relationships of an equal number of Archaea, Bacteria, and Eukarya (34 each) and at least 5 viruses from each known viral family/order (a total of 266 viruses belonging to 87 ICTV families) (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). These trees included each major phyla/group in the same proportion that was present in the original 5,080-dataset comprising 1,620 cellular organisms and 3,460 viruses. In comparison, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) extracted 17 species each from Archaea, Bacteria, and Eukarya, and only 9 viruses from our data matrix to produce 60-taxon trees without explaining any taxon selection rationale. Absence of close-relatives in trees could lead to unrealistic and arbitrary groupings and topological distortions that increase phylogenetic error (Heath et al., <xref ref-type="bibr" rid="B43">2008</xref>), as observed in the 60-taxon trees of Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) but not in our 368-taxon trees (<bold>Figure 7</bold> in Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>) or even in Harish et al. trees (Figure <xref ref-type="supplementary-material" rid="SM1">S3</xref> in Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>) when they restored the full taxon cellular set.</p></list-item>
<list-item><p><italic>Genome size definition:</italic> Genome size cannot be defined by FSF <italic>use</italic> when exploring a putative SGA artifact because our phylogenomic data matrices build evolutionary trees from FSF <italic>reuse</italic> (i.e., abundance or redundant count of FSFs in taxa). In other words, a single FSF could be present multiple times in the same genome owing to well-known evolutionary processes such as gene duplication, amplification and HGT (Nasir et al., <xref ref-type="bibr" rid="B78">2014b</xref>), their multiplicity contributing to overall genome size. Moreover, organisms that are related by a relatively recent common ancestor will likely have similar FSF abundance profiles compared to organisms separated by large evolutionary distances (emphasizing the need for broader and inclusive taxon sampling). In addition, gene loss and reductive evolution, which can occur both in free-living and parasitic/obligate parasitic organisms (and viruses) (Dufresne et al., <xref ref-type="bibr" rid="B25">2005</xref>; McCutcheon and von Dohlen, <xref ref-type="bibr" rid="B71">2011</xref>), can decrease FSF <italic>use</italic>. The interplay between FSF <italic>use</italic> (the domain vocabulary) and FSF <italic>reuse</italic> (the proteomic use of the domain vocabulary) of <italic>total</italic> (i.e., the entire repertoire) or <italic>universal</italic> (i.e., ABEV) FSFs contributes meaningful information to our data matrices (Figure <xref ref-type="fig" rid="F2">2</xref>) and <italic>neither of the two alone can define genome size for predicting taxa placement in trees</italic>. Thus, Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) definition of genome size is ill defined.</p></list-item>
<list-item><p><italic>Universal characters:</italic> Only <italic>universal</italic> ABEV FSFs were kept in the phylogenomic data matrix for tree reconstruction purposes (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Although <italic>use</italic> and <italic>reuse</italic> of <italic>total</italic> and <italic>universal</italic> FSFs are positively and strongly correlated, indicating a link between protein fold innovation and abundance (Figure <xref ref-type="fig" rid="F2">2</xref>), there are interesting and significant differences. For example, <italic>Emiliania huxleyi</italic> encodes a total of 963 FSFs, out of which 378 (39%) are ABEV (Table <xref ref-type="supplementary-material" rid="SM3">S2</xref>). This organism has the highest number of distinct <italic>universal</italic> FSFs among all sampled eukaryotes, even greater than <italic>Mus musculus</italic> (370 FSFs) and <italic>Homo sapiens</italic> (369). However, in terms of <italic>total</italic> FSFs, <italic>E. huxleyi</italic> encodes the 11th &#x0201C;largest&#x0201D; proteome in eukaryotes harboring 963 FSFs (Table <xref ref-type="supplementary-material" rid="SM3">S2</xref>). Similarly, the bacterium <italic>Sorangium cellulosum</italic> encodes 371 distinct <italic>universal</italic> FSFs, exceeding the ABEV <italic>use</italic> of all eukaryotic proteomes except <italic>E. huxleyi</italic> (Table <xref ref-type="supplementary-material" rid="SM3">S2</xref>). Because it is the <italic>universal</italic> set, and specifically FSF <italic>reuse</italic>, that is included in the phylogenomic data matrix, defining organism genome size by <italic>total</italic> FSF <italic>use</italic> (or even ABEV <italic>use</italic>; Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>) would be incorrect. Furthermore, we observed lack of correlation between ABEV FSF <italic>use</italic> and genome size for cellular organisms (Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>), which indicates that using total FSF <italic>use</italic> as extrapolation of our <italic>universal</italic> FSF set is a misleading proxy for genome size.</p></list-item>
</list>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>FSF <italic>use</italic> (occurrence) and <italic>reuse</italic> (abundance) are strongly correlated. Scatter log-log plots reveal a strong correlation between FSF <italic>use</italic> and FSF <italic>reuse</italic> for <italic>total</italic> <bold>(A)</bold> and <italic>universal</italic> ABEV FSF <bold>(B)</bold> sets for 368-taxon trees (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Viruses (266), Archaea (34), Bacteria (34), and Eukarya (34) are colored red, black, blue, and green, respectively. Each of these supergroups has its own power law regime that complies with a four-regime Heaps law of vocabulary growth. Individual regimes are indicated with numbers and follow <italic>V</italic> &#x0007E; <italic>N</italic><sup>&#x003B2;</sup> relationships, with <italic>V</italic> representing FSF vocabulary size (<italic>use</italic>) and <italic>N</italic> representing FSF database size (<italic>reuse</italic>) in proteomes. Their fits to linear regression models using ordinary least squares and the estimation of the Heaps exponent &#x003B2; are described in Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>.</p></caption>
<graphic xlink:href="fmicb-08-01178-g0002.tif"/>
</fig>
</sec>
<sec>
<title>No &#x0201C;small genome attraction&#x0201D; (SGA) artifact</title>
<p>Our 368-taxon ToL (Figure <xref ref-type="fig" rid="F3">3A</xref>) dissected organisms and viruses into four supergroups (see also Figure 7 in Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Importantly, there was no mixing of taxa from different supergroups in the ToL despite of considerable overlap in FSF <italic>use</italic> and <italic>reuse</italic>, especially among cellular organisms (Figure <xref ref-type="fig" rid="F2">2</xref>, examples above). The ToL revealed that taxa recognized their true evolutionary relatives thanks to the complex interplay between FSF <italic>use</italic> and <italic>reuse</italic>, which acts as composite variable (e.g., an archaeon encoding 100 FSFs will still be distinguished from a bacterium encoding 100 FSFs as the two organisms will likely have different FSF <italic>reuse</italic> and will also differ in the composition of the 100-FSF set). Labeling the phylogenetic positions of the &#x0201C;smallest&#x0201D; proteomes in our trees (defined by ABEV FSF <italic>use</italic> and <italic>reuse</italic>) confirmed that the smallest genomes were not attracted toward the root. For example, among the 102-cellular taxa used in our ToL (Figure <xref ref-type="fig" rid="F3">3A</xref>), the euryarchaeote <italic>Ignicoccus hospitalis</italic> was the smallest proteome either by <italic>universal</italic> FSF <italic>use</italic> (<italic>n</italic> &#x0003D; 213 ABEV FSFs) or <italic>reuse</italic> (868). The archaeon however did not appear at the root of the cellular subtree but appeared at a rather well derived position within the archaeal subtree (Figure <xref ref-type="fig" rid="F3">3A</xref>, see the black asterisk). Even the smallest virus in our dataset (the 1.7 kb bat cyclovirus encoding a single FSF and harboring a ssDNA genome) did not appear with basal RNA viruses but clustered with its closest evolutionary relative, the Dragonfly cyclovirus at the more derived positions (Figure <xref ref-type="fig" rid="F3">3A</xref>, red asterisk). Similarly, <italic>Ashbya gossypii</italic> was the smallest eukaryotic proteome (<italic>use</italic> &#x0003D; 326 FSFs, <italic>reuse</italic> &#x0003D; 3,217 FSFs) but was not the most basal eukaryote within the eukaryal subtree (the most basal was <italic>Cyanidioschyzon merolae, use</italic> &#x0003D; 331, <italic>reuse</italic> &#x0003D; 3,507), although it appeared in basal positions (Figure <xref ref-type="fig" rid="F3">3A</xref>, green asterisk). In turn, the bacterial proteome with lowest FSF <italic>use</italic> (<italic>Lactobacillus delbrueckii</italic>, 261 FSFs) was not the smallest with FSF <italic>reuse</italic> (<italic>Aquifex aeolicus</italic>, 1,155 FSFs).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Trees of proteomes are robust and insensitive to the effects of genome size but sensitive to holobiont relationships defining taxa. <bold>(A)</bold> The single most parsimonious tree (taxa &#x0003D; 368; characters &#x0003D; 442; length &#x0003D; 45,935, retention index &#x0003D; 0.83, <italic>g</italic><sub>1</sub> &#x0003D; &#x02212;0.31) describing the evolution of 102 cellular organisms (34 each from Archaea, Bacteria, and Eukarya) and 266 viruses (sampled at least 5 viruses from each family/order) (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). The smallest proteomes for cells (<italic>I. hospitalis</italic> and <italic>A. gossypii</italic>; black and green asterisks) and viruses (bat cycloviruses; red asterisk) are indicated. The names of taxa are not shown because they would not be visible. Instead, the positions of terminals were colored according to supergroup, green (Eukarya), blue (Bacteria), black (Archaea) and red (viruses). <bold>(B)</bold> A strict consensus of two most parsimonious trees (length &#x0003D; 46,781, retention index &#x0003D; 0.83, <italic>g</italic><sub>1</sub> &#x0003D; &#x02212;19.81 and &#x02212;19.82) built using phylogenomic data from the 368 proteomes of panel <bold>(A)</bold> plus the proteomes from the two extremely reduced <italic>R. prowazekii</italic> and <italic>N. equitans</italic> (gray circles and asterisks). While no major topological distortions are observed, the consensus tree losses resolution at its base.</p></caption>
<graphic xlink:href="fmicb-08-01178-g0003.tif"/>
</fig>
<p>Importantly, topological distortions do not appear in our ToLs (Figure <xref ref-type="fig" rid="F3">3</xref>) despite FSF <italic>use-reuse</italic> value overlaps (Figure <xref ref-type="fig" rid="F2">2</xref>) negating the existence of proteome-size dependent taxa clustering. This is showcased by the observation that the addition of the extremely reduced proteomes of <italic>Rickettsia prowazeki</italic> (Bacteria, <italic>use</italic> &#x0003D; 201, <italic>reuse</italic> &#x0003D; 626) and <italic>Nanoarchaeum equitans</italic> (Archaea, <italic>use</italic> &#x0003D; 131, <italic>reuse</italic> &#x0003D; 345) that caused topological distortions and mixing of archaeal and bacteria taxa in the 60-taxon trees of Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) had no such effects on either the crown of 368-taxon trees (Figure <xref ref-type="fig" rid="F3">3B</xref>) or even when Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) restored the full taxon set of 102 cellular organisms (Figure <xref ref-type="supplementary-material" rid="SM1">S3</xref> in Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>). We emphasize that Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) did not increase sampling of viral taxa from 9 to 266. It is interesting to note that <italic>N. equitans</italic> encodes a proteome even smaller than some &#x0201C;giant&#x0201D; viruses such as <italic>Acanthamoeba polyphaga mimivirus</italic> (<italic>use</italic> &#x0003D; 149, <italic>reuse</italic> &#x0003D; 508) and <italic>Megavirus chilensis</italic> (146, 581) but does not cause any distortions by mixing with viral taxa. The exercise therefore confirms that the smallest proteomes do not &#x0201C;fight&#x0201D; for the basal positions in trees. Instead, they recognize their true evolutionary relatives during exhaustive tree optimization of information in ABEV FSF <italic>use</italic> and <italic>reuse</italic> values that are encoded in the evolutionary data matrix. The addition of <italic>R</italic>. <italic>prowazekii</italic> and <italic>N. equitans</italic> however reduced support of phylogenetic relationships at the base of our ToLs (Figure <xref ref-type="fig" rid="F3">3B</xref>), an expected outcome when adding &#x0201C;rogue&#x0201D; taxa known to assume varying positions in sets of optimal trees (Thorley and Wilkinson, <xref ref-type="bibr" rid="B95">1999</xref>). In brief, the ToLs strongly negate arbitrary groupings of taxa based on genome size.</p>
<p>To empirically demonstrate the absence of a systemic SGA artifact, we plotted the &#x0201C;node distance&#x0201D; (<italic>nd</italic>) from the root to each terminal node (i.e., taxa) of the ToL&#x02014;on a scale from 0 (most basal) to 1 (most recent)&#x02014;against ABEV FSF <italic>use</italic> and <italic>reuse</italic> of supergroup taxa (Figure <xref ref-type="fig" rid="F4">4</xref>). The <italic>nd</italic> variable describes on a relative scale how evolutionarily derived is each taxon in the tree. The plots revealed substantial scatter, especially in viruses, and genome-size independent clustering of cellular proteomes indicating an absence of systemic SGA (Figure <xref ref-type="fig" rid="F4">4</xref>). For example, despite comparable FSF <italic>use-reuse</italic> between archaeal and bacterial proteomes, bacterial proteomes occupied a similar <italic>nd</italic> range with eukaryotic proteomes albeit harboring big differences between their <italic>use-reuse</italic> values (see also different slopes between Bacteria and Eukarya in Figure <xref ref-type="fig" rid="F2">2</xref>). However, a generic tendency of increase in proteome growth (mediated by both gains and losses of FSF domains throughout the evolutionary timeline, Nasir et al., <xref ref-type="bibr" rid="B78">2014b</xref>) is obvious but reflects the strong link between protein fold innovation and abundance (i.e., FSF <italic>use-reuse</italic>) that exists for both viral and cellular proteomes and is discovered by our reconstructions. For example, many bacterial proteomes overlap archaeal proteomes in FSF <italic>use</italic> and <italic>reuse</italic>, and so do many bacterial and eukaryal proteomes (Figure <xref ref-type="fig" rid="F4">4</xref>). However, their placement in the trees is at well-derived positions and comparable to eukaryotic taxa rather than archaeal taxa with their lower <italic>nd</italic> values.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Scatter plots describe the relationship between ABEV FSF <italic>use</italic> <bold>(A)</bold> and <italic>reuse</italic> <bold>(B)</bold> and node distance (<italic>nd</italic>) for the 368-taxon ToL (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Data points for different supergroups are colored green (Eukarya), blue (Bacteria), black (Archaea) and red (viruses). The black line describes the nature of the relationship, as determined by the Locally Weighted Regression Scatter Plot Smoothing (LOWESS) method, which obtains a smoothed curve by fitting successive regression functions (<italic>q</italic> &#x0003D; 0.1, <italic>i</italic> &#x0003D; 100). The plot reveals high scatter, especially toward smaller <italic>nd</italic> values and clustering of bacterial and eukaryal taxa in the same <italic>nd</italic> range despite harboring big differences in FSF <italic>use</italic> and <italic>reuse</italic>.</p></caption>
<graphic xlink:href="fmicb-08-01178-g0004.tif"/>
</fig>
<p>Next, we performed a simple test for the existence of the alleged SGA that was inspired by the Siddal and Whiting test of the long branch attraction (LBA) artifact (Siddal and Whiting, <xref ref-type="bibr" rid="B90">1999</xref>). The test evaluates clades influenced by putative LBA by removing (for example) one of the two long branched taxa from the phylogenetic tree. Under LBA, such removals are expected to change the topology of the tree, as the branch attracted to the putative long branch is now free to occupy its correct phylogenetic position (reviewed in Bergsten, <xref ref-type="bibr" rid="B10">2005</xref>). To extrapolate this logic, if a small-sized genome attracts another small-sized genome, then removal of the offending genome will restore the attracted genome to its accurate (different) phylogenetic position on the tree. To test, we selected 2 primates and 2 ascomycetes from Eukarya, 2 Crenarchaeota and 2 Euryarchaeota from Archaea, 2 Gamma-proteobacteria and 2 Firmicutes from Bacteria, and 2 mimiviridae and 2 phycodnaviridae from viruses (the 4444 dataset). We intentionally kept organisms and viruses of known taxonomies in the data matrix to observe any topological distortions influenced by taxa removal during tree reconstructions. Taxa were labeled both by <italic>use</italic> and <italic>reuse</italic> of ABEV FSFs (Figure <xref ref-type="fig" rid="F5">5</xref>). In the first reconstruction, we recovered the four-supergroup ToL without any topological mixing (Figure <xref ref-type="fig" rid="F5">5</xref>, tree <italic>a</italic>). Remarkably, FSF <italic>use</italic> and <italic>reuse</italic> of <italic>Exiguobacterium sibiricum</italic> (Firmicute) were either comparable or significantly lower to the <italic>use</italic> and <italic>reuse</italic> of the two euryarchaeotes included in the tree (309 and 2,158 vs. 307 and 2,638 and 308 and 2,290), respectively. Still, <italic>E. sibiricum</italic> clustered with its Firmicute relative, <italic>Bacilus subtilis</italic>, with good bootstrap (BS) support (72%). Nevertheless, applying the Siddal and Whiting test, we next removed the smallest viral proteomes sequentially, <italic>Ostreococcus tauri virus 2, Ostreococcus tauri virus OsV5, Acanthamoeba polyphaga moumovirus</italic>, and <italic>Acanthamoeba polyphaga mimivirus</italic> (Figure <xref ref-type="fig" rid="F5">5</xref>, trees <italic>b</italic> through <italic>e</italic>). None of the exclusions changed either the clustering patterns or tree topology indicating that the alleged SGA did not exist and that clustering of viral and prokaryotic proteomes toward the root of the ToL resulted from character change (FSF abundance) optimization in trees, not from properties of the ill-defined genome size.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Testing the SGA artifact with the Siddal and Whiting (<xref ref-type="bibr" rid="B90">1999</xref>) approach. A single most parsimonious phylogenomic tree (<italic>a</italic>) describes the evolutionary relationships between four proteomes sampled each from viruses, Archaea, Bacteria, and Eukarya. Taxa are colored as previously described. Numbers on branches indicated BS support values (%). Single most parsimonious trees <italic>b</italic> through <italic>e</italic> were recovered after successive elimination of the smallest viral proteomes. TL, tree length; RI, retention index.</p></caption>
<graphic xlink:href="fmicb-08-01178-g0005.tif"/>
</fig>
</sec>
<sec>
<title>Taxon definitions and leaf stabilities prompt exclusion of cellular endosymbionts and inclusion of viruses in ToLs</title>
<p>Our practice of excluding cellular endosymbionts was interpreted as avoidance of genome size attraction artifacts (Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>), when in reality our intention was to exclude organisms with ill-defined hologenomes of holobiont collectives (the host and its associated organismal communities), which are known to complicate definitions of taxa (Zilber-Rosenberg and Rosenberg, <xref ref-type="bibr" rid="B104">2008</xref>; Keeling, <xref ref-type="bibr" rid="B52">2011</xref>). No such exclusion was extended to the viral supergroup since one hallmark of viruses is harboring a life cycle with strict dependence on a cellular host (see below). We previously confirmed that cellular endosymbionts and obligate parasites harbor an FSF domain repertoire that is distinct from the other members of their respective superkingdoms (Nasir et al., <xref ref-type="bibr" rid="B80">2011</xref>). Cellular organisms committed to obligate parasitism show an increase in informational domains that is sometimes offset by loss of metabolic domains. This unique signature is conserved among nearly all known endosymbionts (Nasir et al., <xref ref-type="bibr" rid="B80">2011</xref>) and distinguishes these organisms from other members of their respective superkingdom. The existence of two unique signature FSF repertoires in cellular organisms (i.e., of free-living organisms and endosymbionts) creates conflict when the two lifestyles are considered together in genome-composition phylogenies. It leads to distortions when endosymbionts from different superkingdoms cluster together irrespective of their taxonomic affiliation). In turn, there are no &#x0201C;free-living&#x0201D; viruses and this conflict does not exist in the virosphere.</p>
<p>Viruses are also different from cellular endosymbionts in their FSF composition profile (Figure <xref ref-type="fig" rid="F6">6</xref>) and hence do not cause any distortions to the cellular subtrees (Figure <xref ref-type="fig" rid="F3">3</xref>). Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) disregarded the rationale and added questionable taxa to their data matrices. These taxa were likely &#x0201C;cherry-picked&#x0201D; from extreme proteomic outliers and sometimes even outside our initial sampling (e.g., <italic>Cand</italic>. Nausia deltocephalinicola). For example, <italic>Cand</italic>. Tremblaya princeps included in their trees (Figure 2 in Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>) is part of a three-pronged endosymbiotic organismal system (McCutcheon and von Dohlen, <xref ref-type="bibr" rid="B71">2011</xref>). Its genome encodes only 55 <italic>universal</italic> FSFs. It is not considered an independent organism since it depends on its host (<italic>Planococcus citri</italic>) and its endosymbiont (<italic>Cand</italic>. Moranella endobia) to synthesize essential metabolites (L&#x000F3;pez-Madrigal et al., <xref ref-type="bibr" rid="B66">2011</xref>). Similarly, <italic>Cand</italic>. N. deltocephalinicola is an obligate endosymbiont of leafhoppers, which harbors the smallest known bacterial genome (Bennett and Moran, <xref ref-type="bibr" rid="B8">2013</xref>) and encodes only 53 <italic>universal</italic> FSFs. These extreme proteomic outliers do not bias tree reconstructions because of their genome size nor induce &#x0201C;<italic>grossly erroneous rootings,&#x0201D;</italic> as suggested by Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>). Instead, their hologenomes arise from relatively modern genomic exchanges and recruitments likely resulting from complex trade-off relationships that complicate the dissection of their evolutionary origin and their definition as single valid taxon in the phylogenetic data matrices. Phylogenetically, they represent problematic taxa that should be excluded from analysis pending further understanding of their genetic makeup. The intentional inclusion of problematic taxa is expected to generate biased reconstructions (e.g., see Wilkinson et al., <xref ref-type="bibr" rid="B103">2000</xref> for a dinosaur phylogeny example and the detection of problematic taxa with double decay analysis).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Cellular endosymbionts differ from free-living organisms and viruses in their FSF composition profiles. Annotation of FSF domains into one of the seven major functional categories (<italic>Metabolism, Information, Intracellular Processes, Extracellular Processes, Regulation, General</italic>, and <italic>Other</italic>) for archaeal, bacterial, eukaryal, and viral proteomes sampled in our study (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>) and for nine viral and three extremely reduced cellular proteomes included by Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) in their reconstructions <italic>Cand</italic>. Nausia deltocephalinicola was not part of our reconstructions (encodes only 55 <italic>universal</italic> FSFs). Obligate endosymbionts or parasites often increase the repertoire of informational FSF domains, as showcased by <italic>Cand</italic>. Tremblaya included by Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>), and for 311 other known obligate and facultative parasitic organisms in (Figure 3 in Nasir et al., <xref ref-type="bibr" rid="B80">2011</xref>). Functional scheme as defined by Christine Vogel in SUPERFAMILY database (<ext-link ext-link-type="uri" xlink:href="http://supfam.org/SUPERFAMILY/function.html">http://supfam.org/SUPERFAMILY/function.html</ext-link>). Category <italic>Other</italic> includes proteins with either unknown or viral functions. <italic>General</italic> includes proteins involved in binding to small molecules, ligands, and lipids, and structural proteins. Numbers in parenthesis indicate total number of proteomes included in the FSF profile representation.</p></caption>
<graphic xlink:href="fmicb-08-01178-g0006.tif"/>
</fig>
<p>In the absence of tree statistics, it is impossible to evaluate the effect of progressive inclusion of extremely-reduced obligate parasitic taxa on the reconstructions of Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>). We therefore performed a series of tests to determine if &#x0201C;rogue&#x0201D; taxon addition affected the support of unrooted phylogenies (Figure <xref ref-type="fig" rid="F7">7</xref>, Table <xref ref-type="supplementary-material" rid="SM4">S3</xref>). In unrooted trees, the smallest phylogenetic statement is the relationship of a quartet of leaves. When examining BS-resampled phylogenies, the frequency of alternative resolved quartets provides measures of support for the position of each leaf and the accuracy of the tree (Thorley and Wilkinson, <xref ref-type="bibr" rid="B95">1999</xref>). These BS-based leaf stability (LS) indices describe phylogenetic instabilities that often result from either insufficient samplings or conflicting data. Since the genomic census is exhaustive, the culprit of LS varying scores can be character incongruence imposed by problems in the definition of taxa and characters. An unstable leaf can lower the LS scores of the other leaves and affect the overall LS of the taxon set by either occurring in unstable quartets (direct effects) or by lowering the stability of quartets in which it does not occur (indirect effects) when there is character conflict. Figure <xref ref-type="fig" rid="F7">7A</xref> shows a 20-taxon strict consensus tree with equal representation of supergroup taxa from 2,000 BS replicates used as a control (C). BS replicates were also generated for all 5 possible permutations of the free-living <italic>Acidobacterium capsulatum</italic> control and the obligate endoparasite <italic>R. prowazekii</italic> with the taxon set of the corresponding bacterial supergroup. These replicates were used to evaluate LS measures (Figure <xref ref-type="fig" rid="F7">7B</xref>, Table <xref ref-type="supplementary-material" rid="SM4">S3</xref>). Remarkably, LS indices from <italic>R. prowazekii</italic> permutations were significantly more variable and globally lower than those of <italic>A. capsulatum</italic>, explaining the reduced support of phylogenetic relationships we observed at the base of our ToL when the obligate parasites were added (Figure <xref ref-type="fig" rid="F3">3B</xref>). Similar results were obtained when alternative tree statistics such as LS difference and LS entropy were compared (Table <xref ref-type="supplementary-material" rid="SM4">S3</xref>) indicating the potentially &#x0201C;rogue&#x0201D; <italic>R. prowazekii</italic> taxon could be excluded from tree reconstructions for better and reliable recovery of evolutionary relationships. Explicitly Agree (EA) similarity, the proportion of quartets including the leaf that are resolved and of the same type in the trees, describe the similarity of the position of leaves (Estabrook, <xref ref-type="bibr" rid="B26">1992</xref>). EA values increase with the putatively rogue <italic>R. prowazekii</italic> taxon (Table <xref ref-type="supplementary-material" rid="SM4">S3</xref>). Thus, their addition decreases leaf stability while at the same time resulting in similar leaf positions. Finally, the RogueNaRok algorithm (Aberer et al., <xref ref-type="bibr" rid="B1">2013</xref>) also indicated that the <italic>R. prowazekii</italic> taxon was rogue and was a candidate for pruning.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Obligate parasitic taxa destabilize leaves of trees. <bold>(A)</bold> Leaf stabilities (LS maximum) were calculated with RadCon (Thorley and Page, <xref ref-type="bibr" rid="B94">2000</xref>) from 2,000 unrooted BS trees. LS values are ordered in the table <bold>(A)</bold> according to the most informative strict reduced consensus (SRC) tree (33.54 bits) out of a set of 5 SRC trees, which matches the strict component consensus (consensus efficiency &#x0003D; 0.555) derived from the unrooted trees. <bold>(B)</bold> LS values are visualized as violin plots. Violin plot is a combination of the box plot (the black rectangle with white circle representing group median) and density plot on each side (yellow) reflecting data distribution. The spread of LS values was calculated for the control set (C) and all possible permutations of free-living <italic>Acidobacterium capsulatum</italic> (A1&#x02013;A5) and the obligate endoparasite <italic>R. prowazekii</italic> (R1&#x02013;R5) with individual taxa of the corresponding bacterial superkingdoms (identified with numbers following taxon labels). The density trace is plotted symmetrically around the boxplots. White circles are group medians. Asterisks are distributions significantly different from control C (Wilcoxon rank sum test, two-tailed, <italic>P</italic> &#x0003C; 0.01).</p></caption>
<graphic xlink:href="fmicb-08-01178-g0007.tif"/>
</fig>
<p>Given that the persistence of viruses as a supergroup depends on viral interactions with cellular hosts, considerations of lifestyle and taxon definition alone cannot be used to exclude viruses in phylogenomic reconstructions. Cellular dependency is a necessary condition for the propagation of all viruses (with no exceptions), which generally occurs through lysis, exocytosis and transport (Nasir et al., <xref ref-type="bibr" rid="B79">2017</xref>). Viruses can also engage in host-specific dependency and dormancy interactions via symbiosis and latency (e.g., polydnaviruses and wasps behaving as holobionts; Federici and Bigot, <xref ref-type="bibr" rid="B29">2003</xref>). However, cellular dependencies could result in viruses acting as rogue taxa in phylogenetic reconstructions. We therefore tested the impact of including viruses on the stability of ToL topologies. Figure <xref ref-type="fig" rid="F8">8</xref> shows that the reconstruction of 24-taxon unrooted BS trees with 8 taxa each for Archaea, Bacteria and Eukarya, but no viruses (the dataset 8880, Figure <xref ref-type="fig" rid="F8">8A</xref>) had LS indices that were not significantly different (LS<sub>maximum</sub>, <italic>P</italic> &#x0003D; 0.98 LS<sub>difference</sub>, <italic>P</italic> &#x0003D; 0.61; LS<sub>entropy</sub>, <italic>P</italic> &#x0003D; 0.60) from those where the most &#x0201C;stable&#x0201D; cellular organisms were replaced by 6 viral taxa to produce a balanced 4-supergroup BS set (dataset 6666, Figure <xref ref-type="fig" rid="F8">8B</xref>). Thus, LS distributions show that viruses and cellular organisms are equally stable in ToLs (Figure <xref ref-type="fig" rid="F8">8C</xref>). To further inspect the two BS tree sets, we measured taxon instability indices (TII), which compute the variation of pair-wise patristic distances between taxon pairs across all trees (Maddison and Maddison, <xref ref-type="bibr" rid="B70">2001</xref>). TII also evaluates leaf stabilities and the impact of rogue taxa (Aberer et al., <xref ref-type="bibr" rid="B1">2013</xref>). Figure <xref ref-type="fig" rid="F8">8D</xref> shows that the 8880 unrooted BS trees gain a 37% significant decrease (<italic>P</italic> &#x0003C; 0.01) in taxonomic instability by replacements with the balanced 6666 BS set (Table <xref ref-type="table" rid="T2">2</xref>). In addition, none of the viruses that were added were considered rogue taxa and candidates for pruning by the RogueNaRok algorithm (Aberer et al., <xref ref-type="bibr" rid="B1">2013</xref>). Therefore, and contrary to the claims of Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>), phylogenetic stability provides one more reason to include viruses in ToLs.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Viruses stabilize leaves of trees. <bold>(A)</bold> A single most parsimonious phylogenomic tree (length &#x0003D; 13,004, retention index &#x0003D; 0.61) reconstructed from the genomic abundance census of 442 <italic>universal</italic> FSFs (432 parsimony informative characters) in 24 proteomes selected equally from Archaea (black), Bacteria (blue), and Eukarya (green) (the 8880 dataset). The most stable taxa in each superkingdoms, as indicated by TII values (Table <xref ref-type="table" rid="T2">2</xref>), are labeled with an asterisk. <bold>(B)</bold> A single most parsimonious phylogenomic tree (length &#x0003D; 12,033, retention index &#x0003D; 0.70) reconstructed from the genomic abundance census of 442 <italic>universal</italic> FSFs (428 parsimony informative characters) in 24 proteomes selected equally from viruses (red), Archaea (black), Bacteria (blue), and Eukarya (green) after replacing the most stable cellular taxa in <bold>(A)</bold> with viruses (the 6666 dataset). <bold>(C)</bold> A comparison of various LS statistics between the 8880 and 6666 BS tree datasets, as displayed by violin plots. None of the comparisons were statistically significant (Wilcoxon rank sum test, two-tailed). <bold>(D)</bold> Comparison of TII distribution for the 8880 dataset against the 6666 dataset, as displayed by violin plots. Inclusion of viral taxa significantly reduces overall tree instability. Asterisk indicates significant mean difference (Wilcoxon rank sum test, two-tailed, <italic>P</italic> &#x0003C; 0.01).</p></caption>
<graphic xlink:href="fmicb-08-01178-g0008.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Inclusion of viral taxa decreases tree instability.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>8880</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>6666</bold></th>
<th valign="top" align="left"><bold>Decrease (%)</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Taxon</bold></th>
<th valign="top" align="center"><bold>TII</bold></th>
<th valign="top" align="left"><bold>Taxon</bold></th>
<th valign="top" align="center"><bold>TII</bold></th>
<th/>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic><bold>Acidobacterium capsulatum</bold></italic></td>
<td valign="top" align="center">339984.56</td>
<td valign="top" align="left"><italic><bold>Acanthamoeba polyphaga mimivirus</bold></italic></td>
<td valign="top" align="center">176059.76</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Archaeoglobus fulgidus</italic></td>
<td valign="top" align="center">275171.68</td>
<td valign="top" align="left"><italic>Archaeoglobus fulgidus</italic></td>
<td valign="top" align="center">276206.40</td>
<td valign="top" align="center">&#x02212;0.004</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Arcobacter butzleri</italic></td>
<td valign="top" align="center">389177.39</td>
<td valign="top" align="left"><italic>Arcobacter butzleri</italic></td>
<td valign="top" align="center">260465.25</td>
<td valign="top" align="center">33.07</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Burkholderia</italic> sp.</td>
<td valign="top" align="center">384956.57</td>
<td valign="top" align="left"><italic>Burkholderia</italic> sp.</td>
<td valign="top" align="center">202131.54</td>
<td valign="top" align="center">47.49</td>
</tr>
<tr>
<td valign="top" align="left"><italic><bold>Chlorobium phaeobacteroides</bold></italic></td>
<td valign="top" align="center">348612.99</td>
<td valign="top" align="left"><italic><bold>Bovine coronavirus</bold></italic></td>
<td valign="top" align="center">139006.79</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Daphnia pulex</italic></td>
<td valign="top" align="center">296748.86</td>
<td valign="top" align="left"><italic>Daphnia pulex</italic></td>
<td valign="top" align="center">51657.62</td>
<td valign="top" align="center">82.59</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Emiliania huxleyi</italic></td>
<td valign="top" align="center">252024.48</td>
<td valign="top" align="left"><italic>Emiliania huxleyi</italic></td>
<td valign="top" align="center">86941.27</td>
<td valign="top" align="center">65.50</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Gluconacetobacter diazotrophicus</italic></td>
<td valign="top" align="center">351756.39</td>
<td valign="top" align="left"><italic>Gluconacetobacter diazotrophicus</italic></td>
<td valign="top" align="center">208054.12</td>
<td valign="top" align="center">40.85</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Gramella forsetii</italic></td>
<td valign="top" align="center">367648.99</td>
<td valign="top" align="left"><italic>Gramella forsetii</italic></td>
<td valign="top" align="center">172381.65</td>
<td valign="top" align="center">53.11</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Haloarcula marismortui</italic></td>
<td valign="top" align="center">245672.98</td>
<td valign="top" align="left"><italic>Haloarcula marismortui</italic></td>
<td valign="top" align="center">227995.40</td>
<td valign="top" align="center">7.20</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Haloquadratum walsbyi</italic></td>
<td valign="top" align="center">244554.08</td>
<td valign="top" align="left"><italic>Haloquadratum walsbyi</italic></td>
<td valign="top" align="center">227292.68</td>
<td valign="top" align="center">7.06</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Lottia gigantea</italic></td>
<td valign="top" align="center">244638.31</td>
<td valign="top" align="left"><italic>Lottia gigantea</italic></td>
<td valign="top" align="center">51716.23</td>
<td valign="top" align="center">78.86</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Methanoculleus marisnigri</italic></td>
<td valign="top" align="center">216223.26</td>
<td valign="top" align="left"><italic>Methanoculleus marisnigri</italic></td>
<td valign="top" align="center">193300.26</td>
<td valign="top" align="center">10.60</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Methanosarcina mazei</italic></td>
<td valign="top" align="center">218019.48</td>
<td valign="top" align="left"><italic>Methanosarcina mazei</italic></td>
<td valign="top" align="center">194245.12</td>
<td valign="top" align="center">10.90</td>
</tr>
<tr>
<td valign="top" align="left"><italic><bold>Mus musculus</bold></italic></td>
<td valign="top" align="center">221980.99</td>
<td valign="top" align="left"><italic><bold>Megavirus chilensis</bold></italic></td>
<td valign="top" align="center">176092.12</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Nectria haematococca</italic></td>
<td valign="top" align="center">278079.67</td>
<td valign="top" align="left"><italic>Nectria haematococca</italic></td>
<td valign="top" align="center">115236.83</td>
<td valign="top" align="center">58.56</td>
</tr>
<tr>
<td valign="top" align="left"><italic><bold>Pan troglodytes</bold></italic></td>
<td valign="top" align="center">239186.33</td>
<td valign="top" align="left"><italic><bold>Pandoravirus dulcis</bold></italic></td>
<td valign="top" align="center">145974.04</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left"><italic><bold>Pyrococcus horikoshii</bold></italic></td>
<td valign="top" align="center">151131.73</td>
<td valign="top" align="left"><italic><bold>Pandoravirus salinus</bold></italic></td>
<td valign="top" align="center">143402.03</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Roseiflexus castenholzii</italic></td>
<td valign="top" align="center">454093.65</td>
<td valign="top" align="left"><italic>Roseiflexus castenholzii</italic></td>
<td valign="top" align="center">216056.64</td>
<td valign="top" align="center">52.42</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Sorghum bicolor</italic></td>
<td valign="top" align="center">271355.69</td>
<td valign="top" align="left"><italic>Sorghum bicolor</italic></td>
<td valign="top" align="center">102020.45</td>
<td valign="top" align="center">62.40</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Sulfolobus tokodaii</italic></td>
<td valign="top" align="center">208389.88</td>
<td valign="top" align="left"><italic>Sulfolobus tokodaii</italic></td>
<td valign="top" align="center">365974.42</td>
<td valign="top" align="center">&#x02212;75.62</td>
</tr>
<tr>
<td valign="top" align="left"><italic><bold>Thermococcus kodakarensis</bold></italic></td>
<td valign="top" align="center">151131.73</td>
<td valign="top" align="left"><italic><bold>Sweet potato chlorotic stunt virus</bold></italic></td>
<td valign="top" align="center">139006.79</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Thermosipho melanesiensis</italic></td>
<td valign="top" align="center">350181.51</td>
<td valign="top" align="left"><italic>Thermosipho melanesiensis</italic></td>
<td valign="top" align="center">308662.81</td>
<td valign="top" align="center">11.86</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Xenopus laevis</italic></td>
<td valign="top" align="center">254966.77</td>
<td valign="top" align="left"><italic>Xenopus laevis</italic></td>
<td valign="top" align="center">63294.76</td>
<td valign="top" align="center">75.18</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Comparison of TII values for the &#x0201C;8880&#x0201D; BS tree dataset with taxa comprising 8 proteomes each from Archaea, Bacteria, and Eukarya against the &#x0201C;6666&#x0201D; that includes 6 proteomes each from Archaea, Bacteria, Eukarya, and viruses. For the construction of the 6666 dataset, two most stable taxa from each of Archaea, Bacteria, and Eukarya were replaced with viral proteomes (highlighted in bold) used by Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) in their trees. The last column indicates percentage decrease when comparing TII for the 8880 against the 6666 dataset and is only meaningful for unchanged taxa in both experiments</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Multidimensional scaling challenges the SGA artifact but supports the gradual evolutionary accretion of structural domains in proteomes</title>
<p>In addition to comparative genomics and phylogenomics data matrices, the virus-early evolutionary scenario was also supported by a 3D evolutionary projection of viral and cellular proteomes treated as biological systems (Figure 8 in Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). The overall age of each system is determined by the ages of its individual component parts (FSFs, in this case) derived from a ToD describing the evolution of FSFs, which was previously linked to the geological record through a molecular clock of protein folds (Wang et al., <xref ref-type="bibr" rid="B98">2011</xref>). The evoPCO analysis combines the power of cladistics and phenetics and produces a multidimensional view of evolutionary relationships among molecular systems such as proteomes. There are two main advantages of evoPCO: (i) there is no genome-size related variable in the data matrix as FSF abundances are replaced by their relative ages (i.e., evolutionary origin of the FSFs as inferred from <italic>nd</italic> or timelines calibrated in billions of years), and (ii) the method ensures that the fundamental assumption of character independence in phylogenetic tree reconstruction remains intact (Huelsenbeck and Nielsen, <xref ref-type="bibr" rid="B48">1999</xref>; Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Figure <xref ref-type="fig" rid="F9">9A</xref> shows an evoPCO analysis plot explaining in its first three major axes 85% variability in evolutionary distances between 368 cellular and viral proteomes. The plot revealed four distinct temporal clouds of proteomes for viruses, Archaea, Bacteria and Eukarya (Figure <xref ref-type="fig" rid="F9">9</xref>) and considerable scatter, with patterns resembling those of the <italic>nd</italic> plots previously described (Figure <xref ref-type="fig" rid="F4">4</xref>). For example, the &#x0201C;<italic>Megavirales</italic>&#x0201D; group (<italic>nd</italic> &#x0003D; 0.44&#x02013;0.51) with the largest viral proteomes was clearly dissected from the main viral cloud. Its terminal placement suggests the late appearance of &#x0201C;giant viruses&#x0201D; (La Scola et al., <xref ref-type="bibr" rid="B61">2003</xref>; Philippe et al., <xref ref-type="bibr" rid="B84">2013</xref>; Legendre et al., <xref ref-type="bibr" rid="B62">2014</xref>, <xref ref-type="bibr" rid="B63">2015</xref>) in evolution. The placement of supergroup clouds in the evoPCO plot relative to the proteome of the last universal common ancestor of cells reconstructed from Kim and Caetano-Anoll&#x000E9;s (<xref ref-type="bibr" rid="B53">2011</xref>) provided time directionality in the plot, which supported the early rise of viruses, followed by Archaea and a group of Bacteria and Eukarya, in that order. This matches evolutionary patterns of the rooted ToL (Figure <xref ref-type="fig" rid="F3">3</xref>) and indicates that the topology of the ToL is not due to an artifact induced by genome size because the evoPCO plot relies exclusively on the individual ages of the universal ABEV FSFs of proteomes. These ages cannot be distorted by an SGA artifact since they are derived from a ToD, a phylogenomic tree that describes the evolution of individual structural domains.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>The space of ages of FSF structural domains reveals supergroups as distinct clouds and global evolutionary tendencies of growth in proteomes. <bold>(A)</bold> An evolutionary principal coordinate (evoPCO) analysis plot portrays in its first three axes (85% variability explained) the evolutionary distances between cellular and viral proteomes [taxa &#x0003D; 368, characters &#x0003D; 442 <italic>universal</italic> FSFs, character states &#x0003D; occurrence <sup>&#x0002A;</sup> (1&#x02212;<italic>nd</italic>)]. <bold>(B)</bold> The most important evoPCO component plotted against <italic>universal</italic> ABEV FSF <italic>reuse</italic> in logarithm scale. The reconstructed proteome of the last common ancestor of modern cells was added as reference to infer the direction of evolutionary change (Kim and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B53">2011</xref>). <italic>a</italic>, Lassa virus; <italic>b</italic>, Ancestor; c, <italic>Pandoravisus salinus</italic>; <italic>d, Pandoravirus dulcis</italic>; <italic>e, Acanthamoeba polyphaga mimivirus</italic>; <italic>f</italic>, <italic>Megavirus chilensis</italic>; <italic>g, Megavirus iba</italic>; <italic>h, Ignicoccus hospitalis</italic>; <italic>i, Haloarcula marismortui</italic>; <italic>j, Lactobacillus delbrueckii</italic>; <italic>k, Sorangium cellulosum</italic>; <italic>l, Ashbya gossypii</italic>; <italic>m, Emiliana huxleyi</italic>.</p></caption>
<graphic xlink:href="fmicb-08-01178-g0009.tif"/>
</fig>
<p>To confirm, we studied global patterns of accretion of structural domains by tracing proteome size in the evoPCO analysis plot for each major axis. Figure <xref ref-type="fig" rid="F9">9B</xref> shows the most important evoPCO component (responsible for 80% of variation) plotted against ABEV FSF <italic>reuse</italic>. We found a gradual increase of the genome size proxy as one travels in time through each axis of the temporal clouds. This confirms that the global tendencies of genome growth we have observed arise from the evolutionary accretion of novel structural domains in the protein world. This is in line with the prevalence of domain gains over domain losses derived from character state reconstructions along the branches of ToLs that describe proteome evolution (Nasir et al., <xref ref-type="bibr" rid="B78">2014b</xref>).</p>
</sec>
<sec>
<title>The heaps law of language and the evolutionary growth of proteome size</title>
<p>While linguistic metaphors have dominated molecular biology since the discovery of DNA, there are striking similarities in the complexity of natural human languages and those of protein and nucleic acid macromolecules (Searls, <xref ref-type="bibr" rid="B88">2002</xref>). This has prompted the use of linguistic theory to explain the modular makeup of proteins (Gimona, <xref ref-type="bibr" rid="B38">2006</xref>). For example, the combination of structural domains in multi-domain proteins resembles the combination of atomic linguistic units (morphemes) that form higher-level units such as words or phrases (lexemes). Remarkably, protein structure complies with a number of language laws, most prominently the Zipf law, the statistical paradigm of linguistics (Zipf, <xref ref-type="bibr" rid="B105">1949</xref>). The Zipf law is a power law that links the rank of a word with its frequency. This link can be presented as a probability density distribution <italic>P</italic>(<italic>k</italic>) &#x0007E; <italic>k</italic><sup>&#x02212;&#x003B3;</sup>, where <italic>P</italic>(<italic>k</italic>) is the probability that a word be present <italic>j</italic> times in a text and &#x003B3; is an exponent that approximates 2. The Zipf law explains patterns of occurrence of Pfam domains in proteins that match words in Shakespeare&#x00027;s <italic>Romeo and Juliet</italic> (Searls, <xref ref-type="bibr" rid="B88">2002</xref>). The law is a special case of the scale-free distribution that it explains, which pervades the rich-get-richer behavior of connections in many biological networks, including those describing metabolism and protein-protein interactions (Barab&#x000E1;si, <xref ref-type="bibr" rid="B7">2009</xref>). The Zipf law is followed by structural domains at fold and FSF levels (Qian et al., <xref ref-type="bibr" rid="B85">2001</xref>; Caetano-Anolles and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B14">2003</xref>), with &#x003B3; decay values of &#x0007E;2 for Bacteria and Archaea and &#x0007E;1.4 for Eukarya (Caetano-Anolles and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B14">2003</xref>) matching values for the English and Chinese languages, respectively (Li et al., <xref ref-type="bibr" rid="B65">2016</xref>). Domain structure is also subject to functional type laws that link two kinds of variables. The combination of domains in mutidomain proteins follows the Menzerath-Altmann (MA) law of language distilled by the motto: &#x0201C;<italic>the greater the whole, the smaller its constituents&#x0201D;</italic> (Shahzad et al., <xref ref-type="bibr" rid="B89">2015</xref>). The law governs the size of domains in proteins and expresses a diminishing return tendency associated with trade-offs between economy of matter-energy and information in domain makeup. Both, the Zipf and MA laws describe &#x0201C;principles of least effort&#x0201D; that lessen costs of communication or information in any system.</p>
<p>We now show that the FSF <italic>use</italic> and <italic>reuse</italic> plots of Figure <xref ref-type="fig" rid="F2">2</xref> comply with another important law that links language properties to time, with time expressed as accumulating innovation, the Heaps law. This law describes how vocabulary sizes (<italic>V</italic>) are concave increasing power laws of text database sizes <italic>N</italic>, with <italic>V</italic> &#x0007E; <italic>N</italic><sup>&#x003B2;</sup>, where &#x003B2; represents the Heaps exponent (Heaps, <xref ref-type="bibr" rid="B42">1978</xref>). The signature of the law is sublinear growth (&#x003B2; &#x0003C; 1), which is typical of &#x0201C;economies of scale&#x0201D; showing increasingly marginal returns for new vocabulary innovations. Note that the Heaps law can be interpreted in the context of a Zipf distribution when &#x003B2; &#x0003D; 1/&#x003B3;, that this relationship has been empirically confirmed under asymptotic conditions, that vocabulary and database size are proportional to time, and that constituents of vocabularies can be constant over centuries if they represent &#x0201C;kernel&#x0201D; words that appear with high frequency (Petersen et al., <xref ref-type="bibr" rid="B82">2012</xref>; Gerlach and Altmann, <xref ref-type="bibr" rid="B37">2013</xref>). These properties have interesting implications for proteome growth. For example, the study of deviations in tail distributions linked to the Heap law regression can estimate if a pan-genome representing a gene or domain core shared between a group of organisms will continue to expand when more genomes are explored, defining &#x0201C;open&#x0201D; or &#x0201C;closed&#x0201D; pangenomic repertoires (Tettelin et al., <xref ref-type="bibr" rid="B93">2005</xref>; Koehorst et al., <xref ref-type="bibr" rid="B56">2016</xref>). When these tail distribution deviations were offset in a study of a large body of English text, Ferrer i Cancho and Sol&#x000E9; (<xref ref-type="bibr" rid="B31">2001</xref>) discovered that the probability density function showed two scaling regimes. The steepest regime followed a Zipf law characterizing a &#x0201C;kernel&#x0201D; lexicon of frequently used words. The other regime characterized an &#x0201C;unlimited&#x0201D; lexicon of growing words of less frequent use. The two-regime Zipf distribution translates into a two-regime Heaps law with &#x003B2; exponents close to 1 for the kernel and 0.4&#x02013;0.7 for the unlimited lexicon of a number of Indo-European languages, with exponent variation reflecting differences in language organization. These regimes showcase a decreasing marginal need for new words and a slowdown (cooling) of linguistic evolution (Petersen et al., <xref ref-type="bibr" rid="B82">2012</xref>). Recent studies of languages with limited dictionary sizes such as Chinese, Japanese, and Korean (Petersen et al., <xref ref-type="bibr" rid="B82">2012</xref>; L&#x000FC; et al., <xref ref-type="bibr" rid="B67">2013</xref>) have shown multi-regime Heaps laws. A recent study shows Chinese text follows a 3-regime Heaps law with &#x003B2; scaling exponents of 1, 0.7, and 0.3 for increasing text lengths, which is explained by a stochastic feedback model of vocabulary growth driven by two probabilities, one for the reuse of frequently used words and the other for the rise of word novelties (Li et al., <xref ref-type="bibr" rid="B65">2016</xref>).</p>
<p>Remarkably, the FSF <italic>use</italic> and <italic>reuse</italic> log-log plots of Figure <xref ref-type="fig" rid="F2">2</xref> show not two but four distinct power law patterns suggestive of four regimes of slowdown of vocabulary growth, each corresponding to the proteomes of viruses, Archaea, Bacteria and Eukarya, in that order (fittings in log-log plots are shown in Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>). Table <xref ref-type="table" rid="T3">3</xref> describes how the vocabulary of <italic>total</italic> and ABEV FSF domains scales with corresponding proteomic datasets with decreasing &#x003B2;, ranging from exponents of &#x0007E;1 for viruses to approximating 0 for Eukarya (Figure <xref ref-type="fig" rid="F4">4</xref>). Thus, viral proteomes use a very ancient kernel-like vocabulary with &#x003B2; exponents of 0.81 approaching unity but not far from the second regime of languages with limited vocabularies (&#x003B2; &#x0003D; 0.7&#x02013;0.77, Petersen et al., <xref ref-type="bibr" rid="B82">2012</xref>; L&#x000FC; et al., <xref ref-type="bibr" rid="B67">2013</xref>). This ancestral kernel is then expanded successively by growing vocabularies with slowdowns in the proteomes of Archaea and Bacteria and to an extreme in the proteomes of Eukarya, as these gradually appeared in evolution. The values of &#x003B2; for the proteomes of Archaea (&#x003B2; &#x0003D; 0.36&#x02013;0.40) are not far away from those of English text corpora (&#x003B2; &#x0003D; 0.4&#x02013;0.7), such as the Gutenberg Project e-book collection (&#x003B2; &#x0003D; 0.45, Tria et al., <xref ref-type="bibr" rid="B96">2014</xref>). The values of &#x003B2; for the proteomes of Bacteria (&#x003B2; &#x0003D; 0.19&#x02013;0.26) match those of the third regime of Chinese language (Petersen et al., <xref ref-type="bibr" rid="B82">2012</xref>; Li et al., <xref ref-type="bibr" rid="B65">2016</xref>).</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Scaling exponents summarizing the Heaps law for the four distinct regimes that correspond to viruses and the cellular superkingdoms (see also Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>FSF set</bold></th>
<th valign="top" align="left"><bold>Regime</bold></th>
<th valign="top" align="center"><bold>&#x003B2;</bold></th>
<th valign="top" align="center"><bold>R<sup>2</sup></bold></th>
<th valign="top" align="center"><bold><italic>F</italic></bold></th>
<th valign="top" align="center"><bold><italic>P</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ABEV</td>
<td valign="top" align="left">1-Viruses</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">4,243</td>
<td valign="top" align="center">2.2E-16</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">2-Archaea</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">160</td>
<td valign="top" align="center">5.5E-14</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3- Bacteria</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">259</td>
<td valign="top" align="center">2.2E-16</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">4-Eukarya</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">32</td>
<td valign="top" align="center">2.8E-6</td>
</tr>
<tr>
<td valign="top" align="left">Total</td>
<td valign="top" align="left">1-Viruses</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">3,874</td>
<td valign="top" align="center">2.2E-16</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">2-Archaea</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">233</td>
<td valign="top" align="center">3.0E-16</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">3-Bacteria</td>
<td valign="top" align="center">0.26</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">182</td>
<td valign="top" align="center">9.6E-15</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">4-Eukarya</td>
<td valign="top" align="center">0.12</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">108</td>
<td valign="top" align="center">9.3E-12</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Linear relationships were tested with the F statistics and coefficients of determination (R<sup>2</sup>)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Since Figures <xref ref-type="fig" rid="F4">4</xref>, <xref ref-type="fig" rid="F9">9</xref> place proteomic growth within a temporal framework, combining those results with the growth and scaling patterns of Figure <xref ref-type="fig" rid="F2">2</xref> confirm that the dynamic process of vocabulary growth of structural domains can be described in static terms with the Heaps law, with growth of database size measured as collection of FSF abundances of proteomes. This property matches the evolutionary growth of languages, derived from the analysis of hundreds of years of text corpora (two centuries in Petersen et al., <xref ref-type="bibr" rid="B82">2012</xref>) showing the growth dynamic and the static scaling patterns of word innovation are linked. Our results also show that kernels of <italic>total</italic> and ABEV FSF vocabularies exist for the proteomes of each supergroup of life that are constant over billions of years. These kernels of FSFs frequently found in proteomes are complemented with a growing set of FSF vocabularies. However, as time progresses there is a slowdown of domain innovation that can be illustrated by the decreasing Heaps exponents of the power law regimes. This outcome probably stems from economies of scales manifesting at the molecular level, as we have shown for the combination of domains in multi-domain proteins following a MA law of decreasing returns (Shahzad et al., <xref ref-type="bibr" rid="B89">2015</xref>). It also likely results from &#x0201C;semantic compression,&#x0201D; a process of compacting vocabulary with time by reducing language heterogeneity without affecting its semantics (conveying a same message with a smaller number of words, Chomsky, <xref ref-type="bibr" rid="B16">1995</xref>; Sayood and Khalid, <xref ref-type="bibr" rid="B87">2006</xref>).</p>
<p>While there are a number of Heaps-like scaling relationships in the vocabulary of genomes that appear universal, some reflecting the scaling of number of genes in different functional categories as a function of genome size (Molina and van Nimwegen, <xref ref-type="bibr" rid="B72">2009</xref>), the link between dynamic and static properties of the models must always be confirmed with phylogenetic methods. We recently built global dynamic models for the evolution of structural domains that used birth-death differential equations with global abundances of domains as state variables without the need to capture the distribution of domains in proteomes (Tal et al., <xref ref-type="bibr" rid="B92">2016</xref>). We fitted the models to data from ToDs assuming that only transitions present in the trees were possible between fold structures and that branches emerged directly from a trunk. We found that parameters of growth of domains within FSFs (FSF <italic>reuse</italic>) and diversification of FSFs (FSF <italic>use</italic>) showed emergent biphasic patterns with opposing trends, i.e., increases in FSF innovation were always counterbalanced by decreases in growth of FSF abundance, and vice versa, with the growth of the many more recent FSFs offsetting the growth of the older FSFs (Tal et al., <xref ref-type="bibr" rid="B92">2016</xref>). Since the model is global and independent of the existence of proteomes, simulations suggest a frustrated and complex interplay of growth and diversification of domain structures in the protein world that emerges from organismal diversification but is not a consequence of proteome size. This complements the findings of proteome size mappings of evoPCO plots (Figure <xref ref-type="fig" rid="F9">9</xref>) and the links of a Heaps law with history that we have formalized.</p>
</sec>
<sec>
<title>Phylogenetic tracings support the cellular origins of viral lineages</title>
<p>Our phylogenomic tracings support the primordial cellular origin of viruses and the gradual rise of molecular diversity in proteome evolution (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). The first of the four regimes of the Heaps law (the kernel regime) corresponds to the viral group (Table <xref ref-type="table" rid="T3">3</xref>) and phylogenetic tracings confirm that scaling is historical (Figures <xref ref-type="fig" rid="F4">4</xref>, <xref ref-type="fig" rid="F9">9</xref>). Comparative genomics provides additional evidence: the remarkably large number of <italic>universal</italic> FSFs that are widespread in cellular and viral proteomes (22% of <italic>total</italic> FSFs) and harbor ancient proteins associated with cell membranes supports the ancient domain kernel. Similarly, the existence of V<sub><italic>abe</italic></sub> FSFs (<italic>n</italic> &#x0003D; 68) in archaeoviruses, bacterioviruses, and eukaryoviruses also indicates that viral lineages existed prior to cellular diversification. This pushes viral origins back to ancient cells harboring segmented RNA genomes (since viruses with these features were basal in our ToL, Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>) from which modern viral lineages originated either via &#x0201C;escape&#x0201D; or &#x0201C;reduction&#x0201D; (Hendrix et al., <xref ref-type="bibr" rid="B44">2000</xref>; Forterre, <xref ref-type="bibr" rid="B33">2006</xref>; Holmes, <xref ref-type="bibr" rid="B46">2011a</xref>; Forterre and Krupovic, <xref ref-type="bibr" rid="B35">2012</xref>), albeit the reduction scenario was relatively better supported by our data and also by the discovery of giant viruses that overlap cellular endosymbionts and parasitic species in genome and particle sizes (La Scola et al., <xref ref-type="bibr" rid="B61">2003</xref>; Philippe et al., <xref ref-type="bibr" rid="B84">2013</xref>; Legendre et al., <xref ref-type="bibr" rid="B62">2014</xref>, <xref ref-type="bibr" rid="B63">2015</xref>) evolving in a similar way (Claverie and Abergel, <xref ref-type="bibr" rid="B19">2013</xref>).</p>
<p>Historically, however, the origin of viral lineages prior to the ancestors of Archaea, Bacteria, and Eukarya has been taken with skepticism as viruses by definition must reproduce inside their cellular hosts and are tightly associated with proteins (i.e., capsids) thus requiring ribosome-encoding cells for reproduction. However, virus-early scenarios do not mean &#x0201C;virus-first&#x0201D; in evolution (as interpreted by Harish et al., <xref ref-type="bibr" rid="B40">2016</xref>), but only prior to the last universal common ancestor of modern cells (Forterre, <xref ref-type="bibr" rid="B32">2005</xref>). This ancestor itself had many cellular ancestors that should better be referred to as &#x0201C;ancient&#x0201D; or &#x0201C;primordial&#x0201D; cells. Indeed, fossil records have indicated existence of primordial cells early in evolution (Javaux et al., <xref ref-type="bibr" rid="B50">2010</xref>; Wacey et al., <xref ref-type="bibr" rid="B97">2011</xref>). In other words, a distinction between ancient and modern cells is necessary for broader understanding of virus-early scenarios and to overcome roadblocks preventing acceptance of viruses as major players in the evolutionary biology of cells. To quote Forterre (<xref ref-type="bibr" rid="B34">2016</xref>), &#x0201C;<italic>The confusion between &#x02018;cells&#x02019; and &#x02018;modern cells&#x02019; (the descendants of the last universal common ancestor) is another major drawback in discussions about the origin of viruses&#x0201D;</italic> (Forterre, <xref ref-type="bibr" rid="B34">2016</xref>). Thus, our conjecture simply triggers atypical thinking about viral origins and evolution, which may be timely given how the discovery of giant viruses has broken multiple epistemological barriers (Claverie and Abergel, <xref ref-type="bibr" rid="B20">2016</xref>).</p>
<p>Finally, viruses have been routinely considered as &#x0201C;pickpockets&#x0201D; of cellular genomes (Moreira and Lopez-Garcia, <xref ref-type="bibr" rid="B73">2009</xref>). This claim however greatly underestimates virus-cell interactions and has been challenged by several independent analyses confirming the existence of an abundance of virus-specific genes in viral lineages (Daubin et al., <xref ref-type="bibr" rid="B24">2003</xref>; Cortez et al., <xref ref-type="bibr" rid="B23">2009</xref>) and from endogenous integrated viral-like elements in cellular genomes (Katzourakis and Gifford, <xref ref-type="bibr" rid="B51">2010</xref>; Cornelis et al., <xref ref-type="bibr" rid="B22">2012</xref>) suggesting that gene flow from viruses-to-cells likely exceeds gene transfer from cells-to-viruses (reviewed by Forterre, <xref ref-type="bibr" rid="B34">2016</xref>, see also Claverie and Abergel, <xref ref-type="bibr" rid="B20">2016</xref>). In brief, our evolutionary model is biphasic in nature and reconstructs an early &#x0201C;cell-like&#x0201D; phase in viral evolution distinguished from modern viral lineages. Interestingly, the cell-like phase in viral evolution can be restored today when viruses take over cellular machinery and produce viral factories that resemble cell-like organelles (Claverie, <xref ref-type="bibr" rid="B18">2006</xref>) or when they endogenize cellular genomes either in the form of integrated elements or plasmids (Weiss, <xref ref-type="bibr" rid="B99">2006</xref>; Holmes, <xref ref-type="bibr" rid="B47">2011b</xref>).</p>
</sec>
<sec>
<title>Synthesis</title>
<p>Here we show that Harish et al. (<xref ref-type="bibr" rid="B40">2016</xref>) failed to challenge the virus-early scenario that is supported by our phylogenomic data-driven retrodictive exploration (Nasir and Caetano-Anoll&#x000E9;s, <xref ref-type="bibr" rid="B74">2015</xref>). Their claim that our rooting approach attracts the proteomes of organisms (and viruses) with small genomes to the base of rooted trees does not hold in light of our demonstrations because tree topology is established <italic>prior</italic> to rooting and character polarization. Furthermore, they asserted that our ToLs were rooted <italic>a priori</italic> with an indirect method and an outgroup taxon they interpreted as an ancestor, when in reality we root our ToLs <italic>a posteriori</italic> using a direct method that follows Weston&#x00027;s generality criterion (Weston, <xref ref-type="bibr" rid="B100">1988</xref>, <xref ref-type="bibr" rid="B101">1994</xref>). They utilized <italic>total</italic> FSF <italic>use</italic> as proxy for genome size while our phylogenomic data matrices optimize both <italic>universal</italic> FSF <italic>use</italic> and <italic>reuse</italic> during unrooted tree reconstruction. Their trees are not supported by tree metrics of any kind (in addition to several other inaccuracies) and are derived from a subset of our data matrices (representing only 16% of our taxa) that were selected (apparently) without a rationale to showcase desired topologies. In contrast, we show that proteome size tracings along historical evoPCO projections and ToLs derived from a universal biology of evolutionarily conserved protein folds not only controvert unfounded phylogenetic attractions but reveal a hidden interplay between protein fold innovation and abundance. This interplay holds true for simpler viruses and Archaea to more complex Bacteria and Eukarya. Remarkably, it materializes in a multi-regime Heap&#x00027;s law of vocabulary growth (Figure <xref ref-type="fig" rid="F2">2</xref>) that makes explicit the axiom of historical continuity that is a cornerstone of evolutionary thinking and ToL reconstruction.</p>
</sec>
</sec>
<sec sec-type="materials and methods" id="s3">
<title>Materials and methods</title>
<p>Phylogenomic data and reconstruction methods follow Nasir and Caetano-Anoll&#x000E9;s (<xref ref-type="bibr" rid="B74">2015</xref>). In brief, a census of structural domains in proteomes defined a phylogenetic data matrix of FSF <italic>reuse</italic>, which was normalized, encoded and used to build most parsimonious phylogenetic trees using PAUP<sup>&#x0002A;</sup> (Swofford, <xref ref-type="bibr" rid="B91">2002</xref>). Optimal trees were rooted using Weston&#x00027;s generality criterion implemented with the Lundberg method (Lundberg, <xref ref-type="bibr" rid="B68">1972</xref>), which polarizes character state change without specification of an outgroup or ancestor. Rogue taxa identification and TII calculations were performed using RogueNaRok (Aberer et al., <xref ref-type="bibr" rid="B1">2013</xref>). LS measurements and Explicitly Agree (EA) similarities were calculated with RadCon (Thorley and Page, <xref ref-type="bibr" rid="B94">2000</xref>). EvoPCO analysis was performed using Excel XLSTAT plugin as described in Nasir and Caetano-Anoll&#x000E9;s (<xref ref-type="bibr" rid="B74">2015</xref>). Since proteomic make up involves a collective of FSFs of different ages, we use <italic>nd</italic> values of age derived from a ToD to transform an FSF occurrence (FSF <italic>use</italic>) matrix into an FSF occurrence<sup>&#x0002A;</sup>(1&#x02212;<italic>nd</italic>) matrix. This makes it possible to study a multidimensional space of &#x0201C;reverse&#x0201D; evolutionary ages of domains without losing information of FSF of very ancient origin or introducing biases from FSF absences. Euclidean distances describing dissimilarities between proteomes were calculated and the distance matrices were used to calculate the first three principal coordinates describing maximum variability in data. These three most significant loadings described how FSF parts contributed to the history of proteome systems.</p>
</sec>
<sec id="s4">
<title>Author contributions</title>
<p>AN, KK and GCA contributed to the design, experimentation, and analysis of the study, drafted, edited, improved, and finalized the manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The reviewer CH and handling Editor declared their shared affiliation, and the handling Editor states that the process nevertheless met the standards of a fair and objective review.</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="supplementary-material" id="s5">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fmicb.2017.01178/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fmicb.2017.01178/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table1.PDF" id="SM2" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table2.XLSX" id="SM3" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table3.XLSX" id="SM4" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aberer</surname> <given-names>A. J.</given-names></name> <name><surname>Krompass</surname> <given-names>D.</given-names></name> <name><surname>Stamatakis</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Pruning rogue taxa improves phylogenetic accuracy: an efficient algorithm and webservice</article-title>. <source>Syst. Biol.</source> <volume>62</volume>, <fpage>162</fpage>&#x02013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/sys078</pub-id><pub-id pub-id-type="pmid">22962004</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abergel</surname> <given-names>C.</given-names></name> <name><surname>Legendre</surname> <given-names>M.</given-names></name> <name><surname>Claverie</surname> <given-names>J.-M.</given-names></name></person-group> (<year>2015</year>). <article-title>The rapidly expanding universe of giant viruses: Mimivirus, Pandoravirus, Pithovirus and Mollivirus</article-title>. <source>FEMS Microbiol. Rev.</source> <volume>39</volume>, <fpage>779</fpage>&#x02013;<lpage>796</lpage>. <pub-id pub-id-type="doi">10.1093/femsre/fuv037</pub-id><pub-id pub-id-type="pmid">26391910</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abrescia</surname> <given-names>N. G. A.</given-names></name> <name><surname>Bamford</surname> <given-names>D. H.</given-names></name> <name><surname>Grimes</surname> <given-names>J. M.</given-names></name> <name><surname>Stuart</surname> <given-names>D. I.</given-names></name></person-group> (<year>2012</year>). <article-title>Structure unifies the viral universe</article-title>. <source>Annu. Rev. Biochem.</source> <volume>81</volume>, <fpage>795</fpage>&#x02013;<lpage>822</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-biochem-060910-095130</pub-id><pub-id pub-id-type="pmid">22482909</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andreeva</surname> <given-names>A.</given-names></name> <name><surname>Howorth</surname> <given-names>D.</given-names></name> <name><surname>Chandonia</surname> <given-names>J. M.</given-names></name> <name><surname>Brenner</surname> <given-names>S. E.</given-names></name> <name><surname>Hubbard</surname> <given-names>T. J.</given-names></name> <name><surname>Chothia</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Data growth and its impact on the SCOP database: new developments</article-title>. <source>Nucleic Acids Res.</source> <volume>36</volume>, <fpage>D419</fpage>&#x02013;<lpage>D425</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkm993</pub-id><pub-id pub-id-type="pmid">18000004</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bamford</surname> <given-names>D. H.</given-names></name></person-group> (<year>2003</year>). <article-title>Do viruses form lineages across different domains of life?</article-title> <source>Res. Microbiol.</source> <volume>154</volume>, <fpage>231</fpage>&#x02013;<lpage>236</lpage>. <pub-id pub-id-type="doi">10.1016/S0923-2508(03)00065-2</pub-id><pub-id pub-id-type="pmid">12798226</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Bandea</surname> <given-names>C. I.</given-names></name></person-group> (<year>2009</year>). <article-title>The origin and evolution of viruses as molecular organisms</article-title>, <source>Nature Proceedings</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://precedings.nature.com/documents/3886/version/1">http://precedings.nature.com/documents/3886/version/1</ext-link></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barab&#x000E1;si</surname> <given-names>A.-L.</given-names></name></person-group> (<year>2009</year>). <article-title>Scale-free networks: A decade and beyond</article-title>. <source>Science</source> <volume>325</volume>, <fpage>412</fpage>&#x02013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1126/science.1173299</pub-id><pub-id pub-id-type="pmid">19628854</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bennett</surname> <given-names>G. M.</given-names></name> <name><surname>Moran</surname> <given-names>N. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Small, smaller, smallest: the origins and evolution of ancient dual symbioses in a Phloem-feeding insect</article-title>. <source>Genome Biol. Evol.</source> <volume>5</volume>, <fpage>1675</fpage>&#x02013;<lpage>1688</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evt118</pub-id><pub-id pub-id-type="pmid">23918810</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benson</surname> <given-names>S. D.</given-names></name> <name><surname>Bamford</surname> <given-names>J. K. H.</given-names></name> <name><surname>Bamford</surname> <given-names>D. H.</given-names></name> <name><surname>Burnett</surname> <given-names>R. M.</given-names></name></person-group> (<year>2004</year>). <article-title>Does common architecture reveal a viral lineage spanning all three domains of life?</article-title> <source>Mol. Cell</source> <volume>16</volume>, <fpage>673</fpage>&#x02013;<lpage>685</lpage>. <pub-id pub-id-type="doi">10.1016/j.molcel.2004.11.016</pub-id><pub-id pub-id-type="pmid">15574324</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergsten</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>A review of long-branch attraction</article-title>. <source>Cladistics</source> <volume>21</volume>, <fpage>163</fpage>&#x02013;<lpage>193</lpage>. <pub-id pub-id-type="doi">10.1111/j.1096-0031.2005.00059.x</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brower</surname> <given-names>A. V. Z.</given-names></name> <name><surname>de Pinna</surname> <given-names>M. C. C.</given-names></name></person-group> (<year>2012</year>). <article-title>Homology and errors</article-title>. <source>Cladistics</source> <volume>28</volume>, <fpage>529</fpage>&#x02013;<lpage>538</lpage>. <pub-id pub-id-type="doi">10.1111/j.1096-0031.2012.00398.x</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bryant</surname> <given-names>H. N.</given-names></name></person-group> (<year>1997</year>). <article-title>Hypothetical ancestors and rooting in cladistic analysis</article-title>. <source>Cladistics</source> <volume>13</volume>, <fpage>337</fpage>&#x02013;<lpage>348</lpage>. <pub-id pub-id-type="doi">10.1111/j.1096-0031.1997.tb00323.x</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bryant</surname> <given-names>H. N.</given-names></name></person-group> (<year>2001</year>). <article-title>Character polarity and the rooting of cladograms</article-title> in <source>The Character Concept in Evolutionary Biology</source>, ed <person-group person-group-type="editor"><name><surname>Wagner</surname> <given-names>G.</given-names></name></person-group> (<publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>319</fpage>&#x02013;<lpage>337</lpage>.</citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caetano-Anolles</surname> <given-names>G.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>D.</given-names></name></person-group> (<year>2003</year>). <article-title>An evolutionarily structured universe of protein architecture</article-title>. <source>Genome Res.</source> <volume>13</volume>, <fpage>1563</fpage>&#x02013;<lpage>1571</lpage>. <pub-id pub-id-type="doi">10.1101/gr.1161903</pub-id><pub-id pub-id-type="pmid">12840035</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name> <name><surname>Nasir</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Benefits of using molecular structure and abundance in phylogenomic analysis</article-title>. <source>Front. Genet.</source> <volume>3</volume>:<fpage>172</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2012.00172</pub-id><pub-id pub-id-type="pmid">22973296</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><collab>Chomsky</collab></person-group> (<year>1995</year>). <source>The Minimalist Program (Current Studies in Linguistics).</source> <publisher-loc>Cambridge</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chothia</surname> <given-names>C.</given-names></name> <name><surname>Lesk</surname> <given-names>A. M.</given-names></name></person-group> (<year>1986</year>). <article-title>The relation between the divergence of sequence and structure in proteins</article-title>. <source>EMBO J.</source> <volume>5</volume>, <fpage>823</fpage>&#x02013;<lpage>826</lpage>. <pub-id pub-id-type="pmid">3709526</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Claverie</surname> <given-names>J.-M.</given-names></name></person-group> (<year>2006</year>). <article-title>Viruses take center stage in cellular evolution</article-title>. <source>Genome Biol.</source> <volume>7</volume>:<fpage>110</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2006-7-6-110</pub-id><pub-id pub-id-type="pmid">16787527</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Claverie</surname> <given-names>J. M.</given-names></name> <name><surname>Abergel</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>Open questions about giant viruses</article-title>. <source>Adv. Virus Res.</source> <volume>85</volume>, <fpage>25</fpage>&#x02013;<lpage>56</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-408116-1.00002-1</pub-id><pub-id pub-id-type="pmid">23439023</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Claverie</surname> <given-names>J.-M.</given-names></name> <name><surname>Abergel</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>Giant viruses: the difficult breaking of multiple epistemological barriers</article-title>. <source>Stud. Hist. Philos. Biol. Biomed. Sci.</source> <volume>59</volume>, <fpage>89</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/j.shpsc.2016.02.015</pub-id><pub-id pub-id-type="pmid">26972873</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Claverie</surname> <given-names>J. M.</given-names></name> <name><surname>Ogata</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Ten good reasons not to exclude giruses from the evolutionary picture</article-title>. <source>Nat. Rev.</source> <volume>7</volume>:<fpage>615</fpage>; author reply 615. <pub-id pub-id-type="doi">10.1038/nrmicro2108-c3</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cornelis</surname> <given-names>G.</given-names></name> <name><surname>Heidmann</surname> <given-names>O.</given-names></name> <name><surname>Bernard-Stoecklin</surname> <given-names>S.</given-names></name> <name><surname>Reynaud</surname> <given-names>K.</given-names></name> <name><surname>Veron</surname> <given-names>G.</given-names></name> <name><surname>Mulot</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Ancestral capture of syncytin-Car1, a fusogenic endogenous retroviral envelope gene involved in placentation and conserved in Carnivora</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>109</volume>, <fpage>E432</fpage>&#x02013;<lpage>E441</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1115346109</pub-id><pub-id pub-id-type="pmid">22308384</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cortez</surname> <given-names>D.</given-names></name> <name><surname>Forterre</surname> <given-names>P.</given-names></name> <name><surname>Gribaldo</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>A hidden reservoir of integrative elements is the major source of recently acquired foreign genes and ORFans in archaeal and bacterial genomes</article-title>. <source>Genome Biol.</source> <volume>10</volume>:<fpage>R65</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2009-10-6-r65</pub-id><pub-id pub-id-type="pmid">19531232</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Daubin</surname> <given-names>V.</given-names></name> <name><surname>Lerat</surname> <given-names>E.</given-names></name> <name><surname>Perri&#x000E8;re</surname> <given-names>G.</given-names></name> <name><surname>Sueoka</surname> <given-names>N.</given-names></name> <name><surname>Grantham</surname> <given-names>R.</given-names></name> <name><surname>Gautier</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>The source of laterally transferred genes in bacterial genomes</article-title>. <source>Genome Biol.</source> <volume>4</volume>:<fpage>R57</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2003-4-9-r57</pub-id><pub-id pub-id-type="pmid">12952536</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dufresne</surname> <given-names>A.</given-names></name> <name><surname>Garczarek</surname> <given-names>L.</given-names></name> <name><surname>Partensky</surname> <given-names>F.</given-names></name> <name><surname>Lawrence</surname> <given-names>J.</given-names></name> <name><surname>Roth</surname> <given-names>J.</given-names></name> <name><surname>Andersson</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Accelerated evolution associated with genome reduction in a free-living prokaryote</article-title>. <source>Genome Biol.</source> <volume>6</volume>:<fpage>R14</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2005-6-2-r14</pub-id><pub-id pub-id-type="pmid">15693943</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Estabrook</surname> <given-names>G. F.</given-names></name></person-group> (<year>1992</year>). <article-title>Evaluating undirected positional congruence of individual taxa between two estimates of the phylogenetic tree for a group of taxa</article-title>. <source>Syst. Biol.</source> <volume>41</volume>:<fpage>172</fpage>. <pub-id pub-id-type="doi">10.2307/2992519</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Farris</surname> <given-names>J. S.</given-names></name></person-group> (<year>1970</year>). <article-title>Methods for computing Wagner trees</article-title>. <source>Syst. Zool.</source> <volume>19</volume>, <fpage>83</fpage>&#x02013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/19.1.83</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Farris</surname> <given-names>J. S.</given-names></name></person-group> (<year>1972</year>). <article-title>Estimating phylogenetic trees from distance matrices</article-title>. <source>Am. Nat.</source> <volume>106</volume>, <fpage>645</fpage>&#x02013;<lpage>668</lpage>. <pub-id pub-id-type="doi">10.1086/282802</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Federici</surname> <given-names>B. A.</given-names></name> <name><surname>Bigot</surname> <given-names>Y.</given-names></name></person-group> (<year>2003</year>). <article-title>Origin and evolution of polydnaviruses by symbiogenesis of insect DNA viruses in endoparasitic wasps</article-title>. <source>J. Insect. Physiol.</source> <volume>49</volume>, <fpage>419</fpage>&#x02013;<lpage>432</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-1910(03)00059-3</pub-id><pub-id pub-id-type="pmid">12770621</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Felsenstein</surname> <given-names>J.</given-names></name></person-group> (<year>1983</year>). <article-title>Methods for inferring phylogenies: a statistical view</article-title>, in <source>Numerical Taxonomy</source>, ed <person-group person-group-type="editor"><name><surname>Felsenstein</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>), <fpage>315</fpage>&#x02013;<lpage>334</lpage>.</citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferrer i Cancho</surname> <given-names>R.</given-names></name> <name><surname>Sol&#x000E9;</surname> <given-names>R. V.</given-names></name></person-group> (<year>2001</year>). <article-title>Two regimes in the frequency of words and the origins of complex lexicons: Zipf&#x00027;s law revisited</article-title>. <source>J. Quant. Linguist.</source> <volume>8</volume>, <fpage>165</fpage>&#x02013;<lpage>173</lpage>. <pub-id pub-id-type="doi">10.1076/jqul.8.3.165.4101</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Forterre</surname> <given-names>P.</given-names></name></person-group> (<year>2005</year>). <article-title>The two ages of the RNA world, and the transition to the DNA world: a story of viruses and cells</article-title>. <source>Biochimie</source> <volume>87</volume>, <fpage>793</fpage>&#x02013;<lpage>803</lpage>. <pub-id pub-id-type="doi">10.1016/j.biochi.2005.03.015</pub-id><pub-id pub-id-type="pmid">16164990</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Forterre</surname> <given-names>P.</given-names></name></person-group> (<year>2006</year>). <article-title>The origin of viruses and their possible roles in major evolutionary transitions</article-title>. <source>Virus Res.</source> <volume>117</volume>, <fpage>5</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1016/j.virusres.2006.01.010</pub-id><pub-id pub-id-type="pmid">16476498</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Forterre</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>To be or not to be alive: how recent discoveries challenge the traditional definitions of viruses and life</article-title>. <source>Stud. Hist. Philos. Biol. Biomed. Sci.</source> <volume>59</volume>, <fpage>100</fpage>&#x02013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1016/j.shpsc.2016.02.013</pub-id><pub-id pub-id-type="pmid">26996409</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Forterre</surname> <given-names>P.</given-names></name> <name><surname>Krupovic</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>The origin of virions and virocells: the escape hypothesis revisited</article-title>, in <source>Viruses Essential Agents of Life</source>, ed <person-group person-group-type="editor"><name><surname>Witzany</surname> <given-names>G.</given-names></name></person-group> (<publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>43</fpage>&#x02013;<lpage>60</lpage>.</citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fox</surname> <given-names>N. K.</given-names></name> <name><surname>Brenner</surname> <given-names>S. E.</given-names></name> <name><surname>Chandonia</surname> <given-names>J. M.</given-names></name></person-group> (<year>2014</year>). <article-title>SCOPe: structural classification of proteins&#x02013;extended, integrating SCOP and ASTRAL data and classification of new structures</article-title>. <source>Nucleic Acids Res.</source> <volume>42</volume>, <fpage>D304</fpage>&#x02013;<lpage>D309</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkt1240</pub-id><pub-id pub-id-type="pmid">24304899</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gerlach</surname> <given-names>M.</given-names></name> <name><surname>Altmann</surname> <given-names>E. G.</given-names></name></person-group> (<year>2013</year>). <article-title>Stochastic model for the vocabulary growth in natural languages</article-title>. <source>Phys. Rev. X</source> <volume>3</volume>:<fpage>021006</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevX.3.021006</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gimona</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>Protein linguistics &#x02014; a grammar for modular protein assembly?</article-title> <source>Nat. Rev. Mol. Cell Biol.</source> <volume>7</volume>, <fpage>68</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1038/nrm1785</pub-id><pub-id pub-id-type="pmid">16493414</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gough</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>Convergent evolution of domain architectures (is rare)</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>1464</fpage>&#x02013;<lpage>1471</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti204</pub-id><pub-id pub-id-type="pmid">15585523</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harish</surname> <given-names>A.</given-names></name> <name><surname>Abroi</surname> <given-names>A.</given-names></name> <name><surname>Gough</surname> <given-names>J.</given-names></name> <name><surname>Kurland</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>Did viruses evolve as a distinct supergroup from common ancestors of cells?</article-title> <source>Genome Biol. Evol.</source> <volume>8</volume>, <fpage>2474</fpage>&#x02013;<lpage>2481</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evw175</pub-id><pub-id pub-id-type="pmid">27497315</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harish</surname> <given-names>A.</given-names></name> <name><surname>Tunlid</surname> <given-names>A.</given-names></name> <name><surname>Kurland</surname> <given-names>C. G.</given-names></name></person-group> (<year>2013</year>). <article-title>Rooted phylogeny of the three superkingdoms</article-title>. <source>Biochimie</source> <volume>95</volume>, <fpage>1593</fpage>&#x02013;<lpage>1604</lpage>. <pub-id pub-id-type="doi">10.1016/j.biochi.2013.04.016</pub-id><pub-id pub-id-type="pmid">23669449</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Heaps</surname> <given-names>H. S.</given-names></name></person-group> (<year>1978</year>). <source>Information Retrieval, Computational and Theoretical Aspects.</source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Academic Press</publisher-name>.</citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heath</surname> <given-names>T. A.</given-names></name> <name><surname>Hedtke</surname> <given-names>S. M.</given-names></name> <name><surname>Hillis</surname> <given-names>D. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Taxon sampling and the accuracy of phylogenetic analyses</article-title>. <source>J. Syst. Evol.</source> <volume>46</volume>, <fpage>239</fpage>&#x02013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.3724/SP.J.1002.2008.08016</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hendrix</surname> <given-names>R. W.</given-names></name> <name><surname>Lawrence</surname> <given-names>J. G.</given-names></name> <name><surname>Hatfull</surname> <given-names>G. F.</given-names></name> <name><surname>Casjens</surname> <given-names>S.</given-names></name></person-group> (<year>2000</year>). <article-title>The origins and ongoing evolution of viruses</article-title>. <source>Trends Microbiol.</source> <volume>8</volume>, <fpage>504</fpage>&#x02013;<lpage>508</lpage>. <pub-id pub-id-type="doi">10.1016/S0966-842X(00)01863-1</pub-id><pub-id pub-id-type="pmid">11121760</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hillis</surname> <given-names>D. M.</given-names></name> <name><surname>Bull</surname> <given-names>J. J.</given-names></name> <name><surname>White</surname> <given-names>M. E.</given-names></name> <name><surname>Badgett</surname> <given-names>M. R.</given-names></name> <name><surname>Molineux</surname> <given-names>I. J.</given-names></name></person-group> (<year>1992</year>). <article-title>Experimental phylogenetics: generation of a known phylogeny</article-title>. <source>Science</source> <volume>255</volume>, <fpage>589</fpage>&#x02013;<lpage>592</lpage>. <pub-id pub-id-type="doi">10.1126/science.1736360</pub-id><pub-id pub-id-type="pmid">1736360</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holmes</surname> <given-names>E. C.</given-names></name></person-group> (<year>2011a</year>). <article-title>What does virus evolution tell us about virus origins?</article-title> <source>J. Virol.</source> <volume>85</volume>, <fpage>5247</fpage>&#x02013;<lpage>5251</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.02203-10</pub-id><pub-id pub-id-type="pmid">21450811</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holmes</surname> <given-names>E. C.</given-names></name></person-group> (<year>2011b</year>). <article-title>The evolution of endogenous viral elements</article-title>. <source>Cell Host Microbe</source> <volume>10</volume>, <fpage>368</fpage>&#x02013;<lpage>377</lpage>. <pub-id pub-id-type="doi">10.1016/j.chom.2011.09.002</pub-id><pub-id pub-id-type="pmid">22018237</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huelsenbeck</surname> <given-names>J. P.</given-names></name> <name><surname>Nielsen</surname> <given-names>R.</given-names></name></person-group> (<year>1999</year>). <article-title>Effect of nonindependent substitution on phylogenetic accuracy</article-title>. <source>Syst. Biol.</source> <volume>48</volume>, <fpage>317</fpage>&#x02013;<lpage>328</lpage>. <pub-id pub-id-type="pmid">12066710</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Illerg&#x000E5;rd</surname> <given-names>K.</given-names></name> <name><surname>Ardell</surname> <given-names>D. H.</given-names></name> <name><surname>Elofsson</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Structure is three to ten times more conserved than sequence&#x02013;a study of structural response in protein cores</article-title>. <source>Proteins</source> <volume>77</volume>, <fpage>499</fpage>&#x02013;<lpage>508</lpage>. <pub-id pub-id-type="doi">10.1002/prot.22458</pub-id><pub-id pub-id-type="pmid">19507241</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Javaux</surname> <given-names>E. J.</given-names></name> <name><surname>Marshall</surname> <given-names>C. P.</given-names></name> <name><surname>Bekker</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Organic-walled microfossils in 3.2-billion-year-old shallow-marine siliciclastic deposits</article-title>. <source>Nature</source> <volume>463</volume>, <fpage>934</fpage>&#x02013;<lpage>938</lpage>. <pub-id pub-id-type="doi">10.1038/nature08793</pub-id><pub-id pub-id-type="pmid">20139963</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katzourakis</surname> <given-names>A.</given-names></name> <name><surname>Gifford</surname> <given-names>R. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Endogenous viral elements in animal genomes</article-title>. <source>PLoS Genet.</source> <volume>6</volume>:<fpage>e1001191</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1001191</pub-id><pub-id pub-id-type="pmid">21124940</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Keeling</surname> <given-names>P. J.</given-names></name></person-group> (<year>2011</year>). <article-title>Endosymbiosis: bacteria sharing the load</article-title>. <source>Curr. Biol.</source> <volume>21</volume>, <fpage>R623</fpage>&#x02013;<lpage>R624</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2011.06.061</pub-id><pub-id pub-id-type="pmid">21855000</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <article-title>The proteomic complexity and rise of the primordial ancestor of diversified life</article-title>. <source>BMC Evol. Biol.</source> <volume>11</volume>:<fpage>140</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-11-140</pub-id><pub-id pub-id-type="pmid">21612591</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>The evolutionary history of protein fold families and proteomes confirms that the archaeal ancestor is more ancient than the ancestors of other superkingdoms</article-title>. <source>BMC Evol. Biol.</source> <volume>12</volume>:<fpage>13</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-12-13</pub-id><pub-id pub-id-type="pmid">22284070</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>The importance of using realistic evolutionary models for retrodicting proteomes</article-title>. <source>Biochimie</source> <volume>99</volume>, <fpage>129</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1016/j.biochi.2013.11.019</pub-id><pub-id pub-id-type="pmid">24316279</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koehorst</surname> <given-names>J. J.</given-names></name> <name><surname>Saccenti</surname> <given-names>E.</given-names></name> <name><surname>Schaap</surname> <given-names>P. J.</given-names></name> <name><surname>Martins dos Santos</surname> <given-names>V. A. P.</given-names></name> <name><surname>Suarez-Diez</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Protein domain architectures provide a fast, efficient and scalable alternative to sequence-based methods for comparative functional genomics</article-title>. <source>F1000Research</source> <volume>5</volume>:<fpage>1987</fpage>. <pub-id pub-id-type="doi">10.12688/f1000research.9416.1</pub-id><pub-id pub-id-type="pmid">27703668</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koonin</surname> <given-names>E. V.</given-names></name> <name><surname>Dolja</surname> <given-names>V. V.</given-names></name> <name><surname>Krupovic</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Origins and evolution of viruses of eukaryotes: the ultimate modularity</article-title>. <source>Virology</source> <volume>479</volume>, <fpage>2</fpage>&#x02013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1016/j.virol.2015.02.039</pub-id><pub-id pub-id-type="pmid">25771806</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koonin</surname> <given-names>E. V.</given-names></name> <name><surname>Senkevich</surname> <given-names>T. G.</given-names></name> <name><surname>Dolja</surname> <given-names>V. V.</given-names></name></person-group> (<year>2006</year>). <article-title>The ancient Virus World and evolution of cells</article-title>. <source>Biol. Direct</source> <volume>1</volume>:<fpage>29</fpage>. <pub-id pub-id-type="doi">10.1186/1745-6150-1-29</pub-id><pub-id pub-id-type="pmid">16984643</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koonin</surname> <given-names>E. V.</given-names></name> <name><surname>Senkevich</surname> <given-names>T. G.</given-names></name> <name><surname>Dolja</surname> <given-names>V. V.</given-names></name></person-group> (<year>2009</year>). <article-title>Compelling reasons why viruses are relevant for the origin of cells</article-title>. <source>Nat. Rev.</source> <volume>7</volume>:<fpage>615</fpage>; author reply 615. <pub-id pub-id-type="doi">10.1038/nrmicro2108-c5</pub-id><pub-id pub-id-type="pmid">19561624</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krupovic</surname> <given-names>M.</given-names></name> <name><surname>Bamford</surname> <given-names>D. H.</given-names></name></person-group> (<year>2011</year>). <article-title>Double-stranded DNA viruses: 20 families and only five different architectural principles for virion assembly</article-title>. <source>Curr. Opin. Virol.</source> <volume>1</volume>, <fpage>118</fpage>&#x02013;<lpage>124</lpage>. <pub-id pub-id-type="doi">10.1016/j.coviro.2011.06.001</pub-id><pub-id pub-id-type="pmid">22440622</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>La Scola</surname> <given-names>B.</given-names></name> <name><surname>Audic</surname> <given-names>S.</given-names></name> <name><surname>Robert</surname> <given-names>C.</given-names></name> <name><surname>Jungang</surname> <given-names>L.</given-names></name> <name><surname>de Lamballerie</surname> <given-names>X.</given-names></name> <name><surname>Drancourt</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>A giant virus in amoebae</article-title>. <source>Science</source> <volume>299</volume>:<fpage>2033</fpage>. <pub-id pub-id-type="doi">10.1126/science.1081867</pub-id><pub-id pub-id-type="pmid">12663918</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Legendre</surname> <given-names>M.</given-names></name> <name><surname>Bartoli</surname> <given-names>J.</given-names></name> <name><surname>Shmakova</surname> <given-names>L.</given-names></name> <name><surname>Jeudy</surname> <given-names>S.</given-names></name> <name><surname>Labadie</surname> <given-names>K.</given-names></name> <name><surname>Adrait</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Thirty-thousand-year-old distant relative of giant icosahedral DNA viruses with a pandoravirus morphology</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>111</volume>, <fpage>4274</fpage>&#x02013;<lpage>4279</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1320670111</pub-id><pub-id pub-id-type="pmid">24591590</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Legendre</surname> <given-names>M.</given-names></name> <name><surname>Lartigue</surname> <given-names>A.</given-names></name> <name><surname>Bertaux</surname> <given-names>L.</given-names></name> <name><surname>Jeudy</surname> <given-names>S.</given-names></name> <name><surname>Bartoli</surname> <given-names>J.</given-names></name> <name><surname>Lescot</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>In-depth study of <italic>Mollivirus sibericum</italic>, a new 30,000-y-old giant virus infecting Acanthamoeba</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>E5327</fpage>&#x02013;<lpage>E5335</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1510795112</pub-id><pub-id pub-id-type="pmid">26351664</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Leibniz</surname> <given-names>G. W.</given-names></name></person-group> (<year>1687</year>). <source>Letter to Bayle: Extrait d&#x00027;une Lettre de M. L. sur un Principe G&#x000E9;n&#x000E9;ral, utile a explication des Loix de la Nature, par la Consideration de la Sagesse Divine; pour servir de R&#x000E9;plique &#x000E0; la R&#x000E9;ponse du R. P. M. Nouvelles de la Republique des Lettres.</source> <fpage>744</fpage>&#x02013;<lpage>753</lpage>.</citation></ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Lin</surname> <given-names>R.</given-names></name> <name><surname>Bian</surname> <given-names>C.</given-names></name> <name><surname>Ma</surname> <given-names>Q. D. Y.</given-names></name> <name><surname>Ivanov</surname> <given-names>P. C.</given-names></name> <name><surname>Makse</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Model of the dynamic construction process of texts and scaling laws of words organization in language systems</article-title>. <source>PLoS ONE</source> <volume>11</volume>:<fpage>e0168971</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0168971</pub-id><pub-id pub-id-type="pmid">28006026</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>L&#x000F3;pez-Madrigal</surname> <given-names>S.</given-names></name> <name><surname>Latorre</surname> <given-names>A.</given-names></name> <name><surname>Porcar</surname> <given-names>M.</given-names></name> <name><surname>Moya</surname> <given-names>A.</given-names></name> <name><surname>Gil</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Complete genome sequence of &#x0201C;Candidatus Tremblaya princeps&#x0201D; strain PCVAL, an intriguing translational machine below the living-cell status</article-title>. <source>J. Bacteriol.</source> <volume>193</volume>, <fpage>5587</fpage>&#x02013;<lpage>5588</lpage>. <pub-id pub-id-type="doi">10.1128/J.B.05749-11</pub-id><pub-id pub-id-type="pmid">21914892</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>L&#x000FC;</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.-K.</given-names></name> <name><surname>Zhou</surname> <given-names>T.</given-names></name></person-group> (<year>2013</year>). <article-title>Deviation of Zipf&#x00027;s and Heaps&#x00027; laws in human languages with limited dictionary sizes</article-title>. <source>Sci. Rep.</source> <volume>3</volume>, <fpage>8028</fpage>&#x02013;<lpage>8033</lpage>. <pub-id pub-id-type="doi">10.1038/srep01082</pub-id><pub-id pub-id-type="pmid">23378896</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lundberg</surname> <given-names>J. G.</given-names></name></person-group> (<year>1972</year>). <article-title>Wagner networks and ancestors</article-title>. <source>Syst. Zool.</source> <volume>21</volume>, <fpage>398</fpage>&#x02013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/21.4.398</pub-id></citation></ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lundin</surname> <given-names>D.</given-names></name> <name><surname>Poole</surname> <given-names>A. M.</given-names></name> <name><surname>Sjoberg</surname> <given-names>B.-M.</given-names></name> <name><surname>Hogbom</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Use of structural phylogenetic networks for classification of the ferritin-like superfamily</article-title>. <source>J. Biol. Chem.</source> <volume>287</volume>, <fpage>20565</fpage>&#x02013;<lpage>20575</lpage>. <pub-id pub-id-type="doi">10.1074/jbc.M112.367458</pub-id><pub-id pub-id-type="pmid">22535960</pub-id></citation></ref>
<ref id="B70">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Maddison</surname> <given-names>W.</given-names></name> <name><surname>Maddison</surname> <given-names>D.</given-names></name></person-group> (<year>2001</year>). <source>Mesquite: A Modular System for Evolutionary Analysis</source>. Version 3.10. Available online at: <ext-link ext-link-type="uri" xlink:href="http://mesquiteproject.org">http://mesquiteproject.org</ext-link></citation></ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCutcheon</surname> <given-names>J. P.</given-names></name> <name><surname>von Dohlen</surname> <given-names>C. D.</given-names></name></person-group> (<year>2011</year>). <article-title>An interdependent metabolic patchwork in the nested symbiosis of mealybugs</article-title>. <source>Curr. Biol.</source> <volume>21</volume>, <fpage>1366</fpage>&#x02013;<lpage>1372</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2011.06.051</pub-id><pub-id pub-id-type="pmid">21835622</pub-id></citation></ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Molina</surname> <given-names>N.</given-names></name> <name><surname>van Nimwegen</surname> <given-names>E.</given-names></name></person-group> (<year>2009</year>). <article-title>Scaling laws in functional genome content across prokaryotic clades and lifestyles</article-title>. <source>Trends Genet.</source> <volume>25</volume>, <fpage>243</fpage>&#x02013;<lpage>247</lpage>. <pub-id pub-id-type="doi">10.1016/j.tig.2009.04.004</pub-id><pub-id pub-id-type="pmid">19457568</pub-id></citation></ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moreira</surname> <given-names>D.</given-names></name> <name><surname>Lopez-Garcia</surname> <given-names>P.</given-names></name></person-group> (<year>2009</year>). <article-title>Ten reasons to exclude viruses from the tree of life</article-title>. <source>Nat. Rev.</source> <volume>7</volume>, <fpage>306</fpage>&#x02013;<lpage>311</lpage>. <pub-id pub-id-type="doi">10.1038/nrmicro2108</pub-id><pub-id pub-id-type="pmid">19270719</pub-id></citation></ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>A phylogenomic data-driven exploration of viral origins and evolution</article-title>. <source>Sci. Adv.</source> <volume>1</volume>:<fpage>e1500527</fpage>. <pub-id pub-id-type="doi">10.1126/sciadv.1500527</pub-id><pub-id pub-id-type="pmid">26601271</pub-id></citation></ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Forterre</surname> <given-names>P.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anolles</surname> <given-names>G.</given-names></name></person-group> (<year>2014a</year>). <article-title>The distribution and impact of viral lineages in domains of life</article-title>. <source>Front. Microbiol.</source> <volume>5</volume>:<fpage>194</fpage>. <pub-id pub-id-type="doi">10.3389/fmicb.2014.00194</pub-id><pub-id pub-id-type="pmid">24817866</pub-id></citation></ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anolles</surname> <given-names>G.</given-names></name></person-group> (<year>2012a</year>). <article-title>Giant viruses coexisted with the cellular ancestors and represent a distinct supergroup along with superkingdoms Archaea, Bacteria and Eukarya</article-title>. <source>BMC Evol. Biol.</source> <volume>12</volume>:<fpage>156</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-12-156</pub-id><pub-id pub-id-type="pmid">22920653</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2012b</year>). <article-title>Viral evolution: primordial cellular origins and late adaptation to parasitism</article-title>. <source>Mob. Genet. Elements.</source> <volume>2</volume>, <fpage>247</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.4161/mge.22797</pub-id><pub-id pub-id-type="pmid">23550145</pub-id></citation></ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2014b</year>). <article-title>Global patterns of protein domain gain and loss in superkingdoms</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003452</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003452</pub-id><pub-id pub-id-type="pmid">24499935</pub-id></citation></ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Long-term evolution of viruses: a Janus-faced balance</article-title>. <source>BioEssays</source>. [Epub ahead of print]. <pub-id pub-id-type="doi">10.1002/bies.201700026</pub-id><pub-id pub-id-type="pmid">28621804</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Naeem</surname> <given-names>A.</given-names></name> <name><surname>Khan</surname> <given-names>M. J.</given-names></name> <name><surname>Lopez-Nicora</surname> <given-names>H. D.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <article-title>Annotation of protein domains reveals remarkable conservation in the functional make up of proteomes across superkingdoms</article-title>. <source>Genes (Basel).</source> <volume>2</volume>, <fpage>869</fpage>&#x02013;<lpage>911</lpage>. <pub-id pub-id-type="doi">10.3390/genes2040869</pub-id><pub-id pub-id-type="pmid">24710297</pub-id></citation></ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasir</surname> <given-names>A.</given-names></name> <name><surname>Sun</surname> <given-names>F. J.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Caetano-Anolles</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Untangling the origin of viruses and their impact on cellular evolution</article-title>. <source>Ann. N.Y. Acad. Sci.</source> <volume>1341</volume>, <fpage>61</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1111/nyas.12735</pub-id><pub-id pub-id-type="pmid">25758413</pub-id></citation></ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Petersen</surname> <given-names>A. M.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. N.</given-names></name> <name><surname>Havlin</surname> <given-names>S.</given-names></name> <name><surname>Stanley</surname> <given-names>H. E.</given-names></name> <name><surname>Perc</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Languages cool as they expand: allometric scaling and the decreasing need for new words</article-title>. <source>Sci. Rep.</source> <volume>2</volume>, <fpage>721</fpage>&#x02013;<lpage>725</lpage>. <pub-id pub-id-type="doi">10.1038/srep00943</pub-id><pub-id pub-id-type="pmid">23230508</pub-id></citation></ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Philippe</surname> <given-names>H.</given-names></name> <name><surname>Laurent</surname> <given-names>J.</given-names></name></person-group> (<year>1998</year>). <article-title>How good are deep phylogenetic trees?</article-title> <source>Curr. Opin. Genet. Dev.</source> <volume>8</volume>, <fpage>616</fpage>&#x02013;<lpage>623</lpage>. <pub-id pub-id-type="pmid">9914208</pub-id></citation></ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Philippe</surname> <given-names>N.</given-names></name> <name><surname>Legendre</surname> <given-names>M.</given-names></name> <name><surname>Doutre</surname> <given-names>G.</given-names></name> <name><surname>Cout&#x000E9;</surname> <given-names>Y.</given-names></name> <name><surname>Poirot</surname> <given-names>O.</given-names></name> <name><surname>Lescot</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Pandoraviruses: amoeba viruses with genomes up to 2.5 Mb reaching that of parasitic eukaryotes</article-title>. <source>Science</source> <volume>341</volume>, <fpage>281</fpage>&#x02013;<lpage>286</lpage>. <pub-id pub-id-type="doi">10.1126/science.1239181</pub-id><pub-id pub-id-type="pmid">23869018</pub-id></citation></ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname> <given-names>J.</given-names></name> <name><surname>Luscombe</surname> <given-names>N. M.</given-names></name> <name><surname>Gerstein</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model</article-title>. <source>J. Mol. Biol.</source> <volume>313</volume>, <fpage>673</fpage>&#x02013;<lpage>681</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.2001.5079</pub-id><pub-id pub-id-type="pmid">11697896</pub-id></citation></ref>
<ref id="B86">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Raoult</surname> <given-names>D.</given-names></name> <name><surname>Forterre</surname> <given-names>P.</given-names></name></person-group> (<year>2008</year>). <article-title>Redefining viruses: lessons from Mimivirus</article-title>. <source>Nat. Rev.</source> <volume>6</volume>, <fpage>315</fpage>&#x02013;<lpage>319</lpage>. <pub-id pub-id-type="doi">10.1038/nrmicro1858</pub-id><pub-id pub-id-type="pmid">18311164</pub-id></citation></ref>
<ref id="B87">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sayood</surname> <given-names>K.</given-names></name> <name><surname>Khalid</surname></name></person-group> (<year>2006</year>). <source>Introduction to Data Compression</source>. <publisher-loc>Waltham, MA</publisher-loc>: <publisher-name>Morgan Kaufmann</publisher-name>.</citation></ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Searls</surname> <given-names>D. B.</given-names></name></person-group> (<year>2002</year>). <article-title>The language of genes</article-title>. <source>Nature</source> <volume>420</volume>, <fpage>211</fpage>&#x02013;<lpage>217</lpage>. <pub-id pub-id-type="doi">10.1038/nature01255</pub-id><pub-id pub-id-type="pmid">12432405</pub-id></citation></ref>
<ref id="B89">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shahzad</surname> <given-names>K.</given-names></name> <name><surname>Mittenthal</surname> <given-names>J. E.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>The organization of domains in proteins obeys Menzerath-Altmann&#x00027;s law of language</article-title>. <source>BMC Syst. Biol.</source> <volume>9</volume>:<fpage>44</fpage>. <pub-id pub-id-type="doi">10.1186/s12918-015-0192-9</pub-id><pub-id pub-id-type="pmid">26260760</pub-id></citation></ref>
<ref id="B90">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siddal</surname> <given-names>M. E.</given-names></name> <name><surname>Whiting</surname> <given-names>M. F.</given-names></name></person-group> (<year>1999</year>). <article-title>Long-branch abstractions</article-title>. <source>Cladistics</source> <volume>15</volume>, <fpage>9</fpage>&#x02013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1111/j.1096-0031.1999.tb00391.x</pub-id></citation></ref>
<ref id="B91">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Swofford</surname> <given-names>D. L.</given-names></name></person-group> (<year>2002</year>). <source>Phylogenomic Analysis Using Parsimony and Other Programs (PAUP<sup>&#x0002A;</sup>) Ver 4.0b10</source>. <publisher-loc>Sunderland, MA</publisher-loc>: <publisher-name>Sinauer</publisher-name>.</citation></ref>
<ref id="B92">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tal</surname> <given-names>G.</given-names></name> <name><surname>Boca</surname> <given-names>S. M.</given-names></name> <name><surname>Mittenthal</surname> <given-names>J.</given-names></name> <name><surname>Caetano-Anoll&#x000E9;s</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>A dynamic model for the evolution of protein structure</article-title>. <source>J. Mol. Evol.</source> <volume>82</volume>, <fpage>230</fpage>&#x02013;<lpage>243</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-016-9740-1</pub-id><pub-id pub-id-type="pmid">27146880</pub-id></citation></ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tettelin</surname> <given-names>H.</given-names></name> <name><surname>Masignani</surname> <given-names>V.</given-names></name> <name><surname>Cieslewicz</surname> <given-names>M. J.</given-names></name> <name><surname>Donati</surname> <given-names>C.</given-names></name> <name><surname>Medini</surname> <given-names>D.</given-names></name> <name><surname>Ward</surname> <given-names>N. L.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Genome analysis of multiple pathogenic isolates of <italic>Streptococcus agalactiae</italic>: implications for the microbial &#x0201C;pan-genome.&#x0201D;</article-title> <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>102</volume>, <fpage>13950</fpage>&#x02013;<lpage>13955</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0506758102</pub-id><pub-id pub-id-type="pmid">16172379</pub-id></citation></ref>
<ref id="B94">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thorley</surname> <given-names>J. L.</given-names></name> <name><surname>Page</surname> <given-names>R. D.</given-names></name></person-group> (<year>2000</year>). <article-title>RadCon: phylogenetic tree comparison and consensus</article-title>. <source>Bioinformatics</source> <volume>16</volume>, <fpage>486</fpage>&#x02013;<lpage>487</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/16.5.486</pub-id><pub-id pub-id-type="pmid">10871273</pub-id></citation></ref>
<ref id="B95">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thorley</surname> <given-names>J. L.</given-names></name> <name><surname>Wilkinson</surname> <given-names>M.</given-names></name></person-group> (<year>1999</year>). <article-title>Testing the phylogenetic stability of early tetrapods</article-title>. <source>J. Theor. Biol.</source> <volume>200</volume>, <fpage>343</fpage>&#x02013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1006/jtbi.1999.0999</pub-id><pub-id pub-id-type="pmid">10527723</pub-id></citation></ref>
<ref id="B96">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tria</surname> <given-names>F.</given-names></name> <name><surname>Loreto</surname> <given-names>V.</given-names></name> <name><surname>Servedio</surname> <given-names>V. D. P.</given-names></name> <name><surname>Strogatz</surname> <given-names>S. H.</given-names></name></person-group> (<year>2014</year>). <article-title>The dynamics of correlated novelties</article-title>. <source>Sci. Rep.</source> <volume>4</volume>, <fpage>721</fpage>&#x02013;<lpage>723</lpage>. <pub-id pub-id-type="doi">10.1038/srep05890</pub-id><pub-id pub-id-type="pmid">25080941</pub-id></citation></ref>
<ref id="B97">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wacey</surname> <given-names>D.</given-names></name> <name><surname>Kilburn</surname> <given-names>M. R.</given-names></name> <name><surname>Saunders</surname> <given-names>M.</given-names></name> <name><surname>Cliff</surname> <given-names>J.</given-names></name> <name><surname>Brasier</surname> <given-names>M. D.</given-names></name></person-group> (<year>2011</year>). <article-title>Microfossils of sulphur-metabolizing cells in 3.40 billion-year-old rocks of Western Australia</article-title>. <source>Nat. Geosci</source> <volume>4</volume>, <fpage>698</fpage>&#x02013;<lpage>702</lpage>. <pub-id pub-id-type="doi">10.1038/ngeo1238</pub-id></citation></ref>
<ref id="B98">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Jiang</surname> <given-names>Y.-Y.</given-names></name> <name><surname>Kim</surname> <given-names>K. M.</given-names></name> <name><surname>Qu</surname> <given-names>G.</given-names></name> <name><surname>Ji</surname> <given-names>H.-F.</given-names></name> <name><surname>Mittenthal</surname> <given-names>J. E.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>A universal molecular clock of protein folds and its power in tracing the early history of aerobic metabolism and planet oxygenation</article-title>. <source>Mol. Biol. Evol.</source> <volume>28</volume>, <fpage>567</fpage>&#x02013;<lpage>582</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msq232</pub-id><pub-id pub-id-type="pmid">20805191</pub-id></citation></ref>
<ref id="B99">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weiss</surname> <given-names>R. A.</given-names></name></person-group> (<year>2006</year>). <article-title>The discovery of endogenous retroviruses</article-title>. <source>Retrovirology</source> <volume>3</volume>:<fpage>67</fpage>. <pub-id pub-id-type="doi">10.1186/1742-4690-3-67</pub-id><pub-id pub-id-type="pmid">17018135</pub-id></citation></ref>
<ref id="B100">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Weston</surname> <given-names>P. H.</given-names></name></person-group> (<year>1988</year>). <article-title>Indirect and direct methods in systematics</article-title>, in <source>Ontogeny and Systematics</source>, ed <person-group person-group-type="editor"><name><surname>Humphries</surname> <given-names>C. J.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Columbia University Press</publisher-name>), <fpage>27</fpage>&#x02013;<lpage>56</lpage>.</citation></ref>
<ref id="B101">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Weston</surname> <given-names>P. H.</given-names></name></person-group> (<year>1994</year>). <article-title>Methods for rooting cladistic trees</article-title>, in <source>Models in Phylogeny Reconstruction</source>, eds <person-group person-group-type="editor"><name><surname>Scotland</surname> <given-names>R. W.</given-names></name> <name><surname>Siebert</surname> <given-names>D. J.</given-names></name> <name><surname>Williams</surname> <given-names>D. M.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>), <fpage>125</fpage>&#x02013;<lpage>155</lpage>.</citation></ref>
<ref id="B102">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wheeler</surname> <given-names>W.</given-names></name></person-group> (<year>2012</year>). <source>Systematics : A Course of Lectures</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x00026; SonsWiley-Blackwell</publisher-name>.</citation></ref>
<ref id="B103">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wilkinson</surname> <given-names>M.</given-names></name> <name><surname>Thorley</surname> <given-names>J. L.</given-names></name> <name><surname>Upchurch</surname> <given-names>P.</given-names></name></person-group> (<year>2000</year>). <article-title>A chain is no stronger than its weakest link: double decay analysis of phylogenetic hypotheses</article-title>. <source>Syst. Biol.</source> <volume>49</volume>, <fpage>754</fpage>&#x02013;<lpage>776</lpage>. <pub-id pub-id-type="doi">10.1080/106351500750049815</pub-id><pub-id pub-id-type="pmid">12116438</pub-id></citation></ref>
<ref id="B104">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zilber-Rosenberg</surname> <given-names>I.</given-names></name> <name><surname>Rosenberg</surname> <given-names>E.</given-names></name></person-group> (<year>2008</year>). <article-title>Role of microorganisms in the evolution of animals and plants: the hologenome theory of evolution</article-title>. <source>FEMS Microbiol. Rev.</source> <volume>32</volume>, <fpage>723</fpage>&#x02013;<lpage>735</lpage>. <pub-id pub-id-type="doi">10.1111/j.1574-6976.2008.00123.x</pub-id><pub-id pub-id-type="pmid">18549407</pub-id></citation></ref>
<ref id="B105">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zipf</surname> <given-names>G. K.</given-names></name></person-group> (<year>1949</year>). <source>Human Behavior and the Principle of Least Effort</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Addison-Wesley Press</publisher-name>.</citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> Research was supported by grants from the National Science Foundation (OISE-1132791) and the National Institute of Food and Agriculture (ILLU-802-909 and ILLU-483-625) to GCA, from the Marine Biotechnology Program (PJT200620, Genome Analysis of Marine Organisms and Development of Functional Applications) funded by Ministry of Oceans and Fisheries, Korea to KK, and from the Higher Education Commission, Start-up Research Grant Program (Project No: 21-519/SRGP/R &#x00026; D/HEC/2014), Pakistan to AN.</p>
</fn>
</fn-group>
</back>
</article>