<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2017.02151</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Evolutionary Analysis of HIV-1 Pol Proteins Reveals Representative Residues for Viral Subtype Differentiation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Nagata</surname> <given-names>Shohei</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/466722/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Imai</surname> <given-names>Junnosuke</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Makino</surname> <given-names>Gakuto</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Tomita</surname> <given-names>Masaru</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2503/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Kanai</surname> <given-names>Akio</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/17800/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute for Advanced Biosciences, Keio University</institution>, <addr-line>Tsuruoka</addr-line>, <country>Japan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Faculty of Environment and Information Studies, Keio University</institution>, <addr-line>Fujisawa</addr-line>, <country>Japan</country></aff>
<aff id="aff3"><sup>3</sup><institution>Systems Biology Program, Graduate School of Media and Governance, Keio University</institution>, <addr-line>Fujisawa</addr-line>, <country>Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Guenther Witzany, Telos - Philosophische Praxis, Austria</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Hirotaka Ode, Nagoya Medical Center (NHO), Japan; Masako Nomaguchi, Tokushima University Graduate School of Medical Sciences, Japan</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Akio Kanai <email>akio&#x00040;sfc.keio.ac.jp</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Virology, a section of the journal Frontiers in Microbiology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>11</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>2151</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>08</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>10</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Nagata, Imai, Makino, Tomita and Kanai.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Nagata, Imai, Makino, Tomita and Kanai</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>RNA viruses have been used as model systems to understand the patterns and processes of molecular evolution because they have high mutation rates and are genetically diverse. <italic>Human immunodeficiency virus 1</italic> (HIV-1), the etiological agent of acquired immune deficiency syndrome, is highly genetically diverse, and is classified into several groups and subtypes. However, it has been difficult to use its diverse sequences to establish the overall phylogenetic relationships of different strains or the trends in sequence conservation with the construction of phylogenetic trees. Our aims were to systematically characterize HIV-1 subtype evolution and to identify the regions responsible for HIV-1 subtype differentiation at the amino acid level in the Pol protein, which is often used to classify the HIV-1 subtypes. In this study, we systematically characterized the mutation sites in 2,052 Pol proteins from HIV-1 group M (144 subtype A; 1,528 subtype B; 380 subtype C), using sequence similarity networks. We also used spectral clustering to group the sequences based on the network graph structures. A stepwise analysis of the cluster hierarchies allowed us to estimate a possible evolutionary pathway for the Pol proteins. The subtype A sequences also clustered according to when and where the viruses were isolated, whereas both the subtype B and C sequences remained as single clusters. Because the Pol protein has several functional domains, we identified the regions that are discriminative by comparing the structures of the domain-based networks. Our results suggest that sequence changes in the RNase H domain and the reverse transcriptase (RT) connection domain are responsible for the subtype classification. By analyzing the different amino acid compositions at each site in both domain sequences, we found that a few specific amino acid residues (i.e., M357 in the RT connection domain and Q480, Y483, and L491 in the RNase H domain) represent the differences among the subtypes. These residues were located on the surface of the RT structure and in the vicinity of the amino acid sites responsible for RT enzymatic activity or function.</p>
</abstract>
<kwd-group>
<kwd>HIV-1</kwd>
<kwd>bioinformatics</kwd>
<kwd>pol protein</kwd>
<kwd>protein domain</kwd>
<kwd>network analysis</kwd>
<kwd>molecular evolution</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="1"/>
<equation-count count="4"/>
<ref-count count="54"/>
<page-count count="10"/>
<word-count count="7642"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p><italic>Human Immunodeficiency Virus 1</italic> (HIV-1) is a retrovirus, a specific type of RNA virus that has been widely used as a model system for studying the molecular evolution of life because it is highly adaptive and highly genetically diverse. HIV-1 has a single-stranded RNA genome and synthesizes double-stranded DNA based on its RNA genome using reverse transcriptase (RT), which is retained within the viral particle after it enters the target cell. The HIV-1 genome contains nine genes: <italic>gag</italic>, encoding the structural proteins involved in viral particle formation; <italic>env</italic>, encoding the envelope protein; <italic>pol</italic>, encoding the enzymes for replication (protease, RT, RNase H, integrase); <italic>tat</italic> and <italic>rev</italic>, involved in the regulation of gene expression; and <italic>vif</italic>, <italic>vpr, vpu</italic>, and <italic>nef</italic>, which are accessory genes required for optimal viral replication <italic>in vivo</italic>.</p>
<p>Molecular phylogenies have shown that HIV-1 arose in humans by cross-species infection from chimpanzees at the beginning of the twentieth century (Sharp and Hahn, <xref ref-type="bibr" rid="B46">2010</xref>), and the infection has spread worldwide since the latter half of the twentieth century. This lineage, which is the predominant lineage throughout the world, is called group M and is classified into nine subtypes based on their phylogenetic relationships: subtypes A, B, C, D, F, G, H, J, and K. This genetic diversity is mainly attributed to an error-prone RT (Preston et al., <xref ref-type="bibr" rid="B39">1988</xref>) and the genetic recombination mechanism of retroviruses (Hu and Temin, <xref ref-type="bibr" rid="B21">1990</xref>). Recombination occurs frequently between the same subtypes or between different subtypes, and plays an important role in the diversification of HIV-1 (Rambaut et al., <xref ref-type="bibr" rid="B40">2004</xref>).</p>
<p>The rates of disease progression and transmission differ according to the HIV-1 subtype involved, and it is thought that these differences contribute to differences in the prevalence and expansion of the subtypes. Several studies have reported that subtype D infections have a faster disease progression rate than subtype A infections (Kaleebu et al., <xref ref-type="bibr" rid="B24">2002</xref>; Vasan et al., <xref ref-type="bibr" rid="B51">2006</xref>; Baeten et al., <xref ref-type="bibr" rid="B6">2007</xref>; Kiwanuka et al., <xref ref-type="bibr" rid="B28">2008</xref>; Ng et al., <xref ref-type="bibr" rid="B35">2013</xref>); the transmissibility of subtype C is greater than that of subtype A or D (Renjifo et al., <xref ref-type="bibr" rid="B41">2004</xref>); the replication capacity of subtype C is lower than that of the other group M subtypes (Abraha et al., <xref ref-type="bibr" rid="B1">2009</xref>; Kiguoya et al., <xref ref-type="bibr" rid="B27">2017</xref>); and RT activity during replication differs between subtypes B and C (Armstrong et al., <xref ref-type="bibr" rid="B5">2009</xref>; Iordanskiy et al., <xref ref-type="bibr" rid="B22">2010</xref>). Several studies have also detected sequence differences in the HIV-1 proteases and the active N-terminal regions of RT and integrase (Gordon et al., <xref ref-type="bibr" rid="B17">2003</xref>; Kantor et al., <xref ref-type="bibr" rid="B25">2005</xref>; Rhee et al., <xref ref-type="bibr" rid="B42">2006</xref>; Myers and Pillay, <xref ref-type="bibr" rid="B33">2008</xref>). The Pol protein, which contains these regions, is thought to be associated with the differences in the replication capacity and disease progression of the different subtypes (Ng et al., <xref ref-type="bibr" rid="B35">2013</xref>). However, the functional regions or amino acid residues in each viral protein that correspond to subtype differentiation have not been clarified.</p>
<p>The HIV-1 subtypes have usually been classified according to phylogenetic trees based on nucleotide or protein sequences of the viral core genes (<italic>gag, pol</italic>, and <italic>env</italic>; Castro-Nallar et al., <xref ref-type="bibr" rid="B9">2012</xref>) and the clade relationships established (Robertson et al., <xref ref-type="bibr" rid="B43">2000</xref>). Phylogenetic trees reflect the bifurcating phylogenetic relationships of sequences, but the construction of exact trees is difficult when the sequences contain intrasubtype recombinants, which occur frequently in HIV-1 (Posada et al., <xref ref-type="bibr" rid="B38">2002</xref>; Arenas and Posada, <xref ref-type="bibr" rid="B4">2010</xref>). Therefore, we constructed a sequence similarity network (SSN), a weighted undirected graph based on sequence similarities, to visualize the sequence space and observe the positional relationships among the subtypes in various regions of the HIV-1 genome.</p>
<p>Our aims were to systematically characterize the evolution of the HIV-1 subtypes, and to clarify the sequence regions that are responsible for the differentiation of the viral subtypes. In this study, we analyzed the mutation sites in 2,052 Pol proteins from HIV-1 group M using SSNs. Because the Pol protein is often used for group or subtype classification, we determined the overall positional relationships among the subtypes based on the Pol sequences. We then compared the structures of the domain-based networks to identify the regions that characterize the subtypes. The amino acid sites corresponding to the different subtypes were specified and mapped to the three-dimensional structure of the protein. We discuss the possible implications of these results in light of the kinds of regions that have changed during the adaptation of the virus in its spread throughout the world.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and methods</title>
<sec>
<title>Data sources</title>
<p>The near-complete genome sequences and their attributions (sampling year, sampling region, and subtype/sub-subtype of the viral sequence) of HIV-1 group M subtypes A (which consists of sub-subtypes A1 and A2), B, and C were downloaded from the HIV Sequence Database at Los Alamos (<ext-link ext-link-type="uri" xlink:href="http://www.hiv.lanl.gov">http://www.hiv.lanl.gov</ext-link>, last accessed August 2014). We used these three subtypes because the genetic distances between their sequences are almost equivalent (Robertson et al., <xref ref-type="bibr" rid="B43">2000</xref>) and the number of sequences registered in the database is large enough for our analysis. After the intersubtype recombinants and truncated sequences were excluded, 2,052 Pol protein sequences (144 subtype A; 1,528 subtype B; 330 subtype C) were obtained. The regional breakdown of the datasets is shown in Table <xref ref-type="table" rid="T1">1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Geographic regional breakdown of HIV-1 datasets used in this study.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="6" style="border-bottom: thin solid #000000;"><bold>Number of sequences</bold></th>
<th/>
</tr>
<tr>
<th valign="top" align="left"><bold>Subtype</bold></th>
<th valign="top" align="center"><bold>Africa</bold></th>
<th valign="top" align="center"><bold>Asia</bold></th>
<th valign="top" align="center"><bold>Central and South America</bold></th>
<th valign="top" align="center"><bold>Europe</bold></th>
<th valign="top" align="center"><bold>North America</bold></th>
<th valign="top" align="center"><bold>Oceania</bold></th>
<th valign="top" align="center"><bold>Total</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">A</td>
<td valign="top" align="center">73 (1)</td>
<td valign="top" align="center">38 (1)</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">31</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">144 (2)</td>
</tr>
<tr>
<td valign="top" align="left">B</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">222</td>
<td valign="top" align="center">149</td>
<td valign="top" align="center">148</td>
<td valign="top" align="center">985</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center">1,528</td>
</tr>
<tr>
<td valign="top" align="left">C</td>
<td valign="top" align="center">338</td>
<td valign="top" align="center">28</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">380</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Subtype A can further be divided into sub-subtypes A1 and A2; the numbers in parentheses show the number of sub-subtype A2 sequences</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Network analysis based on sequence similarities</title>
<p>The sequence similarity scores were calculated to construct a weighted undirected graph (SSN). The similarity scores (Basic Local Alignment Search Tool [BLAST] bit scores; Altschul et al., <xref ref-type="bibr" rid="B2">1990</xref>) for all the HIV-1 Pol protein sequences were calculated with an all-against-all BLASTP (BLAST 2.2.31&#x0002B;) analysis (Altschul et al., <xref ref-type="bibr" rid="B3">1997</xref>; Camacho et al., <xref ref-type="bibr" rid="B8">2009</xref>), with a cut-off <italic>E</italic>-value of &#x02264; 1e&#x02212;5. Using the BLAST bit scores, the sequence similarities were normalized to 0.0&#x02013;1.0, with the following equation (Dufour et al., <xref ref-type="bibr" rid="B12">2010</xref>; Matsui et al., <xref ref-type="bibr" rid="B30">2013</xref>):</p>
<disp-formula id="E1"><mml:math id="M1"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where sim(<italic>x</italic>,<italic>y</italic>) represents the normalized sequence similarity between two sequences <italic>x</italic> and <italic>y</italic>. If the score was 1.0, the pair was deemed to be identical. A weighted undirected graph was constructed based on the scores of all the pairs of sequences, and the edges were weighted with the scores. We set a threshold sequence identity value and connected the nodes when the sequence identity exceeded the threshold. The threshold to be used was determined by comparing the networks constructed with an incremental series of threshold values. We constructed SSNs of both the full-length Pol protein sequences and the functional domain sequences within the Pol protein. The constructed networks were visualized with Cytoscape 3.4.0 (Shannon et al., <xref ref-type="bibr" rid="B45">2003</xref>), with a force-directed layout.</p>
</sec>
<sec>
<title>Clustering based on the network structure</title>
<p>Spectral clustering (Paccanaro et al., <xref ref-type="bibr" rid="B36">2006</xref>), a clustering method that divides data into clusters based on the structure of a network graph, was performed with SCPS 0.9.5 (Nepusz et al., <xref ref-type="bibr" rid="B34">2010</xref>) for the networks constructed from the full-length Pol protein sequence. With this clustering algorithm, we analyzed the factors (sampling year, sampling region, and subtype) that affected the mutations in the Pol proteins by gradually changing the number of divisions.</p>
</sec>
<sec>
<title>Extraction of functional domain sequences in the Pol protein</title>
<p>Based on the HIV-1 group M subtype B reference strain HXB2 (GenBank accession: <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="K03455">K03455</ext-link>), information on the functional domains was obtained from the Swiss-Prot database (<ext-link ext-link-type="uri" xlink:href="http://www.uniprot.org/">http://www.uniprot.org/</ext-link>; The UniProt Consortium, <xref ref-type="bibr" rid="B10">2017</xref>), a high-quality annotated protein sequence knowledgebase. The dataset sequences were then aligned to the HXB2 Pol sequence using MAFFT L-INS-i 7.245 (Katoh and Standley, <xref ref-type="bibr" rid="B26">2013</xref>) to extract the functional domains. Note that the amino acid residues mentioned in this study are numbered according to the HXB2 sequence.</p>
</sec>
<sec>
<title>Construction of phylogenetic trees from functional domain sequences of Pol protein</title>
<p>We randomly selected 10 sequences from each subtype (subtype A, B, or C) to represent each functional domain sequence. A multiple-sequence alignment of each domain was created with MAFFT L-INS-i 7.245 (Katoh and Standley, <xref ref-type="bibr" rid="B26">2013</xref>), and maximum likelihood phylogenetic trees were constructed with RAxML 8.2.9 (Stamatakis, <xref ref-type="bibr" rid="B48">2014</xref>; GAMMA model with 1,000 bootstrap replicates). The calculated trees were visualized with FigTree 1.4.2 (<ext-link ext-link-type="uri" xlink:href="http://tree.bio.ed.ac.uk/software/">http://tree.bio.ed.ac.uk/software/</ext-link>).</p>
</sec>
<sec>
<title>Calculation of cumulative relative entropy (CRE)</title>
<p>To identify the sites that characterize the differences between each subtype of HIV-1, we calculated CRE (Hannenhalli and Russell, <xref ref-type="bibr" rid="B19">2000</xref>) for each amino acid site in the Pol protein. We calculated the amino acid compositions of the three subtypes (A, B, and C) at each site in a multiple alignment. The <italic>hmmbuild</italic> program of HMMER 3.1b2 (Eddy, <xref ref-type="bibr" rid="B14">1998</xref>) was used to build profile P of the alignment. The weighting method of Henikoff and Henikoff (<xref ref-type="bibr" rid="B20">1994</xref>) was used for the residue counts. The relative entropy (Shannon, <xref ref-type="bibr" rid="B44">1996</xref>; Durbin et al., <xref ref-type="bibr" rid="B13">1998</xref>) of position <italic>i</italic> for subtype <inline-formula><mml:math id="M2"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> with respect to the entropy of that position for subtype was calculated. If <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mtext>RE</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the relative entropy of <inline-formula><mml:math id="M4"><mml:msubsup><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> with respect to P<sub><italic>i</italic></sub>:</p>
<disp-formula id="E2"><mml:math id="M5"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>x</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Note that RE is greater than or equal to zero, and is exactly zero when the two distributions are identical (Durbin et al., <xref ref-type="bibr" rid="B13">1998</xref>). To estimate the role of alignment position <italic>i</italic> in characterizing the HIV-1 subtypes, CRE<sub><italic>i</italic></sub> was calculated as:</p>
<disp-formula id="E3"><mml:math id="M6"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>b</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>s</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>R</mml:mi><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The CREs for all the positions were converted into Z-scores based on the distribution of the entropies within a sequence alignment. Let &#x003BC; and &#x003C3; be the means and standard deviations of the CREs of all positions, then the Z-score for position <italic>i</italic> is calculated as:</p>
<disp-formula id="E4"><mml:math id="M7"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>Z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We expect a position with a high Z-score to be important in characterizing the subtypes. We calculated the CREs, and positions with Z-scores of CRE &#x02265; 3.0 were defined as &#x0201C;high-CRE&#x0201D; positions. All multiple sequence alignments used to calculate CRE were constructed with MAFFT L-INS-i 7.245 (Katoh and Standley, <xref ref-type="bibr" rid="B26">2013</xref>). The generated alignments were visualized and sequence conservation was calculated with Jalview 2.9.0b2 (Waterhouse et al., <xref ref-type="bibr" rid="B54">2009</xref>).</p>
</sec>
<sec>
<title>Protein conformation analysis of RT</title>
<p>The functional domains characterizing the subtypes and the high-CRE residues were mapped onto the structure of HIV-1 RT. The protein structural data (PDB ID: <ext-link ext-link-type="PDB" xlink:href="1REV">1REV</ext-link>; <ext-link ext-link-type="PDB" xlink:href="3KJV">3KJV</ext-link> for RT complexed with the DNA duplex) were obtained from the Protein Data Bank (<ext-link ext-link-type="uri" xlink:href="http://www.rcsb.org/pdb">http://www.rcsb.org/pdb</ext-link>; Bernstein et al., <xref ref-type="bibr" rid="B7">1977</xref>), a database of experimentally determined protein structures, and visualized with UCSF Chimera 1.10.2 (Pettersen et al., <xref ref-type="bibr" rid="B37">2004</xref>).</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<sec>
<title>Comparison of thousands of HIV-1 Pol sequences based on a network analysis</title>
<p>The classification of and relationships between each HIV-1 subtype were determined by constructing networks based on the amino acid sequence similarities of the Pol polyprotein (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">1</xref> and Figure <xref ref-type="fig" rid="F1">1</xref>). The SSN is a graphical representation of the similarities between sequences. Each sequence is indicated by a point (node) and the similarity between the sequences is represented by the length of the line (edge) connecting the points. The smaller the distance between the nodes, the greater the degree of similarity between the sequences. We used subtypes A, B, and C from HIV-1 group M in the present analysis. Subtypes A, B, and C clearly form distinct groups when their sequence similarities are analyzed (Robertson et al., <xref ref-type="bibr" rid="B43">2000</xref>) and more of these sequences are registered in the database than those of other subtypes, so we assumed that enough sequences were available for our purpose.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Stepwise analysis of the SSNs of HIV-1 Pol protein. In total, 2,052 amino acid sequences of the HIV-1 Pol protein were used to construct networks based on sequence similarities (SSNs). Nodes (colored dots) represent the Pol protein sequences and the edge lengths represent the sequence similarities. <bold>(A)</bold> Nodes are colored according to the HIV-1 subtype. <bold>(B&#x02013;E)</bold> To investigate the possible process of Pol protein evolution, a total of 2,052 Pol sequences, used to construct the network structure shown in panel A, were clustered into 2, 3, 4, or 5 groups with the spectral clustering method by changing the number of divisions in a sequential order, and are colored according to cluster. A stepwise analysis of the hierarchy of the clusters allowed the possible evolutionary pathways of the Pol proteins to be estimated. Panel <bold>(E)</bold> shows the attributes (subtype, region of sampling) of the sequences in each cluster.</p></caption>
<graphic xlink:href="fmicb-08-02151-g0001.tif"/>
</fig>
<p>When constructing an SSN, the network structure changes according to the threshold value of the sequence identity used when connecting the edges (Fujishima et al., <xref ref-type="bibr" rid="B16">2008</xref>). Therefore, we first constructed a series of networks by gradually changing the threshold value and compared their structures (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">1</xref>). In networks in which the edges were connected with sequence identities &#x02265;80%, the three subtypes of HIV-1 were not well-separated and formed one large network (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">1A</xref>). When the sequence identity threshold was &#x02265;92%, each of the three subtypes was properly separated. Because the nodes were still connected under this threshold, the relative positional relationships among subtypes were determined (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">1B</xref>). However, as the threshold value became much stricter, the connections between the subtypes were broken, and with a sequence identity threshold of &#x02265;96%, more than 130 graphs were generated, so that it was impossible to determine the exact positional relationships among the subtypes. Therefore, we adopted a threshold value for sequence identity of &#x02265;92% for the subsequent analysis.</p>
<p>In Figure <xref ref-type="fig" rid="F1">1A</xref>, the nodes are colored to represent the three subtypes shown in Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">1B</xref>, and provides an overall view of the sequence similarities between and/or within each subtype. To confirm the subtype groupings in the same SSN and to clarify the attributions obtained from the database (sampling year, sampling region, and viral sequence subtype), we used the spectral clustering method, which divides sequence data into clusters based on the structure of a network graph. This methodology enables the estimation of the process of subtype differentiation according to the order of cluster division when dividing clusters step by step. Therefore, we changed the clustering division numbers to two clusters (Figure <xref ref-type="fig" rid="F1">1B</xref>), three clusters (Figure <xref ref-type="fig" rid="F1">1C</xref>), four clusters (Figure <xref ref-type="fig" rid="F1">1D</xref>), and five clusters (Figure <xref ref-type="fig" rid="F1">1E</xref>). At the two-cluster stage (Figure <xref ref-type="fig" rid="F1">1B</xref>), subtype B and subtypes A and C formed specific groups. Subtypes A and C were then differentiated into different groups in the three-cluster stage (Figure <xref ref-type="fig" rid="F1">1C</xref>). These evolutionary steps are consistent with the order of subtype differentiation reported in a previous study (Castro-Nallar et al., <xref ref-type="bibr" rid="B9">2012</xref>), and the elements of the clusters and each subtype in the three-cluster stage showed a high coincidence ratio (99.4%), indicating that this network analysis is a suitable technique for classifying these subtypes (Figures <xref ref-type="fig" rid="F1">1A,C</xref>). When the number of divisions was set to 4, the sequence group corresponding to subtype A was further divided into two clusters, which exactly matched sub-subtypes A1 and A2, respectively (Figure <xref ref-type="fig" rid="F1">1D</xref>). With five clusters, subtype A1 was further divided into an additional two clusters, consisting mainly of the sequences from Africa or from Asia and Europe (Figure <xref ref-type="fig" rid="F1">1E</xref>).</p>
</sec>
<sec>
<title>Domain-based network analysis shows that RNase H domain and RT connection domain are important for subtype differentiation</title>
<p>To analyze the regions in the HIV-1 Pol polyprotein that are responsible for HIV-1 subtype differentiation, we constructed SSNs based on each functional domain (Figure <xref ref-type="fig" rid="F2">2</xref>). Figure <xref ref-type="fig" rid="F2">2A</xref> shows the eight domains present in the HIV-1 Pol polyprotein (<bold>a</bold>, retroviral aspartyl protease, residues 61&#x02013;153; <bold>b</bold>, RT (RNA-dependent DNA polymerase), residues 218&#x02013;388; <bold>c</bold>, RT thumb domain, residues 396&#x02013;458; <bold>d</bold>, RT connection domain, residues 473&#x02013;573; <bold>e</bold>, RNase H, residues 591&#x02013;711; <bold>f</bold>, integrase zinc-binding domain, residues 723&#x02013;759; <bold>g</bold>, integrase core domain, residues 770&#x02013;875; and <bold>h</bold>, integrase DNA binding domain, residues 936&#x02013;982). When we compared the SSNs constructed for each domain, the nodes of the three subtypes were mixed in the networks of the RT (RNA-dependent DNA polymerase) region (Figure <xref ref-type="fig" rid="F2">2Bb</xref>) and the integrase DNA-binding domain (Figure <xref ref-type="fig" rid="F2">2Bh</xref>) because of the high sequence conservation in these regions. In the RT thumb domain (Figure <xref ref-type="fig" rid="F2">2Bc</xref>), the integrase zinc-binding domain (Figure <xref ref-type="fig" rid="F2">2Bf</xref>), and the integrase core domain (Figure <xref ref-type="fig" rid="F2">2Bg</xref>), the nodes for subtype A and C were mixed, although those for subtype B and subtypes A and C were separated. Therefore, we consider these regions unsuitable for distinguishing these subtypes. In contrast, the nodes of the three subtypes were well-separated in another three domains: the retroviral aspartyl protease region (Figure <xref ref-type="fig" rid="F2">2Ba</xref>), the RT connection domain (Figure <xref ref-type="fig" rid="F2">2Bd</xref>), and the RNase H domain (Figure <xref ref-type="fig" rid="F2">2Be</xref>). Among these, we consider the retroviral aspartyl protease region inappropriate for distinguishing the subtypes because the nodes are dispersed compared with those of the other two regions. Therefore, we conclude that the RT connection domain and the RNase H domain represent the differences among the HIV-1 subtypes. To observe effects of non-domain regions on subtype classification, we extracted five regions lying between the eight domains of the Pol protein (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">2Ai</xref>&#x02013;<xref ref-type="supplementary-material" rid="SM1">m</xref>) and constructed the corresponding SSNs (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">2B</xref>). However, in these networks of non-domain regions, the boundaries of the three subtypes were indefinable. In particular, for the region shown in Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">2Ai</xref>, hundreds of graphs were generated, which did not provide clear indices for distinguishing the subtypes.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Comparison of the SSNs of functional domains of the HIV-1 Pol protein. <bold>(A)</bold> Schematic representation of the positions and lengths of the functional domains of the HIV-1 Pol protein: <bold>a</bold>, retroviral aspartyl protease; <bold>b</bold>, RT (RNA-dependent DNA polymerase); <bold>c</bold>, RT thumb domain; <bold>d</bold>, RT connection domain; <bold>e</bold>, RNase H; <bold>f</bold>, integrase zinc-binding domain; <bold>g</bold>, integrase core domain; and <bold>h</bold>, integrase DNA-binding domain. <bold>(B)</bold> SSNs of each functional domain region of the HIV-1 Pol protein (2,052 sequences) were created and colored according to subtype. Nodes represent the sequences of each functional domain and the edge lengths represent the sequence similarities. Symbols <bold>(a&#x02013;h)</bold> correspond to those in panel <bold>(A)</bold>.</p></caption>
<graphic xlink:href="fmicb-08-02151-g0002.tif"/>
</fig>
<p>To confirm these results, we analyzed each selected domain phylogenetically, with the maximum likelihood method. The results supported our conclusions based on our domain-based SSNs (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">3</xref>). In the phylogenetic trees for the RT connection domain (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">3A</xref>) and the RNase H domain (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">3B</xref>), branches of the three subtypes are clearly separated. However, in other domains, such as the RT (RNA-dependent DNA polymerase) domain (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">3C</xref>) and integrase core domain (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">3D</xref>), there are ambiguous boundaries among the nodes corresponding to each subtype. In particular, the nodes of the three subtypes are mixed in the clades for the integrase zinc-binding domain (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">3E</xref>). Furthermore, to quantitatively identify the domains of the sequences that characterize the differences among subtypes, CRE was calculated for each site after the sequences were aligned. This index takes a large value when the difference between an amino acid distribution in one subtype and that in the other subtypes is large. The mean CRE value was calculated for each domain of the Pol protein. The regions with the highest average CRE values were the RT connection domain and the RNase H domain (Figure <xref ref-type="fig" rid="F3">3</xref>). This supports the results based on the network structures shown in Figure <xref ref-type="fig" rid="F2">2B</xref>. The mean CRE value for the RT connection domain was statistically significantly larger than those for the other regions.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Mean cumulative relative entropy (CRE) values for each functional domain of HIV-1 Pol protein. The x-axis represents each functional domain region, corresponding to those in Figure <xref ref-type="fig" rid="F2">2A</xref>: a, retroviral aspartyl protease; b, RT (RNA-dependent DNA polymerase); c, RT thumb domain; d, RT connection domain; e, RNase H; f, integrase zinc-binding domain; g, integrase core domain; and h, integrase DNA-binding domain. The y-axis represents the mean CRE value per site in each domain region. Groups were compared with the Mann&#x02013;Whitney <italic>U-</italic>test (<sup>&#x0002A;</sup><italic>P</italic> &#x0003C; 0.05; <sup>&#x0002A;&#x0002A;</sup><italic>P</italic> &#x0003C; 0.01; <sup>&#x0002A;&#x0002A;&#x0002A;</sup><italic>P</italic> &#x0003C; 0.001). Error bars represent standard errors of the means and asterisks indicate the significance of differences between the RT connection domain or RNase H domain and each region.</p></caption>
<graphic xlink:href="fmicb-08-02151-g0003.tif"/>
</fig>
</sec>
<sec>
<title>Mapping the amino acid residues that are crucial for subtype specification</title>
<p>To clarify the changes in the amino acid residues that distinguish the three subtypes, the CRE value was calculated for each site in both the RT connection domain and RNase H domain. The sites with CRE-derived Z-scores &#x02265; 3.0 (see section Materials and Methods) were then identified on an amino acid sequence alignments (Figure <xref ref-type="fig" rid="F4">4</xref>). High-CRE sites occurred at position 357 in the RT connection domain (Figure <xref ref-type="fig" rid="F4">4A</xref>) and at positions 480, 483, and 491 in the RNase H domain (Figure <xref ref-type="fig" rid="F4">4B</xref>). At position 357 (d1 site on Figure <xref ref-type="fig" rid="F4">4A</xref>) in the RT connection domain, lysine (K) occurs in 100% of subtype A sequences, methionine (M) in the majority (70%) of subtype B sequences, and methionine (M) and arginine (R) in about 50% each of the subtype C sequences, so the amino acid distribution at this site differs in the three subtypes. Meanwhile, in the RNase H domain, the amino acid compositions of subtypes B and C are similar at position 480 (e1), those of subtypes A and B are similar at position 483 (e2), and those of subtypes A and C are similar at position 491 (e3). In the RT connection domain, the subtypes can be distinguished at one site, whereas in the RNase H domain, the subtypes can be distinguished at three sites. These three sites are situated in close proximity in the RNase H domain. In terms of the sequence conservation in both domains, a region with low conservation is distributed throughout the domains, but these regions are not necessarily high-CRE regions. Therefore, we conclude that low sequence conservation does not generate the differences between the subtypes.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Multiple amino acid sequence alignments of the RT connection domain and RNase H domain of the Pol protein. A total of 30 sequences, 10 sequences randomly selected from each subtypes (subtypes A, B, and C), were aligned with the MAFFT L-INS-i algorithm. <bold>(A)</bold> Sequence alignment around the sites in the RT connection domain that characterize the subtypes. <bold>(B)</bold> Sequence alignment around the sites in the RNase H domain that characterize the subtypes. For each amino acid position, the conservation scores according to Livingstone and Barton (<xref ref-type="bibr" rid="B29">1993</xref>) are shown as 12 ranks (from 0 to 11). Identical amino acids (rank 11) are indicated by asterisks, and partly conserved amino acids (rank 10) are indicated by a plus symbol. Orange triangles at the top of each alignment indicate the sites with CRE Z-scores &#x02265; 3.0 (designated d1 and e1-3). CRE was calculated based on all the sequences in the dataset. Sequence conservation scores and Z-scores of CRE for each site are shown together at the bottom of each alignment.</p></caption>
<graphic xlink:href="fmicb-08-02151-g0004.tif"/>
</fig>
<p>To localize the regions that characterize the subtypes on the tertiary structure of RT, the aforementioned two domains and the high-CRE sites were mapped onto the three-dimensional conformation of HIV-1 RT (Figure <xref ref-type="fig" rid="F5">5</xref>). HIV-1 RT is a heterodimer, consisting of p66, which contains the RNase H domain, and p51 without the RNase H domain. The high-CRE sites M357 in the p66 chain of the RT connection domain and Q480, Y483, and L491 in the RNase H domain are physically close (Figure <xref ref-type="fig" rid="F5">5A</xref>). These amino acid residues were also located on the surface of RT in a representation of the molecular surface structure (Figure <xref ref-type="fig" rid="F5">5B</xref> and see section Discussion).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Mapping the amino acid residues corresponding to the differences among HIV-1 subtypes onto RT. The structure of the HIV-1 RT p66/p51 heterodimer is shown in <bold>(A)</bold> as a ribbon diagram, and in <bold>(B)</bold> as a molecular surface representation. The RT connection domain (on p66 and p51) and the RNase H domain (on p66) are colored light blue and light pink, respectively. High-CRE positions d1 and e1-3 are colored blue and magenta, respectively (see also Figure <xref ref-type="fig" rid="F4">4</xref>). Note that the positions are represented on a space-filling model in panel <bold>(A)</bold>. PDB ID: <ext-link ext-link-type="PDB" xlink:href="1REV">1REV</ext-link>.</p></caption>
<graphic xlink:href="fmicb-08-02151-g0005.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>In this study, we successfully visualized the similarity relationships of thousands of sequences and summarized these data in SSNs. In the network constructed using the full-length amino acid sequence of the HIV-1 Pol protein, the sequence relationships of the three HIV-1 group M subtypes were visualized, and stepwise clustering divided the subtype A sequence into further clusters (Figure <xref ref-type="fig" rid="F1">1</xref>). This cluster division corresponded to the different regions in which the viruses were sampled. Therefore, although HIV-1 is prevalent throughout the world, sequence changes arise particularly frequently in epidemic regions, and the network structure properly reflected these differences in their sequences. By comparing the network structures of each domain and the mean CRE values, we identified the RT connection domain and the RNase H domain of RT as characterizing the differences among the HIV-1 subtypes. By calculating the CREs for the amino acid sequences of the connection domain and the RNase H domain, we identified one amino acid residue in the connection domain and three residues in the RNase H domain as subtype-characterizing sites.</p>
<p>Using SSNs, we have previously determined the similarity relationships among a huge number of sequences with complex phylogenetic relationships and have discussed their evolution, including the bacterial CRP/FNR transcriptional regulator superfamily (Matsui et al., <xref ref-type="bibr" rid="B30">2013</xref>) and the novel tRNA genes that have expanded in certain species of eukaryotes (Hamashima et al., <xref ref-type="bibr" rid="B18">2015</xref>). In these reports, we estimated a possible evolutionary pathway by observing the phylogenies inferred from the stepwise analysis of cluster hierarchies. Although, those studies dealt mainly with the amino acid sequences of whole proteins or the nucleotide sequences of whole tRNA genes, in the present study, we constructed SSNs based on domain-level sequences for the first time. In Figure <xref ref-type="fig" rid="F2">2Bd</xref>, the RT connection domain of subtype A is divided further into two different groups. As described below, at least at the domain level, the difference in the RT activity corresponds to the difference between subtypes B and C. Thus, we speculated that there might be a difference in the RT enzymatic activity between these two groups of the RT connection domain of subtype A. Our results suggest that comparing the networks based on individual protein domains allows not only the detection of subtype differences, but also the functional divergence of the domains analyzed.</p>
<p>We excluded the retroviral aspartyl protease region from the analysis for the reasons described above (see section Domain-Based Network Analysis Shows That RNase H Domain and RT Connection Domain Are Important for Subtype Differentiation). However, comparisons of the network structures (Figure <xref ref-type="fig" rid="F2">2</xref>) and the mean CRE values (Figure <xref ref-type="fig" rid="F3">3</xref>) for each domain indicated that the retroviral aspartyl protease region is also a subtype-distinguishing region, in addition to the RT connection domain and RNase H domain. HIV-1 protease has been reported to have different activities and different target cleavage sites, predominantly in subtypes C and B (Velazquez-Campoy et al., <xref ref-type="bibr" rid="B52">2001</xref>; de Oliveira et al., <xref ref-type="bibr" rid="B11">2003</xref>). Together with the thumb domain, the RT connection domain forms a binding cleft and bridges both the N-terminal polymerase activity region and the C-terminal RNase H domain. RNase H is an enzyme that specifically degrades the RNA strand of DNA/RNA complexes during reverse transcription. It has been reported that differences in replication capacity between subtypes B and C are derived from the differences between the RT connection domain and the RNase H domain (Iordanskiy et al., <xref ref-type="bibr" rid="B22">2010</xref>). This supported our current observations at least at the domain level. In addition, mutations in the connection domain and the RNase H domain are known to change the sensitivity of the virus to anti-HIV-1 drugs (RT inhibitors; Julias et al., <xref ref-type="bibr" rid="B23">2003</xref>; Men&#x000E9;ndez-Arias et al., <xref ref-type="bibr" rid="B31">2011</xref>), and these are the same domains of the Pol protein that characterize the subtypes identified in this study.</p>
<p>At the amino acid level, the sites responsible for viral subtype differentiation do not perfectly match the drug-resistance mutations (Ehteshami and G&#x000F6;tte, <xref ref-type="bibr" rid="B15">2008</xref>; von Wyl et al., <xref ref-type="bibr" rid="B53">2010</xref>; Men&#x000E9;ndez-Arias et al., <xref ref-type="bibr" rid="B31">2011</xref>), but are located very close to them in the RT structure (Supplementary Figures <xref ref-type="supplementary-material" rid="SM1">4A,B</xref>). In particular, the drug-resistance mutations R356K, R358K, and A360V (Ehteshami and G&#x000F6;tte, <xref ref-type="bibr" rid="B15">2008</xref>; von Wyl et al., <xref ref-type="bibr" rid="B53">2010</xref>; Men&#x000E9;ndez-Arias et al., <xref ref-type="bibr" rid="B31">2011</xref>) are located on the surface of the RT domain, close to M357, the position with the highest CRE in the connection domain. It is possible that mutations strongly associated with drug resistance are also closely related to the activity of RT, and the domain structure may be greatly altered and its activity reduced by a mutation that changes an amino acid to one with dissimilar biochemical properties. The three resistance mutations are conservative, including from arginine (R) to lysine (K) or from alanine (A) to valine (V). Therefore, we speculate that the HIV-1 subtypes were differentiated by the accumulation of mutations in the surface region where the sequence conservation is low, but at positions located very close to the critical amino acid residues required for enzymatic activity, changing the local structure and modulating the enzyme&#x00027;s activity. In contrast, RT is a heterodimer comprised of subunits p66 and p51. The larger p66 subunit contains the RNase H domain and the catalytic region with the main polymerase activities, whereas the smaller p51 subunit mainly plays a structural role (Telesnitsky and Goff, <xref ref-type="bibr" rid="B49">1997</xref>). In this context, further analysis is required to identify the aspect (activity or structure) of the protein that most strongly affects subtype differentiation. Mutation Q509L, a drug-resistance mutation site in the RNase H domain (Ehteshami and G&#x000F6;tte, <xref ref-type="bibr" rid="B15">2008</xref>; Men&#x000E9;ndez-Arias et al., <xref ref-type="bibr" rid="B31">2011</xref>), is not physically close to any of the high-CRE sites (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">4A</xref>). Instead, three amino acid residues critical for RNase H activity, D443, E478, and D498, are located together in the RNase H domain close to the high-CRE regions (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">4C</xref>). Mutation E478, in particular, is located very close to Q480, one of three high-CRE sites. We also noted that in the structure of RT complexed with the DNA duplex, M357 is physically close to the DNA molecule but does not interact directly with it (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S5</xref>). Again, these results suggest that the region that characterizes the subtypes is located in the vicinity of the amino acid site responsible for its enzyme activity or RT function.</p>
<p>By mapping the high-CRE sites onto sequence alignments, we found that locations with low sequence conservation do not necessarily characterize the differences in the subtypes (Figure <xref ref-type="fig" rid="F4">4</xref>). The amino acid residues at the high-CRE sites are conserved within each subtype, but differ between subtypes, so these sites are not fully conserved through all HIV-1 subtypes. This suggests that a region in which CRE is high can accommodate incoming mutations but has functional constraints that do not allow completely random mutations. We found that three HIV-1 subtypes can be distinguished by one amino acid residue (position 357) in the RT connection domain. However, each of the three amino acid residues (positions 480, 483, or 491) in the RNase H domain can be distinguished in only two of the three subtypes, indicating that all three amino acid residues need to be considered to effectively classify the subtypes in this case. We cannot completely exclude the possibility that the accumulation of mutations in these subtype-characterizing regions of the HIV-1 Pol protein is caused by genetic drift. However, it is possible that these mutations are adaptions to the environment at places throughout the world in which HIV-1 is prevalent. The internal environments of various hosts are considered to differ among regions, based on race, the immune system, and the indigenous microbial flora. As seen in Figure <xref ref-type="fig" rid="F1">1E</xref>, HIV-1 even differs between regions in which the same subtype predominates in the populations. Therefore, we suggest that the virus does not mutate precisely in genomic regions encoding enzyme activities but in neighboring regions. This modulates the enzyme functions to allow the adaptation of the virus to geographic regional differences it encounters in areas of prevalence.</p>
<p>Many currently emerging viruses that cause pandemics throughout the world are RNA viruses characterized by high mutation rates, including <italic>Ebola virus, Zika virus</italic>, and others. Genome analyses have already shown that as viral infections spread, mutations accumulate in the viral genomic sequences, causing them to differ in different endemic areas (Simon-Loriere et al., <xref ref-type="bibr" rid="B47">2015</xref>; Tong et al., <xref ref-type="bibr" rid="B50">2015</xref>; Metsky et al., <xref ref-type="bibr" rid="B32">2017</xref>). Our research provides a molecular basis for HIV-1 evolution and subtype differentiation, and should extend our understanding of the evolution and differentiation of other RNA viruses, including emerging viruses.</p>
</sec>
<sec id="s5">
<title>Author contributions</title>
<p>SN and AK conceived and designed the study, and SN wrote the manuscript. SN, JI, and GM performed the analyses and interpreted the data. AK and MT edited the manuscript. AK supervised the project. All the authors have read and approved the final manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>We thank Dr. Motomu Matsui, Dr. Haruo Suzuki, Dr. Yasuhiro Naito and Mr. Satoshi Tamaki for their constructive suggestions and discussions. We also thank all the members of the RNA Group at the Institute for Advanced Biosciences, Keio University, Japan, for their insightful discussions.</p>
</ack>
<sec sec-type="supplementary-material" id="s6">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmicb.2017.02151/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmicb.2017.02151/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Image1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abraha</surname> <given-names>A.</given-names></name> <name><surname>Nankya</surname> <given-names>I. L.</given-names></name> <name><surname>Gibson</surname> <given-names>R.</given-names></name> <name><surname>Demers</surname> <given-names>K.</given-names></name> <name><surname>Tebit</surname> <given-names>D. M.</given-names></name> <name><surname>Johnston</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>CCR5- and CXCR4-tropic subtype C human immunodeficiency virus type 1 isolates have a lower level of pathogenic fitness than other dominant group M subtypes: implications for the epidemic</article-title>. <source>J. Virol.</source> <volume>83</volume>, <fpage>5592</fpage>&#x02013;<lpage>5605</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.02051-08</pub-id><pub-id pub-id-type="pmid">19297481</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Gish</surname> <given-names>W.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <name><surname>Myers</surname> <given-names>E. W.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Basic local alignment search tool</article-title>. <source>J. Mol. Biol.</source> <volume>215</volume>, <fpage>403</fpage>&#x02013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id><pub-id pub-id-type="pmid">2231712</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S.</given-names></name> <name><surname>Madden</surname> <given-names>T.</given-names></name> <name><surname>Schaffer</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>1997</year>). <article-title>Gapped BLAST and PSI- BLAST: a new generation of protein database search programs</article-title>. <source>Nucleic Acids Res.</source> <volume>25</volume>, <fpage>3389</fpage>&#x02013;<lpage>3402</lpage>. <pub-id pub-id-type="doi">10.1093/nar/25.17.3389</pub-id><pub-id pub-id-type="pmid">9254694</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arenas</surname> <given-names>M.</given-names></name> <name><surname>Posada</surname> <given-names>D.</given-names></name></person-group> (<year>2010</year>). <article-title>The effect of recombination on the reconstruction of ancestral sequences</article-title>. <source>Genetics</source> <volume>184</volume>, <fpage>1133</fpage>&#x02013;<lpage>1139</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.109.113423</pub-id><pub-id pub-id-type="pmid">20124027</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Armstrong</surname> <given-names>K. L.</given-names></name> <name><surname>Lee</surname> <given-names>T.-H.</given-names></name> <name><surname>Essex</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Replicative capacity differences of thymidine analog resistance mutations in subtype B and C human immunodeficiency virus type 1</article-title>. <source>J. Virol.</source> <volume>83</volume>, <fpage>4051</fpage>&#x02013;<lpage>4059</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.02645-08</pub-id><pub-id pub-id-type="pmid">19225005</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baeten</surname> <given-names>J. M.</given-names></name> <name><surname>Chohan</surname> <given-names>B.</given-names></name> <name><surname>Lavreys</surname> <given-names>L.</given-names></name> <name><surname>Chohan</surname> <given-names>V.</given-names></name> <name><surname>McClelland</surname> <given-names>R. S.</given-names></name> <name><surname>Certain</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>HIV-1 subtype D infection is associated with faster disease progression than subtype A in spite of similar plasma HIV-1 loads</article-title>. <source>J. Infect. Dis.</source> <volume>195</volume>, <fpage>1177</fpage>&#x02013;<lpage>1180</lpage>. <pub-id pub-id-type="doi">10.1086/512682</pub-id><pub-id pub-id-type="pmid">17357054</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>F. C.</given-names></name> <name><surname>Koetzle</surname> <given-names>T. F.</given-names></name> <name><surname>Williams</surname> <given-names>G. J.</given-names></name> <name><surname>Meyer</surname> <given-names>E. E.</given-names> <suffix>Jr.</suffix></name> <name><surname>Brice</surname> <given-names>M. D.</given-names></name> <name><surname>Rodgers</surname> <given-names>J. R.</given-names></name> <etal/></person-group>. (<year>1977</year>). <article-title>The protein data bank: a computer-based archival fole for macromolecular structures</article-title>. <source>J. Mol. Biol.</source> <volume>112</volume>, <fpage>535</fpage>&#x02013;<lpage>542</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(77)80200-3</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Camacho</surname> <given-names>C.</given-names></name> <name><surname>Coulouris</surname> <given-names>G.</given-names></name> <name><surname>Avagyan</surname> <given-names>V.</given-names></name> <name><surname>Ma</surname> <given-names>N.</given-names></name> <name><surname>Papadopoulos</surname> <given-names>J.</given-names></name> <name><surname>Bealer</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>BLAST&#x0002B;: architecture and applications</article-title>. <source>BMC Bioinformatics</source> <volume>10</volume>:<fpage>421</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-10-421</pub-id><pub-id pub-id-type="pmid">20003500</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Castro-Nallar</surname> <given-names>E.</given-names></name> <name><surname>P&#x000E9;rez-Losada</surname> <given-names>M.</given-names></name> <name><surname>Burton</surname> <given-names>G. F.</given-names></name> <name><surname>Crandall</surname> <given-names>K. A.</given-names></name></person-group> (<year>2012</year>). <article-title>The evolution of HIV: inferences using phylogenetics</article-title>. <source>Mol. Phylogenet. Evol.</source> <volume>62</volume>, <fpage>777</fpage>&#x02013;<lpage>792</lpage>. <pub-id pub-id-type="doi">10.1016/j.ympev.2011.11.019</pub-id><pub-id pub-id-type="pmid">22138161</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><collab>The UniProt Consortium</collab></person-group> (<year>2017</year>). <article-title>UniProt: the universal protein knowledgebase</article-title>. <source>Nucl. Acids Res</source>. <volume>45</volume>, <fpage>D158</fpage>&#x02013;<lpage>D169</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw1099</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Oliveira</surname> <given-names>T.</given-names></name> <name><surname>Engelbrecht</surname> <given-names>S.</given-names></name> <name><surname>Janse van Rensburg</surname> <given-names>E.</given-names></name> <name><surname>Gordon</surname> <given-names>M.</given-names></name> <name><surname>Bishop</surname> <given-names>K.</given-names></name> <name><surname>Zur Megede</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Variability at human immunodeficiency virus type 1 subtype C protease cleavage sites: an indication of viral fitness?</article-title> <source>J. Virol.</source> <volume>77</volume>, <fpage>9422</fpage>&#x02013;<lpage>9430</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.77.17.9422-9430.2003</pub-id><pub-id pub-id-type="pmid">12915557</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dufour</surname> <given-names>Y. S.</given-names></name> <name><surname>Kiley</surname> <given-names>P. J.</given-names></name> <name><surname>Donohue</surname> <given-names>T. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Reconstruction of the core and extended regulons of global transcription factors</article-title>. <source>PLoS Genet.</source> <volume>6</volume>:<fpage>e1001027</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1001027</pub-id><pub-id pub-id-type="pmid">20661434</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Durbin</surname> <given-names>R.</given-names></name> <name><surname>Eddy</surname> <given-names>S. R.</given-names></name> <name><surname>Krogh</surname> <given-names>A.</given-names></name> <name><surname>Mitchison</surname> <given-names>G.</given-names></name></person-group> (<year>1998</year>). <source>Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eddy</surname> <given-names>S.</given-names></name></person-group> (<year>1998</year>). <article-title>Profile hidden Markov models</article-title>. <source>Bioinformatics</source> <volume>14</volume>, <fpage>755</fpage>&#x02013;<lpage>763</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/14.9.755</pub-id><pub-id pub-id-type="pmid">9918945</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ehteshami</surname> <given-names>M.</given-names></name> <name><surname>G&#x000F6;tte</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>Effects of mutations in the connection and RNase H domains of HIV-1 reverse transcriptase on drug susceptibility</article-title>. <source>AIDS Rev.</source> <volume>10</volume>, <fpage>224</fpage>&#x02013;<lpage>235</lpage>. <pub-id pub-id-type="pmid">19092978</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fujishima</surname> <given-names>K.</given-names></name> <name><surname>Sugahara</surname> <given-names>J.</given-names></name> <name><surname>Tomita</surname> <given-names>M.</given-names></name> <name><surname>Kanai</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>Sequence evidence in the archaeal genomes that tRNAs emerged through the combination of ancestral genes as 5&#x02032; and 3&#x02032; tRNA halves</article-title>. <source>PLoS ONE</source> <volume>3</volume>:<fpage>e1622</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0001622</pub-id><pub-id pub-id-type="pmid">18286179</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gordon</surname> <given-names>M.</given-names></name> <name><surname>De Oliveira</surname> <given-names>T.</given-names></name> <name><surname>Bishop</surname> <given-names>K.</given-names></name> <name><surname>Coovadia</surname> <given-names>H. M.</given-names></name> <name><surname>Madurai</surname> <given-names>L.</given-names></name> <name><surname>Engelbrecht</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Molecular characteristics of human immunodeficiency virus type 1 subtype C viruses from KwaZulu-Natal, South Africa: implications for vaccine and antiretroviral control strategies</article-title>. <source>Society</source> <volume>77</volume>, <fpage>2587</fpage>&#x02013;<lpage>2599</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.77.4.2587-2599.2003</pub-id><pub-id pub-id-type="pmid">12551997</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hamashima</surname> <given-names>K.</given-names></name> <name><surname>Tomita</surname> <given-names>M.</given-names></name> <name><surname>Kanai</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Expansion of noncanonical V-Arm-Containing tRNAs in eukaryotes. <italic>Mol. Biol</italic></article-title>. <source>Evol.</source> <volume>33</volume>, <fpage>530</fpage>&#x02013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msv253</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hannenhalli</surname> <given-names>S. S.</given-names></name> <name><surname>Russell</surname> <given-names>R. B.</given-names></name></person-group> (<year>2000</year>). <article-title>Analysis and prediction of functional sub-types from protein sequence alignments</article-title>. <source>J. Mol. Biol.</source> <volume>303</volume>, <fpage>61</fpage>&#x02013;<lpage>76</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.2000.4036</pub-id><pub-id pub-id-type="pmid">11021970</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henikoff</surname> <given-names>S.</given-names></name> <name><surname>Henikoff</surname> <given-names>J. G.</given-names></name></person-group> (<year>1994</year>). <article-title>Position-based sequence weights</article-title>. <source>J. Mol. Biol.</source> <volume>243</volume>, <fpage>574</fpage>&#x02013;<lpage>578</lpage>. <pub-id pub-id-type="doi">10.1016/0022-2836(94)90032-9</pub-id><pub-id pub-id-type="pmid">7966282</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>W. S.</given-names></name> <name><surname>Temin</surname> <given-names>H. M.</given-names></name></person-group> (<year>1990</year>). <article-title>Retroviral recombination and reverse transcription</article-title>. <source>Science</source> <volume>250</volume>, <fpage>1227</fpage>&#x02013;<lpage>1233</lpage>. <pub-id pub-id-type="doi">10.1126/science.1700865</pub-id><pub-id pub-id-type="pmid">1700865</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iordanskiy</surname> <given-names>S.</given-names></name> <name><surname>Waltke</surname> <given-names>M.</given-names></name> <name><surname>Feng</surname> <given-names>Y.</given-names></name> <name><surname>Wood</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>Subtype-associated differences in HIV-1 reverse transcription affect the viral replication</article-title>. <source>Retrovirology</source> <volume>7</volume>:<fpage>85</fpage>. <pub-id pub-id-type="doi">10.1186/1742-4690-7-85</pub-id><pub-id pub-id-type="pmid">20939905</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Julias</surname> <given-names>J. G.</given-names></name> <name><surname>McWilliams</surname> <given-names>M. J.</given-names></name> <name><surname>Sarafianos</surname> <given-names>S. G.</given-names></name> <name><surname>Alvord</surname> <given-names>W. G.</given-names></name> <name><surname>Arnold</surname> <given-names>E.</given-names></name> <name><surname>Hughes</surname> <given-names>S. H.</given-names></name></person-group> (<year>2003</year>). <article-title>Mutation of amino acids in the connection domain of human immunodeficiency virus type 1 reverse transcriptase that contact the template-primer affects RNase H activity</article-title>. <source>J. Virol.</source> <volume>77</volume>, <fpage>8548</fpage>&#x02013;<lpage>8554</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.77.15.8548-8554.2003</pub-id><pub-id pub-id-type="pmid">12857924</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaleebu</surname> <given-names>P.</given-names></name> <name><surname>French</surname> <given-names>N.</given-names></name> <name><surname>Mahe</surname> <given-names>C.</given-names></name> <name><surname>Yirrell</surname> <given-names>D.</given-names></name> <name><surname>Watera</surname> <given-names>C.</given-names></name> <name><surname>Lyagoba</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2002</year>). <article-title>Effect of human immunodeficiency virus (HIV) type 1 envelope subtypes A and D on disease progression in a large cohort of HIV-1&#x02013;positive persons in uganda</article-title>. <source>J. Infect. Dis.</source> <volume>185</volume>, <fpage>1244</fpage>&#x02013;<lpage>1250</lpage>. <pub-id pub-id-type="doi">10.1086/340130</pub-id><pub-id pub-id-type="pmid">12001041</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kantor</surname> <given-names>R.</given-names></name> <name><surname>Katzenstein</surname> <given-names>D. A.</given-names></name> <name><surname>Efron</surname> <given-names>B.</given-names></name> <name><surname>Carvalho</surname> <given-names>A. P.</given-names></name> <name><surname>Wynhoven</surname> <given-names>B.</given-names></name> <name><surname>Cane</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Impact of HIV-1 subtype and antiretroviral therapy on protease and reverse transcriptase genotype: results of a global collaboration</article-title>. <source>PLoS Med.</source> <volume>2</volume>:<fpage>e112</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pmed.0020112</pub-id><pub-id pub-id-type="pmid">15839752</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katoh</surname> <given-names>K.</given-names></name> <name><surname>Standley</surname> <given-names>D. M.</given-names></name></person-group> (<year>2013</year>). <article-title>MAFFT multiple sequence alignment software version 7: improvements in performance and usability</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>772</fpage>&#x02013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mst010</pub-id><pub-id pub-id-type="pmid">23329690</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kiguoya</surname> <given-names>M. W.</given-names></name> <name><surname>Mann</surname> <given-names>J. K.</given-names></name> <name><surname>Chopera</surname> <given-names>D.</given-names></name> <name><surname>Gounder</surname> <given-names>K.</given-names></name> <name><surname>Lee</surname> <given-names>G. Q.</given-names></name> <name><surname>Hunt</surname> <given-names>P. W.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Subtype-specific differences in gag-protease-driven replication capacity are consistent with inter-subtype differences in HIV-1 disease progression</article-title>. <source>J. Virol.</source> <volume>91</volume>:<fpage>e00253</fpage>-e17. <pub-id pub-id-type="doi">10.1128/JVI.00253-17</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kiwanuka</surname> <given-names>N.</given-names></name> <name><surname>Laeyendecker</surname> <given-names>O.</given-names></name> <name><surname>Robb</surname> <given-names>M.</given-names></name> <name><surname>Kigozi</surname> <given-names>G.</given-names></name> <name><surname>Arroyo</surname> <given-names>M.</given-names></name> <name><surname>McCutchan</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Effect of human immunodeficiency virus Type 1 (HIV-1) subtype on disease progression in persons from Rakai, Uganda, with incident HIV-1 infection</article-title>. <source>J. Infect. Dis.</source> <volume>197</volume>, <fpage>707</fpage>&#x02013;<lpage>713</lpage>. <pub-id pub-id-type="doi">10.1086/527416</pub-id><pub-id pub-id-type="pmid">18266607</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Livingstone</surname> <given-names>C. D.</given-names></name> <name><surname>Barton</surname> <given-names>C. J.</given-names></name></person-group> (<year>1993</year>). <article-title>Protein sequence alignments: a strategy for the hierarchical analysis of sequence conservation</article-title>. <source>Cabios</source> <volume>9</volume>, <fpage>745</fpage>&#x02013;<lpage>756</lpage>. <pub-id pub-id-type="pmid">8143162</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matsui</surname> <given-names>M.</given-names></name> <name><surname>Tomita</surname> <given-names>M.</given-names></name> <name><surname>Kanai</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Comprehensive computational analysis of bacterial CRP/FNR superfamily and its target motifs reveals stepwise evolution of transcriptional networks</article-title>. <source>Genome Biol. Evol</source>. <volume>5</volume>, <fpage>267</fpage>&#x02013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evt004</pub-id><pub-id pub-id-type="pmid">23315382</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Men&#x000E9;ndez-Arias</surname> <given-names>L.</given-names></name> <name><surname>Betancor</surname> <given-names>G.</given-names></name> <name><surname>Matamoros</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>HIV-1 reverse transcriptase connection subdomain mutations involved in resistance to approved non-nucleoside inhibitors</article-title>. <source>Antiviral Res.</source> <volume>92</volume>, <fpage>139</fpage>&#x02013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1016/j.antiviral.2011.08.020</pub-id><pub-id pub-id-type="pmid">21896288</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Metsky</surname> <given-names>H. C.</given-names></name> <name><surname>Matranga</surname> <given-names>C. B.</given-names></name> <name><surname>Wohl</surname> <given-names>S.</given-names></name> <name><surname>Schaffner</surname> <given-names>S. F.</given-names></name> <name><surname>Freije</surname> <given-names>C. A.</given-names></name> <name><surname>Winnicki</surname> <given-names>S. M.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Zika virus evolution and spread in the Americas</article-title>. <source>Nature</source> <volume>546</volume>, <fpage>411</fpage>&#x02013;<lpage>415</lpage>. <pub-id pub-id-type="doi">10.1038/nature22402</pub-id><pub-id pub-id-type="pmid">28538734</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Myers</surname> <given-names>R. E.</given-names></name> <name><surname>Pillay</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Analysis of natural sequence variation and covariation in human immunodeficiency virus type 1 integrase</article-title>. <source>J. Virol.</source> <volume>82</volume>, <fpage>9228</fpage>&#x02013;<lpage>9235</lpage>. <pub-id pub-id-type="doi">10.1128/JVI.01535-07</pub-id><pub-id pub-id-type="pmid">18596095</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nepusz</surname> <given-names>T.</given-names></name> <name><surname>Sasidharan</surname> <given-names>R.</given-names></name> <name><surname>Paccanaro</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>SCPS: a fast implementation of a spectral method for detecting protein families on a genome-wide scale</article-title>. <source>BMC Bioinformatics</source> <volume>11</volume>:<fpage>120</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-11-120</pub-id><pub-id pub-id-type="pmid">20214776</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ng</surname> <given-names>O. T.</given-names></name> <name><surname>Laeyendecker</surname> <given-names>O.</given-names></name> <name><surname>Redd</surname> <given-names>A. D.</given-names></name> <name><surname>Munshaw</surname> <given-names>S.</given-names></name> <name><surname>Grabowski</surname> <given-names>M. K.</given-names></name> <name><surname>Paquet</surname> <given-names>A. C.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>HIV type 1 polymerase gene polymorphisms are associated with phenotypic differences in replication capacity and disease progression</article-title>. <source>J. Infect. Dis.</source> <volume>209</volume>, <fpage>66</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1093/infdis/jit425</pub-id><pub-id pub-id-type="pmid">23922373</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paccanaro</surname> <given-names>A.</given-names></name> <name><surname>Casbon</surname> <given-names>J. A.</given-names></name> <name><surname>Saqi</surname> <given-names>M. A.</given-names></name></person-group> (<year>2006</year>). <article-title>Spectral clustering of protein sequences</article-title>. <source>Nucleic Acids Res</source>. <volume>34</volume>, <fpage>1571</fpage>&#x02013;<lpage>1580</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkj515</pub-id><pub-id pub-id-type="pmid">16547200</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pettersen</surname> <given-names>E. F.</given-names></name> <name><surname>Goddard</surname> <given-names>T. D.</given-names></name> <name><surname>Huang</surname> <given-names>C. C.</given-names></name> <name><surname>Couch</surname> <given-names>G. S.</given-names></name> <name><surname>Greenblatt</surname> <given-names>D. M.</given-names></name> <name><surname>Meng</surname> <given-names>E. C.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>UCSF chimera - a visualization system for exploratory research and analysis</article-title>. <source>J. Comput. Chem.</source> <volume>25</volume>, <fpage>1605</fpage>&#x02013;<lpage>1612</lpage>. <pub-id pub-id-type="doi">10.1002/jcc.20084</pub-id><pub-id pub-id-type="pmid">15264254</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Posada</surname> <given-names>D.</given-names></name> <name><surname>Crandall</surname> <given-names>A. K.</given-names></name> <name><surname>Crandall</surname> <given-names>K. A.</given-names></name></person-group> (<year>2002</year>). <article-title>The effect of recombination on the accuracy of phylogeny estimation</article-title>. <source>J. Mol. Evol.</source> <volume>54</volume>, <fpage>396</fpage>&#x02013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-001-0034-9</pub-id><pub-id pub-id-type="pmid">11847565</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Preston</surname> <given-names>B. D.</given-names></name> <name><surname>Poiesz</surname> <given-names>B. J.</given-names></name> <name><surname>Loeb</surname> <given-names>L. A.</given-names></name></person-group> (<year>1988</year>). <article-title>Fidelity of HIV-1 reverse transcriptase</article-title>. <source>Science</source> <volume>242</volume>, <fpage>1168</fpage>&#x02013;<lpage>1171</lpage>. <pub-id pub-id-type="doi">10.1126/science.2460924</pub-id><pub-id pub-id-type="pmid">2460924</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rambaut</surname> <given-names>A.</given-names></name> <name><surname>Posada</surname> <given-names>D.</given-names></name> <name><surname>Crandall</surname> <given-names>K. A.</given-names></name> <name><surname>Holmes</surname> <given-names>E. C.</given-names></name></person-group> (<year>2004</year>). <article-title>The causes and consequences of HIV evolution</article-title>. <source>Nat. Rev. Genet.</source> <volume>5</volume>, <fpage>52</fpage>&#x02013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1038/nrg1246</pub-id><pub-id pub-id-type="pmid">14708016</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Renjifo</surname> <given-names>B.</given-names></name> <name><surname>Gilbert</surname> <given-names>P.</given-names></name> <name><surname>Chaplin</surname> <given-names>B.</given-names></name> <name><surname>Msamanga</surname> <given-names>G.</given-names></name> <name><surname>Mwakagile</surname> <given-names>D.</given-names></name> <name><surname>Fawzi</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>Preferential in-utero transmission of HIV-1 subtype C as compared to HIV-1 subtype A or D</article-title>. <source>AIDS</source> <volume>18</volume>, <fpage>1629</fpage>&#x02013;<lpage>1636</lpage>. <pub-id pub-id-type="doi">10.1097/01.aids.0000131392.68597.34</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rhee</surname> <given-names>S.-Y.</given-names></name> <name><surname>Kantor</surname> <given-names>R.</given-names></name> <name><surname>Katzenstein</surname> <given-names>D. A.</given-names></name> <name><surname>Camacho</surname> <given-names>R.</given-names></name> <name><surname>Morris</surname> <given-names>L.</given-names></name> <name><surname>Sirivichayakul</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>HIV-1 pol mutation frequency by subtype and treatment experience: extension of the HIVseq program to seven non-B subtypes</article-title>. <source>AIDS</source> <volume>20</volume>, <fpage>643</fpage>&#x02013;<lpage>651</lpage>. <pub-id pub-id-type="doi">10.1097/01.aids.0000216363.36786.2b</pub-id><pub-id pub-id-type="pmid">16514293</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robertson</surname> <given-names>D. L.</given-names></name> <name><surname>Anderson</surname> <given-names>J. P.</given-names></name> <name><surname>Bradac</surname> <given-names>J. A.</given-names></name> <name><surname>Carr</surname> <given-names>J. K.</given-names></name> <name><surname>Foley</surname> <given-names>B.</given-names></name> <name><surname>Funkhouser</surname> <given-names>R. K.</given-names></name> <etal/></person-group>. (<year>2000</year>). <article-title>HIV-1 nomenclature proposal</article-title>. <source>Science</source> <volume>288</volume>, <fpage>55</fpage>&#x02013;<lpage>56</lpage>. <pub-id pub-id-type="doi">10.1126/science.288.5463.55d</pub-id><pub-id pub-id-type="pmid">10766634</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shannon</surname> <given-names>C. E.</given-names></name></person-group> (<year>1996</year>). <article-title>The mathematical theory of communication. <italic>MD Comput. Comput. Med</italic></article-title>. <source>Pract</source>. <volume>14</volume>, <fpage>306</fpage>&#x02013;<lpage>317</lpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shannon</surname> <given-names>P.</given-names></name> <name><surname>Markiel</surname> <given-names>A.</given-names></name> <name><surname>Ozier</surname> <given-names>O.</given-names></name> <name><surname>Baliga</surname> <given-names>N. S.</given-names></name> <name><surname>Wang</surname> <given-names>J. T.</given-names></name> <name><surname>Ramage</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Cytoscape: a software environment for integrated models of biomolecular interaction networks</article-title>. <source>Genome Res.</source> <volume>13</volume>, <fpage>2498</fpage>&#x02013;<lpage>2504</lpage>. <pub-id pub-id-type="doi">10.1101/gr.1239303</pub-id><pub-id pub-id-type="pmid">14597658</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharp</surname> <given-names>P. M.</given-names></name> <name><surname>Hahn</surname> <given-names>B. H.</given-names></name></person-group> (<year>2010</year>). <article-title>The evolution of HIV-1 and the origin of AIDS</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>365</volume>, <fpage>2487</fpage>&#x02013;<lpage>2494</lpage>. <pub-id pub-id-type="doi">10.1098/rstb.2010.0031</pub-id><pub-id pub-id-type="pmid">20643738</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simon-Loriere</surname> <given-names>E.</given-names></name> <name><surname>Faye</surname> <given-names>O.</given-names></name> <name><surname>Faye</surname> <given-names>O.</given-names></name> <name><surname>Koivogui</surname> <given-names>L.</given-names></name> <name><surname>Magassouba</surname> <given-names>N.</given-names></name> <name><surname>Keita</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Distinct lineages of ebola virus in Guinea during the 2014 West African epidemic</article-title>. <source>Nature</source> <volume>524</volume>, <fpage>102</fpage>&#x02013;<lpage>104</lpage>. <pub-id pub-id-type="doi">10.1038/nature14612</pub-id><pub-id pub-id-type="pmid">26106863</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stamatakis</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>1312</fpage>&#x02013;<lpage>1313</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu033</pub-id><pub-id pub-id-type="pmid">24451623</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Telesnitsky</surname> <given-names>A.</given-names></name> <name><surname>Goff</surname> <given-names>S.</given-names></name></person-group> (<year>1997</year>). <source>Reverse Transcriptase and the Generation of Retroviral DNA</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Cold Spring Harbor Laboratory Press</publisher-name>.</citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tong</surname> <given-names>Y.-G.</given-names></name> <name><surname>Shi</surname> <given-names>W.-F.</given-names></name> <name><surname>Di Liu</surname></name> <name><surname>Qian</surname> <given-names>J.</given-names></name> <name><surname>Liang</surname> <given-names>L.</given-names></name> <name><surname>Bo</surname> <given-names>X.-C.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Genetic diversity and evolutionary dynamics of Ebola virus in Sierra Leone</article-title>. <source>Nature</source> <volume>524</volume>, <fpage>93</fpage>&#x02013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1038/nature14490</pub-id><pub-id pub-id-type="pmid">25970247</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vasan</surname> <given-names>A.</given-names></name> <name><surname>Renjifo</surname> <given-names>B.</given-names></name> <name><surname>Hertzmark</surname> <given-names>E.</given-names></name> <name><surname>Chaplin</surname> <given-names>B.</given-names></name> <name><surname>Msamanga</surname> <given-names>G.</given-names></name> <name><surname>Essex</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>Different rates of disease progression of HIV type 1 infection in Tanzania based on infecting subtype</article-title>. <source>Clin. Infect. Dis.</source> <volume>42</volume>, <fpage>843</fpage>&#x02013;<lpage>852</lpage>. <pub-id pub-id-type="doi">10.1086/499952</pub-id><pub-id pub-id-type="pmid">16477563</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Velazquez-Campoy</surname> <given-names>A.</given-names></name> <name><surname>Todd</surname> <given-names>M. J.</given-names></name> <name><surname>Vega</surname> <given-names>S.</given-names></name> <name><surname>Freire</surname> <given-names>E.</given-names></name></person-group> (<year>2001</year>). <article-title>Catalytic efficiency and vitality of HIV-1 proteases from African viral subtypes</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>98</volume>, <fpage>6062</fpage>&#x02013;<lpage>6067</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.111152698</pub-id><pub-id pub-id-type="pmid">11353856</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>von Wyl</surname> <given-names>V.</given-names></name> <name><surname>Ehteshami</surname> <given-names>M.</given-names></name> <name><surname>Demeter</surname> <given-names>L. M.</given-names></name> <name><surname>B&#x000FC;rgisser</surname> <given-names>P.</given-names></name> <name><surname>Nijhuis</surname> <given-names>M.</given-names></name> <name><surname>Symons</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>HIV-1 reverse transcriptase connection domain mutations: dynamics of emergence and implications for success of combination antiretroviral therapy</article-title>. <source>Clin. Infect. Dis.</source> <volume>51</volume>, <fpage>620</fpage>&#x02013;<lpage>628</lpage>. <pub-id pub-id-type="doi">10.1086/655764</pub-id><pub-id pub-id-type="pmid">20666602</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Waterhouse</surname> <given-names>A. M.</given-names></name> <name><surname>Procter</surname> <given-names>J. B.</given-names></name> <name><surname>Martin</surname> <given-names>D. M. A.</given-names></name> <name><surname>Clamp</surname> <given-names>M.</given-names></name> <name><surname>Barton</surname> <given-names>G. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Jalview version 2-A multiple sequence alignment editor and analysis workbench</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>1189</fpage>&#x02013;<lpage>1191</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp033</pub-id><pub-id pub-id-type="pmid">19151095</pub-id></citation>
</ref>
</ref-list>
<glossary>
<def-list>
<title>Abbreviations</title>
<def-item><term>HIV-1</term>
<def><p><italic>Human immunodeficiency virus 1</italic></p></def></def-item>
<def-item><term>RT</term>
<def><p>reverse transcriptase</p></def></def-item>
<def-item><term>SSN</term>
<def><p>sequence similarity network</p></def></def-item>
<def-item><term>CRE</term>
<def><p>cumulative relative entropy.</p></def></def-item>
</def-list>
</glossary>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This work was supported, in part, by research funds from the Yamagata Prefectural Government and Tsuruoka City, Japan. The funding bodies played no role in the study design, data collection or analysis, decision to publish, or preparation of the manuscript.</p>
</fn>
</fn-group>
</back>
</article>