<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">766496</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.766496</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>An Information-Entropy Position-Weighted <italic>K</italic>-Mer Relative Measure for Whole Genome Phylogeny Reconstruction</article-title>
<alt-title alt-title-type="left-running-head">Wu et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Measure for Whole Genome Phylogeny Reconstruction</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Wu</surname>
<given-names>Yao-Qun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1460517/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Yu</surname>
<given-names>Zu-Guo</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/873331/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Tang</surname>
<given-names>Run-Bin</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1478156/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Han</surname>
<given-names>Guo-Sheng</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/797721/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Anh</surname>
<given-names>Vo V.</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Hunan Key Laboratory for Computation and Simulation in Science and Engineering and Key Laboratory of Intelligent Computing and Information Processing of Ministry of Education, Xiangtan University, <addr-line>Hunan</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Provincial Key Laboratory of Informational Service for Rural Area of Southwestern Hunan, Shaoyang University, <addr-line>Shaoyang</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Faculty of Science, Engineering and Technology, Swinburne University of Technology, <addr-line>Hawthorn</addr-line>, <addr-line>VIC</addr-line>, <country>Australia</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/560593/overview">Juan Wang</ext-link>, Inner Mongolia University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/562375/overview">Liang Cheng</ext-link>, Harbin Medical University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1465511/overview">Yanjuan Li</ext-link>, Quzhou University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Zu-Guo Yu, <email>yuzuguo@aliyun.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>766496</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Wu, Yu, Tang, Han and Anh.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Wu, Yu, Tang, Han and Anh</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Alignment methods have faced disadvantages in sequence comparison and phylogeny reconstruction due to their high computational costs in handling time and space complexity. On the other hand, alignment-free methods incur low computational costs and have recently gained popularity in the field of bioinformatics. Here we propose a new alignment-free method for phylogenetic tree reconstruction based on whole genome sequences. A key component is a measure called <italic>information-entropy position-weighted k-mer relative measure</italic> (IEPWRMkmer), which combines the position-weighted measure of <italic>k</italic>-mers proposed by our group and the information entropy of frequency of <italic>k</italic>-mers. The Manhattan distance is used to calculate the pairwise distance between species. Finally, we use the Neighbor-Joining method to construct the phylogenetic tree. To evaluate the performance of this method, we perform phylogenetic analysis on two datasets used by other researchers. The results demonstrate that the <italic>IEPWRMkmer</italic> method is efficient and reliable. The source codes of our method are provided at <ext-link ext-link-type="uri" xlink:href="https://github.com/">https://github.com/</ext-link> wuyaoqun37/IEPWRMkmer.</p>
</abstract>
<kwd-group>
<kwd>alignment-free method</kwd>
<kwd>k-mer relative distance</kwd>
<kwd>information entropy</kwd>
<kwd>phylogenetic analysis</kwd>
<kwd>genome</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The reconstruction of a phylogenetic tree is a primary problem in evolutionary biology. Sequence alignment is a key step in the reconstruction, aiming to identify the homology of sequences and uncover phylogenetic relationships in sequences. Traditional sequence comparison is based on pairwise or multiple sequence alignment (<xref ref-type="bibr" rid="B7">Felsenstein and Felenstein, 2004</xref>; <xref ref-type="bibr" rid="B19">Morrison, 2006</xref>) and was implemented by software packages such as BLAST (<xref ref-type="bibr" rid="B1">Altschul et&#x20;al., 1990</xref>), ClustalW (<xref ref-type="bibr" rid="B29">Thompson et&#x20;al., 1994</xref>), and MrBayes (<xref ref-type="bibr" rid="B24">Ronquist et&#x20;al., 2012</xref>). However, the methods based on sequence alignment have some disadvantages, including high computational cost in handling the time and space complexity of the algorithm. Therefore, alignment-free methods have been proposed to overcome these problems (<xref ref-type="bibr" rid="B38">Zielezinski et&#x20;al., 2017</xref>). The computational cost of alignment-free methods is low because they are generally of linear complexity (<xref ref-type="bibr" rid="B8">Fox et&#x20;al<italic>.</italic>, 1977</xref>).</p>
<p>Several alignment-free methods for sequence comparison are based on word counts (<xref ref-type="bibr" rid="B3">Blaisdell, 1986</xref>; <xref ref-type="bibr" rid="B11">H&#xf6;hl et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B31">Wang et&#x20;al., 2016</xref>). A key idea is to use the close distribution of <italic>k</italic>-mers to imply the high correlation degree, hence the similarity of the sequences. The methods have been implemented in software tools, such as FFP (<xref ref-type="bibr" rid="B26">Sims et&#x20;al., 2009</xref>), kWIP (<xref ref-type="bibr" rid="B20">Murray et&#x20;al., 2017</xref>), CVtree (<xref ref-type="bibr" rid="B22">Qi et&#x20;al., 2004</xref>), and DLtree (<xref ref-type="bibr" rid="B32">Wu et&#x20;al., 2017</xref>). Many <italic>k</italic>-mer methods transform the input sequence into a frequency vector of <italic>k</italic>-mers, then define the distance of the sequences by that of the frequency vector of <italic>k</italic>-mers (<xref ref-type="bibr" rid="B22">Qi et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B32">Wu et&#x20;al., 2017</xref>). To reduce the statistical dependence between adjacent word matches, Spaced-Words (<xref ref-type="bibr" rid="B15">Leimeister and Boden, 2014</xref>) proposed to use spaced words, which are defined by patterns of matches without reference to positions. Some alignment-free methods are based on match length, which defines the distance between sequences based on the length of substring matches between two sequences. These include the shortest unique substring method (<xref ref-type="bibr" rid="B9">Haubold et&#x20;al., 2005</xref>), ACS (<xref ref-type="bibr" rid="B30">Ulitsky et&#x20;al<italic>.</italic> 2006</xref>), UA (<xref ref-type="bibr" rid="B5">Comin and Verzotto, 2012</xref>), and ALFRED (<xref ref-type="bibr" rid="B28">Thankachan et&#x20;al<italic>.</italic> 2016</xref>). In addition, graphical representation was used to construct the probability distribution of a DNA sequence (<xref ref-type="bibr" rid="B35">Yu et&#x20;al., 2011</xref>). The chaos game representation transforms the distribution of characters in a DNA sequence into the distribution of nodes in a graph (<xref ref-type="bibr" rid="B10">Hoang et&#x20;al<italic>.</italic> 2016</xref>; <xref ref-type="bibr" rid="B34">Yin, 2017</xref>; <xref ref-type="bibr" rid="B18">Mendizabal-Ruiz et&#x20;al., 2018</xref>). Many researchers considered extracting the position information of a <italic>k</italic>-mer (<xref ref-type="bibr" rid="B12">Huang and Wang, 2011</xref>; <xref ref-type="bibr" rid="B6">Ding et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B27">Tang et&#x20;al., 2014</xref>). <xref ref-type="bibr" rid="B6">Ding et&#x20;al. (2013)</xref> used the average interval distance of normalized <italic>k</italic>-mers to capture evolutionary information for sequence comparison. <xref ref-type="bibr" rid="B27">Tang et&#x20;al. (2014)</xref> presented the average relative distance of normalized <italic>k</italic>-mers to improve the method of <xref ref-type="bibr" rid="B6">Ding et&#x20;al. (2013)</xref>. <xref ref-type="bibr" rid="B17">Ma et&#x20;al. (2020)</xref> proposed the <italic>PWKmer</italic> method, which combines the <italic>k</italic>-mer counts and <italic>k</italic>-mer position distributions for phylogenetic analysis.</p>
<p>In this work, we propose a new alignment-free method which combines the position-weighted measure of <italic>k</italic>-mers proposed by <xref ref-type="bibr" rid="B17">Ma et&#x20;al. (2020)</xref> and the information entropy of frequency of <italic>k</italic>-mers to obtain phylogenetic information for sequence comparison. It is named <italic>information-entropy position-weighted k-mer relative measure</italic> (IEPWRMkmer). To evaluate the performance of this method, we carry out phylogenetic analysis on two data sets used by other researchers.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>Materials and Methods</title>
<sec id="s2-1">
<title>Genomic Datasets</title>
<sec id="s2-1-1">
<title>Dataset 1</title>
<p>The first dataset for analysis consists of the same whole genome DNA sequences of 30 mammalian species studied in <xref ref-type="bibr" rid="B16">Li et&#x20;al. (2001)</xref>, <xref ref-type="bibr" rid="B21">Otu and Sayood (2003)</xref>, and <xref ref-type="bibr" rid="B27">Tang et&#x20;al. (2014)</xref>. The accession numbers, species, and species name are listed in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. All sequences were downloaded from NCBI GenBank.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Names, species, and accession numbers for mitochondrial genomes of 30 mammalian species.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No</th>
<th align="center">Accession no</th>
<th align="center">Species</th>
<th align="center">Sequence name</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">AJ002189</td>
<td align="left">
<italic>Sus scrofa</italic>
</td>
<td align="left">Pig</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">AJ010957</td>
<td align="left">
<italic>Homo sapiens</italic>
</td>
<td align="left">
<italic>Hippopotamus</italic>
</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">AJ001588</td>
<td align="left">
<italic>Pan troglodytes</italic>
</td>
<td align="left">Rabbit</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">U96639</td>
<td align="left">
<italic>Canis familiaris</italic>
</td>
<td align="left">Dog</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">AF010406</td>
<td align="left">
<italic>Ovis aries</italic>
</td>
<td align="left">Sheep</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">V00662</td>
<td align="left">
<italic>Homo sapiens</italic>
</td>
<td align="left">Human</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">U20753</td>
<td align="left">
<italic>Felis catus</italic>
</td>
<td align="left">Cat</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">X72004</td>
<td align="left">
<italic>Halichoerus grypus</italic>
</td>
<td align="left">Gray seal</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">D38115</td>
<td align="left">
<italic>Pongo pygmaeus</italic>
</td>
<td align="left">Orangutan</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">V00654</td>
<td align="left">
<italic>Bos taurus</italic>
</td>
<td align="left">Cow</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">X97337</td>
<td align="left">
<italic>Equus asinus</italic>
</td>
<td align="left">Donkey</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">D38116</td>
<td align="left">
<italic>Pan troglodytes</italic>
</td>
<td align="left">Common chimpanzee</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">D38113</td>
<td align="left">
<italic>Pan paniscus</italic>
</td>
<td align="left">Pigmy chimpanzee</td>
</tr>
<tr>
<td align="left">14</td>
<td align="left">Z29573</td>
<td align="left">
<italic>Didelphis virginiana</italic>
</td>
<td align="left">Opossum</td>
</tr>
<tr>
<td align="left">15</td>
<td align="left">Y10524</td>
<td align="left">
<italic>Macropus robustus</italic>
</td>
<td align="left">Wallaroo</td>
</tr>
<tr>
<td align="left">16</td>
<td align="left">X99256</td>
<td align="left">
<italic>Hylobates lar</italic>
</td>
<td align="left">Gibbon</td>
</tr>
<tr>
<td align="left">17</td>
<td align="left">Y18001</td>
<td align="left">
<italic>Papio hamadryas</italic>
</td>
<td align="left">Baboon</td>
</tr>
<tr>
<td align="left">18</td>
<td align="left">X97336</td>
<td align="left">
<italic>Rhinoceros unicornis</italic>
</td>
<td align="left">Indian rhinoceros</td>
</tr>
<tr>
<td align="left">19</td>
<td align="left">Y07726</td>
<td align="left">
<italic>Ceratotherium simum</italic>
</td>
<td align="left">White rhinoceros</td>
</tr>
<tr>
<td align="left">20</td>
<td align="left">X63726</td>
<td align="left">
<italic>Phoca vitulina</italic>
</td>
<td align="left">Harbor seal</td>
</tr>
<tr>
<td align="left">21</td>
<td align="left">AJ238588</td>
<td align="left">
<italic>Sciurus vulgaris</italic>
</td>
<td align="left">Squirrel</td>
</tr>
<tr>
<td align="left">22</td>
<td align="left">AJ001562</td>
<td align="left">
<italic>Glis glis</italic>
</td>
<td align="left">Fat dormouse</td>
</tr>
<tr>
<td align="left">23</td>
<td align="left">AJ222767</td>
<td align="left">
<italic>Cavia porcellus</italic>
</td>
<td align="left">Guinea pig</td>
</tr>
<tr>
<td align="left">24</td>
<td align="left">X79547</td>
<td align="left">
<italic>Equus caballus</italic>
</td>
<td align="left">Horse</td>
</tr>
<tr>
<td align="left">25</td>
<td align="left">X14848</td>
<td align="left">
<italic>Rattus norvegicus</italic>
</td>
<td align="left">Rat</td>
</tr>
<tr>
<td align="left">26</td>
<td align="left">V00711</td>
<td align="left">
<italic>Mus musculus</italic>
</td>
<td align="left">Mouse</td>
</tr>
<tr>
<td align="left">27</td>
<td align="left">D38114</td>
<td align="left">
<italic>Gorilla gorilla</italic>
</td>
<td align="left">
<italic>Gorilla</italic>
</td>
</tr>
<tr>
<td align="left">28</td>
<td align="left">X61145</td>
<td align="left">
<italic>Balenoptera physalus</italic>
</td>
<td align="left">Fin whale</td>
</tr>
<tr>
<td align="left">29</td>
<td align="left">X72204</td>
<td align="left">
<italic>Balenoptera musculus</italic>
</td>
<td align="left">Blue whale</td>
</tr>
<tr>
<td align="left">30</td>
<td align="left">X83427</td>
<td align="left">
<italic>Ornithorhyncus anatinus</italic>
</td>
<td align="left">Platypus</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2-1-2">
<title>Dataset 2</title>
<p>The second dataset for analysis is the HIV-1 dataset studied in <xref ref-type="bibr" rid="B17">Ma et&#x20;al. (2020)</xref>. This dataset contains 43 HIV genome sequences used in <xref ref-type="bibr" rid="B33">Wu et&#x20;al. (2007)</xref> and a controversial taxonomic sequence used in <xref ref-type="bibr" rid="B4">Chang et&#x20;al. (2014)</xref>. The dataset includes subtypes A, B, C, D, F, G, J, K, and H of the HIV-1 M, O, N groups and the CPZ sequence. The area, accession numbers, and subtypes are listed in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. All these sequences were downloaded from NCBI GenBank.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Accession numbers, subtype, and area for 44&#x20;HIV-1.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">No</th>
<th align="center">Area</th>
<th align="center">Accession no</th>
<th align="center">Subtype</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Belgium (DRC)</td>
<td align="left">AF084936</td>
<td align="left">G</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Finland (Kenya)</td>
<td align="left">AF061641</td>
<td align="left">G</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Sweden (DRC)</td>
<td align="left">AF061642</td>
<td align="left">G</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Belgium</td>
<td align="left">AF190128</td>
<td align="left">H</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Belgium</td>
<td align="left">AF190127</td>
<td align="left">H</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">Cent. Afr. Rep</td>
<td align="left">AF005496</td>
<td align="left">H</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">Tanzania</td>
<td align="left">AF447763</td>
<td align="left">CPZ</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">Cameroon</td>
<td align="left">L20571</td>
<td align="left">O</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">Senegal</td>
<td align="left">AJ302647</td>
<td align="left">O</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">Cameroon</td>
<td align="left">L20587</td>
<td align="left">O</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">Cameroon</td>
<td align="left">AY169812</td>
<td align="left">O</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">India</td>
<td align="left">AF067155</td>
<td align="left">C</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">South Africa</td>
<td align="left">AY772699</td>
<td align="left">C</td>
</tr>
<tr>
<td align="left">14</td>
<td align="left">Ethiopia</td>
<td align="left">U46016</td>
<td align="left">C</td>
</tr>
<tr>
<td align="left">15</td>
<td align="left">Brazil</td>
<td align="left">U52953</td>
<td align="left">C</td>
</tr>
<tr>
<td align="left">16</td>
<td align="left">Cameroon</td>
<td align="left">AY371157</td>
<td align="left">D</td>
</tr>
<tr>
<td align="left">17</td>
<td align="left">DRC</td>
<td align="left">K03454</td>
<td align="left">D</td>
</tr>
<tr>
<td align="left">18</td>
<td align="left">Uganda</td>
<td align="left">U88824</td>
<td align="left">D</td>
</tr>
<tr>
<td align="left">19</td>
<td align="left">Somalia</td>
<td align="left">AF069670</td>
<td align="left">A1</td>
</tr>
<tr>
<td align="left">20</td>
<td align="left">Uganda</td>
<td align="left">AF484509</td>
<td align="left">A1</td>
</tr>
<tr>
<td align="left">21</td>
<td align="left">Uganda</td>
<td align="left">U51190</td>
<td align="left">A1</td>
</tr>
<tr>
<td align="left">22</td>
<td align="left">Kenya</td>
<td align="left">AF004885</td>
<td align="left">A1</td>
</tr>
<tr>
<td align="left">23</td>
<td align="left">DRC</td>
<td align="left">AF286238</td>
<td align="left">A2</td>
</tr>
<tr>
<td align="left">24</td>
<td align="left">Cyprus</td>
<td align="left">AF286237</td>
<td align="left">A2</td>
</tr>
<tr>
<td align="left">25</td>
<td align="left">Sweden</td>
<td align="left">AF082395</td>
<td align="left">J</td>
</tr>
<tr>
<td align="left">26</td>
<td align="left">Sweden</td>
<td align="left">AF082394</td>
<td align="left">J</td>
</tr>
<tr>
<td align="left">27</td>
<td align="left">Cameroon</td>
<td align="left">AJ249239</td>
<td align="left">K</td>
</tr>
<tr>
<td align="left">28</td>
<td align="left">DRC</td>
<td align="left">AJ249235</td>
<td align="left">K</td>
</tr>
<tr>
<td align="left">29</td>
<td align="left">Cameroon</td>
<td align="left">AJ249237</td>
<td align="left">F2</td>
</tr>
<tr>
<td align="left">30</td>
<td align="left">Cameroon</td>
<td align="left">AY371158</td>
<td align="left">F2</td>
</tr>
<tr>
<td align="left">31</td>
<td align="left">Cameroon</td>
<td align="left">AJ249236</td>
<td align="left">F2</td>
</tr>
<tr>
<td align="left">32</td>
<td align="left">Cameroon</td>
<td align="left">AF377956</td>
<td align="left">F2</td>
</tr>
<tr>
<td align="left">33</td>
<td align="left">Finland</td>
<td align="left">AF075703</td>
<td align="left">F1</td>
</tr>
<tr>
<td align="left">34</td>
<td align="left">France</td>
<td align="left">AJ249238</td>
<td align="left">F1</td>
</tr>
<tr>
<td align="left">35</td>
<td align="left">Brazil</td>
<td align="left">AF005494</td>
<td align="left">F1</td>
</tr>
<tr>
<td align="left">36</td>
<td align="left">Belgium (DRC)</td>
<td align="left">AF077336</td>
<td align="left">F1</td>
</tr>
<tr>
<td align="left">37</td>
<td align="left">Cameroon</td>
<td align="left">AJ271370</td>
<td align="left">N</td>
</tr>
<tr>
<td align="left">38</td>
<td align="left">Cameroon</td>
<td align="left">AY532635</td>
<td align="left">N</td>
</tr>
<tr>
<td align="left">39</td>
<td align="left">Cameroon</td>
<td align="left">AJ006022</td>
<td align="left">N</td>
</tr>
<tr>
<td align="left">40</td>
<td align="left">Netherlands</td>
<td align="left">AY423387</td>
<td align="left">B</td>
</tr>
<tr>
<td align="left">41</td>
<td align="left">Thailand</td>
<td align="left">AY173951</td>
<td align="left">B</td>
</tr>
<tr>
<td align="left">42</td>
<td align="left">Australia</td>
<td align="left">Gray seal</td>
<td align="left">B</td>
</tr>
<tr>
<td align="left">43</td>
<td align="left">France</td>
<td align="left">K03455</td>
<td align="left">B</td>
</tr>
<tr>
<td align="left">44</td>
<td align="left">U.S.</td>
<td align="left">AY331295</td>
<td align="left">B</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We use two approaches to validate the method. First, we use the Robinson-Foulds (RF) distance to compare our method with other alignment-free methods. Second, we use the bootstrap method to construct consensus trees and show the stability of the trees obtained by our method.</p>
</sec>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>Methods</title>
<p>Let <italic>S</italic> &#x3d; <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> be a DNA sequence with length <italic>L</italic>, <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>is a <italic>k</italic>-mer, where <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>&#x2208;(A,T,C,G). If the <italic>k</italic>-mer <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> occurs in <italic>S</italic>, we denote by <inline-formula id="inf5">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> the vector composed of the positions of <inline-formula id="inf6">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> in this given sequence and by <inline-formula id="inf7">
<mml:math id="m7">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> its <italic>i</italic>th element. If the <italic>k</italic>-mer <inline-formula id="inf8">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> does not occur in <italic>S</italic>,<inline-formula id="inf9">
<mml:math id="m9">
<mml:mrow>
<mml:mtext>&#xa0;we&#xa0;set&#xa0;&#xa0;</mml:mtext>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>&#x3d;(0). For example, for the DNA sequence GTA&#x200b;ACC&#x200b;TGA&#x200b;ACG&#x200b;TAC&#x200b;TTG&#x200b;GA with length 20, we list all 2-mer position vectors:</p>
<p>
<italic>P</italic>
<sub>
<italic>AA</italic>
</sub>&#x3d;(3,9); <italic>P</italic>
<sub>
<italic>AC</italic>
</sub>&#x3d;(4,10,14); <italic>P</italic>
<sub>
<italic>AG</italic>
</sub>&#x3d; (0); <italic>P</italic>
<sub>
<italic>AT</italic>
</sub>&#x3d; (0); <italic>P</italic>
<sub>
<italic>CA</italic>
</sub>&#x3d;(0); <italic>P</italic>
<sub>
<italic>CC</italic>
</sub>&#x3d;(5); <italic>P</italic>
<sub>
<italic>CG</italic>
</sub>&#x3d;(11); <italic>P</italic>
<sub>
<italic>CT</italic>
</sub>&#x3d;(6,15); <italic>P</italic>
<sub>
<italic>GA</italic>
</sub>&#x3d;(8,19); <italic>P</italic>
<sub>
<italic>GC</italic>
</sub>&#x3d;(0); <italic>P</italic>
<sub>
<italic>GG</italic>
</sub>&#x3d;(18); <italic>P</italic>
<sub>
<italic>GT</italic>
</sub>&#x3d;(1,12); <italic>P</italic>
<sub>
<italic>TA</italic>
</sub>&#x3d;(2,13); <italic>P</italic>
<sub>
<italic>TC</italic>
</sub> &#x3d; 0; <italic>P</italic>
<sub>
<italic>TG</italic>
</sub>&#x3d;(7,17); <italic>P</italic>
<sub>
<italic>TT</italic>
</sub>&#x3d;(16).</p>
<p>In this example, the 2-mers AG, AT, CA, GC, and TC do not appear. For each <italic>k</italic>-mer, its position vector provides its position distribution information in the sequence. One can use the <italic>k</italic>-mer position vectors to reconstruct the DNA sequence (<xref ref-type="bibr" rid="B17">Ma et&#x20;al., 2020</xref>).</p>
<p>
<xref ref-type="bibr" rid="B17">Ma et&#x20;al. (2020)</xref> defined the position-weighted measure <inline-formula id="inf10">
<mml:math id="m10">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of <inline-formula id="inf11">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> based on its position in the sequence as<disp-formula id="e1">
<mml:math id="m12">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2260;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>n</italic> is the length of the vector <inline-formula id="inf12">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. Actually <inline-formula id="inf13">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mi>L</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> means the position weight of <inline-formula id="inf14">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> in the given sequence with length&#x20;<italic>L</italic>.</p>
<p>We denote by <italic>N</italic> the number of sequences in a dataset. In order to characterize the importance of <italic>k</italic>-mers in the whole dataset, we count the number <italic>m</italic> of the sequences that contain a <italic>k</italic>-mer <inline-formula id="inf15">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. Then the occurrence frequency <italic>F</italic>
<inline-formula id="inf16">
<mml:math id="m17">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of this <italic>k</italic>-mer in the whole dataset is defined as <italic>m</italic>/<italic>N</italic>. We introduce the Shannon entropy <italic>H</italic>(<inline-formula id="inf17">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>) of frequency <italic>F</italic>(<inline-formula id="inf18">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>) defined by <xref ref-type="bibr" rid="B20">Murray et&#x20;al. (2017)</xref> as<disp-formula id="e2">
<mml:math id="m20">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>log</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>log</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>F</italic> stands for <italic>F</italic> (<inline-formula id="inf19">
<mml:math id="m21">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>).</p>
<p>In this study, we aim to get more DNA phylogenetic information by combining the above two methods and defining<disp-formula id="e3">
<mml:math id="m22">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>Here, we regard Shannon entropy <italic>H</italic> (<inline-formula id="inf20">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>) as another weight.</p>
<p>For a fixed <italic>K</italic>, there are 4<sup>
<italic>K</italic>
</sup> <italic>k</italic>-mers. For each <italic>k</italic>-mer <inline-formula id="inf21">
<mml:math id="m24">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, we can calculate the corresponding <inline-formula id="inf22">
<mml:math id="m25">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, then arrange 4<sup>
<italic>K</italic>
</sup> of these <inline-formula id="inf23">
<mml:math id="m26">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x22ef;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> to get a feature representation vector (<inline-formula id="inf24">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mn>4</mml:mn>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>) according to the alphabet order of the 4<sup>
<italic>K</italic>
</sup> <italic>k</italic>-mers for each genome.</p>
<p>For two given genome sequences <italic>A</italic> and <italic>B</italic>, we can obtain <inline-formula id="inf25">
<mml:math id="m28">
<mml:mrow>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mtext>A</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> &#x3d; <inline-formula id="inf26">
<mml:math id="m29">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi>A</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mn>2</mml:mn>
<mml:mi>A</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mn>4</mml:mn>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mi>A</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf27">
<mml:math id="m30">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mtext>B</mml:mtext>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi>B</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mn>2</mml:mn>
<mml:mi>B</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mn>4</mml:mn>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mi>B</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> by the method. We use the Manhattan distance to calculate the pairwise distance between these two genome sequences:<disp-formula id="e4">
<mml:math id="m31">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mn>4</mml:mn>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>A</mml:mi>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>B</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>.<label>(4)</label>
</disp-formula>
</p>
<p>For a given dataset, we can derive a distance matrix by <xref ref-type="disp-formula" rid="e4">Eq. 4</xref>. This distance matrix contains the sequence similarity information. After obtaining the distance matrix, we insert it into the mega 7.0 software (<xref ref-type="bibr" rid="B13">Sudhir et&#x20;al., 2016</xref>) and use Neighbor-Joining (NJ) program (<xref ref-type="bibr" rid="B25">Saitou et&#x20;al<italic>.</italic> 1987</xref>) to construct the phylogenetic&#x20;tree.</p>
<sec id="s3-1">
<title>Robinson-Foulds Distance and the Bootstrap Method</title>
<p>We use the Robinson-Foulds (RF) distance (<xref ref-type="bibr" rid="B23">Robinson and Foulds 1981</xref>) to judge the quality of the method. A smaller RF value means a closer distance between the phylogenetic tree and the reference&#x20;tree.</p>
<p>(<xref ref-type="bibr" rid="B37">Yu et&#x20;al., 2010</xref>) proposed a modified version of the bootstrap method to evaluate the reliability of the constructed phylogenetic tree. We also use this method in the present work. Its workflow is as follows: Each row is the feature vector (<inline-formula id="inf28">
<mml:math id="m32">
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mn>4</mml:mn>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> of a species, and each column is the feature value of all genome sequences based on the same <italic>k</italic>-mer. Through random sampling of all columns, in which some columns may be selected many times, while some columns may not be selected at all, we randomly select one column. After 4<sup>
<italic>K</italic>
</sup> times of selection, a new <italic>N</italic>
<inline-formula id="inf29">
<mml:math id="m33">
<mml:mo>&#xd7;</mml:mo>
</mml:math>
</inline-formula>4<sup>K</sup> feature matrix is constructed. Using the new feature matrix, the Manhattan distance of any two rows is calculated to get a new distance matrix. Then we use the NJ method to construct a phylogenetic tree and repeat the above steps 100 times. Finally, a consensus tree is drawn by using consense. exe in the Phylip package. The frequency of a particular branch of a phylogenetic tree can be used as a measure of the stability of this branch.</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>Results</title>
<sec id="s4-1">
<title>Experiment 1</title>
<p>We use the genomes of 30 mammalian species in dataset 1 to construct a phylogenetic tree using ClustalX (<xref ref-type="bibr" rid="B14">Larkin et&#x20;al<italic>.</italic> 2007</xref>) as the reference tree. ClustalX is one of the widely used multiple alignment programs. The result is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1A</xref>. It is seen that rabbit, fat dormouse, squirrel, guinea pig, mouse, rat, platypus, opossum, and wallaroo belong to the rodents group; human, baboon, orangutan, gibbon, gorilla, pigmy chimpanzee, and common chimpanzee belong to the primates group; blue whale, fin whale, hippopotamus, cow, sheep, pig, donkey, horse, Indian-rhinoceros, white rhinoceros, cat, dog, gray seal, and harbor seal belong to the ferungulates group. When <italic>K</italic>&#x20;&#x3c; 5, it is not feasible to construct a phylogenetic tree using our method. When <italic>K</italic>&#x20;&#x3d; 5, 6, the 30 mammals cannot be divided into three groups in our tree. When <italic>K</italic>&#x20;&#x3d; 7, it can be divided into three groups, but the relationship between guinea pig and fat dormouse is not correct. When <italic>K</italic>&#x20;&#x3d; 8, 9, the branches of the tree become correct. We list the RF distances between the phylogenetic tree constructed by our method at <italic>K</italic>&#x20;&#x3d; 5, 6, 7, 8, 9 and the reference tree constructed by ClustalX in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. From <xref ref-type="table" rid="T3">Table&#x20;3</xref>, we can see that the RF distance reaches the minimum when <italic>K</italic>&#x20;&#x3d; 8. We show the phylogenetic tree of <italic>K</italic>&#x20;&#x3d; 8 constructed by our method in <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>. From <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>, we can see that the species in the three main categories are grouped correctly. Primates and ferungulates are closer, and this relationship is consistent with that in <xref ref-type="fig" rid="F1">Figure&#x20;1A</xref>. In terms of branches, monotremes (platypus), marsupials (wallaroo, opossum), murid rodents (mouse, rat), non-murid rodents (guinea pig, squirrel, fat dormouse, rabbit), perissodactyls (white rhinoceros, horse, Indian rhinoceros, donkey), carnivores (harbor seal, dog, gray seal, cat), artiodactyls (sheep, cow, hippopotamus, pig), primates (human, pigmy chimpanzee, common chimpanzee, gorilla, baboon, gibbon, orangutan), and cetaceans (blue whale, fin whale) are grouped into respective taxonomic classes accurately.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>
<bold>(A)</bold> The phylogenetic tree of 30 mammalian species reconstructed by ClustalX. <bold>(B)</bold> The phylogenetic tree of 30 mammalian species at <italic>K</italic>&#x20;&#x3d; 8 based on our method.</p>
</caption>
<graphic xlink:href="fgene-12-766496-g001.tif"/>
</fig>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The RF distance between the phylogenetic tree conducted by our method at K &#x3d; 5,6,7,8,9 and the reference tree conducted by ClustalX.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">
<italic>K</italic>
</th>
<th align="center">5</th>
<th align="center">6</th>
<th align="center">7</th>
<th align="center">8</th>
<th align="center">9</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">RF distance</td>
<td align="char" char=".">38</td>
<td align="char" char=".">28</td>
<td align="char" char=".">22</td>
<td align="char" char=".">8</td>
<td align="char" char=".">10</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="fig" rid="F2">Figure&#x20;2</xref> shows the RF distance between the reference tree constructed by ClustalX and the phylogenetic tree constructed by our method, Tang&#x2019;s method, PWKmer, DLtree, and CVtree on dataset 1. Using our method, when <italic>K</italic>&#x20;&#x3d; 8, the RF distance is 8. The shortest RF distance of DLtree (<italic>K</italic>&#x20;&#x3d; 9) is 10, the shortest distance of CVtree (<italic>K</italic>&#x20;&#x3d; 9) is 16, the shortest distance of Tang&#x2019;s method (<italic>K</italic>&#x20;&#x3d; 7) is 16, and the shortest distance of <italic>PWKmer</italic> (<italic>K</italic>&#x20;&#x3d; 9) is 10. Therefore, the results of our method are closer to those of ClustalX than those of the other methods, which indicates that our method is effective.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The Robinson&#x2013;Foulds distance between the tree reconstructed by ClustalX method and the phylogenetic trees reconstructed by our method (IEPWRMkmer K &#x3d; 8), the CVTree method, the DLTree method, Tang&#x2019;s method (K &#x3d; 7), and the PWKmer method (K &#x3d; 9) on dataset 1 (we used the optimal tree by CVTree and DLTree).</p>
</caption>
<graphic xlink:href="fgene-12-766496-g002.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F3">Figure&#x20;3</xref> shows the consensus tree of 30 mammalian species based on our method. Compared with <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>, 30 mammalian species are divided into the rodents group, the ferungulates group, and the primates group correctly. The support rate is 80% for the rodents group and 100% for both ferungulates and primates groups. Among the branches, marsupials (opossum, wallaroo), carnivores (dog, cat, harbor seal, gray seal), murid roots (rat, mouse), and cetaceans (fin whale, blue whale) are all supported by a 100% rate. In the artiodactyls group (cow, sheep, pig, hippopotamus), pig is separated out of the artiodactyls group, but the support rate is low at 43%. It indicates that the phylogenetic tree constructed by our method is quite robust.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The modified bootstrap consensus tree for <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref> based on 100 replicates.</p>
</caption>
<graphic xlink:href="fgene-12-766496-g003.tif"/>
</fig>
</sec>
<sec id="s4-2">
<title>Experiment 2</title>
<p>The human immunodeficiency viruses (HIV) represent a group of retroviruses, which are not presumed to have originated from human cellular DNA sequences, hence are distinct from endogenous retroviruses (<xref ref-type="bibr" rid="B33">Wu et&#x20;al., 2007</xref>). HIV-1 can be classified into three major phylogenetic groups, namely M (major), N (new), and O (others). Group M is responsible for the HIV pandemic, it is divided into nine subtypes, namely A, B, C, D, F, G, J, K, and H. Based on differential phylogenetic clustering, the subtypes A and F are further divided into sub-subtypes (A1, A2) and (F1, F2), respectively. Groups N and O are derived from other primates and then infect humans. CPZ is a non-human primate virus isolated from chimpanzees, which is closest to human-to-human transmission of&#x20;HIV.</p>
<p>We performed the phylogenetic analysis of 44&#x20;HIV-1 complete genome sequences in dataset 2 using ClustalX and our method. The phylogenetic trees reconstructed by ClustalX and our method (<italic>K</italic>&#x20;&#x3d; 7) are shown in <xref ref-type="fig" rid="F4">Figure&#x20;4A</xref> and <xref ref-type="fig" rid="F4">Figure&#x20;4B</xref>, respectively. From <xref ref-type="fig" rid="F4">Figure&#x20;4B</xref>, we can see that the species from all subtypes can be correctly classified into their groups (A, B, C, D, F, G, J, K, H, O, and M), and CPZ as the reference sequence is separated into the outermost. From the internal branches, both F and A contain two subtypes (F1 and F2) and (A1 and A2), respectively. Our method can separate the two subtypes, and in the branches, both F and A subtypes can be closely grouped together.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>
<bold>(A)</bold> The phylogenetic tree of 44&#x20;HIV-1 genomes reconstructed by ClustalX. <bold>(B)</bold> The phylogenetic tree of 44&#x20;HIV-1 genomes reconstructed by our method (<italic>K</italic>&#x20;&#x3d; 7).</p>
</caption>
<graphic xlink:href="fgene-12-766496-g004.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F5">Figure&#x20;5</xref> shows the RF distances between the reference tree constructed by ClustalX and the phylogenetic trees constructed by our method, Tang&#x2019;s method, PWKmer, DLtree, and CVtree. Using our method, when <italic>K &#x3d;</italic> 7, the RF distance is 10. The shortest RF distance of the DLtree (<italic>K &#x3d;</italic> 11) is 12, the shortest distance of the CVtree (<italic>K &#x3d;</italic> 9) is 16, the shortest distance of the PWKmer (<italic>K&#x20;&#x3d;</italic>9) is 10, and the shortest distance of Tang&#x2019;s method (<italic>K &#x3d;</italic> 9) is 10. Therefore, our method performs better than the DLtree and the CVtree on dataset 2 and has the same performance as Tang&#x2019;s method and PWKmer. The results indicate that our method is quite effective&#x20;again.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The RF distance between the reference tree constructed by Clustalx and the phylogenetic trees constructed by our method (IEPWRMkmer, <italic>K</italic>&#x20;&#x3d; 7), Tang&#x2019;s method (<italic>K</italic>&#x20;&#x3d; 8), the PWKmer method (<italic>K</italic>&#x20;&#x3d; 9), the DLtree method, and the CVtree method. (For the PWKmer method, the DLtree method, and the CVtree method, we chose their optimal classification tree).</p>
</caption>
<graphic xlink:href="fgene-12-766496-g005.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F6">Figure&#x20;6</xref> shows the consensus tree of 44&#x20;HIV-1 based on our method. Comparing with <xref ref-type="fig" rid="F4">Figure&#x20;4B</xref>, all HIV-1 sequences are divided into the M, N, O, and CPZ groups, whose support rate is 100%. From the branch point of view, in group M, the branch support rate of all subtypes is 100%. For subtypes A and F, the subtypes (A1, A2) and (F1 and F2) are clustered with 100% support. It again indicates that the phylogenetic tree constructed by our method is quite robust.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The modified bootstrap consensus tree for <xref ref-type="fig" rid="F4">Figure&#x20;4B</xref> based on 100 replicates.</p>
</caption>
<graphic xlink:href="fgene-12-766496-g006.tif"/>
</fig>
</sec>
<sec id="s4-3">
<title>Estimate of the Optimal Parameter <italic>K</italic>
</title>
<p>Different lengths of <italic>k</italic>-mers contain different phylogenetic information. Short <italic>k</italic>-mers may not contain sufficient DNA sequence information. Long <italic>k</italic>-mers contain sufficient phylogenetic information, but it needs large memory and takes a long time to calculate the distance based on information on long <italic>k</italic>-mers. Therefore, it is also very important to estimate an optimal value of <italic>K</italic> as heralded in (<xref ref-type="bibr" rid="B36">Yu et&#x20;al., 2010</xref>) for the DLTree method and (<xref ref-type="bibr" rid="B22">Qi et&#x20;al., 2004</xref>) for the CVTree method.</p>
<p>In this paper, we propose to use the Shannon entropy of the feature matrix to determine the optimal value of <italic>K</italic>. Using <xref ref-type="disp-formula" rid="e3">Eq. 3</xref>, we can obtain an <italic>N</italic> <inline-formula id="inf30">
<mml:math id="m34">
<mml:mo>&#xd7;</mml:mo>
</mml:math>
</inline-formula>4<sup>
<italic>K</italic>
</sup> feature matrix for a dataset with <italic>N</italic> genomes. Then, we propose to define a scoring strategy as<disp-formula id="e5">
<mml:math id="m35">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mn>4</mml:mn>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>log</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>log</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
<p>The optimal <italic>K</italic> is the value at which <inline-formula id="inf31">
<mml:math id="m36">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> reaches its maximum.</p>
<p>We use <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> to calculate <inline-formula id="inf32">
<mml:math id="m37">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> on datasets 1 and 2 for different <italic>K</italic>. The relationship between <inline-formula id="inf33">
<mml:math id="m38">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <italic>K</italic> is shown in <xref ref-type="fig" rid="F7">Figure&#x20;7</xref> for these two datasets. It is seen that <inline-formula id="inf34">
<mml:math id="m39">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> reaches the largest value when <italic>K</italic>&#x20;&#x3d; 8 on the two datasets. Considering that the larger <italic>K</italic> is, the more memory resources are consumed, we only consider the values near <italic>K</italic>&#x20;&#x3d; 8 (e.g., <italic>K</italic>&#x20;&#x3d; 7, 8, 9). For the 30 mammalian species dataset, we have seen that the phylogenetic tree for <italic>K</italic>&#x20;&#x3d; 8 constructed by our method is closest to the reference tree. The same happened for the HIV-1 dataset with <italic>K</italic>&#x20;&#x3d; 7. The outcomes indicate that <inline-formula id="inf35">
<mml:math id="m40">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> can provide an effective means to estimate the optimal value of&#x20;<italic>K</italic>.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The trend chart of <italic>K</italic> value vs scoring measure <inline-formula id="inf36">
<mml:math id="m41">
<mml:mrow>
<mml:mtext>score</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. The red circles represent the scores of the dataset of 30 mammalian species for different <italic>K</italic> values, and the blue dots represent the scores of the HIV dataset for different <italic>K</italic> values.</p>
</caption>
<graphic xlink:href="fgene-12-766496-g007.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>Conclusion</title>
<p>In this paper, a new alignment-free method is proposed for phylogenetic analysis and sequence comparison based on whole genome sequences. Our method combines the position-weighted measure of <italic>k</italic>-mers and the information entropy of frequency of <italic>k</italic>-mers. We used the Manhattan metric to measure the distance between a pair of sequences and the NJ method to construct the phylogenetic tree. In order to test the effectiveness and reliability of our method, we applied it on two datasets of 30 mammalian species and 44&#x20;HIV-1 genomes. The results demonstrated that the present method is efficient and reliable. A suitable <italic>K</italic> value is important to capture rich phylogenetic information of DNA sequences. In order to choose an optimal <italic>K</italic> value, we proposed a scoring measure based on the information entropy. The obtained results on two real datasets support that the method can capture the <italic>k</italic>-mer distribution information and is effective for whole genome sequence comparison and phylogenetic analysis.</p>
<p>Remark: The method of this paper is derived from the two studies <xref ref-type="bibr" rid="B17">Ma et&#x20;al. (2020)</xref> and Murray et&#x20;al<italic>.</italic> (2017). There are differences between this work and previous works: Tang et&#x20;al. presented the average relative distance for normalized <italic>k</italic>-mers. PWKmer uses the counts and position distributions of <italic>k</italic>-mers to capture more evolutionary information. KWIP (Murray et&#x20;al<italic>.</italic> 2017) uses information entropy to weight the inner product (Si<inline-formula id="inf37">
<mml:math id="m42">
<mml:mo>&#x2217;</mml:mo>
</mml:math>
</inline-formula>Sj), while we use information entropy to weight the relative positions of <italic>k</italic>-mers. KWIP uses a kernel function to calculate the distance, while we use the Manhattan metric to calculate the pairwise distance between species. Here, we claimed that the results obtained by the IEPWRMkmer method are close to those by ClustalX and the IEPWRMkmer is superior to the other distance metrics. We used the phylogenetic tree constructed by ClustalX as the reference tree or standard tree, hence we cannot claim that our method is superior to the ClustalX method.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The genome datasets analyzed for this study can be found in the GenBank <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov/</ext-link>
</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>Y-QW contributed to the conception and design of the study, developed the method, and wrote the manuscript. Z-GY gave the ideas and supervised the project. All authors discussed the results and reviewed the manuscript. All authors have read and agreed to the published version of the manuscript.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by funds from the National Natural Science Foundation of China (grant numbers: 11871061 and 12026213); The National Key Research and Development Program of China (grant number: 2020YFC0832405); Innovation Foundation of Qian Xuesen Laboratory of Space Technology.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Gish</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Myers</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Basic Local Alignment Search Tool</article-title>. <source>J.&#x20;Mol. Biol.</source> <volume>215</volume> (<issue>3</issue>), <fpage>403</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Blaisdell</surname>
<given-names>B. E.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>A Measure of the Similarity of Sets of Sequences Not Requiring Sequence Alignment</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>83</volume> (<issue>14</issue>), <fpage>5155</fpage>&#x2013;<lpage>5159</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.83.14.5155</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A Novel Alignment-free Method for Whole Genome Analysis: Application to HIV-1 Subtyping and HEV Genotyping</article-title>. <source>Inf. Sci.</source> <volume>279</volume>, <fpage>776</fpage>&#x2013;<lpage>784</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2014.04.029</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Comin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Verzotto</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Alignment-free Phylogeny of Whole Genomes Using Underlying Subwords</article-title>. <source>Algorithms Mol. Biol.</source> <volume>7</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1186/1748-7188-7-34</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>A Simple <italic>K</italic>-word Interval Method for Phylogenetic Analysis of DNA Sequences</article-title>. <source>J.&#x20;Theor. Biol.</source> <volume>317</volume>, <fpage>192</fpage>&#x2013;<lpage>199</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2012.10.010</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Felsenstein</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Felenstein</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2004</year>). <source>Inferring Phylogenies</source>. (<publisher-loc>Sunderland, MA</publisher-loc>: <publisher-name>Sinauer Associates</publisher-name>). <pub-id pub-id-type="doi">10.1086/383584</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fox</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Magrum</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Balch</surname>
<given-names>W. E.</given-names>
</name>
<name>
<surname>Wolfe</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Woese</surname>
<given-names>C. R.</given-names>
</name>
</person-group> (<year>1977</year>). <article-title>Classification of Methanogenic Bacteria by 16S Ribosomal RNA Characterization</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>74</volume> (<issue>10</issue>), <fpage>4537</fpage>&#x2013;<lpage>4541</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.74.10.4537</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haubold</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Pierstorff</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>M&#xf6;ller</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wiehe</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Genome Comparison without Alignment Using Shortest Unique Substrings</article-title>. <source>BMC Bioinformatics</source> <volume>6</volume> (<issue>1</issue>), <fpage>123</fpage>&#x2013;<lpage>211</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-6-123</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hoang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yin</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yau</surname>
<given-names>S. S.-T.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Numerical Encoding of DNA Sequences by Chaos Game Representation with Application in Similarity Comparison</article-title>. <source>Genomics</source> <volume>108</volume>, <fpage>134</fpage>&#x2013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2016.08.002</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>H&#xf6;hl</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rigoutsos</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Ragan</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Pattern-based Phylogenetic Distance Estimation and Tree Reconstruction</article-title>. <source>Evol. Bioinformatics</source> <volume>2</volume>, <fpage>359</fpage>&#x2013;<lpage>375</lpage>. <pub-id pub-id-type="doi">10.2174/157489306775330570</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Phylogenetic Analysis of DNA Sequences with a Novel Characteristic Vector</article-title>. <source>J.&#x20;Math. Chem.</source> <volume>49</volume> (<issue>8</issue>), <fpage>1479</fpage>&#x2013;<lpage>1492</lpage>. <pub-id pub-id-type="doi">10.1007/s10910-011-9811-x</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Stecher</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Tamura</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>MEGA7: Molecular Evolutionary Genetics Analysis Version 7.0 for Bigger Datasets</article-title>. <source>Mol. Biol. Evol.</source> <volume>33</volume> (<issue>7</issue>), <fpage>1870</fpage>&#x2013;<lpage>1874</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msw054</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Larkin</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Blackshields</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Brown</surname>
<given-names>N. P.</given-names>
</name>
<name>
<surname>Chenna</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>McGettigan</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>McWilliam</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>Clustal W and Clustal X Version 2.0</article-title>. <source>Bioinformatics</source> <volume>23</volume> (<issue>21</issue>), <fpage>2947</fpage>&#x2013;<lpage>2948</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm404</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Leimeister</surname>
<given-names>C.-A.</given-names>
</name>
<name>
<surname>Boden</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Horwege</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lindner</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Morgenstern</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Fast Alignment-free Sequence Comparison Using Spaced-word Frequencies</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>1991</fpage>&#x2013;<lpage>1999</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu177</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Badger</surname>
<given-names>J.&#x20;H.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Kwong</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kearney</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>An Information-Based Sequence Distance and its Application to Whole Mitochondrial Genome Phylogeny</article-title>. <source>Bioinformatics</source> <volume>17</volume> (<issue>2</issue>), <fpage>149</fpage>&#x2013;<lpage>154</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/17.2.149</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Anh</surname>
<given-names>V. V.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Phylogenetic Analysis of HIV-1 Genomes Based on the Position-Weighted <italic>K</italic>-Mers Method</article-title>. <source>Entropy</source> <volume>22</volume> (<issue>2</issue>), <fpage>255</fpage>. <pub-id pub-id-type="doi">10.3390/e22020255</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mendizabal-Ruiz</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Rom&#xe1;n-God&#xed;nez</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Torres-Ramos</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Salido-Ruiz</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>V&#xe9;lez-P&#xe9;rez</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Morales</surname>
<given-names>J.&#x20;A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Genomic Signal Processing for DNA Sequence Clustering</article-title>. <source>PeerJ</source> <volume>6</volume> (<issue>3</issue>), <fpage>e4264</fpage>. <pub-id pub-id-type="doi">10.7717/peerj.4264</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morrison</surname>
<given-names>D. A.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Multiple Sequence Alignment for Phylogenetic Purposes</article-title>. <source>Aust. Syst. Bot.</source> <volume>19</volume> (<issue>6</issue>), <fpage>479</fpage>&#x2013;<lpage>539</lpage>. <pub-id pub-id-type="doi">10.1071/sb06020</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Murray</surname>
<given-names>K. D.</given-names>
</name>
<name>
<surname>Webers</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ong</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Borevitz</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Warthmann</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>KWIP: The <italic>K</italic>-Mer Weighted Inner Product, a De Novo Estimator of Genetic Similarity</article-title>. <source>Plos Comput. Biol.</source> <volume>13</volume> (<issue>9</issue>), <fpage>e1005727</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1005727</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Otu</surname>
<given-names>H. H.</given-names>
</name>
<name>
<surname>Sayood</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>A New Sequence Distance Measure for Phylogenetic Tree Construction</article-title>. <source>Bioinformatics</source> <volume>19</volume> (<issue>16</issue>), <fpage>2122</fpage>&#x2013;<lpage>2130</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btg295</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hao</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>CVTree: a Phylogenetic Tree Reconstruction Tool Based on Whole Genomes</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume> (<issue>Suppl. l_2</issue>), <fpage>W45</fpage>&#x2013;<lpage>W47</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh362</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Robinson</surname>
<given-names>D. F.</given-names>
</name>
<name>
<surname>Foulds</surname>
<given-names>L. R.</given-names>
</name>
</person-group> (<year>1981</year>). <article-title>Comparison of Phylogenetic Trees</article-title>. <source>Math. Biosciences</source> <volume>53</volume> (<issue>1-2</issue>), <fpage>131</fpage>&#x2013;<lpage>147</lpage>. <pub-id pub-id-type="doi">10.1016/0025-5564(81)90043-2</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ronquist</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Teslenko</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Van Der Mark</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ayres</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Darling</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>H&#xf6;hna</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>MrBayes 3.2: Efficient Bayesian Phylogenetic Inference and Model Choice across a Large Model Space</article-title>. <source>Syst. Biol.</source> <volume>61</volume> (<issue>3</issue>), <fpage>539</fpage>&#x2013;<lpage>542</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/sys029</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saitou</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Nei</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1987</year>). <article-title>The Neighbor-Joining Method: a New Method for Reconstructing Phylogenetic Trees</article-title>. <source>Mol. Biol. Evol.</source> <volume>4</volume> (<issue>4</issue>), <fpage>406</fpage>&#x2013;<lpage>425</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a040454</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sims</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Jun</surname>
<given-names>S.-R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S.-H.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Alignment-free Genome Comparison with Feature Frequency Profiles (FFP) and Optimal Resolutions</article-title>. <source>Pnas</source> <volume>106</volume> (<issue>8</issue>), <fpage>2677</fpage>&#x2013;<lpage>2682</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0813249106</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hua</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A Novel <italic>K</italic>-word Relative Measure for Sequence Comparison</article-title>. <source>Comput. Biol. Chem.</source> <volume>53</volume>, <fpage>331</fpage>&#x2013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiolchem.2014.10.007</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thankachan</surname>
<given-names>S. V.</given-names>
</name>
<name>
<surname>Chockalingam</surname>
<given-names>S. P.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Apostolico</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Aluru</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>ALFRED: a Practical Method for Alignment-free Distance Computation</article-title>. <source>J.&#x20;Comput. Biol.</source> <volume>23</volume> (<issue>6</issue>), <fpage>452</fpage>&#x2013;<lpage>460</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2015.0217</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thompson</surname>
<given-names>J.&#x20;D.</given-names>
</name>
<name>
<surname>Higgins</surname>
<given-names>D. G.</given-names>
</name>
<name>
<surname>Gibson</surname>
<given-names>T. J.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>CLUSTAL W: Improving the Sensitivity of Progressive Multiple Sequence Alignment through Sequence Weighting, Position-specific gap Penalties and Weight Matrix Choice</article-title>. <source>Nucl. Acids Res.</source> <volume>22</volume> (<issue>22</issue>), <fpage>4673</fpage>&#x2013;<lpage>4680</lpage>. <pub-id pub-id-type="doi">10.1093/nar/22.22.4673</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ulitsky</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Burstein</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tuller</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chor</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>The Average Common Substring Approach to Phylogenomic Reconstruction</article-title>. <source>J.&#x20;Comput. Biol.</source> <volume>13</volume> (<issue>2</issue>), <fpage>336</fpage>&#x2013;<lpage>350</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2006.13.336</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lei</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Effect of K-Tuple Length on Sample-Comparison with High-Throughput Sequencing Data</article-title>. <source>Biochem. Biophysical Res. Commun.</source> <volume>469</volume> (<issue>4</issue>), <fpage>1021</fpage>&#x2013;<lpage>1027</lpage>. <pub-id pub-id-type="doi">10.1016/j.bbrc.2015.11.094</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z.-G.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>DLTree: Efficient and Accurate Phylogeny Reconstruction Using the Dynamical Language Method</article-title>. <source>Bioinformatics</source> <volume>33</volume> (<issue>14</issue>), <fpage>2214</fpage>&#x2013;<lpage>2215</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx158</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>X.-F.</given-names>
</name>
<name>
<surname>Hoang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Goebel</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Nucleotide Composition String Selection in HIV-1 Subtyping Using Whole Genomes</article-title>. <source>Bioinformatics</source> <volume>23</volume> (<issue>14</issue>), <fpage>1744</fpage>&#x2013;<lpage>1752</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm248</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yin</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Encoding and Decoding DNA Sequences by Integer Chaos Game Representation</article-title>. <source>J.&#x20;Comput. Biol.</source> <volume>26</volume> (<issue>2</issue>), <fpage>143</fpage>&#x2013;<lpage>151</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2018.0173</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yau</surname>
<given-names>S. S.-T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>DNA Sequence Comparison by a Novel Probabilistic Method</article-title>. <source>Inf. Sci.</source> <volume>181</volume> (<issue>8</issue>), <fpage>1484</fpage>&#x2013;<lpage>1492</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2010.12.010</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>Z.-G.</given-names>
</name>
<name>
<surname>Chu</surname>
<given-names>K. H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Anh</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>L.-Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R. W.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Whole-proteome Phylogeny of Large dsDNA Viruses and Parvoviruses through a Composition Vector Method Related to Dynamical Language Model</article-title>. <source>BMC Evol. Biol.</source> <volume>10</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-10-192</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>Z.-G.</given-names>
</name>
<name>
<surname>Zhan</surname>
<given-names>X.-W.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>G.-S.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R. W.</given-names>
</name>
<name>
<surname>Anh</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Chu</surname>
<given-names>K. H.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Proper Distance Metrics for Phylogenetic Analysis Using Complete Genomes without Sequence Alignment</article-title>. <source>Ijms</source> <volume>11</volume> (<issue>3</issue>), <fpage>1141</fpage>&#x2013;<lpage>1154</lpage>. <pub-id pub-id-type="doi">10.3390/ijms11031141</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zielezinski</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Vinga</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Almeida</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Karlowski</surname>
<given-names>W. M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Alignment-free Sequence Comparison: Benefits, Applications, and Tools</article-title>. <source>Genome Biol.</source> <volume>18</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1186/s13059-017-1319-7</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>