<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1091575</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2023.1091575</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Next-generation development and application of codon model in evolution</article-title>
<alt-title alt-title-type="left-running-head">Gupta and Vadde</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fgene.2023.1091575">10.3389/fgene.2023.1091575</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Gupta</surname>
<given-names>Manoj Kumar</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1328003/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Vadde</surname>
<given-names>Ramakrishna</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/892047/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Department of Biotechnology &#x26; Bioinformatics</institution>, <institution>Yogi Vemana University</institution>, <addr-line>Kadapa</addr-line>, <addr-line>Andhra Pradesh</addr-line>, <country>India</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/285087/overview">Yassine Souilmi</ext-link>, University of Adelaide, Australia</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2130633/overview">Laurent Gu&#xe9;guen</ext-link>, UMR5558 Biom&#xe9;trie et Biologie Evolutive (LBBE), France</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2158987/overview">Purnachandra Ganji</ext-link>, University of Alabama at Birmingham, United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/677807/overview">Madhu Sudhana Saddala</ext-link>, The Johns Hopkins Hospital, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Ramakrishna Vadde, <email>vrkrishna70@gmail.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Evolutionary and Population Genetics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1091575</elocation-id>
<history>
<date date-type="received">
<day>07</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>17</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Gupta and Vadde.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Gupta and Vadde</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>To date, numerous nucleotide, amino acid, and codon substitution models have been developed to estimate the evolutionary history of any sequence/organism in a more comprehensive way. Out of these three, the codon substitution model is the most powerful. These models have been utilized extensively to detect selective pressure on a protein, codon usage bias, ancestral reconstruction and phylogenetic reconstruction. However, due to more computational demanding, in comparison to nucleotide and amino acid substitution models, only a few studies have employed the codon substitution model to understand the heterogeneity of the evolutionary process in a genome-scale analysis. Hence, there is always a question of how to develop more robust but less computationally demanding codon substitution models to get more accurate results. In this review article, the authors attempted to understand the basis of the development of different types of codon-substitution models and how this information can be utilized to develop more robust but less computationally demanding codon substitution models. The codon substitution model enables to detect selection regime under which any gene or gene region is evolving, codon usage bias in any organism or tissue-specific region and phylogenetic relationship between different lineages more accurately than nucleotide and amino acid substitution models. Thus, in the near future, these codon models can be utilized in the field of conservation, breeding and medicine.</p>
</abstract>
<kwd-group>
<kwd>codon substitution models</kwd>
<kwd>mechanistic model</kwd>
<kwd>empirical model</kwd>
<kwd>semi-emperical models</kwd>
<kwd>evolution</kwd>
<kwd>phylogenetic reconstruction</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The continuous growth of DNA and protein data has provided an opportunity to infer their function and their evolutionary history of any sequence/organisms in a more comprehensive way (<xref ref-type="bibr" rid="B10">Anisimova and Liberles, 2007</xref>; <xref ref-type="bibr" rid="B114">Miyazawa, 2011a</xref>; <xref ref-type="bibr" rid="B46">Dufresne and Jeffery, 2011</xref>; <xref ref-type="bibr" rid="B68">Gupta and Vadde, 2019a</xref>; <xref ref-type="bibr" rid="B72">Gupta et al., 2019</xref>; <xref ref-type="bibr" rid="B62">Gouda et al., 2020</xref>; <xref ref-type="bibr" rid="B73">Gupta et al., 2021a</xref>; <xref ref-type="bibr" rid="B74">Gupta et al., 2021b</xref>; <xref ref-type="bibr" rid="B75">Gupta et al., 2021c</xref>; <xref ref-type="bibr" rid="B76">Gupta et al., 2021d</xref>; <xref ref-type="bibr" rid="B77">Gupta et al., 2021e;</xref> <xref ref-type="bibr" rid="B29">Chu et al., 2021</xref>). Population genetics and phylogenetics are two of the most important subfields for inferring the evolutionary history of any sequences/organisms (<xref ref-type="bibr" rid="B81">Haubold, 2014</xref>). While phylogeny approaches infer the evolution of species and higher taxonomic orders, population genetics approaches are generally used for understanding the evolution of the groups below the species level (<xref ref-type="bibr" rid="B81">Haubold, 2014</xref>). It is pertinent to note that there is only one diagram, a phylogeny, that appears in Darwin&#x2019;s seminal work, &#x201c;The Origin of Species.&#x201d; This, in turn, indicates that phylogenies are the core metaphor of evolutionary biology, and efforts to create them is as old as the evolutionary science field itself. Phylogenies relation are inferred by comparing homologous characteristics that differ (<xref ref-type="bibr" rid="B81">Haubold, 2014</xref>). A phylogenetic tree, often known as an evolutionary tree, is a diagrammatic depiction of the evolutionary relationship between different species (<xref ref-type="bibr" rid="B75">Gupta et al., 2021c</xref>). All phylogenetic tree analysis is based on certain implicit/explicit hypothetical models that make the complex biological process into simpler form (<xref ref-type="bibr" rid="B57">Gatto et al., 2007</xref>; <xref ref-type="bibr" rid="B183">Zaheri et al., 2014</xref>). However, the validity of certain models could be plausibly challenged while analyzing real data. For instance, the JC69 model hypothesizes that the rate of nucleotide substitution is the same for all pairs of the four nucleotides, namely, guanine (G), cytosine (C), adenine (A), and thymine (T) (<xref ref-type="bibr" rid="B94">Jukes and Cantor, 1969</xref>). However, in reality, there are numerous mutations, and some mutations are less tolerated in comparison to others (<xref ref-type="bibr" rid="B107">Li&#xf2; and Goldman, 1998</xref>; <xref ref-type="bibr" rid="B28">Choudhuri, 2014</xref>). Nevertheless, all models share a common assumption, i.e., the Markov property (<xref ref-type="bibr" rid="B107">Li&#xf2; and Goldman, 1998</xref>). In probability theory, any stochastic process has the Markov property if the probability distribution of future states of the process is dependent only on the present state (<xref ref-type="bibr" rid="B66">Gudivada et al., 2015</xref>). Markov property has been widely utilized in population genetics research to understand the change in gene frequencies in small populations affected <italic>via</italic> genetic drift (<xref ref-type="bibr" rid="B163">Watterson, 1996</xref>).</p>
<p>Based on sequence type, all these substitution models can be broadly classified as nucleotides, amino acid, and codon substitution models (<xref ref-type="bibr" rid="B15">Arenas, 2015b</xref>). The parameter space dimension for these models varies from 4 &#xd7; 4 nucleotide substitution model to 20 &#xd7; 20 amino acid substitution models and, finally, to 61 &#xd7; 61 codon substitution models (where stop codons are generally omitted). Because of only four states and small physiochemical differences between base properties, nucleotide substitution models are easily modeled <italic>via</italic> Markov models (<xref ref-type="bibr" rid="B181">Yang, 2006</xref>). However, as natural selection functions mostly at the protein level, estimating evolutionary history based on nucleotide substitution models can sometimes be misleading (<xref ref-type="bibr" rid="B146">Shapiro et al., 2006</xref>; <xref ref-type="bibr" rid="B144">Seo and Kishino, 2008</xref>). Both amino acid or codon substitution models consider protein-coding sequences and thus, an evolutionary distance estimated <italic>via</italic> them is more accurate than the evolutionary distance estimated through nucleotide substitution models (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). Nevertheless, due to the complex physiochemical relationship between amino acids, it is often difficult to predict the substitution rate between amino acids in a small set of the original dataset. Hence, the substitution rate in amino acid substitution models is generally estimated from pre-defined empirical data sets (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>), which in turn may predict evolutionary history less accurately.</p>
<p>The codon substitution models are especially interesting for protein-coding genes because they consider both mutational propensities at the nucleotide level and selective pressure on amino acid substitutes as well as genetic code for estimating evolutionary distance (<xref ref-type="bibr" rid="B150">Sullivan and Joyce, 2005</xref>). Additionally, amino acid substitution models can estimate only purifying selection acting on each site of sequence, whereas codon substitution models can estimate both purifying as well as positive Darwinian selection (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>). Even for highly divergent species, phylogenetic trees constructed <italic>via</italic> codon models were reported to be more accurate than the phylogenetic tree constructed through the amino acid substitution model (<xref ref-type="bibr" rid="B183">Zaheri et al., 2014</xref>). Though these models were neglected initially, tracing phylogenetic relationships between populations and traits/diseases <italic>via</italic> codon substitution models is increasing nowadays in evolutionary medicine research (<xref ref-type="bibr" rid="B65">Grunspan et al., 2017</xref>). Thus, the codon substitution model is more powerful than nucleotide (<xref ref-type="bibr" rid="B96">Kimura, 1980</xref>; <xref ref-type="bibr" rid="B80">Hasegawa et al., 1985</xref>; <xref ref-type="bibr" rid="B154">Tamura and Nei, 1993</xref>) and amino acid (<xref ref-type="bibr" rid="B35">Dayhoff et al., 1978</xref>; <xref ref-type="bibr" rid="B91">Jones et al., 1992</xref>; <xref ref-type="bibr" rid="B2">Adachi and Hasegawa, 1996</xref>) substitution models. However, as codon substitution models are computationally more demanding, their usage is minimal. Hence, it needs to develop more robust but less computationally demanding codon substitution models for reconstructing evolutionary history from sequence data. In this review article, authors made an attempt to understand the basis of the development of different types of codon-substitution models and how this information can be utilized to develop more robust and less computationally demanding codon substitution models for more accurate phylogeny as well as understand the evolutionary history of any sequences or organisms. In the near future, these models can be applied in the field of conservation, breeding and medicine.</p>
</sec>
<sec id="s2">
<title>Basis of development of substitution models</title>
<p>Evolution is generally considered a stochastic process by which the DNA segment can be either inserted or deleted, or duplicated, or recombination may take place (<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, 2012</xref>). The most frequent events during evolution are point mutation, which may have either no effect or a small effect or change protein function completely. If this point mutation becomes fixed, either due to genetic drift/positive selection, it is called substitution (<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, 2012</xref>). The probability by which a new base gets fixed in a population is dependent on the accompanying modification in the species&#x2019; fitness (<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, 2012</xref>). To date, numerous hypothetical substitution models have been proposed to understand the mechanism associated with substitutions in either nucleotide or amino acid sequences (<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, 2012</xref>). Though some of these models are more complex than others, all substitution models share a common assumption, i.e., the Markov property (<xref ref-type="bibr" rid="B107">Li&#xf2; and Goldman, 1998</xref>). Utilizing the Markov property, these substitution models (Markov model) estimate the probabilities for possible temporal or sequential DNA or protein sequences in any individual or species. They also enable us to detect a preference of any sequences towards the GC or AT content (<xref ref-type="bibr" rid="B94">Jukes and Cantor, 1969</xref>).</p>
<p>Each Markov model has some parameters, which, in evolutionary biology, either represent the substitution rate or from which the substitution rate can be derived (<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, 2012</xref>). These parameters and hypotheses related to these parameters are often estimated <italic>via</italic> Maximum Likelihood (ML) approaches (<xref ref-type="bibr" rid="B181">Yang, 2006</xref>). ML detects optimum parameters associated with the occurrence of data (for instance, a group of nucleotide sequences) under a given phylogenetic tree and specific evolutionary model (<xref ref-type="bibr" rid="B53">Fisher, 1925</xref>; <xref ref-type="bibr" rid="B50">Edwards, 1972</xref>). DART (DNA, Amino acid, and RNA Tests) (<xref ref-type="bibr" rid="B85">Holmes and Rubin, 2002</xref>) and PAML (&#x201c;Phylogenetic Analysis by Maximum Likelihood&#x201d;) (<xref ref-type="bibr" rid="B182">Yang, 2007</xref>) are two of the most widely used software for estimating phylogenetic ML. Apart from the phylogenetic tree, the ML value can also be utilized for answering several other biological questions. For instance, the identification of the most extreme transition/transversion mutational biases in a set of aligned sequences and sites that are evolving under the greatest selective constraints (<xref ref-type="bibr" rid="B53">Fisher, 1925</xref>; <xref ref-type="bibr" rid="B50">Edwards, 1972</xref>; <xref ref-type="bibr" rid="B51">Felsenstein and Felenstein, 2004</xref>).</p>
<p>These days, codon substitution models are also developed in a Bayesian framework. Apart from the likelihood function, Bayesian inference also describes a prior likelihood distribution on the model parameter (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Thus, the main objective of the Bayesian inference is to compute the posterior distribution (which is proportional to the likelihood multiplied by the prior), acquire random samples from this posterior density by means of Monte Carlo (MC) methods, and estimate mean and other numerous quantities on the basis of the sample obtained (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Theoretically, any ML model can be converted into a Bayesian model just by adding a prior distribution. The MC methods employed in Bayesian inference are often very powerful in exploring entirely novel codon substitution models, which is not possible by using classical numerical optimization approaches. For instance, the likelihood function, which is to be numerically assessed point-wise, as well as optimized with regard to its parameters, is itself a complicated integral over random variables (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). This integral is present analytically available only for the simplest models. On the contrary, when complicated models are considered, analyticity gets depleted, and the classical numerical optimization approaches employed in recent ML fail. In contrast, MC sampling approaches permit numerous algorithmic tricks, for instance, data augmentation as well as parameter expansion, which in turn improvise the requirement for obvious analytical integration over incomplete observations or over auxiliary variables. Nevertheless, MC methods are very demanding in the context of both code development as well as computational cost (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Thus, the number of Bayesian MCMC methods developed to date is very few and still, we have to wait for the joint availability of huge amounts of inter-specific sequence data as well as more powerful computational facilities for understanding their potential in a more comprehensive way (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Thus, the Markov property or Bayesian framework is the basis for the development of almost all substitution models developed to-date.</p>
</sec>
<sec id="s3">
<title>Substitution models</title>
<p>Based on sequence type, substitution models can be broadly classified as (a) nucleotide, (b) amino acid, and (c) codon substitution models.</p>
<sec id="s3-1">
<title>Nucleotide substitution model</title>
<p>JC69 model is the simplest model of nucleotide substitution (<xref ref-type="bibr" rid="B107">Li&#xf2; and Goldman, 1998</xref>). It is based on two simple assumptions. The first assumption is that each residue of DNA is equally likely to change to any of the other three nucleotide bases. The second assumption is that all four bases have the same frequency. Hence, the rate of transition is equal to the rate of transversions (<xref ref-type="bibr" rid="B127">Pevsner, 2015</xref>). Because of its simple assumptions, the JC69 model is unlikely to be applicable in most of the data sets and works reasonably only in closely related sequences. Though it can be utilized in distantly related sequences, the correction made can sometimes be too significant to be reliable. As this model involves a single parameter, <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, for both the rate of substitution for each nucleotide (3 <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> per unit time) and the rate of substitution in each of the three possible directions of change (<inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:mfenced open="" close=")" separators="|">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
</inline-formula>, it is called a one-parameter model. Kimura 2 Parameter (K80) is an extension of the JC69 model (<xref ref-type="bibr" rid="B94">Jukes and Cantor, 1969</xref>). As in real data, transversions generally occur at lower rates than transitions; (<xref ref-type="bibr" rid="B96">Kimura, 1980</xref>) proposed a model, which assumes that the transitions rate is different from the transversions rate. However, like the JC69 model, Kimura also assumes that all four bases have the same frequency.</p>
<p>Later several more robust nucleotide substitution models, like, F81 (<xref ref-type="bibr" rid="B52">Felsenstein, 1981</xref>), HKY85 (<xref ref-type="bibr" rid="B80">Hasegawa et al., 1985</xref>) and TN93 (<xref ref-type="bibr" rid="B154">Tamura and Nei, 1993</xref>) models, were developed. The F81 model was developed by American scientist (<xref ref-type="bibr" rid="B52">Felsenstein, 1981</xref>). Unlike Jukes &#x26; Cantor or Kimura-2 parameters, the F81 model assumes that the base frequency of all bases is different. However, like Jukes &#x26; Cantor, the F81 model assumes that the base substitution occurs with equal probability (<xref ref-type="bibr" rid="B52">Felsenstein, 1981</xref>). Hasegawa-Kishino-Yano 85 (HKY85) model assumes unequal base frequencies as well as different substitution rates between transversions and transitions (<xref ref-type="bibr" rid="B80">Hasegawa et al., 1985</xref>). <xref ref-type="bibr" rid="B154">Tamura and Nei&#x2019;s 1993</xref> (TN93) model assumes unequal base frequencies, but all transversions are assumed to take place at an equal rate, but the transition rate between purine differs from that of pyrimidine (<xref ref-type="bibr" rid="B154">Tamura and Nei, 1993</xref>).</p>
<p>For the first time in 1986, Simon Tavar&#xe9; described a general independent, finite-sites, neutral, and time-reversible model called the general time-reversible <bold>(</bold>GTR<bold>)</bold> model or the general reversible <bold>(</bold>REV<bold>)</bold> model, which assumes different substitution rates for each pair of nucleotide and unequal base frequencies (<xref ref-type="bibr" rid="B156">Tavar&#xe9;, 1986</xref>). Additionally, the rate of variation across sites (&#x2b;G) (<xref ref-type="bibr" rid="B177">Yang, 1994a</xref>) and/or a proportion of invariable sites (&#x2b;I) (<xref ref-type="bibr" rid="B148">Shoemaker and Fitch, 1989</xref>) can also be included in any model. Recently several other DNA substitution models comprising of non-stationary (nucleotide composition can change over time) and non-reversible (asymmetric) matrices (<xref ref-type="bibr" rid="B22">Boussau and Gouy, 2006</xref>; <xref ref-type="bibr" rid="B90">Jayaswal et al., 2011</xref>) or even involving neighbor interactions (<xref ref-type="bibr" rid="B109">Lunter and Hein, 2004</xref>) were developed for inferring phylogenetic trees more accurately.</p>
</sec>
<sec id="s3-2">
<title>Amino acid substitution model</title>
<p>The two commonly used amino acid substitution matrices are the PAM matrices (<xref ref-type="bibr" rid="B35">Dayhoff et al., 1978</xref>) and the Blocks amino acid substitution matrices (BLOSUM) matrices. Margaret Dayhoff and the team aligned closely related protein sequences of seventy-one groups (<xref ref-type="bibr" rid="B35">Dayhoff et al., 1978</xref>). As all the sequences were closely related homologs, mutation detected in them were less likely to change the function of the protein, and hence the matrix designed was named PAM, which is an abbreviated form of Percent Accepted Mutations, where &#x201c;accepted&#x201d; designate the mutation favored <italic>via</italic> natural selection in the sequence (<xref ref-type="bibr" rid="B172">Xiong, 2006</xref>). The PAM matrices were generated on the basis of the evolutionary divergence amongst sequences of the same group. For instance, one PAM unit is described as 1% of amino acids have been modified (<xref ref-type="bibr" rid="B172">Xiong, 2006</xref>) and PAM60 is generated when the PAM1 matrix is multiplied by itself sixty times. Thus, PAM with a lower serial number is suitable for aligning closely related sequences, and PAM with a higher serial number is suitable for divergent sequences (<xref ref-type="bibr" rid="B172">Xiong, 2006</xref>). Later, Jones and the team utilized PAM matrices and developed a more advanced replacement matrix, namely the JTT model based on a large sequences dataset. After constructing a phylogenetic tree of each protein family, this method identified sequence pairs that are &#x3e;85% identical and nearest-neighbors. Further, it also calculated the evolutionary distance among them. This pair of sequences were subsequently removed for avoiding recounting modifications on any given branch of a phylogeny. Likewise, this complete process was repeated for all such pairs of sequences in all protein families until the JTT matrix was finally developed (<xref ref-type="bibr" rid="B91">Jones et al., 1992</xref>).</p>
<p>BLOSUM is constructed on the basis of &#x3e;2,000 conserved amino acid arrangements representing 500 groups of diverse protein sequences (<xref ref-type="bibr" rid="B82">Henikoff and Henikoff, 1992</xref>). Unlike PAM matrices, the BLOSUM matrices indicate the actual identity percentage amongst sequences selected for constructing the matrices (<xref ref-type="bibr" rid="B82">Henikoff and Henikoff, 1992</xref>). For instance, BLOSUM52 represents that sequences nominated for generating matrix share an average identity value of 52%. Hence, a higher BLOSUM number represents less divergent sequences. As PAM matrices, except PAM1, are generated from an evolutionary model and BLOSUM matrices are generated from direct observations, PAM matrices have more evolutionary meaning as compared to the BLOSUM matrices. Hence, PAM matrices are generally utilized for reconstructing phylogenetic trees. Nevertheless, due to the mathematical extrapolation technique utilized, the PAM matrices are less realistic for divergent sequences. The BLOSUM matrices are generated from local sequence alignments of conserved sequence blocks, while the PAM1 matrix is generated based on the global alignment of full-length sequences comprising both variable and conserved regions. Hence, BLOSUM matrices are more advantageous during database searching as well as finding conserved domains within proteins (<xref ref-type="bibr" rid="B82">Henikoff and Henikoff, 1992</xref>).</p>
<p>Later, (<xref ref-type="bibr" rid="B2">Adachi and Hasegawa, 1996</xref>), Yanga &#x26; team (<xref ref-type="bibr" rid="B175">Yang et al., 1998</xref>), and Adachi &#x26; team (<xref ref-type="bibr" rid="B3">Adachi et al., 2000</xref>) utilized the Maximum Likelihood method for developing vertebrate mitochondrial, mammalian mitochondrial (mtMAM), and chloroplast sequences (cpREV) specific amino acid replacement models, respectively. As Adachi &#x26; team, Yanga &#x26; team, and Adachi &#x26; team utilized only 20, 23 and 10 sequences, respectively, for constructing an amino acid replacement model, the accuracy of their matrices is always under question (<xref ref-type="bibr" rid="B164">Whelan and Goldman, 2001</xref>). Later Whelan and Goldman combined the best attributes of both Maximum Likelihood (ML) and counting methods for developing a more powerful amino acid replacement model from an extensive database of different globular protein families (<xref ref-type="bibr" rid="B164">Whelan and Goldman, 2001</xref>). Recently, Le and the team have also developed a more robust amino acid substitution model for metazoan mitochondrial (mtMet), vertebrate mitochondrial (mtVer) and invertebrate mitochondrial (mtInv) (<xref ref-type="bibr" rid="B103">Le et al., 2017</xref>). Amino acid substitution models have also been developed for Influenza virus (FLU) (<xref ref-type="bibr" rid="B32">Dang et al., 2010</xref>), HIV between-patient matrix HIV-Bm (HIVb) (<xref ref-type="bibr" rid="B119">Nickle et al., 2007</xref>), HIV within-patient matrix HIV-Wm (HIVw) (<xref ref-type="bibr" rid="B119">Nickle et al., 2007</xref>), arthropod mitochondrial (mtART) (<xref ref-type="bibr" rid="B1">Abascal et al., 2007</xref>), retrovirus (rtREV) (<xref ref-type="bibr" rid="B40">Dimmic et al., 2002</xref>) and general &#x2018;Variable Time&#x2019; matrix (VT) (<xref ref-type="bibr" rid="B117">M&#xfc;ller and Vingron, 2000</xref>).</p>
</sec>
<sec id="s3-3">
<title>Codon substitution models</title>
<p>A codon is a continuous three DNA/RNA bases sequences, which encodes a specific amino acid or stop signal during protein synthesis. As there are four different nucleotides, there are only 64 possible codons. Out of these 64, only 61 code for specific amino acids, while rest three codes act as a stop codon. Since there are only 20 amino acids, more than one codon encodes one amino acid. This degeneracy property of genetic code enables us to distinguish between synonymous (do not alter encoded amino acid) and non-synonymous (alter encoded amino acid) substitution at the nucleotide level (<xref ref-type="bibr" rid="B176">Yang et al., 2000</xref>). Codon models are generally utilized to estimate evolutionary pressures on proteins across divergent lineages <italic>via</italic> comparing the ratio of substitution rates at non-synonymous (dN) and synonymous sites (dS) in the protein-coding regions (&#x3c9; &#x3d; dN/dS). Employing synonymous polymorphisms as a proxy of neutral diversity, one can estimate if non-synonymous polymorphisms are hindered or favored by natural selection. In the neutral evolving genes, the fixation rate of non-synonymous and synonymous mutation will be the same (&#x3c9; &#x3d; 1). During negative (purifying) selection, the non-synonymous mutation is not favored by natural selection and thus is eliminated, causing the fixation rate of non-synonymous mutation to be lower than the synonymous rate (&#x3c9;&#x3c; 1). During positive (adaptive) selection, the non-synonymous mutation is favored <italic>via</italic> Darwinian selection, thereby causing the fixation rate of non-synonymous mutation to be higher than the synonymous rate (&#x3c9;&#x3e; 1) (<xref ref-type="bibr" rid="B70">Gupta and Vadde, 2020</xref>).</p>
<p>A study reported that ancient proteins are under strong purifying selection, while newly developed proteins are under positive selection (<xref ref-type="bibr" rid="B161">Vishnoi et al., 2010</xref>). As newly developed young genes perform either highly specialized (if generated <italic>de novo</italic> or <italic>via</italic> horizontal transfer) or redundancy (if generated <italic>via</italic> duplication) functions, they are more at risk of either losing their function or gaining novel functions in succeeding lineages (<xref ref-type="bibr" rid="B41">Domazet-Loso and Tautz, 2003</xref>; <xref ref-type="bibr" rid="B33">Daubin and Ochman, 2004</xref>; <xref ref-type="bibr" rid="B168">Wolf et al., 2009</xref>; <xref ref-type="bibr" rid="B161">Vishnoi et al., 2010</xref>). Though initially, young genes experience a large number of adaptive mutations, the substitution of some of the mutations will slowly optimize the function of the gene in due course of time and hence the supply of new adaptive mutations will also reduce; hence, &#x3c9; value of a young gene will decline over time (<xref ref-type="bibr" rid="B161">Vishnoi et al., 2010</xref>; <xref ref-type="bibr" rid="B116">Moutinho et al., 2022</xref>). On the contrary, functions of old genes, like diabetic genes, are highly optimized and they are likely to have already exhausted all beneficial mutations by recent times and; thus, they are expected to evolve under negative selection and fix only neutral and/or nearly neutral mutations (<xref ref-type="bibr" rid="B161">Vishnoi et al., 2010</xref>).</p>
<p>Though &#x3c9; was originally designed for detecting selective pressure acting on a protein across divergent lineage, &#x3c9; can also be utilized for detecting selective pressure acting on a protein in a single population (<xref ref-type="bibr" rid="B100">Kryazhimskiy and Plotkin, 2008</xref>). However, selective pressure estimated <italic>via</italic> &#x3c9; on sequences sampled from a single population differs from that of the divergent lineages. For instance, though &#x3c9;&#x3c; 1 is a clear signature of negative selection across divergent lineages, weak negative or strong positive selection between population samples is also expected to produce &#x3c9;&#x3c; 1 (<xref ref-type="bibr" rid="B136">Roumagnac et al., 2006</xref>; <xref ref-type="bibr" rid="B86">Holt et al., 2008</xref>; <xref ref-type="bibr" rid="B100">Kryazhimskiy and Plotkin, 2008</xref>). Strong positive selection in a population will generate speedy sweeps at selected sites (but not at neutral sites, which are presumed to be independent). Thus, two individuals from the same population under strong positive selection are likely to contain identical alleles at each selected site, generating &#x3c9;&#x3c; 1 (<xref ref-type="bibr" rid="B100">Kryazhimskiy and Plotkin, 2008</xref>).</p>
</sec>
</sec>
<sec id="s4">
<title>Approaches to estimate selective pressure on the coding region of a gene</title>
<p>To date, numerous methods have been developed for estimating selective pressure on the coding region of a gene. Most models consider numerous factors like codon biases and variation amongst sites to estimate selective pressure more accurately. Initial models were designed to estimate global &#x3c9; for the entire sequence or for subsequences utilizing a sliding window approach. However, in reality, &#x3c9; varies amongst each amino acid site in sequence data or amongst each branch in a phylogeny. Recently more advanced approaches were developed to predict &#x3c9; per amino acid site (<xref ref-type="bibr" rid="B190">Yang, 2002</xref>; <xref ref-type="bibr" rid="B191">Suzuki, 2004b</xref>), which enable the identification of single sites under positive selection in spite of low global &#x3c9; value for the entire protein. All these models can be broadly classified as mechanistic, empirical and semi-empirical codon substitution models.</p>
<sec id="s4-1">
<title>Mechanistic codon substitution models</title>
<p>Mechanistic codon models detect selective pressure on the coding region of a gene utilizing a finite set of parameters, for instance, synonymous/non-synonymous rate ratio, transversion/transition rate ratio, and codon frequencies at equilibrium. As the mechanistic codon model utilizes a finite set of parameters, it is also known as the parametric codon substitution model<bold>.</bold> Mechanistic codon models focus mainly on silent-transversion, silent-transition, replacement-transversion rates, and replacement-transitions amongst sense codons and codon frequencies. Considering all parameters in a single codon model will be computationally more demanding. Thus to avoid this problem, several mechanistic codon models have been developed to date. Each mechanistic codon substitution models have distinctive parameters that differentiate the substitution rate at the nucleotide level and selective pressure at the protein level. Thus, each mechanistic model has the capacity to estimate selective forces acting on any protein in their unique way (<xref ref-type="bibr" rid="B166">Whelan et al., 2001</xref>; <xref ref-type="bibr" rid="B38">Delport et al., 2009</xref>). If selective pressure at the protein level is not considered, codon models will be equivalent to nucleotide substitution models. If the substitution rate at nucleotide is not considered, the codon model will be equivalent to amino acid substitution models (<xref ref-type="bibr" rid="B114">Miyazawa, 2011a</xref>). Several studies utilizing a large set of protein-coding sequences reported that codon substitution models are statistically more powerful than nucleotide and amino acid models (<xref ref-type="bibr" rid="B145">Seo and Kishino, 2009</xref>; <xref ref-type="bibr" rid="B115">Miyazawa, 2011b</xref>). However, the codon model having a larger substitution rate was reported to be equivalent to the amino acid substitution model (<xref ref-type="bibr" rid="B144">Seo and Kishino, 2008</xref>).</p>
<p>The first two mechanistic codon substitution models (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>; <xref ref-type="bibr" rid="B118">Muse and Gaut, 1994</xref>) were capable of estimating only the global &#x3c9; of the coding region of a gene. These two models considered transition/transversion ratio and codon frequencies for estimating &#x3c9; (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>; <xref ref-type="bibr" rid="B118">Muse and Gaut, 1994</xref>). Besides, Goldman and Yang (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>) also considered replacement probabilities amongst amino acids on the basis of the Grantham physicochemical distance matrix (<xref ref-type="bibr" rid="B64">Grantham, 1974</xref>). Later Nielsen and Yang (<xref ref-type="bibr" rid="B120">Nielsen and Yang, 1998</xref>) &#x26; Yang and the team (<xref ref-type="bibr" rid="B176">Yang et al., 2000</xref>) developed more robust mechanistic Bayesian models individually. It is pertinent to note that in the models developed by Goldman and Yang (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>) and Nielsen and Yang (<xref ref-type="bibr" rid="B120">Nielsen and Yang, 1998</xref>), the rate of substitution is proportional to the frequency of the target codon (which is not very mechanistic), and later many models employes these models to &#x201c;explain&#x201d; the stationary distribution in codons, whereas in the model developed by Muse and Gaut (<xref ref-type="bibr" rid="B118">Muse and Gaut, 1994</xref>), it is proportional to the target nucleotide, which is much more mechanistic considering the mutation process, and based on which later different type of mutation-selection (MutSel) model was developed [described below].</p>
<p>In 1994, Yang (<xref ref-type="bibr" rid="B178">Yang, 1994b</xref>) developed two approximation approaches for Maximum Likelihood phylogenetic estimation, which allow for varying substitution rates across nucleotide sites. The first, known as the &#x201c;discrete gamma model,&#x201d; approximates the gamma distribution by using many rate categories with equal probability for each category. Each category&#x2019;s mean is employed to depict all of the rates within that category. This method&#x2019;s performance has been shown to be rather acceptable, with four such categories seeming to be adequate to achieve both an optimum or near-optimal fit by the model to the data, as well as an acceptable approximation to the continuous distribution. The second strategy, dubbed the &#x201c;fixed-rates model,&#x201d; divides sites into multiple groups based on the rates anticipated by the star tree. When evaluating alternative tree topologies, sites in various classes are considered to evolve at these constant rates. Analyses of the data sets indicated that this approach might yield good results; however, it seems to have certain aspects with a least-squares pairwise comparison (<xref ref-type="bibr" rid="B178">Yang, 1994b</xref>). These models, however, overlook the fact that substitution rates of each amino acid differ distinctly. For instance, as only one transversion is required to convert the phenylalanine codon (UUU) into a leucine codon (UUG) as well as the tryptophan codon (UGG) into a leucine codon (UUG), they consider their substitution rate to be same (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>). But in reality, the probability of occurring a former event is approximately 5 times higher than a later event (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>).</p>
<p>Considering this lacuna, for the first time, in 2004, Whelan and Goldman developed a complete parametric model that considers several instantaneous substitutions (<xref ref-type="bibr" rid="B165">Whelan and Goldman, 2004</xref>). This model estimated substitution rate matrice for single-, double- and triple nucleotide mutation individually utilizing transition to transversion ratio and equilibrated frequency of mutated nucleotides. Later, these three matrices were joined together to estimate the general codon rate matrix. This method is reported to estimate the likelihood of parameters more accurately in comparison to other mechanistic models (<xref ref-type="bibr" rid="B165">Whelan and Goldman, 2004</xref>). Double and triple nucleotide substitutions are reported to occur through the mechanistic process, like during repairing DNA break (<xref ref-type="bibr" rid="B139">Sakofsky et al., 2014</xref>) or error-prone polymerase activity (<xref ref-type="bibr" rid="B79">Harris and Nielsen, 2014</xref>). Although double and triple substitutions rates are predicted to be two to three orders of magnitude lower than single substitutions (<xref ref-type="bibr" rid="B149">Smith et al., 2003</xref>; <xref ref-type="bibr" rid="B165">Whelan and Goldman, 2004</xref>; <xref ref-type="bibr" rid="B155">Tamuri et al., 2012</xref>), the model which included double and triple substitutions were reported to fit better in real data. Later several different models were developed by Doron-Faigenboim and Pupko (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>), Kosiol, Holmes and Goldman (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>), De Maio &#x26; team (<xref ref-type="bibr" rid="B37">De Maio et al., 2013</xref>), Miyazawa (<xref ref-type="bibr" rid="B115">Miyazawa, 2011b</xref>), Zoller and Schneider (<xref ref-type="bibr" rid="B188">Zoller and Schneider, 2012</xref>), Zaheri, Dib and Salamin (<xref ref-type="bibr" rid="B183">Zaheri et al., 2014</xref>), Venkat &#x26; team (<xref ref-type="bibr" rid="B160">Venkat et al., 2018</xref>) and Jones &#x26; team (<xref ref-type="bibr" rid="B93">Jones et al., 2018</xref>), which included double and triple substitution between codon.</p>
<p>Later, the model developed <italic>via</italic> Goldman and Yang (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>) was modified to include various nucleotide models (<xref ref-type="bibr" rid="B130">Pond et al., 2005</xref>; <xref ref-type="bibr" rid="B128">Pond and Frost, 2005</xref>; <xref ref-type="bibr" rid="B12">Arenas and Posada, 2014</xref>), estimate &#x3c9; variation across sites (<xref ref-type="bibr" rid="B182">Yang, 2007</xref>) and branches (<xref ref-type="bibr" rid="B182">Yang, 2007</xref>; <xref ref-type="bibr" rid="B48">Dutheil et al., 2012</xref>). In the model developed by Pond and Muse, they consider the possibility of site-to-site variation in synonymous and non-synonymous substitution rates in protein-coding DNA sequences and observed that within-gene variability in synonymous substitution rates is common (<xref ref-type="bibr" rid="B129">Pond and Muse, 2005</xref>). Another model developed by <xref ref-type="bibr" rid="B111">Mayrose et al. (2007)</xref> that uses two hidden Markov models and function on the spatial dimension. First and the second model depicts the dependency between adjacent non-synonymous and synonymous rates rates, respectively. They demonstrate that taking into consideration synonymous rate variability and dependence substantially improves the accuracy of &#x3c9; estimate, in particular for positively selected sites. In some models, codons were partitioned based on the physicochemical properties of the encoded amino acids (e.g., polarity or charge) (<xref ref-type="bibr" rid="B138">Sainudiin et al., 2005</xref>; <xref ref-type="bibr" rid="B169">Wong et al., 2006</xref>), codon bias (<xref ref-type="bibr" rid="B174">Yang and Nielsen, 2008</xref>) or the effects of GC contents (<xref ref-type="bibr" rid="B113">Misawa, 2011</xref>). Models in which codons were partitioned based on the physicochemical properties of the encoded amino acids explicitly parameterized physiochemical constraints due to non-synonymous substitution. The model developed by Goldman &#x26; Yang (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>) and Yang, Nielsen &#x26; Hasegawa (<xref ref-type="bibr" rid="B175">Yang et al., 1998</xref>) applied mathematical functions for modeling association amongst physiochemical properties and &#x3c9; parameter. Yang (<xref ref-type="bibr" rid="B180">Yang, 2000</xref>) permitted the effect of the physicochemical property to fluctuate among sites. Sainudiin &#x26; team (<xref ref-type="bibr" rid="B138">Sainudiin et al., 2005</xref>) and Wong &#x26; team (<xref ref-type="bibr" rid="B169">Wong et al., 2006</xref>) developed two separate models that at first divided non-synonymous substitutions into small groups in accordance with the pre-defined physiochemical property. As the main objective of these two models is to examine the impact of certain physicochemical properties of amino acids on the structure and function of a protein, their parameterization is focused on comparing the property-modifying substitutions rate with the property-conserving substitutions rate. Conant and Stadler (<xref ref-type="bibr" rid="B31">Conant and Stadler, 2009</xref>) estimated multiple amino acid properties <italic>via</italic> modeling exchangeability amongst non-synonymous codons as a linear combination of five pre-specified measures of physiochemical properties. This model enabled us to investigate the association between selection pressure and physicochemical properties while avoiding over parameterization of the codon model.</p>
<p>In 2008, Yang and Nielsen (<xref ref-type="bibr" rid="B174">Yang and Nielsen, 2008</xref>) developed the FMutSel model in which the amino acid frequencies are determined by the functional requirements of the protein (<xref ref-type="bibr" rid="B134">Rodrigue et al., 2008</xref>; <xref ref-type="bibr" rid="B17">Beaulieu et al., 2019</xref>). In the FMutSel model, each codon was allocated a fitness parameter. Dissimilarities in fitness parameters amongst two codons are utilized for specifying substitution rates in the Markov matrix <italic>via</italic> altering the rates specified by the standard mutation models (<xref ref-type="bibr" rid="B174">Yang and Nielsen, 2008</xref>). The FMutSel/FMutSel0 model combination has only been implemented in PAML4 with the M0 and M3 models so far. Model M0 implies that &#x3c9; across all branches and sites is constant, whereas Model M3 allows &#x3c9; to vary between sites (<xref ref-type="bibr" rid="B44">Du et al., 2014</xref>). Likewise, in 2010, Rodrigue and the team developed a complex extension of this model in which site-specific amino acid propensity scores are utilized for estimating scaled selection coefficients, which in turn was utilized for identifying substitution rates (<xref ref-type="bibr" rid="B135">Rodrigue et al., 2010</xref>). In 2013, De Maio and team (<xref ref-type="bibr" rid="B37">De Maio et al., 2013</xref>) reported that when some models were employed to compute &#x3c9; heterogeneity on data, where both multiple-nonsynonymous rates and double &#x26; triple codon modification occur, they yielded high false-positive rates. Recently Venkat &#x26; team (<xref ref-type="bibr" rid="B160">Venkat et al., 2018</xref>) reported that when branch-site codon models are employed in branch-specific tests to detect positive selection, double modification may cause high false-positive rates. To avoid this problem, recently, Dunn and the team developed a statistically more powerful general-purpose parametric modeling framework for codons (<xref ref-type="bibr" rid="B47">Dunn et al., 2019</xref>). By including information about all possible instantaneous codon substitutions, along with instantaneous double and triple nucleotide substitution and multiple non-synonymous rates, both accuracy, as well as statistical power was highly improved (<xref ref-type="bibr" rid="B47">Dunn et al., 2019</xref>).</p>
</sec>
<sec id="s4-2">
<title>Empirical codon substitution model</title>
<p>Though empirical codon models are highly useful in understanding protein evolution as well as in phylogenetic applications, only a few models have been developed to date (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>). Substitution rates amongst codons were empirically determined to utilize a large set of protein-coding sequences (<xref ref-type="bibr" rid="B142">Schneider et al., 2005</xref>; <xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>). Unlike mechanistic models, empirical codon substitution cannot distinguish between the substitution rate at the nucleotide level and selective pressure at the protein level (<xref ref-type="bibr" rid="B114">Miyazawa, 2011a</xref>). Thus, there is no parameter except codon frequencies for tailoring of each protein family (<xref ref-type="bibr" rid="B114">Miyazawa, 2011a</xref>). Delport and team reported that empirical substitution matrices represent average propensities of substitutions across several protein families <italic>via</italic> sacrificing gene-level resolution (<xref ref-type="bibr" rid="B39">Delport et al., 2010</xref>).</p>
<p>For the first time in 1990, Sch&#xf6;niger and the team constructed counted codon-codon substitutions matrix on the basis of &#x223c;800 pairwise alignments of 41 actin genes (<xref ref-type="bibr" rid="B143">Sch&#xf6;niger et al., 1990</xref>). However, due to a lack of a better electronic facility, this matrix lost its fame in a short interval of time (<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, 2012</xref>). Additionally, as this matrix was developed based on a small number of sequences, it was less reliable. Later in 2005, Schneider and the team (<xref ref-type="bibr" rid="B142">Schneider et al., 2005</xref>) developed another empirical codon model utilizing a somewhat similar approach employed by Gonnet and the team (<xref ref-type="bibr" rid="B61">Gonnet et al., 1992</xref>) for constructing an amino acid substitution matrix. Since the development of the first amino acid substitution matrix (<xref ref-type="bibr" rid="B35">Dayhoff et al., 1978</xref>), it was for the first time Goonet &#x26; team (<xref ref-type="bibr" rid="B61">Gonnet et al., 1992</xref>) and Jones &#x26; team (<xref ref-type="bibr" rid="B91">Jones et al., 1992</xref>) developed two individual models based on the sufficiently large amount of sequences and thus are more reliable. In 2007, Kosiol and the team (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>) developed an empirical codon substitution matrix utilizing an extensive database of protein-coding DNA sequences. They reported that the accuracy of the model gets significantly improved by considering instantaneous double and triple substitution. Additionally, the amino acid encoded by each codon, associations amongst codons, and physicochemical properties of amino acids is key factors for driving the process of codon evolution (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>). Empirical codon substitution matrix is reported to outperform mechanistic codon substitution matrix when utilized in likelihood-based phylogenetic analysis (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>). Empirical codon models can also be utilized to detect different lineages sampled in a single phylogenomic dataset (<xref ref-type="bibr" rid="B37">De Maio et al., 2013</xref>) rather than depending on a general sequence database for instance Pandit database (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>). In 2014, Bloom developed another novel model that depicts the experimental determination of a parameter-free evolutionary model <italic>via</italic> deep sequencing, mutagenesis, and functional selection (<xref ref-type="bibr" rid="B21">Bloom, 2014</xref>). Employing this, Bloom build a model of influenza nucleoprotein evolution that represents the gene phylogeny in a far better way as compared to earlier existing models with nearly hundreds of free parameters. He also emphasized that the data provided by these types of high-throughput experiments can significantly improve the accuracy of both phylogenetic as well as genetic studies.</p>
</sec>
<sec id="s4-3">
<title>Semi-empirical</title>
<p>The semi-empirical codon substitution matrix is often called a mixed empirical codon substitution matrix because it combines empirical substitution rate with mechanistic parameters of codon evolution (<xref ref-type="bibr" rid="B47">Dunn et al., 2019</xref>). For the first time, Doron-Faigenboim &#x26; Pupko combined existing empirical amino acid substitution matrices with mechanistic parameters (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>). They assumed that the substitution rate between non-synonymous substitution amongst codons was equal to the pre-estimated amino-substitution rate, which was obtained by utilizing 189 parameters and a huge amount of amino acid sequences (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>). Later Kosiol and team, utilizing 1830 codon substitution parameters and large datasets, developed the first fully empirical codon model and then appended those models with mechanistic parameters for codon evolution (<xref ref-type="bibr" rid="B99">Kosiol et al., 2007</xref>). Subsequently, De Maio and team (<xref ref-type="bibr" rid="B37">De Maio et al., 2013</xref>), developed another model with almost the same accuracy but was less complex than the model developed by Doron-Faigenboim &#x26; Pupko (<xref ref-type="bibr" rid="B42">Doron-Faigenboim and Pupko, 2007</xref>). The empirical matrices in those studies denote a wide range of amino acid change propensity (<xref ref-type="bibr" rid="B47">Dunn et al., 2019</xref>). Later, Zoller &#x26; Schneider (<xref ref-type="bibr" rid="B188">Zoller and Schneider, 2012</xref>) and Miyazawa (<xref ref-type="bibr" rid="B114">Miyazawa, 2011a</xref>) developed different methods for tailoring information contained in an empirical substitution matrix to a specific dataset, and the benefit of these two approaches is that they can easily distinguish between substitution rate at the nucleotide level and selective pressure at the protein level.</p>
</sec>
</sec>
<sec id="s5">
<title>Applications of the codon substitution model</title>
<p>Codon substitution models are mainly utilized in detecting selective pressure on a protein, codon usage bias, ancestral reconstruction, and phylogenetic reconstruction. All the applications available in recent literature are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Applications of codon substitution models.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">S. No</th>
<th align="center">Applications</th>
<th align="center">References</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="left">Identifying heterogeneous selection pressure at amino acid sites</td>
<td align="left">
<xref ref-type="bibr" rid="B176">Yang et al. (2000)</xref>
</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">Identifying molecular adaptation at individual sites along specific lineages</td>
<td align="left">
<xref ref-type="bibr" rid="B173">Yang and Nielsen, (2002)</xref>
</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Phylogenetic reconstruction</td>
<td align="left">
<xref ref-type="bibr" rid="B133">Ren et al. (2005)</xref>
</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Codon usage bias</td>
<td align="left">
<xref ref-type="bibr" rid="B170">Wu et al. (2007)</xref>, <xref ref-type="bibr" rid="B174">Yang and Nielsen, (2008)</xref>, <xref ref-type="bibr" rid="B185">Zhao et al. (2016)</xref>
</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Reconstructing ancestral coding sequences</td>
<td align="left">
<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, (2009)</xref>
</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">Molecular dating &#x26; functional analysis</td>
<td align="left">
<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, (2012)</xref>
</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">Evolution of sexual chromosomes, gene families, host-pathogen interactions or regulatory networks</td>
<td align="left">
<xref ref-type="bibr" rid="B23">Cannarozzi and Schneider, (2012)</xref>
</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">Identification and estimation of conservation at synonymous sites</td>
<td align="left">
<xref ref-type="bibr" rid="B137">Rubinstein et al. (2012)</xref>
</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">Detect pathogen evolutionary rate variation</td>
<td align="left">
<xref ref-type="bibr" rid="B16">Baele et al. (2016)</xref>
</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">Identify antibody lineage</td>
<td align="left">
<xref ref-type="bibr" rid="B84">Hoehn et al. (2017)</xref>
</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">Detect evolutionary histories under time-dependent substitution rates</td>
<td align="left">
<xref ref-type="bibr" rid="B112">Membrebe et al. (2019)</xref>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-1">
<title>Studying selective pressure on a protein</title>
<p>Recent advancements in high throughput sequencing technologies have enabled the generation of a huge amount of sequence data (<xref ref-type="bibr" rid="B71">Gupta et al., 2017</xref>; <xref ref-type="bibr" rid="B69">Gupta and Vadde, 2019b</xref>; <xref ref-type="bibr" rid="B70">2020</xref>). This enormous amount of sequence data provides an opportunity to detect a direct association between selective pressure and the function of any protein in a more comprehensive way (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). Codon models are commonly utilized for identifying candidate genes and their variants under positive selection (<xref ref-type="bibr" rid="B125">Ouyang and Liang, 2007</xref>; <xref ref-type="bibr" rid="B126">Parto and Lartillot, 2018</xref>; <xref ref-type="bibr" rid="B47">Dunn et al., 2019</xref>). Initially, codon models presume that synonymous and non-synonymous substitution rates among sites as well as throughout the phylogenetic history, are constant (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). Though the majority of proteins are evolved under purifying selection, the positive selection may affect a few lineages. During this adaptive evolution, only a few protein sites have the capability to increase protein fitness during amino acid substitution (<xref ref-type="bibr" rid="B132">Pupko and Galtier, 2002</xref>). Hence, these codon model approaches presuming constant selective pressure over time as well as across sites lack the power to detect genes evolving under positive selection (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). Subsequently, several situations of variation in the selective pressure was included with the model developed by Muse &#x26; Gaut (<xref ref-type="bibr" rid="B118">Muse and Gaut, 1994</xref>) and Goldman &#x26; Yang (<xref ref-type="bibr" rid="B60">Goldman and Yang, 1994</xref>). These models were later utilised extensively to detect positive selection by likelihood ratio test comparing two nested models. One model (null hypothesis) do not permit positive selection while other model (alternative hypothesis) permit positive positive selection. Positive selection is identified when model permitting sites/lineages under positive selection fits data significantly better than the model restricting the site/lineages under positive selection. Nevertheless, if few parameters become invaluable or because of boundary problems, the asymptomatic null distribution may differ from the standard (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>).</p>
<p>Additionally, the codon substitution model can be utilized for detecting site-specific positive selection in proteins. Later, this information can be used for testing the biological hypothesis through laboratory experiments (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). For instance, in 2005, Swayer and the team reported that a small portion of TRIM5&#x3b1;, an immune defense protein, was recognized to be under positive selection. Later, through functional analysis, they confirmed the significance of the peptide segment in species-specific viral inhibition (<xref ref-type="bibr" rid="B140">Sawyer et al., 2005</xref>). The conditional selection model developed by Chen and the team may be utilized particularly for detecting interaction amongst sites during drug resistance (<xref ref-type="bibr" rid="B27">Chen and Lee, 2006</xref>). Considering this earlier, we also employed a phylogenetic approach implemented in the PAML&#x2019;s CODEML modeling tool to identify the kind of selection operating on T2D genes in the Drosophila genus. The data showed that the gene sequences encoding T2D are evolving under purifying selection. However, few membrane protein sites, including those encoded by CG8051, ZnT35C, and kar, are substantially evolving under positive selection. This may be due to adaptive evolution in response to changes in the niche, food, or other environmental conditions (<xref ref-type="bibr" rid="B70">Gupta and Vadde, 2020</xref>). Thus, the identification of selective pressure <italic>via</italic> codon substitution models may provide detailed insight into disease progression, pathogenic drug resistance, and epidemic dynamics (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>).</p>
</sec>
<sec id="s5-2">
<title>Codon usage bias</title>
<p>Gene expression is modulated through transcription (DNA to mRNA) and translation (mRNA to protein) mechanisms (<xref ref-type="bibr" rid="B186">Zhou et al., 2016</xref>). Promoter strength &#x26; RNA stability are mainly responsible for mRNA concentrations and transcript levels &#x26; protein stability is responsible for protein concentrations in any cell (<xref ref-type="bibr" rid="B89">Ikemura, 1985</xref>; <xref ref-type="bibr" rid="B147">Sharp et al., 1986</xref>). During translation, the information is transmitted as codons. This genetic code is degenerate in nature, i.e., except for tryptophan and methionine, more than one codon (synonymous codons) can encode a single amino acid (<xref ref-type="bibr" rid="B162">Wang et al., 2018</xref>). In coding sequences of many organisms, these synonymous codons are utilized at unequal frequencies (<xref ref-type="bibr" rid="B25">Chakraborty et al., 2017</xref>). This phenomenon is called codon usage bias. Preferred codons are more frequently utilized in highly expressed genes (<xref ref-type="bibr" rid="B186">Zhou et al., 2016</xref>). The degree of codon usage bias differs amongst genes &#x26; species and is mainly affected <italic>via</italic> neutral selection, directional mutation, tRNA abundance (<xref ref-type="bibr" rid="B123">Olejniczak and Uhlenbeck, 2006</xref>), selection for efficient translation initiation (<xref ref-type="bibr" rid="B184">Zalucki et al., 2007</xref>), gene length (<xref ref-type="bibr" rid="B151">Sun et al., 2009</xref>), an expression level (<xref ref-type="bibr" rid="B83">Hiraoka et al., 2009</xref>), DNA replication initiation site (<xref ref-type="bibr" rid="B87">Huang et al., 2009</xref>), <italic>etc.</italic> Codon usage bias can also be utilized in detecting phylogenetic trees amongst species (<xref ref-type="bibr" rid="B170">Wu et al., 2007</xref>; <xref ref-type="bibr" rid="B185">Zhao et al., 2016</xref>). In 2016, SENCA (site evolution of nucleotides, codons, and amino acids), a codon substitution model, was developed that distinctly describes (a) preferences amongst synonymous codons, (b) amino acids, and (c) nucleotide processes that apply on all sequence sites such as the mutational bias (<xref ref-type="bibr" rid="B131">Pouyet et al., 2016</xref>). This model assumes that the vast majority of synonymous substitutions are not neutral and can predict more accurate estimates of selection in comparison to more traditional codon sequence models (<xref ref-type="bibr" rid="B131">Pouyet et al., 2016</xref>).</p>
</sec>
<sec id="s5-3">
<title>Ancestral reconstruction</title>
<p>Codon substitution models are also utilized for reconstructing ancestral coding sequences through parsimony and Maximum Likelihood approaches (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). These ancestral sequences can further be utilized to detect alterations that have been experienced in every branch of phylogeny and at each individual site of the gene sequence. Several studies have utilized ancestral state information to understand protein evolution and episodic or lineage-specific base composition (<xref ref-type="bibr" rid="B108">Long and Langley, 1993</xref>; <xref ref-type="bibr" rid="B7">Akashi, 1996</xref>; <xref ref-type="bibr" rid="B49">Eanes et al., 1996</xref>; <xref ref-type="bibr" rid="B54">Fitch et al., 1997</xref>; <xref ref-type="bibr" rid="B153">Takano-Shimizu, 2001</xref>). For instance, the evolution of steroid receptors (<xref ref-type="bibr" rid="B159">Thornton et al., 2003</xref>) and ancestral archosaur visual pigment rhodopsin (<xref ref-type="bibr" rid="B26">Chang et al., 2002</xref>). Ancestral sequence reconstruction is also employed in studying HIV evolution (<xref ref-type="bibr" rid="B56">Gaschen et al., 2002</xref>), protein engineering (<xref ref-type="bibr" rid="B30">Cole and Gaucher, 2011</xref>), and understanding variation in DNA turnover because of indels and substitutions amongst eutherian mammalian lineages (<xref ref-type="bibr" rid="B20">Blanchette et al., 2004</xref>). Additionally, numerous population genetic tests depend on this ancestral reconstruction to understand the impact of natural selection on the functional classes of mutations or genetic regions (<xref ref-type="bibr" rid="B6">Akashi, 1995</xref>; <xref ref-type="bibr" rid="B157">Templeton, 1996</xref>; <xref ref-type="bibr" rid="B8">Akashi, 1999</xref>; <xref ref-type="bibr" rid="B152">Suzuki and Gojobori, 1999</xref>) and also identify coevolving nucleotides/amino acids (<xref ref-type="bibr" rid="B124">Osada and Akashi, 2012</xref>; <xref ref-type="bibr" rid="B105">Liao et al., 2013</xref>).</p>
</sec>
<sec id="s5-4">
<title>Phylogenetic reconstruction</title>
<p>Codon models reconstruct phylogenetic trees by considering genetic code and the rate of non-synonymous &#x26; synonymous base substitutions. In almost every protein-coding gene, the incidence of non-synonymous substitution is less and is mainly involved in early divergence. Synonymous substitutions are higher and are responsible for recent divergence. By considering this information, the codon models may be utilized in reconstructing phylogenetic trees more accurately (<xref ref-type="bibr" rid="B133">Ren et al., 2005</xref>). Earlier studies have reported that though nucleotide substitution models are modified to accommodating differences in the evolutionary dynamics at three codon positions (<xref ref-type="bibr" rid="B179">Yang, 1996</xref>), the accuracy of this model is lower as compared to codon models. Nevertheless, due to the lack of efficient codon-based tree search methods, tree inference from coding sequence data is generally performed under DNA and AA (amino acid) models. Because of the 61 &#xd7; 61 matrix, tree generation utilizing codon models is computationally more demanding. To date, no efficient methods have been developed for phylogeny reconstruction utilizing the codon substitution model on a large dataset. For the small dataset, phylogeny can be constructed using CODEML from the PAML package (<xref ref-type="bibr" rid="B182">Yang, 2007</xref>). However, Yang has implemented a heuristic algorithm in PAML, which is not the most efficient approach. One possible way to reconstruct an efficient phylogenetic tree is by initially generating numerous phylogenetic trees utilizing both DNA and amino acid substitution models. Later, these trees can be utilized for constructing more accurate trees under efficient Maximum Likelihood (ML) heuristics under codon models (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>).</p>
<p>Another significant approach in the reconstruction of the phylogenetic tree is by implementing codon models with a Bayesian framework and sampling topological space with an efficient Markov chain Monte Carlo (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>). By using the Bayesian framework, we can either get similar or even better tree topology in comparison to ML approaches. The main reason for achieving better tree topology <italic>via</italic> the Bayesian framework is that ML approaches search for a single best tree while the Bayesian framework scan cluster of best trees. The benefits of the Bayesian framework can also be explained <italic>via</italic> the matter of probability. The best tree generated <italic>via</italic> ML may &#x223c;90% probability of demonstrating the real information. On the contrary, the Bayesian framework generates hundreds/thousands of near-optimal/optimal having &#x223c;90% probability representing the real information. Hence, the phylogenetic tree generated <italic>via</italic> the Bayesian framework is more realistic than the phylogenetic tree generated <italic>via</italic> ML methods. However, in this Bayesian framework, the rate of substitution will differ for each three codon sites because it considers different data partitions. Thus, codon model usage may serve as an important asset while comparing several candidate trees inferred under either DNA or amino acid models (<xref ref-type="bibr" rid="B9">Anisimova and Kosiol, 2009</xref>).</p>
</sec>
</sec>
<sec id="s6">
<title>Limitation and development of next-generation codon model</title>
<p>Though various codon models develop to date provide researchers with a more powerful bioinformatics toolbox, these models&#x2019; enormous exchangeability matrices (61 &#xd7; 61, excluding stop codons) make implementation difficult (<xref ref-type="bibr" rid="B15">Arenas, 2015b</xref>). Thus, the development of next-generation development of the codon model with significant attention to model choice as well as the implantation assumptions is highly demanded (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). This can be achieved by using a substantial quantity of data and a considerable amount of computing power. Fortunately, efforts to optimize codon-based algorithms are developing new evolutionary tools for simulating (<xref ref-type="bibr" rid="B55">Fletcher and Yang, 2009</xref>; <xref ref-type="bibr" rid="B13">Arenas, 2012</xref>) and analyzing (<xref ref-type="bibr" rid="B58">Gil et al., 2013</xref>; <xref ref-type="bibr" rid="B189">Zoller et al., 2015</xref>) the codons evolution, even though additional research is needed in this area (<xref ref-type="bibr" rid="B15">Arenas, 2015b</xref>). In addition to the development of new empirical models, these models may follow two fascinating trends; First, evaluate heterogeneity throughout the sequence and across time since various sites/regions and time periods may evolve differently under distinct models (<xref ref-type="bibr" rid="B14">Arenas, 2015a</xref>; <xref ref-type="bibr" rid="B189">Zoller et al., 2015</xref>). It is important to remember that these partition methods may be highly realistic, for instance, by using distinct models for coding and non-coding regions. And it is well-known that &#x3c9; estimations may be influenced by codon models that are based on differences in codon frequency among sites. Thus, there is a need for programs and methods that can determine which codon substitution model works best in a given codon region and time scale (<xref ref-type="bibr" rid="B15">Arenas, 2015b</xref>). The second possible trend may be the integration of protein structure data into codon models. Codon models could take into account information about the proteins&#x2019; functions and their folding stability (<xref ref-type="bibr" rid="B63">Grahnen et al., 2011</xref>; <xref ref-type="bibr" rid="B106">Liberles et al., 2012</xref>). But if the protein structure changes over time or if more than one protein structure is required to depict the encoded proteins in the dataset, these implementations would incur high computational costs (<xref ref-type="bibr" rid="B15">Arenas, 2015b</xref>).</p>
<p>Considering the limitation above and with the aim to develop more robust codon-substitution models, in 2010, Zoller &#x26; Schneider (<xref ref-type="bibr" rid="B187">Zoller and Schneider, 2010</xref>) investigated 3,666 codon substitution matrices for detecting the most vital parameters of any codon model. They employed principal component analysis (PCA) to identify the numerous substitution rates that may co-vary across diverse genes. Each individual 3,666 matrices were estimated employing &#x201c;<italic>XRate&#x201d;</italic> from a single multiple sequence alignment generated from Mammalian coding sequences. Irrespective of large variance related to parameters computed from very less data, PCA analysis was able to capture a few significant factors. As per PCA analysis, one of the most important parameters in any codon substitution model is the &#x3c9; value. Amusingly, the substitutions in serine demand two nucleotide alterations and a transitional non-synonymous modification were grouped together with the non-synonymous substitutions. The second most key parameter detected is the ratio amongst substitutions having only one nucleotide dissimilarity and those with two/three dissimilarities. Interestingly, this parameter is not considered in any of the codon-substitution models developed to date. As PCA analysis determines factors that differ maximum in any dataset, there might be an evolutionary use that affects the multi-nucleotide substitutions number that might get stable during coding sequence evolution (<xref ref-type="bibr" rid="B187">Zoller and Schneider, 2010</xref>). However, this method was unable to detect other important parameters associated with codon substitution models.</p>
<p>Another study reported that, even though we assume phylogeny on which molecular evolution is modeled is a more appropriate representation of the evolutionary history of any lineage/taxa, but this might not be true in the case of a small dataset or if recombination has been ignored while generating tree topology (<xref ref-type="bibr" rid="B38">Delport et al., 2009</xref>). It is possible to include such uncertainty in tree topology <italic>via</italic> Bayesian methods (<xref ref-type="bibr" rid="B176">Yang et al., 2000</xref>). For example, MrBayes employed codon substitution models for generating tree topology (<xref ref-type="bibr" rid="B88">Huelsenbeck and Ronquist, 2001</xref>). These methods relax the assumption that a specific tree is correct but not the assumption that a correct, though unknown, tree exists. One of the probable solutions for the recombination problem is the introduction of population genetics approximation within the coalescent which co-estimates recombination rate and selective pressure (<xref ref-type="bibr" rid="B167">Wilson and McVean, 2005</xref>). Another solution is the identification of recombination breakpoints as well as the prediction of a distinct phylogeny for each individual recombinant. Parameters of these codon models are consecutively calculated in the usual way, except that phylogenies, as well as branch lengths, are partition-specific, while the remaining parameters are shared across all segments (<xref ref-type="bibr" rid="B141">Scheffler et al., 2006</xref>). It is also highly advisable to incorporate different synonymous rates in each recombinant because recombination may also lead to differences in synonymous rates (<xref ref-type="bibr" rid="B141">Scheffler et al., 2006</xref>). Software, namely, genetic algorithm for recombination detection (GARD), is one of the suitable algorithms for the detection of individual adaptive evolving sites in recombination sequences (<xref ref-type="bibr" rid="B97">Kosakovsky Pond et al., 2006</xref>). It is pertinent to note that codon models, particularly those that take rate variation into account, may tolerate modest amounts of recombination (<xref ref-type="bibr" rid="B11">Anisimova et al., 2003</xref>; <xref ref-type="bibr" rid="B141">Scheffler et al., 2006</xref>). False positive rate estimates may be inflated, however, if recombination rates in such models are very high. Thus, positive selection predictions should be regarded with care for genes with the greatest recombination rates (<xref ref-type="bibr" rid="B34">Davydov et al., 2019</xref>).</p>
<p>Earlier several studies have also reported that though codon models developed to date are extensively employed for estimating selective pressure on the gene(s) (<xref ref-type="bibr" rid="B110">MacCallum and Hill, 2006</xref>) and scanning genes under positive selection (<xref ref-type="bibr" rid="B104">Li et al., 2010</xref>), most of these models generally aimed at investigating the recurrent diversifying selection. Considering this, a few definite models were also developed for investigating the directional selection and were employed on viral data (<xref ref-type="bibr" rid="B98">Kosakovsky Pond et al., 2008</xref>; <xref ref-type="bibr" rid="B101">Lacerda et al., 2010</xref>). Nevertheless, these directional models are not time-reversible. Recent advancement in sequence technologies enables enormous sequence growth and the development of empirical codon models. Even successful attempts were made to combine empirical estimates along with conventional parameters (<xref ref-type="bibr" rid="B167">Wilson and McVean, 2005</xref>).</p>
<p>In addition to phylogeny, codon substitution models can also be employed for studying synonymous codon bias, which may develop because of optimizing for translational kinetics, efficiency, and robustness. Selection against the non-optimal codons often causes a negative correlation amongst synonymous substitution rates and codon bias (<xref ref-type="bibr" rid="B5">Akashi and Eyre-Walker, 1998</xref>). Nevertheless, codon bias is generally investigated with different codon adaptation indexes on the basis of single sequences instead of estimating <italic>via</italic> multiple sequence alignment and other parameters of a substitution model. Markov models having fewer states, for instance, codons translated <italic>via</italic> distinct tRNAs, can be employed for studying codon usage as well as asymmetric selective effects (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). On the other hand, mutation and selection may be modeled distinctly for investigating the effects of mutational biases and translational selection (<xref ref-type="bibr" rid="B121">Nielsen and Yang, 2003</xref>). Employing such models in 2007, Nielsen and the team investigated the evolution of codon usage over time (<xref ref-type="bibr" rid="B122">Nielsen et al., 2007</xref>). In another study, (<xref ref-type="bibr" rid="B174">Yang and Nielsen, 2008</xref>) computed optimal codon frequencies as well as mutational bias parameters across multiple species and genes. Further, LRT amongst pairs of nested selection mutation models can be employed for investigating if the codon bias is because of the mutational bias only. This model was further designed to include site specific amino acid profiles, which in turn provide an attractive substitute for fixed as well as random effects models (<xref ref-type="bibr" rid="B135">Rodrigue et al., 2010</xref>). Utilizing the Dirichlet process, site-profiles were fitted to the dataset in the Bayesian framework.</p>
<p>One of the underlying presumptions of the codon substitution model is that the rate of codon change is a product of the mutation fixation probability and the mutation rate (<xref ref-type="bibr" rid="B95">Kimura, 1962</xref>); this, in turn, forms a significant connection to the population genetic theory. Thus, we may also employ codon substitution models for estimating relationships amongst interspecific and population parameters, e.g., the scaled selection coefficient (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Given the importance of codon-based models for detecting diversifying positive selection, previous research has focused on two aspects of codon-based models that are important for population genetic interpretations of diversifying positive selection (<xref ref-type="bibr" rid="B158">Thorne et al., 2012</xref>). At first, diversified positive selection is a kind of positive selection which is often referred to an allele having a fitness advantage. When alleles&#x2019; relative fitnesses are largely consistent across environments, the presence of positive selection is determined by the alleles involved in the substitution rather than the codon position and/or lineage influenced by substitution. On the other hand, codon-based substitution models often seek to identify instances when non-synonymous mutations are beneficial independent of the specific alleles present before and after the mutation. Secondly, diversifying positive selection within codon-based substitution models should be interpreted with care while analysing population genetics. Even though several parameterizations of codon-based models having diversifying positive selection have been developed, they seems have this simple model for substitution rates, as depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Simple model for substitution rates, where Rij is a nonsynonymous rate, u is a proportionality constant and &#x03BC;ij is the rate at which i mutates to j.</p>
</caption>
<graphic xlink:href="fgene-14-1091575-g001.tif"/>
</fig>
<p>Where u is a proportionality constant and <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>
<sub>ij</sub> is the rate at which i mutates to j. A population genetic interpretation of a non-synonymous rate R<sub>ij</sub> would therefore have <italic>&#x3c9;</italic> proportional to P (Z<sub>ij</sub>), which is the fixation probability approximation by Kimura (<xref ref-type="bibr" rid="B95">Kimura, 1962</xref>). One possible way to achieve this is to have all non-synonymous modifications be neutral with respect to selection; however, this would result in &#x3c9;&#x3d; 1, which would negate the necessity inclusion for the <italic>&#x3c9;</italic> parameter. The relative fitness of alleles might also be determined by whether they represent a novel mutation. This would mean that differences in fitness across alleles have nothing to do with the DNA that code for them.</p>
<p>In 2003, Nielsen et al. tried to develop such a model. Interestingly, they allow for variation amongst codon sites. For non-synonymous modifications affecting a specific codon position in a certain lineage, &#x3c9; was considered to be independent of the decoded amino acids before and after the modification. Since &#x3c9; was independent of the amino acids involved in the change, Nielsen et al. were able to derive stationary sequence distributions that were independent of the &#x3c9; value. Since the stationary distribution does not change with codon locations and stationarity can be presumed if the &#x3c9; value for a branch refers to a small or large population, inferences can be derived more straightforward (<xref ref-type="bibr" rid="B121">Nielsen and Yang, 2003</xref>). It is pertinent to note that inference of stationary distribution was also possible before as in (<xref ref-type="bibr" rid="B121">Nielsen and Yang, 2003</xref>), however not much studies have been done. Earlier, Halpern and Bruno (<xref ref-type="bibr" rid="B78">Halpern and Bruno, 1998</xref>) also developed the MutSel model to unmask the mechanistic, population-genetic explanation of evolution. In this method, a nucleotide mutation model that is the same for all sites is combined with fixation probability calculated from site-specific vectors of fitness coefficients under the assumption of a Wright-Fisher population with mutation and selection (<xref ref-type="bibr" rid="B92">Jones et al., 2017</xref>). This framework offers a systematic approach to generating realistic sequence alignments that are capable for detecting positive selection by directly relating &#x3c9; to fitness differences across amino acids. By forcing changes in fitness coefficients at predetermined sites and branches, extensions of the MutSel model (<xref ref-type="bibr" rid="B43">dos Reis, 2015</xref>) can also capture episodic positive selection. Irrespective of all these advancements, implementation of population genetic theory in the codon models is still in the infancy stage (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). One important challenge is how to differentiate between episodic changes in fitness landscapes and shifting balance in the model. Positive selection <italic>via</italic> shifting balance is an autonomous, unpredictable, and site-specific mechanism. So, the key question is, how often is shifting balance in real-world data? (<xref ref-type="bibr" rid="B158">Thorne et al., 2012</xref>; <xref ref-type="bibr" rid="B92">Jones et al., 2017</xref>). Another difficulty is posed by the fact that mutation-selection equilibrium may be disrupted by a wide variety of population genetic processes and how to include all these parameters in the model (<xref ref-type="bibr" rid="B158">Thorne et al., 2012</xref>).</p>
<p>There is also still scope for Monte Carlo approach development. Specifically, to date, &#x201c;data-augmentation-based&#x201d; methods have received very little attention in terms of codon substitution model development. This &#x201c;data-augmentation-based&#x201d; is though short-lived but has computational benefits (<xref ref-type="bibr" rid="B36">de Koning et al., 2010</xref>). For instance, thermodynamic integration is computationally expensive and, hence, not much used in molecular evolutionary or Bayesian phylogenetic applications. This is why the harmonic mean estimator (HME), which has an infinite variance and produces less reliable results (<xref ref-type="bibr" rid="B102">Lartillot and Philippe, 2006</xref>), is still widely used. Advancement in this direction, nevertheless, is also in full swing. For example, in 2011, Xie and the team (<xref ref-type="bibr" rid="B171">Xie et al., 2011</xref>) developed a more robust method, namely, the &#x201c;stepping-stone method&#x201d;, on the basis of similar concepts, though employing a discrete path in preference to a continuous one. In the near future, there is also scope for combining &#x201c;thermodynamic-based&#x201d; methods with &#x201c;data-augmentation based&#x201d; approaches. The &#x201c;stepping-stone&#x201d; approaches, along with other recently developed computational methods, may also contribute significantly to developing Bayes factor, thereby providing a wide-range evaluation of the performance of numerous different codon substitution modeling methods.</p>
<p>In 2010, Du and the team proposed new codon-based ancestral reconstruction approaches that permit to examine changes in codon usage bias in rhodopsin, which in turn might be responsible for shifts in the visual ecology within the early mammals (<xref ref-type="bibr" rid="B45">Du, 2010</xref>). Using the same approach, they observed an evolutionary trend towards enhanced GC-ending codons at three early mammalians, i.e., therian, placental and mammalian lineages of rhodopsin. However, they also proposed that there is still scope for incorporating a Bayesian distribution of different ancestral states while estimating the Akashi ratio for calculating deviations from equilibrium codon usage, as well as simulations for accessing the significance of the deviations detected for rhodopsin (<xref ref-type="bibr" rid="B45">Du, 2010</xref>).</p>
<p>In one study, authors proposed that augmenting codon model application along with information obtained from other approaches, for instance, population genetics, coalescence, and HMMs may enable us to understand the evolution of the complex system in a more comprehensive way. For instance, in 2010, Gilbert and Parker proposed a codon substitution model that can be used extensively to study the origin of fungal diseases; specifically, that are associated with crops (<xref ref-type="bibr" rid="B59">Gilbert and Parker, 2010</xref>). When any fungi are exposed to a novel environment in a new host, they evolve very fast. Using these new codon models, we can predict pesticide targets on the basis of the nature of selection acting on crucial genes. These models can also be employed for investigating the novel function of regulatory genes as well as networks and important pathways associated with pathogenesis (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Recently, several other studies have also proposed a new hypothesis in the context of intracellular pathogens (<xref ref-type="bibr" rid="B24">Casadevall, 2008</xref>). As per that hypothesis, fungi become intracellular pathogens <italic>via</italic> dual-use traits evolution. For instance, genes originally associated with escaping amoeba predation consequently became advantageous and helped in invading animal or plant cells (e.g. adhesins, toxins, efflux pumps, and injectors, among others) (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Codon models can also be employed for tracing selective pressure acting on dual traits under diverse circumstances (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>).</p>
<p>Some researchers have also proposed that functional divergence of proteins subsequently after some events, for instance, gene duplication, may also result in complex sequence evolution, which is poorly described <italic>via</italic> presently available &#x201c;branch-site&#x201d; codon models (<xref ref-type="bibr" rid="B10">Anisimova and Liberles, 2007</xref>; <xref ref-type="bibr" rid="B18">Benner, 2012</xref>). On the contrary, recently developed clade models, Clade model C (CmC) &#x26; Clade model D (CmD) (present in the CODEML utility of the PAML software package (<xref ref-type="bibr" rid="B182">Yang, 2007</xref>), are a collection of flexible &#x201c;codon-substitution&#x201d; models comprised of both &#x201c;among-lineage&#x201d; as well as &#x201c;among-site&#x201d; variation in selective pressure, which in turn can be an effective tool for investigating signatures of functional divergence amongst clades (<xref ref-type="bibr" rid="B19">Bielawski and Yang, 2004</xref>). To date, the clade models have been utilized for studying functional divergence in numerous gene families, e.g., &#x3b2;-globins (<xref ref-type="bibr" rid="B4">Aguileta et al., 2004</xref>) and vertebrate Troponin C (<xref ref-type="bibr" rid="B19">Bielawski and Yang, 2004</xref>).</p>
<p>When augmented with EB site assignment methods, these clade models may also provide an opportunity to unmask the molecular bases of functional diversification, as well as help in understanding biochemical analyses of homologous yet functionally divergent proteins (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). However, these clade models are still in the infancy phase and further research is required to establish actual power as well as accuracy while dealing with complex forms of divergence among clades (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). Nevertheless, one most important limitations of the present clade models is the absence of incorporation of &#x201c;among-site rate variation&#x201d; within <italic>&#x3c9;</italic>. At present, both CmC and CmD presume only one site class for which <italic>&#x3c9;</italic> either decreases or increases (but not both). But in reality, a large number of complex divergence scenarios are possible. For example, a few sites present within the divergent clade may switch to neutral from purifying class, while others may switch in the opposite direction (<xref ref-type="bibr" rid="B18">Benner, 2012</xref>). If such a scenario exists, novel approaches for detecting might be necessary as like the &#x2018;switching&#x2019; codon models developed <italic>via</italic> Guindon and the team (<xref ref-type="bibr" rid="B67">Guindon et al., 2004</xref>).</p>
<p>Thus, by augmenting new parameters to existing codon substitution models or by designing novel algorithms, we can develop more robust and less computationally demanding codon substitution models for more accurate phylogeny as well as understanding the evolutionary history of any sequences or organisms.</p>
</sec>
<sec sec-type="conclusion" id="s7">
<title>Conclusion</title>
<p>wing to the presence of the huge amount of genomic sequences due to recent advancements in technology, it is easy to understand the evolutionary history of any sequences or organisms in a far better way. Phylogenetic analysis utilizing nucleotide/amino acid/codon substitution models are the most powerful tool for unraveling the evolutionary history of genomic sequences/organisms. However, in comparison with nucleotide and amino acid models, the codon substitution model is more powerful. These models have been utilized extensively to detect selective pressure on a protein, codon usage bias, ancestral reconstruction and phylogenetic reconstruction. However, most of the codon substitution models are still in their infancy stage and deserve further attention. On the downside, the presence of a large variety of models and each considering different biological factors, enhances the margin for misinterpretation. The biological meaning of certain parameters may differ amongst models and thus, model selection procedures also deserve greater attention. Additionally, due to more computational demanding, in comparison to nucleotide and amino acid substitution matrices, only a few studies have employed the codon substitution model to understand the heterogeneity of the evolutionary process in genome-scale analyses. Thus, there is still scope for developing more robust and less computationally demanding codon models. Authors believe that a more robust codon substitution model can be developed considering parameters like the size and structure of the population across time and uncertainty in the ancestral state during estimation. Additionally, results obtained from these models, when combined with other multidisciplinary approaches, like epidemiology, physiology, and molecular biology, are most likely to detect selective pressure on a protein, codon usage bias, ancestral reconstruction and phylogenetic reconstruction in a more comprehensive way. Thus, it seems clear that, in the near future, research on substitution models requires the design and development of more sophisticated as well as realistic substitution models. For instance, the development of codon models with more relaxing assumptions like temporal heterogeneity in both mutational as well as selective processes. Additional effort is also being required to evaluate, compare and apply these newly developed models with real large datasets. As the codon substitution model enables to detect selection regime under which any gene or gene region is evolving, codon usage bias in any organisms or tissue-specific region and phylogenetic relationship between different lineages more accurately than nucleotide and amino acid substitution models, in the near future, these codon models can be utilized in the field of conservation, breeding and medicine.</p>
</sec>
</body>
<back>
<sec id="s8">
<title>Author contributions</title>
<p>MG and RV conceived and designed the study. MG wrote the manuscript under the guidance of RV. Both authors have read, edited, commented on and approved the final manuscript.</p>
</sec>
<ack>
<p>Authors thank reviewers for providing their insightful comments and suggestions. Authors would also like to thank Jan Benzenberg, Incident Manager, Carrier Radio Access Networks (Germany), and MakerSpace Bonn e.V (Germany), for providing computational facility for the literature survey.</p>
</ack>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abascal</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Posada</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zardoya</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>MtArt: A new model of amino acid replacement for arthropoda</article-title>. <source>Mol. Biol. Evol.</source> <volume>24</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msl136</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adachi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hasegawa</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Model of amino acid substitution in proteins encoded by mitochondrial DNA</article-title>. <source>J. Mol. Evol.</source> <volume>42</volume>, <fpage>459</fpage>&#x2013;<lpage>468</lpage>. <pub-id pub-id-type="doi">10.1007/BF02498640</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adachi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Waddell</surname>
<given-names>P. J.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Hasegawa</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Plastid genome phylogeny and a model of amino acid substitution for proteins encoded by chloroplast DNA</article-title>. <source>J. Mol. Evol.</source> <volume>50</volume>, <fpage>348</fpage>&#x2013;<lpage>358</lpage>. <pub-id pub-id-type="doi">10.1007/s002399910038</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aguileta</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bielawski</surname>
<given-names>J. P.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Gene conversion and functional divergence in the beta-globin gene family</article-title>. <source>J. Mol. Evol.</source> <volume>59</volume>, <fpage>177</fpage>&#x2013;<lpage>189</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-004-2612-0</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akashi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Eyre-Walker</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Translational selection and molecular evolution</article-title>. <source>Curr. Opin. Genet. Dev.</source> <volume>8</volume>, <fpage>688</fpage>&#x2013;<lpage>693</lpage>. <pub-id pub-id-type="doi">10.1016/S0959-437X(98)80038-5</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akashi</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Inferring weak selection from patterns of polymorphism and divergence at" silent" sites in Drosophila DNA</article-title>. <source>Genetics</source> <volume>139</volume>, <fpage>1067</fpage>&#x2013;<lpage>1076</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/139.2.1067</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akashi</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Molecular evolution between <italic>Drosophila melanogaster</italic> and <italic>D. simulans</italic> reduced codon bias, faster rates of amino acid substitution, and larger proteins in <italic>D. melanogaster</italic>
</article-title>. <source>Genetics</source> <volume>144</volume>, <fpage>1297</fpage>&#x2013;<lpage>1307</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/144.3.1297</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akashi</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Inferring the fitness effects of DNA mutations from polymorphism and divergence data: Statistical power to detect directional selection under stationarity and free recombination</article-title>. <source>Genetics</source> <volume>151</volume>, <fpage>221</fpage>&#x2013;<lpage>238</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/151.1.221</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Anisimova</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kosiol</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Investigating protein-coding sequence evolution with probabilistic codon substitution models</article-title>. <source>Mol. Biol. Evol.</source> <volume>26</volume>, <fpage>255</fpage>&#x2013;<lpage>271</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msn232</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Anisimova</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Liberles</surname>
<given-names>D. A.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>The quest for natural selection in the age of comparative genomics</article-title>. <source>Heredity</source> <volume>99</volume>, <fpage>567</fpage>&#x2013;<lpage>579</lpage>. <pub-id pub-id-type="doi">10.1038/sj.hdy.6801052</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Anisimova</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Effect of recombination on the accuracy of the likelihood method for detecting positive selection at amino acid sites</article-title>. <source>Genetics</source> <volume>164</volume>, <fpage>1229</fpage>&#x2013;<lpage>1236</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/164.3.1229</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arenas</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Posada</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Simulation of genome-wide evolution under heterogeneous substitution models and complex multispecies coalescent histories</article-title>. <source>Mol. Biol. Evol.</source> <volume>31</volume>, <fpage>1295</fpage>&#x2013;<lpage>1301</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msu078</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arenas</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Simulation of molecular data under diverse evolutionary scenarios</article-title>. <source>PLOS Comput. Biol.</source> <volume>8</volume>, <fpage>e1002495</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002495</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arenas</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015a</year>). <article-title>Advances in computer simulation of genome evolution: Toward more realistic evolutionary genomics analysis by approximate bayesian computation</article-title>. <source>J. Mol. Evol.</source> <volume>80</volume>, <fpage>189</fpage>&#x2013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-015-9673-0</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arenas</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015b</year>). <article-title>Trends in substitution models of molecular evolution</article-title>. <source>Front. Genet.</source> <volume>6</volume>, <fpage>319</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2015.00319</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baele</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Suchard</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Bielejec</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Lemey</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Bayesian codon substitution modelling to identify sources of pathogen evolutionary rate variation</article-title>. <source>Microb. Genomics</source> <volume>2</volume>, <fpage>e000057</fpage>. <pub-id pub-id-type="doi">10.1099/mgen.0.000057</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beaulieu</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>O&#x2019;Meara</surname>
<given-names>B. C.</given-names>
</name>
<name>
<surname>Zaretzki</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Landerer</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gilchrist</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Population genetics based phylogenetics under stabilizing selection for an optimal amino acid sequence: A nested modeling approach</article-title>. <source>Mol. Biol. Evol.</source> <volume>36</volume>, <fpage>834</fpage>&#x2013;<lpage>851</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msy222</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Benner</surname>
<given-names>S. A.</given-names>
</name>
</person-group> (<year>2012</year>). <source>Use of codon models in molecular dating and functional analysis</source>. <publisher-name>Oxford University Press</publisher-name>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.oxfordscholarship.com/view/10.1093/acprof:osobl/9780199601165.001.0001/acprof-9780199601165-chapter-10">https://www.oxfordscholarship.com/view/10.1093/acprof:osobl/9780199601165.001.0001/acprof-9780199601165-chapter-10</ext-link> (Accessed May 25, 2019)</comment>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bielawski</surname>
<given-names>J. P.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>A maximum likelihood method for detecting functional divergence at individual codon sites, with application to gene family evolution</article-title>. <source>J. Mol. Evol.</source> <volume>59</volume>, <fpage>121</fpage>&#x2013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-004-2597-8</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Blanchette</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>E. D.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Haussler</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Reconstructing large regions of an ancestral mammalian genome <italic>in silico</italic>
</article-title>. <source>Genome Res.</source> <volume>14</volume>, <fpage>2412</fpage>&#x2013;<lpage>2423</lpage>. <pub-id pub-id-type="doi">10.1101/gr.2800104</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bloom</surname>
<given-names>J. D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>An experimentally determined evolutionary model dramatically improves phylogenetic fit</article-title>. <source>Mol. Biol. Evol.</source> <volume>31</volume>, <fpage>1956</fpage>&#x2013;<lpage>1978</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msu173</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boussau</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gouy</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Efficient likelihood computations with nonreversible models of evolution</article-title>. <source>Syst. Biol.</source> <volume>55</volume>, <fpage>756</fpage>&#x2013;<lpage>768</lpage>. <pub-id pub-id-type="doi">10.1080/10635150600975218</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Cannarozzi</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <source>Codon evolution: mechanisms and models</source>. <publisher-loc>Oxford, United Kingdom</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Casadevall</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Evolution of intracellular pathogens</article-title>. <source>Annu. Rev. Microbiol.</source> <volume>62</volume>, <fpage>19</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.micro.61.080706.093305</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chakraborty</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nag</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Mazumder</surname>
<given-names>T. H.</given-names>
</name>
<name>
<surname>Uddin</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Codon usage pattern and prediction of gene expression level in Bungarus species</article-title>. <source>Gene</source> <volume>604</volume>, <fpage>48</fpage>&#x2013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1016/j.gene.2016.11.023</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chang</surname>
<given-names>B. S. W.</given-names>
</name>
<name>
<surname>J&#xf6;nsson</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kazmi</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Donoghue</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Sakmar</surname>
<given-names>T. P.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Recreating a functional ancestral archosaur visual pigment</article-title>. <source>Mol. Biol. Evol.</source> <volume>19</volume>, <fpage>1483</fpage>&#x2013;<lpage>1489</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a004211</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Distinguishing HIV-1 drug resistance, accessory, and viral fitness mutations using conditional selection pressure analysis of treated versus untreated patient samples</article-title>. <source>Biol. Direct</source> <volume>1</volume>, <fpage>14</fpage>. <pub-id pub-id-type="doi">10.1186/1745-6150-1-14</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Choudhuri</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Chapter 9 - phylogenetic Analysis&#x2a;&#x2a;The opinions expressed in this chapter are the author&#x2019;s own and they do not necessarily reflect the opinions of the FDA, the DHHS, or the Federal Government</article-title>,&#x201d; in <source>Bioinformatics for beginners</source>. Editor <person-group person-group-type="editor">
<name>
<surname>Choudhuri</surname>
<given-names>S.</given-names>
</name>
</person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>209</fpage>&#x2013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-410471-6.00009-8</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Koeken</surname>
<given-names>V. A. C. M.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multi-omics approaches in immunological research</article-title>. <source>Front. Immunol.</source> <volume>12</volume>, <fpage>668045</fpage>. <pub-id pub-id-type="doi">10.3389/fimmu.2021.668045</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cole</surname>
<given-names>M. F.</given-names>
</name>
<name>
<surname>Gaucher</surname>
<given-names>E. A.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Utilizing natural diversity to evolve protein function: applications towards thermostability</article-title>. <source>Curr. Opin. Chem. Biol.</source> <volume>15</volume>, <fpage>399</fpage>&#x2013;<lpage>406</lpage>. <pub-id pub-id-type="doi">10.1016/j.cbpa.2011.03.005</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Conant</surname>
<given-names>G. C.</given-names>
</name>
<name>
<surname>Stadler</surname>
<given-names>P. F.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Solvent exposure imparts similar selective pressures across a range of yeast proteins</article-title>. <source>Mol. Biol. Evol.</source> <volume>26</volume>, <fpage>1155</fpage>&#x2013;<lpage>1161</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msp031</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dang</surname>
<given-names>C. C.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>Q. S.</given-names>
</name>
<name>
<surname>Gascuel</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>V. S.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>FLU, an amino acid substitution model for influenza proteins</article-title>. <source>BMC Evol. Biol.</source> <volume>10</volume>, <fpage>99</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-10-99</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Daubin</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ochman</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Bacterial genomes as new gene homes: the genealogy of ORFans in <italic>E. coli</italic>
</article-title>. <source>Genome Res.</source> <volume>14</volume>, <fpage>1036</fpage>&#x2013;<lpage>1042</lpage>. <pub-id pub-id-type="doi">10.1101/gr.2231904</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davydov</surname>
<given-names>I. I.</given-names>
</name>
<name>
<surname>Salamin</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Robinson-Rechavi</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Large-scale comparative analysis of codon models accounting for protein and nucleotide selection</article-title>. <source>Mol. Biol. Evol.</source> <volume>36</volume>, <fpage>1316</fpage>&#x2013;<lpage>1332</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msz048</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dayhoff</surname>
<given-names>M. O.</given-names>
</name>
<name>
<surname>Schwartz</surname>
<given-names>R. M.</given-names>
</name>
<name>
<surname>Orcutt</surname>
<given-names>B. C.</given-names>
</name>
</person-group> (<year>1978</year>). <article-title>22 a model of evolutionary change in proteins</article-title>. <source>Atlas Protein Seq. Struct.</source> <volume>5</volume>, <fpage>345</fpage>&#x2013;<lpage>352</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Koning</surname>
<given-names>A. P. J.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Pollock</surname>
<given-names>D. D.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Rapid likelihood analysis on large phylogenies using partial sampling of substitution histories</article-title>. <source>Mol. Biol. Evol.</source> <volume>27</volume>, <fpage>249</fpage>&#x2013;<lpage>265</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msp228</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>De Maio</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Holmes</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Schl&#xf6;tterer</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kosiol</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Estimating empirical codon hidden Markov models</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>725</fpage>&#x2013;<lpage>736</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mss266</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Delport</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Scheffler</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Seoighe</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Models of coding sequence evolution</article-title>. <source>Brief. Bioinform.</source> <volume>10</volume>, <fpage>97</fpage>&#x2013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbn049</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Delport</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Scheffler</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Gravenor</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Muse</surname>
<given-names>S. V.</given-names>
</name>
<name>
<surname>Pond</surname>
<given-names>S. K.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Benchmarking multi-rate codon models</article-title>. <source>PLOS ONE</source> <volume>5</volume>, <fpage>e11587</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0011587</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dimmic</surname>
<given-names>M. W.</given-names>
</name>
<name>
<surname>Rest</surname>
<given-names>J. S.</given-names>
</name>
<name>
<surname>Mindell</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Goldstein</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>rtREV: An amino acid substitution matrix for inference of retrovirus and reverse transcriptase phylogeny</article-title>. <source>J. Mol. Evol.</source> <volume>55</volume>, <fpage>65</fpage>&#x2013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-001-2304-y</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Domazet-Loso</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tautz</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>An evolutionary analysis of orphan genes in Drosophila</article-title>. <source>Genome Res.</source> <volume>13</volume>, <fpage>2213</fpage>&#x2013;<lpage>2219</lpage>. <pub-id pub-id-type="doi">10.1101/gr.1311003</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doron-Faigenboim</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pupko</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>A combined empirical and mechanistic codon model</article-title>. <source>Mol. Biol. Evol.</source> <volume>24</volume>, <fpage>388</fpage>&#x2013;<lpage>397</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msl175</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>dos Reis</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>How to calculate the non-synonymous to synonymous rate ratio of protein-coding genes under the Fisher&#x2013;Wright mutation&#x2013;selection framework</article-title>. <source>Biol. Lett.</source> <volume>11</volume>, <fpage>20141031</fpage>. <pub-id pub-id-type="doi">10.1098/rsbl.2014.1031</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dungan</surname>
<given-names>S. Z.</given-names>
</name>
<name>
<surname>Sabouhanian</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>B. S.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Selection on synonymous codons in mammalian rhodopsins: a possible role in optimizing translational processes</article-title>. <source>BMC Evol. Biol.</source> <volume>14</volume>, <fpage>96</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-14-96</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Investigating molecular evolution of rhodopsin using likelihood/bayesian phylogenetic methods</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://tspace.library.utoronto.ca/handle/1807/24561">https://tspace.library.utoronto.ca/handle/1807/24561</ext-link> (Accessed December 23, 2019)</comment>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dufresne</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Jeffery</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A guided tour of large genome size in animals: what we know and where we are heading</article-title>. <source>Chromosome Res. Int. J. Mol. Supramol. Evol. Asp. Chromosome Biol.</source> <volume>19</volume>, <fpage>925</fpage>&#x2013;<lpage>938</lpage>. <pub-id pub-id-type="doi">10.1007/s10577-011-9248-x</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dunn</surname>
<given-names>K. A.</given-names>
</name>
<name>
<surname>Kenney</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Bielawski</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Improved inference of site-specific positive selection under a generalized parametric codon model when there are multinucleotide mutations and multiple nonsynonymous rates</article-title>. <source>BMC Evol. Biol.</source> <volume>19</volume>, <fpage>22</fpage>. <pub-id pub-id-type="doi">10.1186/s12862-018-1326-7</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dutheil</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Galtier</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Romiguier</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Douzery</surname>
<given-names>E. J. P.</given-names>
</name>
<name>
<surname>Ranwez</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Boussau</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Efficient selection of branch-specific models of sequence evolution</article-title>. <source>Mol. Biol. Evol.</source> <volume>29</volume>, <fpage>1861</fpage>&#x2013;<lpage>1874</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mss059</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eanes</surname>
<given-names>W. F.</given-names>
</name>
<name>
<surname>Kirchner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yoon</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Biermann</surname>
<given-names>C. H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>McCartney</surname>
<given-names>M. A.</given-names>
</name>
<etal/>
</person-group> (<year>1996</year>). <article-title>Historical selection, amino acid polymorphism and lineage-specific divergence at the G6pd locus in <italic>Drosophila melanogaster</italic> and <italic>D. simulans</italic>
</article-title>. <source>Genetics</source> <volume>144</volume>, <fpage>1027</fpage>&#x2013;<lpage>1041</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/144.3.1027</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Edwards</surname>
<given-names>A. W. F.</given-names>
</name>
</person-group> (<year>1972</year>). <source>Likelihood</source>. <publisher-loc>Cambridge, United Kingdom</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B51">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Felsenstein</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Felenstein</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2004</year>). <source>Inferring phylogenies</source>. <publisher-loc>MA</publisher-loc>: <publisher-name>Sinauer associates Sunderland</publisher-name>.</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Felsenstein</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1981</year>). <article-title>Evolutionary trees from DNA sequences: A maximum likelihood approach</article-title>. <source>J. Mol. Evol.</source> <volume>17</volume>, <fpage>368</fpage>&#x2013;<lpage>376</lpage>. <pub-id pub-id-type="doi">10.1007/BF01734359</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fisher</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>1925</year>). <article-title>Theory of statistical estimation</article-title>. <source>Math. Proc. Camb. Philos. Soc.</source> <volume>22</volume>, <fpage>700</fpage>&#x2013;<lpage>725</lpage>. <pub-id pub-id-type="doi">10.1017/S0305004100009580</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fitch</surname>
<given-names>W. M.</given-names>
</name>
<name>
<surname>Bush</surname>
<given-names>R. M.</given-names>
</name>
<name>
<surname>Bender</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Cox</surname>
<given-names>N. J.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Long term trends in the evolution of H (3) HA1 human influenza type A</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>94</volume>, <fpage>7712</fpage>&#x2013;<lpage>7718</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.94.15.7712</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fletcher</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>INDELible: A flexible simulator of biological sequence evolution</article-title>. <source>Mol. Biol. Evol.</source> <volume>26</volume>, <fpage>1879</fpage>&#x2013;<lpage>1888</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msp098</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gaschen</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yusim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Foley</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Lang</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2002</year>). <article-title>Diversity considerations in HIV-1 vaccine selection</article-title>. <source>Science</source> <volume>296</volume>, <fpage>2354</fpage>&#x2013;<lpage>2360</lpage>. <pub-id pub-id-type="doi">10.1126/science.1070441</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gatto</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Catanzaro</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Milinkovitch</surname>
<given-names>M. C.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Assessing the applicability of the GTR nucleotide substitution model through simulations</article-title>. <source>Evol. Bioinforma. Online</source> <volume>2</volume>, <fpage>117693430600200</fpage>&#x2013;<lpage>155</lpage>. <pub-id pub-id-type="doi">10.1177/117693430600200020</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gil</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zanetti</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Zoller</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Anisimova</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>CodonPhyML: Fast maximum likelihood phylogeny estimation under codon substitution models</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>1270</fpage>&#x2013;<lpage>1280</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mst034</pub-id>
</citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gilbert</surname>
<given-names>G. S.</given-names>
</name>
<name>
<surname>Parker</surname>
<given-names>I. M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Rapid evolution in a plant-pathogen interaction and the consequences for introduced host species</article-title>. <source>Evol. Appl.</source> <volume>3</volume>, <fpage>144</fpage>&#x2013;<lpage>156</lpage>. <pub-id pub-id-type="doi">10.1111/j.1752-4571.2009.00107.x</pub-id>
</citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>A codon-based model of nucleotide substitution for protein-coding DNA sequences</article-title>. <source>Mol. Biol. Evol.</source> <volume>11</volume>, <fpage>725</fpage>&#x2013;<lpage>736</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a040153</pub-id>
</citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gonnet</surname>
<given-names>G. H.</given-names>
</name>
<name>
<surname>Cohen</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Benner</surname>
<given-names>S. A.</given-names>
</name>
</person-group> (<year>1992</year>). <article-title>Exhaustive matching of the entire protein sequence database</article-title>. <source>Science</source> <volume>256</volume>, <fpage>1443</fpage>&#x2013;<lpage>1445</lpage>. <pub-id pub-id-type="doi">10.1126/science.1604319</pub-id>
</citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Parida</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mohapatra</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Characterization of haplotypes and single nucleotide polymorphisms associated with Gn1a for high grain number formation in rice plant</article-title>. <source>Genomics</source> <volume>112</volume>, <fpage>2647</fpage>&#x2013;<lpage>2657</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2020.02.016</pub-id>
</citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grahnen</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Nandakumar</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Kubelka</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liberles</surname>
<given-names>D. A.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Biophysical and structural considerations for protein sequence evolution</article-title>. <source>BMC Evol. Biol.</source> <volume>11</volume>, <fpage>361</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2148-11-361</pub-id>
</citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grantham</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>1974</year>). <article-title>Amino acid difference formula to help explain protein evolution</article-title>. <source>Science</source> <volume>185</volume>, <fpage>862</fpage>&#x2013;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.1126/science.185.4154.862</pub-id>
</citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grunspan</surname>
<given-names>D. Z.</given-names>
</name>
<name>
<surname>Nesse</surname>
<given-names>R. M.</given-names>
</name>
<name>
<surname>Barnes</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Brownell</surname>
<given-names>S. E.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Core principles of evolutionary medicine: A delphi study</article-title>. <source>Evol. Med. Public Health</source> <volume>2018</volume>, <fpage>13</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1093/emph/eox025</pub-id>
</citation>
</ref>
<ref id="B66">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gudivada</surname>
<given-names>V. N.</given-names>
</name>
<name>
<surname>Rao</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Raghavan</surname>
<given-names>V. V.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Chapter 9 - big data driven natural language processing research and applications</article-title>,&#x201d; in <source>Handbook of statistics</source>. <source>Big data analytics</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Govindaraju</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Raghavan</surname>
<given-names>V. V.</given-names>
</name>
<name>
<surname>Rao</surname>
<given-names>C. R.</given-names>
</name>
</person-group> (<publisher-name>Elsevier</publisher-name>), <fpage>203</fpage>&#x2013;<lpage>238</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-444-63492-4.00009-5</pub-id>
</citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guindon</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Rodrigo</surname>
<given-names>A. G.</given-names>
</name>
<name>
<surname>Dyer</surname>
<given-names>K. A.</given-names>
</name>
<name>
<surname>Huelsenbeck</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Modeling the site-specific variation of selection patterns along lineages</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>101</volume>, <fpage>12957</fpage>&#x2013;<lpage>12962</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0402177101</pub-id>
</citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2019a</year>). <article-title>Genetic basis of adaptation and maladaptation via balancing selection</article-title>. <source>Zoology</source> <volume>136</volume>, <fpage>125693</fpage>. <pub-id pub-id-type="doi">10.1016/j.zool.2019.125693</pub-id>
</citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2019b</year>). <article-title>Identification and characterization of differentially expressed genes in type 2 diabetes using <italic>in silico</italic> approach</article-title>. <source>Comput. Biol. Chem.</source> <volume>79</volume>, <fpage>24</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiolchem.2019.01.010</pub-id>
</citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Divergent evolution and purifying selection of the type 2 diabetes gene sequences in Drosophila: a phylogenomic study</article-title>. <source>Genetica</source> <volume>148</volume>, <fpage>269</fpage>&#x2013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1007/s10709-020-00101-7</pub-id>
</citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Behara</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>
<italic>In silico</italic> analysis of differential gene expressions in biliary stricture and hepatic carcinoma</article-title>. <source>Gene</source> <volume>597</volume>, <fpage>49</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1016/j.gene.2016.10.032</pub-id>
</citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>De novo assembly and characterization of transcriptome towards understanding molecular mechanism associated with MYMIV-resistance in Vigna mungo-A computational study</article-title>. <source>BioRxiv</source>, <fpage>844639</fpage>. <pub-id pub-id-type="doi">10.1101/844639</pub-id>
</citation>
</ref>
<ref id="B73">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sabarinathan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dash</surname>
<given-names>G. K.</given-names>
</name>
<name>
<surname>Rajesh</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2021a</year>). &#x201c;<article-title>3000 genome project: A brief insight</article-title>,&#x201d; in <source>Bioinformatics in rice research: Theories and techniques</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
</person-group> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>89</fpage>&#x2013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-16-3993-7_5</pub-id>
</citation>
</ref>
<ref id="B74">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sabarinathan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Dash</surname>
<given-names>G. K.</given-names>
</name>
<name>
<surname>Ponnana</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2021b</year>). &#x201c;<article-title>Brief insight into the evolutionary history and domestication of wild rice relatives</article-title>,&#x201d; in <source>Bioinformatics in rice research: Theories and techniques</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
</person-group> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>71</fpage>&#x2013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-16-3993-7_4</pub-id>
</citation>
</ref>
<ref id="B75">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sabarinathan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rajesh</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Pati</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2021c</year>). &#x201c;<article-title>Phylogenetic analysis</article-title>,&#x201d; in <source>Bioinformatics in rice research: Theories and techniques</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
</person-group> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>179</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-16-3993-7_9</pub-id>
</citation>
</ref>
<ref id="B76">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sabarinathan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021d</year>). &#x201c;<article-title>Mapping algorithms in high-throughput sequencing</article-title>,&#x201d; in <source>Bioinformatics in rice research: Theories and techniques</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
</person-group> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>305</fpage>&#x2013;<lpage>323</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-16-3993-7_14</pub-id>
</citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Vemula</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Donde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gouda</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Behera</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Vadde</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021e</year>). <article-title>
<italic>In-silico</italic> approaches to detect inhibitors of the human severe acute respiratory syndrome coronavirus envelope protein ion channel</article-title>. <source>J. Biomol. Struct. Dyn.</source> <volume>39</volume>, <fpage>2617</fpage>&#x2013;<lpage>2627</lpage>. <pub-id pub-id-type="doi">10.1080/07391102.2020.1751300</pub-id>
</citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Halpern</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Bruno</surname>
<given-names>W. J.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Evolutionary distances for protein-coding sequences: modeling site-specific residue frequencies</article-title>. <source>Mol. Biol. Evol.</source> <volume>15</volume>, <fpage>910</fpage>&#x2013;<lpage>917</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a025995</pub-id>
</citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harris</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Error-prone polymerase activity causes multinucleotide mutations in humans</article-title>. <source>Genome Res.</source> <volume>24</volume>, <fpage>1445</fpage>&#x2013;<lpage>1454</lpage>. <pub-id pub-id-type="doi">10.1101/gr.170696.113</pub-id>
</citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hasegawa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kishino</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yano</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>Dating of the human-ape splitting by a molecular clock of mitochondrial DNA</article-title>. <source>J. Mol. Evol.</source> <volume>22</volume>, <fpage>160</fpage>&#x2013;<lpage>174</lpage>. <pub-id pub-id-type="doi">10.1007/BF02101694</pub-id>
</citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haubold</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Alignment-free phylogenetics and population genetics</article-title>. <source>Brief. Bioinform.</source> <volume>15</volume>, <fpage>407</fpage>&#x2013;<lpage>418</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbt083</pub-id>
</citation>
</ref>
<ref id="B82">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Henikoff</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Henikoff</surname>
<given-names>J. G.</given-names>
</name>
</person-group> (<year>1992</year>). <article-title>Amino acid substitution matrices from protein blocks</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>89</volume>, <fpage>10915</fpage>&#x2013;<lpage>10919</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.89.22.10915</pub-id>
</citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hiraoka</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kawamata</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Haraguchi</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chikashige</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Codon usage bias is correlated with gene expression levels in the fission yeast <italic>Schizosaccharomyces pombe</italic>
</article-title>. <source>Genes Cells Devoted Mol. Cell. Mech.</source> <volume>14</volume>, <fpage>499</fpage>&#x2013;<lpage>509</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2443.2009.01284.x</pub-id>
</citation>
</ref>
<ref id="B84">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hoehn</surname>
<given-names>K. B.</given-names>
</name>
<name>
<surname>Lunter</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Pybus</surname>
<given-names>O. G.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A phylogenetic codon substitution model for antibody lineages</article-title>. <source>Genetics</source> <volume>206</volume>, <fpage>417</fpage>&#x2013;<lpage>427</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.116.196303</pub-id>
</citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Holmes</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Rubin</surname>
<given-names>G. M.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>An expectation maximization algorithm for training hidden substitution models</article-title>. <source>J. Mol. Biol.</source> <volume>317</volume>, <fpage>753</fpage>&#x2013;<lpage>764</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.2002.5405</pub-id>
</citation>
</ref>
<ref id="B86">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Holt</surname>
<given-names>K. E.</given-names>
</name>
<name>
<surname>Parkhill</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mazzoni</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Roumagnac</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Weill</surname>
<given-names>F.-X.</given-names>
</name>
<name>
<surname>Goodhead</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2008</year>). <article-title>High-throughput sequencing provides insights into genome variation and evolution in Salmonella Typhi</article-title>. <source>Nat. Genet.</source> <volume>40</volume>, <fpage>987</fpage>&#x2013;<lpage>993</lpage>. <pub-id pub-id-type="doi">10.1038/ng.195</pub-id>
</citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Koonin</surname>
<given-names>E. V.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Przytycka</surname>
<given-names>T. M.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Selection for minimization of translational frameshifting errors as a factor in the evolution of codon usage</article-title>. <source>Nucleic Acids Res.</source> <volume>37</volume>, <fpage>6799</fpage>&#x2013;<lpage>6810</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkp712</pub-id>
</citation>
</ref>
<ref id="B88">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huelsenbeck</surname>
<given-names>J. P.</given-names>
</name>
<name>
<surname>Ronquist</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>MRBAYES: Bayesian inference of phylogenetic trees</article-title>. <source>Bioinforma. Oxf. Engl.</source> <volume>17</volume>, <fpage>754</fpage>&#x2013;<lpage>755</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/17.8.754</pub-id>
</citation>
</ref>
<ref id="B89">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ikemura</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>Codon usage and tRNA content in unicellular and multicellular organisms</article-title>. <source>Mol. Biol. Evol.</source> <volume>2</volume>, <fpage>13</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a040335</pub-id>
</citation>
</ref>
<ref id="B90">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jayaswal</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Jermiin</surname>
<given-names>L. S.</given-names>
</name>
<name>
<surname>Poladian</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Robinson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Two stationary nonhomogeneous Markov models of nucleotide sequence evolution</article-title>. <source>Syst. Biol.</source> <volume>60</volume>, <fpage>74</fpage>&#x2013;<lpage>86</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/syq076</pub-id>
</citation>
</ref>
<ref id="B91">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>D. T.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>W. R.</given-names>
</name>
<name>
<surname>Thornton</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>1992</year>). <article-title>The rapid generation of mutation data matrices from protein sequences</article-title>. <source>Comput. Appl. Biosci.</source> <volume>8</volume>, <fpage>275</fpage>&#x2013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/8.3.275</pub-id>
</citation>
</ref>
<ref id="B92">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>C. T.</given-names>
</name>
<name>
<surname>Youssef</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Susko</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bielawski</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Shifting balance on a static mutation&#x2013;selection landscape: A novel scenario of positive selection</article-title>. <source>Mol. Biol. Evol.</source> <volume>34</volume>, <fpage>391</fpage>&#x2013;<lpage>407</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msw237</pub-id>
</citation>
</ref>
<ref id="B93">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>C. T.</given-names>
</name>
<name>
<surname>Youssef</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Susko</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bielawski</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Phenomenological load on model parameters can lead to false biological conclusions</article-title>. <source>Mol. Biol. Evol.</source> <volume>35</volume>, <fpage>1473</fpage>&#x2013;<lpage>1488</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msy049</pub-id>
</citation>
</ref>
<ref id="B94">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jukes</surname>
<given-names>T. H.</given-names>
</name>
<name>
<surname>Cantor</surname>
<given-names>C. R.</given-names>
</name>
</person-group> (<year>1969</year>). &#x201c;<article-title>CHAPTER 24 - evolution of protein molecules</article-title>,&#x201d; in <source>Mammalian protein metabolism</source>. Editor <person-group person-group-type="editor">
<name>
<surname>Munro</surname>
<given-names>H. N.</given-names>
</name>
</person-group> (<publisher-name>Academic Press</publisher-name>), <fpage>21</fpage>&#x2013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1016/B978-1-4832-3211-9.50009-7</pub-id>
</citation>
</ref>
<ref id="B95">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kimura</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1962</year>). <article-title>On the probability of fixation of mutant genes in a population</article-title>. <source>Genetics</source> <volume>47</volume>, <fpage>713</fpage>&#x2013;<lpage>719</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/47.6.713</pub-id>
</citation>
</ref>
<ref id="B96">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kimura</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1980</year>). <article-title>A simple method for estimating evolutionary rates of base substitutions through comparative studies of nucleotide sequences</article-title>. <source>J. Mol. Evol.</source> <volume>16</volume>, <fpage>111</fpage>&#x2013;<lpage>120</lpage>. <pub-id pub-id-type="doi">10.1007/BF01731581</pub-id>
</citation>
</ref>
<ref id="B97">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kosakovsky Pond</surname>
<given-names>S. L.</given-names>
</name>
<name>
<surname>Posada</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gravenor</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Woelk</surname>
<given-names>C. H.</given-names>
</name>
<name>
<surname>Frost</surname>
<given-names>S. D. W.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>GARD: a genetic algorithm for recombination detection</article-title>. <source>Bioinforma. Oxf. Engl.</source> <volume>22</volume>, <fpage>3096</fpage>&#x2013;<lpage>3098</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl474</pub-id>
</citation>
</ref>
<ref id="B98">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kosakovsky Pond</surname>
<given-names>S. L.</given-names>
</name>
<name>
<surname>Poon</surname>
<given-names>A. F. Y.</given-names>
</name>
<name>
<surname>Leigh Brown</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Frost</surname>
<given-names>S. D. W.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>A maximum likelihood method for detecting directional evolution in protein sequences and its application to influenza A virus</article-title>. <source>Mol. Biol. Evol.</source> <volume>25</volume>, <fpage>1809</fpage>&#x2013;<lpage>1824</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msn123</pub-id>
</citation>
</ref>
<ref id="B99">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kosiol</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Holmes</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>An empirical codon model for protein sequence evolution</article-title>. <source>Mol. Biol. Evol.</source> <volume>24</volume>, <fpage>1464</fpage>&#x2013;<lpage>1479</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msm064</pub-id>
</citation>
</ref>
<ref id="B100">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kryazhimskiy</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Plotkin</surname>
<given-names>J. B.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>The Population Genetics of dN/dS</article-title>. <source>PLoS Genet.</source> <volume>4</volume>, <fpage>e1000304</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1000304</pub-id>
</citation>
</ref>
<ref id="B101">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lacerda</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Scheffler</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Seoighe</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Epitope discovery with phylogenetic hidden Markov models</article-title>. <source>Mol. Biol. Evol.</source> <volume>27</volume>, <fpage>1212</fpage>&#x2013;<lpage>1220</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msq008</pub-id>
</citation>
</ref>
<ref id="B102">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lartillot</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Philippe</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Computing Bayes factors using thermodynamic integration</article-title>. <source>Syst. Biol.</source> <volume>55</volume>, <fpage>195</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1080/10635150500433722</pub-id>
</citation>
</ref>
<ref id="B103">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Le</surname>
<given-names>V. S.</given-names>
</name>
<name>
<surname>Dang</surname>
<given-names>C. C.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>Q. S.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Improved mitochondrial amino acid substitution models for metazoan evolutionary studies</article-title>. <source>BMC Evol. Biol.</source> <volume>17</volume>, <fpage>136</fpage>. <pub-id pub-id-type="doi">10.1186/s12862-017-0987-y</pub-id>
</citation>
</ref>
<ref id="B104">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>The sequence and de novo assembly of the giant panda genome</article-title>. <source>Nature</source> <volume>463</volume>, <fpage>311</fpage>&#x2013;<lpage>317</lpage>. <pub-id pub-id-type="doi">10.1038/nature08696</pub-id>
</citation>
</ref>
<ref id="B105">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liao</surname>
<given-names>H.-X.</given-names>
</name>
<name>
<surname>Lynch</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Alam</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Boyd</surname>
<given-names>S. D.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Co-evolution of a broadly neutralizing HIV-1 antibody and founder virus</article-title>. <source>Nature</source> <volume>496</volume>, <fpage>469</fpage>&#x2013;<lpage>476</lpage>. <pub-id pub-id-type="doi">10.1038/nature12053</pub-id>
</citation>
</ref>
<ref id="B106">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liberles</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Teichmann</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Bahar</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bastolla</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Bloom</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bornberg-Bauer</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>The interface of protein structure, protein biophysics, and molecular evolution</article-title>. <source>Protein Sci.</source> <volume>21</volume>, <fpage>769</fpage>&#x2013;<lpage>785</lpage>. <pub-id pub-id-type="doi">10.1002/pro.2071</pub-id>
</citation>
</ref>
<ref id="B107">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li&#xf2;</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Models of molecular evolution and phylogeny</article-title>. <source>Genome Res.</source> <volume>8</volume>, <fpage>1233</fpage>&#x2013;<lpage>1244</lpage>. <pub-id pub-id-type="doi">10.1101/gr.8.12.1233</pub-id>
</citation>
</ref>
<ref id="B108">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Long</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Langley</surname>
<given-names>C. H.</given-names>
</name>
</person-group> (<year>1993</year>). <article-title>Natural selection and the origin of jingwei, a chimeric processed functional gene in Drosophila</article-title>. <source>Science</source> <volume>260</volume>, <fpage>91</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1126/science.7682012</pub-id>
</citation>
</ref>
<ref id="B109">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lunter</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hein</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>A nucleotide substitution model with nearest-neighbour interactions</article-title>. <source>Bioinforma. Oxf. Engl.</source> <volume>20</volume>, <fpage>i216</fpage>&#x2013;<lpage>223</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bth901</pub-id>
</citation>
</ref>
<ref id="B110">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>MacCallum</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hill</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Being positive about selection</article-title>. <source>PLoS Biol.</source> <volume>4</volume>, <fpage>e87</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pbio.0040087</pub-id>
</citation>
</ref>
<ref id="B111">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mayrose</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Doron-Faigenboim</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bacharach</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Pupko</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Towards realistic codon models: among site variability and dependency of synonymous and non-synonymous rates</article-title>. <source>Bioinforma. Oxf. Engl.</source> <volume>23</volume>, <fpage>i319</fpage>&#x2013;<lpage>327</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm176</pub-id>
</citation>
</ref>
<ref id="B112">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Membrebe</surname>
<given-names>J. V.</given-names>
</name>
<name>
<surname>Suchard</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Rambaut</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Baele</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lemey</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Bayesian inference of evolutionary histories under time-dependent substitution rates</article-title>. <source>Mol. Biol. Evol.</source> <volume>36</volume>, <fpage>1793</fpage>&#x2013;<lpage>1803</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msz094</pub-id>
</citation>
</ref>
<ref id="B113">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Misawa</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A codon substitution model that incorporates the effect of the GC contents, the gene density and the density of CpG islands of human chromosomes</article-title>. <source>BMC Genomics</source> <volume>12</volume>, <fpage>397</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-12-397</pub-id>
</citation>
</ref>
<ref id="B114">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Miyazawa</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011a</year>). <article-title>Advantages of a mechanistic codon substitution model for evolutionary analysis of protein-coding sequences</article-title>. <source>PLOS ONE</source> <volume>6</volume>, <fpage>e28892</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0028892</pub-id>
</citation>
</ref>
<ref id="B115">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Miyazawa</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011b</year>). <article-title>Selective constraints on amino acids estimated by a mechanistic codon substitution model with multiple nucleotide changes</article-title>. <source>PLoS One</source> <volume>6</volume>, <fpage>e17244</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0017244</pub-id>
</citation>
</ref>
<ref id="B116">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moutinho</surname>
<given-names>A. F.</given-names>
</name>
<name>
<surname>Eyre-Walker</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dutheil</surname>
<given-names>J. Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Strong evidence for the adaptive walk model of gene evolution in Drosophila and Arabidopsis</article-title>. <source>PLOS Biol.</source> <volume>20</volume>, <fpage>e3001775</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pbio.3001775</pub-id>
</citation>
</ref>
<ref id="B117">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>M&#xfc;ller</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Vingron</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Modeling amino acid replacement</article-title>. <source>J. Comput. Biol.</source> <volume>7</volume>, <fpage>761</fpage>&#x2013;<lpage>776</lpage>. <pub-id pub-id-type="doi">10.1089/10665270050514918</pub-id>
</citation>
</ref>
<ref id="B118">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Muse</surname>
<given-names>S. V.</given-names>
</name>
<name>
<surname>Gaut</surname>
<given-names>B. S.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>A likelihood approach for comparing synonymous and nonsynonymous nucleotide substitution rates, with application to the chloroplast genome</article-title>. <source>Mol. Biol. Evol.</source> <volume>11</volume>, <fpage>715</fpage>&#x2013;<lpage>724</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a040152</pub-id>
</citation>
</ref>
<ref id="B119">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nickle</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Heath</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jensen</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Gilbert</surname>
<given-names>P. B.</given-names>
</name>
<name>
<surname>Mullins</surname>
<given-names>J. I.</given-names>
</name>
<name>
<surname>Pond</surname>
<given-names>S. L. K.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>HIV-specific probabilistic models of protein evolution</article-title>. <source>PLOS ONE</source> <volume>2</volume>, <fpage>e503</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0000503</pub-id>
</citation>
</ref>
<ref id="B120">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Likelihood models for detecting positively selected amino acid sites and applications to the HIV-1 envelope gene</article-title>. <source>Genetics</source> <volume>148</volume>, <fpage>929</fpage>&#x2013;<lpage>936</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/148.3.929</pub-id>
</citation>
</ref>
<ref id="B121">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Estimating the distribution of selection coefficients from phylogenetic data with applications to mitochondrial and viral DNA</article-title>. <source>Mol. Biol. Evol.</source> <volume>20</volume>, <fpage>1231</fpage>&#x2013;<lpage>1239</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msg147</pub-id>
</citation>
</ref>
<ref id="B122">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bauer DuMont</surname>
<given-names>V. L.</given-names>
</name>
<name>
<surname>Hubisz</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Aquadro</surname>
<given-names>C. F.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Maximum likelihood estimation of ancestral codon usage bias parameters in Drosophila</article-title>. <source>Mol. Biol. Evol.</source> <volume>24</volume>, <fpage>228</fpage>&#x2013;<lpage>235</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msl146</pub-id>
</citation>
</ref>
<ref id="B123">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olejniczak</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Uhlenbeck</surname>
<given-names>O. C.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>tRNA residues that have coevolved with their anticodon to ensure uniform and accurate codon recognition</article-title>. <source>Biochimie</source> <volume>88</volume>, <fpage>943</fpage>&#x2013;<lpage>950</lpage>. <pub-id pub-id-type="doi">10.1016/j.biochi.2006.06.005</pub-id>
</citation>
</ref>
<ref id="B124">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Osada</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Akashi</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Mitochondrial-nuclear interactions and accelerated compensatory evolution: evidence from the primate cytochrome C oxidase complex</article-title>. <source>Mol. Biol. Evol.</source> <volume>29</volume>, <fpage>337</fpage>&#x2013;<lpage>346</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msr211</pub-id>
</citation>
</ref>
<ref id="B125">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ouyang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>Detecting positively selected sites from amino acid sequences: An implicit codon model</article-title>,&#x201d; in <conf-name>2007 29th Annual International Conference of the IEEE Engineering in Medicine and Biology Society</conf-name>, <fpage>5302</fpage>&#x2013;<lpage>5306</lpage>. <pub-id pub-id-type="doi">10.1109/IEMBS.2007.4353538</pub-id>
</citation>
</ref>
<ref id="B126">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Parto</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lartillot</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Molecular adaptation in Rubisco: Discriminating between convergent evolution and positive selection using mechanistic and classical codon models</article-title>. <source>PLOS ONE</source> <volume>13</volume>, <fpage>e0192697</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0192697</pub-id>
</citation>
</ref>
<ref id="B127">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pevsner</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <source>Bioinformatics and functional genomics</source>. <publisher-loc>Sussex, United Kingdom</publisher-loc>: <publisher-name>John Wiley &#x26; Sons</publisher-name>.</citation>
</ref>
<ref id="B128">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pond</surname>
<given-names>S. L. K.</given-names>
</name>
<name>
<surname>Frost</surname>
<given-names>S. D. W.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>A genetic algorithm approach to detecting lineage-specific variation in selection pressure</article-title>. <source>Mol. Biol. Evol.</source> <volume>22</volume>, <fpage>478</fpage>&#x2013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msi031</pub-id>
</citation>
</ref>
<ref id="B129">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pond</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Muse</surname>
<given-names>S. V.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Site-to-Site variation of synonymous substitution rates</article-title>. <source>Mol. Biol. Evol.</source> <volume>22</volume>, <fpage>2375</fpage>&#x2013;<lpage>2385</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msi232</pub-id>
</citation>
</ref>
<ref id="B130">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pond</surname>
<given-names>S. L. K.</given-names>
</name>
<name>
<surname>Frost</surname>
<given-names>S. D. W.</given-names>
</name>
<name>
<surname>Muse</surname>
<given-names>S. V.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>HyPhy: hypothesis testing using phylogenies</article-title>. <source>Bioinforma. Oxf. Engl.</source> <volume>21</volume>, <fpage>676</fpage>&#x2013;<lpage>679</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti079</pub-id>
</citation>
</ref>
<ref id="B131">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pouyet</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bailly-Bechet</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mouchiroud</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gu&#xe9;guen</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>SENCA: A multilayered codon model to study the origins and dynamics of codon usage</article-title>. <source>Genome Biol. Evol.</source> <volume>8</volume>, <fpage>2427</fpage>&#x2013;<lpage>2441</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evw165</pub-id>
</citation>
</ref>
<ref id="B132">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pupko</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Galtier</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>A covarion-based method for detecting molecular adaptation: application to the evolution of primate mitochondrial genomes</article-title>. <source>Proc. Biol. Sci.</source> <volume>269</volume>, <fpage>1313</fpage>&#x2013;<lpage>1316</lpage>. <pub-id pub-id-type="doi">10.1098/rspb.2002.2025</pub-id>
</citation>
</ref>
<ref id="B133">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ren</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Tanaka</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>An empirical examination of the utility of codon-substitution models in phylogeny reconstruction</article-title>. <source>Syst. Biol.</source> <volume>54</volume>, <fpage>808</fpage>&#x2013;<lpage>818</lpage>. <pub-id pub-id-type="doi">10.1080/10635150500354688</pub-id>
</citation>
</ref>
<ref id="B134">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rodrigue</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Lartillot</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Philippe</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Bayesian comparisons of codon substitution models</article-title>. <source>Genetics</source> <volume>180</volume>, <fpage>1579</fpage>&#x2013;<lpage>1591</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.108.092254</pub-id>
</citation>
</ref>
<ref id="B135">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rodrigue</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Philippe</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Lartillot</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Mutation-selection models of coding sequence evolution with site-heterogeneous amino acid fitness profiles</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>107</volume>, <fpage>4629</fpage>&#x2013;<lpage>4634</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0910915107</pub-id>
</citation>
</ref>
<ref id="B136">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Roumagnac</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Weill</surname>
<given-names>F.-X.</given-names>
</name>
<name>
<surname>Dolecek</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Baker</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Brisse</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chinh</surname>
<given-names>N. T.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>Evolutionary history of Salmonella typhi</article-title>. <source>Science</source> <volume>314</volume>, <fpage>1301</fpage>&#x2013;<lpage>1304</lpage>. <pub-id pub-id-type="doi">10.1126/science.1134933</pub-id>
</citation>
</ref>
<ref id="B137">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rubinstein</surname>
<given-names>N. D.</given-names>
</name>
<name>
<surname>Pupko</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Cannarozzi</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Detection and analysis of conservation at synonymous sites</article-title>. <source>Codon Evol. Mech. Models</source> <volume>218</volume>, <fpage>228</fpage>.</citation>
</ref>
<ref id="B138">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sainudiin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>W. S. W.</given-names>
</name>
<name>
<surname>Yogeeswaran</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Nasrallah</surname>
<given-names>J. B.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Detecting site-specific physicochemical selective pressures: applications to the class I HLA of the human major histocompatibility complex and the SRK of the plant sporophytic self-incompatibility system</article-title>. <source>J. Mol. Evol.</source> <volume>60</volume>, <fpage>315</fpage>&#x2013;<lpage>326</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-004-0153-1</pub-id>
</citation>
</ref>
<ref id="B139">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sakofsky</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Roberts</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Malc</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mieczkowski</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Resnick</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Gordenin</surname>
<given-names>D. A.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Break-induced replication is a source of mutation clusters underlying kataegis</article-title>. <source>Cell Rep.</source> <volume>7</volume>, <fpage>1640</fpage>&#x2013;<lpage>1648</lpage>. <pub-id pub-id-type="doi">10.1016/j.celrep.2014.04.053</pub-id>
</citation>
</ref>
<ref id="B140">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sawyer</surname>
<given-names>S. L.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>L. I.</given-names>
</name>
<name>
<surname>Emerman</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Malik</surname>
<given-names>H. S.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Positive selection of primate TRIM5alpha identifies a critical species-specific retroviral restriction domain</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>102</volume>, <fpage>2832</fpage>&#x2013;<lpage>2837</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0409853102</pub-id>
</citation>
</ref>
<ref id="B141">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Scheffler</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Seoighe</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Robust inference of positive selection from recombining coding sequences</article-title>. <source>Bioinforma. Oxf. Engl.</source> <volume>22</volume>, <fpage>2493</fpage>&#x2013;<lpage>2499</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl427</pub-id>
</citation>
</ref>
<ref id="B142">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schneider</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cannarozzi</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Gonnet</surname>
<given-names>G. H.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Empirical codon substitution matrix</article-title>. <source>BMC Bioinforma.</source> <volume>6</volume>, <fpage>134</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-6-134</pub-id>
</citation>
</ref>
<ref id="B143">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sch&#xf6;niger</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hofacker</surname>
<given-names>G. L.</given-names>
</name>
<name>
<surname>Borstnik</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Stochastic traits of molecular evolution&#x2014;acceptance of point mutations in native actin genes</article-title>. <source>J. Theor. Biol.</source> <volume>143</volume>, <fpage>287</fpage>&#x2013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-5193(05)80031-1</pub-id>
</citation>
</ref>
<ref id="B144">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Seo</surname>
<given-names>T.-K.</given-names>
</name>
<name>
<surname>Kishino</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Synonymous substitutions substantially improve evolutionary inference from highly diverged proteins</article-title>. <source>Syst. Biol.</source> <volume>57</volume>, <fpage>367</fpage>&#x2013;<lpage>377</lpage>. <pub-id pub-id-type="doi">10.1080/10635150802158670</pub-id>
</citation>
</ref>
<ref id="B145">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Seo</surname>
<given-names>T.-K.</given-names>
</name>
<name>
<surname>Kishino</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Statistical comparison of nucleotide, amino acid, and codon substitution models for evolutionary analysis of protein-coding sequences</article-title>. <source>Syst. Biol.</source> <volume>58</volume>, <fpage>199</fpage>&#x2013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/syp015</pub-id>
</citation>
</ref>
<ref id="B146">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shapiro</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Rambaut</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Drummond</surname>
<given-names>A. J.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Choosing appropriate substitution models for the phylogenetic analysis of protein-coding sequences</article-title>. <source>Mol. Biol. Evol.</source> <volume>23</volume>, <fpage>7</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msj021</pub-id>
</citation>
</ref>
<ref id="B147">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sharp</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Tuohy</surname>
<given-names>T. M. F.</given-names>
</name>
<name>
<surname>Mosurski</surname>
<given-names>K. R.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>Codon usage in yeast: cluster analysis clearly differentiates highly and lowly expressed genes</article-title>. <source>Nucleic Acids Res.</source> <volume>14</volume>, <fpage>5125</fpage>&#x2013;<lpage>5143</lpage>. <pub-id pub-id-type="doi">10.1093/nar/14.13.5125</pub-id>
</citation>
</ref>
<ref id="B148">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shoemaker</surname>
<given-names>J. S.</given-names>
</name>
<name>
<surname>Fitch</surname>
<given-names>W. M.</given-names>
</name>
</person-group> (<year>1989</year>). <article-title>Evidence from nuclear sequences that invariable sites should be considered when sequence divergence is calculated</article-title>. <source>Mol. Biol. Evol.</source> <volume>6</volume>, <fpage>270</fpage>&#x2013;<lpage>289</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a040550</pub-id>
</citation>
</ref>
<ref id="B149">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Smith</surname>
<given-names>N. G. C.</given-names>
</name>
<name>
<surname>Webster</surname>
<given-names>M. T.</given-names>
</name>
<name>
<surname>Ellegren</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>A low rate of simultaneous double-nucleotide mutations in primates</article-title>. <source>Mol. Biol. Evol.</source> <volume>20</volume>, <fpage>47</fpage>&#x2013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msg003</pub-id>
</citation>
</ref>
<ref id="B150">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sullivan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Joyce</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Model selection in phylogenetics</article-title>. <source>Annu. Rev. Ecol. Evol. Syst.</source> <volume>36</volume>, <fpage>445</fpage>&#x2013;<lpage>466</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.ecolsys.36.102003.152633</pub-id>
</citation>
</ref>
<ref id="B151">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Murphy</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X. S.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>D. W.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Analysis of codon usage on Wolbachia pipientis wMel genome</article-title>. <source>Sci. China C Life Sci.</source> <volume>39</volume>, <fpage>948</fpage>&#x2013;<lpage>953</lpage>.</citation>
</ref>
<ref id="B191">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Suzuki</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>New methods for detecting positive selection at single amino acid sites</article-title>. <source>J. Mol. Evol.</source> <volume>59</volume>, <fpage>11</fpage>&#x2013;<lpage>19</lpage>.</citation>
</ref>
<ref id="B152">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Suzuki</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Gojobori</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>A method for detecting positive selection at single amino acid sites</article-title>. <source>Mol. Biol. Evol.</source> <volume>16</volume>, <fpage>1315</fpage>&#x2013;<lpage>1328</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a026042</pub-id>
</citation>
</ref>
<ref id="B153">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Takano-Shimizu</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Local changes in GC/AT substitution biases and in crossover frequencies on Drosophila chromosomes</article-title>. <source>Mol. Biol. Evol.</source> <volume>18</volume>, <fpage>606</fpage>&#x2013;<lpage>619</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a003841</pub-id>
</citation>
</ref>
<ref id="B154">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tamura</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Nei</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1993</year>). <article-title>Estimation of the number of nucleotide substitutions in the control region of mitochondrial DNA in humans and chimpanzees</article-title>. <source>Mol. Biol. Evol.</source> <volume>10</volume>, <fpage>512</fpage>&#x2013;<lpage>526</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a040023</pub-id>
</citation>
</ref>
<ref id="B155">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tamuri</surname>
<given-names>A. U.</given-names>
</name>
<name>
<surname>dos Reis</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Goldstein</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Estimating the distribution of selection coefficients from phylogenetic data using sitewise mutation-selection models</article-title>. <source>Genetics</source> <volume>190</volume>, <fpage>1101</fpage>&#x2013;<lpage>1115</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.111.136432</pub-id>
</citation>
</ref>
<ref id="B156">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tavar&#xe9;</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>Some probabilistic and statistical problems in the analysis of DNA sequences</article-title>. <source>Lect. Math. Life Sci.</source> <volume>17</volume>, <fpage>57</fpage>&#x2013;<lpage>86</lpage>.</citation>
</ref>
<ref id="B157">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Templeton</surname>
<given-names>A. R.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Contingency tests of neutrality using intra/interspecific gene trees: The rejection of neutrality for the evolution of the mitochondrial cytochrome oxidase II gene in the hominoid primates</article-title>. <source>Genetics</source> <volume>144</volume>, <fpage>1263</fpage>&#x2013;<lpage>1270</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/144.3.1263</pub-id>
</citation>
</ref>
<ref id="B158">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Thorne</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>Lartillot</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Rodrigue</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>S. C.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Codon models as a vehicle for reconciling population genetics with inter-specific sequence data</article-title>,&#x201d; in <source>Codon evolution: Mechanisms and models</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Cannarozzi</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>A.</given-names>
</name>
</person-group> (<publisher-name>Oxford University Press</publisher-name>). <pub-id pub-id-type="doi">10.1093/acprof:osobl/9780199601165.003.0007</pub-id>
</citation>
</ref>
<ref id="B159">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thornton</surname>
<given-names>J. W.</given-names>
</name>
<name>
<surname>Need</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Crews</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Resurrecting the ancestral steroid receptor: ancient origin of estrogen signaling</article-title>. <source>Science</source> <volume>301</volume>, <fpage>1714</fpage>&#x2013;<lpage>1717</lpage>. <pub-id pub-id-type="doi">10.1126/science.1086185</pub-id>
</citation>
</ref>
<ref id="B160">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Venkat</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hahn</surname>
<given-names>M. W.</given-names>
</name>
<name>
<surname>Thornton</surname>
<given-names>J. W.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Multinucleotide mutations cause false inferences of lineage-specific positive selection</article-title>. <source>Nat. Ecol. Evol.</source> <volume>2</volume>, <fpage>1280</fpage>&#x2013;<lpage>1288</lpage>. <pub-id pub-id-type="doi">10.1038/s41559-018-0584-5</pub-id>
</citation>
</ref>
<ref id="B161">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vishnoi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kryazhimskiy</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bazykin</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Hannenhalli</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Plotkin</surname>
<given-names>J. B.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Young proteins experience more variable selection pressures than old proteins</article-title>. <source>Genome Res.</source> <volume>20</volume>, <fpage>1574</fpage>&#x2013;<lpage>1581</lpage>. <pub-id pub-id-type="doi">10.1101/gr.109595.110</pub-id>
</citation>
</ref>
<ref id="B162">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Saeed</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Genome-wide analysis of codon usage bias in four sequenced cotton species</article-title>. <source>PLOS ONE</source> <volume>13</volume>, <fpage>e0194372</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0194372</pub-id>
</citation>
</ref>
<ref id="B163">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Watterson</surname>
<given-names>G. A.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Motoo kimura&#x2019;s use of diffusion theory in population genetics</article-title>. <source>Theor. Popul. Biol.</source> <volume>49</volume>, <fpage>154</fpage>&#x2013;<lpage>188</lpage>. <pub-id pub-id-type="doi">10.1006/tpbi.1996.0010</pub-id>
</citation>
</ref>
<ref id="B164">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whelan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>A general empirical model of protein evolution derived from multiple protein families using a maximum-likelihood approach</article-title>. <source>Mol. Biol. Evol.</source> <volume>18</volume>, <fpage>691</fpage>&#x2013;<lpage>699</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a003851</pub-id>
</citation>
</ref>
<ref id="B165">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whelan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Estimating the frequency of events that cause multiple-nucleotide changes</article-title>. <source>Genetics</source> <volume>167</volume>, <fpage>2027</fpage>&#x2013;<lpage>2043</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.103.023226</pub-id>
</citation>
</ref>
<ref id="B166">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whelan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li&#xf2;</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Molecular phylogenetics: State-of-the-art methods for looking into the past</article-title>. <source>Trends Genet. TIG</source> <volume>17</volume>, <fpage>262</fpage>&#x2013;<lpage>272</lpage>. <pub-id pub-id-type="doi">10.1016/s0168-9525(01)02272-7</pub-id>
</citation>
</ref>
<ref id="B167">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wilson</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>McVean</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Estimating diversifying selection and functional constraint in the presence of recombination</article-title>. <source>Genetics</source> <volume>172</volume>, <fpage>1411</fpage>&#x2013;<lpage>1425</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.105.044917</pub-id>
</citation>
</ref>
<ref id="B168">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolf</surname>
<given-names>Y. I.</given-names>
</name>
<name>
<surname>Novichkov</surname>
<given-names>P. S.</given-names>
</name>
<name>
<surname>Karev</surname>
<given-names>G. P.</given-names>
</name>
<name>
<surname>Koonin</surname>
<given-names>E. V.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>The universal distribution of evolutionary rates of genes and distinct characteristics of eukaryotic genes of different apparent ages</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>106</volume>, <fpage>7273</fpage>&#x2013;<lpage>7280</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0901808106</pub-id>
</citation>
</ref>
<ref id="B169">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wong</surname>
<given-names>W. S. W.</given-names>
</name>
<name>
<surname>Sainudiin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Identification of physicochemical selective pressure on protein encoding nucleotide sequences</article-title>. <source>BMC Bioinforma.</source> <volume>7</volume>, <fpage>148</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-7-148</pub-id>
</citation>
</ref>
<ref id="B170">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>X.-M.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>S.-F.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>D.-M.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Y.-P.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>F.-C.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>The analysis method and progress in the study of codon bias</article-title>. <source>Yi Chuan Hered.</source> <volume>29</volume>, <fpage>420</fpage>&#x2013;<lpage>426</lpage>. <pub-id pub-id-type="doi">10.1360/yc-007-0420</pub-id>
</citation>
</ref>
<ref id="B171">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lewis</surname>
<given-names>P. O.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kuo</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>M.-H.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Improving marginal likelihood estimation for Bayesian phylogenetic model selection</article-title>. <source>Syst. Biol.</source> <volume>60</volume>, <fpage>150</fpage>&#x2013;<lpage>160</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/syq085</pub-id>
</citation>
</ref>
<ref id="B172">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Xiong</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2006</year>). <source>Essential bioinformatics</source>. <publisher-loc>Cambridge, United Kingdom</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B173">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Codon-substitution models for detecting molecular adaptation at individual sites along specific lineages</article-title>. <source>Mol. Biol. Evol.</source> <volume>19</volume>, <fpage>908</fpage>&#x2013;<lpage>917</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a004148</pub-id>
</citation>
</ref>
<ref id="B190">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Inference of selection from multiple species alignments</article-title>. <source>Curr. Opin. Genet. Dev.</source> <volume>12</volume>, <fpage>688</fpage>&#x2013;<lpage>694</lpage>.</citation>
</ref>
<ref id="B174">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Mutation-selection models of codon substitution and their use to estimate selective strengths on codon usage</article-title>. <source>Mol. Biol. Evol.</source> <volume>25</volume>, <fpage>568</fpage>&#x2013;<lpage>579</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msm284</pub-id>
</citation>
</ref>
<ref id="B175">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Hasegawa</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Models of amino acid substitution and applications to mitochondrial protein evolution</article-title>. <source>Mol. Biol. Evol.</source> <volume>15</volume>, <fpage>1600</fpage>&#x2013;<lpage>1611</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a025888</pub-id>
</citation>
</ref>
<ref id="B176">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Goldman</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Pedersen</surname>
<given-names>A. M.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Codon-substitution models for heterogeneous selection pressure at amino acid sites</article-title>. <source>Genetics</source> <volume>155</volume>, <fpage>431</fpage>&#x2013;<lpage>449</lpage>. <pub-id pub-id-type="doi">10.1093/genetics/155.1.431</pub-id>
</citation>
</ref>
<ref id="B177">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1994a</year>). <article-title>Estimating the pattern of nucleotide substitution</article-title>. <source>J. Mol. Evol.</source> <volume>39</volume>, <fpage>105</fpage>&#x2013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1007/BF00178256</pub-id>
</citation>
</ref>
<ref id="B178">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1994b</year>). <article-title>Maximum likelihood phylogenetic estimation from DNA sequences with variable rates over sites: Approximate methods</article-title>. <source>J. Mol. Evol.</source> <volume>39</volume>, <fpage>306</fpage>&#x2013;<lpage>314</lpage>. <pub-id pub-id-type="doi">10.1007/BF00160154</pub-id>
</citation>
</ref>
<ref id="B179">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Maximum-likelihood models for combined analyses of multiple sequence data</article-title>. <source>J. Mol. Evol.</source> <volume>42</volume>, <fpage>587</fpage>&#x2013;<lpage>596</lpage>. <pub-id pub-id-type="doi">10.1007/BF02352289</pub-id>
</citation>
</ref>
<ref id="B180">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Relating physicochemical properties of amino acids to variable nucleotide substitution patterns among sites</article-title>. <source>Pac. Symp. Biocomput. Pac. Symp. Biocomput.</source> <volume>1999</volume>, <fpage>81</fpage>&#x2013;<lpage>92</lpage>.</citation>
</ref>
<ref id="B181">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2006</year>). <source>Computational molecular evolution</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>OUP</publisher-name>.</citation>
</ref>
<ref id="B182">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>PAML 4: phylogenetic analysis by maximum likelihood</article-title>. <source>Mol. Biol. Evol.</source> <volume>24</volume>, <fpage>1586</fpage>&#x2013;<lpage>1591</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msm088</pub-id>
</citation>
</ref>
<ref id="B183">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zaheri</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Dib</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Salamin</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A generalized mechanistic codon model</article-title>. <source>Mol. Biol. Evol.</source> <volume>31</volume>, <fpage>2528</fpage>&#x2013;<lpage>2541</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msu196</pub-id>
</citation>
</ref>
<ref id="B184">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zalucki</surname>
<given-names>Y. M.</given-names>
</name>
<name>
<surname>Power</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Jennings</surname>
<given-names>M. P.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Selection for efficient translation initiation biases codon usage at second amino acid position in secretory proteins</article-title>. <source>Nucleic Acids Res.</source> <volume>35</volume>, <fpage>5748</fpage>&#x2013;<lpage>5754</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkm577</pub-id>
</citation>
</ref>
<ref id="B185">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Analysis of codon usage bias of envelope glycoprotein genes in nuclear polyhedrosis virus (NPV) and its relation to evolution</article-title>. <source>BMC Genomics</source> <volume>17</volume>, <fpage>677</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-016-3021-7</pub-id>
</citation>
</ref>
<ref id="B186">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Dang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Codon usage is an important determinant of gene expression levels largely through its effects on transcription</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>113</volume>, <fpage>E6117</fpage>&#x2013;<lpage>E6125</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1606724113</pub-id>
</citation>
</ref>
<ref id="B187">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zoller</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Empirical analysis of the most relevant parameters of codon substitution models</article-title>. <source>J. Mol. Evol.</source> <volume>70</volume>, <fpage>605</fpage>&#x2013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1007/s00239-010-9356-9</pub-id>
</citation>
</ref>
<ref id="B188">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zoller</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>A new semiempirical codon substitution model based on principal component analysis of mammalian sequences</article-title>. <source>Mol. Biol. Evol.</source> <volume>29</volume>, <fpage>271</fpage>&#x2013;<lpage>277</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msr198</pub-id>
</citation>
</ref>
<ref id="B189">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zoller</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Boskova</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Anisimova</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Maximum-likelihood tree estimation using codon substitution models with multiple partitions</article-title>. <source>Mol. Biol. Evol.</source> <volume>32</volume>, <fpage>2208</fpage>&#x2013;<lpage>2216</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msv097</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>