<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2022.875202</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Genetic Divergence of Lineage-Specific Tandemly Duplicated Gene Clusters in Four Diploid Potato Genotypes</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Bonthala</surname>
<given-names>Venkata Suresh</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/245203/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Stich</surname>
<given-names>Benjamin</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="aff3" ref-type="aff"><sup>3</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/499492/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute of Quantitative Genetics and Genomics of Plants, Heinrich Heine University of D&#x00FC;sseldorf</institution>, <addr-line>D&#x00FC;sseldorf</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Max Planck Institute for Plant Breeding Research</institution>, <addr-line>K&#x00F6;ln</addr-line>, <country>Germany</country></aff>
<aff id="aff3"><sup>3</sup><institution>Cluster of Excellence on Plant Sciences, From Complex Traits Towards Synthetic Modules</institution>, <addr-line>D&#x00FC;sseldorf</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Nunzio D'Agostino, University of Naples Federico II, Italy</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Alfonso Del Rio, University of Wisconsin-Madison, United States; Nicholas Louis Panchy, The University of Tennessee, Knoxville, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Venkata Suresh Bonthala, <email>bonthala@hhu.de</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Plant Bioinformatics, a section of the journal Frontiers in Plant Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>875202</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Bonthala and Stich.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Bonthala and Stich</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Potato (<italic>Solanum tuberosum</italic> L.) is the most important non-grain food crop. Tandem duplication significantly contributes to genome evolution. The objectives of this study were to (i) identify tandemly duplicated genes and compare their genomic distributions across potato genotypes, (ii) investigate the bias in functional specificities, (iii) explore the relationships among coding sequence, promoter and expression divergences associated with tandemly duplicated genes, (iv) examine the role of tandem duplication in generating and expanding lineage-specific gene families, (v) investigate the evolutionary forces affecting tandemly duplicated genes, and (vi) assess the similarities and differences with respect to above mentioned aspects between cultivated genotypes and their wild-relative. In this study, we used well-annotated and chromosome-scale <italic>de novo</italic> genome assemblies of multiple potato genotypes. Our results showed that tandemly duplicated genes are abundant and dispersed through the genome. We found that several functional specificities, such as disease resistance, stress-tolerance, and biosynthetic pathways of tandemly duplicated genes were differentially enriched across multiple potato genomes. Our results indicated the existence of a significant correlation among expression, promoter, and protein divergences in tandemly duplicated genes. We found about one fourth of tandemly duplicated gene clusters as lineage-specific among multiple potato genomes, and these tended to localize toward centromeres and revealed distinct selection signatures and expression patterns. Furthermore, our results showed that a majority of duplicated genes were retained through sub-functionalization followed by genetic redundancy, while only a small fraction of duplicated genes was retained though neo-functionalization. The lineage-specific expansion of gene families by tandem duplication coupled with functional bias might have significantly contributed to potato&#x2019;s genotypic diversity, and, thus, to adaption to environmental stimuli.</p>
</abstract>
<kwd-group>
<kwd>tandem duplication</kwd>
<kwd>lineage-specific duplicated genes</kwd>
<kwd>gene expression</kwd>
<kwd>whole-genome duplication</kwd>
<kwd>agronomic traits</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="77"/>
<page-count count="17"/>
<word-count count="11953"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>Gene duplication is thought to have significantly contributed to the evolution of genetic and morphological diversity, and speciation in eukaryotes (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>). Plant genomes contain a significant proportion of duplicated genes that are formed by various mechanisms, including single-gene duplications and larger chromosomal regions or whole-genome duplication (WGD or polyploidization; <xref ref-type="bibr" rid="ref19">Freeling and Thomas, 2006</xref>; <xref ref-type="bibr" rid="ref16">Flagel and Wendel, 2009</xref>). WGD is prevalent in the plant kingdom and involves duplication of all nuclear genes of an organism. This in turn leads to a sudden increase in both genome size and the entire gene set (<xref ref-type="bibr" rid="ref44">Moghe and Shiu, 2014</xref>; <xref ref-type="bibr" rid="ref58">Salman-Minkov et al., 2016</xref>). Many angiosperm lineages experienced repeated WGD events throughout their evolutionary history, and genome sequencing continues to report new events in various plant species (<xref ref-type="bibr" rid="ref59">Schmutz et al., 2010</xref>; <xref ref-type="bibr" rid="ref76">Zhang et al., 2011</xref>; <xref ref-type="bibr" rid="ref39">Li et al., 2015</xref>). Recent WGD that have occurred in lineages of crop species, including soybean, wheat, and cotton, have contributed to important traits, such as nodulation and oil production (<xref ref-type="bibr" rid="ref59">Schmutz et al., 2010</xref>), grain quality (<xref ref-type="bibr" rid="ref76">Zhang et al., 2011</xref>), and spinnable fibers (<xref ref-type="bibr" rid="ref39">Li et al., 2015</xref>), respectively.</p>
<p>In addition to WGD, single-gene duplications, such as tandem, proximal, dispersed, DNA-transposed, and retrotransposed duplications are also prevalent in plant genomes. These contribute to the expansion and evolution of multigene families (<xref ref-type="bibr" rid="ref6">Cannon et al., 2004</xref>; <xref ref-type="bibr" rid="ref18">Freeling, 2009</xref>; <xref ref-type="bibr" rid="ref51">Qiao et al., 2018</xref>, <xref ref-type="bibr" rid="ref50">2019</xref>). Tandemly duplicated genes (TDG) are present next to the original copy or are intervened by several unrelated genes in the same genomic neighborhoods and often occur as a result of unequal crossing over followed by inversions or transposon activities (<xref ref-type="bibr" rid="ref18">Freeling, 2009</xref>). Furthermore, these genes are found to be scattered throughout the genome but the majority tend to localize toward terminal regions of the chromosomes (<xref ref-type="bibr" rid="ref31">Jiang et al., 2013</xref>; <xref ref-type="bibr" rid="ref36">Kono et al., 2018</xref>; <xref ref-type="bibr" rid="ref51">Qiao et al., 2018</xref>). TDGs exhibit bias in functional specificities to generate functional novelties in the genome (<xref ref-type="bibr" rid="ref31">Jiang et al., 2013</xref>; <xref ref-type="bibr" rid="ref51">Qiao et al., 2018</xref>). In addition, tandem duplication generates lineage-specific gene duplicates followed by their expansion among evolutionarily closed (<xref ref-type="bibr" rid="ref36">Kono et al., 2018</xref>) and distant species (<xref ref-type="bibr" rid="ref24">Hanada et al., 2008</xref>) for adaptive response to environmental stimuli. However, it is unclear whether tandem duplication creates bias in functional specificities and lineage-specific gene duplicates between cultivated species and their wild relatives.</p>
<p>Tandemly duplicated genes may experience different evolutionary fates such as (i) loss of one of duplicated gene copy <italic>via</italic> pseudogenization and/or accumulation of deleterious mutations (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>), (ii) retention of both duplicated genes due to selection for genetic redundancy that may be beneficial (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>), (iii) retention of duplicated genes simply because there has been insufficient time for one copy to be removed/mutated or due to genetic drift (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>), and (iv) retention of both duplicated genes because of a selective advantage either due to the existing or the novel functions (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>). The retention of both duplicated genes because of selective advantage due to the existing functions can be explained by gene dosage (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>), sub-functionalization (<xref ref-type="bibr" rid="ref17">Force et al., 1999</xref>), dosage balance (<xref ref-type="bibr" rid="ref19">Freeling and Thomas, 2006</xref>), and paralog interference (<xref ref-type="bibr" rid="ref2">Baker et al., 2013</xref>) models. Similarly, both duplicated genes can be retained because of selective advantage due to novel functions and can be explained by neo-functionalization (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>) and escape from adaptive conflict (<xref ref-type="bibr" rid="ref9">Des Marais and Rausher, 2008</xref>) models. Among the above-mentioned models, both sub- and neo-functionalization provide testable hypotheses suggesting that the sub-functionalized gene copies show divergence in expression across tissues and are expected to undergo purifying selection (i.e., <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003C;&#x2009;1) because the functions of ancestral gene have become divided among the daughter copies (<xref ref-type="bibr" rid="ref17">Force et al., 1999</xref>; <xref ref-type="bibr" rid="ref10">Duarte et al., 2006</xref>; <xref ref-type="bibr" rid="ref7">Cusack and Wolfe, 2007</xref>; <xref ref-type="bibr" rid="ref43">Ma et al., 2015</xref>), whereas neo-functionalized gene copies undergo positive selection (i.e., <italic>K</italic><sub>a</sub>/<italic>K<sub>s</sub></italic>&#x2009;&#x003E;&#x2009;1) because gain of a novel function by one of gene copy that contributes to better fitness (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>; <xref ref-type="bibr" rid="ref3">Blanc and Wolfe, 2004</xref>). Based on these testable hypotheses, <xref ref-type="bibr" rid="ref57">Roulin et al., 2013</xref> unraveled the contribution of sub- and neo-functionalization in retention of duplicated genes in soybean, and found that 50% of paralogs have undergone expression sub-functionalization, while a small fraction of paralogs has been neo-functionalized. However, it is unclear whether sub- or/and neo-functionalization play a role in retention of TDGs between cultivated species and their wild relatives. In addition, it is also unclear whether tandem duplication creates different proportion of duplicated genes between cultivated species and their wild relatives.</p>
<p>Potato (<italic>Solanum tuberosum</italic>. L) is a highly heterozygous autotetraploid species, and is the world&#x2019;s most important non-grain food crop with a worldwide production of 370 million metric tons <italic>per annum</italic> (<xref ref-type="bibr" rid="ref14">FAO, 2019</xref>). <xref ref-type="bibr" rid="ref66">Wang et al. (2018)</xref> focused on comparative analysis of DNA methylation patterns between duplicated genes of potato and tomato, and found DNA methylation divergence between duplicated genes. Recently, <xref ref-type="bibr" rid="ref50">Qiao et al. (2019)</xref> investigated the signatures of selection, expression divergence, and gene conversion underlying evolution of duplicated genes across 141 plant species including potato. However, these two studies did not address various aspects associated with TDGs in potato. This includes the genomic distribution, bias in functional specificities, the influence of evolutionary forces, and relationships among coding sequence, promoter and expression divergences associated with TDGs. Further, the role of tandem duplication in generating lineage-specific gene families and their expansions, as well as the factors contributing to the retention of TDGs, were not studied in potato yet.</p>
<p>The objectives of our study were to (i) identify TDGs and compare their genomic distributions across potato genotypes, (ii) investigate the bias in functional specificities, (iii) explore the relationships among coding sequence, promoter, and expression divergences of TDGs, (iv) examine the role of tandem duplication in generating and expanding lineage-specific gene families, (v) investigate the evolutionary forces affecting the TDGs, and (vi) assess the similarities and differences with respect to above mentioned aspects between cultivated genotypes and their wild-relative.</p>
</sec>
<sec id="sec2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="sec3">
<title>Data Sources</title>
<p>Thousands of potato cultivars exists and most of them are tetraploid (2n&#x2009;=&#x2009;4x&#x2009;=&#x2009;48; <xref ref-type="bibr" rid="ref13">FAO, 2008</xref>). However, chromosome level genome assemblies for tetraploid potato clones became only recently available (<xref ref-type="bibr" rid="ref26">Hoopes et al., 2022</xref>), after the analyses for this study were finalized. Instead, we used four tuber-bearing diploid potato clones belonging to cultivated, non-cultivated, and wild potato species for which chromosome-level genome assemblies are available. The cultivated potato <italic>Solanum tuberosum</italic> ssp. <italic>tuberosum</italic> L. is represented in our study by a diploid clone (hereafter referred to as dAg) which was derived from the tetraploid elite potato cultivar Agria (tAg; <xref ref-type="bibr" rid="ref20">Freire et al., 2021</xref>). The non-cultivated potato clones include <italic>Solanum tuberosum</italic> L. DM1-3516 R44 (hereafter referred to as DM), which is a doubled monoploid clone derived from the group Phureja (<xref ref-type="bibr" rid="ref47">Pham et al., 2020</xref>) and <italic>Solanum tuberosum</italic> L. RH89-039-16 (hereafter referred to as RH) which is a diploid clone derived from a cross between a dihaploid and a diploid potato (<xref ref-type="bibr" rid="ref78">Zhou et al., 2020</xref>). The wild clone in our study is <italic>Solanum chacoense</italic> M6 (hereafter referred to as M6; <xref ref-type="bibr" rid="ref38">Leisner et al., 2018</xref>). Both the sequence (genome and gene) and annotation files for DM (version 6.1), RH, and M6 were downloaded from <ext-link xlink:href="http://potato.plantbiology.msu.edu" ext-link-type="uri">http://potato.plantbiology.msu.edu</ext-link> and for dAg was obtained from <xref ref-type="bibr" rid="ref20">Freire et al. (2021)</xref>. In addition, we obtained transposable elements (TEs) annotation for dAg and DM from the above-mentioned sources. Due to the lack or absence of chromosome-level TE annotation for M6 and RH, we excluded TE annotation for these two genotypes from the analyses.</p>
</sec>
<sec id="sec4">
<title>Improving the Gene Annotation for dAg</title>
<p>In this study, we improved the existing gene annotation for dAg using the PASA pipeline v2.5.0 (<xref ref-type="bibr" rid="ref21">Haas et al., 2003</xref>) followed by classifying the resulting annotation into full-length and partial gene models using AGAT v0.8.0 (<xref ref-type="bibr" rid="ref8">Dainat et al., 2022</xref>). In that procedure, we used transcriptome datasets generated as part of dAg genome sequencing (<xref ref-type="bibr" rid="ref20">Freire et al., 2021</xref>) to improve the existing gene annotations.</p>
</sec>
<sec id="sec5">
<title>Functional Annotation, Orthology Prediction, and Filtering Transposons</title>
<p>For reasons of consistency, we have performed functional annotation for the longest isoform of high-confidence genes of all four potato genomes using the AHRD pipeline.<xref rid="fn0004" ref-type="fn"><sup>1</sup></xref> Orthologs among the four potato genomes were predicted by feeding protein sequences of the longest isoform of high-confidence genes to OrthoFinder v2.5.4 (<xref ref-type="bibr" rid="ref12">Emms and Kelly, 2019</xref>). As, high-confidence genes, we considered those genes for which expression/functional evidence was available and that had full-length without internal stop-codons. Hence, annotated partial/pseudogenes were ignored in our study. Furthermore, we annotated transposon (TE) or TE-related genes in all four potato genomes using the approach described by <xref ref-type="bibr" rid="ref30">Jayakodi et al. (2020)</xref>. Briefly, this approach involves two stages to annotate TEs in all annotated genes. The first stage involves searching for keywords and PFAM IDs related to TEs in the functional annotation obtained from AHRD and classify each gene as either TE or non-TE gene. The second stage involves combining the orthogroup (OG) information for each gene obtained from OrthoFinder with the curated genes of the first stage. In the last step, we evaluated whether each OG is classified as non-TE OG based on the criteria that the OG contains less than 30% of TE genes and the mean AHRD score of OG is &#x2265;2, otherwise the OG is classified as TE OG.</p>
</sec>
<sec id="sec6">
<title>Identification of TDG Clusters</title>
<p>Tandemly duplicated genes clusters were identified among the non-TE genes of each potato genome separately using the methodology described by <xref ref-type="bibr" rid="ref30">Jayakodi et al. (2020)</xref>. In our study, we restricted our analyses to non-TE genes with a known chromosomal location. Briefly, this approach involves the identification of homologous gene pairs present on the same chromosome using all vs. all BlastN (<xref ref-type="bibr" rid="ref1">Altschul et al., 1990</xref>) of coding sequences (CDS) of the longest iso-form of high-confidence gene models. This is followed by finding TDG clusters based on the following thresholds: <italic>e</italic>-value cut off of 1e<sup>&#x2212;10</sup>, bit score ratio of &#x2265;30%, and coverage of both query and subjects of &#x2265;50%. In this study, we defined TDGs as the homologous genes present on the same chromosome which are intervened by up to 10 genes. TDGs (or TDG pair) correspond to single genes (or gene pairs) that belong to a TDG cluster. A TDG cluster corresponds to a group of TDG pairs.</p>
<p>Further, the variation in density of TDGs across the respective genomes were explained by fitting a general linear model against the density of various genomic features such as all non-TE genes, DNA transposable elements (TEs), and RNA TEs using R v3.6.1.<xref rid="fn0005" ref-type="fn"><sup>2</sup></xref> In the next step, the residuals of these models were tested against a uniform distribution using a Kolmogorov&#x2013;Smirnov (KS) test in order to evaluate whether the TDGs were distributed uniformly across the genome after adjusting for the distribution effects of the above-mentioned genomic features. In addition, Pearson&#x2019;s correlation coefficients between density of TDGs and the above-mentioned genomic features were computed. All density calculations were performed in 1&#x2009;Mb windows across the genome. For TDG also, 1.5&#x2009;Mb windows were considered. Density was defined as the proportion of bases in each window that corresponded to the respective genomic feature. TDG clusters were categorized based on their level of sharing across four potato genomes into core, i.e., present in all four potato genomes, shared, i.e., present in more than one potato genome but absent in at least one potato genome, and private (or lineage-specific) clusters, i.e., present in a single potato genome.</p>
</sec>
<sec id="sec7">
<title>Enrichment of Pfam Domains and Gene Ontology Terms Among TDGs</title>
<p>Both Pfam domain and Gene Ontology (GO) term information was extracted from the output of AHRD for each potato genome. For each detected Pfam domain, we calculated the number of proteins present among the proteins of TDGs followed by performing the protein domain enrichment analysis using Fisher exact test (<xref ref-type="bibr" rid="ref15">Fisher, 1992</xref>). FDR corrected values of <italic>p</italic>&#x2009;&#x003C;&#x2009;0.05 were considered as significant. Enrichment of GO terms was performed using GOATOOLS (<xref ref-type="bibr" rid="ref35">Klopfenstein et al., 2018</xref>).</p>
</sec>
<sec id="sec8">
<title>Gene Expression Quantification and Estimating Expression Divergence</title>
<p>For this analysis, publicly available RNA-Seq datasets from NCBI SRA<xref rid="fn0006" ref-type="fn"><sup>3</sup></xref> were used (<xref ref-type="supplementary-material" rid="SM1">Supplementary Tables S5</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">S8</xref>). The raw-reads were filtered for low-quality and trimmed adapter sequences using Trimmomatic v0.39 (<xref ref-type="bibr" rid="ref4">Bolger et al., 2014</xref>) with the following parameters: (1) adapters were removed using for pair-end: ILLUMINACLIP:TruSeq3-PE.fa:2:30:10:8:true, for single-end: ILLUMINACLIP:TruSeq3-SE.fa:2:30:10:8; (2) removing leading and trailing low-quality or <italic>N</italic> bases using LEADING:3 TRAILING:3; (3) scanning the read with a four-base wide sliding window, cutting when the average quality per base drops below 20 (SLIDINGWINDOW:4:20); and (4) selecting reads with at least 50 (for pair-end: MINLEN:50) or 36 bases long (for single-end: MINLEN:36). Only, RNA-Seq datasets with at least 50% of high-quality reads after trimming were chosen for further analysis. The high-quality reads were used to estimate transcripts abundances on a gene level (i.e., averaging across the alleles present at the respective gene) using Kalisto v0.46.1 (<xref ref-type="bibr" rid="ref5">Bray et al., 2016</xref>) with default parameters (for single-ends: -l 200 -s 20) and obtained transcript per million (TPM) values. For all examined four genomes, we classified a gene as expressed, if TPM was &#x003E;0.5 in at least one RNA-Seq dataset, otherwise classified as unexpressed. In the next step, we evaluated expression breadth, i.e., the number of RNA-Seq datasets in which the gene was expressed. Further, the expressed genes were used to calculate the expression level and expression specificity. The expression level was defined as the mean value of TPM across RNA-Seq datasets for each gene. The expression specificity was measured as described by <xref ref-type="bibr" rid="ref72">Yang and Gaut (2011)</xref> and ranged from 0 to 1, with a higher value indicating higher specificity, i.e., higher variation in expression across RNA-Seq datasets. If a gene is expressed in a single library only, the expression specificity is 1. In contrast, if a gene is expressed equally in all RNA-Seq datasets, the expression specificity is 0.</p>
</sec>
<sec id="sec9">
<title>Collection of Promoter Sequences and Estimating Promoter Divergence</title>
<p>For each potato genome and each gene, we considered the non-overlapping 1&#x2009;Kb sequence upstream of the transcription start site (TSS) as putative promoter sequence and retrieved it from the respective potato genomes using BEDTools v2.27.1 (<xref ref-type="bibr" rid="ref52">Quinlan and Hall, 2010</xref>). The promoter sequences for each tandem duplicate gene pair were aligned using the matcher program of EMBOSS v6.6.0.0 (<xref ref-type="bibr" rid="ref56">Rice et al., 2000</xref>) to compute promoter similarity (Ps). We excluded promoter sequences that contained unknown nucleotides (N).</p>
</sec>
<sec id="sec10">
<title>Computing <italic>K</italic><sub>a</sub>, <italic>K</italic><sub>s</sub>, and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> Values</title>
<p>For each TDG pair, the amino acid sequences were aligned using MAFFT v7.453 (<xref ref-type="bibr" rid="ref34">Katoh and Standley, 2013</xref>) followed by reverse translation into nucleotide alignment using PAL2NAL v14 (<xref ref-type="bibr" rid="ref61">Suyama et al., 2006</xref>). Finally, the nucleotide alignments were used to compute <italic>K</italic><sub>a</sub>, <italic>K</italic><sub>s</sub>, and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> values using the Gamma-MYN method of <italic>K</italic><sub>a</sub><italic>K</italic><sub>s</sub>_Calculator v2.0 (<xref ref-type="bibr" rid="ref67">Wang et al., 2010</xref>). As the high number of reversions or multiple substitutions at synonymous sites reduces accuracy and reliability for rate estimation (<xref ref-type="bibr" rid="ref64">Vanneste et al., 2013</xref>), we excluded <italic>K</italic><sub>s</sub> values &#x003E;2 from the analysis.</p>
</sec>
<sec id="sec11">
<title>Correlation Analyses Among Promoter, Protein, and Expression Divergences</title>
<p>Pearson&#x2019;s correlation coefficient (<italic>r</italic>) was calculated across all TDG pairs between (i) expression and promoter divergence, (ii) expression divergence and age of duplicate pairs, (iii) expression divergence and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> ratios, and (iv) promoter divergence and age of duplicate pairs using SciPy library (<xref ref-type="bibr" rid="ref65">Virtanen et al., 2020</xref>) in Python v3.8.5.<xref rid="fn0007" ref-type="fn"><sup>4</sup></xref></p>
</sec>
</sec>
<sec id="sec12" sec-type="results">
<title>Results</title>
<sec id="sec13">
<title>Annotation of Transposon-Related Genes in Potato Genomes</title>
<p>The recently published gene annotation of the diploid clone dAg derived from the elite variety Agria does not contain information about isoforms and partial gene models. Hence, we first improved the existing gene annotation of dAg using the PASA pipeline (<xref ref-type="bibr" rid="ref21">Haas et al., 2003</xref>) by utilizing the available Iso-Seq data of Agria (tAg; <xref ref-type="bibr" rid="ref20">Freire et al., 2021</xref>) and obtained 44,464 gene models with 58,734 isoforms. We then filtered out partial gene models using AGAT tools (<xref ref-type="bibr" rid="ref8">Dainat et al., 2022</xref>) and obtained 39,088 full-length gene models with 53,352 isoforms (referred as full-length set in <xref rid="tab1" ref-type="table">Table 1A</xref>). This new annotation contains a significantly higher number of gene models than the high-confidence annotation of DM v6.1 (<xref ref-type="bibr" rid="ref47">Pham et al., 2020</xref>), and a slightly higher number of gene models than the annotation of M6 (<xref ref-type="bibr" rid="ref38">Leisner et al., 2018</xref>) and RH (<xref ref-type="bibr" rid="ref78">Zhou et al., 2020</xref>; <xref rid="tab1" ref-type="table">Table 1B</xref>).</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption><p><bold>(A)</bold> Improved gene annotation of dAg; <bold>(B)</bold> Non-TE genes of potato genomes.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Genomic feature</th>
<th align="center" valign="top">Old annotation</th>
<th align="center" valign="top">Working set</th>
<th align="center" valign="top">Full-length set</th>
<th align="center" valign="top">Representative set</th>
<th align="center" valign="top">Partial set</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" colspan="6">Table 1A</td>
</tr>
<tr>
<td align="left" valign="bottom"># Genes</td>
<td align="center" valign="bottom">44,952</td>
<td align="center" valign="bottom">44,464</td>
<td align="center" valign="bottom">39,088</td>
<td align="center" valign="bottom">39,088</td>
<td align="center" valign="bottom">5,382</td>
</tr>
<tr>
<td align="left" valign="top"># mRNAs</td>
<td align="center" valign="top">44,952</td>
<td align="center" valign="top">58,734</td>
<td align="center" valign="top">53,352</td>
<td align="center" valign="top">39,088</td>
<td align="center" valign="top">5,382</td>
</tr>
<tr>
<td align="left" valign="top"># CDSs</td>
<td align="center" valign="top">220,904</td>
<td align="center" valign="top">354,174</td>
<td align="center" valign="top">332,602</td>
<td align="center" valign="top">195,622</td>
<td align="center" valign="top">21,572</td>
</tr>
<tr>
<td align="left" valign="top"># Exons</td>
<td align="center" valign="top">226,161</td>
<td align="center" valign="top">374,932</td>
<td align="center" valign="top">353,161</td>
<td align="center" valign="top">201,589</td>
<td align="center" valign="top">21,771</td>
</tr>
<tr>
<td align="left" valign="top"># Introns</td>
<td align="center" valign="top">NA</td>
<td align="center" valign="top">316,198</td>
<td align="center" valign="top">299,809</td>
<td align="center" valign="top">162,501</td>
<td align="center" valign="top">16,389</td>
</tr>
<tr>
<td align="left" valign="top"># 5`-UTRs</td>
<td align="center" valign="top">17,300</td>
<td align="center" valign="top">44,596</td>
<td align="center" valign="top">43,643</td>
<td align="center" valign="top">20,385</td>
<td align="center" valign="top">953</td>
</tr>
<tr>
<td align="left" valign="top"># 3`-UTRs</td>
<td align="center" valign="top">16,118</td>
<td align="center" valign="top">39,268</td>
<td align="center" valign="top">39,049</td>
<td align="center" valign="top">19,432</td>
<td align="center" valign="top">219</td>
</tr>
<tr>
<td align="left" valign="top" colspan="6">Table 1B</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Potato Genotype</td>
<td align="center" valign="top"># Genes before TE Filtering</td>
<td align="center" valign="top"># Genes after TE filtering</td>
<td align="center" valign="top" colspan="2">% Genes TEs</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">dAg</td>
<td align="center" valign="top">39,088</td>
<td align="center" valign="top">33,934</td>
<td align="center" valign="top" colspan="2">13.19</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">DM</td>
<td align="center" valign="top">32,917</td>
<td align="center" valign="top">31,494</td>
<td align="center" valign="top" colspan="2">4.32</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">M6</td>
<td align="center" valign="top">37,740</td>
<td align="center" valign="top">35,330</td>
<td align="center" valign="top" colspan="2">6.39</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">RH</td>
<td align="center" valign="top">37,115</td>
<td align="center" valign="top">31,249</td>
<td align="center" valign="top" colspan="2">15.8</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Total</td>
<td align="center" valign="top">146,860</td>
<td align="center" valign="top">132,007</td>
<td align="center" valign="top" colspan="2">10.11</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Genes left after TE filtering were used for all down-stream analyses. The working set, Full-length set, Representative set, and Partial sets indicate all annotated gene models, gene models with full length genes, gene models with longest isoforms, and gene models with partial genes, respectively.</p>
</table-wrap-foot>
</table-wrap>
<p>To ensure that the results of our analyses can be compared across all four potato genomes, the functional annotation for all four potato genomes was performed using the AHRD pipeline.<xref rid="fn0008" ref-type="fn"><sup>5</sup></xref> This pipeline assigns a quality score in the form of a three-character string, where each character is either &#x201C;&#x002A;&#x201D; if respective criteria is met or &#x201C;-&#x201D; otherwise, for each annotated gene to indicate how confident the assigned annotation is. The &#x201C;&#x002A;&#x201D; in first position indicates bit score, and <italic>e</italic>-value of the blast result are &#x003E;50 and 1e<sup>&#x2212;10</sup>, respectively. The &#x201C;&#x002A;&#x201D; in second position indicates overlap of the blast result is &#x003E;60%, and the &#x201C;&#x002A;&#x201D; in third position indicates top token score of assigned Human-Readable-Description is &#x003E;0.5. In our study, we selected annotation with a quality score of at least two stars as best annotation, i.e., at least two out of three criteria should meet by the annotated gene. Overall, AHRD assigned functions to 84.31% of the genes of all four potato genomes with at least two stars, whereas, AHRD was unable to annotate 13.13% of the genes of all four potato genomes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S1</xref>). We also estimated orthologs and orthogroups among the four potato genomes using OrthoFinder (<xref ref-type="bibr" rid="ref12">Emms and Kelly, 2019</xref>) and obtained 28,647 orthogroups representing 93.5% of the genes of all potato genomes. A total of 16,107 orthogroups contained genes from all four potato genomes, of which 9,563 were single-copy orthogroups (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S2</xref>). Finally, we annotated 10.11% of the genes (<xref rid="tab1" ref-type="table">Table 1B</xref>) of all four potato genomes as TE or TE-related genes using the approach of <xref ref-type="bibr" rid="ref30">Jayakodi et al. (2020)</xref>.</p>
</sec>
<sec id="sec14">
<title>Systematic Identification of TDG Clusters in Four Potato Genomes</title>
<p>Tandemly duplicated genes clusters were identified among the non-TE genes of each potato genome. In total, 2,090, 1,867, 1,661, and 1,832 TDG clusters were identified in dAg, DM, M6, and RH, respectively. The percentage of annotated genes in TDG clusters of the total number of genes were 18.67, 18.52, 16.83, and 16.06% in dAg, DM, M6, and RH, respectively (<xref rid="tab2" ref-type="table">Table 2</xref>). The availability of multiple high-quality <italic>de novo</italic> potato genome assemblies allowed us to determine the consistency of various characteristics of TDG clusters across potato genomes. The identified TDGs were dispersed throughout the genome and shown to have a similar distribution across the four potato genomes (<xref rid="fig1" ref-type="fig">Figure 1</xref>). Moreover, the density distribution of TDGs with both 1 and 1.5&#x2009;Mb sliding-windows resulted in same density distribution pattern across the four potato genomes (<xref rid="fig1" ref-type="fig">Figure 1</xref>; <xref ref-type="supplementary-material" rid="SM5">Supplementary Figure S5</xref>). Overall, the density of TDGs across the genome was significantly (value of <italic>p</italic>&#x2009;&#x003C;&#x2009;2.2e<sup>&#x2212;16</sup>) associated with the density of non-TE genes across the respective potato genomes. The correlation coefficients were 0.6, 0.68, 0.61, and 0.63 for dAg, DM, M6, and RH, respectively. The same observation was made for the densities of both DNA and RNA TEs. After correcting for these densities using a general linear model, the density of TDG showed a significant (value of <italic>p</italic>&#x2009;&#x003C;&#x2009;2.2e<sup>&#x2212;16</sup>) deviation from a uniform distribution in each potato genome. The KS test statistics (<italic>D</italic>) are 0.38, 0.38, 0.41, and 0.57 for dAg, DM, M6, and RH, respectively. Chromosome 1 of all potato genomes harbored the highest number of TDG clusters, while Chromosomes 6, 5, 8, and 7 harbored the lowest number of TDG clusters in dAg, DM, M6, and RH, respectively (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S3</xref>). A similar distribution in terms of the size of TDG clusters was observed across the four potato genomes, and the majority of the TDG clusters comprised two genes (<xref rid="fig2" ref-type="fig">Figure 2A</xref>). Further, the majority of TDG pairs within TDG clusters did not contain intervening genes, and moreover, DM contained the highest number of TDG pairs without intervening genes among the four potato genomes (<xref rid="fig2" ref-type="fig">Figure 2B</xref>). For dAg, a higher proportion of TDGs with two exons was observed compared to the other three genomes. For the latter, the proportion of TDG with single exons was higher compared to that of non-tandemly duplicated genes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S1</xref>).</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption><p>Summary of identified putative tandemly duplicated gene clusters in potato genomes.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Description</th>
<th align="center" valign="top">dAg</th>
<th align="center" valign="top">DM</th>
<th align="center" valign="top">M6</th>
<th align="center" valign="top">RH</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Number of non-TE genes</td>
<td align="center" valign="top">33,934</td>
<td align="center" valign="top">31,494</td>
<td align="center" valign="top">35,330</td>
<td align="center" valign="top">31,249</td>
</tr>
<tr>
<td align="left" valign="top">Number of non-TE genes (Known Location)</td>
<td align="center" valign="top">32,555</td>
<td align="center" valign="top">31,410</td>
<td align="center" valign="top">28,210</td>
<td align="center" valign="top">31,249</td>
</tr>
<tr>
<td align="left" valign="top">Number of TDG Clusters</td>
<td align="center" valign="top">2090</td>
<td align="center" valign="top">1867</td>
<td align="center" valign="top">1,661</td>
<td align="center" valign="top">1832</td>
</tr>
<tr>
<td align="left" valign="top">Number of genes in TDG Clusters</td>
<td align="center" valign="top">6,078</td>
<td align="center" valign="top">5,817</td>
<td align="center" valign="top">4,748</td>
<td align="center" valign="top">5,018</td>
</tr>
<tr>
<td align="left" valign="top">Percentage of genes in TDG clusters</td>
<td align="center" valign="top">18.67</td>
<td align="center" valign="top">18.52</td>
<td align="center" valign="top">16.83</td>
<td align="center" valign="top">16.06</td>
</tr>
<tr>
<td align="left" valign="top">Number of Orthogroups</td>
<td align="center" valign="top">2,654</td>
<td align="center" valign="top">2,576</td>
<td align="center" valign="top">2,218</td>
<td align="center" valign="top">2,352</td>
</tr>
<tr>
<td align="left" valign="top">Percent of TDG clusters with two genes</td>
<td align="center" valign="top">66.89</td>
<td align="center" valign="top">60.47</td>
<td align="center" valign="top">64.90</td>
<td align="center" valign="top">68.72</td>
</tr>
<tr>
<td align="left" valign="top">Largest TDG cluster</td>
<td align="center" valign="top">36</td>
<td align="center" valign="top">31</td>
<td align="center" valign="top">23</td>
<td align="center" valign="top">21</td>
</tr>
<tr>
<td align="left" valign="top">Percent of TDGs with Pfam domains</td>
<td align="center" valign="top">83.42</td>
<td align="center" valign="top">91.47</td>
<td align="center" valign="top">86.73</td>
<td align="center" valign="top">78.68</td>
</tr>
<tr>
<td align="left" valign="top">Number of Pfam Protein Domains identified in TDGs</td>
<td align="center" valign="top">917</td>
<td align="center" valign="top">805</td>
<td align="center" valign="top">706</td>
<td align="center" valign="top">756</td>
</tr>
<tr>
<td align="left" valign="top">Number of Pfam Protein Domains Enriched in TDGs</td>
<td align="center" valign="top">46</td>
<td align="center" valign="top">69</td>
<td align="center" valign="top">60</td>
<td align="center" valign="top">61</td>
</tr>
<tr>
<td align="left" valign="top">Number of Pfam Domains Enriched in TDGs with <italic>K<sub>a</sub></italic>/<italic>K<sub>s</sub></italic>&#x2009;&#x003E;&#x2009;1</td>
<td align="center" valign="top">30</td>
<td align="center" valign="top">46</td>
<td align="center" valign="top">30</td>
<td align="center" valign="top">24</td>
</tr>
<tr>
<td align="left" valign="top">Number of GO terms identified in TDGs</td>
<td align="center" valign="top">1901</td>
<td align="center" valign="top">2097</td>
<td align="center" valign="top">1783</td>
<td align="center" valign="top">1,634</td>
</tr>
<tr>
<td align="left" valign="top">Significantly Enriched BP terms in TDGs</td>
<td align="center" valign="top">190</td>
<td align="center" valign="top">230</td>
<td align="center" valign="top">199</td>
<td align="center" valign="top">154</td>
</tr>
<tr>
<td align="left" valign="top">Significantly Enriched BP terms in TDGs with <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1</td>
<td align="center" valign="top">108</td>
<td align="center" valign="top">162</td>
<td align="center" valign="top">127</td>
<td align="center" valign="top">74</td>
</tr>
<tr>
<td align="left" valign="top">Significantly Enriched MF terms in TDGs</td>
<td align="center" valign="top">135</td>
<td align="center" valign="top">165</td>
<td align="center" valign="top">159</td>
<td align="center" valign="top">102</td>
</tr>
<tr>
<td align="left" valign="top">Significantly Enriched CC terms in TDGs</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">15</td>
<td align="center" valign="top">14</td>
<td align="center" valign="top">15</td>
</tr>
<tr>
<td align="left" valign="top">Percent of TDGs expressed (TPM&#x2009;&#x003E;&#x2009;0.5)</td>
<td align="center" valign="top">68.82</td>
<td align="center" valign="top">81.14</td>
<td align="center" valign="top">76.98</td>
<td align="center" valign="top">74.35</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption><p>Distribution of density of tandemly duplicated genes per 1 Mb across each potato genome. Square boxes between rows do not correspond to sequence alignment.</p></caption>
<graphic xlink:href="fpls-13-875202-g001.tif"/>
</fig>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption><p><bold>(A)</bold> Distribution of cluster sizes vs. percent of clusters across four potato genomes and <bold>(B)</bold> Number of intervening genes vs. number of tandem duplicated gene pairs across four potato genomes.</p></caption>
<graphic xlink:href="fpls-13-875202-g002.tif"/>
</fig>
</sec>
<sec id="sec15">
<title>Evolutionary Forces Affecting TDGs</title>
<p>Tandemly duplicated genes account for about 18% of the total non-TE genes. Thus, it would be interesting to gain insights into the evolutionary forces that affect the TDGs. Therefore, we examined the sequence divergence in TDGs of each potato genome by estimating <italic>K</italic><sub>a</sub> (number of substitutions per nonsynonymous site), <italic>K</italic><sub>s</sub> (number of substitutions per synonymous site), and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> ratios, and compared their distributions across the four potato genomes. We observed pronounced peaks at 0.1, between 0.1 and 0.15, as well as between 0.3 and 0.4 for <italic>K</italic><sub>a</sub>, <italic>K</italic><sub>s</sub>, and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> distributions, respectively, for all potato genomes (<xref rid="fig3" ref-type="fig">Figures 3A</xref>&#x2013;<xref rid="fig3" ref-type="fig">C</xref>). Further, all four potato genomes showed a higher <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> values compared to <italic>K</italic><sub>a</sub> values (<xref rid="fig3" ref-type="fig">Figures 3D</xref>,<xref rid="fig3" ref-type="fig">F</xref>). An average of 92.97% TDG pairs showed <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003C;&#x2009;1.0 (i.e., negative or purifying selection), while an average of 6.47% showed a <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> value &#x003E;1.0 (i.e., positive selection). In addition, we observed that the cultivated potato genotype dAg contained the highest number of TDG pairs (916) that were under positive selection (i.e., <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1), while the wild potato genotype M6 contained the least number (345; <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S4</xref>).</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption><p><bold>(A&#x2013;C)</bold> Distribution of <italic>K</italic><sub>a</sub>, <italic>K</italic><sub>s</sub>, and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> for tandemly duplicated genes in DM, M6, RH, and dAg, and sorted into bins of width&#x2009;=&#x2009;0.1. <bold>(D&#x2013;F)</bold> Evolutionary patterns of tandemly duplicated genes in DM, M6, RH, and dAg.</p></caption>
<graphic xlink:href="fpls-13-875202-g003.tif"/>
</fig>
<p>Furthermore, we investigated the functional specificities of the TDGs by an enrichment analysis to answer whether evolutionary forces drive TDGs in potato toward a specific biological function. First, we used the Pfam protein domain information for an enrichment analysis (DEA). A total of 83.42, 91.47, 86.73, and 78.68% of the identified TDGs contain Pfam domains in dAg, DM, M6, and RH, respectively. Further, a total of 917, 805, 706, and 756 unique Pfam domains were identified in TDGs of dAg, DM, M6, and RH, respectively. Across the four potato genomes, TDGs showed a similar distribution of the number of protein domains they harbor. The majority of TDGs contained a single Pfam domain only (<xref ref-type="supplementary-material" rid="SM2">Supplementary Figure S2</xref>). The DEA identified that 46, 69, 60, and 61 protein domains in TDGs of dAg, DM, M6, and RH, respectively, were significantly (FDR&#x2009;&#x003C;&#x2009;0.05 and a minimum number of 10 TDGs/protein domain) over-represented. The most important protein domains that were enriched included NB-ARC, leucine-rich repeat, pathogenesis-related proteins, UDP-glucosyl transferase, glutathione S-transferase, and auxin-responsive protein IPP transferase. Interestingly, all the enriched protein domains were also differentially enriched (<italic>z</italic>-score) across the four potato genomes (<xref rid="tab2" ref-type="table">Table 2</xref> and <xref ref-type="supplementary-material" rid="SM3">Supplementary Figure S3</xref>). Further, a total of 65.21, 66.66, 50, and 39.34% of enriched protein domains were present in positively selected (<italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1) TDGs of dAg, DM, M6, and RH, respectively (<xref rid="fig4" ref-type="fig">Figure 4A</xref>). As alternative approach to identify functional specificities of the TDGs, we performed a GO-term enrichment analysis (GOEA) to identify over-represented gene ontology (GO) terms in TDGs. GOEA identified 190, 230, 199, and 154 statistically significant (FDR&#x2009;&#x003C;&#x2009;0.01) biological processes (BP) in dAg, DM, M6, and RH, respectively (<xref rid="tab2" ref-type="table">Table 2</xref>). Further, a total of 56.84, 70.43, 63.81, and 48.05% of enriched BP were present in positively selected (<italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1) TDGs of dAg, DM, M6, and RH, respectively. The majority of the top 30 BP were associated with defense responses against various biotic conditions (bacteria, fungus, and virus), and stress responses against various abiotic (hypoxia, cadmium, heat, light, and UV-B) stress conditions (<xref rid="fig4" ref-type="fig">Figure 4B</xref>). In line with the domain enrichment, the enriched BP were also differentially enriched (fold enrichment) across four potato genomes (<xref rid="fig4" ref-type="fig">Figure 4B</xref>).</p>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption><p><bold>(A)</bold> Enriched Pfam protein domains in tandemly duplicated genes which are under positive selection (i.e., <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1) in DM, M6, RH, and dAg. <bold>(B)</bold> Top 30 GO biological processes enriched in tandemly duplicated genes which are under positive selection (<italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1) in DM, M6, RH, and dAg.</p></caption>
<graphic xlink:href="fpls-13-875202-g004.tif"/>
</fig>
</sec>
<sec id="sec16">
<title>Expression Divergence Between TDGs</title>
<p>In this study, we examined patterns of expression divergence between TDGs in four potato genomes using publicly available RNA-Seq datasets of the respective potato genomes except for dAg. In the public domain, only one RNA-Seq dataset was available for tAg, the tetraploid ancestor of dAg, but the dataset did not pass our selection criteria after trimming out low-quality reads to include in the analysis. Consequently, the RNA-Seq datasets generated under various stress conditions (drought, salt, heat, and cold) belonging to different potato cultivars were used to estimate expression of dAg (<xref ref-type="supplementary-material" rid="SM1">Supplementary Tables S5</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">S8</xref>). We used log10 transformed transcripts per million (TPM) values obtained from Kallisto (<xref ref-type="bibr" rid="ref5">Bray et al., 2016</xref>) as a proxy for expression levels. In the next step, we classified each TDG as expressed if TPM&#x2009;&#x003E;&#x2009;0.5 in at least one RNA-Seq dataset, otherwise classified as unexpressed. Based on this criterion, we found that 68.82, 81.14, 76.98, and 74.35% of TDGs were expressed in dAg, DM, M6, and RH, respectively (<xref rid="tab2" ref-type="table">Table 2</xref>). Across all potato genomes, TDGs showed higher expression specificities than non-TDGs (<xref rid="fig5" ref-type="fig">Figure 5A</xref>), while both expression breadth and levels were lower than non-TDGs (<xref rid="fig5" ref-type="fig">Figures 5B</xref>,<xref rid="fig5" ref-type="fig">C</xref>).</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption><p>Expression patterns of tandemly duplicated genes in DM, M6, RH, and dAg. <bold>(A)</bold> Expression Specificity, <bold>(B)</bold> Expression Breadth (%), and <bold>(C)</bold> Expression Level.</p></caption>
<graphic xlink:href="fpls-13-875202-g005.tif"/>
</fig>
<p>For each potato genome, we classified each TDG pair as expressed TDG pair if both gene copies were expressed, otherwise classified as unexpressed TDG pair. Based on this criterion, we found 46.57, 67.24, 64.29, and 69.08% of TDG pairs were classified as &#x201C;expressed TDG pairs&#x201D; in dAg, DM, M6, and RH, respectively (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S9</xref>). Further, for each potato genome, we selected expressed TDG pairs and calculated Pearson&#x2019;s correlation coefficient (r) between expression profiles of both gene copies. As comparison, we calculated r for the same number of randomly selected non-TDG pairs. The TDG pairs of DM, RH, and dAg revealed the same approximate normal distribution of correlation coefficients as the control sample (<xref rid="fig6" ref-type="fig">Figures 6A</xref>,<xref rid="fig6" ref-type="fig">C</xref>,<xref rid="fig6" ref-type="fig">D</xref>), while the TDG pairs of M6 showed for both a distribution of the correlation coefficients that deviated from a normal distribution (<xref rid="fig6" ref-type="fig">Figure 6B</xref>). Further, the 95% quantile of the correlation coefficient <italic>r</italic> of randomly selected gene pairs was used as threshold for determining if the two gene copies of a TDG pair have diverged expression, i.e., if <italic>r</italic>&#x2009;&#x003E;&#x2009;= 95% quantile, then the TDG pair have a similar expression, while <italic>r</italic>&#x2009;&#x003C;&#x2009;95% quantile, then the TDG pair have diverged expression. Based on this criterion, 79.56, 69.95, 74.55, and 73.66% of expressed TDG pairs showed diverged expression in dAg, DM, M6, and RH, respectively (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S9</xref>).</p>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption><p>Distributions of Pearson&#x2019;s correlation coefficient (<italic>r</italic>) between the expression profiles of two copies derived from tandem duplication in <bold>(A)</bold> DM, <bold>(B)</bold> M6, <bold>(C)</bold> RH, and <bold>(D)</bold> dAg. The dashed green line indicates the 95% quantile of <italic>r</italic> distribution of random gene pairs in respective genotypes. <italic>r</italic>&#x2009;&#x003C;&#x2009;95% quantile indicates that the gene pairs have diverged expression, while <italic>r</italic>&#x2009;&#x2265;&#x2009;95% quantile indicates that the gene pairs have conserved expression in respective potato genomes.</p></caption>
<graphic xlink:href="fpls-13-875202-g006.tif"/>
</fig>
</sec>
<sec id="sec17">
<title>Promoter Divergence Between TDGs</title>
<p>As shown above, the TDG pairs exhibited significant transcriptional divergences that prompted us to undertake a systematic investigation of variation present in their promoters. In order to do that we retrieved non-overlapping 1&#x2009;Kb sequence upstream of the transcription start site for each gene of a TDG pair as a putative promoter sequence and measured promoter similarity (Ps). We measured Ps for the same number of randomly selected non-TE gene pairs of respective potato genomes to represent the background level of Ps that is expected to be observed by chance. We found a similar distribution in Ps of tandemly duplicated gene pairs across four potato genomes (<xref rid="fig7" ref-type="fig">Figure 7</xref>). On average, Ps for randomly selected gene pairs was 0.159, 0.152, 0.146, and 0.139%, while Ps for TDG pairs was 0.34, 0.29, 0.25, and 0.31% in dAg, DM, M6, and RH, respectively (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S10</xref>). As mentioned above, 95% quantile in the Ps of randomly selected gene pairs was used to classify the promoter sequences of TDG pairs as either conserved or diverged. Based on this criterion, 66.06, 70.83, 75.61, and 66.59% of TDG pairs showed diverged promoters in dAg, DM, M6, and RH, respectively (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S10</xref>).</p>
<fig position="float" id="fig7">
<label>Figure 7</label>
<caption><p>Distributions of promoter similarities (Ps) between TDG pairs in <bold>(A)</bold> DM, <bold>(B)</bold> M6, <bold>(C)</bold> RH, and <bold>(D)</bold> dAg. The dashed green line indicates 95% quantile of Ps distribution of random gene pairs. Promoter pairs with Ps&#x2009;&#x003C;&#x2009;95% quantile indicates that the promoters of the gene pairs shown to diverged, while Ps&#x2009;&#x2265;&#x2009;95% quantile indicates that the promoters of the gene pair shown to conserved in respective potato genomes.</p></caption>
<graphic xlink:href="fpls-13-875202-g007.tif"/>
</fig>
</sec>
<sec id="sec18">
<title>Correlation of Expression, Promoter and Protein Divergence in TDGs</title>
<p>As shown above, expression, age of TDG pairs measured as <italic>K</italic><sub>s</sub>, and their associated promoters exhibited significant similarities as well as divergences in TDGs. It would be interesting to know the correlations among them. Therefore, to test whether the divergence of promoter similarities correlates with the expression divergence, we computed Pearson correlation coefficients (<italic>r</italic>) between the expressed TDG pairs against their promoter similarities. The correlation coefficients were low but significantly positive and ranged for the four genomes from 0.09 (value of <italic>p</italic>&#x2009;=&#x2009;6.73&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;10</sup>) for dAg to 0.19 (value of <italic>p</italic>&#x2009;=&#x2009;2.54&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;39</sup>) for RH (<xref rid="fig8" ref-type="fig">Figure 8A</xref>). Further, to test whether promoter divergence correlates with coding sequence divergence measured as <italic>K</italic><sub>s</sub>, we computed Pearson correlation coefficients between <italic>K</italic><sub>s</sub> of TDG pairs and their respective promoter similarities. The correlation coefficients were low but significantly negative and ranged from &#x2212;0.28 (value of <italic>p</italic>&#x2009;=&#x2009;1.93&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;198</sup>) for DM to &#x2212;0.23 (value of <italic>p</italic>&#x2009;=&#x2009;7.47&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;78</sup>) for M6 (<xref rid="fig8" ref-type="fig">Figure 8B</xref>). However, the distribution of promoter divergence across <italic>K</italic><sub>s</sub> (<xref rid="fig8" ref-type="fig">Figure 8B</xref>) suggests a continuous expression divergence has occurred over the evolutionary time (<italic>K</italic><sub>s</sub>) between TDG pairs. To verify this, we computed Pearson correlation coefficients between <italic>K</italic><sub>s</sub> of TDG pairs and their expression. Similarly, a continuous expression divergence over the time between TDG pairs occurred with a significantly negative correlation ranging from &#x2212;0.17 (value of <italic>p</italic>&#x2009;=&#x2009;4.97&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;30</sup>) for M6 to &#x2212;0.07 (value of <italic>p</italic>&#x2009;=&#x2009;6.72&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;07</sup>) for RH (<xref rid="fig8" ref-type="fig">Figure 8C</xref>). Further, we computed Pearson correlation coefficients between <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> ratios of TDG pairs and their expression to reveal the type of selection that caused the divergence in expression between the duplicated genes of a TDG pair. The results indicated that the divergence in gene expression between duplicated genes across TDG pairs was due to purifying selection (<xref rid="fig8" ref-type="fig">Figure 8D</xref>). Only for M6 a positive correlation of 0.08 (value of <italic>p</italic>&#x2009;=&#x2009;5.5&#x2009;&#x00D7;&#x2009;10<sup>&#x2212;8</sup>) was observed.</p>
<fig position="float" id="fig8">
<label>Figure 8</label>
<caption><p><bold>(A)</bold> Promoter associated with gene expression correlation (<italic>r</italic>). Gene expression correlation (Pearson <italic>r</italic>) on <italic>x</italic>-axis was plotted against promoter similarity within TDGs in DM, M6, RH, and dAg. <bold>(B)</bold> Promoter similarity coupled to protein divergence correlations. Protein divergence time <italic>K</italic><sub>s</sub> on <italic>x</italic>-axis plotted against promoter similarity within TDGs in DM, M6, RH, and dAg. <bold>(C)</bold> Protein divergence time <italic>K</italic><sub>s</sub> uncoupled to gene expression (Pearson <italic>r</italic>). Gene expression correlation (Pearson <italic>r</italic>) on <italic>y</italic>-axis was plotted against protein divergence time <italic>K</italic><sub>s</sub> within TDGs in DM, M6, RH, and dAg. <bold>(D)</bold> <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> uncoupled to gene expression (Pearson <italic>r</italic>). Gene expression correlation (Pearson <italic>r</italic>) on <italic>y</italic>-axis was plotted against <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> within TDGs in DM, M6, RH, and dAg.</p></caption>
<graphic xlink:href="fpls-13-875202-g008.tif"/>
</fig>
</sec>
<sec id="sec19">
<title>Core, Shared, and Private TDG Clusters</title>
<p>A total of 7,450 TDG clusters were identified across four potato genomes. To determine if TDG clusters were shared across the four potato genomes, we used the orthology information that linked the non-TE gene models of all four potato genomes and categorized them into core, shared, and private (or lineage-specific). Based on this categorization, on average, 25.02, 29.94, and 45.03% of all TDG clusters were private, core, and shared clusters, respectively, across the four potato genotypes (<xref rid="fig9" ref-type="fig">Figure 9A</xref>; <xref rid="tab3" ref-type="table">Table 3</xref>). In addition, the non-cultivated potato genotype DM contained the highest proportion of shared clusters (51.96%), while the cultivated potato genotype dAg contained the lowest proportion of shared clusters (40.57%). Conversely, the cultivated genotype dAg contained the highest proportion of private clusters (32.92%), while the non-cultivated potato genotype DM contained the lowest proportion of private clusters (18.37%; <xref rid="tab3" ref-type="table">Table 3</xref>). An average of 52.04% of Pfam protein domains enriched in all TDG clusters was present in private clusters. In addition, the private clusters of the cultivated potato genotype dAg showed with the highest proportion of enriched Pfam protein domains (about 70%), while the private clusters of the non-cultivated potato genotype RH revealed the lowest proportion of enriched Pfam protein domains (about 41%). Furthermore, an average of 62.24% of Pfam protein domains which were enriched within the private clusters were present in positively selected TDGs (<italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1). In addition, the cultivated potato genotype dAg contained a high proportion of Pfam protein domains (about 66%) which were enriched within the private clusters and that were present in positively selected TDGs (<italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1). In contrast, the wild potato genotype contained the lowest proportion of Pfam protein domains (about 59%) which were enriched within the private clusters and that were present in positively selected TDGs (<italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1; <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S11</xref>). In general, private clusters showed low-expression specificities but higher expression breadth compared to shared and core clusters. The private clusters were shown to have lower <italic>K</italic><sub>a</sub> and <italic>K</italic><sub>s</sub> values but higher <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> values compared to shared and core clusters (<xref rid="fig9" ref-type="fig">Figure 9C</xref>). Further, we observed that the majority of private clusters localized toward the centromere, while the majority of core and shared clusters tended to localize toward the ends of chromosome-arms (<xref ref-type="supplementary-material" rid="SM4">Supplementary Figure S4</xref>).</p>
<fig position="float" id="fig9">
<label>Figure 9</label>
<caption><p><bold>(A)</bold> Distribution of Private (Red), Shared (Black), and Core (Blue). Tandemly duplicated clusters across potato genomes <italic>via</italic> orthogroups. <bold>(B)</bold> Expression patterns of private, shared, and core TDG clusters. <bold>(C)</bold> Evolutionary patterns of private, shared, and core TDG clusters.</p></caption>
<graphic xlink:href="fpls-13-875202-g009.tif"/>
</fig>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption><p>Number of TDG shared across potato genotypes.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Genotype</th>
<th align="center" valign="top"># TDG clusters</th>
<th align="center" valign="top"># Private</th>
<th align="center" valign="top">Private (%)</th>
<th align="center" valign="top"># Core</th>
<th align="center" valign="top">Core (%)</th>
<th align="center" valign="top"># Shared</th>
<th align="center" valign="top">Shared (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">dAg</td>
<td align="center" valign="top">2,090</td>
<td align="center" valign="top">688</td>
<td align="center" valign="top">32.92</td>
<td align="center" valign="top">554</td>
<td align="center" valign="top">26.51</td>
<td align="center" valign="top">848</td>
<td align="center" valign="top">40.57</td>
</tr>
<tr>
<td align="left" valign="top">DM</td>
<td align="center" valign="top">1,867</td>
<td align="center" valign="top">343</td>
<td align="center" valign="top">18.37</td>
<td align="center" valign="top">554</td>
<td align="center" valign="top">29.67</td>
<td align="center" valign="top">970</td>
<td align="center" valign="top">51.96</td>
</tr>
<tr>
<td align="left" valign="top">M6</td>
<td align="center" valign="top">1,661</td>
<td align="center" valign="top">350</td>
<td align="center" valign="top">21.07</td>
<td align="center" valign="top">554</td>
<td align="center" valign="top">33.35</td>
<td align="center" valign="top">757</td>
<td align="center" valign="top">45.57</td>
</tr>
<tr>
<td align="left" valign="top">RH</td>
<td align="center" valign="top">1832</td>
<td align="center" valign="top">508</td>
<td align="center" valign="top">27.73</td>
<td align="center" valign="top">554</td>
<td align="center" valign="top">30.24</td>
<td align="center" valign="top">770</td>
<td align="center" valign="top">42.03</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="sec20" sec-type="discussions">
<title>Discussion</title>
<sec id="sec21">
<title>Genome-Wide Identification of TDGs in Potatoes</title>
<p>Tandem duplications are widespread in plant genomes and contribute significantly to the evolution of genomes. By using the available well-annotated multiple <italic>de novo</italic> genome assemblies of potatoes, we observed that the TDGs in potato were dispersed throughout the genome with similar distribution across four potato genomes. TDGs accounted for about 18% of all non-TE genes in potatoes. This number is considerably higher than in rice (about 15.1%; <xref ref-type="bibr" rid="ref31">Jiang et al., 2013</xref>), maize (average of 10.3%; <xref ref-type="bibr" rid="ref36">Kono et al., 2018</xref>), and pear (about 11.1%; <xref ref-type="bibr" rid="ref51">Qiao et al., 2018</xref>). The differences in proportions of TDGs among potatoes, rice, maize, and pear might be generated by species-specific gene duplications as observed across 141 plant species (<xref ref-type="bibr" rid="ref50">Qiao et al., 2019</xref>). The distribution and chromosomal localization of TDGs of potato observed in our study are similar to the TDGs of rice (<xref ref-type="bibr" rid="ref31">Jiang et al., 2013</xref>) and maize (<xref ref-type="bibr" rid="ref36">Kono et al., 2018</xref>). Further, our results indicate that there is variation among potato genotypes for the content of TDGs in the genome (<xref rid="tab2" ref-type="table">Table 2</xref>). The higher number of TDGs in dAg compared to the other three genotypes is likely due to the availability of a more accurate and larger genome assembly. For example, more than 20% of annotated non-TE genes of M6 were present on the unknown chromosome, and hence these genes were excluded from the analysis which might in turn lead to the identification of the lowest number of TDGs among the four potato genomes. However, further studies with larger number of genotypes are required in order to link the observed variation in the content of TDGs in the genome to the history of the examined genetic material.</p>
</sec>
<sec id="sec22">
<title>Differential Enrichment of Functional Specificities of TDGs</title>
<p>Gene duplication is a mechanism that creates functional innovation and novelty in the genome. Here, we explored the relationships between TDGs and functional specificities across cultivated and wild genotypes of potatoes. Protein domain enrichment revealed enrichment of several important protein domains related to genes involved in disease resistance [NB-ARC (PF00931) and leucine-rich repeat (PF00560; <xref ref-type="bibr" rid="ref32">Jupe et al., 2012</xref>; <xref ref-type="bibr" rid="ref49">Prakash et al., 2020</xref>); pathogenesis-related proteins (PF00407; <xref ref-type="bibr" rid="ref37">Lakhssassi et al., 2020</xref>)], stress-responsive [UDP-glucosyl transferase (PF00201; <xref ref-type="bibr" rid="ref53">Rehman et al., 2018</xref>), glutathione S-transferase (PF02798; <xref ref-type="bibr" rid="ref28">Islam et al., 2018</xref>)], auxin-responsive protein (PF02519; <xref ref-type="bibr" rid="ref29">Jain et al., 2006</xref>), and various biosynthetic pathways [IPP transferase (PF01715; <xref ref-type="bibr" rid="ref41">Lindner et al., 2014</xref>)] in TDGs across four potato genomes. Our results highlighted that the potato inhibitor 1 family protein domain (PF00280) containing genes were enriched only in the wild potato genotype (i.e., M6) and were shown to be under positive selection, i.e., <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1 (<xref rid="fig4" ref-type="fig">Figure 4A</xref>). The potato inhibitor 1 family protein domain containing genes are naturally occurring plant serine proteinase inhibitors. They act as both endogenous and defense-related plant regulators in potato under wounding and nematode infection (<xref ref-type="bibr" rid="ref63">Turr&#x00E0; et al., 2009</xref>). In line with protein domain enrichment, GO enrichment also revealed that these TDGs were involved in biological processes related to defense responses against various pathogens, such as bacteria, fungi, and virus, stress responses against various abiotic stress conditions (such as light, auxin, cadmium, heat, and UV), and biosynthetic pathways (lignin, saponin, phenylpropanoid, di-terpenoid, flavonoid, glucosinolate, and indole). Moreover, both protein domains and GO processes were differentially enriched across four potato genomes and TDGs encoding these functional specificities were under positive selection, i.e., <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1 (<xref rid="fig4" ref-type="fig">Figures 4A</xref>,<xref rid="fig4" ref-type="fig">B</xref>). These results suggested that the bias in functional specificities coupled with positive selection might play an important role in the retention of TDGs in potatoes (<xref ref-type="bibr" rid="ref60">Shiu et al., 2006</xref>; <xref ref-type="bibr" rid="ref54">Ren et al., 2014</xref>). This finding is consistent with a previous study that found that the retention of TDGs favors genes involved in certain important functions to maintain the fitness of the organism (<xref ref-type="bibr" rid="ref3">Blanc and Wolfe, 2004</xref>; <xref ref-type="bibr" rid="ref11">Edger and Pires, 2009</xref>).</p>
</sec>
<sec id="sec23">
<title>Rapid Sequence, Expression, and Regulatory Divergences Among TDG Pairs</title>
<p>Potato underwent at least two rounds of genome duplication, 185 and 67 million years ago (<xref ref-type="bibr" rid="ref48">Potato Genome Sequencing Consortium et al., 2011</xref>), and retained 6,078, 5,817, 4,748, and 5,018 TDGs for dAg, DM, M6, and RH, respectively. Our study highlighted a number of striking patterns in sequence, expression, and regulatory divergences between gene copies of TDG pairs across four potato genomes (<xref rid="fig3" ref-type="fig">Figures 3</xref>, <xref rid="fig5" ref-type="fig">5</xref>&#x2013;<xref rid="fig8" ref-type="fig">8</xref>). Based on these patterns, we propose and distinguish multiple models such as sub-functionalization (<xref ref-type="bibr" rid="ref17">Force et al., 1999</xref>), genetic redundancy (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>), and neo-functionalization (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>) that may contribute to the retention of TDGs in potato genomes.</p>
<sec id="sec24">
<title>Sub-Functionalization</title>
<p>In general, our results indicate that the TDGs were expressed in all potato genotypes in a lower number of samples but with higher expression specificities than non-TDGs in all potato genotypes (<xref rid="fig5" ref-type="fig">Figures 5A</xref>,<xref rid="fig5" ref-type="fig">B</xref>) and this in turn indicates that the duplicated genes functions in specific tissues. In addition, we found that an average of 74.43% of expressed TDG pairs showed divergence in expression (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>), regardless of the age of duplication, across four potato genotypes. As we already excluded annotated partial or pseudogenes from the dataset, the divergence in expression may not be due to pseudogenization. The expression divergence is consistent with expression divergence between duplicated gene copies of TDG pairs of Arabidopsis thaliana (<xref ref-type="bibr" rid="ref22">Haberer et al., 2004</xref>), <italic>Glycine</italic> max (<xref ref-type="bibr" rid="ref57">Roulin et al., 2013</xref>), and the D-genome of <italic>Gossypium raimondii</italic> (<xref ref-type="bibr" rid="ref55">Renny-Byfield et al., 2014</xref>). We also found that an average of 92.3% of expressed TDG pairs which showed divergence in expression have <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003C;&#x2009;1.0 (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>), indicative of a purifying selective pressure at the nucleotide level across four potato genotypes. In addition, either substantially a weak (for M6) or negative correlations (for dAg, DM, and RH) were observed between expression divergence and <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub> ratios (<xref rid="fig8" ref-type="fig">Figure 8D</xref>), suggesting that sub-functionalization of duplicated genes across tissues has been occurred to retain the duplicated gene copies. These results are consistent with the retention of duplicated genes through sub-functionalization in <italic>Glycine max</italic> (<xref ref-type="bibr" rid="ref57">Roulin et al., 2013</xref>). Furthermore, the vast majority of the above retained TDG pairs (an average of 84.33%) had identical annotation in terms of protein domains between gene copies of respective TDG pairs (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>). These results together indicated that the retention of duplicated genes occurred through sub-functionalization, where the partitioning of an ancestral gene into daughter genes across tissues implies that both daughter genes must remain functionally (<xref ref-type="bibr" rid="ref17">Force et al., 1999</xref>; <xref ref-type="bibr" rid="ref40">Lynch and Force, 2000</xref>). The results also demonstrated that the sub-functionalization might have been established after polyploidization in potato, and it was maintained over time (<xref rid="fig8" ref-type="fig">Figure 8C</xref>), as observed in <italic>Glycine max</italic> (<xref ref-type="bibr" rid="ref57">Roulin et al., 2013</xref>) and cotton (<xref ref-type="bibr" rid="ref001">Chaudhary et al., 2009</xref>).</p>
<p>The divergence in expression and there with sub-functionalization could be due to divergence in promoter sequences of the respective duplicated gene copies of TDG pairs (<xref ref-type="bibr" rid="ref33">Katju and Lynch, 2003</xref>). In line with this explanation, we observed a divergence in promoter sequences of an average of 50.95% of expressed TDG pairs (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>). However, only weak correlations between divergence of promoter sequences and expression levels were observed across four potato genotypes (<xref rid="fig8" ref-type="fig">Figure 8A</xref>). These results are consistent with a previous study in <italic>Arabidopsis thaliana</italic> (<xref ref-type="bibr" rid="ref22">Haberer et al., 2004</xref>) and suggested that even small changes in the promoter sequences could be sufficient for sub- or neo-functionalization. Our results also indicated that the expression of TDGs might be regulated by trans-acting factors (<xref ref-type="bibr" rid="ref74">Yvert et al., 2003</xref>).</p>
</sec>
<sec id="sec25">
<title>Genetic Redundancy</title>
<p>Despite the overall pattern of expression divergence between duplicated gene copies of TDG pairs, for 25.6% of the expressed TDG pairs, a strong similarity in the expression profiles was observed. Of these TDG pairs, the vast majority (an average of 87.13%) was under purifying selection across four potato genotypes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>). Furthermore, with an average of 86.5% of the above retained expressed TDG pairs had an identical annotation in terms of protein domains between duplicated genes of TDG pairs (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>). These results together suggested that the duplicated genes of similarly expressed TDG pairs might have been retained through selection for genetic redundancy that may be beneficial in a way that is similar to a fail-safe in engineered systems (<xref ref-type="bibr" rid="ref23">Hanada et al., 2009</xref>; <xref ref-type="bibr" rid="ref75">Zhang, 2012</xref>; <xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>). Alternatively, these TDG pairs might have been retained simply because there has been insufficient time for one copy to be removed or mutated or because they are evolving close to neutrally (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>).</p>
</sec>
<sec id="sec26">
<title>Neo-Functionalization</title>
<p>Protein domain analysis performed on TDGs showed that an average of 77.56% of expressed TDG pairs contained identical protein domains between duplicated gene copies of TDG pairs across four potato genotypes. We observed that about an average of ~1% of expressed TDG pairs showed a different protein domain composition with <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1.0 between duplicated gene copies of TDG pairs across four potato genotypes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>). These results suggested that this ~1% of TDG pairs were retained through neo-functionalization, where both duplicate gene copies were retained because of a gain of novel functions that contributes to better fitness post duplication (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>). This observation is consistent with the retention of a small fraction (4%) of duplicated gene pairs through neo-functionalization in <italic>Glycine max</italic> (<xref ref-type="bibr" rid="ref57">Roulin et al., 2013</xref>). Furthermore, our results also highlighted that an average of 75.91% of the retained TDG pairs showed divergence in expression (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S12</xref>). Overall, these results indicated that the divergence in expression, different protein functions, and positive selective pressure combinedly accounted for the neo-functionalization of those average of ~1% of TDG pairs across potato genotypes. Furthermore, we found a total of 27 enriched protein domains were present in the neo-functionalized TDG pairs across four potato genotypes and these enriched protein domains were mainly involved in important biological processes such as disease resistance (NB-ARC: PF00931, leucine-rich repeat: PF13855 and PF00560, and Rx N-terminal domain: PF18052; <xref ref-type="bibr" rid="ref32">Jupe et al., 2012</xref>; <xref ref-type="bibr" rid="ref49">Prakash et al., 2020</xref>); self-incompatibility (S-locus glycoprotein domain: PF00954; <xref ref-type="bibr" rid="ref68">Xing et al., 2013</xref>); seedling development, senescence and pathogen resistance (F-box domain: PF00646; <xref ref-type="bibr" rid="ref69">Xu et al., 2009</xref>; <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S13</xref>).</p>
</sec>
</sec>
<sec id="sec27">
<title>Private TDG Clusters Across Four Potato Genotypes</title>
<p>Based on the orthology information, we found a significant proportion (an average of 25.02% of all TDG clusters) of private or lineage-specific TDG clusters across four potato genotypes (<xref rid="fig9" ref-type="fig">Figure 9A</xref>; <xref rid="tab3" ref-type="table">Table 3</xref>). The majority of them localized in pericentromeric regions which was not observed for core and shared clusters (<xref ref-type="supplementary-material" rid="SM4">Supplementary Figure S4</xref>). The reason for this observation might be the same that is responsible for an over-representation of presence absence variation (PAV) genes in <italic>Arabidopsis thaliana</italic> (<xref ref-type="bibr" rid="ref62">Tan et al., 2012</xref>) in pericentromeric regions. The low extent of recombination in those regions of the genome might prevent the spread of present TDGs in a population, and thus the TDGs remains private. Our results highlighted that the tandem duplication generates a significantly varying proportion of private clusters across four potato genomes (<xref rid="tab3" ref-type="table">Table 3</xref>). In addition, we found that the cultivated genotype dAg contained a high proportion of enriched Pfam protein domains which are present in positively selected TDGs of private clusters compared to non-cultivated as well wild potato genotypes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S11</xref>). This observation might be due to breeder&#x2019;s selection to combine positive alleles for many traits. These results may also indicate that the tandem duplication generates lineage-specific TDGs with functional bias between evolutionarily closed species, such as the four potato genotypes, which is similar to that of generation of lineage-specific TDGs with functional bias between evolutionarily distant plant species (<xref ref-type="bibr" rid="ref24">Hanada et al., 2008</xref>).</p>
<p>In general, our results highlight that the private TDG clusters showed a lower expression specificity and higher expression breadth compare to the shared and core clusters (<xref rid="fig9" ref-type="fig">Figure 9B</xref>). Our observation indicates that the private clusters were involved in tissue-specific functional specificities. This result is in contrast to results of legume species (<xref ref-type="bibr" rid="ref70">Xu et al., 2018</xref>) where private TDGs showed higher expression specificity and lower expression breadth. The reason for that remains elusive. Furthermore, an average of 30.99% of private TDG pairs showed divergence in expression and have <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003C;&#x2009;1.0 (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S14</xref>) indicating purifying selection at the nucleotide level which in turn suggests that sub-functionalization of duplicated genes across tissues has been occurred to retain the duplicated genes across four potato genotypes (<xref ref-type="bibr" rid="ref17">Force et al., 1999</xref>; <xref ref-type="bibr" rid="ref40">Lynch and Force, 2000</xref>). In addition, a vast majority of these retained private TDG pairs (an average of 84.29%) had an identical annotation in terms of protein domains between duplicated gene copies of respective TDG pairs (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S14</xref>). These results together reinforced that the retention of a majority of TDGs of private clusters were occurred through sub-functionalization.</p>
<p>In addition, an average of 19.05% of private TDG pairs showed similarity in expression profiles, of which a majority of them (an average of 61.6%) are under purifying selective pressure (i.e., <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003C;&#x2009;1.0), and a vast majority of them (an average of 87.62%) contained identical annotation in terms of protein domains between duplicated genes of respective private TDG pairs, across four potato genotypes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S14</xref>). These results indicated that these private TDG pairs might have been retained through selection for genetic redundancy that may be beneficial in a way that is similar to fail-safe in an engineered system (<xref ref-type="bibr" rid="ref23">Hanada et al., 2009</xref>; <xref ref-type="bibr" rid="ref75">Zhang, 2012</xref>; <xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>). Alternatively, these private TDG pairs might have been retained simply because there has been insufficient time for one copy to be removed, because they are evolving relatively neutrally (<xref ref-type="bibr" rid="ref46">Panchy et al., 2016</xref>).</p>
<p>We also found that an average of 0.97% only of private TDG pairs have different annotation in terms of protein domain composition with <italic>K</italic><sub>a</sub>/<italic>K</italic><sub>s</sub>&#x2009;&#x003E;&#x2009;1.0 between gene copies of respective TDG pairs across four potato genotypes (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S14</xref>). These results indicate that these 0.97% of private TDG pairs might have been retained through neo-functionalization (<xref ref-type="bibr" rid="ref45">Ohno, 1970</xref>). In addition, a total of eight enriched Pfam protein domains are present in these retained 0.96% of private TDG pairs in all potato genotypes and these enriched protein domains were mainly involved in important biological processes such as disease resistance (NB-ARC domain: PF00931; Leucine rich repeat: PF13855; <xref ref-type="bibr" rid="ref32">Jupe et al., 2012</xref>; <xref ref-type="bibr" rid="ref49">Prakash et al., 2020</xref>) and self-incompatibility (S-locus glycoprotein domain: PF00954; <xref ref-type="bibr" rid="ref68">Xing et al., 2013</xref>; <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S15</xref>).</p>
</sec>
<sec id="sec28">
<title>Lineage-Specific Expansion of Gene Families and Species Divergences</title>
<p>Our results indicated that the tandem duplication contributed to lineage-specific expansion of several gene families across potato genotypes. For example, NBS-LRR, Cytochrome P450, UDP-glucosyl transferase, and 2OG-Fe (II) oxygenase gene families were differentially expanded by tandem duplication across potato genotypes (<xref rid="fig4" ref-type="fig">Figure 4A</xref>). Furthermore, the GO enrichment revealed a functional bias of TDGs across the four potato genotypes (<xref rid="fig4" ref-type="fig">Figure 4B</xref>). This is supported by recent studies of specific gene families in potato (<xref ref-type="bibr" rid="ref25">Herath and Verchot, 2020</xref>; <xref ref-type="bibr" rid="ref42">Liu et al., 2020</xref>; <xref ref-type="bibr" rid="ref73">Yang et al., 2020</xref>; <xref ref-type="bibr" rid="ref71">Xuanyuan et al., 2022</xref>) and provided an important source for genetic diversity in plants for adaptive evolution against various environmental stimuli. These results are similar to a previous study conducted on two maize genotypes (such as B73 and PH207) where more than 49% of B73&#x2019;s and 40% of PH207&#x2019;s TDGs were lineage-specific (<xref ref-type="bibr" rid="ref36">Kono et al., 2018</xref>). Furthermore, the importance of lineage-specific expansion of TDGs was also studied in <italic>A. thaliana</italic> against various abiotic stress stimuli and found a strong correlation between tandem duplication and abiotic stress conditions (<xref ref-type="bibr" rid="ref24">Hanada et al., 2008</xref>). Thus, the lineage-specific expansion of gene families by tandem duplication coupled with functional bias might significantly contribute to potato&#x2019;s genotypic diversity. However, to understand their effect on phenotypic characters requires further research.</p>
</sec>
</sec>
<sec id="sec29" sec-type="conclusions">
<title>Conclusion</title>
<p>By investigating the divergence in sequence, functional and transcriptional features of TDGs across four diploid potato genomes, we found that after at least two rounds of genome duplication, a large proportion of TDGs were retained through sub-functionalization. Sub-functionalization, by keeping both copies of the same gene, may pave an intermediate step to neo-functionalization for some genes, which is supported by a very small fraction of neo-functionalized duplicated TDGs in potatoes. In addition, TDGs contributed to lineage-specific expansion of several gene families for adaptive changes. These results show that evolution of functions and fates of genes after tandem duplication is a complex process which drives the evolution of gene duplication in association with expression, as well as the duplicated and/or retention of genes with specific functions. In addition, we found variation within TDGs among cultivated, non-cultivated and wild potato genotypes in terms of bias in functional specificities, proportion of lineage-specific clusters, diverged expression and promoter similarities.</p>
</sec>
<sec id="sec30" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link xlink:href="https://www.figshare.com/s/7dee6e184ab4cc666976" ext-link-type="uri">https://www.figshare.com/s/7dee6e184ab4cc666976</ext-link>.</p>
</sec>
<sec id="sec31">
<title>Author Contributions</title>
<p>VSB conceived, designed, performed the experiments and data analysis, and wrote the manuscript. BS contributed to data analysis and manuscript writing. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec34" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack>
<p>The authors acknowledge the computational infrastructure and support provided by the Center for Information and Media Technology at Heinrich Heine University D&#x00FC;sseldorf, and the German Network for Bioinformatics Infrastructure (de.NBI, <ext-link xlink:href="https://www.denbi.de/" ext-link-type="uri">https://www.denbi.de/</ext-link>) that contributed to the research results reported within this study.</p>
</ack>
<sec id="sec33" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/fpls.2022.875202/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/fpls.2022.875202/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.ZIP" id="SM1" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"><label>Supplementary Figure S1</label><caption><p>Distribution of number of exons within TDGs.</p></caption></supplementary-material>
<supplementary-material xlink:href="Data_Sheet_1.ZIP" id="SM2" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"><label>Supplementary Figure S2</label><caption><p>Distribution of number of Pfam domains in TDGs.</p></caption></supplementary-material>
<supplementary-material xlink:href="Data_Sheet_1.ZIP" id="SM3" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"><label>Supplementary Figure S3</label><caption><p>Enrichment of Pfam protein domains in all TDGs.</p></caption></supplementary-material>
<supplementary-material xlink:href="Data_Sheet_1.ZIP" id="SM4" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"><label>Supplementary Figure S4</label><caption><p>Distribution of core, shared, and private tandemly duplicated gene clusters across the genomes of <bold>(A)</bold> DM, <bold>(B)</bold> M6, <bold>(C)</bold> RH, and <bold>(D)</bold> dAg.</p></caption></supplementary-material>
<supplementary-material xlink:href="Data_Sheet_1.ZIP" id="SM5" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink"><label>Supplementary Figure S5</label><caption><p>Distribution of density of tandemly duplicated genes per 1.5Mb across each potato genome. Square boxes between rows do not correspond to sequence alignment</p></caption></supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Gish</surname> <given-names>W.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <name><surname>Myers</surname> <given-names>E. W.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Basic local alignment search tool</article-title>. <source>J. Mol. Biol.</source> <volume>215</volume>, <fpage>403</fpage>&#x2013;<lpage>410</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id>, PMID: <pub-id pub-id-type="pmid">2231712</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baker</surname> <given-names>C. R.</given-names></name> <name><surname>Hanson-Smith</surname> <given-names>V.</given-names></name> <name><surname>Johnson</surname> <given-names>A. D.</given-names></name></person-group> (<year>2013</year>). <article-title>Following gene duplication, paralog interference constrains transcriptional circuit evolution</article-title>. <source>Science</source> <volume>342</volume>, <fpage>104</fpage>&#x2013;<lpage>108</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.1240810</pub-id>, PMID: <pub-id pub-id-type="pmid">24092741</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blanc</surname> <given-names>G.</given-names></name> <name><surname>Wolfe</surname> <given-names>K. H.</given-names></name></person-group> (<year>2004</year>). <article-title>Functional divergence of duplicated genes formed by polyploidy during Arabidopsis evolution</article-title>. <source>Plant Cell</source> <volume>16</volume>, <fpage>1679</fpage>&#x2013;<lpage>1691</lpage>. doi: <pub-id pub-id-type="doi">10.1105/tpc.021410</pub-id>, PMID: <pub-id pub-id-type="pmid">15208398</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bolger</surname> <given-names>A. M.</given-names></name> <name><surname>Lohse</surname> <given-names>M.</given-names></name> <name><surname>Usadel</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>Trimmomatic: a flexible trimmer for Illumina sequence data</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>2114</fpage>&#x2013;<lpage>2120</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btu170</pub-id>, PMID: <pub-id pub-id-type="pmid">24695404</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bray</surname> <given-names>N. L.</given-names></name> <name><surname>Pimentel</surname> <given-names>H.</given-names></name> <name><surname>Melsted</surname> <given-names>P.</given-names></name> <name><surname>Pachter</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Near-optimal probabilistic RNA-seq quantification</article-title>. <source>Nat. Biotechnol.</source> <volume>34</volume>, <fpage>525</fpage>&#x2013;<lpage>527</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nbt.3519</pub-id>, PMID: <pub-id pub-id-type="pmid">27043002</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cannon</surname> <given-names>S. B.</given-names></name> <name><surname>Mitra</surname> <given-names>A.</given-names></name> <name><surname>Baumgarten</surname> <given-names>A.</given-names></name> <name><surname>Young</surname> <given-names>N. D.</given-names></name> <name><surname>May</surname> <given-names>G.</given-names></name></person-group> (<year>2004</year>). <article-title>The roles of segmental and tandem gene duplication in the evolution of large gene families in <italic>Arabidopsis thaliana</italic></article-title>. <source>BMC Plant Biol.</source> <volume>4</volume>:<fpage>10</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1471-2229-4-10</pub-id>, PMID: <pub-id pub-id-type="pmid">15171794</pub-id></citation></ref>
<ref id="ref001"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chaudhary</surname> <given-names>B.</given-names></name> <name><surname>Flagel</surname> <given-names>L.</given-names></name> <name><surname>Stupar</surname> <given-names>R. M.</given-names></name> <name><surname>Udall</surname> <given-names>J. A.</given-names></name> <name><surname>Verma</surname> <given-names>N.</given-names></name> <name><surname>Springer</surname> <given-names>N. M.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Reciprocal silencing, transcriptional bias and functional divergence of homeologs in polyploid cotton (gossypium)</article-title>. <source>Genetics</source> <volume>182</volume>, <fpage>503</fpage>&#x2013;<lpage>517</lpage>.</citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cusack</surname> <given-names>B. P.</given-names></name> <name><surname>Wolfe</surname> <given-names>K. H.</given-names></name></person-group> (<year>2007</year>). <article-title>When gene marriages don't work out: divorce by subfunctionalization</article-title>. <source>Trends Genet.</source> <volume>23</volume>, <fpage>270</fpage>&#x2013;<lpage>272</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tig.2007.03.010</pub-id>, PMID: <pub-id pub-id-type="pmid">17418444</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Dainat</surname> <given-names>J.</given-names></name> <name><surname>Here&#x00F1;&#x00FA;</surname> <given-names>D.</given-names></name> <collab id="coll1">LucileSol</collab> <collab id="coll2">pascal-git</collab></person-group> (<year>2022</year>). NBISweden/AGAT: AGAT-v0.8.1.</citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Des Marais</surname> <given-names>D. L.</given-names></name> <name><surname>Rausher</surname> <given-names>M. D.</given-names></name></person-group> (<year>2008</year>). <article-title>Escape from adaptive conflict after duplication in an anthocyanin pathway gene</article-title>. <source>Nature</source> <volume>454</volume>, <fpage>762</fpage>&#x2013;<lpage>765</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature07092</pub-id>, PMID: <pub-id pub-id-type="pmid">18594508</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duarte</surname> <given-names>J. M.</given-names></name> <name><surname>Cui</surname> <given-names>L.</given-names></name> <name><surname>Wall</surname> <given-names>P. K.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Leebens-Mack</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>Expression pattern shifts following duplication indicative of subfunctionalization and neofunctionalization in regulatory genes of Arabidopsis</article-title>. <source>Mol. Biol. Evol.</source> <volume>23</volume>, <fpage>469</fpage>&#x2013;<lpage>478</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/msj051</pub-id>, PMID: <pub-id pub-id-type="pmid">16280546</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edger</surname> <given-names>P. P.</given-names></name> <name><surname>Pires</surname> <given-names>J. C.</given-names></name></person-group> (<year>2009</year>). <article-title>Gene and genome duplications: the impact of dosage-sensitivity on the fate of nuclear genes</article-title>. <source>Chromos. Res.</source> <volume>17</volume>, <fpage>699</fpage>&#x2013;<lpage>717</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10577-009-9055-9</pub-id>, PMID: <pub-id pub-id-type="pmid">19802709</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Emms</surname> <given-names>D. M.</given-names></name> <name><surname>Kelly</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>OrthoFinder: phylogenetic orthology inference for comparative genomics</article-title>. <source>Genome Biol.</source> <volume>20</volume>:<fpage>238</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13059-019-1832-y</pub-id>, PMID: <pub-id pub-id-type="pmid">31727128</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll3">FAO</collab></person-group> (<year>2008</year>). Statistical data. Rome.</citation></ref>
<ref id="ref14"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll4">FAO</collab></person-group> (<year>2019</year>). Statistical data. Rome.</citation></ref>
<ref id="ref15"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>R. A.</given-names></name></person-group> (<year>1992</year>). &#x201C;<article-title>Statistical methods for research workers</article-title>,&#x201D; in: <source>Breakthroughs in Statistics. Springer Series in Statistics</source>. eds. <person-group person-group-type="editor"><name><surname>Kotz</surname> <given-names>S.</given-names></name> <name><surname>Johnson</surname> <given-names>N. L.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>).</citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Flagel</surname> <given-names>L. E.</given-names></name> <name><surname>Wendel</surname> <given-names>J. F.</given-names></name></person-group> (<year>2009</year>). <article-title>Gene duplication and evolutionary novelty in plants</article-title>. <source>New Phytol.</source> <volume>183</volume>, <fpage>557</fpage>&#x2013;<lpage>564</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1469-8137.2009.02923.x</pub-id>, PMID: <pub-id pub-id-type="pmid">19555435</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Force</surname> <given-names>A.</given-names></name> <name><surname>Lynch</surname> <given-names>M.</given-names></name> <name><surname>Pickett</surname> <given-names>F. B.</given-names></name> <name><surname>Amores</surname> <given-names>A.</given-names></name> <name><surname>Yan</surname> <given-names>Y. L.</given-names></name> <name><surname>Postlethwait</surname> <given-names>J.</given-names></name></person-group> (<year>1999</year>). <article-title>Preservation of duplicate genes by complementary, degenerative mutations</article-title>. <source>Genetics</source> <volume>151</volume>, <fpage>1531</fpage>&#x2013;<lpage>1545</lpage>. doi: <pub-id pub-id-type="doi">10.1093/genetics/151.4.1531</pub-id>, PMID: <pub-id pub-id-type="pmid">10101175</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Freeling</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Bias in plant gene content following different sorts of duplication: tandem, whole-genome, segmental, or by transposition</article-title>. <source>Annu. Rev. Plant Biol.</source> <volume>60</volume>, <fpage>433</fpage>&#x2013;<lpage>453</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev.arplant.043008.092122</pub-id>, PMID: <pub-id pub-id-type="pmid">19575588</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Freeling</surname> <given-names>M.</given-names></name> <name><surname>Thomas</surname> <given-names>B. C.</given-names></name></person-group> (<year>2006</year>). <article-title>Gene-balanced duplications, like tetraploidy, provide predictable drive to increase morphological complexity</article-title>. <source>Genome Res.</source> <volume>16</volume>, <fpage>805</fpage>&#x2013;<lpage>814</lpage>. doi: <pub-id pub-id-type="doi">10.1101/gr.3681406</pub-id>, PMID: <pub-id pub-id-type="pmid">16818725</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Freire</surname> <given-names>R.</given-names></name> <name><surname>Weisweiler</surname> <given-names>M.</given-names></name> <name><surname>Guerreiro</surname> <given-names>R.</given-names></name> <name><surname>Baig</surname> <given-names>N.</given-names></name> <name><surname>H&#x00FC;tte</surname> <given-names>B.</given-names></name> <name><surname>Obeng-Hinneh</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Chromosome-scale reference genome assembly of a diploid potato clone derived from an elite variety</article-title>. <source>G3</source> <volume>11</volume>:<fpage>jkab330</fpage>. doi: <pub-id pub-id-type="doi">10.1093/g3journal/jkab330</pub-id>, PMID: <pub-id pub-id-type="pmid">34534288</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haas</surname> <given-names>B. J.</given-names></name> <name><surname>Delcher</surname> <given-names>A. L.</given-names></name> <name><surname>Mount</surname> <given-names>S. M.</given-names></name> <name><surname>Wortman</surname> <given-names>J. R.</given-names></name> <name><surname>Smith</surname> <given-names>R. K.</given-names></name> <name><surname>Hannick</surname> <given-names>L. I.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies</article-title>. <source>Nucleic Acids Res.</source> <volume>31</volume>, <fpage>5654</fpage>&#x2013;<lpage>5666</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkg770</pub-id>, PMID: <pub-id pub-id-type="pmid">14500829</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haberer</surname> <given-names>G.</given-names></name> <name><surname>Hindemitt</surname> <given-names>T.</given-names></name> <name><surname>Meyers</surname> <given-names>B. C.</given-names></name> <name><surname>Mayer</surname> <given-names>K. E. X.</given-names></name></person-group> (<year>2004</year>). <article-title>Transcriptional similarities, dissimilarities, and conservation of cis-elements in duplicated genes of Arabidopsis</article-title>. <source>Plant Physiol.</source> <volume>136</volume>, <fpage>3009</fpage>&#x2013;<lpage>3022</lpage>. doi: <pub-id pub-id-type="doi">10.1104/pp.104.046466</pub-id>, PMID: <pub-id pub-id-type="pmid">15489284</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanada</surname> <given-names>K.</given-names></name> <name><surname>Kuromori</surname> <given-names>T.</given-names></name> <name><surname>Myouga</surname> <given-names>F.</given-names></name> <name><surname>Toyoda</surname> <given-names>T.</given-names></name> <name><surname>Li</surname> <given-names>W. H.</given-names></name> <name><surname>Shinozaki</surname> <given-names>K.</given-names></name></person-group> (<year>2009</year>). <article-title>Evolutionary persistence of functional compensation by duplicate genes in Arabidopsis</article-title>. <source>Genome Biol. Evol.</source> <volume>1</volume>, <fpage>409</fpage>&#x2013;<lpage>414</lpage>. doi: <pub-id pub-id-type="doi">10.1093/gbe/evp043</pub-id>, PMID: <pub-id pub-id-type="pmid">20333209</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanada</surname> <given-names>K.</given-names></name> <name><surname>Zou</surname> <given-names>C.</given-names></name> <name><surname>Lehti-Shiu</surname> <given-names>M. D.</given-names></name> <name><surname>Shinozaki</surname> <given-names>K.</given-names></name> <name><surname>Shiu</surname> <given-names>S. H.</given-names></name></person-group> (<year>2008</year>). <article-title>Importance of lineage-specific expansion of plant tandem duplicates in the adaptive response to environmental stimuli</article-title>. <source>Plant Physiol.</source> <volume>148</volume>, <fpage>993</fpage>&#x2013;<lpage>1003</lpage>. doi: <pub-id pub-id-type="doi">10.1104/pp.108.122457</pub-id>, PMID: <pub-id pub-id-type="pmid">18715958</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herath</surname> <given-names>V.</given-names></name> <name><surname>Verchot</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Insight into the bZIP gene family in <italic>Solanum tuberosum</italic>: genome and transcriptome analysis to understand the roles of gene diversification in spatiotemporal gene expression and function</article-title>. <source>Int. J. Mol. Sci.</source> <volume>22</volume>:<fpage>253</fpage>. doi: <pub-id pub-id-type="doi">10.3390/ijms22010253</pub-id>, PMID: <pub-id pub-id-type="pmid">33383823</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoopes</surname> <given-names>G.</given-names></name> <name><surname>Meng</surname> <given-names>X.</given-names></name> <name><surname>Hamilton</surname> <given-names>J. P.</given-names></name> <name><surname>Achakkagari</surname> <given-names>S. R.</given-names></name> <name><surname>de Alves Freitas Guesdes</surname> <given-names>F.</given-names></name> <name><surname>Finkers</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Phased, chromosome-scale genome assemblies of tetraploid potato reveal a complex genome, transcriptome, and predicted proteome landscape underpinning genetic diversity</article-title>. <source>Mol. Plant</source> <volume>15</volume>, <fpage>P520</fpage>&#x2013;<lpage>P536</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.molp.2022.01.003</pub-id>, PMID: <pub-id pub-id-type="pmid">35026436</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Islam</surname> <given-names>M. S.</given-names></name> <name><surname>Choudhury</surname> <given-names>M.</given-names></name> <name><surname>Majlish</surname> <given-names>A. N. K.</given-names></name> <name><surname>Islam</surname> <given-names>T.</given-names></name> <name><surname>Ghosh</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Comprehensive genome-wide analysis of glutathione S-transferase gene family in potato (<italic>Solanum tuberosum</italic> L.) and their expression profiling in various anatomical tissues and perturbation conditions</article-title>. <source>Gene</source> <volume>639</volume>, <fpage>149</fpage>&#x2013;<lpage>162</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.gene.2017.10.007</pub-id>, PMID: <pub-id pub-id-type="pmid">28988961</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jain</surname> <given-names>M.</given-names></name> <name><surname>Tyagi</surname> <given-names>A. K.</given-names></name> <name><surname>Khurana</surname> <given-names>J. P.</given-names></name></person-group> (<year>2006</year>). <article-title>Genome-wide analysis, evolutionary expansion, and expression of early auxin-responsive SAUR gene family in rice (Oryza sativa)</article-title>. <source>Genomics</source> <volume>88</volume>, <fpage>360</fpage>&#x2013;<lpage>371</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ygeno.2006.04.008</pub-id>, PMID: <pub-id pub-id-type="pmid">16707243</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jayakodi</surname> <given-names>M.</given-names></name> <name><surname>Padmarasu</surname> <given-names>S.</given-names></name> <name><surname>Haberer</surname> <given-names>G.</given-names></name> <name><surname>Bonthala</surname> <given-names>V. S.</given-names></name> <name><surname>Gundlach</surname> <given-names>H.</given-names></name> <name><surname>Monat</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>The barley pan-genome reveals the hidden legacy of mutation breeding</article-title>. <source>Nature</source> <volume>588</volume>, <fpage>284</fpage>&#x2013;<lpage>289</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41586-020-2947-8</pub-id>, PMID: <pub-id pub-id-type="pmid">33239781</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>S. Y.</given-names></name> <name><surname>Gonz&#x00E1;lez</surname> <given-names>J. M.</given-names></name> <name><surname>Ramachandran</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Comparative genomic and transcriptomic analysis of tandemly and segmentally duplicated genes in rice</article-title>. <source>PLoS One</source> <volume>8</volume>:<fpage>e63551</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0063551</pub-id>, PMID: <pub-id pub-id-type="pmid">23696832</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jupe</surname> <given-names>F.</given-names></name> <name><surname>Pritchard</surname> <given-names>L.</given-names></name> <name><surname>Etherington</surname> <given-names>G. J.</given-names></name> <name><surname>MacKenzie</surname> <given-names>K.</given-names></name> <name><surname>Cock</surname> <given-names>P. J. A.</given-names></name> <name><surname>Wright</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Identification and localisation of the NB-LRR gene family within the potato genome</article-title>. <source>BMC Genomics</source> <volume>13</volume>:<fpage>75</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1471-2164-13-75</pub-id>, PMID: <pub-id pub-id-type="pmid">22336098</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katju</surname> <given-names>V.</given-names></name> <name><surname>Lynch</surname> <given-names>M.</given-names></name></person-group> (<year>2003</year>). <article-title>The structure and early evolution of recently arisen gene duplicates in the <italic>Caenorhabditis elegans</italic> genome</article-title>. <source>Genetics</source> <volume>165</volume>, <fpage>1793</fpage>&#x2013;<lpage>1803</lpage>. doi: <pub-id pub-id-type="doi">10.1093/genetics/165.4.1793</pub-id>, PMID: <pub-id pub-id-type="pmid">14704166</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katoh</surname> <given-names>K.</given-names></name> <name><surname>Standley</surname> <given-names>D. M.</given-names></name></person-group> (<year>2013</year>). <article-title>MAFFT multiple sequence alignment software version 7: improvements in performance and usability</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>772</fpage>&#x2013;<lpage>780</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/mst010</pub-id>, PMID: <pub-id pub-id-type="pmid">23329690</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klopfenstein</surname> <given-names>D. V.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Pedersen</surname> <given-names>B. S.</given-names></name> <name><surname>Ram&#x00ED;rez</surname> <given-names>F.</given-names></name> <name><surname>Vesztrocy</surname> <given-names>A. W.</given-names></name> <name><surname>Naldi</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>GOATOOLS: a python library for gene ontology analyses</article-title>. <source>Sci. Rep.</source> <volume>8</volume>:<fpage>10872</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-018-28948-z</pub-id>, PMID: <pub-id pub-id-type="pmid">30022098</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kono</surname> <given-names>T. J. Y.</given-names></name> <name><surname>Brohammer</surname> <given-names>A. B.</given-names></name> <name><surname>McGaugh</surname> <given-names>S. E.</given-names></name> <name><surname>Hirsch</surname> <given-names>C. N.</given-names></name></person-group> (<year>2018</year>). <article-title>Tandem duplicate genes in maize are abundant and date to two distinct periods of time</article-title>. <source>G3</source> <volume>8</volume>, <fpage>3049</fpage>&#x2013;<lpage>3058</lpage>. doi: <pub-id pub-id-type="doi">10.1534/g3.118.200580</pub-id>, PMID: <pub-id pub-id-type="pmid">30030405</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lakhssassi</surname> <given-names>N.</given-names></name> <name><surname>Piya</surname> <given-names>S.</given-names></name> <name><surname>Bekal</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Bergounioux</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A pathogenesis-related protein GmPR08-bet VI promotes a molecular interaction between the GmSHMT08 and GmSNAP18 in resistance to Heterodera glycines</article-title>. <source>Plant Biotechnol. J.</source> <volume>18</volume>, <fpage>1810</fpage>&#x2013;<lpage>1829</lpage>. doi: <pub-id pub-id-type="doi">10.1111/pbi.13343</pub-id>, PMID: <pub-id pub-id-type="pmid">31960590</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leisner</surname> <given-names>C. P.</given-names></name> <name><surname>Hamilton</surname> <given-names>J. P.</given-names></name> <name><surname>Crisovan</surname> <given-names>E.</given-names></name> <name><surname>Manrique-Carpintero</surname> <given-names>N. C.</given-names></name> <name><surname>Marand</surname> <given-names>A. P.</given-names></name> <name><surname>Newton</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Genome sequence of M6, a diploid inbred clone of the high-glycoalkaloid-producing tuber-bearing potato species <italic>Solanum chacoense</italic>, reveals residual heterozygosity</article-title>. <source>Plant J.</source> <volume>94</volume>, <fpage>562</fpage>&#x2013;<lpage>570</lpage>. doi: <pub-id pub-id-type="doi">10.1111/tpj.13857</pub-id>, PMID: <pub-id pub-id-type="pmid">29405524</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>F.</given-names></name> <name><surname>Fan</surname> <given-names>G.</given-names></name> <name><surname>Lu</surname> <given-names>C.</given-names></name> <name><surname>Xiao</surname> <given-names>G.</given-names></name> <name><surname>Zou</surname> <given-names>C.</given-names></name> <name><surname>Kohel</surname> <given-names>R. J.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Genome sequence of cultivated upland cotton (<italic>Gossypium hirsutum</italic> TM-1) provides insights into genome evolution</article-title>. <source>Nat. Biotechnol.</source> <volume>33</volume>, <fpage>524</fpage>&#x2013;<lpage>530</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nbt.3208</pub-id>, PMID: <pub-id pub-id-type="pmid">25893780</pub-id></citation></ref>
<ref id="ref41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lindner</surname> <given-names>A. C.</given-names></name> <name><surname>Lang</surname> <given-names>D.</given-names></name> <name><surname>Seifert</surname> <given-names>M.</given-names></name> <name><surname>Podle&#x0161;&#x00E1;kov&#x00E1;</surname> <given-names>K.</given-names></name> <name><surname>Nov&#x00E1;k</surname> <given-names>O.</given-names></name> <name><surname>Strnad</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Isopentenyltransferase-1 (IPT1) knockout in Physcomitrella together with phylogenetic analyses of IPTs provide insights into evolution of plant cytokinin biosynthesis</article-title>. <source>J. Exp. Bot.</source> <volume>65</volume>, <fpage>2533</fpage>&#x2013;<lpage>2543</lpage>. doi: <pub-id pub-id-type="doi">10.1093/jxb/eru142</pub-id>, PMID: <pub-id pub-id-type="pmid">24692654</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Coulter</surname> <given-names>J. A.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Meng</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Genome-wide identification and analysis of the Q-type C2H2 gene family in potato (<italic>Solanum tuberosum</italic> L.)</article-title>. <source>Int. J. Biol. Macromol.</source> <volume>153</volume>, <fpage>327</fpage>&#x2013;<lpage>340</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ijbiomac.2020.03.022</pub-id>, PMID: <pub-id pub-id-type="pmid">32145229</pub-id></citation></ref>
<ref id="ref40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lynch</surname> <given-names>M.</given-names></name> <name><surname>Force</surname> <given-names>A.</given-names></name></person-group> (<year>2000</year>). <article-title>The probability of duplicate gene preservation by subfunctionalization</article-title>. <source>Genetics</source> <volume>154</volume>, <fpage>459</fpage>&#x2013;<lpage>473</lpage>.</citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Zhong</surname> <given-names>Y.</given-names></name> <name><surname>Geng</surname> <given-names>F.</given-names></name> <name><surname>Cramer</surname> <given-names>G. R.</given-names></name> <name><surname>Cheng</surname> <given-names>Z. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Subfunctionalization of cation/proton antiporter 1 genes in grapevine in response to salt stress in different organs</article-title>. <source>Hortic Res.</source> <volume>2</volume>:<fpage>15031</fpage>. doi: <pub-id pub-id-type="doi">10.1038/hortres.2015.31</pub-id>, PMID: <pub-id pub-id-type="pmid">26504576</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moghe</surname> <given-names>G. D.</given-names></name> <name><surname>Shiu</surname> <given-names>S. H.</given-names></name></person-group> (<year>2014</year>). <article-title>The causes and molecular consequences of polyploidy in flowering plants</article-title>. <source>Ann. N. Y. Acad. Sci.</source> <volume>1320</volume>, <fpage>16</fpage>&#x2013;<lpage>34</lpage>. doi: <pub-id pub-id-type="doi">10.1111/nyas.12466</pub-id>, PMID: <pub-id pub-id-type="pmid">24903334</pub-id></citation></ref>
<ref id="ref45"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ohno</surname> <given-names>S.</given-names></name></person-group> (<year>1970</year>). <source>Evolution by Gene Duplication</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer-Verlag</publisher-name>.</citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Panchy</surname> <given-names>N.</given-names></name> <name><surname>Lehti-Shiu</surname> <given-names>M.</given-names></name> <name><surname>Shiu</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Evolution of gene duplication in plants</article-title>. <source>Plant Physiol.</source> <volume>171</volume>, <fpage>2294</fpage>&#x2013;<lpage>2316</lpage>. doi: <pub-id pub-id-type="doi">10.1104/pp.16.00523</pub-id>, PMID: <pub-id pub-id-type="pmid">27288366</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pham</surname> <given-names>G. M.</given-names></name> <name><surname>Hamilton</surname> <given-names>J. P.</given-names></name> <name><surname>Wood</surname> <given-names>J. C.</given-names></name> <name><surname>Burke</surname> <given-names>J. T.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Vaillancourt</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Construction of a chromosome-scale long-read reference genome assembly for potato</article-title>. <source>GigaScience</source> <volume>9</volume>:<fpage>giaa100</fpage>. doi: <pub-id pub-id-type="doi">10.1093/gigascience/giaa100</pub-id>, PMID: <pub-id pub-id-type="pmid">32964225</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="journal"><person-group person-group-type="author"><collab id="coll5">Potato Genome Sequencing Consortium</collab> <name><surname>Xu</surname> <given-names>X.</given-names></name> <name><surname>Pan</surname> <given-names>S.</given-names></name> <name><surname>Cheng</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name> <name><surname>Mu</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Genome sequence and analysis of the tuber crop potato</article-title>. <source>Nature</source> <volume>475</volume>, <fpage>189</fpage>&#x2013;<lpage>195</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature10158</pub-id>, PMID: <pub-id pub-id-type="pmid">21743474</pub-id></citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prakash</surname> <given-names>C.</given-names></name> <name><surname>Trognitz</surname> <given-names>F. C.</given-names></name> <name><surname>Venhuizen</surname> <given-names>P.</given-names></name> <name><surname>von Haeseler</surname> <given-names>A.</given-names></name> <name><surname>Trognitz</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>A compendium of genome-wide sequence reads from NBS (nucleotide binding site) domains of resistance genes in the common potato</article-title>. <source>Sci. Rep.</source> <volume>10</volume>:<fpage>11392</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-020-67848-z</pub-id>, PMID: <pub-id pub-id-type="pmid">32647195</pub-id></citation></ref>
<ref id="ref50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiao</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Yin</surname> <given-names>H.</given-names></name> <name><surname>Qi</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Gene duplication and evolution in recurring polyploidization-diploidization cycles in plants</article-title>. <source>Genome Biol.</source> <volume>20</volume>:<fpage>38</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13059-019-1650-2</pub-id>, PMID: <pub-id pub-id-type="pmid">30791939</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiao</surname> <given-names>X.</given-names></name> <name><surname>Yin</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>R.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Different modes of gene duplication show divergent evolutionary patterns and contribute differently to the expansion of gene families involved in important fruit traits in pear (<italic>Pyrus bretschneideri</italic>)</article-title>. <source>Front. Plant Sci.</source> <volume>9</volume>:<fpage>161</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpls.2018.00161</pub-id>, PMID: <pub-id pub-id-type="pmid">29487610</pub-id></citation></ref>
<ref id="ref52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quinlan</surname> <given-names>A. R.</given-names></name> <name><surname>Hall</surname> <given-names>I. M.</given-names></name></person-group> (<year>2010</year>). <article-title>BEDTools: a flexible suite of utilities for comparing genomic features</article-title>. <source>Bioinformatics</source> <volume>26</volume>, <fpage>841</fpage>&#x2013;<lpage>842</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btq033</pub-id>, PMID: <pub-id pub-id-type="pmid">20110278</pub-id></citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rehman</surname> <given-names>H. M.</given-names></name> <name><surname>Nawaz</surname> <given-names>M. A.</given-names></name> <name><surname>Shah</surname> <given-names>Z. H.</given-names></name> <name><surname>Ludwig-M&#x00FC;ller</surname> <given-names>J.</given-names></name> <name><surname>Chung</surname> <given-names>G.</given-names></name> <name><surname>Ahmad</surname> <given-names>M. Q.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Comparative genomic and transcriptomic analyses of Family-1 UDP glycosyltransferase in three brassica species and Arabidopsis indicates stress-responsive regulation</article-title>. <source>Sci. Rep.</source> <volume>8</volume>:<fpage>6237</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-018-24308-z</pub-id>, PMID: <pub-id pub-id-type="pmid">29666382</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>L. L.</given-names></name> <name><surname>Liu</surname> <given-names>Y. J.</given-names></name> <name><surname>Liu</surname> <given-names>H. J.</given-names></name> <name><surname>Qian</surname> <given-names>T. T.</given-names></name> <name><surname>Qi</surname> <given-names>L. W.</given-names></name> <name><surname>Wang</surname> <given-names>X. R.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Subcellular relocalization and positive selection play key roles in the retention of duplicate genes of populus class III peroxidase family</article-title>. <source>Plant Cell</source> <volume>26</volume>, <fpage>2404</fpage>&#x2013;<lpage>2419</lpage>. doi: <pub-id pub-id-type="doi">10.1105/tpc.114.124750</pub-id>, PMID: <pub-id pub-id-type="pmid">24934172</pub-id></citation></ref>
<ref id="ref55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Renny-Byfield</surname> <given-names>S.</given-names></name> <name><surname>Gallagher</surname> <given-names>J. P.</given-names></name> <name><surname>Grover</surname> <given-names>C. E.</given-names></name> <name><surname>Szadkowski</surname> <given-names>E.</given-names></name> <name><surname>Page</surname> <given-names>J. T.</given-names></name> <name><surname>Udall</surname> <given-names>J. A.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Ancient gene duplicates in Gossypium (cotton) exhibit near-complete expression divergence</article-title>. <source>Genome Biol. Evol.</source> <volume>6</volume>, <fpage>559</fpage>&#x2013;<lpage>571</lpage>. doi: <pub-id pub-id-type="doi">10.1093/gbe/evu037</pub-id>, PMID: <pub-id pub-id-type="pmid">24558256</pub-id></citation></ref>
<ref id="ref56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rice</surname> <given-names>P.</given-names></name> <name><surname>Longden</surname> <given-names>L.</given-names></name> <name><surname>Bleasby</surname> <given-names>A.</given-names></name></person-group> (<year>2000</year>). <article-title>EMBOSS: the European molecular biology open software suite</article-title>. <source>Trends Genet.</source> <volume>16</volume>, <fpage>276</fpage>&#x2013;<lpage>277</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0168-9525(00)02024-2</pub-id>, PMID: <pub-id pub-id-type="pmid">10827456</pub-id></citation></ref>
<ref id="ref57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roulin</surname> <given-names>A.</given-names></name> <name><surname>Auer</surname> <given-names>P. L.</given-names></name> <name><surname>Libault</surname> <given-names>M.</given-names></name> <name><surname>Schlueter</surname> <given-names>J.</given-names></name> <name><surname>Farmer</surname> <given-names>A.</given-names></name> <name><surname>May</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>The fate of duplicated genes in a polyploid plant genome</article-title>. <source>Plant J.</source> <volume>73</volume>, <fpage>143</fpage>&#x2013;<lpage>153</lpage>. doi: <pub-id pub-id-type="doi">10.1111/tpj.12026</pub-id>, PMID: <pub-id pub-id-type="pmid">22974547</pub-id></citation></ref>
<ref id="ref58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salman-Minkov</surname> <given-names>A.</given-names></name> <name><surname>Sabath</surname> <given-names>N.</given-names></name> <name><surname>Mayrose</surname> <given-names>I.</given-names></name></person-group> (<year>2016</year>). <article-title>Whole-genome duplication as a key factor in crop domestication</article-title>. <source>Nat. Plants</source> <volume>2</volume>:<fpage>16115</fpage>. doi: <pub-id pub-id-type="doi">10.1038/nplants.2016.115</pub-id>, PMID: <pub-id pub-id-type="pmid">27479829</pub-id></citation></ref>
<ref id="ref59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmutz</surname> <given-names>J.</given-names></name> <name><surname>Cannon</surname> <given-names>S. B.</given-names></name> <name><surname>Schlueter</surname> <given-names>J.</given-names></name> <name><surname>Ma</surname> <given-names>J.</given-names></name> <name><surname>Mitros</surname> <given-names>T.</given-names></name> <name><surname>Nelson</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Genome sequence of the palaeopolyploid soybean</article-title>. <source>Nature</source> <volume>463</volume>, <fpage>178</fpage>&#x2013;<lpage>183</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature08670</pub-id>, PMID: <pub-id pub-id-type="pmid">20075913</pub-id></citation></ref>
<ref id="ref60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shiu</surname> <given-names>S. H.</given-names></name> <name><surname>Byrnes</surname> <given-names>J. K.</given-names></name> <name><surname>Pan</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>W. H.</given-names></name></person-group> (<year>2006</year>). <article-title>Role of positive selection in the retention of duplicate genes in mammalian genomes</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>103</volume>, <fpage>2232</fpage>&#x2013;<lpage>2236</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.0510388103</pub-id>, PMID: <pub-id pub-id-type="pmid">16461903</pub-id></citation></ref>
<ref id="ref61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Suyama</surname> <given-names>M.</given-names></name> <name><surname>Torrents</surname> <given-names>D.</given-names></name> <name><surname>Bork</surname> <given-names>P.</given-names></name></person-group> (<year>2006</year>). <article-title>PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments</article-title>. <source>Nucleic Acids Res.</source> <volume>34</volume>, <fpage>W609</fpage>&#x2013;<lpage>W612</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkl315</pub-id>, PMID: <pub-id pub-id-type="pmid">16845082</pub-id></citation></ref>
<ref id="ref62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>S.</given-names></name> <name><surname>Zhong</surname> <given-names>Y.</given-names></name> <name><surname>Hou</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>Tian</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Variation of presence/absence genes among Arabidopsis populations</article-title>. <source>BMC Evol. Biol.</source> <volume>12</volume>:<fpage>86</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1471-2148-12-86</pub-id>, PMID: <pub-id pub-id-type="pmid">22697058</pub-id></citation></ref>
<ref id="ref63"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turr&#x00E0;</surname> <given-names>D.</given-names></name> <name><surname>Bellin</surname> <given-names>D.</given-names></name> <name><surname>Lorito</surname> <given-names>M.</given-names></name> <name><surname>Gebhardt</surname> <given-names>C.</given-names></name></person-group> (<year>2009</year>). <article-title>Genotype-dependent expression of specific members of potato protease inhibitor gene families in different tissues and in response to wounding and nematode infection</article-title>. <source>J. Plant Physiol.</source> <volume>166</volume>, <fpage>762</fpage>&#x2013;<lpage>774</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jplph.2008.10.005</pub-id>, PMID: <pub-id pub-id-type="pmid">19095329</pub-id></citation></ref>
<ref id="ref64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vanneste</surname> <given-names>K.</given-names></name> <name><surname>van de Peer</surname> <given-names>Y.</given-names></name> <name><surname>Maere</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Inference of genome duplications from age distributions revisited</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>177</fpage>&#x2013;<lpage>190</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/mss214</pub-id>, PMID: <pub-id pub-id-type="pmid">22936721</pub-id></citation></ref>
<ref id="ref65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Virtanen</surname> <given-names>P.</given-names></name> <name><surname>Gommers</surname> <given-names>R.</given-names></name> <name><surname>Oliphant</surname> <given-names>T. E.</given-names></name> <name><surname>Haberland</surname> <given-names>M.</given-names></name> <name><surname>Reddy</surname> <given-names>T.</given-names></name> <name><surname>Cournapeau</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>SciPy 1.0: fundamental algorithms for scientific computing in python</article-title>. <source>Nat. Methods</source> <volume>17</volume>, <fpage>261</fpage>&#x2013;<lpage>272</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41592-019-0686-2</pub-id>, PMID: <pub-id pub-id-type="pmid">32015543</pub-id></citation></ref>
<ref id="ref66"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Xie</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Lan</surname> <given-names>B.</given-names></name> <name><surname>You</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Comparative epigenomics reveals evolution of duplicated genes in potato and tomato</article-title>. <source>Plant J.</source> <volume>93</volume>, <fpage>460</fpage>&#x2013;<lpage>471</lpage>. doi: <pub-id pub-id-type="doi">10.1111/tpj.13790</pub-id>, PMID: <pub-id pub-id-type="pmid">29178145</pub-id></citation></ref>
<ref id="ref67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>KaKs_Calculator 2.0: a toolkit incorporating gamma-series methods and sliding window strategies</article-title>. <source>Genom. Proteom. Bioinform.</source> <volume>8</volume>, <fpage>77</fpage>&#x2013;<lpage>80</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S1672-0229(10)60008-3</pub-id>, PMID: <pub-id pub-id-type="pmid">20451164</pub-id></citation></ref>
<ref id="ref68"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xing</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>Evolution of S-domain receptor-like kinases in land plants and origination of S-locus receptor kinases in Brassicaceae</article-title>. <source>BMC Evol. Biol.</source> <volume>13</volume>:<fpage>69</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1471-2148-13-69</pub-id>, PMID: <pub-id pub-id-type="pmid">23510165</pub-id></citation></ref>
<ref id="ref69"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>G.</given-names></name> <name><surname>Ma</surname> <given-names>H.</given-names></name> <name><surname>Nei</surname> <given-names>M.</given-names></name> <name><surname>Kong</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Evolution of F-box genes in plants: different modes of sequence divergence and their relationships with functional diversification</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>106</volume>, <fpage>835</fpage>&#x2013;<lpage>840</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.0812043106</pub-id>, PMID: <pub-id pub-id-type="pmid">19126682</pub-id></citation></ref>
<ref id="ref70"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Nadon</surname> <given-names>B. D.</given-names></name> <name><surname>do Kim</surname> <given-names>K.</given-names></name> <name><surname>Jackson</surname> <given-names>S. A.</given-names></name></person-group> (<year>2018</year>). <article-title>Genetic and epigenetic divergence of duplicate genes in two legume species</article-title>. <source>Plant Cell Environ.</source> <volume>41</volume>, <fpage>2033</fpage>&#x2013;<lpage>2044</lpage>. doi: <pub-id pub-id-type="doi">10.1111/pce.13127</pub-id>, PMID: <pub-id pub-id-type="pmid">29314059</pub-id></citation></ref>
<ref id="ref71"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xuanyuan</surname> <given-names>G.</given-names></name> <name><surname>Lian</surname> <given-names>Q.</given-names></name> <name><surname>Jia</surname> <given-names>R.</given-names></name> <name><surname>Du</surname> <given-names>M.</given-names></name> <name><surname>Kang</surname> <given-names>L.</given-names></name> <name><surname>Pu</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Genome-wide screening and identification of nuclear factor-Y family genes and exploration their function on regulating abiotic and biotic stress in potato (<italic>Solanum tuberosum</italic> L.)</article-title>. <source>Gene</source> <volume>812</volume>:<fpage>146089</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.gene.2021.146089</pub-id>, PMID: <pub-id pub-id-type="pmid">34896520</pub-id></citation></ref>
<ref id="ref72"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Gaut</surname> <given-names>B. S.</given-names></name></person-group> (<year>2011</year>). <article-title>Factors that contribute to variation in evolutionary rate among Arabidopsis genes</article-title>. <source>Mol. Biol. Evol.</source> <volume>28</volume>, <fpage>2359</fpage>&#x2013;<lpage>2369</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/msr058</pub-id>, PMID: <pub-id pub-id-type="pmid">21389272</pub-id></citation></ref>
<ref id="ref73"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Qin</surname> <given-names>M.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Genome-wide identification and expression analysis of the class III peroxidase gene family in potato (<italic>Solanum tuberosum</italic> L.)</article-title>. <source>Front. Genet.</source> <volume>11</volume>:<fpage>593577</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fgene.2020.593577</pub-id>, PMID: <pub-id pub-id-type="pmid">33343634</pub-id></citation></ref>
<ref id="ref74"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yvert</surname> <given-names>G.</given-names></name> <name><surname>Brem</surname> <given-names>R. B.</given-names></name> <name><surname>Whittle</surname> <given-names>J.</given-names></name> <name><surname>Akey</surname> <given-names>J. M.</given-names></name> <name><surname>Foss</surname> <given-names>E.</given-names></name> <name><surname>Smith</surname> <given-names>E. N.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Trans-acting regulatory variation in <italic>Saccharomyces cerevisiae</italic> and the role of transcription factors</article-title>. <source>Nat. Genet.</source> <volume>35</volume>, <fpage>57</fpage>&#x2013;<lpage>64</lpage>. doi: <pub-id pub-id-type="doi">10.1038/ng1222</pub-id>, PMID: <pub-id pub-id-type="pmid">12897782</pub-id></citation></ref>
<ref id="ref75"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Genetic redundancies and their evolutionary maintenance</article-title>. <source>Adv. Exp. Med. Biol.</source> <volume>751</volume>, <fpage>279</fpage>&#x2013;<lpage>300</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-1-4614-3567-9_13</pub-id>, PMID: <pub-id pub-id-type="pmid">22821463</pub-id></citation></ref>
<ref id="ref76"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Belcram</surname> <given-names>H.</given-names></name> <name><surname>Gornicki</surname> <given-names>P.</given-names></name> <name><surname>Charles</surname> <given-names>M.</given-names></name> <name><surname>Just</surname> <given-names>J.</given-names></name> <name><surname>Huneau</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Duplication and partitioning in evolution and function of homoeologous Q loci governing domestication characters in polyploid wheat</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>108</volume>, <fpage>18737</fpage>&#x2013;<lpage>18742</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.1110552108</pub-id>, PMID: <pub-id pub-id-type="pmid">22042872</pub-id></citation></ref>
<ref id="ref78"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Q.</given-names></name> <name><surname>Tang</surname> <given-names>D.</given-names></name> <name><surname>Huang</surname> <given-names>W.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Hamilton</surname> <given-names>J. P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Haplotype-resolved genome analyses of a heterozygous diploid potato</article-title>. <source>Nat. Genet.</source> <volume>52</volume>, <fpage>1018</fpage>&#x2013;<lpage>1023</lpage>. doi: <pub-id pub-id-type="doi">10.1038/s41588-020-0699-x</pub-id>, PMID: <pub-id pub-id-type="pmid">32989320</pub-id></citation></ref></ref-list>
<fn-group>
<fn id="fn0004"><p><sup>1</sup><ext-link xlink:href="https://github.com/groupschoof/AHRD" ext-link-type="uri">https://github.com/groupschoof/AHRD</ext-link></p></fn>
<fn id="fn0005"><p><sup>2</sup><ext-link xlink:href="https://www.r-project.org" ext-link-type="uri">https://www.r-project.org</ext-link></p></fn>
<fn id="fn0006"><p><sup>3</sup><ext-link xlink:href="https://www.ncbi.nlm.nih.gov/sra" ext-link-type="uri">https://www.ncbi.nlm.nih.gov/sra</ext-link></p></fn>
<fn id="fn0007"><p><sup>4</sup><ext-link xlink:href="https://www.python.org" ext-link-type="uri">https://www.python.org</ext-link></p></fn>
<fn id="fn0008"><p><sup>5</sup><ext-link xlink:href="https://github.com/groupschoof/AHRD" ext-link-type="uri">https://github.com/groupschoof/AHRD</ext-link></p></fn>
</fn-group>
</back>
</article>