<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="brief-report" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Ecol. Evol.</journal-id>
<journal-title>Frontiers in Ecology and Evolution</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Ecol. Evol.</abbrev-journal-title>
<issn pub-type="epub">2296-701X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fevo.2023.1112636</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Ecology and Evolution</subject>
<subj-group>
<subject>Perspective</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The geography of genetic data: Current status and future perspectives</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Peng</surname>
<given-names>Xin</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Qiang</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Cheng</surname>
<given-names>Zhentao</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Huang</surname>
<given-names>Xiaolei</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/423108/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>State Key Laboratory of Ecological Pest Control for Fujian and Taiwan Crops, College of Plant Protection, Fujian Agriculture and Forestry University</institution>, <addr-line>Fuzhou</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Fujian Provincial Key Laboratory of Insect Ecology, Fujian Agriculture and Forestry University</institution>, <addr-line>Fuzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Peter Convey, British Antarctic Survey (BAS), United Kingdom</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Jiufeng Wei, Shanxi Agricultural University, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Xiaolei Huang, &#x02709; <email>huangxl@fafu.edu.cn</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Biogeography and Macroecology, a section of the journal Frontiers in Ecology and Evolution</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>11</volume>
<elocation-id>1112636</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Peng, Li, Cheng and Huang.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Peng, Li, Cheng and Huang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>The biogeography field benefits more and more from the growth and application of genetic data such as nucleotide sequences and whole genomes. It has been perceived by scientists that genetic data may be imbalanced among different geographical regions and taxonomic groups. However, the lack of empirical evidence prevents the understanding of current data volume and distribution of genetic data. Based on the construction of a dataset including records for 365 millions of nucleotide sequences of Animalia, Plantae, and Fungi kingdoms, 6 millions of COI sequences of insects, 77 thousands of COI sequences of mammals, 220 thousands of rbcl sequences of Magnoliopsida, and 44 thousands of ITS sequences of Dothideomycetes, here we present evidence on geographical and taxonomical imbalance of the genetic data, identify major gaps and inappropriate practices in the production, application and sharing of genetic data. We then discuss our perspectives on how to fill up gaps and improve the quantity and quality of genetic data.</p>
</abstract>
<kwd-group>
<kwd>biodiversity</kwd>
<kwd>biogeography</kwd>
<kwd>DNA barcode</kwd>
<kwd>genomics</kwd>
<kwd>population genetics</kwd>
</kwd-group>
<contract-num rid="cn1">32270499</contract-num>
<contract-sponsor id="cn1">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="5"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="25"/>
<page-count count="7"/>
<word-count count="3948"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>As genetic variations present within species, genetic diversity is one of the most important components of biological diversity. The maintenance of genetic diversity makes it possible for species to adapt under environmental changes and positively affects ecosystem function and resilience. Genetic diversity and evolutionary process can be presented and analyzed by different kinds of genetic data such as sequences of genes and genomes. As the most widely used genetic data, sequence data can help identify species and infer the relationship among biological groups (<xref ref-type="bibr" rid="ref24">Tautz et al., 2002</xref>; <xref ref-type="bibr" rid="ref8">Hebert et al., 2003</xref>; <xref ref-type="bibr" rid="ref1">Alamouti et al., 2011</xref>), and together with species distributions, can provide important information for quantifying the geographical distribution of genetic variation within and among species and the evolution process of current biodiversity patterns (<xref ref-type="bibr" rid="ref13">Ma et al., 2012</xref>; <xref ref-type="bibr" rid="ref17">Miraldo et al., 2016</xref>; <xref ref-type="bibr" rid="ref25">Toczydlowski et al., 2021</xref>). Large scale geographic data analysis provides an effective way to understand the extensive impact of geographic, geological and climatic changes on species distribution (<xref ref-type="bibr" rid="ref7">Guralnick and Hill, 2009</xref>; <xref ref-type="bibr" rid="ref19">Pope et al., 2015</xref>; <xref ref-type="bibr" rid="ref18">Pelletier et al., 2022</xref>). The integration of sequence and distribution data can help reveal mechanisms underlying species spatial patterns and the historical evolutionary processes (<xref ref-type="bibr" rid="ref2">Avise, 2000</xref>, <xref ref-type="bibr" rid="ref3">2009</xref>), as well as determine conservation units (<xref ref-type="bibr" rid="ref6">Guo et al., 2019</xref>).</p>
<p>The development of sequencing technology in past decades has promoted the accumulation of genetic data (<xref ref-type="bibr" rid="ref23">Sanger et al., 1977</xref>), and with the rise of the second and third generation sequencing technologies, genome sequencing becomes cheaper, faster and more efficient (<xref ref-type="bibr" rid="ref14">Mardis, 2008a</xref>,<xref ref-type="bibr" rid="ref15">b</xref>; <xref ref-type="bibr" rid="ref22">Rhoads and Au, 2015</xref>). A large number of living species have been sequenced and placed in the context of tree of life (<xref ref-type="bibr" rid="ref9">Hedges et al., 2015</xref>; <xref ref-type="bibr" rid="ref10">Hinchliff et al., 2015</xref>). Genetic data are being archived at an alarming rate in public databases (<xref ref-type="bibr" rid="ref25">Toczydlowski et al., 2021</xref>). The GenBank,<xref rid="fn0004" ref-type="fn"><sup>1</sup></xref> a comprehensive repository for genetic data (<xref ref-type="bibr" rid="ref4">Benson et al., 2012</xref>), and the Barcode of Life Data System (BOLD)<xref rid="fn0005" ref-type="fn"><sup>2</sup></xref> mainly archiving DNA barcode sequences of metazoans (<xref ref-type="bibr" rid="ref21">Ratnasingham and Hebert, 2007</xref>), are two examples. As of October 2022, more than 2.4 billion sequences have been available for download in the GenBank database, and more than 12 million barcode sequences are available in the BOLD database. However, although it has been perceived by many scientists that genetic data may be imbalanced among different geographical regions and taxonomic groups, the investigation on current data volume and distribution of genetic data itself and its application is very limited. To reveal the current status and possible existing problems of genetic data, here we present results of a comprehensive analysis and identify major gaps and inappropriate practices in the production, application and sharing of genetic data, and discuss our perspectives on how to improve the quantity and quality of genetic data in the future.</p>
</sec>
<sec id="sec2">
<title>Genetic data volume</title>
<p>In order to explore the current data volume and distribution of genetic data of major biological groups in the public databases, we compiled statistics of the number of taxa (phylum, class, order, family, genus, and species) under the Animalia, Plantae, and Fungi based on the Catalog of Life database (CoL)<xref rid="fn0006" ref-type="fn"><sup>3</sup></xref>, and searched the amount of genetic data (gene and genome sequences) existing in the GenBank and BOLD databases (<xref rid="tab1" ref-type="table">Table 1</xref>). The Animalia has the highest number of species and genetic data, followed by the Plantae and Fungi. <xref rid="fig1" ref-type="fig">Figure 1</xref> shows the number of species/genus and gene sequences at the class level of the three major biological groups. In the Animalia, Insecta has the highest species richness, including 953,381 named species, accounting for 70.25% of the reported species in the Animalia. But its sequence data exhibits obvious difference between different databases. In the GenBank database, sequence data of Insecta only accounts for 17.94% (39,967,813) of that of the Animalia, while in the BOLD database, barcode sequences of Insecta account for 87.53% (9,810,338) of that of the Animalia. Although the species number of Mammalia only accounts for 0.44% (6,025) of the Animalia, this group has the highest amount of sequence data (84,446,193, 37.90%) in the GenBank. In the Plantae, Magnoliopsida has the highest numbers of genus (10,954) and sequences (85,734,539 in GenBank, 317,394 in BOLD), accounting for 53.14% (CoL), 68.25% (GenBank) and 61.19% (BOLD) of the total numbers of genus and sequences in the Plantae, respectively. In the Fungi, the classes with abundant species and sequence data include Dothideomycetes, Sordariomycetes, and Agaricomycetes. We then analyzed the data volume of whole genomes at the class level of different biological groups in the GenBank (<xref rid="fig2" ref-type="fig">Figure 2</xref>). In the Animalia, Insecta has the most genome data (4,691, 31.66%), while Mammalia has only 1,708 genomes (11.53%), which may be due to the small number of Mammalian species. In the Plantae, Magnoliopsida has the highest number of genomes (8,939, 73.68%), followed by Liliopsida (2,136, 17.60%). Among Fungi classes, Sordariomycetes has the most genomes (981, 25.78%).</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption><p>Number of taxa and genetic data in Animalia, Plantea, and Fungi (September 2022).</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Kingdom</th>
<th align="center" valign="top">Phylum</th>
<th align="center" valign="top">Class</th>
<th align="center" valign="top">Order</th>
<th align="center" valign="top">Family</th>
<th align="center" valign="top">Genus</th>
<th align="center" valign="top">Species</th>
<th align="center" valign="top">Gene sequence (GenBank)</th>
<th align="center" valign="top">Genome sequence (GenBank)</th>
<th align="center" valign="top">Barcode sequence (BOLD)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Animalia</td>
<td align="center" valign="top">24</td>
<td align="center" valign="top">99</td>
<td align="center" valign="top">617</td>
<td align="center" valign="top">8,224</td>
<td align="center" valign="top">146,950</td>
<td align="center" valign="top">1,357,163</td>
<td align="center" valign="top">222,798,176</td>
<td align="center" valign="top">14,816</td>
<td align="center" valign="top">11,207,875</td>
</tr>
<tr>
<td align="left" valign="top">Plantea</td>
<td align="center" valign="top">8</td>
<td align="center" valign="top">39</td>
<td align="center" valign="top">211</td>
<td align="center" valign="top">1,002</td>
<td align="center" valign="top">20,613</td>
<td align="center" valign="top">377,980</td>
<td align="center" valign="top">125,614,830</td>
<td align="center" valign="top">12,133</td>
<td align="center" valign="top">518,739</td>
</tr>
<tr>
<td align="left" valign="top">Fungi</td>
<td align="center" valign="top">7</td>
<td align="center" valign="top">52</td>
<td align="center" valign="top">236</td>
<td align="center" valign="top">836</td>
<td align="center" valign="top">9,709</td>
<td align="center" valign="top">140,587</td>
<td align="center" valign="top">16,656,872</td>
<td align="center" valign="top">3,806</td>
<td align="center" valign="top">166,459</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption><p>Statistics of species/genus number and gene sequences at class level in Animalia, Plantae and Fungi.</p></caption>
<graphic xlink:href="fevo-11-1112636-g001.tif"/>
</fig>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption><p>Statistics of species/genus number and whole genomes at class level in Animalia, Plantae and Fungi.</p></caption>
<graphic xlink:href="fevo-11-1112636-g002.tif"/>
</fig>
</sec>
<sec id="sec3">
<title>Geographic distribution of genetic data</title>
<p>In order to explore the spatial pattern of genetic data, we selected the data of most commonly used molecular markers for four classes with rich genetic data as representatives. In total, 2,565,994 and 6,496,753 cytochrome oxidase subunit I (COI) sequences of Insecta, 220,532 and 98,061 rubisco large subunit (rbcl) sequences of Magnoliopsida, 46,521 and 77,439 COI sequences of Mammalia and 44,778 and 13,505 internally transcribed spacer (ITS) sequences of Dothideomycetes were downloaded from the GenBank and BOLD databases, respectively (access on September 21, 2022). It was found that 81.68% (GenBank) and 94.29% (BOLD) COI sequence data of Insecta were with longitude and latitude information, and that of rbcl of Magnoliopsida, COI of Mammalia and ITS of Dothideomycetes were 23.44% (GenBank) and 54.96% (BOLD), 52.87% (GenBank) and 42.70% (BOLD), and 8.07% (GenBank) and 8.60% (BOLD), respectively. After data screening, we obtained 8,390,003 longitude and latitude records from the GenBank and BOLD databases. We then used ArcMap v10.7 to present the global distributions of the sequence data with 4&#x00B0; grid maps (<xref rid="fig3" ref-type="fig">Figure 3</xref>).</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption><p>The geographical patterns of sequence data of Insecta, Mammalia, Magnoliopsida and Dothideomycetes in the GenBank and BOLD database.</p></caption>
<graphic xlink:href="fevo-11-1112636-g003.tif"/>
</fig>
<p>The sequences of Insecta are mainly distributed in North America and Europe, but less distributed in Africa and South America. For the GenBank data, two grids near Toronto, Canada have very high amount of sequences (No. 2, 113620 and No. 3, 174876), and one grid near Banff National Park, Canada (No. 1, 140420) also has high data volume. The BOLD data shows that the grid near Guanacaste National Park in Costa Rica has the highest number of sequences (No. 7, 1870986). The Mammalia COI sequences are mainly distributed in central and south America and southeast Asia, with most data in the grid near Guyana (GenBank: No. 9, 6227; BOLD: No. 10, 6498), but less in Africa and Australia. The sequence data of Magnoliopsida are mainly distributed in North America and South Africa, but less in South America. The grid near Guanacast National Park in Costa Rica has the highest amount of rbcl sequences of Magnoliopsida (GenBank: No. 11, 4483; BOLD: No. 12, 4185). The global distribution of Dothideomycetes ITS sequences is significantly different between databases. GenBank data of Dothideomycetes are mainly distributed in Benin, Africa (No. 13, 845) and Europe, while BOLD data are mainly in the United States and the grid near Texas has most sequences (No. 14, 165).</p>
<p>Obviously, the geographical distribution of sequence data of different biological groups is unbalanced. Due to the subjectivity of selection of sampling sites, a large number of sequence data are concentrated in a few geographical coordinate points. Most of these geographical coordinates are distributed in habitats with high biodiversity, such as national parks and natural reserves. For examples, in the GenBank database, 32,525 Insecta COI sequences are assigned with a single longitude and latitude point (49.001&#x00B0;N, 106.557&#x00B0;W) of Canada Grasslands National Park East Block; 327 Mammalia COI sequences are from a single longitude and latitude point (0.65&#x00B0;S, 76.45&#x00B0;W) in Ecuador Yasuni National Park; and 845 Dothideomycetes ITS sequences are from a longitude and latitude point (9.75&#x00B0;N, 2.2&#x00B0;E) in the African Benin national forest. In the BOLD database, 851,876 Insecta COI sequences are with the same longitude and latitude (10.763&#x00B0;N, 85.334&#x00B0;W) in Linkondela Beha Volcano National Park in Costa Rica, and 1,680 Magnoliopsida rbcl sequences are from a single longitude and latitude point (1.85&#x00B0;S, 102.65&#x00B0;E) in Bukit Duabelas National Park in Indonesia. However, we also found that some longitude and latitude points were located in schools or scientific research institutions rather than natural sampling sites. For examples, for the GenBank data, several longitude and latitude points located near the University of Guelph (43.528&#x00B0;N, 80.229&#x00B0;W; 43.537&#x00B0;N, 80.134&#x00B0;W; 43.5187&#x00B0;N, 80.1709&#x00B0;W; 43.5282&#x00B0;N, 80.229&#x00B0;W; 43.54&#x00B0;N, 80.14&#x00B0;W) contribute 71,616 Insecta COI sequences; and 64,680 Insecta COI sequences are with same longitude and latitude (27.4447&#x00B0;S, 54.9403&#x00B0;W) displaying as Antonio Ramos Research Center in Argentina. Also for the BOLD data, 42,998 Insecta COI sequences are with a same longitude and latitude (22.4685&#x00B0;N, 91.7808&#x00B0;E) as Chittagong University, Bangladesh, and 24,262 Insecta COI sequences with the longitude and latitude (3.1295&#x00B0;N, 101.657&#x00B0;E) of the University of Malaysia in Kuala Lumpur.</p>
</sec>
<sec id="sec4">
<title>Published papers using genetic data</title>
<p>The bibliometric analysis of scientific papers using genetic data can provide information on the application of genetic data. We did subject retrieval (on September 10, 2022) by using &#x201C;Genetic data or Genetic diversity or Molecular phylogeny&#x201D; as retrieval formula in the core database of Web of Science. After duplications were removed by Citespace v6.1.2,<xref rid="fn0007" ref-type="fn"><sup>4</sup></xref> a total of 369,900 valid papers are obtained. It is obvious that the annual numbers of published articles using genetic data were small before 2002, while grew rapidly after 2002 and increased year by year (<xref rid="fig4" ref-type="fig">Figure 4</xref>).</p>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption><p>Growth chart of published papers using genetic data.</p></caption>
<graphic xlink:href="fevo-11-1112636-g004.tif"/>
</fig>
<p>Using VOSviewer<xref rid="fn0008" ref-type="fn"><sup>5</sup></xref> and Scimago graphica<xref rid="fn0009" ref-type="fn"><sup>6</sup></xref> software, the national distribution and international cooperation relationship of papers using genetic data were shown in <xref rid="fig5" ref-type="fig">Figure 5</xref>. The size of nodes represents the number of articles published by a country. The larger the node, the more articles published by that country. The color of the node represents the number of cooperation times with other countries, the closer the node color is to red, the more cooperation times with other countries. The lines connecting nodes represent cooperation in paper publication between countries. The thicker the line between two countries, the more cooperation they have. The results show that the United States has the highest number of international co-authored papers and the most frequent cooperative relationships with other countries. Other countries with relatively high number of international co-authored papers include the United Kingdom, China, Germany, France, Italy, Australia, Spain, Netherlands, and Canada. These countries are not only the main force of doing related research using genetic data, but also have extensive scientific cooperation with many countries.</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption><p>Top 10 countries published international co-authored WoS papers using genetic data.</p></caption>
<graphic xlink:href="fevo-11-1112636-g005.tif"/>
</fig>
</sec>
<sec id="sec5">
<title>Problems in genetic data</title>
<p>In general, in the Animalia, Plantae and Fungi, most of the genetic data are concentrated in a few classes. For example, the Insecta and Mammalia account for 55.84% (GenBank) and 88.77% (BOLD) of the total genetic data in the Animalia. Magnoliopsida account for 68.25% (GenBank) and 61.19% (BOLD) of the total genetic data of the Plantae. Although hundreds of millions of sequence data have been accumulated in the public databases, they are still concentrated in biological groups closely related to human production and life, which makes genetic data accumulation for some biological classes with relatively rich species diversity is still insufficient. For examples, although there are 650 named species in Scaphopoda in the Animalia, the number of sequences is only 128 (GenBank) and 262 (BOLD); although there are 2,395 named species in Laboulbeniomycetes in the Fungi, the number of nucleotide sequences is only 957 (GenBank) and 7 (BOLD); and although there are 102 named species in Andreaeopsida in the Plantae, the number of sequences is only 421 (GenBank) and 96 (BOLD). Obviously, we still suffer a lack of accumulation of genetic data for many biological groups.</p>
<p>While South America, Africa and Australia have very rich species richness, the genetic data are relatively scarce in these regions. For example, although South America has very rich Magnoliopsida species, the rbcl sequence data are rarely distributed in this region. Similarly, compared with the species richness of Mammalia in Africa, their COI data are not much. In addition, most of the genetic data are concentrated in a few geographical grids, as we previously mentioned. Therefore, current sequence data show a disproportional distribution globally. Many geographical regions still lack the accumulation of genetic data, which may hinder us from conducting large-scale biogeography researches using genetic data.</p>
<p>In recent years, studies have reported the metadata deficiency of genetic data in public databases (<xref ref-type="bibr" rid="ref20">Rajesh et al., 2021</xref>), such as that related to the collection location (<xref ref-type="bibr" rid="ref25">Toczydlowski et al., 2021</xref>). For example, <xref ref-type="bibr" rid="ref16">Marques et al. (2013)</xref> found that only 7% of barcode sequences in the then GenBank database contain longitude and latitude information. <xref ref-type="bibr" rid="ref5">Gratton et al. (2017)</xref> found that only 6.2% of GenBank tetrapod accessions include locality data. Obviously, the lack of geographic coordinates will greatly hinder the effective use of genetic data in public databases, and clarifying the status of geographic coordinates of genetic data in different biological groups will provide reference for us to improve practices in data sharing. In our analysis, the geographical coordinate missing of the ITS sequences of Dothideomycetes was the most serious, with only 8.07% (GenBank) and 8.60% (BOLD) sequences having longitudes and latitudes. As a class with high species richness and relatively more genetic data, Dothideomycetes may reflect a general phenomenon of the lack of geographical coordinates in the Fungi. Magnoliopsida is the class with richest species and sequence data in the Plantae, but only 23.44% (GenBank) and 54.97% (BOLD) of rbcl sequences have longitude and latitude information. To our surprise, only about 18.32% (GenBank) and 5.71% (BOLD) of the Insecta COI sequences in are missing their geographic coordinates. However, problems may still exist for these insect geographical coordinates. A large amount of sequence data were found locating at research institutions rather than natural habitats, which cannot correctly reflect the real distribution of organisms in the natural environment. The subsequent use and analysis of these genetic data will be greatly hindered.</p>
</sec>
<sec id="sec6">
<title>Future perspectives</title>
<p>We are now entering the third decade of the 21st century. Over the past 30&#x2009;years, thanks to the efforts of many scientific research institutions and researchers, huge genetic data resources have been accumulated in public databases, making us enter an &#x201C;age of big genetic data.&#x201D; These genetic data have helped researchers make extraordinary achievements in many fields such as ecology, evolution and biogeography. However, there are still major gaps in the public genetic data and inappropriate practices in sharing of genetic data need to be improved. In the future, we need to encourage genetic data accumulation for previously neglected biological groups, especially those with relatively higher species richness but less genetic data volume, and scientific research investment should be increased in areas where genetic data are scarce. These efforts will greatly enrich the global genetic resource pool and provide strong support for in-depth understanding of biological and geographical evolution of species. Furthermore, we should also resolve the loss and error of geographical coordinates of genetic data by improving data practices (<xref ref-type="bibr" rid="ref12">Huang and Qiao, 2011</xref>; <xref ref-type="bibr" rid="ref11">Huang et al., 2012</xref>). Carefully archiving genetic data and relevant geographical information (e.g., natural sampling localities) by researcher will leave more valuable legacy for future research. Public databases should also update their data policy to improve the quality of metadata. In addition, increasing national research input and international cooperation with other countries will effectively help resolve imbalance in distribution of genetic data and research, especially in Africa and South America.</p>
</sec>
<sec id="sec7" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="sec8">
<title>Author contributions</title>
<p>XH conceptualized the study. XP, QL, and ZC collected and analyzed the data. XP and XH wrote the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec9" sec-type="funding-information">
<title>Funding</title>
<p>This research was supported by National Natural Science Foundation of China (Grant number: 32270499).</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alamouti</surname> <given-names>S. M.</given-names></name> <name><surname>Wang</surname> <given-names>V.</given-names></name> <name><surname>Diguistini</surname> <given-names>S.</given-names></name> <name><surname>Six</surname> <given-names>D. L.</given-names></name> <name><surname>Bohlmann</surname> <given-names>J.</given-names></name> <name><surname>Hamelin</surname> <given-names>R. C.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Gene genealogies reveal cryptic species and host preferences for the pine fungal pathogen <italic>Grosmannia clavigera</italic></article-title>. <source>Mol. Ecol.</source> <volume>20</volume>, <fpage>2581</fpage>&#x2013;<lpage>2602</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1365-294X.2011.05109.x</pub-id>, PMID: <pub-id pub-id-type="pmid">21557782</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Avise</surname> <given-names>J.</given-names></name></person-group> (<year>2000</year>). <source>Phylogeography: The History and Formation of Species. Vol. 214</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Harvard</publisher-name>. <fpage>47</fpage>&#x2013;<lpage>48</lpage>.</citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Avise</surname> <given-names>J. C.</given-names></name></person-group> (<year>2009</year>). <article-title>Phylogeography: retrospect and prospect</article-title>. <source>J. Biogeogr.</source> <volume>36</volume>, <fpage>3</fpage>&#x2013;<lpage>15</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1365-2699.2008.02032.x</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benson</surname> <given-names>D. A.</given-names></name> <name><surname>Karsch-Mizrachi</surname> <given-names>I.</given-names></name> <name><surname>Clark</surname> <given-names>K.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name> <name><surname>Ostell</surname> <given-names>J.</given-names></name> <name><surname>Sayers</surname> <given-names>E. W.</given-names></name></person-group> (<year>2012</year>). <article-title>GenBank</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>, <fpage>D48</fpage>&#x2013;<lpage>D53</lpage>. doi: <pub-id pub-id-type="doi">10.1093/nar/gkr1202</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gratton</surname> <given-names>P.</given-names></name> <name><surname>Marta</surname> <given-names>S.</given-names></name> <name><surname>Bocksberger</surname> <given-names>G.</given-names></name> <name><surname>Winter</surname> <given-names>M.</given-names></name> <name><surname>Trucchi</surname> <given-names>E.</given-names></name> <name><surname>K&#x00FC;hl</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>A world of sequences: can we use georeferenced nucleotide databases for a robust automated phylogeography?</article-title> <source>J. Biogeogr.</source> <volume>44</volume>, <fpage>475</fpage>&#x2013;<lpage>486</lpage>. doi: <pub-id pub-id-type="doi">10.1111/jbi.12786</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Wei</surname> <given-names>K.</given-names></name> <name><surname>Ji</surname> <given-names>W.</given-names></name> <name><surname>Yan</surname> <given-names>R.</given-names></name> <name><surname>Wei</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Phylogeography of the threatened tetraploid fish, <italic>Schizothorax waltoni</italic>, in the Yarlung Tsangpo River on the southern Qinghai-Tibet plateau: implications for conservation</article-title>. <source>Sci. Rep.</source> <volume>9</volume>:<fpage>2704</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-019-39128-y</pub-id>, PMID: <pub-id pub-id-type="pmid">30804376</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guralnick</surname> <given-names>R.</given-names></name> <name><surname>Hill</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Biodiversity informatics: automated approaches for documenting global biodiversity patterns and processes</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>421</fpage>&#x2013;<lpage>428</lpage>. doi: <pub-id pub-id-type="doi">10.1093/bioinformatics/btn659</pub-id>, PMID: <pub-id pub-id-type="pmid">19129210</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hebert</surname> <given-names>P. D. N.</given-names></name> <name><surname>Cywinska</surname> <given-names>A.</given-names></name> <name><surname>Ball</surname> <given-names>S. L.</given-names></name> <name><surname>DeWaard</surname> <given-names>J. R.</given-names></name></person-group> (<year>2003</year>). <article-title>Biological identifications through DNA barcodes</article-title>. <source>Biol. Sci.</source> <volume>270</volume>, <fpage>313</fpage>&#x2013;<lpage>321</lpage>. doi: <pub-id pub-id-type="doi">10.1098/rspb.2002.2218</pub-id>, PMID: <pub-id pub-id-type="pmid">12614582</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hedges</surname> <given-names>S. B.</given-names></name> <name><surname>Marin</surname> <given-names>J.</given-names></name> <name><surname>Suleski</surname> <given-names>M.</given-names></name> <name><surname>Paymer</surname> <given-names>M.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Tree of life reveals clock-like speciation and diversification</article-title>. <source>Mol. Biol. Evol.</source> <volume>32</volume>, <fpage>835</fpage>&#x2013;<lpage>845</lpage>. doi: <pub-id pub-id-type="doi">10.1093/molbev/msv037</pub-id>, PMID: <pub-id pub-id-type="pmid">25739733</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinchliff</surname> <given-names>C. E.</given-names></name> <name><surname>Smith</surname> <given-names>S. A.</given-names></name> <name><surname>Allman</surname> <given-names>J. F.</given-names></name> <name><surname>Burleigh</surname> <given-names>J. G.</given-names></name> <name><surname>Chaudhary</surname> <given-names>R.</given-names></name> <name><surname>Coghill</surname> <given-names>L. M.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Synthesis of phylogeny and taxonomy into a comprehensive tree of life</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>112</volume>, <fpage>12764</fpage>&#x2013;<lpage>12769</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.1423041112</pub-id>, PMID: <pub-id pub-id-type="pmid">26385966</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Hawkins</surname> <given-names>B. A.</given-names></name> <name><surname>Lei</surname> <given-names>F.</given-names></name> <name><surname>Miller</surname> <given-names>G. L.</given-names></name> <name><surname>Favret</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Willing or unwilling to share primary biodiversity data: results and implications of an international survey</article-title>. <source>Conserv. Lett.</source> <volume>5</volume>, <fpage>399</fpage>&#x2013;<lpage>406</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1755-263X.2012.00259.x</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Qiao</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <article-title>Biodiversity databases should gain support from journals</article-title>. <source>Trends Ecol. Evol.</source> <volume>26</volume>, <fpage>377</fpage>&#x2013;<lpage>378</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tree.2011.05.006</pub-id>, PMID: <pub-id pub-id-type="pmid">21665319</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>P.</given-names></name> <name><surname>Jiang</surname> <given-names>F.</given-names></name> <name><surname>Chapuis</surname> <given-names>M.</given-names></name> <name><surname>Shali</surname> <given-names>Y.</given-names></name> <name><surname>Sword</surname> <given-names>G. A.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Mitochondrial genomes reveal the global phylogeography and dispersal routes of the migratory locust</article-title>. <source>Mol. Ecol.</source> <volume>21</volume>, <fpage>4344</fpage>&#x2013;<lpage>4358</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1365-294X.2012.05684.x</pub-id>, PMID: <pub-id pub-id-type="pmid">22738353</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mardis</surname> <given-names>E. R.</given-names></name></person-group> (<year>2008a</year>). <article-title>Next-generation DNA sequencing methods</article-title>. <source>Annu. Rev. Genomics Hum. Genet.</source> <volume>9</volume>, <fpage>387</fpage>&#x2013;<lpage>402</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev.genom.9.081307.164359</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mardis</surname> <given-names>E. R.</given-names></name></person-group> (<year>2008b</year>). <article-title>The impact of next-generation sequencing technology on genetics</article-title>. <source>Trends Genet.</source> <volume>24</volume>, <fpage>133</fpage>&#x2013;<lpage>141</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.tig.2007.12.007</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marques</surname> <given-names>A. C.</given-names></name> <name><surname>Maronna</surname> <given-names>M. M.</given-names></name> <name><surname>Collins</surname> <given-names>A. G.</given-names></name></person-group> (<year>2013</year>). <article-title>Putting GenBank data on the map</article-title>. <source>Science</source> <volume>341</volume>:<fpage>1341</fpage>. doi: <pub-id pub-id-type="doi">10.1126/science.341.6152.1341-a</pub-id>, PMID: <pub-id pub-id-type="pmid">24052287</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miraldo</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Borregaard</surname> <given-names>M. K.</given-names></name> <name><surname>Fl&#x00F3;rez-Rodr&#x00ED;guez</surname> <given-names>A.</given-names></name> <name><surname>Gopalakrishnan</surname> <given-names>S.</given-names></name> <name><surname>Rizvanovic</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>An Anthropocene map of genetic diversity</article-title>. <source>Science</source> <volume>353</volume>, <fpage>1532</fpage>&#x2013;<lpage>1535</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.aaf4381</pub-id>, PMID: <pub-id pub-id-type="pmid">27708102</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pelletier</surname> <given-names>T. A.</given-names></name> <name><surname>Parsons</surname> <given-names>D. J.</given-names></name> <name><surname>Decker</surname> <given-names>S. K.</given-names></name> <name><surname>Crouch</surname> <given-names>S.</given-names></name> <name><surname>Franz</surname> <given-names>E.</given-names></name> <name><surname>Ohrstrom</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>phylogatR: Phylogeographic data aggregation and repurposing</article-title>. <source>Mol. Ecol. Resour.</source> <volume>22</volume>, <fpage>2830</fpage>&#x2013;<lpage>2842</lpage>. doi: <pub-id pub-id-type="doi">10.1111/1755-0998.13673</pub-id>, PMID: <pub-id pub-id-type="pmid">35748425</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pope</surname> <given-names>L. C.</given-names></name> <name><surname>Liggins</surname> <given-names>L.</given-names></name> <name><surname>Keyse</surname> <given-names>J.</given-names></name> <name><surname>Carvalho</surname> <given-names>S. B.</given-names></name> <name><surname>Riginos</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Not the time or the place: the missing spatio-temporal link in publicly available genetic data</article-title>. <source>Mol. Ecol.</source> <volume>24</volume>, <fpage>3802</fpage>&#x2013;<lpage>3809</lpage>. doi: <pub-id pub-id-type="doi">10.1111/mec.13254</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rajesh</surname> <given-names>A.</given-names></name> <name><surname>Chang</surname> <given-names>Y.</given-names></name> <name><surname>Abedalthagafi</surname> <given-names>M. S.</given-names></name> <name><surname>Wong-Beringer</surname> <given-names>A.</given-names></name> <name><surname>Love</surname> <given-names>M. I.</given-names></name> <name><surname>Mangul</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Improving the completeness of public metadata accompanying omics studies</article-title>. <source>Genome Biol.</source> <volume>22</volume>:<fpage>106</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13059-021-02332-z</pub-id>, PMID: <pub-id pub-id-type="pmid">33858487</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ratnasingham</surname> <given-names>S.</given-names></name> <name><surname>Hebert</surname> <given-names>P. D. N.</given-names></name></person-group> (<year>2007</year>). <article-title>Bold: the barcode of life data system (http://www.barcodinglife.org)</article-title>. <source>Mol. Ecol. Notes</source> <volume>7</volume>, <fpage>355</fpage>&#x2013;<lpage>364</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1471-8286.2007.01678.x</pub-id>, PMID: <pub-id pub-id-type="pmid">18784790</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rhoads</surname> <given-names>A.</given-names></name> <name><surname>Au</surname> <given-names>K. F.</given-names></name></person-group> (<year>2015</year>). <article-title>PacBio sequencing and its applications</article-title>. <source>Genomics Proteomics Bioinf.</source> <volume>13</volume>, <fpage>278</fpage>&#x2013;<lpage>289</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.gpb.2015.08.002</pub-id>, PMID: <pub-id pub-id-type="pmid">26542840</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanger</surname> <given-names>F.</given-names></name> <name><surname>Nicklen</surname> <given-names>S.</given-names></name> <name><surname>Coulson</surname> <given-names>A. R.</given-names></name></person-group> (<year>1977</year>). <article-title>DNA sequencing with chain terminating inhibitors</article-title>. <source>PNAS</source> <volume>74</volume>, <fpage>5463</fpage>&#x2013;<lpage>5467</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.74.12.5463</pub-id>, PMID: <pub-id pub-id-type="pmid">271968</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tautz</surname> <given-names>D.</given-names></name> <name><surname>Arctander</surname> <given-names>P.</given-names></name> <name><surname>Minelli</surname> <given-names>A.</given-names></name> <name><surname>Thomas</surname> <given-names>R. H.</given-names></name> <name><surname>Vogler</surname> <given-names>A. P.</given-names></name></person-group> (<year>2002</year>). <article-title>DNA points the way ahead in taxonomy</article-title>. <source>Nature</source> <volume>418</volume>:<fpage>479</fpage>. doi: <pub-id pub-id-type="doi">10.1038/418479a</pub-id>, PMID: <pub-id pub-id-type="pmid">12152050</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toczydlowski</surname> <given-names>R. H.</given-names></name> <name><surname>Liggins</surname> <given-names>L.</given-names></name> <name><surname>Gaither</surname> <given-names>M. R.</given-names></name> <name><surname>Anderson</surname> <given-names>T. J.</given-names></name> <name><surname>Barton</surname> <given-names>R. L.</given-names></name> <name><surname>Berg</surname> <given-names>J. T.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Poor data stewardship will hinder global genetic diversity surveillance</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>118</volume>:<fpage>e2107934118</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.2107934118</pub-id>, PMID: <pub-id pub-id-type="pmid">34404731</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn0004"><p><sup>1</sup><ext-link xlink:href="http://www.ncbi.nlm.nih.gov/genbank" ext-link-type="uri">www.ncbi.nlm.nih.gov/genbank</ext-link></p></fn>
<fn id="fn0005"><p><sup>2</sup><ext-link xlink:href="http://www.boldsystems.org" ext-link-type="uri">www.boldsystems.org</ext-link></p></fn>
<fn id="fn0006"><p><sup>3</sup><ext-link xlink:href="http://www.catalogueoflife.org" ext-link-type="uri">www.catalogueoflife.org</ext-link></p></fn>
<fn id="fn0007"><p><sup>4</sup><ext-link xlink:href="https://citespace.podia.com/" ext-link-type="uri">https://citespace.podia.com/</ext-link></p></fn>
<fn id="fn0008"><p><sup>5</sup><ext-link xlink:href="https://www.vosviewer.com/" ext-link-type="uri">https://www.vosviewer.com/</ext-link></p></fn>
<fn id="fn0009"><p><sup>6</sup><ext-link xlink:href="https://graphica.app/" ext-link-type="uri">https://graphica.app/</ext-link></p></fn>
</fn-group>
</back>
</article>
