<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2021.752863</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Superiority Verification of Deep Learning in the Identification of Medicinal Plants: Taking <italic>Paris polyphylla</italic> var. <italic>yunnanensis</italic> as an Example</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Yue</surname> <given-names>JiaQi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>WanYi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wang</surname> <given-names>YuanZhong</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/489981/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Medicinal Plants Research Institute, Yunnan Academy of Agricultural Sciences</institution>, <addr-line>Kunming</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>College of Traditional Chinese Medicine, Yunnan University of Chinese Medicine</institution>, <addr-line>Kunming</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Kioumars Ghamkhar, AgResearch Ltd., New Zealand</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Jingli Lu, AgResearch Ltd., New Zealand; Ke Han, Harbin University of Commerce, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: YuanZhong Wang, <email>boletus@126.com</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Technical Advances in Plant Science, a section of the journal Frontiers in Plant Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>09</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>752863</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2021 Yue, Li and Wang.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Yue, Li and Wang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Medicinal plants have a variety of values and are an important source of new drugs and their lead compounds. They have played an important role in the treatment of cancer, AIDS, COVID-19 and other major and unconquered diseases. However, there are problems such as uneven quality and adulteration. Therefore, it is of great significance to find comprehensive, efficient and modern technology for its identification and evaluation to ensure quality and efficacy. In this study, deep learning, which is superior to conventional identification techniques, was extended to the identification of the part and region of the medicinal plant <italic>Paris polyphylla</italic> var. <italic>yunnanensis</italic> from the perspective of spectroscopy. Two pattern recognition models, partial least squares discriminant analysis (PLS-DA) and support vector machine (SVM), were established, and the overall discrimination performance of the three types of models was compared. In addition, we also compared the effects of different sample sizes on the discriminant performance of the models for the first time to explore whether the three models had sample size dependence. The results showed that the deep learning model had absolute superiority in the identification of medicinal plant. It was almost unaffected by factors such as data type and sample size. The overall identification ability was significantly better than the PLS-DA and SVM models. This study verified the superiority of the deep learning from examples, and provided a practical reference for related research on other medicinal plants.</p>
</abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>identification research</kwd>
<kwd>medicinal plant</kwd>
<kwd><italic>Paris polyphylla</italic> var. <italic>yunnanensis</italic></kwd>
<kwd>superiority verification</kwd>
<kwd>ResNet</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="3"/>
<equation-count count="5"/>
<ref-count count="38"/>
<page-count count="15"/>
<word-count count="8670"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="S1">
<title>Introduction</title>
<p>Medicinal plants are a kind of highly exploitable plants with various values such as medicinal edible ecology. Their research has become the latest source for the emergence of new drugs (<xref ref-type="bibr" rid="B14">Newman and Cragg, 2015</xref>). The development potential of the international market for the utilization of medicinal plants is huge, and countries all over the world generally attach importance to its research in order to better transform and utilize medicinal plants, solve the problem of human survival resource shortage, and improve human health (<xref ref-type="bibr" rid="B9">Jamshidi-Kia et al., 2018</xref>). Medicinal plants have a wide range of sources. Due to differences in regional natural conditions, climatic conditions, flora and natural resources, they present a unique distribution with great differences in quantity and type (<xref ref-type="bibr" rid="B3">Deng et al., 2016</xref>). Many factors have different degrees of influence on the quality of medicinal plants. Therefore, the use of comprehensive, efficient, and modern technical means to clarify the region and part of medicinal plants has far-reaching significance for quality and efficacy.</p>
<p>Traditional identification and evaluation techniques for medicinal plants mainly include the technology of DNA barcoding, macroscopic identification, microscopic identification, chromatography, spectroscopy, etc. (<xref ref-type="bibr" rid="B24">Pang et al., 2011</xref>; <xref ref-type="bibr" rid="B26">Pei et al., 2020</xref>; <xref ref-type="bibr" rid="B13">Liu et al., 2021</xref>). Among them, spectroscopy has the advantages of simplicity, speed, economy, and high throughput, which can fully characterize the chemical information of samples with complex mixed systems (<xref ref-type="bibr" rid="B25">Pasquini, 2018</xref>). The identification research of medicinal plants mostly uses spectroscopy combined with chemometrics. Among them, the partial least square discriminant analysis (PLS-DA) and support vector machine (SVM) have excellent performance, and have been successfully applied to the identification and evaluation of a variety of medicinal plants, including species identification, origin identification, age identification, part identification, adulteration identification, etc. (<xref ref-type="bibr" rid="B12">Liu et al., 2020</xref>; <xref ref-type="bibr" rid="B29">Shen et al., 2020</xref>; <xref ref-type="bibr" rid="B33">Wang et al., 2020</xref>) <xref ref-type="bibr" rid="B36">Yang and Wang (2018)</xref> compared the effects of PLS-DA and SVM on the identification of <italic>P. polyphylla</italic> var. <italic>yunnanensis</italic> from different regions based on infrared spectroscopy and ultraviolet spectroscopy data. It is found that both models have higher recognition performance, and the accuracy of SVM is higher than that of PLS-DA.</p>
<p>In addition, two-dimensional correlation spectroscopy (2DCOS) is also a powerful tool for identification evaluation. This technology fully combines the advantages of computational chemistry, statistics, spectroscopy and computer science to increase the spectral resolution and enrich the information carried by the spectrum by increasing the dimension (<xref ref-type="bibr" rid="B16">Noda, 1989</xref>, <xref ref-type="bibr" rid="B18">1993</xref>). In recent years, reports on the research and application of 2DCOS technology are increasing year by year, covering drug metabolism, drug toxicology, drug structure-activity relationship, traditional Chinese medicine, etc. (<xref ref-type="bibr" rid="B19">Noda, 2004</xref>, <xref ref-type="bibr" rid="B20">2014</xref>, <xref ref-type="bibr" rid="B21">2016</xref>; <xref ref-type="bibr" rid="B11">Li et al., 2014</xref>). Based on years of research, <xref ref-type="bibr" rid="B30">Sun et al. (2003)</xref> wrote a book called <italic>&#x201C;Atlas of Two-dimensional Correlation Infrared Spectroscopy for Traditional Chinese Medicine Identification,&#x201D;</italic> which contains the 2DCOS spectra of more than 300 kinds of traditional Chinese medicine, providing a reference for the identification research of related traditional Chinese medicine. However, the artificial identification and analysis of 2DCOS spectra has limitations in time, technology, and experience. Moreover, interdisciplinary research has become a current hot spot and also the trend of future scientific research field. Therefore, it is necessary to combine 2DCOS with more modern, convenient and intelligent technical means of other disciplines to realize the rapid identification of medicinal plants.</p>
<p>Deep learning is the main research method used in the development of artificial intelligence research at the present stage, which has unique advantages in image classification and object recognition (<xref ref-type="bibr" rid="B10">LeCun et al., 2015</xref>; <xref ref-type="bibr" rid="B7">Houssein et al., 2021</xref>). Combining it with 2DCOS images for the identification of medicinal plants can take advantage of the respective advantages of the two technologies and greatly improve the efficiency of identification and analysis. Deep learning combined with 2DCOS seems to show superior performance in many aspects than traditional spectroscopy combined with chemometrics in identifying medicinal plants (<xref ref-type="bibr" rid="B4">Dong et al., 2020</xref>). For example, deep learning can achieve good identification without complex spectral preprocessing, and there is no need to manually extract features in the modeling process, which greatly improves efficiency and reduces various risks caused by human factors (<xref ref-type="bibr" rid="B6">Grinblat et al., 2016</xref>). However, these conclusions are all based on theories or the application of a single method, and there has been no actual comparison and discussion on them.</p>
<p><italic>Paris polyphylla</italic> var. <italic>yunnanensis</italic> (PPY), as the original plant of the precious Chinese medicine Paridis Rhizoma, is a medicinal plant resource with a representative and global influence (<xref ref-type="bibr" rid="B2">Cunningham et al., 2018</xref>). In the market, there are more than 80 commonly used Chinese patent medicines with Paridis Rhizoma as the main raw material, and 107 pharmaceutical companies are involved in the production, which are distributed in 23 provinces of China. They have significant clinical efficacy and economic value (<xref ref-type="bibr" rid="B31">Tao et al., 2020</xref>). At present, domestic and foreign scholars have conducted a lot of research on PPY, but the research on the resources evaluation is still in a situation where there are results but no conclusions, and they are all based on the traditional medicinal rhizoma. Moreover, studying the above-ground parts of PPY can promote the development and utilization of non-medicinal parts, and improve economic benefits (<xref ref-type="bibr" rid="B38">Zhao et al., 2021</xref>). Besides, there is currently no research on the use of deep learning combined with 2DCOS to identify the parts and regions of PPY.</p>
<p>In conclusion, taking PPY as an example, two pattern recognition models of PLS-DA and SVM, and a deep learning model of Residual neural network (ResNet) were established in this study to explore and verify whether deep learning combined with 2DCOS has advantages in the identification of medicinal plant resources. In order to increase comparability and credibility, we simultaneously identified and evaluated PPY samples of different regions and parts. In addition, we also compared the impact of different sample sizes on model identification performance to explore whether the three models are dependent on sample size. This research not only provided a reasonable, standardized, fast and effective method for the identification of regions and parts of PPY, but also verified the superiority of the deep learning model in the identification of medicinal plants and the response of the three models to sample size. This is conducive to the development and utilization of advanced deep learning models such as ResNet in other fields.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2.SS1">
<title>Sample Information</title>
<p>A total of 772 individuals were collected in 12 sampling sites in central, northwest, southeast, southwest and western Yunnan (<xref ref-type="fig" rid="F1">Figure 1</xref>). All samples were identified as <italic>Paris polyphylla</italic> var. <italic>yunnanensis</italic> by Professor Hang Jin from the Institute of Medicinal Plants, Yunnan Academy of Agricultural Sciences. Some samples are shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. Afterward, all the samples were cleaned and divided into four parts: rhizome, stem, leaf and fibrous root. Then the samples were dried to a constant weight at 50&#x00B0;C in an electric thermostatic drying oven. Next, the samples were passed through a 100-mesh sieve. Finally, the fine powders were stored in self-sealed bags and kept in a dry environment away from light for subsequent analysis. The detailed information of the samples is shown in <xref ref-type="supplementary-material" rid="DS1">Supplementary Table 1</xref>. There are a total of 772 rhizomes, all of which were used for regions identification analysis. Rhizome (G: 142), stem (J: 107), leaf (Y: 137), and fibrous root (XG: 107) from Dehong and Yuxi were selected for identification of parts.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Location distribution of <italic>Paris polyphylla</italic> var. <italic>yunnanensis</italic> samples in western, central, northwest, southwest and southeast of Yunnan.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Sample picture of the planting site, whole plant and rhizome of <italic>Paris polyphylla</italic> var. <italic>yunnanensis.</italic></p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g002.tif"/>
</fig>
</sec>
<sec id="S2.SS2">
<title>FT-MIR Spectra Acquisition</title>
<p>The Fourier transform mid-infrared spectra were collected by a Fourier transform infrared spectrometer equipped with an attenuated total reflection accessory (Perkin Elmer, Norwalk, CT, United States). Sample powder (2 &#x00B1; 0.2 mg) was placed in the center of the metal ring (ZnSe crystal surface), and the manometer knob was adjusted to a uniform progress bar of 131 &#x00B1; 1 to form sample powder sheets with the same thickness. The infrared spectrum scanning range was set to be 4,000&#x2013;550 cm<sup>&#x2013;1</sup> with a spectral resolution of 4 cm<sup>&#x2013;1</sup>. Sixteen times of scanning were carried out, and each sample was measured in parallel for three times. Finally, the average spectrum was taken. Before the sample scanning, the infrared spectrum of the blank crystal surface is collected, and the interference of air and the scattering spectrum of the crystal part was deducted. During the spectrum measurement, keep the laboratory temperature at 25&#x00B0;C and the relative air humidity at 30%.</p>
</sec>
<sec id="S2.SS3">
<title>Data Processing and Exploratory Analysis</title>
<p>Although the spectral data preprocessing and the characteristic variable selection have been proved by previous studies to be effective for optimizing identification model (<xref ref-type="bibr" rid="B23">Obaid et al., 2019</xref>), the complex data preprocessing process will greatly reduce the recognition efficiency. Moreover, the preprocessing methods and characteristic variable selection methods used for different data sets cannot be unified, which requires a lot of time and resource costs to verify. Therefore, this study directly used original spectral data for subsequent identification analysis without considering data preprocessing and characteristic variable selection, so as to fairly compare the recognition performance of the three types of models and verify whether the ResNet model has advantages in the identification research.</p>
<p>In addition, in order to explore the impact of sample size on the recognition ability of the three types of models, we divided the data sets of region and part into low sample size group (10%), medium sample size group (50%), and high sample size group (100%), and the percentage in parentheses is the proportion of each group of samples (<xref ref-type="supplementary-material" rid="DS1">Supplementary Table 2</xref>). The Kennard-stone algorithm was performed to divide the data of all groups into training set (2/3) and test set (1/3), which was directly used to build PLS-DA and SVM models. The data for establishing the ResNet model is the 2DCOS images of all groups, and the generation method is shown in the following section.</p>
<p>Exploratory analysis used the unsupervised analysis method of t-distributed stochastic neighbor embedding (t-SNE) to summarize the distribution of grouped samples in a multivariate space. By identifying the distribution trend of samples, high-dimensional data can be visualized as data points in two-dimensional or three-dimensional graphs. The above process was completed by MATLAB software.</p>
</sec>
<sec id="S2.SS4">
<title>Two-Dimensional Correlation Spectroscopy Spectra Image Acquisition</title>
<p>The generalized two-dimensional correlation spectrum is an effective method to improve spectral resolution and solve spectral overlap by designing disturbance variables, which is obtained by discrete generalized 2DCOS algorithm. Its dynamic spectrum is expressed as <bold><italic>S</italic></bold>, and the expression is as follows, where <italic>v</italic> is variable and <italic>t</italic> is the external disturbance (<xref ref-type="bibr" rid="B22">Noda, 2018</xref>).</p>
<disp-formula id="S2.E1"><label>(1)</label><mml:math id="M1" display="block"><mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo rspace="5.3pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="8.1pt">=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable displaystyle="true" rowspacing="0pt"><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mo>&#x22C5;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>The synchronous spectral intensity &#x03A6;(<bold><italic>v</italic></bold><sub>1</sub>, <bold><italic>v</italic></bold><sub>2</sub>) is equal to the cross product of the dynamic spectral intensity at (<italic>v</italic><sub>1</sub>, <italic>v</italic><sub>2</sub>). The asynchronous spectral intensity &#x03A8;(<italic>v</italic><sub>1</sub>, <italic>v</italic><sub>2</sub>) is equal to the cross product of the Hilbert-Noda matrix defined as <italic>N</italic><sub><italic>jk</italic></sub> for the dynamic spectral intensity at (<italic>v</italic><sub>1</sub>, <italic>v</italic><sub>2</sub>). Their expressions are as follows:</p>
<disp-formula id="S2.E2"><label>(2)</label><mml:math id="M2" display="block"><mml:mrow><mml:mrow><mml:mpadded lspace="2.8pt" width="+2.8pt"><mml:mtext>&#x03A6;</mml:mtext></mml:mpadded><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo rspace="5.3pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.3pt">=</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mi>S</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E3"><label>(3)</label><mml:math id="M3" display="block"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo rspace="5.3pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.3pt">=</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mi>S</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E4"><label>(4)</label><mml:math id="M4" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mpadded width="+2.8pt"><mml:mtable displaystyle="true" rowspacing="0pt"><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo mathvariant="italic" separator="true">&#x2003;&#x2002;&#x2006;</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mrow><mml:mpadded width="+8.3pt"><mml:mstyle displaystyle="false"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x03C0;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:mpadded><mml:mi>j</mml:mi></mml:mrow><mml:mo>&#x2260;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mpadded><mml:mi/></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The product of a pair of synchronous and asynchronous correlation intensities can obtain the integrated two-dimensional correlation intensity, which is expressed asI(<italic>v</italic><sub>1</sub>, <italic>v</italic><sub>2</sub>) (<xref ref-type="bibr" rid="B1">Chen et al., 2018</xref>).</p>
<disp-formula id="S2.E5"><label>(5)</label><mml:math id="M5" display="block"><mml:mtable><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mrow><mml:mtext>I</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Spectral data matrix S(m &#x00D7; n) contains two spectra, the first is the average FT-MIR of each class, and the second is the <italic>i</italic>th FT-MIR spectra of each class. The synchronous 2DCOS spectra, asynchronous 2DCOS spectra and integrative 2DCOS (i2DCOS) spectra for the <italic>i</italic>th sample of each category can be obtained by equation (2), (3) and (4). In order to reduce the amount of calculation, save computer resources and speed up the calculation efficiency, the fingerprint area of 1,750&#x2013;550 cm<sup>&#x2013;1</sup> was selected, and the synchronous 2DCOS, asynchronous 2DCOS and i2DCOS spectral images were automatically generated by the software Matlab2017b. The image size can be chosen according to the processing power of the computer (32 &#x00D7; 32 pixel, 64 &#x00D7; 64 pixel and 128 &#x00D7; 128 pixel), and the generated 2DCOS images were stored in JPEG image format with the size as 64 &#x00D7; 64 pixel in the corresponding folder for building ResNet model. Using the Kennard-stone algorithm, all datasets were divided into training set (60%), test set (30%), and external validation set (10%). The process of generating all types of 2DCOS spectra images is shown in <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 1</xref>.</p>
</sec>
<sec id="S2.SS5">
<title>Partial Least Squares Discrimination Analysis</title>
<p>Partial least squares discriminant analysis is a linear supervised classification method established on the basis of the standard PLS regression algorithm. It searches for the variable with the largest covariance of the classification matrix Y from the variable matrix X. Y is divided into two categories, where Y = 1 represents that the sample belongs to a specific category, and Y = 0 represents that the sample does not belong to a specific category. Finally, the probability of each sample classified into each category is obtained. In the calculation, the observed X matrix is transformed into a set of several intermediate linear latent variables (LVs). The first n LVs are selected according to the maximum eigenvalue greater than 1. The statistical parameters of accuracy, model fitting determination coefficient R<sup>2</sup>, Q<sup>2</sup>, root mean square error of estimation (RMSEE), root mean square error of cross validation (RMSECV), and root mean square error of prediction (RMSEP) are used to evaluate the performance of the model. Permutation test was performed on the established model with a total of 50 iterations. And according to the R<sup>2</sup>-intercept and Q<sup>2</sup>-intercept results, the fitting degree of the model was verified. The process of establishing PLS-DA model was carried out on SIMCA-P+14.1 software.</p>
</sec>
<sec id="S2.SS6">
<title>Support Vector Machine</title>
<p>Support vector machine is a supervised pattern recognition method that can identify unknown samples and has the ability to analyze the data with high collinearity and high noise. The libsvm-3.20 toolbox developed by the Institute of Industrial Engineering, National Taiwan University, Lin Zhiren, etc., was used to establish SVM discriminant models to identify the region and part of <italic>P. polyphylla</italic> var. <italic>yunnanensis</italic>. The 1,789 data points of the original FT-MIR spectra were used as the X variable, and the classification labels were used as the <italic>Y</italic> variable. The training set was used to establish discriminant models, and the text set was used to externally verify the accuracy of models. The best kernel functions <italic>c</italic> and <italic>g</italic> were obtained by cross validation of grid search method. The SVM models were implemented using Matlab software.</p>
</sec>
<sec id="S2.SS7">
<title>Residual Neural Network</title>
<p>In this study, a 12-layer ResNet was established with a weight attenuation coefficient &#x03BB; of 0.0001 and a learning rate of 0.01. <xref ref-type="supplementary-material" rid="DS1">Supplementary Table 3</xref> showed the ResNet network parameter configuration. The model was completed by the anaconda data processing hardware platform, and MXNet was selected as the deep learning framework. The model contains two kinds of residual block, namely the identity residual block (<xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 2</xref>) and the convolutional residual block (<xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 3</xref>). The block is selected according to whether the dimensions of the input and output are consistent. When the dimensions of the input and output are the same, the identity residual block is used to build the model. When the input and output dimensions are inconsistent, we introduce the convolutional residual block with a convolution kernel size of 1 &#x00D7; 1 to match the dimensions of the input and output. The model structure is shown in <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 4</xref>, where the input data is synchronous 2DCOS, asynchronous 2DCOS and i2DCOS spectral images. The identification flow chart of ResNet is shown in <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 5</xref>. The training set is used to train the model. The Stochastic Gradient Descent (SGD) method is used to find the optimal parameters for minimizing the loss function value to obtain the optimal model. The test set is used to verify whether the performance of the final model is optimal. The external validation set is used to verify the generalization ability of the model.</p>
</sec>
</sec>
<sec sec-type="results" id="S3">
<title>Results and Discussion</title>
<sec id="S3.SS1">
<title>FT-MIR Spectra Analysis</title>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> shows the average FT-MIR spectra of four parts and five regions of PPY. 3,350, 2,940, 1,645, 1,387, 1,069, 931, 581 cm<sup>&#x2013;1</sup> are the main characteristic absorption peaks of PPY samples. The absorption peak of O-H stretching vibration is mainly near 3,350 cm<sup>&#x2013;1</sup> (<xref ref-type="bibr" rid="B27">Pei et al., 2018</xref>). The absorbance intensity around 2,940 cm<sup>&#x2013;1</sup> is related to the stretching vibration of C-H absorption of lipids (<xref ref-type="bibr" rid="B28">Pei et al., 2019</xref>). The absorption peak at 1,645 cm<sup>&#x2013;1</sup> is assigned to the C = C and C = O stretching vibration of steroid saponin and flavonoid (<xref ref-type="bibr" rid="B34">Wu et al., 2019</xref>). The absorption peak near 1,387 cm<sup>&#x2013;1</sup> is -CH<sub>3</sub> symmetrical bending vibration (<xref ref-type="bibr" rid="B37">Yang et al., 2019</xref>). In the region of 1,300&#x2013;550 cm<sup>&#x2013;1</sup>, the absorption peaks correspond to the stretching vibration peak of C-O and the bending vibration of O-H, which belong to substances such as sugars and saponins (<xref ref-type="bibr" rid="B35">Wu et al., 2018</xref>). It is concluded that the main components in the plant of PPY are flavonoids, starch and glycosides.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Averaged raw spectra of <italic>Paris polyphylla</italic> var. <italic>yunnanensis.</italic> <bold>(A)</bold> parts; <bold>(B)</bold> regions. The G, J, Y, and XG represent the rhizome (G), stem (J), leaf (Y) and fibrous root (XG), respectively.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g003.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="F3">Figure 3A</xref>, the absorption peak intensity of rhizome, stem, leaf and fibrous root is significantly different, especially the absorption peak in the band of 4,000&#x2013;1,200 cm<sup>&#x2013;1</sup>. On the whole, the order of absorption intensity of four parts is Y &#x003E; J &#x003E; XG &#x003E; G. It may imply that the distribution and content of active components in different parts of PPY are significantly different, and the components content of non-medicinal parts (Y, J, and XG) may be higher than the medicinal parts (G), which is nearly consistent with the research results of <xref ref-type="bibr" rid="B5">Feng et al. (2015)</xref>. However, the differences of peak shape and absorption intensity in different regions (<xref ref-type="fig" rid="F3">Figure 3B</xref>) are much lower than those in different parts, which indicates that the differences within individuals may be greater than the differences between individuals, and it&#x2019;s easier to identify parts than regions. Nonetheless, further modeling analysis and more studies are needed to support this conclusion.</p>
</sec>
<sec id="S3.SS2">
<title>The Two-Dimensional Correlation Spectroscopy Spectra Images</title>
<p>In this study, a total of 6,135 2DCOS images were drawn, including synchronous 2DCOS, asynchronous 2DCOS and i2DCOS images of PPY in different parts (<xref ref-type="fig" rid="F4">Figure 4</xref>) and different regions (<xref ref-type="fig" rid="F5">Figure 5</xref>). The synchronous 2DCOS images are symmetric along diagonals, and the correlation peaks may appear on or off the diagonal. The correlation peak on the diagonal line is called the auto peak, which is expressed as the value of the auto-correlation function of spectral intensity change (<xref ref-type="bibr" rid="B8">Huang et al., 2003</xref>). The peaks on both sides of the diagonal are called cross peaks and represent synchronous changes of spectral signals at different wavenumbers. The asynchronous 2DCOS images characterize the asynchronous characteristics of the absorption intensity measured at two different wavenumbers. It is anti-symmetric on both sides of the diagonal, and it has only cross peaks and no automatic peaks (<xref ref-type="bibr" rid="B17">Noda, 1990</xref>). The i2DCOS is defined as the product of the synchronous and asynchronous two-dimensional correlation intensities. It can provide correlation spectra with equal resolution, and its characteristics are clearer than asynchronous 2DCOS (<xref ref-type="bibr" rid="B32">van der Maaten and Hinton, 2008</xref>). By comparing the synchronous, asynchronous and integrated 2DCOS, it is not difficult to see that the colors and lines of the synchronous images are clearer and richer, and it is easy to analyze the differences and intensity changes of auto peaks and cross peaks between different samples. However, asynchronous and integrated images are complex and changeable, and cannot be distinguished by naked eyes. This may be caused by the complex characteristics of traditional Chinese medicine. In addition, the 2DCOS images of different parts has more significant differences than that of different regions, which is consistent with the results presented by the one-dimensional spectral analysis.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>The synchronous, asynchronous and integrated 2DCOS images of parts. <bold>(A)</bold> rhizome; <bold>(B)</bold> stem; <bold>(C)</bold> leaf; <bold>(D)</bold> fibrous root. Asys images are i2DCOS images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g004.tif"/>
</fig>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>The synchronous, asynchronous and integrated 2DCOS images of regions. <bold>(A)</bold> central; <bold>(B)</bold> northwest; <bold>(C)</bold> southeast; <bold>(D)</bold> southwest; <bold>(E)</bold> western. Asys images are i2DCOS images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g005.tif"/>
</fig>
<p>In summary, synchronous 2DCOS has better performance of visual recognition. Different parts are easier to distinguish than different regions. Although 2DCOS overcame the shortcomings of one-dimensional spectral peak overlap and improved its apparent resolution, it was very difficult to recognize different parts and regions by visual analysis alone, so we need to rely on machine learning methods.</p>
</sec>
<sec id="S3.SS3">
<title>Exploratory Analysis of t-Distributed Stochastic Neighbor Embedding</title>
<p>As a relatively novel non-parametric dimensionality reduction technology, t-SNE can visualize high-dimensional data to obtain the position of each data point on a two-dimensional or three-dimensional map. Its focus is to maintain the basic structure of the data matrix to reveal outliers or similarities and differences between groups of observed variables. As shown in <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 6</xref>, t-SNE was used in this study to conduct a preliminary visual evaluation of the spectral data sets. The ellipses in the figure represented the detailed trends of different types of samples. <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 6A</xref> showed the distribution of FT-MIR data sets of different parts, in which there were obvious outliers in both fibrous roots and roots. But in general, most samples could be clustered according to different category, and a few samples were mixed together. <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 6B</xref> showed the distribution of FT-MIR data sets of different regions, which formed a sharp contrast with the data set of different regions. The samples from the five regions were almost completely blended together. The two-dimensional visual results showed that the FT-MIR information of PPY samples in different regions was relatively similar, and it is not easy to distinguish. The results of these exploratory data analysis were consistent with the results of spectrum analysis, that is, the difference between different parts of PPY was higher than that of different regions. Obviously, in the process of data visualization, the vast majority of samples cannot be classified according to their pre-identified labels of different sources. Therefore, further in-depth modeling analysis should be considered.</p>
</sec>
<sec id="S3.SS4">
<title>Discrimination Results of Partial Least Squares-Discriminant Analysis Model</title>
<p>The PLS-DA models for the parts and regions of PPY based on different sample size data sets were, respectively, established. <xref ref-type="table" rid="T1">Table 1</xref> lists all the model parameters and the results of discrimination accuracy. From the table, we can clearly know that the models of different parts, different regions and different sample sizes have significant differences in the identification ability and model performance. In addition, in order to assess whether the PLS-DA model has an over-fitting problem, a permutation test was performed on all models. Generally, if the intercept of R<sup>2</sup> is less than 0.4, there is no risk of over-fitting. <xref ref-type="supplementary-material" rid="DS1">Supplementary Figure 7</xref> shows the results of the permutation test of five classification models (PLS-DA model cannot be established based on the low sample size data of the region). The results show that the R<sup>2</sup> intercepts of the five models are all less than 0.4, and there is no risk of over-fitting. The confusion matrices of the established PLS-DA models based on the data set of parts and regions are shown in <xref ref-type="supplementary-material" rid="DS1">Supplementary Tables 4</xref>, <xref ref-type="supplementary-material" rid="DS1">5</xref>, respectively.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Parameters for PLS-DA models in parts and regions discrimination based on three levels of data sets.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Data</td>
<td valign="top" align="center">Model</td>
<td valign="top" align="center">LVs</td>
<td valign="top" align="center">R<sup>2</sup></td>
<td valign="top" align="center">Q<sup>2</sup></td>
<td valign="top" align="center">RMSEE</td>
<td valign="top" align="center">RMSECV</td>
<td valign="top" align="center">RMSEP</td>
<td valign="top" align="center" colspan="2">Accuracy (%)<hr/></td>
</tr>
<tr>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td valign="top" align="center">Training set</td>
<td valign="top" align="center">Test set</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Parts</bold></td>
<td valign="top" align="center">PLS-DA-L</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.198</td>
<td valign="top" align="center">0.159</td>
<td valign="top" align="center">0.374135</td>
<td valign="top" align="center">0.37687</td>
<td valign="top" align="center">0.295164</td>
<td valign="top" align="center">51.52</td>
<td valign="top" align="center">55.56</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">PLS-DA-M</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">0.899</td>
<td valign="top" align="center">0.831</td>
<td valign="top" align="center">0.143237</td>
<td valign="top" align="center">0.167712</td>
<td valign="top" align="center">0.0758287</td>
<td valign="top" align="center">99.39</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">PLS-DA-H</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">0.918</td>
<td valign="top" align="center">0.887</td>
<td valign="top" align="center">0.120499</td>
<td valign="top" align="center">0.138129</td>
<td valign="top" align="center">0.0669199</td>
<td valign="top" align="center">99.39</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Regions</bold></td>
<td valign="top" align="center">PLS-DA-L</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">PLS-DA-M</td>
<td valign="top" align="center">14</td>
<td valign="top" align="center">0.584</td>
<td valign="top" align="center">0.333</td>
<td valign="top" align="center">0.349237</td>
<td valign="top" align="center">0.441024</td>
<td valign="top" align="center">0.325103</td>
<td valign="top" align="center">87.92</td>
<td valign="top" align="center">88.46</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">PLS-DA-H</td>
<td valign="top" align="center">20</td>
<td valign="top" align="center">0.698</td>
<td valign="top" align="center">0.544</td>
<td valign="top" align="center">0.266242</td>
<td valign="top" align="center">0.351347</td>
<td valign="top" align="center">0.266231</td>
<td valign="top" align="center">95.34</td>
<td valign="top" align="center">92.22</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>First of all, from the models based on different parts of the data set, we can see that the R<sup>2</sup> and Q<sup>2</sup> of the PLS-DA-L model are only 0.198 and 0.159, respectively, which are both lower than 0.5, and the recognition accuracy of the test set is only 55.56%. Therefore, the model based on the low sample size data set has poor performance and low discrimination ability, and cannot realize the discrimination of different parts of PPY. The PLS-DA-M and PLS-DA-H models based on the data sets of parts have high R<sup>2</sup> and Q<sup>2</sup> values greater than 0.8 and low RMSEE, RMSECV and RMSEP values. The accuracy of the test sets of the two models is 100%, which has a very good recognition performance.</p>
<p>Secondly, as shown in the table, the PLS-DA-L model based on regions data cannot be fitted. This result may be related to the amount of data being too small or the data is not preprocessed. Although the PLS-DA-M model has a test set accuracy rate of 88.46%, the model performance is poor with low Q<sup>2</sup> and high RMSEE, RMSECV, and RMSEP values. The PLS-DA-H model is better than the low sample size model and the medium sample size model in terms of model performance and recognition accuracy, so that it can well identify PPY in different regions.</p>
<p>Finally, from the perspective of sample size, whether it is PLS-DA models based on part data or models based on region data, the recognition performance is dependent on the sample size. And it shows that the larger the sample size, the better the model performance and the stronger the recognition ability. However, with the increase of the sample size, the recognition efficiency of the models will be greatly reduced. In addition, through comparison, it can be concluded that the PLS-DA models based on part data is better than that based on region data, regardless of model parameter results or recognition accuracy.</p>
</sec>
<sec id="S3.SS5">
<title>Discrimination Results of Support Vector Machine Model</title>
<p>Support vector machine is a supervised classification tool. It searches for the optimal separation hyperplane between different data categories by maximizing the distance between the classification hyperplane and various sample points. SVM contains two parameters, <italic>c</italic> is used as a penalty parameter, which can control the generalization ability of the model and reduce the over-fitting phenomenon, and the kernel function parameter <italic>g</italic> is related to the stability of the model. <xref ref-type="supplementary-material" rid="DS1">Supplementary Figures 8</xref>, <xref ref-type="supplementary-material" rid="DS1">9</xref> are the optimal separation hyperplane graph and classification result graph of the SVM model based on parts and regions data, respectively. The detailed results of the six SVM models are shown in <xref ref-type="table" rid="T2">Table 2</xref>. Best <italic>c</italic> and Best <italic>g</italic>, respectively, represent the best penalty parameter and kernel function parameter of the model.</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>The accuracy of SVM models for parts and regions identification based on three levels of data sets.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Data</td>
<td valign="top" align="center">Model</td>
<td valign="top" align="center">Best <italic>c</italic></td>
<td valign="top" align="center">Best <italic>g</italic></td>
<td valign="top" align="center" colspan="2">Accuracy (%)<hr/></td>
</tr>
<tr>
<td/>
<td/>
<td/>
<td/>
<td valign="top" align="center">Training set</td>
<td valign="top" align="center">Test set</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Parts</bold></td>
<td valign="top" align="center">SVM-L</td>
<td valign="top" align="center">2,048.00</td>
<td valign="top" align="center">0.000043</td>
<td valign="top" align="center">72.73</td>
<td valign="top" align="center">100.00</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">SVM-M</td>
<td valign="top" align="center">181.02</td>
<td valign="top" align="center">0.00069</td>
<td valign="top" align="center">98.18</td>
<td valign="top" align="center">100.00</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">SVM-H</td>
<td valign="top" align="center">5.66</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">99.39</td>
<td valign="top" align="center">100.00</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Regions</bold></td>
<td valign="top" align="center">SVM-L</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">0.10</td>
<td valign="top" align="center">0.00</td>
<td valign="top" align="center">46.15</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">SVM-M</td>
<td valign="top" align="center">11,585.24</td>
<td valign="top" align="center">0.00017</td>
<td valign="top" align="center">87.92</td>
<td valign="top" align="center">92.31</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">SVM-H</td>
<td valign="top" align="center">46,340.95</td>
<td valign="top" align="center">0.000031</td>
<td valign="top" align="center">94.17</td>
<td valign="top" align="center">97.28</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The accuracy difference between the training set and the test set of the SVM-L model based on part data and region data is more than 20%, while the accuracy of the training set and the test set of the SVM-M and SVM-H models based on part data and region data is less than 5%. This shows that the reliability of the SVM models established with low sample size data is poor. The SVM-M and SVM-H models based on part data both have high identification accuracy and low Best <italic>c</italic> value, so the model performance is good and have the ability to identify different parts of PPY. However, although the SVM-M and SVM-H models based on region data have high identification accuracy, their Best <italic>c</italic> values are abnormally high, indicating that the performance of the two models is poor and there may be over-fitting, which can&#x2019;t well identify the PPY in different regions. The above results show that although a larger sample size can improve the identification accuracy of the SVM model, the establishment of a high-performance model cannot be achieved for data that has not been preprocessed and has small differences between different categories. In addition, as with the results of the PLS-DA model, it is easier to identify the parts of PPY than the regions.</p>
<p>In conclusion, although the SVM model has the advantage of solving the problems of small sample, nonlinear and high-dimensional data (<xref ref-type="bibr" rid="B15">Noble, 2006</xref>), the unpreprocessed small sample data in this study is not applicable to the SVM model, indicating that data preprocessing is very necessary to improve the discrimination performance of traditional models such as SVM. In addition, a larger sample size increases the over-fitting risk of SVM model while improving the recognition accuracy, which leads to poor model performance and low reliability.</p>
</sec>
<sec id="S3.SS6">
<title>Discrimination Results of Residual Neural Network Model</title>
<p>In this research, ResNet models based on 2DCOS images (including synchronous, asynchronous and integrated images) of FT-MIR were established. <xref ref-type="fig" rid="F6">Figures 6</xref>, <xref ref-type="fig" rid="F7">7</xref> are the results of 18 ResNet models based on the data sets of parts and regions, respectively, showing the accuracy curves and cross-entropy cost function curves. The accuracy curves, includes the training set and the test set, were used to evaluate the discrimination ability of the model. The closer its value is to 1, the stronger the discrimination ability of the model. The cross-entropy loss function was used to explain the convergence effect of the model. The closer its value is to zero, the better the convergence effect of the model. In addition, the external validation set was classified using the models established above, and the classification result of the external validation set of different parts and regions was shown in the confusion matrix in <xref ref-type="supplementary-material" rid="DS1">Supplementary Figures 10</xref>, <xref ref-type="supplementary-material" rid="DS1">11</xref>, respectively. External validation is used to judge and evaluate the pros and cons of the model to ensure the stability of the established model. <xref ref-type="table" rid="T3">Table 3</xref> summarized the result parameters of all models, including accuracy (training set, test set and external validation set), epoch, and loss value.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>The accuracy curves and cross-entropy cost function of ResNet models based on part data with different sample size. L, low sample size; M, medium sample size; H, high sample size.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g006.tif"/>
</fig>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>The accuracy curves and cross-entropy cost function of ResNet models based on region data with different sample size. L, low sample size; M, medium sample size; H, high sample size.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g007.tif"/>
</fig>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>The accuracy of ResNet models for parts and regions identification based on three levels of data sets.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Data</td>
<td valign="top" align="center">Code</td>
<td valign="top" align="center">Type</td>
<td valign="top" align="center">Epoch</td>
<td valign="top" align="center">Loss value</td>
<td valign="top" align="center" colspan="3">Accuracy<hr/></td>
</tr>
<tr>
<td/>
<td/>
<td/>
<td/>
<td/>
<td valign="top" align="center">Train (%)</td>
<td valign="top" align="center">Test (%)</td>
<td valign="top" align="center">External validation (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Parts</bold></td>
<td valign="top" align="center">Resnet-L</td>
<td valign="top" align="center"><bold>Synchronous</bold></td>
<td valign="top" align="center"><bold>29</bold></td>
<td valign="top" align="center"><bold>0.091</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asynchronous</td>
<td valign="top" align="center">49</td>
<td valign="top" align="center">0.102</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asys</td>
<td valign="top" align="center">49</td>
<td valign="top" align="center">0.219</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Resnet-M</td>
<td valign="top" align="center"><bold>Synchronous</bold></td>
<td valign="top" align="center"><bold>29</bold></td>
<td valign="top" align="center"><bold>0.012</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asynchronous</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">0.021</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">88</td>
<td valign="top" align="center">87.5</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asys</td>
<td valign="top" align="center">49</td>
<td valign="top" align="center">0.021</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Resnet-H</td>
<td valign="top" align="center"><bold>Synchronous</bold></td>
<td valign="top" align="center"><bold>29</bold></td>
<td valign="top" align="center"><bold>0.009</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asynchronous</td>
<td valign="top" align="center">49</td>
<td valign="top" align="center">0.027</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">89</td>
<td valign="top" align="center">90</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asys</td>
<td valign="top" align="center">49</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">81</td>
<td valign="top" align="center">88</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Regions</bold></td>
<td valign="top" align="center">Resnet-L</td>
<td valign="top" align="center"><bold>Synchronous</bold></td>
<td valign="top" align="center"><bold>29</bold></td>
<td valign="top" align="center"><bold>0.114</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asynchronous</td>
<td valign="top" align="center">47</td>
<td valign="top" align="center">0.248</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">25</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asys</td>
<td valign="top" align="center">69</td>
<td valign="top" align="center">0.132</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">54</td>
<td valign="top" align="center">37.5</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Resnet-M</td>
<td valign="top" align="center"><bold>Synchronous</bold></td>
<td valign="top" align="center"><bold>29</bold></td>
<td valign="top" align="center"><bold>0.030</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asynchronous</td>
<td valign="top" align="center">49</td>
<td valign="top" align="center">0.088</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">62</td>
<td valign="top" align="center">56.4</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asys</td>
<td valign="top" align="center">69</td>
<td valign="top" align="center">0.045</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">61</td>
<td valign="top" align="center">66.7</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Resnet-H</td>
<td valign="top" align="center"><bold>Synchronous</bold></td>
<td valign="top" align="center"><bold>27</bold></td>
<td valign="top" align="center"><bold>0.009</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
<td valign="top" align="center"><bold>100</bold></td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asynchronous</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">0.011</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">63</td>
<td valign="top" align="center">62.7</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Asys</td>
<td valign="top" align="center">47</td>
<td valign="top" align="center">0.020</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">55</td>
<td valign="top" align="center">64</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Note: The bold value are the optimal results of models under the certain data set.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>Comparing the models based on synchronous, asynchronous and integrated 2DCOS images, we can get that the model of synchronous 2DCOS images has the best discrimination effect, and the accuracy of the training set, test set and external verification set is 100%. The modeling results are consistent with the results of image vision analysis, that is, the synchronized 2DCOS images have clearer characteristic peaks and can better characterize different types of samples. Comparing the models with low, medium and high sample sizes showed that the ResNet model had no dependence on the sample size, and there was no obvious rule between the identification accuracy and the sample size. However, too small sample size will lead to poor performance and over-fitting of model. This result can be derived from the identification results of low sample size models based on asynchronous and integrated 2DCOS images. The difference of identification accuracy between the external validation set and the test set was large, and the loss value of models was significantly higher than that of the medium sample size and high sample size models. In addition, the accuracy curves of the training set and test set of the medium sample size and high sample size models showed a consistent upward trend, which also showed that these two types of models had no risk of over-fitting and were robust. However, the accuracy curve of the training set and the test set of the low sample size model had a poor consistency in the upward trend, even for the optimal model of synchronous 2DCOS images, which indicated that the low sample size would reduce the performance of the ResNet model. Finally, on the whole, the recognition effect of the ResNet model based on the part data set was better than that of the ResNet model based on the region data set.</p>
<p>In summary, the recognition accuracy of the models based on synchronous 2DCOS images is the best, which is almost not affected by sample size, part, region and other factors, and is most suitable for the identification of medicinal plants. However, too small sample size does have a small negative impact on the performance of the ResNet model. Therefore, it is worth thinking about how to use an appropriate method to solve the negative impact of low samples on model performance. This is conducive to solving the identifying problem of research subjects with a small sample size. These research objects have very limited data, and it is expensive or impossible to obtain more data, such as scarce and precious animal and plant resources.</p>
</sec>
<sec id="S3.SS7">
<title>Comparison Analysis of Models</title>
<p>Partial least squares discriminant analysis, SVM, and ResNet models showed significant differences in their ability to identify the parts and regions of the PPY, the responses to different sample sizes, and the comprehensive performance of models. As shown in <xref ref-type="fig" rid="F8">Figure 8</xref>, we have made a visual comparison of three type of models.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption><p>Comparison of the overall identification performance of PLS-DA, SVM and ResNet models. <bold>(A)</bold> parts; <bold>(B)</bold> regions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-12-752863-g008.tif"/>
</fig>
<p>In terms of the identification ability of parts and regions, the three types of models show consistent results, that is, the identification ability of parts is better than that of regions, which indicates that the difference of parts data of PPY is greater than that of regions data. This result implies that the difference in component within the sample may be greater than that between samples. This causes us to think about the resource evaluation and the effective development and utilization of the non-medicinal parts of PPY. In addition to the evaluation of the advantages and disadvantages of the medicinal parts between individuals in different origins, the development and utilization of non-medicinal parts within individuals is also very worthy of attention.</p>
<p>From the perspective of different sample sizes, the three models have different responses to low, medium, and high sample size data. The PLS-DA model has a very significant sample size dependence. As the sample size increases, the discrimination ability and the performance of the model have been significantly improved. It can be concluded that the overall performance of the PLS-DA model is positively correlated with the sample size. This result is confirmed by two types of models based on part and region data, which greatly reduces the chance. There is a certain correlation between the merits and demerits of SVM model and the sample size, but not a complete positive or negative correlation. The identification accuracy of the model increases with the increase of the sample size, while the performance of the model based on region data evaluated by parameters will deteriorate with the increase of the sample size. It can be concluded from this study that there are two important factors affecting the overall performance of SVM model, one is the quality of data itself, the other is the sample size. The ResNet model based on the synchronous 2DCOS images has a very perfect overall discrimination performance, both in terms of the discrimination accuracy and the model parameters. It is not limited by the sample size and is almost unaffected by the data itself. Whether it is based on easy-to-identify part data or region data with small differences, it can achieve 100% recognition accuracy.</p>
<p>In summary, the PLS-DA model has the strongest dependence on the sample size, followed by SVM, and the ResNet model based on synchronized 2DCOS images has almost no dependence on the sample size. In addition, the traditional pattern recognition model is also affected by the quality of data itself. Therefore, the ResNet model based on synchronized 2DCOS images occupies an absolute advantage in the identification of medicinal plants. The model is universal and does not require preprocessing or artificial extraction of characteristic variables. It has good discrimination accuracy regardless of the sample size or the quality of the data.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="S4">
<title>Conclusion</title>
<p>In this study, we used three kinds of models to identify the part and region of PPY. PLS-DA and SVM are traditional pattern recognition models, which have been widely used in the past research. ResNet model is a representative dominant model in deep learning. The effects of different types of data and different sample sizes on the discrimination ability and performance of the three models were discussed without any data preprocessing. By comparing the ability of the traditional model and the deep learning model for the identification of PPY, we found that the identification performance of PLS-DA and SVM models was easily affected by the data type, sample size and other factors, and the overall identification ability of both models was not as good as the ResNet model based on synchronous 2DCOS images. Different from the previous single theory or single model analysis, this study verified the superiority of deep learning model in the identification research of medicinal plant resources from the actual and multiple perspectives.</p>
</sec>
<sec sec-type="data-availability" id="S5">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusion of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="S6">
<title>Author Contributions</title>
<p>JY: conceptualization, software, formal analysis, writing&#x2013;original draft preparation, and writing&#x2014;review and editing. WL: methodology, resources, and software. YW: supervision, project administration, and funding acquisition. All authors have read and agreed to the published version of the manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="S7">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="S8">
<title>Funding</title>
<p>This work was supposed by National Natural Science Foundation of China (Grant Number: 31860584), the Reserve Talents of Young and Middle-Aged Academic Leaders in Yunnan Province (Grant Number: 202005AC160032), and the Major Science and Technology Projects in Yunnan Province of &#x201C;Digital Development and Application of Biological Resources&#x201D; (Grant Number: 202002AA100007).</p>
</sec>
<sec id="S9" sec-type="supplementary material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpls.2021.752863/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpls.2021.752863/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.docx" id="DS1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J. B.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Rong</surname> <given-names>L. X.</given-names></name> <name><surname>Wang</surname> <given-names>J. J.</given-names></name></person-group> (<year>2018</year>). <article-title>Integrative two-dimensional correlation spectroscopy (i2DCOS) for the intuitive identification of adulterated herbal materials.</article-title> <source><italic>J. Mol. Struct.</italic></source> <volume>1163</volume> <fpage>327</fpage>&#x2013;<lpage>335</lpage>. <pub-id pub-id-type="doi">10.1016/j.molstruc.2018.02.061</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cunningham</surname> <given-names>A. B.</given-names></name> <name><surname>Brinckmann</surname> <given-names>J. A.</given-names></name> <name><surname>Bi</surname> <given-names>Y. F.</given-names></name> <name><surname>Pei</surname> <given-names>S. J.</given-names></name> <name><surname>Schippmann</surname> <given-names>U.</given-names></name> <name><surname>Luo</surname> <given-names>P.</given-names></name><etal/></person-group> (<year>2018</year>). <article-title>Paris in the spring: a review of the trade, conservation and opportunities in the shift from wild harvest to cultivation of Paris polyphylla (<italic>Trilliaceae</italic>).</article-title> <source><italic>J. Ethnopharmacol.</italic></source> <volume>222</volume> <fpage>208</fpage>&#x2013;<lpage>216</lpage>. <pub-id pub-id-type="doi">10.1016/j.jep.2018.04.048</pub-id> <pub-id pub-id-type="pmid">29727736</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>Q.</given-names></name> <name><surname>Lang</surname> <given-names>T.</given-names></name> <name><surname>Xia</surname> <given-names>J. X.</given-names></name></person-group> (<year>2016</year>). <article-title>Present situation and development of medicinal plant resources&#x2019; utilization.</article-title> <source><italic>J. MUC</italic></source> <volume>25</volume> <fpage>55</fpage>&#x2013;<lpage>59</lpage>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>J. E.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Zuo</surname> <given-names>Z. T.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep learning for geographical discrimination of Panax notoginseng with directly near-infrared spectra image.</article-title> <source><italic>Chemometr. Intell. Lab. Syst</italic>.</source> <volume>197</volume>:<issue>103913</issue>. <pub-id pub-id-type="doi">10.1016/j.chemolab.2019.103913</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feng</surname> <given-names>L. L.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>H. F.</given-names></name> <name><surname>Zhang</surname> <given-names>C. G.</given-names></name></person-group> (<year>2015</year>). <article-title>Quality evaluation of Paris polyphylla var. yunnanensis and accumulation law analysis of its steroidal saponins.</article-title> <source><italic>Chin. J. Exp. Tradit. Med. Formul.</italic></source> <volume>21</volume> <fpage>41</fpage>&#x2013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.13422/j.cnki.syfjx.2015130041</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grinblat</surname> <given-names>G. L.</given-names></name> <name><surname>Uzal</surname> <given-names>L. C.</given-names></name> <name><surname>Larese</surname> <given-names>M. G.</given-names></name> <name><surname>Granitto</surname> <given-names>P. M.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep learning for plant identification using vein morphological patterns.</article-title> <source><italic>Comput. Electron. Agric.</italic></source> <volume>127</volume> <fpage>418</fpage>&#x2013;<lpage>424</lpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2016.07.003</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Houssein</surname> <given-names>E. H.</given-names></name> <name><surname>Emam</surname> <given-names>M. M.</given-names></name> <name><surname>Ali</surname> <given-names>A. A.</given-names></name> <name><surname>Suganthan</surname> <given-names>P. N.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep and machine learning techniques for medical imaging-based breast cancer: a comprehensive review.</article-title> <source><italic>Expert Syst. Appl.</italic></source> <volume>167</volume>:<issue>114161</issue>. <pub-id pub-id-type="doi">10.1016/j.eswa.2020.114161</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Malkov</surname> <given-names>S.</given-names></name> <name><surname>Coleman</surname> <given-names>M.</given-names></name> <name><surname>Painter</surname> <given-names>P.</given-names></name></person-group> (<year>2003</year>). <article-title>Application of two-dimensional correlation infrared spectroscopy to the study of miscible polymer blends.</article-title> <source><italic>Macromolecules</italic></source> <volume>36</volume> <fpage>8156</fpage>&#x2013;<lpage>8163</lpage>. <pub-id pub-id-type="doi">10.1021/ma0259463</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jamshidi-Kia</surname> <given-names>F.</given-names></name> <name><surname>Lorigooini</surname> <given-names>Z.</given-names></name> <name><surname>Amini-Khoei</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>Medicinal plants: past history and future perspective.</article-title> <source><italic>J. Herbmed Pharmacol.</italic></source> <volume>7</volume> <fpage>1</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.15171/jhp.2018.01</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning.</article-title> <source><italic>Nature</italic></source> <volume>521</volume> <fpage>436</fpage>&#x2013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id> <pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J. R.</given-names></name> <name><surname>Sun</surname> <given-names>S. Q.</given-names></name> <name><surname>Wang</surname> <given-names>X. X.</given-names></name> <name><surname>Xu</surname> <given-names>C. H.</given-names></name> <name><surname>Chen</surname> <given-names>J. B.</given-names></name> <name><surname>Zhou</surname> <given-names>Q.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Differentiation of five species of Danggui raw materials by FTIR combined with 2D-COS IR.</article-title> <source><italic>J. Mol. Struct.</italic></source> <volume>1069</volume> <fpage>229</fpage>&#x2013;<lpage>235</lpage>. <pub-id pub-id-type="doi">10.1016/j.molstruc.2014.03.067</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Zuo</surname> <given-names>Z. T.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name> <name><surname>Xu</surname> <given-names>F. R.</given-names></name></person-group> (<year>2020</year>). <article-title>A fast multi-source information fusion strategy based on FTIR spectroscopy for geographical authentication of wild Gentiana rigescens.</article-title> <source><italic>Microchem. J.</italic></source> <volume>159</volume>:<issue>105360</issue>. <pub-id pub-id-type="doi">10.1016/j.microc.2020.105360</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z. M.</given-names></name> <name><surname>Yang</surname> <given-names>M. Q.</given-names></name> <name><surname>Zuo</surname> <given-names>Y. M.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name> <name><surname>Zhang</surname> <given-names>J. Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Fraud detection of herbal medicines based on modern analytical technologies combine with chemometrics approach: a review.</article-title> <source><italic>Crit. Rev. Anal. Chem.</italic></source> [Epub Online ahead of print]. <pub-id pub-id-type="doi">10.1080/10408347.2021.1905503</pub-id> <pub-id pub-id-type="pmid">33840329</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newman</surname> <given-names>D. J.</given-names></name> <name><surname>Cragg</surname> <given-names>G. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Natural products as sources of new drugs from 1981 to 2014.</article-title> <source><italic>J. Nat. Prod.</italic></source> <volume>79</volume> <fpage>629</fpage>&#x2013;<lpage>661</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jnatprod.5b01055</pub-id> <pub-id pub-id-type="pmid">26852623</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noble</surname> <given-names>W. S.</given-names></name></person-group> (<year>2006</year>). <article-title>What is a support vector machine?</article-title> <source><italic>Nat. Biotechnol.</italic></source> <volume>24</volume> <fpage>1565</fpage>&#x2013;<lpage>1567</lpage>. <pub-id pub-id-type="doi">10.1038/nbt1206-1565</pub-id> <pub-id pub-id-type="pmid">17160063</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>1989</year>). <article-title>Two-dimensional infrared spectroscopy.</article-title> <source><italic>J. Am. Chem. Soc.</italic></source> <volume>111</volume> <fpage>8116</fpage>&#x2013;<lpage>8118</lpage>. <pub-id pub-id-type="doi">10.1021/ja00203a008</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>1990</year>). <article-title>Two-dimensional infrared (2D IR) spectroscopy: theory and applications.</article-title> <source><italic>Appl. Spectrosc.</italic></source> <volume>44</volume> <fpage>550</fpage>&#x2013;<lpage>561</lpage>. <pub-id pub-id-type="doi">10.1366/0003702904087398</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>1993</year>). <article-title>Generalized two-dimensional correlation method applicable to infrared, raman, and other types of spectroscopy.</article-title> <source><italic>Appl. Spectrosc.</italic></source> <volume>47</volume> <fpage>1329</fpage>&#x2013;<lpage>1336</lpage>. <pub-id pub-id-type="doi">10.1366/0003702934067694</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>2004</year>). <article-title>Advances in two-dimensional correlation spectroscopy.</article-title> <source><italic>Vib. Spectrosc.</italic></source> <volume>36</volume> <fpage>143</fpage>&#x2013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1016/j.vibspec.2003.12.016</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>2014</year>). <article-title>Frontiers of two-dimensional correlation spectroscopy. Part 1. New concepts and noteworthy developments.</article-title> <source><italic>J. Mol. Struct.</italic></source> <volume>1069</volume> <fpage>3</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1016/j.molstruc.2014.01.025</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>2016</year>). <article-title>Techniques useful in two-dimensional correlation and codistribution spectroscopy (2DCOS and 2DCDS) analyses.</article-title> <source><italic>J. Mol. Struct.</italic></source> <volume>1124</volume> <fpage>29</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1016/j.molstruc.2016.01.089</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noda</surname> <given-names>I.</given-names></name></person-group> (<year>2018</year>). <article-title>Two-trace two-dimensional (2T2D) correlation spectroscopy-a method for extracting useful information from a pair of spectra.</article-title> <source><italic>J. Mol. Struct.</italic></source> <volume>1160</volume> <fpage>471</fpage>&#x2013;<lpage>478</lpage>. <pub-id pub-id-type="doi">10.1016/j.molstruc.2018.01.091</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Obaid</surname> <given-names>H. S.</given-names></name> <name><surname>Dheyab</surname> <given-names>S. A.</given-names></name> <name><surname>Sabry</surname> <given-names>S. S.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>The Impact of data pre-processing techniques and dimensionality reduction on the accuracy of machine learning</article-title>,&#x201D; in <source><italic>9th Annual Information Technology, Electromechanical Engineering and Microelectronics Conference</italic></source>, (<publisher-loc>Jaipur, India</publisher-loc>: <publisher-name>IEEE</publisher-name>).</citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>X. H.</given-names></name> <name><surname>Song</surname> <given-names>J. Y.</given-names></name> <name><surname>Zhu</surname> <given-names>Y. J.</given-names></name> <name><surname>Xu</surname> <given-names>H. X.</given-names></name> <name><surname>Huang</surname> <given-names>L. F.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Applying plant DNA barcodes for Rosaceae species identification.</article-title> <source><italic>Cladistics</italic></source> <volume>27</volume> <fpage>165</fpage>&#x2013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1111/j.1096-0031.2010.00328.x</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pasquini</surname> <given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>Near infrared spectroscopy: a mature analytical technique with new perspectives - a review.</article-title> <source><italic>Anal. Chim. Acta</italic></source> <volume>1026</volume> <fpage>8</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1016/j.aca.2018.04.004</pub-id> <pub-id pub-id-type="pmid">29852997</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pei</surname> <given-names>Y. F.</given-names></name> <name><surname>Zhang</surname> <given-names>Q. Z.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Application of authentication evaluation techniques of ethnobotanical medicinal plant genus Paris: a review.</article-title> <source><italic>Crit. Rev. Anal. Chem.</italic></source> <volume>50</volume> <fpage>405</fpage>&#x2013;<lpage>423</lpage>. <pub-id pub-id-type="doi">10.1080/10408347.2019.1642734</pub-id> <pub-id pub-id-type="pmid">31357868</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pei</surname> <given-names>Y. F.</given-names></name> <name><surname>Zhang</surname> <given-names>Q. Z.</given-names></name> <name><surname>Zuo</surname> <given-names>Z. T.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2018</year>). <article-title>Comparison and identification for rhizomes and leaves of Paris yunnanensis based on Fourier transform mid-infrared spectroscopy combined with chemometrics.</article-title> <source><italic>Molecules</italic></source> <volume>23</volume>:<issue>3343</issue>. <pub-id pub-id-type="doi">10.3390/molecules23123343</pub-id> <pub-id pub-id-type="pmid">30563007</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pei</surname> <given-names>Y. F.</given-names></name> <name><surname>Zuo</surname> <given-names>Z. T.</given-names></name> <name><surname>Zhang</surname> <given-names>Q. Z.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2019</year>). <article-title>Data fusion of Fourier transform mid-infrared (MIR) and near-infrared (NIR) spectroscopies to identify geographical origin of wild Paris polyphylla var. yunnanensis.</article-title> <source><italic>Molecules</italic></source> <volume>24</volume>:<issue>2559</issue>. <pub-id pub-id-type="doi">10.3390/molecules24142559</pub-id> <pub-id pub-id-type="pmid">31337084</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>T.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Discrimination of Gentiana and its related species using IR spectroscopy combined with feature selection and stacked generalization.</article-title> <source><italic>Molecules</italic></source> <volume>25</volume>:<issue>1442</issue>. <pub-id pub-id-type="doi">10.3390/molecules25061442</pub-id> <pub-id pub-id-type="pmid">32210010</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>S. Q.</given-names></name> <name><surname>Zhou</surname> <given-names>Q.</given-names></name> <name><surname>Qin</surname> <given-names>Z.</given-names></name></person-group> (<year>2003</year>). <source><italic>Atlas Of Two-Dimensional Correlation Information Spectroscopy For Traditional Chinese Medicine Identification.</italic></source> <publisher-loc>Beijing</publisher-loc>: <publisher-name>Chemical Industry Press</publisher-name>.</citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tao</surname> <given-names>A. E.</given-names></name> <name><surname>Zhao</surname> <given-names>F. Y.</given-names></name> <name><surname>Li</surname> <given-names>R. S.</given-names></name> <name><surname>Qian</surname> <given-names>J. F.</given-names></name> <name><surname>Xia</surname> <given-names>C. L.</given-names></name></person-group> (<year>2020</year>). <article-title>Industrialization condition and development strategy of Paridis Rhizoma.</article-title> <source><italic>Chin. Tradit. Herb. Drugs</italic></source> <volume>51</volume> <fpage>4809</fpage>&#x2013;<lpage>4815</lpage>. <pub-id pub-id-type="doi">10.7501/j.issn.0253-2670.2020.18.026</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>van der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Visualizing data using t-SNE.</article-title> <source><italic>J. Mach. Learn. Res.</italic></source> <volume>9</volume> <fpage>2579</fpage>&#x2013;<lpage>2605</lpage>.</citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>H. Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Authentication of Dendrobium Officinale from similar species with infrared and ultraviolet-visible spectroscopies with data visualization and mining.</article-title> <source><italic>Anal. Lett.</italic></source> <volume>53</volume> <fpage>1774</fpage>&#x2013;<lpage>1793</lpage>. <pub-id pub-id-type="doi">10.1080/00032719.2020.1719126</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>X. M.</given-names></name> <name><surname>Zhang</surname> <given-names>Q. Z.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2019</year>). <article-title>Traceability the provenience of cultivated Paris polyphylla Smith var. ynnanensis using ATR-FTIR spectroscopy combined with chemometrics.</article-title> <source><italic>Spectrochim. Acta A Mol. Biomol. Spectrosc.</italic></source> <volume>212</volume> <fpage>132</fpage>&#x2013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1016/j.saa.2019.01.008</pub-id> <pub-id pub-id-type="pmid">30639599</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Zuo</surname> <given-names>Z. T.</given-names></name> <name><surname>Xu</surname> <given-names>F. R.</given-names></name> <name><surname>Wnag</surname> <given-names>Y. Z.</given-names></name> <name><surname>Zhang</surname> <given-names>J. Y.</given-names></name><etal/></person-group> (<year>2018</year>). <article-title>Rapid discrimination of the different processed Paris poly phylla var. yunnanensis with infrared spectroscopy combined with chemometrics.</article-title> <source><italic>Spectrosc. Spect. Anal.</italic></source> <volume>38</volume> <fpage>1101</fpage>&#x2013;<lpage>1106</lpage>. <pub-id pub-id-type="doi">10.3964/j.issn.1000-0593201804-1101-06</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y. G.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2018</year>). <article-title>Characterization of Paris polyphylla var. yunnanensis by infrared and ultraviolet spectroscopies with chemometric data fusion.</article-title> <source><italic>Anal. Lett.</italic></source> <volume>51</volume> <fpage>1730</fpage>&#x2013;<lpage>1742</lpage>. <pub-id pub-id-type="doi">10.1080/00032719.2017.1385618</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y. G.</given-names></name> <name><surname>Zhao</surname> <given-names>Y. L.</given-names></name> <name><surname>Zuo</surname> <given-names>Z. T.</given-names></name> <name><surname>Wang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2019</year>). <article-title>Determination of total flavonoids for Paris Polyphylla var. yunnanensis in different geographical origins using UV and FT-IR spectroscopy.</article-title> <source><italic>J. AOAC Int.</italic></source> <volume>102</volume> <fpage>45 7</fpage>&#x2013;<lpage>464</lpage>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>F. Y.</given-names></name> <name><surname>Tao</surname> <given-names>A. E.</given-names></name> <name><surname>Guan</surname> <given-names>X.</given-names></name> <name><surname>Qian</surname> <given-names>J. F.</given-names></name> <name><surname>Xia</surname> <given-names>C. L.</given-names></name></person-group> (<year>2021</year>). <article-title>Research progress on chemical constituents, pharmacological effects and resource utilization modes of non-medicinal parts of Paridis Rhizoma.</article-title> <source><italic>Chin. Tradit. Herb. Drugs</italic></source> <volume>52</volume> <fpage>2449</fpage>&#x2013;<lpage>2457</lpage>. <pub-id pub-id-type="doi">10.7501/j.issn.0253-2670.2021.08.030</pub-id></citation></ref>
</ref-list>
</back>
</article>
