<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Pharmacol.</journal-id>
<journal-title>Frontiers in Pharmacology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Pharmacol.</abbrev-journal-title>
<issn pub-type="epub">1663-9812</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1504230</article-id>
<article-id pub-id-type="doi">10.3389/fphar.2025.1504230</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Pharmacology</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>One-class modeling for verification of botanical identity: a review</article-title>
<alt-title alt-title-type="left-running-head">Harnly</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fphar.2025.1504230">10.3389/fphar.2025.1504230</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Harnly</surname>
<given-names>James</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2854763/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/resources/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/Writing - review &#x26; editing/"/>
</contrib>
</contrib-group>
<aff>
<institution>Methods and Applications Food Composition Lab</institution>, <institution>Beltsville Human Nutrition Research Center</institution>, <institution>Agricultural Research Service</institution>, <institution>U.S. Department of Agriculture</institution>, <addr-line>Beltsville</addr-line>, <addr-line>MD</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/494152/overview">Karim Hosni</ext-link>, Institut National de Recherche et d&#x2019;Analyse Physico-Chimique (INRAP), Tunisia</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/211045/overview">Marjan Vracko</ext-link>, National Institute of Chemistry, Slovenia</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2634816/overview">Bob Allkin</ext-link>, Royal Botanic Gardens, United Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: James Harnly, <email>James.Harnly@usda.gov</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>03</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>16</volume>
<elocation-id>1504230</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>09</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>02</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Harnly.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Harnly</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>One-class modeling is a supervised multivariate botanical identification method based on principal component analysis (PCA) that constructs a model based only on the characteristics of the reference samples and uses the Q statistic as a combined metric. Test samples are judged to be similar (authentic) if their combined metric falls within the model limits or different (adulterated or contaminated) if the metric falls outside the model limits. This review initially considers three major factors affecting identification: the number of variables (univariate versus multivariate), the number of classes (one-class versus multi-class), and the type of analysis (quantitative versus qualitative). Multivariate analysis is commonly used for identification, providing a broader coverage of the identity specifications of the samples. With a combined metric, multivariate methods are analogous to univariate methods. One-class modeling and multi-class modeling employ different approaches for identification with one-class modeling being more flexible. While most methods to date have had a quantitative basis, qualitative methods are possible. This review focuses on multivariate, one-class modeling based on PCA. Examples are presented for the application of one-class modeling to identification of American ginseng (<italic>Panax quinquefolius</italic>), <italic>Echinacea purpurea</italic>, Black Cohosh (<italic>Actaea racemosa</italic>), and Maca (<italic>Lepidium meyenii</italic>). These examples demonstrate the utility and flexibility of one-class modeling.</p>
</abstract>
<abstract abstract-type="graphical">
<title>Graphical Abstract</title>
<p>
<graphic xlink:href="FPHAR_fphar-2025-1504230_wc_abs.tif"/>
</p>
</abstract>
<kwd-group>
<kwd>PCA</kwd>
<kwd>one-class modeling</kwd>
<kwd>review</kwd>
<kwd>authentication</kwd>
<kwd>multivariate analysis</kwd>
</kwd-group>
<contract-sponsor id="cn001">Office of Dietary Supplements<named-content content-type="fundref-id">10.13039/100000063</named-content>
</contract-sponsor>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Ethnopharmacology</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The goal of botanical identification is conceptually simple as shown by the graphical abstract; determine whether a test sample has features that are similar to those for a set of reference samples. This approach assumes that valid reference materials exist, and that identification can be made through a direct comparison of the test sample with a set of reference materials. The features used for a comparison can be sensory, morphological, microscopic, genetic, or chemical. The reference samples must be authenticated botanical reference materials or vouchered materials obtained from reliable sources. The authenticity of the reference materials is critical to the validity of the comparison. Multiple reference samples are needed to account for genetic, environmental, and processing variability. A single reference material from a metrological source can indicate similarity but fails to provide any information on biological variability.</p>
<p>Classically, herbalists, botanists, and taxonomists have authenticated plant materials based on monographs that describe a series of observations and tests developed using vouchered samples. In general, this approach requires access to the full plant material and flower. At one time, all commercial botanical materials were wild crafted by an expert and supplied directly to the user or to a local distributor. Today, botanical supplements are a multi-billion dollar industry and the link between the supplier and the manufacturer is much more tenuous. Manufacturers of botanical supplements are often faced with the problem of verifying that the barrel of brown powder delivered on the back dock by their supplier is really <italic>ginkgo biloba</italic>.</p>
<p>Commercial botanical supplements are classified as a food and are only regulated by FDA&#x2019;s current good manufacturing practices (cGMPs) which require the manufacturer to describe the steps they took to test the purity of their ingredients and products (<xref ref-type="bibr" rid="B2">Food and Drug Administration and Health and Human Services, 2024</xref>). Since the ingredients have often lost their morphological integrity (e.g., they may be sold as powders or extracts), chemical methods are most frequently used for identification although genetic methods are becoming more popular with recent advances in technology. Chemical methods provide quantitative multivariate data in the form of chromatograms or spectra that can be used for targeted or non-targeted analysis (<xref ref-type="bibr" rid="B16">Nichani et al., 2023a</xref>; <xref ref-type="bibr" rid="B15">Nichani et al., 2023b</xref>). However, these methods can only be used for identification by comparison to appropriate reference materials.</p>
<p>The need for identification methods led AOAC International to develop &#x201c;Guidelines for Validation of Botanical Identification Methods&#x201d; in 2012 (<xref ref-type="bibr" rid="B4">Harnly, 2012</xref>). The Guidelines describe a probability of identification (POI) method, establish a well defined nomenclature (<xref ref-type="table" rid="T1">Table 1</xref>), and describe several basic principals. First, the guidelines recognize the use of both quantitative chemical methods and qualitative morphological methods. Second, the guidelines assume a multivariate analysis and specify that the chosen method must reduce multiple observations or measurements to a combined metric. Next, the metric is used to generate a binary response, &#x201c;yes&#x201d; the sample is authentic or &#x2018;no&#x201d; it is adulterated. Finally, the POI method is a two-class analysis method that requires comparison of an authentic and an adulterated sample (<xref ref-type="table" rid="T2">Table 2</xref>) (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). Since the Guidelines are not method specific, 60 analyses of the authentic and adulterated samples are required to guarantee 95% confidence in discriminating between the two populations (<xref ref-type="table" rid="T3">Table 3</xref>). A companion paper to the Guidelines presented an example of a quantitative chemical analysis using mass spectrometric data (<xref ref-type="bibr" rid="B11">LaBudde and Harnly, 2012</xref>). Unfortunately, an example based on a qualitative morphological analysis was not given and has not been forth coming.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Glossary (<xref ref-type="bibr" rid="B11">LaBudde and Harnly, 2012</xref>).</p>
</caption>
<table>
<tbody valign="top">
<tr>
<td align="left">Botanical: Of, or relating to, plants or botany. May also include algae and fungi. May refer to the whole plant, a part of the plant (e.g., bark, woods, leaves, stems, roots, rhizomes, flowers, fruits, seeds, etc.), or an extract of the parts.<break/>Botanical identification method (BIM): A method that establishes identity specifications for a botanical material and determines YES, the test material is a true example of the target botanical material or NO, it is not the target botanical. <break/>Combined Metric: (analytical parameter in previous papers) A measured or computed analytical value used to determine whether the test material matches the target material. The combined metric may be based on morphological features, genetic sequences, chromatographic patterns, spectral patterns, or any other metric appropriate for the target material.<break/>Exclusivity Panel: A list of practically obtainable botanical materials that that are expected to give a negative result when tested by the BIM.<break/>Identity specification (IS): The morphological, genetic, chemical, or other characteristics that define a target botanical material. Specifications may include, but are not limited to, data from macroscopic, microscopic, genetic (e.g., DNA sequencing), chromatographic fingerprinting (e.g., capillary electrophoresis, gas chromatography, liquid chromatography, or thin-layer chromatography), and spectral fingerprinting (e.g., infrared, near-infrared, nuclear magnetic resonance, ultraviolet/visible absorbance, or mass spectrometry) methods.<break/>Inclusivity Panel: A list of practically obtainable botanical materials that are expected to give a positive result when tested by the BIM.<break/>Non-target botanical material: Any botanical material that does not meet the identity specification.<break/>Probability of identification (POI): The expected or observed fraction of test portions at a given concentration that give a positive result when tested by the BIM.<break/>Sample: A small portion or quantity, taken from a population or lot that is ideally a representative selection of the whole.<break/>Sensitivity: Ability of a BIM to correctly identify variants of the target material that meet the identity specification.<break/>Specificity: Ability of a BIM to correctly reject nontarget botanical materials.<break/>Standard inferior test material (SITM): A botanical material mixture that has the maximum concentration of target material that is considered unacceptable, as specified by the SMPRs. The BIM must reject this material. <break/>Standard method performance requirements (SMPRs): Performance requirements based on the fitness-for-purpose statement for each method. For BIMs, the SMPRs should include the physical form of the sample, the ISF, the ESF, the SSTM, the SITM, the number of samples for the inclusivity/ exclusivity panels, and the desired probability and confidence limits for the method.<break/>Standard superior test material (SSTM): A botanical material mixture that has the minimum acceptable concentration of the target material, as specified by the SMPR. The BIM must accept this material.<break/>Target botanical material: The botanical material of interest as described in the identity specification.</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Chemometric methods (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Approach</th>
<th align="center">Method<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</th>
<th align="center">Supervision</th>
<th align="center">Classes</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Exploratory</td>
<td align="center">PCA</td>
<td align="center">None</td>
<td align="center">None</td>
</tr>
<tr>
<td rowspan="2" align="left">One Class Modeling (soft modeling)</td>
<td align="center">PCA</td>
<td align="center">ID 1 class</td>
<td align="center">1</td>
</tr>
<tr>
<td align="center">SIMCA</td>
<td align="center">ID all classes</td>
<td align="center">2 or more</td>
</tr>
<tr>
<td rowspan="5" align="left">Two Class Classification (hard modeling)</td>
<td align="center">LDA</td>
<td align="center">ID all classes</td>
<td align="center">2 or more</td>
</tr>
<tr>
<td align="center">QDA</td>
<td align="center">ID all classes</td>
<td align="center">2 or more</td>
</tr>
<tr>
<td align="center">PLS-DA</td>
<td align="center">ID all classes</td>
<td align="center">2 or more</td>
</tr>
<tr>
<td align="center">SVM</td>
<td align="center">ID all classes</td>
<td align="center">2 or more</td>
</tr>
<tr>
<td align="center">ANN</td>
<td align="center">ID all classes</td>
<td align="center">2 or more</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="Tfn1">
<label>
<sup>a</sup>
</label>
<p>ANN, Artificial neural networks; LDA, Linear discriminate analysis; PLS-DA, Partial least squares-discrimination analysis; PCA, Principal component analysis; QDA, Quadratic discriminate analysis; SIMCA, Soft independent modeling of class analogy; SVM, Support vector machines.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Sample size required for proportion (<xref ref-type="bibr" rid="B4">Harnly, 2012</xref>).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Minimum</th>
<th align="center">Number</th>
<th align="center">Number</th>
<th align="center">1-Sided</th>
<th align="center">2-Sided</th>
<th align="center">2-Sided</th>
<th align="center">Effective</th>
</tr>
<tr>
<th align="center">Probability</th>
<th align="center">tests</th>
<th align="center">failures</th>
<th align="center">LCL</th>
<th align="center">LCL</th>
<th align="center">UCL</th>
<th align="center">AOQL</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">50%</td>
<td align="center">10</td>
<td align="center">2</td>
<td align="center">54.1%</td>
<td align="center">49.0%</td>
<td align="center">94.3%</td>
<td align="center">71.7%</td>
</tr>
<tr>
<td align="center">50%</td>
<td align="center">20</td>
<td align="center">6</td>
<td align="center">51.6%</td>
<td align="center">48.1%</td>
<td align="center">85.5%</td>
<td align="center">66.8%</td>
</tr>
<tr>
<td align="center">50%</td>
<td align="center">40</td>
<td align="center">14</td>
<td align="center">52.0%</td>
<td align="center">49.5%</td>
<td align="center">77.9%</td>
<td align="center">63.7%</td>
</tr>
<tr>
<td align="center">50%</td>
<td align="center">80</td>
<td align="center">32</td>
<td align="center">50.8%</td>
<td align="center">49.0%</td>
<td align="center">70.0%</td>
<td align="center">59.5%</td>
</tr>
<tr>
<td align="center">55%</td>
<td align="center">10</td>
<td align="center">1</td>
<td align="center">65.2%</td>
<td align="center">59.6%</td>
<td align="center">100%</td>
<td align="center">79.8%</td>
</tr>
<tr>
<td align="center">55%</td>
<td align="center">20</td>
<td align="center">5</td>
<td align="center">56.8%</td>
<td align="center">53.1%</td>
<td align="center">88.8%</td>
<td align="center">71.0%</td>
</tr>
<tr>
<td align="center">55%</td>
<td align="center">40</td>
<td align="center">12</td>
<td align="center">57.1%</td>
<td align="center">54.6%</td>
<td align="center">85.8%</td>
<td align="center">68.2%</td>
</tr>
<tr>
<td align="center">55%</td>
<td align="center">80</td>
<td align="center">28</td>
<td align="center">55.9%</td>
<td align="center">54.1%</td>
<td align="center">74.5%</td>
<td align="center">64.3%</td>
</tr>
<tr>
<td align="center">60%</td>
<td align="center">10</td>
<td align="center">1</td>
<td align="center">65.2%</td>
<td align="center">59.6%</td>
<td align="center">100%</td>
<td align="center">79.8%</td>
</tr>
<tr>
<td align="center">60%</td>
<td align="center">20</td>
<td align="center">4</td>
<td align="center">62.2%</td>
<td align="center">58.4%</td>
<td align="center">91.9%</td>
<td align="center">75.2%</td>
</tr>
<tr>
<td align="center">60%</td>
<td align="center">40</td>
<td align="center">10</td>
<td align="center">62.4%</td>
<td align="center">59.8%</td>
<td align="center">85.8%</td>
<td align="center">72.8%</td>
</tr>
<tr>
<td align="center">60%</td>
<td align="center">80</td>
<td align="center">24</td>
<td align="center">61.0%</td>
<td align="center">59.2%</td>
<td align="center">78.9%</td>
<td align="center">69.1%</td>
</tr>
<tr>
<td align="center">65%</td>
<td align="center">10</td>
<td align="center">1</td>
<td align="center">65.2%</td>
<td align="center">59.6%</td>
<td align="center">100%</td>
<td align="center">79.8%</td>
</tr>
<tr>
<td align="center">65%</td>
<td align="center">20</td>
<td align="center">3</td>
<td align="center">67.8%</td>
<td align="center">64.0%</td>
<td align="center">94.8%</td>
<td align="center">79.4%</td>
</tr>
<tr>
<td align="center">65%</td>
<td align="center">40</td>
<td align="center">9</td>
<td align="center">65.1%</td>
<td align="center">62.5%</td>
<td align="center">87.7%</td>
<td align="center">75.1%</td>
</tr>
<tr>
<td align="center">65%</td>
<td align="center">80</td>
<td align="center">21</td>
<td align="center">65.0%</td>
<td align="center">63.2%</td>
<td align="center">82.1%</td>
<td align="center">72.7%</td>
</tr>
<tr>
<td align="center">70%</td>
<td align="center">10</td>
<td align="center">0</td>
<td align="center">78.7%</td>
<td align="center">72.2%</td>
<td align="center">100%</td>
<td align="center">86.1%</td>
</tr>
<tr>
<td align="center">70%</td>
<td align="center">20</td>
<td align="center">2</td>
<td align="center">73.8%</td>
<td align="center">69.9%</td>
<td align="center">97.2%</td>
<td align="center">83.6%</td>
</tr>
<tr>
<td align="center">70%</td>
<td align="center">40</td>
<td align="center">7</td>
<td align="center">70.7%</td>
<td align="center">68.0%</td>
<td align="center">91.3%</td>
<td align="center">79.7%</td>
</tr>
<tr>
<td align="center">70%</td>
<td align="center">80</td>
<td align="center">17</td>
<td align="center">70.4%</td>
<td align="center">68.6%</td>
<td align="center">86.3%</td>
<td align="center">77.4%</td>
</tr>
<tr>
<td align="center">75%</td>
<td align="center">10</td>
<td align="center">0</td>
<td align="center">78.7%</td>
<td align="center">72.2%</td>
<td align="center">100%</td>
<td align="center">86.1%</td>
</tr>
<tr>
<td align="center">75%</td>
<td align="center">20</td>
<td align="center">1</td>
<td align="center">80.4%</td>
<td align="center">76.4%</td>
<td align="center">100%</td>
<td align="center">88.2%</td>
</tr>
<tr>
<td align="center">75%</td>
<td align="center">40</td>
<td align="center">5</td>
<td align="center">76.5%</td>
<td align="center">73.9%</td>
<td align="center">94.5%</td>
<td align="center">84.2%</td>
</tr>
<tr>
<td align="center">75%</td>
<td align="center">80</td>
<td align="center">13</td>
<td align="center">75.9%</td>
<td align="center">74.2%</td>
<td align="center">90.3%</td>
<td align="center">82.2%</td>
</tr>
<tr>
<td align="center">80%</td>
<td align="center">20</td>
<td align="center">1</td>
<td align="center">80.4%</td>
<td align="center">76.4%</td>
<td align="center">100%</td>
<td align="center">88.2%</td>
</tr>
<tr>
<td align="center">80%</td>
<td align="center">40</td>
<td align="center">3</td>
<td align="center">82.7%</td>
<td align="center">80.1%</td>
<td align="center">98.6%</td>
<td align="center">88.8%</td>
</tr>
<tr>
<td align="center">80%</td>
<td align="center">80</td>
<td align="center">10</td>
<td align="center">80.2%</td>
<td align="center">78.5%</td>
<td align="center">93.1%</td>
<td align="center">85.8%</td>
</tr>
<tr>
<td align="center">85%</td>
<td align="center">20</td>
<td align="center">0</td>
<td align="center">88.1%</td>
<td align="center">83.9%</td>
<td align="center">100%</td>
<td align="center">91.9%</td>
</tr>
<tr>
<td align="center">85%</td>
<td align="center">40</td>
<td align="center">2</td>
<td align="center">86.0%</td>
<td align="center">83.5%</td>
<td align="center">98.6%</td>
<td align="center">91.1%</td>
</tr>
<tr>
<td align="center">85%</td>
<td align="center">80</td>
<td align="center">6</td>
<td align="center">86.1%</td>
<td align="center">84.6%</td>
<td align="center">96.5%</td>
<td align="center">90.6%</td>
</tr>
<tr>
<td align="center">90%</td>
<td align="center">40</td>
<td align="center">0</td>
<td align="center">93.7%</td>
<td align="center">91.2%</td>
<td align="center">100%</td>
<td align="center">95.6%</td>
</tr>
<tr>
<td align="center">90%</td>
<td align="center">60</td>
<td align="center">2</td>
<td align="center">90.4%</td>
<td align="center">88.6%</td>
<td align="center">99.1%</td>
<td align="center">93.9%</td>
</tr>
<tr>
<td align="center">90%</td>
<td align="center">80</td>
<td align="center">3</td>
<td align="center">91.0%</td>
<td align="center">89.5%</td>
<td align="center">98.7%</td>
<td align="center">94.1%</td>
</tr>
<tr>
<td align="center">95%</td>
<td align="center">60</td>
<td align="center">0</td>
<td align="center">95.7%</td>
<td align="center">94.0%</td>
<td align="center">100%</td>
<td align="center">97.0%</td>
</tr>
<tr>
<td align="center">95%</td>
<td align="center">80</td>
<td align="center">0</td>
<td align="center">96.7%</td>
<td align="center">95.4%</td>
<td align="center">100%</td>
<td align="center">97.7%</td>
</tr>
<tr>
<td align="center">95%</td>
<td align="center">90</td>
<td align="center">1</td>
<td align="center">95.2%</td>
<td align="center">94.0%</td>
<td align="center">100%</td>
<td align="center">97.0%</td>
</tr>
<tr>
<td align="center">98%</td>
<td align="center">130</td>
<td align="center">0</td>
<td align="center">98.0%</td>
<td align="center">97.1%</td>
<td align="center">100%</td>
<td align="center">98.6%</td>
</tr>
<tr>
<td align="center">98%</td>
<td align="center">240</td>
<td align="center">1</td>
<td align="center">98.2%</td>
<td align="center">97.7%</td>
<td align="center">100%</td>
<td align="center">98.8%</td>
</tr>
<tr>
<td align="center">99%</td>
<td align="center">28/0</td>
<td align="center">0</td>
<td align="center">99.0%</td>
<td align="center">98.6%</td>
<td align="center">100%</td>
<td align="center">99.3%</td>
</tr>
<tr>
<td align="center">99%</td>
<td align="center">400</td>
<td align="center">1</td>
<td align="center">99.1%</td>
<td align="center">98.8%</td>
<td align="center">100%</td>
<td align="center">99.4%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Assume: 1. Binary outcome (occur/not occur).</p>
</fn>
<fn>
<p>2. Constant probability of event occurring</p>
</fn>
<fn>
<p>3. Independent trials</p>
</fn>
<fn>
<p>4. Fixed number of trials</p>
</fn>
<fn>
<p>Inference: 95% confidence interval lies entirely at or above the specified minimum.</p>
</fn>
<fn>
<p>Desired: Sample size N.</p>
</fn>
<fn>
<p>Notes: 1. Based on modified Wilson score 1-sided confidence limit.</p>
</fn>
<fn>
<p>2. AOQL &#x3d; Average Outgoing Quality Level</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>It is recognized that the POI method is philosophically and statistically defensible but not practical. There have been very few applications of the POI method to identification problems (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>). Major obstacles are the need to identify each potential adulterant and to test each with 60 analyses. A daunting task. AOAC International has recently charged an expert panel with the development of new guidelines.</p>
<p>A new method has recently been proposed based on one-class modeling using principal component analysis (PCA) (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>; <xref ref-type="bibr" rid="B5">Harnly, 2023</xref>). This approach simplifies the identification process. It builds a single model based on the reference samples and treats every other sample as a potential adulterant. If an unknown sample falls within the model, the sample is judged to be authentic. If it falls outside the model it is deemed to be adulterated. The method is compatible with any multivariate data set and the Q statistic (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>) of PCA provides an inherent combined metric that can be used to determine the confidence limit of the model (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>). The validity of the model can be established by the cross validation. Thus, one-class modeling simplifies identification by eliminating the need to identify every potential adulterant, provides an inherent combined metric for each sample, provides confidence limits based on the number and quality of the reference samples, and provides a binary authentic/non-authentic output based on a statistical analysis.</p>
<p>This review will consider how chemical identification can be based on one or more features (univariate or multivariate), one or more classes of samples (one-class or two-class modeling), and quantitative or qualitative data. Regardless the number of features, classes, or nature of the data, the method generates a combined metric that produces a binary result, &#x201c;yes&#x201d; or &#x201c;no,&#x201d; with regards to authenticity (<xref ref-type="bibr" rid="B4">Harnly, 2012</xref>). This review will focus on the use of one-class modeling of multivariate quantitative data using PCA (<xref ref-type="bibr" rid="B5">Harnly, 2023</xref>). One-class modeling offers an inherently different approach from the POI model and two-class modeling and requires only the identification of the reference samples. In the last 10 years, this approach has been used to authenticate numerous raw botanicals and botanical supplements (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>; <xref ref-type="bibr" rid="B7">Harnly et al., 2016</xref>; <xref ref-type="bibr" rid="B9">Harnly and Upton, 2024</xref>; <xref ref-type="bibr" rid="B8">Harnly et al., 2017</xref>; <xref ref-type="bibr" rid="B3">Geng et al., 2020</xref>). Several examples of the application of one-class modeling to identification problems will be presented.</p>
<sec id="s1-1">
<title>Univariate versus multivariate analyses</title>
<p>Identification methods, as stated, generally employ multivariate analysis. Qualitative morphological methods are inherently multivariate since all the features of a plant are involved in identification. Quantitative chemical methods offer non-targeted multivariate analyses based on many chromatographic and spectral methods such as gas and liquid chromatography (GC and LC) and infrared spectrometry (IR), near IR spectrometry (NIR), mass spectrometry, and nuclear magnetic resonance spectrometry (NMR) (<xref ref-type="bibr" rid="B16">Nichani et al., 2023a</xref>; <xref ref-type="bibr" rid="B15">Nichani et al., 2023b</xref>). These methods may offer hundreds, even thousands (depending on resolution), of variables to characterize a sample. The many variables improve the chances of including important components of chemical identity or observing components that are not present (adulterants or contaminants) in the reference samples. Multivariate methods are less susceptible to fraud. Ideally, a botanical can be characterized with respect to hundreds of components with known structural and concentration variations that will allow the user to discriminate between similar species, samples with intentionally altered compositions, or material substitutes. However, incorporating a large number of variables in a binary decision (authentic or not authentic) is a challenge. Reduction of multiple variables into a combined metric offers the advantage of using classic univariate statistics to make this binary decision.</p>
<p>There are two excellent examples that illustrate the use of univariate statistics for class analysis. First, the limit of detection (LOD) is a classic one-class model (<xref ref-type="fig" rid="F1">Figure 1A</xref>) aimed at determining if a sample signal is a member of the blank signal population (<xref ref-type="bibr" rid="B12">Long and Winefordner, 1983</xref>). In the univariate mode, multiple measurements of the method blank establish the &#x201c;baseline&#x201d; (blank mean) and the standard deviation (distribution). Assuming a normal distribution, it is possible to statistically determine whether a signal is not simply due to random variation of the blank signal. For this purpose, a threshold is usually set at 3 times the baseline standard deviation (3s). This provides 99% confidence that any signal observed above this level is due to the presence of the analyte. Inversely, this threshold confirms that the test signal is not a member of the blank population. This is a form of one-class modeling.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>One class modeling and two class classification, <bold>(A)</bold> Detection limit and <bold>(B)</bold> Students&#x2019; t-test. Applied to univariate tests, the intensity is the measure for a single variable. For multivariate tests, the intensity is the combined metric computed from all variables.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g001.tif"/>
</fig>
<p>Second, the students t test is a well established univariate approach to determining if the means of two populations are similar (<xref ref-type="fig" rid="F1">Figure 1B</xref>) (<xref ref-type="bibr" rid="B14">Moore and McCabe, 1999</xref>). Multiple measurements of the value of interest are used to establish the mean and standard deviation of each population. Assuming both populations are normally distributed, the difference between the means of the populations is evaluated in terms of their standard deviations to establish the statistical significance. Similarly, the F-test can be used for establishing the difference for three or more populations (<xref ref-type="bibr" rid="B14">Moore and McCabe, 1999</xref>). These are forms of two-class or multi-class modeling. This approach is used as the basis of the POI method. However, in that case, the standard deviation is not known for either population, and the reference and adulterated samples must each be run 30 times to establish their distribution.</p>
<p>The diagrams in <xref ref-type="fig" rid="F1">Figure 1</xref> are the same for univariate and multivariate analyses, if the multivariate data is represented by a combined metric (<xref ref-type="bibr" rid="B16">Nichani et al., 2023a</xref>; <xref ref-type="bibr" rid="B4">Harnly, 2012</xref>; <xref ref-type="bibr" rid="B11">LaBudde and Harnly, 2012</xref>; <xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>). That is, all the variables from a multivariate analysis are used to compute a single, combined value or metric for each sample. This combined metric is then processed in exactly the same manner as a univariate value. The baseline value in <xref ref-type="fig" rid="F1">Figure 1A</xref> corresponds to the combined metric mean of the reference samples. In this case, a test sample is judged to be authentic if it falls within the pre-determined level of uncertainty (e.g., within &#xb1;2s). In <xref ref-type="fig" rid="F1">Figure 1B</xref>, the means correspond to the combined metrics for the reference population and the test (potentially adulterated) population. Classic statistics can be used to compute the confidence with which the two populations can be distinguished. When analyzed, the combined metric for the test sample will be judged to belong to either the authentic or test population.</p>
<p>For both the LOD and t-test, the number of measurements, n, for reference and test sample populations is critical to establishing the confidence with which similarity can be determined (<xref ref-type="bibr" rid="B12">Long and Winefordner, 1983</xref>; <xref ref-type="bibr" rid="B14">Moore and McCabe, 1999</xref>). The calculated mean and standard deviation of a population are only estimates of the true mean and standard deviation. As n increases, the level of confidence in the two values increases as does the decision regarding similarity.</p>
</sec>
<sec id="s1-2">
<title>One-class modeling versus multi-class classification</title>
<p>There are numerous chemometric methods for processing multivariate data sets (<xref ref-type="table" rid="T2">Table 2</xref>). Unsupervised analysis requires no user input and reveals naturally occurring patterns. The most widely used unsupervised methods are PCA and hierarchical clustering analysis (HCA). Both serve to reveal sample patterns that may not be obvious. Supervised methods require identification of the classes of samples in the data set. Brereton has divided supervised multivariate analysis into the one-class classifiers and the two- (or multi-) class classifiers and described their fundamental differences (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). Both approaches require a priori identification of the sample classes making them supervised methods.</p>
<p>One-class classifiers or one-class modeling constructs a model for each class based only on the characteristics of that class (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). Comparison between sample classes requires comparison of the models for each class through soft independent modeling class analogy (SIMCA). SIMCA is considered a soft modeling technique as samples can be assigned to a single class, multiple classes, or no class. This is an excellent approach for recognizing outliers. Two-class classifiers or classification builds a single model for all the classes based on the features of all the classes (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). Two-class classification is considered a hard modeling technique since samples are forced into one of the pre-designated classes. For example, partial least squares-discriminant analysis (PLS-DA) of <italic>Gingko biloba</italic> and <italic>Echinacea purpurea</italic> will require a sample of <italic>Actaea racemosa</italic> to be classified as either <italic>Gingko</italic> or <italic>Echinacea</italic>. Two-class classification methods have difficulty dealing with outliers and require re-calculation when additional classes are added.</p>
<p>One-class modeling has historically been used for process control (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). Samples from key locations/times in a successful process are used to build models which can track the fidelity of succeeding processes. The one-class models tell the operator whether the process is within the limits established in previous runs.</p>
<p>There are numerous ways to compute a combined metric from multivariate analyses (<xref ref-type="bibr" rid="B16">Nichani et al., 2023a</xref>; <xref ref-type="bibr" rid="B15">Nichani et al., 2023b</xref>; <xref ref-type="bibr" rid="B4">Harnly, 2012</xref>; <xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). The two statistical measures for evaluating a multivariate PCA model are shown in <xref ref-type="fig" rid="F2">Figure 2</xref> where a single principal component (PC1) is fit to a set of bivariate data (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>; <xref ref-type="bibr" rid="B11">LaBudde and Harnly, 2012</xref>). In this case PC1 is the model, it is a vector which intercepts the greatest variance of the data set (black symbols). The Q statistic is the variance outside the model, the distance of the sample from the model determined by a line from the sample perpendicular to the model. The Hotelling T<sup>2</sup> statistic characterizes the sample variance within the model, the distance from the perpendicular intercept to the model center. This latter distance is analogous to the distance of a sample to the center of a normal distribution for univariate data. <xref ref-type="fig" rid="F2">Figure 2</xref> shows that the Hotelling T<sup>2</sup> statistic would place the test samples (blue symbols) in the same class as the reference samples while the Q statistic shows that they belong in different classes. In general, the Q statistic is much more sensitive to compositional differences and detection of outliers.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Illustration of the Hotelling T<sup>2</sup> statistic and Q statistic for the first principal component (PC1) fit to a bivariate data set (<inline-graphic xlink:href="fphar-16-1504230-fx1.tif"/>) with test samples (<inline-graphic xlink:href="fphar-16-1504230-fx2.tif"/>). The Q statistic characterizes the variance outside the model, the distance from the sample perpendicular to the model, PC1. The Hotelling T<sup>2</sup> statistic characterizes the variance within the model, the distance on PC1 from the sample intercept to the origin (mean centered data). The normal distribution of the Q statistic and the Hotelling T2 statistic are shown in red. A plot of the Q statistic versus the Hotelling T2 statistic is known as an influence plot.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g002.tif"/>
</fig>
<p>One-class modeling using PCA is ideally suited for identification. A model can be built using reference samples, the Q statistic provides an excellent combined metric, confidence limits can be established, and a test sample can be determined to lie within (authentic) or outside (adulterated) the model. One-class modeling is flexible since only the reference data is identified. In the case of two classes of samples, both can be modeled and the models compared or either class can be modeled and the other treated as unknowns (see the example for black cohosh below).</p>
</sec>
</sec>
<sec id="s2">
<title>Quantitative versus qualitative</title>
<p>The facility of computing a combined metric from quantitative data using one-class modeling is readily seen in the preceding section. Both the Hotelling T<sup>2</sup> statistic and the Q statistic are readily derived from chemometric analyses (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>) and provide a combined metric for all variables for each sample. Deriving such a metric for qualitative data is much more challenging (<xref ref-type="bibr" rid="B4">Harnly, 2012</xref>; <xref ref-type="bibr" rid="B17">Sudberg, 2024</xref>). The AOAC Guidelines state that they are intended for all candidate botanical identification methods and that identity specifications can be based on morphological, genetic, chemical, and/or other defining features of the botanical material. At that date it was envisioned that morphological or microscopic methods could be reduced to an algorithm or check list that would provide a suitable combined metric for judging authenticity. The companion paper (<xref ref-type="bibr" rid="B11">LaBudde and Harnly, 2012</xref>) illustrating the application of the POI method was based on quantitative analysis (mass spectral data) and, unfortunately, no example of a qualitative method was included.</p>
<p>There have not been any reports of qualitative identification in the literature. However, <xref ref-type="bibr" rid="B17">Sudberg (2024)</xref> at Alkemist Labs recently described a qualitative method for discriminating between spearmint (<italic>Menta spicata</italic> L.) and peppermint (<italic>Mentha</italic> x <italic>piperita</italic>) at the International Conference on Natural Products Research (Krakow, Poland, 2024) based on Bayes&#x2019; Theorem. This report nicely incorporated the POI principles which are discussed below.</p>
<sec id="s2-1">
<title>Probability of identification (POI)</title>
<p>The AOAC International POI method is based on two-class classification (<xref ref-type="bibr" rid="B1">Brereton, 2009</xref>). The glossary (<xref ref-type="table" rid="T1">Table 1</xref>) shows that the key feature of the method is to identify an authentic sample population that would always give a positive result (inclusivity sampling frame) and a population that would always give a negative result (exclusivity sampling frame). From their respective populations, a standard superior test material (SSTM, representing the minimum acceptable concentration of the target material) and standard inferior test material (SITM, representing the maximum unacceptable concentration of target material) are chosen to be run 30 times each to establish the precision of the method (<xref ref-type="table" rid="T3">Table 3</xref>).</p>
<p>In more general terms, an authentic (acceptable) and an adulterated (not acceptable) sample are chosen to characterize the resolution of the method, i.e., to determine if the method can discriminate between the two levels of adulteration. The adulterated sample is frequently an aliquot of the authentic sample spiked with a known level of adulterant or a different species from the same genus of the authentic sample (<xref ref-type="bibr" rid="B15">Nichani et al., 2023b</xref>; <xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>; <xref ref-type="bibr" rid="B5">Harnly, 2023</xref>; <xref ref-type="bibr" rid="B7">Harnly et al., 2016</xref>; <xref ref-type="bibr" rid="B9">Harnly and Upton, 2024</xref>; <xref ref-type="bibr" rid="B8">Harnly et al., 2017</xref>). Without specifying the method, the 30 repeats of each are used to establish the mean and distribution of the two populations (<xref ref-type="bibr" rid="B4">Harnly, 2012</xref>). The ability to perform 60 analyses with no false positives or negatives suggests a close to baseline separation of the two populations as shown in <xref ref-type="fig" rid="F1">Figure 1B</xref>. <xref ref-type="table" rid="T3">Table 3</xref> shows that 2 failures out of 60 translates to identification at the 90% confidence level 93.9% of the time, whereas no failures in 60 attempts corresponds to separation at the 95% confidence level 97.0% of the time.</p>
<p>The requirement of identifying an SITM for each potential adulterant and analyzing both the SSTM and SITM materials 30 times made the POI method unpopular. Today there is considerable interest in reducing the number of required analyses and expanding qualitative applications. However, it should be recognized that the number of samples analyzed, and the variance associated with the reference samples will always determine the confidence level that can be achieved by the method.</p>
</sec>
<sec id="s2-2">
<title>One-class modeling</title>
<p>One-class modeling has been described in detail in a previous paper (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>) and in the previous section on &#x201c;One-class modeling versus multi-class classification.&#x201d; It makes use of the original requirement of the POI method to acquire an inclusivity panel (<xref ref-type="table" rid="T2">Table 2</xref>) of samples, requires establishing a PCA model, and uses the Q statistic as a combined metric to determine whether any test sample lies within or beyond the specified confidence limit of the model. There is no exclusivity panel, and any combination of unknown samples can be tested against the model. The false positive and negative rate are determined by the confidence level specified by the user. Validation of the model is established using either cross validation or an external validation sets of samples.</p>
<p>Like all statistical evaluations, the degree of confidence is dependent on n, the number of reference samples. In general, the more reference samples used, the more confidence the user has in the average Q value and the standard deviation. However, with respect to botanical materials, the biggest source of variability is the biological variability of the samples. Significant variability has been observed between sources, growing location, growing year, harvests, and from plant-to-plant. The more varied the plant meta-data, the greater the population distribution. As will be shown for the example of <italic>A. racemosa</italic> below, the variance covering three classes of identical BRMs from different sources is greater than the variance for each class.</p>
<p>
<xref ref-type="fig" rid="F3">Figure 3</xref> presents an example of the application of one-class modeling to a collection of Maca (<italic>Lepidium meyenii</italic>) samples consisting of roots collected in Peru and China and Processed Maca supplements from Peru (<xref ref-type="bibr" rid="B3">Geng et al., 2020</xref>). Processed root supplements are heated, extruded, powdered, and sold commercially. The question to be answered was could the processed Maca be distinguished from the root samples. <xref ref-type="fig" rid="F3">Figure 3A</xref> shows an unsupervised PCA score plot for the Maca samples. The three clusters suggest the classes can be distinguished from each other but offer minimal statistics with respect to their differences. <xref ref-type="fig" rid="F3">Figure 3B</xref> shows the results for supervised one class modeling of the processed Peruvian Maca samples. The variable loadings for the processed Maca samples were used to compute scores for the other samples (Peruvian and Chinese roots). <xref ref-type="fig" rid="F3">Figure 3B</xref> shows ellipsoidal 95% confidence limits for each of the classes but it is again obvious that the score plot is a poor display for statistical differences of the classes.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Analysis of processed Maca (<italic>Lepidium meyenii</italic>): <bold>(A)</bold> unsupervised PCA scores plot for all materials, <bold>(B)</bold> PCA scores plot based on one-class modeling of processed Maca from Peru, <bold>(C)</bold> influence plot (Q residuals versus Hotelling T2 residuals) based on one-class modeling of processed Maca from Peru, and <bold>(D)</bold> Q statistic versus sample based on one-class modeling of processed Maca from Peru. Symbols: (<inline-graphic xlink:href="fphar-16-1504230-fx3.tif"/>) processed Maca from Peru, (<inline-graphic xlink:href="fphar-16-1504230-fx4.tif"/>) Maca root from Peru, and (<inline-graphic xlink:href="fphar-16-1504230-fx5.tif"/>) Maca root from China.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g003.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F3">Figure 3C</xref> presents a influence plot that displays the Q residuals (vertically) versus the Hotelling T<sup>2</sup> residuals (horizontally) for a one class model of the processed Maca data, i.e., a influence plot for the data in <xref ref-type="fig" rid="F3">Figure 3B</xref>. The data in <xref ref-type="fig" rid="F3">Figure 3C</xref> were acquired using only one principal component, minimizing the possibility of over fitting. The statistics for the model can be seen much more clearly in the influence plot as can the difference between the Q and Hotelling T<sup>2</sup> statistics. For the Hotelling T<sup>2</sup> statistic, the sensitivity for the processed Maca is 97% and the specificity for the Peruvian and Chinese roots are 64% and 0%, respectively. For the Q statistic, the sensitivity is 93% and the specificity for both roots is 100%. Thus, the Q statistic is much more sensitive to differences in the sample composition. Plotting only the Q value versus individual samples (<xref ref-type="fig" rid="F3">Figure 3D</xref>) presents a further simplified plot. This presentation can also be rotated to view the sample data as a frequency plot as will be shown below for one of the examples.</p>
<p>Choosing the processed Maca as the reference samples is arbitrary. Selection of the reference samples is dependent on the question being asked. The Peruvian or the Chinese root samples could also be chosen as the reference samples. The number of non-reference samples is also arbitrary. In <xref ref-type="fig" rid="F3">Figure 3</xref>, two sets of non-reference samples (Peruvian and Chinese Maca roots) were chosen. Since the one-class model is based solely on the reference samples, the average and distribution of the non-reference samples is not a concern.</p>
<p>In essence, one-class modeling examines the difference associated with every variable in the data set (8). Simplistically, a model based on a mass spectrum of 205 ions requires that all 205 variables for the test material lie within the statistical limits determined for each variable of the reference samples. Practically, some deviation of variable(s) can be tolerated as determined by the limits set by the user. If a test material is to be judged different or adulterated, one (or more) of the variables in the data set must be significantly different. As shown in a previous study (<xref ref-type="bibr" rid="B5">Harnly, 2023</xref>), variables that are autoscaled will have a variance of 1.0. The 205 MS variables will have a total variance of approximately 205. The signal necessary for a variable to have a significantly different variance and have a significant statistical impact on the combined metric can be predicted from the average intensity of the variable (<xref ref-type="bibr" rid="B5">Harnly, 2023</xref>).</p>
<p>An added factor of considerable importance is that the plots in <xref ref-type="fig" rid="F3">Figure 3</xref> can be constructed by anyone using any commercial chemometric platform. Starting with the same data set and using the same pre-processing steps, any commercial platform will produce the same plots. Identification using one-class modeling can be done in any lab without the need for a chemometrics expert.</p>
</sec>
<sec id="s2-3">
<title>Examples of one-class modeling</title>
<sec id="s2-3-1">
<title>American Ginseng (<italic>Panax quinquefolius</italic>)</title>
<p>American Ginseng (PQ) adulterated with varying levels of Asian Ginseng (<italic>Panax ginseng</italic>, or PG) was used to illustrate an application of the POI (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>). For the study, 44 PQ samples were obtained from the Wisconsin Ginseng Board (harvested over 3 years from 20 different farms in Wisconsin) and 8&#xa0;PG samples grown in China and acquired from American Herbal Pharmacopoeia and commercial sources. Samples were analyzed by flow injection mass spectrometry (FIMS) which yielded spectra with more than 1,000 ions.</p>
<p>The SSTM and SITM (<xref ref-type="table" rid="T1">Table 1</xref>) for the POI analysis were chosen as 98% and 90% PQ, respectively. Spectra for the test materials were synthesized mathematically using appropriate ratios of the 100% PQ and 100% PG spectra. In all, 344 (43 PQ &#xd7; 8&#xa0;PG) spectra were generated for each test material. The unsupervised PCA score plot (not shown) showed visual separation but the computed confidence limits did not allow detailed statistical analysis. Ironically, one-class modeling was used to provide detailed statistical analysis of the data although it was not the focus of the paper. The Q statistic served as the combined metric (<xref ref-type="fig" rid="F4">Figure 4</xref>).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Q statistics plot of American ginseng (PQ) adulterated with Asian ginseng (PG) based on one-class modeling of 100% PQ: <bold>(A)</bold> Q statistic plot for (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) 100% PQ, (<inline-graphic xlink:href="fphar-16-1504230-fx7.tif"/>) 100% PG, and (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) 98% PQ; <bold>(B)</bold> Q statistic plot for (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) 100% PQ, (<inline-graphic xlink:href="fphar-16-1504230-fx7.tif"/>) 100% PG, and (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) 90% PQ. The black dashed line is the 95% confidence limit for a one-class model of 100% PQ. The red dashed line is the 95% confidence limit for the one-class model of 98% PQ. Sensitivity for 100% and 98% PQ is 95% and 98%, respectively. Specificity for 100% and 90% PG is 100% and 99%, respectively.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g004.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F4">Figure 4A</xref> shows the Q statistic plot for 100% PQ, 100% PG, and 98% PQ with confidence limits based on one-class modeling for both 100% PQ and 98% PQ. The one-class model for 100% PQ has a sensitivity of 95% (42/44) and a specificity of 100%. The one-class model for 98% PQ, the SSTM, has a sensitivity of 98% (342/344) and a specificity of 100%. <xref ref-type="fig" rid="F4">Figure 4B</xref> shows a similar plot with 90% PQ, the SITM. Based on the one-class model for 98% PQ, the 90% PG has a specificity of 99% (341/344).</p>
<p>A follow up study verified the previous numerical dilution model with physical dilution of 100% PQ with 100% PG (<xref ref-type="bibr" rid="B6">Harnly et al., 2013</xref>). The more labor intensive nature of the physical dilutions restricted the number of samples. Five samples of 100% PQ (arbitrarily selected from the samples in the previous study) were diluted with two samples of PG at ratios of 95:5, 9:1, 8:2, 6:4, 4:6, and 2:8. <xref ref-type="fig" rid="F5">Figure 5A</xref> shows that unsupervised PCA provided a near linear progression from 100% PQ to 100% PG on the X-axis indicating that PQ concentration was the primary source of variance. <xref ref-type="fig" rid="F5">Figure 5B</xref> shows the influence plot (Q statistic versus Hotelling T<sup>2</sup>) based on a one-class model of 100% PQ. While the Hotelling T2 residuals failed to discriminate between the different PQ purities, the Q statistic provided excellent separation of the five PQ concentrations. A plot of the square root of the Q statistic showed a linear relationship with the PQ concentration (plot not shown).</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Adulteration of American ginseng (<italic>Panax quinquefolius</italic>) with Asian ginseng (<italic>P. ginseng</italic>): <bold>(A)</bold> Unsupervised PCA and <bold>(B)</bold> influence plot (Q residuals versus Hotelling T<sup>2</sup> residuals) based on one-class modeling of 100% <italic>P. quinquefolius</italic>. Symbols: (<inline-graphic xlink:href="fphar-16-1504230-fx9.tif"/>) 100%, (<inline-graphic xlink:href="fphar-16-1504230-fx10.tif"/>) 95%, (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) 90%, (<inline-graphic xlink:href="fphar-16-1504230-fx11.tif"/>) 80%, (<inline-graphic xlink:href="fphar-16-1504230-fx12.tif"/>) 60%, (<inline-graphic xlink:href="fphar-16-1504230-fx13.tif"/>) 40%. (<inline-graphic xlink:href="fphar-16-1504230-fx1.tif"/>) 20%, and (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) 0% PQ.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g005.tif"/>
</fig>
<p>An additional follow up study examined the ability of FIMS to discriminate between PQ:PG ratios of 99:1, 98:2, 97:3, and 96:4 (unpublished data) to determine the level of adulteration of PQ that could be detected. <xref ref-type="fig" rid="F6">Figure 6</xref> shows the influence plots (Q statistic versus Hotelling T<sup>2</sup>) for the 4 purity levels of PQ. Once again, Q statistic was more useful than Hotelling T<sup>2</sup> for discriminating between populations. The sensitivity for 100% PQ was 95% and the specificity was 57%, 90%, 99%, and 100% for PG contamination of 1%, 2%, 3%, and 4%, respectively.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Influence plots (Q residuals versus Hotelling T &#x5e; 2 residuals) for <bold>(A)</bold> PQ:PG &#x3d; 99:1, <bold>(B)</bold> PQ:PG &#x3d; 98:2, <bold>(C)</bold> PQ:PG &#x3d; 97:3, and <bold>(D)</bold> PQ:PG &#x3d; 96:4 based on one-class modeling of 100% PQ. Symbols: (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) 100% PQ and (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) 99%, 98%, 97%, and 96% PQ.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g006.tif"/>
</fig>
<p>This study demonstrates the utility of one-class modeling for detecting adulteration. The Q values for physical and mathematical dilution showed excellent agreement. The digital model, based on the pure spectra (100%) of the authentic material and the adulterant, provides a much simpler means of evaluating the ability of the method to detect adulteration.</p>
<p>The POI studies used the term SIMCA (soft independent modeling of class analogy) instead of the one-class modeling. SIMCA, as explained earlier, consists of a series of one-class models. In all the studies reported in this paper, only a single class was used for modeling, a single PCA model. For the ginseng studies and all those that follow, the 95% confidence limit was chosen as the statistical limit. Other statistical limits can be selected by researchers. The Q and Hotelling T<sup>2</sup> statistics were based on only one principal component, thus mitigating the possibility of over fitting. In addition, pre-processing for all the PCA calculations consisted of unit vector normalization of the samples (i.e., the sum of the squares of the sample intensities are set to 1.0) and autoscaling (normalization of each variable by its standard deviation) with mean centering of each variable.</p>
</sec>
<sec id="s2-3-2">
<title>Black cohosh (<italic>Actaea racemosa</italic>)</title>
<p>One-class modeling was used to distinguish <italic>A. racemosa</italic> L. (Ranunculaceae) from other <italic>Actaea</italic> species and commercially available roots and supplements (<xref ref-type="bibr" rid="B7">Harnly et al., 2016</xref>). FIMS and proton nuclear magnetic resonance spectrometry (NMR), two metabolic fingerprinting methods, and DNA sequencing were used to identify and authenticate the <italic>A. racemosa</italic> species. For this study, authentic <italic>A. racemosa</italic> botanical reference materials were acquired from four sources: American Herbal Pharmacopoeia (AHP), Strategic Sourcing (SS), North Carolina Arboretum (NCA), and the National Institutes of Standards and Technology (NIST). The NCA samples were triplicate samples collected from 22 sites across the eastern US.</p>
<p>Initial analyses of the <italic>Actaea</italic> species furnished by AHP using FIMS and NMR gave similar result as shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. PCA score plots (<xref ref-type="fig" rid="F7">Figures 7A, B</xref>) for both methods showed a pattern differentiating <italic>A. racemosa</italic> from the other <italic>Actaea</italic> species. This differentiation was statistically verified by one-class modeling based on <italic>A. racemosa</italic> (<xref ref-type="fig" rid="F7">Figures 7C, D</xref>). DNA sequencing using two independent gene regions (ITS and psbA-trnH) confirmed the metabolic fingerprinting results. Although not the point of this review, DNA sequencing provided identification for root materials but was inconsistent for the supplements. It was assumed that processing of the supplements destroyed the DNA in some cases.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Unsupervised PCA scores plots <bold>(A, B)</bold> and influence plots (Q residuals versus Hotelling T<sup>2</sup> residuals) <bold>(C, D)</bold> for one-class modeling of <italic>A. racemosa</italic> analyzed by FIMS <bold>(A, C)</bold> and NMR <bold>(B, D)</bold>. Symbols: (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) <italic>A. racemosa</italic>, (<inline-graphic xlink:href="fphar-16-1504230-fx12.tif"/>) <italic>A. rubra</italic>, (<inline-graphic xlink:href="fphar-16-1504230-fx13.tif"/>) <italic>A. cimicifuga</italic>, (<inline-graphic xlink:href="fphar-16-1504230-fx14.tif"/>) <italic>A. pachypoda, and</italic> (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) <italic>A. podocarpa</italic>.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g007.tif"/>
</fig>
<p>The simplicity of categorization in <xref ref-type="fig" rid="F7">Figure 7</xref> was complicated when <italic>A. racemosa</italic> reference samples from the four sources were compared. A PCA score plot (<xref ref-type="fig" rid="F8">Figure 8A</xref>) and a influence plot (combined one-class model plots for AHP and SS reference materials) (<xref ref-type="fig" rid="F8">Figure 8B</xref>) demonstrated that the reference materials were statistically different. The plots show three separate groups with the NIST samples in the same quadrant as those from SS. Sensitivity for the AHP and SS samples was 94% and 91%, respectively, and 91% for the un-modeled NCA samples in the upper right quadrant, exceeding the 95% confidence limits for AHP and SS, i.e., outliers compared to the other reference materials. Interestingly, the NCA samples collected from 22 sites across the US showed a similar lack of uniformity. These data suggested that sample handling and storage, local environmental conditions, and/or genetic drift influenced the plant metabolic profiles. It was also suggested that endophytic bacteria might have a significant influence on the metabolic profile.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Comparison of <italic>A. racemosa</italic> reference samples from 4 sources (AHP, SS, NCA, and NIST). <bold>(A)</bold> Unsupervised PCA scores plot and <bold>(B)</bold> influence plot based on one-class modeling of <italic>A. racemosa</italic> from AHP (vertically) and SS (horizontally). Symbols: (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) AHP, (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) SS, (<inline-graphic xlink:href="fphar-16-1504230-fx15.tif"/>) NIST, and (<inline-graphic xlink:href="fphar-16-1504230-fx16.tif"/>) NCA. AHP and SS provided clusters below and to the left of the 95% confidence limits, respectively. NIST samples lay barely below the confidence limit with the AHP samples and NCA samples were outliers (upper right quadrant) for both one-class models.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g008.tif"/>
</fig>
<p>A follow-up study took a closer look at the differences between the reference samples and the effect of pre-processing on PCA and one-class modeling patterns (<xref ref-type="bibr" rid="B9">Harnly and Upton, 2024</xref>). In <xref ref-type="fig" rid="F9">Figure 9</xref>, samples from all four sources were combined to serve as reference samples and used for one-class modeling. The variance of the combined reference materials exceeded that of each modeled individually, i.e., the model for AHP in <xref ref-type="fig" rid="F7">Figure 7C</xref>. Sensitivity was 97% for the combined <italic>A. racemosa</italic> samples and the specificity was 56% for other <italic>Actaea</italic> species, 55% for commercial <italic>Actaea</italic> roots, and 100% for commercial <italic>Actaea</italic> supplements. These data demonstrated the difficulty of differentiating between the <italic>Actaea</italic> species and that some of the commercial root samples are correctly identified as <italic>A. racemosa</italic>. None of the commercial <italic>Actaea</italic> supplements exhibited metabolic profiles similar to the reference samples, most likely due to the influence of sample preparation.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Q residuals plot based on one-class modeling of <italic>A</italic>. <italic>racemosa</italic> reference samples obtained from (<inline-graphic xlink:href="fphar-16-1504230-fx13.tif"/>) American Herbal Pharmacopoeia, (<inline-graphic xlink:href="fphar-16-1504230-fx16.tif"/>) Strategic Sourcing, (<inline-graphic xlink:href="fphar-16-1504230-fx12.tif"/>) North Carolina Arboretum, and (<inline-graphic xlink:href="fphar-16-1504230-fx17.tif"/>) the National Institutes of Health. Non-modeled samples are: (<inline-graphic xlink:href="fphar-16-1504230-fx6.tif"/>) other <italic>Actaea</italic> species from AHP, (<inline-graphic xlink:href="fphar-16-1504230-fx8.tif"/>) commercial <italic>Actaea</italic> roots, and (<inline-graphic xlink:href="fphar-16-1504230-fx11.tif"/>) commercial <italic>Actaea</italic> supplements. The horizontal dashed line provides the 95% confidence limit, p &#x3d; 0.05).</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g009.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F10">Figure 10</xref> shows an alternate perspective for one-class modeling data. If the modeled data in <xref ref-type="fig" rid="F9">Figure 9</xref> are viewed from the side, they can be displayed as frequency plots. This is a more intuitive view, providing the mean and distribution of the populations. <xref ref-type="fig" rid="F10">Figure 10</xref> shows only the frequency plots for <italic>A</italic>. <italic>racemosa</italic> and the other <italic>Actae</italic>a species. Alternate pre-processing methods had little impact on the separation of the two populations. Use of only variables with a low variance provided improved overlap of the <italic>A. racemosa</italic> from the different sources (data not shown) but reduced the ability to discriminate between species (data not shown). Use of variables that enhanced discrimination between species (based on an f-test) also failed to enhance differentiation as the differences between the reference samples from the 4 different sources were also increased, i.e., the frequency plot for <italic>A. racemosa</italic> broadened (data not shown).</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>A side view of <xref ref-type="fig" rid="F6">Figure 6</xref> showing the frequency plot for (&#x2014;) all 4 <italic>A. racemosa</italic> reference samples and (----) other <italic>Actaea</italic> species. The solid vertical line presents the mean for the <italic>A. racemosa</italic> reference samples and the dashed vertical line presents the &#x2b;2s limit.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g010.tif"/>
</fig>
<p>This study demonstrated that not all reference materials are created equal, use of multiple sources for reference materials increases the variance of the model, discriminating between species is difficult because the same metabolites are present but in a slightly different ratios, and finally, there is no simple pre-processing for improving discrimination between species.</p>
</sec>
<sec id="s2-3-3">
<title>Echinacea (<italic>Echinacea purpurea</italic>)</title>
<p>Preparation of commercial supplements from botanical ingredients results in changes in the chemical composition of the supplement. This makes it difficult to authenticate supplements based on the raw ingredients. One approach to comparing the composition of different samples is cross-correlation of the two spectra (<xref ref-type="bibr" rid="B8">Harnly et al., 2017</xref>). Ions found in both the reference samples and supplements spectra will provide an enhanced numerical product while those missing from either will yield a product close to zero. Comparison of the cross-correlated spectra can be achieved using PCA and eliminates the need for a normalization factor. As usual, however, one-class modeling provides a clearer analysis with statistical evaluation.</p>
<p>
<italic>E. purpurea</italic> (L.) Moench aerial samples were chosen as the test material and cross-correlation and one-class modeling were used to distinguish <italic>E. purpurea</italic> (L.) aerial samples from <italic>E. purpurea</italic> and <italic>Echinacea angustifolia</italic> roots and supplements. Authentic <italic>E. purpurea</italic> aerial and root samples and <italic>E. angustifolia</italic> root samples were obtained from Missouri Botanical Gardens, American Herbal Pharmacopoeia, and commercial sources. Twenty-three liquid and solid (tablets and capsules) supplements were purchased locally; 11 were single-ingredient supplements containing only <italic>E. purpurea</italic> and <italic>E. angustifolia</italic> aerial or root material and 13 were mixed or unknown <italic>Echinacea</italic> supplements labeled to contain mixtures of <italic>E. purpurea</italic> with <italic>E. angustifolia</italic> and/or <italic>Echinacea pallida</italic> (Nutt.) and consisting of either aerial or root material. Samples were analyzed by FIMS.</p>
<p>One-class modeling based on <italic>E. purpurea</italic> aerial samples showed that all other species and plant parts were significantly different from <italic>E. purpurea</italic> aerial (<xref ref-type="fig" rid="F11">Figure 11</xref>). Thus, the compositional profiles of <italic>E. purpurea</italic> aerial could be differentiated from that of the <italic>E. purpurea</italic> roots and <italic>E. angustifolia</italic> roots and from any of the single ingredient or mixed supplements. The sensitivity for the model in <xref ref-type="fig" rid="F11">Figure 11</xref> was 99% and the specificity was 98%. These data reaffirm previous observations that most raw materials are not suitable for authenticating supplements. The preparation process for commercial samples (e.g., extraction, back extraction, drying, powdering, and possible addition of other materials) can result in significant differences in the chemical composition and their levels. It should be noted that his does not necessarily reflect on the supplements efficacy.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>A Q statistics plot based on one-class modeling of <italic>Echinacea purpurea</italic> aerial samples (lower left). Labels for single ingredient samples: EPA &#x2013; <italic>E. purpurea</italic> aerial samples, EPAS &#x2013; EPA solid supplements, EPAL - EPA liquid supplements, EPR - E. purpurea root samples, EPRS - <italic>E. purpurea</italic> root solid supplements, EAR &#x2013; <italic>E. angustifolia</italic> root samples, EARS - <italic>E. angustifolia</italic> root solid supplements, and EARL - E. angustifolia root liquid supplements. Mixed supplements (far right) are not individually labeled.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g011.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F12">Figure 12</xref> shows the average EPA spectrum cross-correlated with itself (auto-correlation), each of the EPA solid supplements (EPAS) (EPAS1-EPAS5), and EPA liquid supplements (EPAL). Only EPAS5 appears to be significantly different. One-class modeling of the cross-correlated spectra of the average EPA with each of the individual EPA samples (plot not shown) showed that the EPAS and EPAL spectra, with the exception of EPAS5 were the same as the cross-correlated EPA population. These results provide some interesting conclusions. First, since the correlation spectra for EPAS1-EPAS4, and EPAL were similar, the supplements contained similar extracted components. This means that the starting ingredients and the processes used by the manufacturers were similar. Second, since the auto-correlation spectrum of EPA was similar to the correlation spectra of EPA with EPAS1-EPAS4 and EPAL, the reference samples in this study were similar to the raw ingredients used by the manufacturers. Finally, since the extracted compounds were the same, the analytical extraction process used in this study was similar to the preparation method employed by the manufacturers. These data also indicate that supplement EPAS5 was produced using either a different starting ingredient or a different extraction process.</p>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption>
<p>Spectra of <italic>E. purpurea aerial</italic> (EPA) cross-correlated with itself, <italic>E. purpurea</italic> aerial solid supplements 1&#x2013;5 (EPAS1-EPAS5), and <italic>E. purpurea</italic> aerial liquid supplement sample (EPAL1).</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g012.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F13">Figure 13</xref> shows a one-class model for mixed supplements based on the cross-correlated spectra for EPA, EPAS1-EPAS4, and EPAL. All mixed supplements that were known to contain aerial plant material fell within the 95% confidence level, even if the species were not specified. All supplements with root material, even if the species were not specified, fell outside the 95% confidence limit. The exceptions falling outside the 95% confidence limit were supplements composed of both species and both plant parts. This suggests that perhaps the aerial portion was of insufficient concentration to provide a clear cross-correlated spectrum.</p>
<fig id="F13" position="float">
<label>FIGURE 13</label>
<caption>
<p>Influence plot of mixed supplements spectra cross correlated with the average <italic>E. purpurea</italic> single ingredient spectra based on EPA, EPAS, and EPAL (<xref ref-type="fig" rid="F12">Figure 12</xref>). Labels: EPA &#x2013; <italic>E. purpurea</italic> aerial, EPAS1-EPAS4 &#x2013; EPA solid supplements, EPAL - EPA liquid supplement. Mixed labels have two factors. First factor abbreviations are NS - species no specified, P &#x2013; <italic>E. purpurea</italic>, A- <italic>E. angustifolia</italic>. Second factor abbreviations are NS &#x2013; plant part not specified, A &#x2013; aerial, and R &#x2013; root.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g013.tif"/>
</fig>
<p>The data in <xref ref-type="fig" rid="F12">Figures 12</xref>, <xref ref-type="fig" rid="F13">13</xref> indicate that cross correlation of supplements with the raw ingredient accurately identifies ions found in both spectra. Ions common to both spectra provide an enhanced multiplicative product whereas ions found in only one spectrum provide a product close to zero. Thus, cross-correlation and one-class modeling can identify commonalities of the raw ingredients and supplements.</p>
</sec>
<sec id="s2-3-4">
<title>Maca (<italic>Lepidium meyenii</italic> walpers)</title>
<p>Maca data were presented earlier for the general description of one-class modeling. This study was a collaboration with the American Botanical Council; Gaia Herbs; Hong Kong Baptist University; Charles Sturt University &#x26; Therapeutic Research, TTD International Pty Ltd; and NSF International Authenticity Laboratory (<xref ref-type="bibr" rid="B3">Geng et al., 2020</xref>). There is some question as to whether the plant material is accurately identified as <italic>L. meyenii</italic> Walpers or <italic>Lepidium Peruvianum</italic> Chacon (<xref ref-type="bibr" rid="B13">Meissner et al., 2015</xref>). Samples consisted of 39 commercial maca supplements from 11 manufacturers, 31 unprocessed maca roots grown in Peru and China, and an historic non-tuber maca sample from Peru. Samples were analyzed using FIMS and DNA next-generation sequencing (NGS).</p>
<p>Initial, untargeted PCA placed all the maca samples in 3 classes: commercial maca samples, roots grown in Peru, and roots grown in China (<xref ref-type="fig" rid="F3">Figure 3A</xref>). With one-class modeling the commercial Maca was readily differentiated from the raw root materials (<xref ref-type="fig" rid="F3">Figures 3C, D</xref>). A similar approach, selecting either the Chinese or Peruvian samples for the model, showed that either could be differentiated from the other two classes (data not shown).</p>
<p>One-class modeling was also combined with analysis of variance (ANOVA) to differentiate roots based on country (China and Peru) and color (black, red, and yellow) (<xref ref-type="fig" rid="F14">Figure 14</xref>). ANOVA was used to isolate the mean variance for each experimental factor (country, processing, and color) and cross factor to provide residuals free of the variance of the other factors (<xref ref-type="bibr" rid="B10">Harnly et al., 2014</xref>). One-class modeling of the individual factor residuals provided plots in <xref ref-type="fig" rid="F14">Figures 14B&#x2013;F</xref>. These data show, not surprisingly, that the metabolite composition correlates with country and color.</p>
<fig id="F14" position="float">
<label>FIGURE 14</label>
<caption>
<p>One-class modeling of FIMS fingerprints acquired by negative ionization after processing by multivariate ANOVA showing <bold>(A)</bold> the score plot for black, red, and yellow maca acquired from Peru and China surrounded by <bold>(B&#x2013;F)</bold> single-class models for each of the 5 country-color combinations. Sample icons: <inline-graphic xlink:href="fphar-16-1504230-fx19.tif"/> Peru black, <inline-graphic xlink:href="fphar-16-1504230-fx20.tif"/> Peru red, <inline-graphic xlink:href="fphar-16-1504230-fx18.tif"/> Peru yellow, <inline-graphic xlink:href="fphar-16-1504230-fx21.tif"/> China yellow, and <inline-graphic xlink:href="fphar-16-1504230-fx22.tif"/> China black.</p>
</caption>
<graphic xlink:href="fphar-16-1504230-g014.tif"/>
</fig>
<p>Metabolite profiling using ultra-high performance liquid chromatography-high resolution mass spectrometry (UHPLC-HRMS) in combination with PCA loadings was used to annotate the compounds responsible for differentiating between country and color. Genetically, all samples were confirmed to be similar and to be L. meyenii Walpers based on NGS at 3 gene regions (ITS2, psbA, and trnL) and comparison to recorded sequences of vouchered standards.</p>
<p>The results of this study show that one-class modeling can be used in combination with ANOVA to determine the statistical significance of differences arising from experimental factors associated with the samples. The metadata associated with genetics, country of origin, year of origin, climate, and handling can have a major impact on a plant&#x2019;s metabolite profile and one-class modeling provides a versatile in deconvoluting the interaction of the factors.</p>
</sec>
</sec>
</sec>
<sec sec-type="conclusion" id="s3">
<title>Conclusion</title>
<p>One-class modeling is a versatile method that can be readily applied to any set of multivariate data (targeted or non-targeted) data. This is a supervised method (reference samples must be identified) that develops a model based on the characteristics of a single class of samples, the reference samples. Using PCA, the Q statistic offers an inherent combined metric which can be used to determine if the reference and test samples belong to the same class, i.e., the test sample is authentic or adulterated. It can also be used as an effective tool for determining the similarity of sample metabolites for different species and, with cross-correlation, the similarity of raw ingredients and supplements. Finally, in combination with ANOVA, one-class modeling can be used to determine the significance of experimental factors. The beauty of one-class modeling is that it can be implemented on any commercial chemometrics platform and is applicable to any data file, i.e., chromatograms, spectra, or database, targeted or non-targeted.</p>
</sec>
</body>
<back>
<sec sec-type="author-contributions" id="s4">
<title>Author contributions</title>
<p>JH: Conceptualization, Data curation, Formal Analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing&#x2013;original draft, Writing&#x2013;review and editing.</p>
</sec>
<sec sec-type="funding-information" id="s5">
<title>Funding</title>
<p>This research was supported by the Agricultural Research Service of the U.S. Department of Agriculture and an Interagency Agreement (AOD1906-001-00004) from the Office of Dietary Supplements of the National Institutes of Health, Health and Human Services.</p>
</sec>
<sec sec-type="COI-statement" id="s6">
<title>Conflict of interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s7">
<title>Generative AI statement</title>
<p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brereton</surname>
<given-names>R.</given-names>
</name>
</person-group> <source>Chemometrics for pattern recognition</source>. <publisher-loc>West Sussex, United Kingdom</publisher-loc>: <publisher-name>John Wiley and Sons</publisher-name>. (<year>2009</year>). P.<fpage>177</fpage>&#x2013;<lpage>231</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="web">
<collab>Food and Drug Administration, Health and Human Services</collab> (<year>2024</year>). <article-title>Current good manufacturing practices (CGMPs) for food and dietary supplements</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.fda.gov/food/guidance-regulation-food-and-dietary-supplements/current-good-manufacturing-practices-cgmps-food-and-dietary-supplements">https://www.fda.gov/food/guidance-regulation-food-and-dietary-supplements/current-good-manufacturing-practices-cgmps-food-and-dietary-supplements</ext-link>.</comment>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Geng</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Brand</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Frame</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Meissner</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Characterization of maca (Lepidium meyenii/lepidium peruvianum) using a mass spectral fingerprinting, metabolomic analysis, and genetic sequencing approach</article-title>. <source>Planta Medica</source> <volume>86</volume>, <fpage>674</fpage>&#x2013;<lpage>685</lpage>. <pub-id pub-id-type="doi">10.1055/a-1161-0372</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>AOAC INTERNATIONAL guidelines for validation of botanical identification methods</article-title>. <source>J. AOAC Int.</source> <volume>95</volume>, <fpage>268</fpage>&#x2013;<lpage>272</lpage>. <pub-id pub-id-type="doi">10.5740/jaoacint.11-447</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Botanical authentication using one-class modeling</article-title>. <source>J. AOAC Int.</source> <volume>106</volume>, <fpage>1077</fpage>&#x2013;<lpage>1086</lpage>. <pub-id pub-id-type="doi">10.1093/jaoacint/qsad023</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Harrington</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Probability of identification: adulteration of American ginseng with asian ginseng</article-title>. <source>J. AOAC Int.</source> <volume>96</volume>, <fpage>1258</fpage>&#x2013;<lpage>1265</lpage>. <pub-id pub-id-type="doi">10.5740/jaoacint.13-290</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Colson</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yuk</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Comparison of flow injection MS, NMR, and DNA sequencing: methods for identification and authentication of black cohosh (Actaea racemosa)</article-title>. <source>Planta Medica</source> <volume>82</volume>, <fpage>250</fpage>&#x2013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1055/s-0035-1558113</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Botanical supplements: detecting the transition from ingredient to product</article-title>. <source>J. Food Comp. Anal.</source> <volume>64</volume>, <fpage>85</fpage>&#x2013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1016/j.jfca.2017.06.010</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Upton</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Variation in botanical reference materials: similarity of Actaea racemosa analyzed by flow injection mass spectrometry</article-title>. <source>J. AOAC Int.</source> <volume>107</volume>, <fpage>332</fpage>&#x2013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1093/jaoacint/qsad137</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harnly</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Harrington</surname>
<given-names>P. B.</given-names>
</name>
<name>
<surname>Botros</surname>
<given-names>L. L.</given-names>
</name>
<name>
<surname>Jablonski</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bergana</surname>
<given-names>M. M.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Characterization of near-infrared spectral variance in the authentication of skim and nonfat dry milk powder collection using ANOVA-PCA, pooled-ANOVA, and partial least-squares regression</article-title>. <source>J. Agric. Food Chem.</source> <volume>62</volume>, <fpage>8060</fpage>&#x2013;<lpage>8067</lpage>. <pub-id pub-id-type="doi">10.1021/jf5013727</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LaBudde</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Harnly</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Guidelines for validation of botanical identification methods</article-title>. <source>J. AOAC Int.</source> <volume>95</volume>, <fpage>268</fpage>&#x2013;<lpage>272</lpage>. <pub-id pub-id-type="doi">10.5740/jaoacint.11-266</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Long</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Winefordner</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1983</year>). <article-title>Limit of detection: a closer look at the IUPAC definition</article-title>. <source>Anal. Chem.</source> <volume>55</volume>, <fpage>712A</fpage>&#x2013;<lpage>724A</lpage>. <pub-id pub-id-type="doi">10.1021/ac00258a724</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meissner</surname>
<given-names>H. O.</given-names>
</name>
<name>
<surname>Mscisz</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kedzia</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Pisulewski</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Piatkowska</surname>
<given-names>E. P.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Peruvian maca: two scientific names Lepidium Meyenii Walpers and Lepidium Peruvianum Chacon &#x2013; are they phytochemically-synonymous?</article-title> <source>Int. J. Biomed. Sci.</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.59566/ijbs.2015.11001</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Moore</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>McCabe</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1999</year>). <source>Introduction to the practice of statistics</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>W. H. Freeman</publisher-name>, <fpage>456</fpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nichani</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Uhlig</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Colson</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Hettwer</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Simon</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>B&#xf6;nick</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2023b</year>). <article-title>Development of non-targeted mass spectrometry method for distinguishing spelt and wheat</article-title>. <source>Foods</source> <volume>12</volume>, <fpage>141</fpage>&#x2013;<lpage>157</lpage>. <pub-id pub-id-type="doi">10.3390/foods12010141</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nichani</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Uhlig</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Stoyke</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kemmlein</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ulberth</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Haase</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2023a</year>). <article-title>Essential terminology and considerations for validation of non-targeted methods</article-title>. <source>Food Chem.</source> <volume>17</volume>: <fpage>100538</fpage>&#x2013;<lpage>100612</lpage>. <pub-id pub-id-type="doi">10.1016/j.fochx.2022.100538</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sudberg</surname>
<given-names>S.</given-names>
</name>
</person-group> <article-title>Personal communication</article-title> (<year>2024</year>).</citation>
</ref>
</ref-list>
</back>
</article>