<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2022.863261</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Spectroscopy Approaches for Food Safety Applications: Improving Data Efficiency Using Active Learning and Semi-supervised Learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhang</surname> <given-names>Huanle</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1652173/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wisuthiphaet</surname> <given-names>Nicharee</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1269201/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Cui</surname> <given-names>Hemiao</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Nitin</surname> <given-names>Nitin</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1271984/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Xin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhao</surname> <given-names>Qing</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Computer Science, University of California, Davis</institution>, <addr-line>Davis, CA</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Food Science and Technology, University of California, Davis</institution>, <addr-line>Davis, CA</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Electrical and Computer Engineering, Cornell University</institution>, <addr-line>Ithaca, NY</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Xiangliang Zhang, University of Notre Dame, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Dan Li, National University of Singapore, Singapore; Denis Helic, Graz University of Technology, Austria</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Huanle Zhang <email>dtczhang&#x00040;ucdavis.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to AI in Food, Agriculture and Water, a section of the journal Frontiers in Artificial Intelligence</p></fn></author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>5</volume>
<elocation-id>863261</elocation-id>
<history>
<date date-type="received">
<day>27</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Zhang, Wisuthiphaet, Cui, Nitin, Liu and Zhao.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Zhang, Wisuthiphaet, Cui, Nitin, Liu and Zhao</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>The past decade witnessed rapid development in the measurement and monitoring technologies for food science. Among these technologies, spectroscopy has been widely used for the analysis of food quality, safety, and nutritional properties. Due to the complexity of food systems and the lack of comprehensive predictive models, rapid and simple measurements to predict complex properties in food systems are largely missing. Machine Learning (ML) has shown great potential to improve the classification and prediction of these properties. However, the barriers to collecting large datasets for ML applications still persists. In this paper, we explore different approaches of data annotation and model training to improve data efficiency for ML applications. Specifically, we leverage Active Learning (AL) and Semi-Supervised Learning (SSL) and investigate four approaches: baseline passive learning, AL, SSL, and a hybrid of AL and SSL. To evaluate these approaches, we collect two spectroscopy datasets: predicting plasma dosage and detecting foodborne pathogen. Our experimental results show that, compared to the <italic>de facto</italic> passive learning approach, advanced approaches (AL, SSL, and the hybrid) can greatly reduce the number of labeled samples, with some cases decreasing the number of labeled samples by more than half.</p></abstract>
<kwd-group>
<kwd>food science</kwd>
<kwd>spectroscopy analysis</kwd>
<kwd>machine learning</kwd>
<kwd>data efficiency</kwd>
<kwd>active learning</kwd>
<kwd>semi-supervised learning</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="2"/>
<equation-count count="6"/>
<ref-count count="51"/>
<page-count count="13"/>
<word-count count="8176"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Rapid measurement and monitoring technologies are being developed for diverse applications in food science. The goal of these technologies is to develop predictive relationships that can be used to better monitor and enhance the quality, safety and nutritional properties of food. Among these measurement approaches, spectroscopic analysis has been widely used to analyze food properties. To fully realize the great potential of these technologies, several key barriers need to be overcome before their transfer to industrial applications. The discovery of predictive relationships between the measurements and properties of food systems is one of the key limitations. This limitation results from the complexity of food systems and the lack of comprehensive predictive models that can use rapid and simple measurements to predict complex properties in food systems. ML has shown significant potential to improve the classification and prediction of these properties. However, the barriers to collecting large datasets for ML applications persists.</p>
<p>With the advances in computational capabilities and big data technologies, ML has been applied to a variety of agriculture and food-related fields (Liakos et al., <xref ref-type="bibr" rid="B24">2018</xref>; Tsakanikas et al., <xref ref-type="bibr" rid="B43">2020</xref>; Khullar and Singh, <xref ref-type="bibr" rid="B18">2021</xref>). A prevailing approach is to train an ML model using labeled samples. While unlabeled samples can often be collected at relatively low cost, annotating each sample to create its label is expensive and time-consuming, as it often involves human inspection and in-field experiments. For example, to predict the vineyard yield, robot carrying cameras can be used to collect a large number of unlabeled images, but practitioners still need to manually process each image in order to assign appropriate labels to the images (Ballesteros et al., <xref ref-type="bibr" rid="B4">2020</xref>). The prevailing practice is to randomly select samples and label them <italic>via</italic> costly in-field experiments and human annotation processes, and the ML model is trained only using the labeled samples. This approach results in poor data efficiency (Chapelle et al., <xref ref-type="bibr" rid="B5">2006</xref>; Settles, <xref ref-type="bibr" rid="B39">2012</xref>).</p>
<sec>
<title>1.1. AL and SSL for Data-Efficient Model Training</title>
<p>To improve data efficiency, we explore two advanced model training techniques: AL (Settles, <xref ref-type="bibr" rid="B39">2012</xref>) and SSL (Chapelle et al., <xref ref-type="bibr" rid="B5">2006</xref>). Instead of randomly selecting unlabeled samples for annotation, AL selects samples for annotation based on how informative these samples are to the currently trained ML model. In comparison, SSL exploits unlabeled samples by assigning pseudo-labels to them and trains the ML model using both labeled and pseudo-labeled samples.</p>
<p>In this paper, we study four approaches of data annotation and model training as illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref>. (A) In the baseline passive learning approach, samples are randomly selected from the pool of unlabeled samples for annotation, and the ML model is trained only on the labeled samples. (B) In the AL approach (details in Section 2.1), the selection of unlabeled samples is dependent on the currently trained ML model. Specifically, AL selects the unlabeled samples most useful for training the ML model. Various sampling strategies are explored, which quantify the usefulness of unlabeled samples based on different criteria. The ML model is trained only using the labeled samples. (C) In the SSL approach (details in Section 2.2), unlabeled samples are randomly selected for annotation. Instead of training the ML model using only the labeled samples, SSL assigns pseudo-labels to the unlabeled samples. The ML model is trained using both the labeled and pseudo-labeled samples. Therefore, the number of training samples in SSL equals the total number of samples. The methods of assigning pseudo-labels can be either related to the currently trained ML model (e.g., self-training) or rely only on the relationship among samples (e.g., label spreading). (D) In the hybrid approach that integrates AL and SSL, AL selects an unlabeled sample for annotation, and SSL assigns pseudo-labels to the remaining unlabeled samples in each iteration. The ML model is trained using both labeled and pseudo-labeled samples. In this hybrid approach, AL and SSL interact with each other in the following manner. The sampling strategy in AL is dependent on the ML model, which is trained using labeled samples from human annotation and pseudo-labeled samples from SSL. Meanwhile, labeled samples from AL affect SSL regarding which samples are required for pseudo-annotation.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Four approaches of data annotation and model training. <bold>(A)</bold> Passive learning. <bold>(B)</bold> Active learning. <bold>(C)</bold> Semi-supervised learning. <bold>(D)</bold> Hybrid of active learning and semi-supervised learning.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-863261-g0001.tif"/>
</fig>
<p>To evaluate different approaches for spectroscopic analysis in the food science field, we collect two datasets: plasma dosage classification and foodborne pathogen detection. Atmospheric plasma technologies are being developed as a non-thermal processing technology to improve food safety, reduce the impact on food quality, and improve the sustainability of food processing operations. One of the key challenges in plasma technology applications is the rapid assessment of its efficacy in the sanitation of food contact surfaces and food products. With this motivation, we are developing infra-red spectroscopy methods to aid in assessing the dosimetry of plasma treatment. Similarly, rapid detection of foodborne pathogens is a critical unmet need in food systems. In this direction, we are developing fluorescence spectroscopic methods based on molecular interactions between bacteriophages and their target bacteria to enable specific detection of bacteria.</p>
<p>We apply different ML models for the multi-class classification tasks. Representative methods are considered for AL and SSL. In the experiments, we adopt five-fold cross-validation for both datasets. Our experimental results show that compared to the baseline passive learning approach, the numbers of labeled samples for the plasma dosage classification dataset and the pathogen detection dataset are reduced by more than 50% using the hybrid approach. The promising results demonstrate that AL and SSL based approaches of data annotation and model training effectively improve data efficiency for spectroscopy-based ML applications.</p></sec>
<sec>
<title>1.2. Related Work of Applying AL and SSL in Food Systems</title>
<p>Both AL and SSL have been successfully applied to many domains, such as drug discovery (Naik et al., <xref ref-type="bibr" rid="B34">2013</xref>; Liu et al., <xref ref-type="bibr" rid="B27">2020</xref>), material science (Lookman et al., <xref ref-type="bibr" rid="B30">2019</xref>; Ma et al., <xref ref-type="bibr" rid="B31">2019</xref>), and systems biology (Tamposis et al., <xref ref-type="bibr" rid="B41">2019</xref>; Wang et al., <xref ref-type="bibr" rid="B47">2020</xref>). Research in food systems has adopted AL (a.k.a. optimal experimental design) to optimize non-ML model parameters, with the goal of reducing the number of experiments. It shows promising results of applying AL to applications such as determining partition coefficient in freeze concentration (Munson-McGee, <xref ref-type="bibr" rid="B32">2014a</xref>), tuning micro-extraction in beer (Leca et al., <xref ref-type="bibr" rid="B21">2015</xref>), identifying a rice drying model (Goujot et al., <xref ref-type="bibr" rid="B11">2012</xref>), and developing uniaxial expression of extracting oil from seeds (Munson-McGee, <xref ref-type="bibr" rid="B33">2014b</xref>). This paper differs from existing AL works in food science as we target general ML models while previous works are designed for specific and analytical models of few parameters. Similarly, SSL have successfully been used in food science problems such as determining tomato maturity (Jiang et al., <xref ref-type="bibr" rid="B16">2021</xref>) and tomato-juice freshness (Hong et al., <xref ref-type="bibr" rid="B14">2015</xref>). To the best of our knowledge, this paper is the first work that extensively evaluates different AL and SSL approaches for the ML models with applications in food systems.</p></sec></sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<p>In this section, we first provide preliminaries for AL in Section 2.1 and SSL in Section 2.2. Then, we explain our two datasets in detail in Section 2.3. Last, our experimental setup is presented in Section 2.4. <xref ref-type="table" rid="T1">Table 1</xref> tabulates the notations that are used in this paper for quick reference.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Notations used in this paper.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Notation</bold></th>
<th valign="top" align="left"><bold>Meaning</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M1"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">The set of unlabeled samples</td>
</tr>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M2"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">The set of labeled samples</td>
</tr>
<tr>
<td valign="top" align="left">&#x003B8;</td>
<td valign="top" align="left">The currently trained ML model</td>
</tr>
<tr>
<td valign="top" align="left">&#x003B8;<sup>&#x0002B;</sup></td>
<td valign="top" align="left">An updated ML model by adding a new training sample</td>
</tr>
<tr>
<td valign="top" align="left"><italic>C</italic></td>
<td valign="top" align="left">The number of classes for classification</td>
</tr>
<tr>
<td valign="top" align="left"><italic>x</italic></td>
<td valign="top" align="left">A sample</td>
</tr>
<tr>
<td valign="top" align="left"><italic>y</italic></td>
<td valign="top" align="left">The label of the sample <italic>x</italic></td>
</tr>
<tr>
<td valign="top" align="left"><italic>x</italic><sup>&#x0002A;</sup></td>
<td valign="top" align="left">The sample that has the maximum utility measure</td>
</tr>
<tr>
<td valign="top" align="left"><italic>y</italic><sup>&#x0002A;</sup></td>
<td valign="top" align="left">The label of the sample <italic>x</italic><sup>&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left"><italic>P</italic><sub>&#x003B8;</sub>(<italic>y</italic>|<italic>x</italic>)</td>
<td valign="top" align="left">The probability that the sample <italic>x</italic> belongs to the label <italic>y</italic> under the model &#x003B8;</td>
</tr>
<tr>
<td valign="top" align="left">&#x00177;</td>
<td valign="top" align="left">The most likely label for the sample <italic>x</italic>, i.e., &#x00177; &#x0003D; arg max<sub><italic>y</italic></sub> <italic>P</italic><sub>&#x003B8;</sub>(<italic>y</italic>|<italic>x</italic>)</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec>
<title>2.1. Active Learning: A Primer</title>
<p>AL, also known as query learning/optimal experimental design, is a sub-field of machine learning, which studies ML models that improve with experience and training (Settles, <xref ref-type="bibr" rid="B39">2012</xref>). Compared with passive learning, AL considers the &#x0201C;usefulness&#x0201D; of unlabeled samples for the current ML model. It strategically selects unlabeled samples for annotation to improve data efficiency for ML model training.</p>
<p><xref ref-type="table" rid="T3">Algorithm 1</xref> shows the workflow of the AL-based data annotation. <inline-formula><mml:math id="M3"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M4"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:math></inline-formula> is the set of unlabeled samples and labeled samples, respectively (line 1&#x02013;2). The ML model, denoted by &#x003B8;, is trained on the labeled sample set <inline-formula><mml:math id="M5"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:math></inline-formula> (line 4). The following unlabeled sample, denoted by <italic>x</italic><sup>&#x0002A;</sup>, is selected that has the highest utility measure according to the currently trained model &#x003B8; (line 5). Then, the experiment is conducted for <italic>x</italic><sup>&#x0002A;</sup> to obtain its ground-truth label, denoted by <italic>y</italic><sup>&#x0002A;</sup> (line 6). Since the label for <italic>x</italic><sup>&#x0002A;</sup> is obtained, <italic>x</italic><sup>&#x0002A;</sup> is removed from the pool of unlabeled samples <inline-formula><mml:math id="M6"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:math></inline-formula>, and the sample <italic>x</italic><sup>&#x0002A;</sup> along with its label <italic>y</italic><sup>&#x0002A;</sup> is added to the set of labeled samples <inline-formula><mml:math id="M7"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:math></inline-formula> (line 7). The process repeats until the trained model &#x003B8; has a satisfactory accuracy or does not improve with more labeled samples.</p>
<table-wrap position="float" id="T3">
<label>Algorithm 1</label>
<caption><p>Selection of unlabeled samples for annotation in active learning.</p></caption>
<graphic xlink:href="frai-05-863261-i0001.tif"/>
</table-wrap>
<p>The utility measure is essential for AL algorithms. There are various utility measures to estimate the usefulness of an unlabeled sample to the ML model. They mostly leverage the probability of the current ML model classifying an unlabeled sample <italic>x</italic> to class <italic>y</italic>, i.e., <italic>P</italic><sub>&#x003B8;</sub>(<italic>y</italic>|<italic>x</italic>). We introduce two categories of utility measure methods of AL: uncertainty-based sampling (Sharma and Bilgic, <xref ref-type="bibr" rid="B40">2017</xref>) and minimizing expected errors (Long et al., <xref ref-type="bibr" rid="B29">2015</xref>).</p>
<sec>
<title>2.1.1. Uncertainty Sampling Based Active Learning</title>
<p>The premise of uncertainty sampling is that we can avoid annotating samples that the ML model is confident about and focus instead on the unlabeled samples that confuse the ML model. Least confident and entropy are the most-used metrics for measuring the uncertainty of unlabeled samples.</p>
<list list-type="bullet">
<list-item><p><italic>Least Confident</italic>. For an unlabeled sample <italic>x</italic>, we can apply the currently trained ML model on it, which outputs the probability of the sample belonging to class <italic>y</italic>, i.e., <italic>P</italic><sub>&#x003B8;</sub>(<italic>y</italic>|<italic>x</italic>), where <italic>y</italic> &#x0003D; 1, 2, &#x02026;, <italic>C</italic> and <italic>C</italic> is the total number of classes for classification. Let&#x00027;s denote by &#x00177;, the class that is most likely for the unlabeled sample <italic>x</italic>, i.e., &#x00177; &#x0003D; arg max<sub><italic>y</italic></sub> <italic>P</italic><sub>&#x003B8;</sub>(<italic>y</italic>|<italic>x</italic>). The confidence of the current model for the sample <italic>x</italic> is thus <italic>P</italic><sub>&#x003B8;</sub>(&#x00177;|<italic>x</italic>). The unlabeled sample with the least confidence is selected for annotation, that is,
<disp-formula id="E1"><label>(1)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mtext>arg&#x000A0;max&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x00177;</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></list-item>
<list-item><p><italic>Entropy</italic>. Entropy is an often-used metric for quantifying the uncertainty of a distribution. For an unlabeled sample <italic>x</italic>, its entropy is calculated as <inline-formula><mml:math id="M17"><mml:mo>-</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, which is over the distribution of classes <italic>y</italic> for <italic>x</italic>. The entropy-based uncertainty sampling method selects the sample with the maximum entropy, i.e.,
<disp-formula id="E2"><label>(2)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mtext>arg&#x000A0;max</mml:mtext></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p> </list-item>
</list></sec>
<sec>
<title>2.1.2. Minimizing Expected Error Based Active Learning</title>
<p>This category of AL algorithms aims to select samples to directly increase the model accuracy without relying on the assumption between the sampling strategy and the model performance (e.g., the ML model can avoid annotating confident samples as in the uncertainty sampling). Since the ground-truth label of an unlabeled sample <italic>x</italic> is not available without experiment and annotation, the model accuracy by adding the sample <italic>x</italic> and its label to the training set is unknown. Nonetheless, we can estimate the expected model accuracy when the sample <italic>x</italic> is added for training, thus selecting a sample that minimizes the expected error of the current ML model.</p>
<p>Assume that the unlabeled sample <italic>x</italic> belongs to the class <italic>y</italic>. Let &#x003B8;<sup>&#x0002B;</sup> denote the updated ML model by adding the sample <italic>x</italic> and its imaginary label <italic>y</italic> to the training set. Then, the expected prediction error of the model &#x003B8;<sup>&#x0002B;</sup> can be estimated by applying &#x003B8;<sup>&#x0002B;</sup> to all unlabeled samples in <inline-formula><mml:math id="M19"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:math></inline-formula>, i.e., <inline-formula><mml:math id="M20"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x00177;</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M21"><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x00177;</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is the prediction error for the unlabeled sample <italic>x</italic>&#x02032;. Since the probability that <italic>x</italic> belongs to class <italic>y</italic> is <italic>P</italic><sub>&#x003B8;</sub>(<italic>y</italic>|<italic>x</italic>), the expected prediction error of selecting the unlabeled sample <italic>x</italic> for annotation is <inline-formula><mml:math id="M22"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x00177;</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. Correspondingly, Equation (3) formulates the sampling strategy that minimizes the expected prediction error.
<disp-formula id="E3"><label>(3)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mtext>arg&#x000A0;min</mml:mtext></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x00177;</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
An alternative to minimizing the expected prediction error is to minimize the expected log-loss error. Log-loss error is the <italic>de facto</italic> loss function for training classification models. Therefore, minimizing the expected log-loss error has a strong connection with ML model training. By replacing the prediction error <inline-formula><mml:math id="M24"><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x00177;</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> in Equation (3) with the log-loss error <inline-formula><mml:math id="M25"><mml:mo>-</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, Equation (4) formulates the sampling strategy of minimizing the expected log-loss error.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M26"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>L</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mtext>arg&#x000A0;min</mml:mtext></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec></sec>
<sec>
<title>2.2. Semi-supervised Learning: A Primer</title>
<p>SSL can be applied to exploit unlabeled samples. Typically, an SSL algorithm converts unlabeled samples to pseudo-labeled samples and then fine-trains the ML model using both labeled and pseudo-labeled samples. SSL can be categorized into inductive methods and transductive methods (van Engelen and Hoos, <xref ref-type="bibr" rid="B45">2020</xref>). We introduce self-training (Triguero et al., <xref ref-type="bibr" rid="B42">2015</xref>) and label spreading (Liu et al., <xref ref-type="bibr" rid="B28">2012</xref>), which are representative methods from these two categories.</p>
<sec>
<title>2.2.1. Self-Training Based Semi-supervised Learning</title>
<p>In the self-training based method, the currently trained ML model is applied to pseudo-annotate the unlabeled samples (Triguero et al., <xref ref-type="bibr" rid="B42">2015</xref>). Self-training methods assume that the prediction of the current model tends to be correct, and thus the model can be further fine-trained by leveraging more training samples.</p>
<p><xref ref-type="table" rid="T4">Algorithm 2</xref> shows the workflow of self-training based SSL. It differs from the AL workflow (<xref ref-type="table" rid="T3">Algorithm 1</xref>) in two main parts. (1) The unlabeled sample <italic>x</italic> is pseudo-labeled by the model itself, i.e., &#x003B8;(<italic>x</italic><sup>&#x0002A;</sup>), in self-training, whereas it is human-annotated, i.e., the ground-truth label <italic>y</italic><sup>&#x0002A;</sup>, in AL (line 6 in <xref ref-type="table" rid="T4">Algorithm 2</xref> vs. line 6 in <xref ref-type="table" rid="T3">Algorithm 1</xref>). (2) The most confident unlabeled sample is selected in self-training. In contrast, the most uncertain unlabeled sample is chosen in AL (line 5 in <xref ref-type="table" rid="T4">Algorithm 2</xref> vs. line 5 in <xref ref-type="table" rid="T3">Algorithm 1</xref>). Self-training selects the most confident unlabeled sample to mitigate the error propagation of pseudo-annotation since the model &#x003B8; is consecutively trained on the pseudo-labeled samples. In contrast, AL selects the model&#x00027;s most confusing unlabeled sample for human annotation so that the model can correctly classify misleading samples. The process of unlabeled sample selection, pseudo-annotation, and model training repeats until all the unlabeled samples are pseudo-annotated (line 3). At that time, the model has been trained on all labeled and pseudo-labeled samples.</p>
<table-wrap position="float" id="T4">
<label>Algorithm 2</label>
<caption><p>Self-training based semi-supervised learning.</p></caption>
<graphic xlink:href="frai-05-863261-i0002.tif"/>
</table-wrap>
</sec>
<sec>
<title>2.2.2. Label Spreading Based Semi-supervised Learning</title>
<p>Label spreading based SSL typically adopts a graph structure, where the vertices are the samples (both labeled and unlabeled) and edges exist between neighbor vertices (Liu et al., <xref ref-type="bibr" rid="B28">2012</xref>). The edge weight between two vertices represents the similarity of the corresponding two samples. The goal of label spreading is to assign pseudo-labels to the vertices of unlabeled samples. The labels of <italic>x</italic><sub><italic>i</italic></sub> and <italic>x</italic><sub><italic>j</italic></sub> are likely to be the same if the edge weight <italic>w</italic><sub><italic>ij</italic></sub> between them is large. There are two mainstream methods to define the graph structure and the edge weights:
<list list-type="bullet">
<list-item><p><italic>Fully Connected Graph</italic> (de Sousa et al., <xref ref-type="bibr" rid="B8">2013</xref>). In this graph, all vertices are connected with other vertices, and the edge weight <italic>w</italic><sub><italic>ij</italic></sub> between <italic>x</italic><sub><italic>i</italic></sub> and <italic>x</italic><sub><italic>j</italic></sub> is calculated as
<disp-formula id="E6"><label>(5)</label><mml:math id="M37"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where ||<italic>x</italic><sub><italic>i</italic></sub> &#x02212; <italic>x</italic><sub><italic>j</italic></sub>|| is the Euclidean distance between sample <italic>x</italic><sub><italic>i</italic></sub> and sample <italic>x</italic><sub><italic>j</italic></sub>, and &#x003C3; controls the decreasing rate of the weight over the distance. The edge weight <italic>w</italic><sub><italic>ij</italic></sub> is 1 when <italic>x</italic><sub><italic>i</italic></sub> &#x0003D; <italic>x</italic><sub><italic>j</italic></sub>, and approximately 0 when <italic>x</italic><sub><italic>i</italic></sub> is far from <italic>x</italic><sub><italic>j</italic></sub>. This weight representation is also called Radial Base Function (RBF) kernel or Gaussian kernel.</p>
</list-item>
<list-item><p><italic>k-Nearest-Neighbor (kNN) Graph</italic> (de Sousa et al., <xref ref-type="bibr" rid="B8">2013</xref>). In the kNN graph, each vertex is connected to its <italic>k</italic> nearest vertices. If sample <italic>x</italic><sub><italic>i</italic></sub> and sample <italic>x</italic><sub><italic>j</italic></sub> are connected, the edge weight <italic>w</italic><sub><italic>ij</italic></sub> is 1; otherwise, <italic>w</italic><sub><italic>ij</italic></sub> is 0. The kNN graph automatically adapts to the density of samples: in dense regions, the radius of the kNN neighborhood is small, while in sparse regions, the radius is large.</p></list-item>
</list></p>
<p>The probability of the label for each unlabeled sample is obtained once the graph converges (Zhou et al., <xref ref-type="bibr" rid="B51">2003</xref>). Then, each unlabeled sample is assigned to the label of the highest probability at once. Afterward, the ML model is re-trained using all samples (labeled and pseudo-labeled).</p></sec></sec>
<sec>
<title>2.3. Dataset</title>
<p>We collect two spectroscopy datasets in food science. One dataset predicts the plasma dosage, and the other detects the foodborne pathogen. Our datasets have representative data structures (1D and 2D spectroscopy), and we use different ML models for these two datasets. Our chosen datasets are representative of food safety research and thus our experimental results are applicable to other food safety applications/datasets. We do not present the details of the data collection process, which is not the focus of this paper.</p>
<sec>
<title>2.3.1. Dataset 1: Plasma Dosage Classification</title>
<p>As an emerging nonthermal processing technology, plasma is highly effective in inactivating various types of food pathogens in solutions as well as on food contact surfaces (Liao et al., <xref ref-type="bibr" rid="B26">2017</xref>). One of the main goals of our plasma dosage classification dataset is to evaluate whether the plasma dosage can be predicated using the Fourier-Transform Infrared Spectroscopy (FTIR) spectral response of the DNA sample exposed to the plasma treatment. The dataset was obtained by subjecting the substrate to plasma treatment at various dosage levels and characterizing biochemical changes of the substrate with FTIR. The data comprise absorbance intensities at various IR wavenumbers of the substrate under plasma treatment.</p>
<p>Our 1D dosage classification dataset categorizes the plasma dosage into four classes, where classes 1, 2, 3, and 4 represent plasma injection of 0, 2, 5, and 10 min, respectively. In total, we collect 114 samples, where classes 1, 2, 3, and 4 have 27, 27, 30, 30 samples, respectively. To study the sample distribution over the feature domain, we apply K-Means (Arthur and Vassilvitskii, <xref ref-type="bibr" rid="B2">2007</xref>) to group samples into 4 clusters. <xref ref-type="fig" rid="F2">Figure 2A</xref> shows the clustering result. It indicates that the class 2 and the class 3 samples are well clustered, while many of the class 4 samples are located within cluster 2 and cluster 3 (a PCA visualization is given in <xref ref-type="fig" rid="F3">Figure 3</xref>). <xref ref-type="fig" rid="F2">Figure 2B</xref> illustrates a sample for each class. Each sample contains 1,868 values representing the intensity of the IR frequency reflectance in the range of 400&#x02013;4,000 Hz with a step of 2 Hz.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Two datasets are used in our experiments. The top row shows the plasma dosage classification dataset where <bold>(A)</bold> is the clustering result and <bold>(B)</bold> illustrates a sample for each class. The bottom row shows the pathogen detection dataset where <bold>(C)</bold> is the clustering result and <bold>(D)</bold> depicts a sample for each class.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-863261-g0002.tif"/>
</fig>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Demonstration of the samples selected in the passive learning approach and the active learning approach for the plasma dosage classification dataset. <bold>(A)</bold> Passive learning approach. <bold>(B)</bold> Active learning approach with the entropy-based uncertainty sampling. The 40 randomly selected samples for training an initial ML model are marked with &#x0201C;*&#x0201D;. The next 10 samples are tagged with the order number.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-863261-g0003.tif"/>
</fig></sec>
<sec>
<title>2.3.2. Dataset 2: Foodborne Pathogen Detection</title>
<p>Detection of the bacterial foodborne pathogen is one of the critical processes in the food and agricultural industry (Velusamy et al., <xref ref-type="bibr" rid="B46">2010</xref>). It ensures the safety of food products before distribution and reduces the risk of a foodborne illness outbreak. Among various types of bacteria, <italic>E. coli</italic> has been used as an indicator of the fecal contamination and poor sanitary quality of food and water (Krumperman, <xref ref-type="bibr" rid="B20">1983</xref>). T7 bacteriophage or T7 phage has been used as a tool for <italic>E. coli</italic> detection as it infects explicitly <italic>E. coli</italic> cells results in bacteria cell lysis and release of the cell components along with the amplified T7 phage progenies (Yang et al., <xref ref-type="bibr" rid="B50">2020</xref>). Therefore, bacterial cell lysis and T7 phage amplification can indicate <italic>E. coli</italic> contamination in the samples. Fluorescence EEM spectroscopy is an analytical technique providing multi-dimensional information (Li et al., <xref ref-type="bibr" rid="B22">2020</xref>) that has been used to detect organic materials, including bacterial cells (Nakar et al., <xref ref-type="bibr" rid="B35">2020</xref>).</p>
<p>We explore whether 2D EEM spectroscopy can detect the change of the physical and chemical properties of the samples due to the phage infection of <italic>E. coli</italic>. We use classes 1, 2, 3, and 4 to represent phage infected <italic>E. coli, E. coli</italic> only, phage only, and <italic>Listeria</italic> solutions, respectively. The EEM spectra of the resuspended samples were collected with the wavelength range of 250&#x02013;400 nm for excitation and 260&#x02013;450 nm for emission with 5-nm increments. <xref ref-type="fig" rid="F2">Figure 2C</xref> shows the result of grouping the samples into 4 clusters using K-Means. It indicates that most <italic>E. coli</italic> only and the phage only samples are wrongly categorized. <xref ref-type="fig" rid="F2">Figure 2D</xref> visualizes one sample for each class. In the heat maps, we take the logarithm of the frequency intensity at each frequency pair of the excitation and emission wavelength. In total, we collect 160 samples, and each class has 40 samples. Each sample includes 744 numbers representing the fluorescence response for the excitation-emission pairs.</p></sec></sec>
<sec>
<title>2.4. Experimental Setup</title>
<p>Given the same number of labeled samples, we compare the ML model accuracy when different data annotation and model training approaches are applied. We use the LightGBM multi-class classification model (Ke et al., <xref ref-type="bibr" rid="B17">2017</xref>) for the plasma dosage dataset, and the logistic regression classifier (Dreiseitl and Ohno-Machado, <xref ref-type="bibr" rid="B9">2002</xref>) for the pathogen detection dataset. We select these ML models because they show promising results in our related projects. We adopt five-fold cross-validation for both datasets. Specifically, at each cross-validation round, the training and validation sets comprise the four-fold samples and the remaining one-fold samples, respectively. We use the training set for annotating a given number of samples and apply the trained ML model to predict the validation set. Note that the training set is assumed to be unlabeled at the beginning of each cross-validation, and the model is trained with different approaches (refer to <xref ref-type="fig" rid="F1">Figure 1</xref>). The predictions for all samples are obtained by aggregating the predictions of the ML model for each validation set from the five cross-validations.</p>
<p>To determine the hyper-parameters (e.g., learning rate, regularization) of ML models for a given number of labeled samples, we apply Optuna (Akiba et al., <xref ref-type="bibr" rid="B1">2019</xref>) to the passive learning approach and adopt the same hyper-parameters for all approaches. Therefore, the hyper-parameters favor the passive learning approach in our experiments. Using Optuna, we only need to provide reasonable ranges for hyper-parameters, and Optuna can identify the optimal hyper-parameters. We find that hyper-parameters from Optuna tend to have higher model accuracy than the default hyper-parameters.</p></sec></sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<p>We evaluate the four approaches of data annotation and model training using our two datasets. We warm-start a LightGBM model using 40 randomly selected labeled samples for the plasma dosage dataset and then apply different approaches. Likewise, we first train a logistic regression classification model using 25 randomly selected labeled samples for the pathogen detection dataset before applying different approaches. We run 10 experiments to average the results. Since we adopt the five-fold cross-validation, the maximum number of samples for training in each cross-validation is 90 for the plasma dosage task and 127 for the pathogen detection task.</p>
<sec>
<title>3.1. Results of the Active Learning Approach</title>
<p>We first demonstrate the order of samples selected for annotation in the passive learning and AL approaches. Principal Component Analysis (PCA) (Wold et al., <xref ref-type="bibr" rid="B48">1987</xref>) is applied to project the high-dimensional samples into two components for visualization. <xref ref-type="fig" rid="F3">Figure 3</xref> illustrates a trace for the passive learning approach and the AL approach for the dosage classification dataset, respectively. As <xref ref-type="fig" rid="F3">Figure 3A</xref> illustrates, the selected samples in the passive learning can be close to each other and also near the center of the class (e.g., samples 2, 3, and 7), which are less likely to improve the ML model accuracy. In comparison, as <xref ref-type="fig" rid="F3">Figure 3B</xref> shows, the selected samples in the AL approach are close to the boundaries of different classes. In these boundary areas, the ML model is difficult to classify. Thus, training with the samples in the boundary areas is more likely to increase the ML model accuracy.</p>
<p>In the rest of the experiments for the AL approach, we consider (1) uncertainty sampling methods based on least confident and maximum entropy, and (2) minimizing the expected prediction error and the expected log-loss error. <xref ref-type="fig" rid="F4">Figure 4</xref> plots the ML model accuracy vs. the number of labeled samples. The first row and the second row show the model performance for the dosage classification and the pathogen detection datasets, respectively. We show the standard deviation in shaded colors. We can see that AL approach achieves better performance than the passive learning approach, especially for the foodborne pathogen dataset. For example, the uncertainty-based AL approach reduces the number of labeled to only 40 for the pathogen detection dataset, compared to 80 in the passive learning approach, achieving 50% label data reduction. The sampling strategies of minimizing expected errors depend on the current ML model accuracy to calculate the expected errors, which does not perform well if the current ML model has low prediction accuracy (e.g., for the plasma dosage classification dataset). Nonetheless, we do not see degradation of data efficiency.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>The ML model performance vs. the number of labeled samples in the AL approach. The first row and the second row show the results for the dosage classification dataset and the pathogen detection dataset, respectively. <bold>(A,C)</bold> Uncertainty sampling based on least confident and entropy; <bold>(B,D)</bold> minimizing expected prediction error and expected log-loss error based sampling.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-863261-g0004.tif"/>
</fig></sec>
<sec>
<title>3.2. Results of the Semi-supervised Learning Approach</title>
<p>SSL exploits unlabeled samples by assigning pseudo-labels to them and then trains the ML model using both labeled and pseudo-labeled samples. In the evaluation of the SSL approach, we consider self-training and label spreading. We set the number of neighbors in the kNN kernel to 7 and the &#x003C3; in the RBF kernel to 0.1 for both datasets.</p>
<p><xref ref-type="fig" rid="F5">Figure 5</xref> plots the ML model accuracy vs. the number of labeled samples in the SSL approach. The first row and the second row show the model performance for the plasma and pathogen datasets. We visualize the standard deviation in shaded colors. We can see that the SSL approach can improve data efficiency in most experiment scenarios. The exception is the label-spreading for the pathogen detection dataset, which shows much-degraded performance. It is probably because the pathogen detection dataset is poorly clustered (see <xref ref-type="fig" rid="F2">Figure 2C</xref>), and thus the pseudo-labels from label-spreading are mostly incorrect. In the cases of label spreading for the plasma dataset and self-training for the pathogen dataset, SSL approach improves the model accuracy by about 10% when the label samples are few.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The ML model accuracy vs. the number of labeled samples in the semi-supervised learning approach. The first row and the second row show the results for the dosage classification dataset and the pathogen detection dataset, respectively. <bold>(A,C)</bold> Self-training based on maximum confident and minimum entropy; <bold>(B,D)</bold> label spreading based on kNN kernel and RBF kernel.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-863261-g0005.tif"/>
</fig></sec>
<sec>
<title>3.3. Results of the Hybrid Approach</title>
<p>We measure the performance of the hybrid approach using two combinations: random sampling based self-training and RBF-based self-labeling for SSL, both with entropy-based uncertainty sampling for AL. The &#x003C3; value in the label spreading is set to 0.01 for both datasets.</p>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> shows the ML model accuracy for the hybrid approach, where the top row is for the plasma dosage dataset and the bottom row is for the pathogen detection dataset. As we can see, the hybrid approach dramatically improves the data efficiency for ML model training. For example, <xref ref-type="fig" rid="F6">Figure 6B</xref> shows that the hybrid approach reaches the maximum model accuracy with only 45 labeled samples, in comparison to 85 in the baseline, and thus reduces the amount of labeled data by about 50%. The pathogen detection dataset performs even better with the hybrid approach: 50 labeled samples achieve the same maximum model accuracy as 120 labeled samples in the baseline, reducing the number of labeled samples by about 60%. Even though the self-labeling based SSL works poorly for the pathogen dataset (<xref ref-type="fig" rid="F5">Figure 5D</xref>), the hybrid approach still works better than the baseline: 70 labeled samples vs. 90 labeled samples. Overall, the hybrid approach is promising in improving data efficiency.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The ML model accuracy vs. the number of labeled samples in the hybrid approach. The top row and the bottom row show the results for the plasma dosage dataset and the foodborne pathogen detection dataset. <bold>(A,C)</bold> Uncertainty-based AL with self-training based SSL. <bold>(B,D)</bold> Uncertainty-based AL with self-labeling based SSL.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-863261-g0006.tif"/>
</fig></sec></sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>We explore different approaches of data annotation and model training to improve data efficiency for ML applications. Our datasets are general in food safety research. Therefore, our approaches can be applied to general food safety research. In this section, we discuss practical considerations in applying these advanced approaches.</p>
<sec>
<title>4.1. Practicability of Advanced Approaches</title>
<p>The passive learning approach has been adopted for the majority of projects. An essential question is whether more advanced approaches can be effectively and efficiently applied for new projects, considering the overheads of implementing these approaches. We argue that when the annotation process is costly, the benefits of the reduced number of labeled samples overweigh the implementation overheads. In fact, AL and SSL have been successfully applied to many domains (Lookman et al., <xref ref-type="bibr" rid="B30">2019</xref>; Ma et al., <xref ref-type="bibr" rid="B31">2019</xref>; Tamposis et al., <xref ref-type="bibr" rid="B41">2019</xref>; Liu et al., <xref ref-type="bibr" rid="B27">2020</xref>; Wang et al., <xref ref-type="bibr" rid="B47">2020</xref>). For example, the medicine industry applies AL to discover antagonists for dopamine <italic>D</italic><sub>4</sub> (Reutlinger et al., <xref ref-type="bibr" rid="B38">2014</xref>) and CXC-chemokine (Reker et al., <xref ref-type="bibr" rid="B37">2016</xref>), among the estimated range of 10<sup>30</sup>-10<sup>60</sup> drug-like molecules (Naik et al., <xref ref-type="bibr" rid="B34">2013</xref>). In food systems, development of AL and SSL methods can help address challenges of obtaining labeled samples for ML models as labeling in many food safety applications is time-consuming and labor-intensive.</p></sec>
<sec>
<title>4.2. Determining the Optimal Approach</title>
<p>We introduce three advanced data annotation and model training approaches and compare their performance with the passive learning approach. Our evaluation results show that the hybrid approach can dramatically improve data efficiency. Therefore, we suggest the hybrid approach that leverages both AL and SSL. However, the sampling strategy in AL and the pseudo-labeling method in SSL are crucial for the hybrid approach&#x00027;s performance.</p></sec>
<sec>
<title>4.3. Determining the Optimal Sampling Strategy in Active Learning</title>
<p>There are a great variety of sampling strategies available in AL (Settles, <xref ref-type="bibr" rid="B39">2012</xref>). Although it is not possible to obtain a universally good AL sampling strategy (Dasgupta, <xref ref-type="bibr" rid="B7">2004</xref>), we have demonstrated that by using simple sampling strategies such as uncertainty-based sampling, we achieve better data efficiency than the passive learning approach. Simple sampling strategies based on uncertainty and disagreement are recommended if the domain knowledge about the samples and the problems are not available (Settles, <xref ref-type="bibr" rid="B39">2012</xref>). On the other hand, incorporating domain knowledge into AL can further improve data efficiency. For example, Liang et al. (<xref ref-type="bibr" rid="B25">2020</xref>) proposes an expert-in-the-loop AL framework that utilizes language explanations from domain experts to iteratively distinguish misleading breeds of birds, which outperform baseline models that are trained with 40&#x02013;100% more training samples. Automatic selection of sampling strategies has also been extensively studied, (e.g., Hsu and Lin, <xref ref-type="bibr" rid="B15">2015</xref>; Konyushkova et al., <xref ref-type="bibr" rid="B19">2017</xref>).</p>
<p>In our experiments, we warm-start the initial ML model by randomly selecting and training 40 samples for the dosage classification dataset and 25 samples for the pathogen detection dataset. Another common practice is to switch between random sampling (for exploration) and AL sampling strategy (for exploitation). For example, Wang et al. (<xref ref-type="bibr" rid="B47">2020</xref>) train a predictive ML model of gene expression with 44% fewer data by consecutive switching between random sampling and mutual information based AL sampling strategy.</p></sec>
<sec>
<title>4.4. Determining the Optimal Pseudo-Labeling Method in Semi-supervised Learning</title>
<p>In addition to the self-training and the label-spreading methods, semi-supervised learning includes many other methods such as co-training, boosting, and perturbation-based (van Engelen and Hoos, <xref ref-type="bibr" rid="B45">2020</xref>). Similar to AL, SSL does not guarantee better data efficiency than the passive learning approach (Li and Zhou, <xref ref-type="bibr" rid="B23">2011</xref>). Nonetheless, our experiments demonstrate that simple pseudo-labeling methods such as self-training can greatly improve the model accuracy when the underlying samples can be well-clustered and the number of labeled samples are small. Recent advances in SSL show promising results of the perturbation-based semi-supervised neural networks, which empirically and consistently outperforms the passive learning approach (van Engelen and Hoos, <xref ref-type="bibr" rid="B45">2020</xref>). We leave it as future work to evaluate the perturbation-based methods.</p></sec>
<sec>
<title>4.5. Complexity Analysis</title>
<p>The computation capability is also a factor in deciding the approach of data annotation and model training. We formulate the computation complexity as the number of ML model training and ML model inference, as they are more computation demanding than other operations (e.g., graph construction in label-spreading). We denote the total number of samples by <italic>T</italic>, the number of the initial labeled samples by <italic>B</italic>, and the number of final labeled samples by <italic>F</italic>. <xref ref-type="table" rid="T2">Table 2</xref> tabulates the overheads in different approaches, whereas we ignore the model validation overheads as they are the same for all approaches. If the model is heavy and slow to execute, then the computation-intensive methods, e.g., minimizing expected errors, should be avoided.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Computation overheads of different approaches of data annotation and model training.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Approach</bold></th>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>No of model training</bold></th>
<th valign="top" align="left"><bold>No of model inference</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Passive</td>
<td valign="top" align="left">-</td>
<td valign="top" align="center"><italic>F</italic> &#x02212; <italic>B</italic></td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="2">AL</td>
<td valign="top" align="left">Uncertainty-based</td>
<td valign="top" align="center"><italic>F</italic> &#x02212; <italic>B</italic></td>
<td valign="top" align="center"><inline-formula><mml:math id="M38"><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>

<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Minimize expected errors</td>
<td valign="top" align="center"><inline-formula><mml:math id="M39"><mml:mfrac><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M40"><mml:mo>&#x02248;</mml:mo><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mi>C</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x000D7;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr> <tr>
<td valign="middle" align="left" rowspan="2">SSL</td>
<td valign="top" align="left">Self-training</td>
<td valign="top" align="center"><inline-formula><mml:math id="M41"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M42"><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>

<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Label-spreading</td>
<td valign="top" align="center"><italic>F</italic> &#x02212; <italic>B</italic></td>
<td valign="top" align="center">0</td>
</tr> <tr>
<td valign="middle" align="left" rowspan="2">Hybrid</td>
<td valign="top" align="left">Uncertainty &#x0002B; Self-training</td>
<td valign="top" align="center"><inline-formula><mml:math id="M43"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math id="M44"><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>

<tr>
<td valign="top" align="left">Uncertainty &#x0002B; Label-spreading</td>
<td valign="top" align="center"><italic>F</italic> &#x02212; <italic>B</italic></td>
<td valign="top" align="center"><inline-formula><mml:math id="M45"><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>T</mml:mi><mml:mo>-</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:mi>B</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>C is the number of classes for classification; T is the total number of samples, B is the initial labeled samples, and F is the number of final labeled labels</italic>.</p>
</table-wrap-foot>
</table-wrap></sec>
<sec>
<title>4.6. Extension to Regression Problems</title>
<p>In this paper, we use classification to illustrate different approaches of data annotation and model training. We expect our main conclusion to remain valid for regression problems; that is, advanced approaches can help improve data efficiency for model training compared to the passive learning approach. There are many existing works solving regression problems in AL and/or SSL (Georgios et al., <xref ref-type="bibr" rid="B10">2018</xref>; Wu et al., <xref ref-type="bibr" rid="B49">2019</xref>), which we will explore in our future work.</p></sec>
<sec>
<title>4.7. Applicability for Diverse Applications in Food Quality, Traceability, and Safety</title>
<p>The spectroscopy measurement approaches selected for this study represent two distinct spectral methods, namely IR spectroscopy and fluorescence spectroscopy. IR spectroscopy has been proposed for diverse applications in food quality (van de Voort, <xref ref-type="bibr" rid="B44">1992</xref>), traceability (Hennessy et al., <xref ref-type="bibr" rid="B13">2009</xref>) and food safety (Bagcioglu et al., <xref ref-type="bibr" rid="B3">2019</xref>). These diversity of applications are enabled by the unique ability of IR spectroscopy to detect molecular composition of food materials using non-destructive sampling and rapid data collection. ML approaches can complement the current chemometric methods used for the analysis of IR spectroscopy data. Complementary to IR spectroscopy, fluorescence spectroscopy approaches rely on the photo-active properties of fluorophores in food systems and their relationship with changes in food quality and traceability (Hassoun et al., <xref ref-type="bibr" rid="B12">2019</xref>). Similarly, fluorescence properties have also been used to monitor the presence of bacteria in water samples (Cumberland et al., <xref ref-type="bibr" rid="B6">2012</xref>). However, due to significant interference between the food materials and bacterial components there has been limited applications of fluorescence spectroscopy in food safety. To address some of these limitations, this study evaluated the applications of Fluorescence EEM spectroscopy for food safety. EEM is a technique that allows for the complete, quantitative determination of the fluorescence profile of a given material and has been used for biomedical and material characterization applications (Ramsay et al., <xref ref-type="bibr" rid="B36">2018</xref>). Application of EEM spectroscopy including their integration with ML methods can address some of the key challenges in application of fluorescence spectroscopy for food applications. In addition, our approaches of active learning, semi-supervised learning, and hybrid are general, and thus are expected to be applicable to other research domains.</p></sec></sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>In this paper, we target data efficiency of ML applications for spectroscopy analysis in food science. To mitigate the annotation cost by reducing the number of labeled samples, we explore different approaches of data annotation and model training: passive learning, active learning, semi-supervised learning, and the hybrid of active learning and semi-supervised learning. We evaluate these approaches in two spectroscopy datasets and find that advanced approaches can greatly improve data efficiency for ML model training. These approaches are general and thus can be applied to various ML-based food science research.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation, upon request to NN.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>NN and XL conceptualized and supervised the project. QZ supervised the project. HZ conceptualized and implemented the project. NW collected the pathogen detection dataset and conducted a preliminary machine learning model evaluation. HC collected the plasma dosage classification dataset and conducted a preliminary machine learning model evaluation. All authors contributed to the revision of the manuscript and approved the final submitted version.</p></sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>The work was partially supported by NSF through grants (USDA-020-67021-32855 and OIA-2134901) and USDA National Institute of Food and Agriculture grant 2015-68003-23411.</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Akiba</surname> <given-names>T.</given-names></name> <name><surname>Sano</surname> <given-names>S.</given-names></name> <name><surname>Yanase</surname> <given-names>T.</given-names></name> <name><surname>Ohta</surname> <given-names>T.</given-names></name> <name><surname>Koyama</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Optuna: a next-generation hyperparameter optimization framework</article-title>, in <source>International Conference on Knowledge Discovery and Data Mining (KDD)</source> (<publisher-loc>Anchorage, AK</publisher-loc>), <fpage>2623</fpage>&#x02013;<lpage>2631</lpage>. <pub-id pub-id-type="doi">10.1145/3292500.3330701</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Arthur</surname> <given-names>D.</given-names></name> <name><surname>Vassilvitskii</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <article-title>k-means&#x0002B;&#x0002B;: the advantages of careful seeding</article-title>, in <source>ACM-SIAM Symposium on Discrete algorithms (SODA)</source> (<publisher-loc>New Orleans, LA</publisher-loc>), <fpage>1027</fpage>&#x02013;<lpage>1035</lpage>.</citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bagcioglu</surname> <given-names>M.</given-names></name> <name><surname>Fricker</surname> <given-names>M.</given-names></name> <name><surname>Johler</surname> <given-names>S.</given-names></name> <name><surname>Ehling-Schulz</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Detection and identification of <italic>Bacilus cereus, Bacillus cytotoxicus</italic> and <italic>Bacillus thuringiensis</italic> and <italic>Bacillus mycoides</italic> and <italic>Bacillus weihenstephanensis</italic> via machine learning based FTIR spectroscopy</article-title>. <source>Front. Microbiol</source>. <volume>10</volume>, <fpage>902</fpage>. <pub-id pub-id-type="doi">10.3389/fmicb.2019.00902</pub-id><pub-id pub-id-type="pmid">31105681</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ballesteros</surname> <given-names>R.</given-names></name> <name><surname>Intrigliolo</surname> <given-names>D. S.</given-names></name> <name><surname>Ortega</surname> <given-names>J. F.</given-names></name> <name><surname>Ram&#x000ED;rez-Cuesta</surname> <given-names>J. M.</given-names></name> <name><surname>Buesa</surname> <given-names>I.</given-names></name> <name><surname>Moreno</surname> <given-names>M. A.</given-names></name></person-group> (<year>2020</year>). <article-title>Vineyard yield estimation by combining remote sensing, computer vision and artificial neural network techniques</article-title>. <source>Precis. Agric</source>. <volume>21</volume>, <fpage>1242</fpage>&#x02013;<lpage>1262</lpage>. <pub-id pub-id-type="doi">10.1007/s11119-020-09717-3</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chapelle</surname> <given-names>O.</given-names></name> <name><surname>Scholkopf</surname> <given-names>B.</given-names></name> <name><surname>Zien</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <source>Semi-Supervised Learning</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262033589.001.0001</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cumberland</surname> <given-names>S.</given-names></name> <name><surname>Bridgeman</surname> <given-names>J.</given-names></name> <name><surname>Baker</surname> <given-names>A.</given-names></name> <name><surname>Sterling</surname> <given-names>M.</given-names></name> <name><surname>Ward</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Fluorescence spectroscopy as a tool for determining microbial quality in potable water applications</article-title>. <source>Environ. Technol</source>. <volume>33</volume>, <fpage>687</fpage>&#x02013;<lpage>693</lpage>. <pub-id pub-id-type="doi">10.1080/09593330.2011.588401</pub-id><pub-id pub-id-type="pmid">22629644</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dasgupta</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>Analysis of a greedy active learning strategy</article-title>, in <source>International Conference on Neural Information Processing Systems (NIPS)</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>de Sousa</surname> <given-names>C. A. R.</given-names></name> <name><surname>Rezende</surname> <given-names>S. O.</given-names></name> <name><surname>Batista</surname> <given-names>G. E. A. P. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Influence of graph construction on semi-supervised learning</article-title>, in <source>Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD)</source> (<publisher-loc>Prague</publisher-loc>), <fpage>160</fpage>&#x02013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-40994-3_11</pub-id><pub-id pub-id-type="pmid">35618503</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dreiseitl</surname> <given-names>S.</given-names></name> <name><surname>Ohno-Machado</surname> <given-names>L.</given-names></name></person-group> (<year>2002</year>). <article-title>Logistic regression and artificial neural network classification models: a methodology review</article-title>. <source>J. Biomed. Inform</source>. <volume>35</volume>, <fpage>352</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1016/S1532-0464(03)00034-0</pub-id><pub-id pub-id-type="pmid">12968784</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Georgios</surname> <given-names>K.</given-names></name> <name><surname>Stamatis</surname> <given-names>K.</given-names></name> <name><surname>Sotiris</surname> <given-names>K.</given-names></name> <name><surname>Omiros</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>Semi-supervised regression: a recent review</article-title>. <source>J. Intell. Fuzzy Syst</source>. <volume>35</volume>, <fpage>1483</fpage>&#x02013;<lpage>1500</lpage>. <pub-id pub-id-type="doi">10.3233/JIFS-169689</pub-id><pub-id pub-id-type="pmid">30630546</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goujot</surname> <given-names>D.</given-names></name> <name><surname>Meyer</surname> <given-names>X.</given-names></name> <name><surname>Courtois</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>). <article-title>Identification of a rice drying model with an improved sequential optimal design of experiments</article-title>. <source>J. Process Control</source> <volume>22</volume>, <fpage>95</fpage>&#x02013;<lpage>107</lpage>. <pub-id pub-id-type="doi">10.1016/j.jprocont.2011.10.003</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassoun</surname> <given-names>A.</given-names></name> <name><surname>Sahar</surname> <given-names>A.</given-names></name> <name><surname>Lakhal</surname> <given-names>L.</given-names></name> <name><surname>Ait-Kaddour</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Fluorescence spectroscopy as a rapid and non-destructive method for monitoring quality and authenticity of fish and meat products: impact of different preservation conditions</article-title>. <source>LWT Food Sci. Technol</source>. <volume>103</volume>, <fpage>279</fpage>&#x02013;<lpage>292</lpage>. <pub-id pub-id-type="doi">10.1016/j.lwt.2019.01.021</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hennessy</surname> <given-names>S.</given-names></name> <name><surname>Downey</surname> <given-names>G.</given-names></name> <name><surname>ODonnell</surname> <given-names>C. P.</given-names></name></person-group> (<year>2009</year>). <article-title>confirmation of food origin claims by Fourier transform infrared spectroscopy and chemometrics: extra virgin olive from Liguria</article-title>. <source>J. Agric. Food Chem</source>. <volume>57</volume>, <fpage>1735</fpage>&#x02013;<lpage>1741</lpage>. <pub-id pub-id-type="doi">10.1021/jf803714g</pub-id><pub-id pub-id-type="pmid">19206534</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>E-nose combined with chemometrics to trace tomato-juice quality</article-title>. <source>J. Food Eng</source>. <volume>149</volume>, <fpage>38</fpage>&#x02013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1016/j.jfoodeng.2014.10.003</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hsu</surname> <given-names>W.-N.</given-names></name> <name><surname>Lin</surname> <given-names>H.-T.</given-names></name></person-group> (<year>2015</year>). <article-title>Active learning by learning</article-title>, in <source>Association for the Advancement of Artificial Intelligence (AAAI)</source> (<publisher-loc>Austin, TX</publisher-loc>), <fpage>2659</fpage>&#x02013;<lpage>2665</lpage>.</citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <name><surname>Bian</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Discrimination of tomato maturity using hyperspectral imaging combined with graph-based semi-supervised method considering class probability information</article-title>. <source>Food Anal. Methods</source> <volume>14</volume>, <fpage>968</fpage>&#x02013;<lpage>983</lpage>. <pub-id pub-id-type="doi">10.1007/s12161-020-01955-5</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ke</surname> <given-names>G.</given-names></name> <name><surname>Meng</surname> <given-names>Q.</given-names></name> <name><surname>Finley</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Ma</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>LightGBM: a highly efficient gradient boosting decision tree</article-title>, in <source>International Conference on Neural Information Processing Systems (NIPS)</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>3149</fpage>&#x02013;<lpage>3157</lpage>.</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khullar</surname> <given-names>S.</given-names></name> <name><surname>Singh</surname> <given-names>N.</given-names></name></person-group> (<year>2021</year>). <article-title>Machine learning techniques in river water quality modelling: a research travelogue</article-title>. <source>Water Supply</source> <volume>21</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.2166/ws.2020.277</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Konyushkova</surname> <given-names>K.</given-names></name> <name><surname>Raphael</surname> <given-names>S.</given-names></name> <name><surname>Fua</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Learning active learning from data</article-title>, in <source>Conference on Neural Information Processing Systems (NIPS)</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krumperman</surname> <given-names>P. H.</given-names></name></person-group> (<year>1983</year>). <article-title>Multiple antibiotic resistance indexing of <italic>Escherichia coli</italic> to identify high-risk sources of fecal contamination of food</article-title>. <source>Appl. Environ. Microbiol</source>. <volume>46</volume>, <fpage>165</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1128/aem.46.1.165-170.1983</pub-id><pub-id pub-id-type="pmid">6351743</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leca</surname> <given-names>J. M.</given-names></name> <name><surname>Pereira</surname> <given-names>A. C.</given-names></name> <name><surname>Vieira</surname> <given-names>A. C.</given-names></name> <name><surname>Reis</surname> <given-names>M. S.</given-names></name> <name><surname>Marques</surname> <given-names>J. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Optimal design of experiments applied to headspace solid phase microextraction for the quantification of vicinal diketones in beer through gas chromatography-mass spectrometric detection</article-title>. <source>Anal. Chim. Acta</source> <volume>887</volume>, <fpage>101</fpage>&#x02013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1016/j.aca.2015.06.044</pub-id><pub-id pub-id-type="pmid">26320791</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Gao</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>New advances in fluorescence excitation-emission matrix spectroscopy for the characterization of dissolved organic matter in drinking water treatment: a review</article-title>. <source>Chem. Eng. J</source>. <volume>381</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1016/j.cej.2019.122676</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.-F.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.-H.</given-names></name></person-group> (<year>2011</year>). <article-title>Towards making unlabeled data never hurt</article-title>, in <source>International Conference on Machine Learning (ICML)</source> (<publisher-loc>Bellevue, WA</publisher-loc>), <fpage>1081</fpage>&#x02013;<lpage>1088</lpage>.</citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liakos</surname> <given-names>K. G.</given-names></name> <name><surname>Busato</surname> <given-names>P.</given-names></name> <name><surname>Moshou</surname> <given-names>D.</given-names></name> <name><surname>Pearson</surname> <given-names>S.</given-names></name> <name><surname>Bochtis</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>Machine learning in agriculture: a review</article-title>. <source>Sensors</source> <volume>18</volume>, <fpage>1</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.3390/s18082674</pub-id><pub-id pub-id-type="pmid">30110960</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>W.</given-names></name> <name><surname>Zou</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>ALICE: active learning with contrastive natural language explanations</article-title>, in <source>Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>, <fpage>4380</fpage>&#x02013;<lpage>4391</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2020.emnlp-main.355</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>D.</given-names></name> <name><surname>Xiang</surname> <given-names>Q.</given-names></name> <name><surname>Ahn</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <name><surname>Ye</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Inactivation mechanisms of non-thermal plasma on microbes: a review</article-title>. <source>Food Control</source> <volume>75</volume>, <fpage>83</fpage>&#x02013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1016/j.foodcont.2016.12.021</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>N.</given-names></name> <name><surname>Chen</surname> <given-names>C.-B.</given-names></name> <name><surname>Kumara</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Semi-supervised learning algorithm for identifying high-priority drug-drug interactions through adverse event reports</article-title>. <source>IEEE J. Biomed. Health Inform</source>. <volume>24</volume>, <fpage>57</fpage>&#x02013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1109/JBHI.2019.2932740</pub-id><pub-id pub-id-type="pmid">31395567</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>S.-F.</given-names></name></person-group> (<year>2012</year>). <article-title>Robust and scalable graph-based semisupervised learning</article-title>. <source>Proc. IEEE</source> <volume>100</volume>, <fpage>2624</fpage>&#x02013;<lpage>2638</lpage>. <pub-id pub-id-type="doi">10.1109/JPROC.2012.2197809</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>B.</given-names></name> <name><surname>Bian</surname> <given-names>J.</given-names></name> <name><surname>Chapelle</surname> <given-names>O.</given-names></name> <name><surname>Ya Zhang</surname> <given-names>Y. I.</given-names></name> <name><surname>Chang</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>Active learning for ranking through expected loss optimization</article-title>. <source>IEEE Trans. Knowledge Data Eng</source>. <volume>27</volume>, <fpage>1180</fpage>&#x02013;<lpage>1191</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2014.2365785</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lookman</surname> <given-names>T.</given-names></name> <name><surname>Balachandran</surname> <given-names>P. V.</given-names></name> <name><surname>Xue</surname> <given-names>D.</given-names></name> <name><surname>Yuan</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>Active learning in materials science with emphasis on adaptive sampling using uncertainties for targeted design</article-title>. <source>Comput. Mater</source>. <volume>5</volume>, <fpage>1</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1038/s41524-019-0153-8</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>W.</given-names></name> <name><surname>Cheng</surname> <given-names>F.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Wen</surname> <given-names>Q.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Probabilistic representation and inverse design of metamaterials based on a deep generative model with semi-supervised learning strategy</article-title>. <source>Adv. Mater</source>. <volume>31</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1002/adma.201901111</pub-id><pub-id pub-id-type="pmid">31259443</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Munson-McGee</surname> <given-names>S. H.</given-names></name></person-group> (<year>2014a</year>). <article-title>D- and G-optimal experimental designs for the partition coefficient in freeze concentration</article-title>. <source>J. Food Eng</source>. <volume>121</volume>, <fpage>80</fpage>&#x02013;<lpage>86</lpage>. <pub-id pub-id-type="doi">10.1016/j.jfoodeng.2013.08.018</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Munson-McGee</surname> <given-names>S. H.</given-names></name></person-group> (<year>2014b</year>). <article-title>D-optimal experimental designs for uniaxial expression</article-title>. <source>J. Food Process Eng</source>. <volume>37</volume>, <fpage>248</fpage>&#x02013;<lpage>256</lpage>. <pub-id pub-id-type="doi">10.1111/jfpe.12080</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naik</surname> <given-names>A. W.</given-names></name> <name><surname>Kangas</surname> <given-names>J. D.</given-names></name> <name><surname>Langmead</surname> <given-names>C. J.</given-names></name> <name><surname>Murphy</surname> <given-names>R. F.</given-names></name></person-group> (<year>2013</year>). <article-title>Efficient modeling and active learning discovery of biological responses</article-title>. <source>PLoS ONE</source> <volume>8</volume>, <fpage>e83996</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0083996</pub-id><pub-id pub-id-type="pmid">24358322</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakar</surname> <given-names>A.</given-names></name> <name><surname>Schmilovitch</surname> <given-names>Z.</given-names></name> <name><surname>Vaizel-Ohayon</surname> <given-names>D.</given-names></name> <name><surname>Kroupitski</surname> <given-names>Y.</given-names></name> <name><surname>Borisover</surname> <given-names>M.</given-names></name> <name><surname>Saldinger</surname> <given-names>S. S.</given-names></name></person-group> (<year>2020</year>). <article-title>Quantification of bacteria in water using PLS analysis of emission spectra of fluorescence and excitation-emission matrices</article-title>. <source>Water Res</source>. <volume>169</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/j.watres.2019.115197</pub-id><pub-id pub-id-type="pmid">31670087</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ramsay</surname> <given-names>H.</given-names></name> <name><surname>Simon</surname> <given-names>D. E.</given-names></name> <name><surname>Steele</surname></name> <name><surname>Hebert</surname> <given-names>A.</given-names></name> <name><surname>Oleschuk</surname> <given-names>R. D.</given-names></name> <name><surname>Stamplecoskie</surname> <given-names>K. G.</given-names></name></person-group> (<year>2018</year>). <article-title>The power of fluorescence excitation-emission matrix (EEM) spectroscopy in the identification and characterization of complex mixtures of fluorescent silver clusters</article-title>. <source>RSC Adv</source>. <volume>8</volume>, <fpage>42080</fpage>&#x02013;<lpage>42086</lpage>. <pub-id pub-id-type="doi">10.1039/C8RA08751B</pub-id><pub-id pub-id-type="pmid">35558801</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reker</surname> <given-names>D.</given-names></name> <name><surname>Schneider</surname> <given-names>P.</given-names></name> <name><surname>Schneider</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Multi-objective active machine learning rapidly improves structure-activity models and reveals new protein-protein interaction inhibitors</article-title>. <source>Chem. Sci</source>. <volume>7</volume>, <fpage>3919</fpage>&#x02013;<lpage>3927</lpage>. <pub-id pub-id-type="doi">10.1039/C5SC04272K</pub-id><pub-id pub-id-type="pmid">30155037</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reutlinger</surname> <given-names>M.</given-names></name> <name><surname>Rodrigues</surname> <given-names>T.</given-names></name> <name><surname>Schneider</surname> <given-names>P.</given-names></name> <name><surname>Schneider</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>Multi-objective molecular <italic>de novo</italic> design by adaptive fragment prioritization</article-title>. <source>Angew. Int. Ed. Chem</source>. <volume>53</volume>, <fpage>4244</fpage>&#x02013;<lpage>4248</lpage>. <pub-id pub-id-type="doi">10.1002/anie.201310864</pub-id><pub-id pub-id-type="pmid">24623390</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Settles</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <source>Active Learning</source>. <publisher-loc>San Rafael, CA</publisher-loc>: <publisher-name>Morgan &#x00026; Claypool</publisher-name>. <pub-id pub-id-type="doi">10.2200/S00429ED1V01Y201207AIM018</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharma</surname> <given-names>M.</given-names></name> <name><surname>Bilgic</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Evidence-based uncertainty sampling for active learning</article-title>. <source>Data Mining Knowledge Discov</source>. <volume>31</volume>, <fpage>164</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1007/s10618-016-0460-3</pub-id><pub-id pub-id-type="pmid">35040795</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tamposis</surname> <given-names>I. A.</given-names></name> <name><surname>Tsirigos</surname> <given-names>K. D.</given-names></name> <name><surname>Theodoropoulou</surname> <given-names>M. C.</given-names></name> <name><surname>Kontou</surname> <given-names>P. I.</given-names></name> <name><surname>Bagos</surname> <given-names>P. G.</given-names></name></person-group> (<year>2019</year>). <article-title>Semi-supervised learning of hidden markov models for biological sequence analysis</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>2208</fpage>&#x02013;<lpage>2215</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty910</pub-id><pub-id pub-id-type="pmid">30445435</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Triguero</surname> <given-names>I.</given-names></name> <name><surname>Garcia</surname> <given-names>S.</given-names></name> <name><surname>Herrera</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>Self-labeled techniques for semi-supervised learning: taxonomy, software and empirical study</article-title>. <source>Knowledge Inform. Syst</source>. <volume>42</volume>, <fpage>245</fpage>&#x02013;<lpage>284</lpage>. <pub-id pub-id-type="doi">10.1007/s10115-013-0706-y</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tsakanikas</surname> <given-names>P.</given-names></name> <name><surname>Karnavas</surname> <given-names>A.</given-names></name> <name><surname>Panagou</surname> <given-names>E. Z.</given-names></name> <name><surname>Nychas</surname> <given-names>G. J.</given-names></name></person-group> (<year>2020</year>). <article-title>A machine learning workflow for raw food spectroscopic classification in a future industry</article-title>. <source>Nat. Sci. Rep</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-68156-2</pub-id><pub-id pub-id-type="pmid">32641761</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van de Voort</surname> <given-names>F. R.</given-names></name></person-group> (<year>1992</year>). <article-title>Fourier transform infrared spectroscopy applied to food analysis</article-title>. <source>Food Res. Int</source>. <volume>25</volume>, <fpage>397</fpage>&#x02013;<lpage>403</lpage>. <pub-id pub-id-type="doi">10.1016/0963-9969(92)90115-L</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Engelen</surname> <given-names>J. E.</given-names></name> <name><surname>Hoos</surname> <given-names>H. H.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey on semi-supervised learning</article-title>. <source>Mach. Learn</source>. <volume>109</volume>, <fpage>373</fpage>&#x02013;<lpage>440</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-019-05855-6</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Velusamy</surname> <given-names>V.</given-names></name> <name><surname>Arshak</surname> <given-names>K.</given-names></name> <name><surname>Korostynska</surname> <given-names>O.</given-names></name> <name><surname>Oliwa</surname> <given-names>K.</given-names></name> <name><surname>Adley</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>An overview of foodborne pathogen detection: in the perspective of biosensors</article-title>. <source>Biotechnol. Adv</source>. <volume>28</volume>, <fpage>232</fpage>&#x02013;<lpage>254</lpage>. <pub-id pub-id-type="doi">10.1016/j.biotechadv.2009.12.004</pub-id><pub-id pub-id-type="pmid">20006978</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Rai</surname> <given-names>N.</given-names></name> <name><surname>Pereira</surname> <given-names>B. M. P.</given-names></name> <name><surname>Eetemadi</surname> <given-names>A.</given-names></name> <name><surname>Tagkopoulos</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>Accelerated knowledge discovery from omics data by optimal experimental design</article-title>. <source>Nat. Commun</source>. <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-18785-y</pub-id><pub-id pub-id-type="pmid">33024104</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wold</surname> <given-names>S.</given-names></name> <name><surname>Esbensen</surname> <given-names>K.</given-names></name> <name><surname>Geladi</surname> <given-names>P.</given-names></name></person-group> (<year>1987</year>). <article-title>Principal component analysis</article-title>. <source>Chemometr. Intell. Lab. Syst</source>. <volume>2</volume>, <fpage>37</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/0169-7439(87)80084-9</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>D.</given-names></name> <name><surname>Lin</surname> <given-names>C.-T.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Active learning for regression using greedy sampling</article-title>. <source>Inform. Sci</source>. <volume>474</volume>, <fpage>90</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2018.09.060</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Wisuthiphaet</surname> <given-names>N.</given-names></name> <name><surname>Young</surname> <given-names>G. M.</given-names></name> <name><surname>Nitin</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>Rapid detection of <italic>Escherichia coli</italic> using bacteriophage-induced lysis and image analysis</article-title>. <source>PLoS ONE</source> <volume>15</volume>, <fpage>e0233853</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0233853</pub-id><pub-id pub-id-type="pmid">32502212</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Bousquet</surname> <given-names>O.</given-names></name> <name><surname>Lal</surname> <given-names>T. N.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name> <name><surname>Scholkopf</surname> <given-names>B.</given-names></name></person-group> (<year>2003</year>). <article-title>Learning with local and global consistency</article-title>, in <source>International Conference on Neural Information Processing Systems (NIPS)</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>321</fpage>&#x02013;<lpage>328</lpage>.</citation></ref>
</ref-list>
</back>
</article>