<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Malar.</journal-id>
<journal-title>Frontiers in Malaria</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Malar.</abbrev-journal-title>
<issn pub-type="epub">2813-7396</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmala.2024.1250220</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Malaria</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Metrics to guide development of machine learning algorithms for malaria diagnosis</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Delahunt</surname>
<given-names>Charles B.</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/581272"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Gachuhi</surname>
<given-names>Noni</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2710023/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Horning</surname>
<given-names>Matthew P.</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2346815"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>Global Health Labs</institution>, <addr-line>Bellevue, WA</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Anita Ghansah, Noguchi Memorial Institute for Medical Research, Ghana</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Geoffrey H. Siwo, University of Michigan, United States</p>
<p>Sonal Kale, National Institute of Allergy and Infectious Diseases (NIH), United States</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Charles B. Delahunt, <email xlink:href="mailto:charles.delahunt@ghlabs.org">charles.delahunt@ghlabs.org</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>04</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>2</volume>
<elocation-id>1250220</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>03</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2024 Delahunt, Gachuhi and Horning</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Delahunt, Gachuhi and Horning</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Automated malaria diagnosis is a difficult but high-value target for machine learning (ML), and effective algorithms could save many thousands of children&#x2019;s lives. However, current ML efforts largely neglect crucial use case constraints and are thus not clinically useful. Two factors in particular are crucial to developing algorithms translatable to clinical field settings: (i) clear understanding of the clinical needs that ML solutions must accommodate; and (ii) task-relevant metrics for guiding and evaluating ML models. Neglect of these factors has seriously hampered past ML work on malaria, because the resulting algorithms do not align with clinical needs. In this paper we address these two issues in the context of automated malaria diagnosis via microscopy on Giemsa-stained blood films. The intended audience are ML researchers as well as anyone evaluating the performance of ML models for malaria. First, we describe why domain expertise is crucial to effectively apply ML to malaria, and list technical documents and other resources that provide this domain knowledge. Second, we detail performance metrics tailored to the clinical requirements of malaria diagnosis, to guide development of ML models and evaluate model performance through the lens of clinical needs (versus a generic ML lens). We highlight the importance of a patient-level perspective, interpatient variability, false positive rates, limit of detection, and different types of error. We also discuss reasons why ROC curves, AUC, and F1, as commonly used in ML work, are poorly suited to this context. These findings also apply to other diseases involving parasite loads, including neglected tropical diseases (NTDs) such as schistosomiasis.</p>
</abstract>
<kwd-group>
<kwd>malaria</kwd>
<kwd>NTDs</kwd>
<kwd>schistosomiasis</kwd>
<kwd>metrics</kwd>
<kwd>machine learning</kwd>
<kwd>sensitivity</kwd>
<kwd>specificity</kwd>
<kwd>limit of detection</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="1"/>
<equation-count count="11"/>
<ref-count count="53"/>
<page-count count="14"/>
<word-count count="8495"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Case Management</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Malaria and some neglected tropical diseases (e.g., schistosomiasis) involve parasite loads that can be detected in microscopy images of a substrate (e.g., blood or filtered urine). They are thus amenable, though difficult, targets for automated diagnosis via machine learning (ML) methods. These diseases are also very high-value ML targets: They are serious global health challenges affecting hundreds of millions of people, especially children, in underserved populations (<xref ref-type="bibr" rid="B47">WHO, 2019</xref>; <xref ref-type="bibr" rid="B4">BMGF, 2023</xref>; <xref ref-type="bibr" rid="B50">WHO, 2023</xref>). Because microscopy on Giemsa-stained blood films is a widespread clinical diagnostic, effective automated ML systems have great potential benefit since they can naturally fit into clinical workflow and enable hard-pressed clinics to treat more patients. In addition, drug resistance sentinel sites have a heavy demand for parasite quantitation on Giemsa-stained blood films, a use case well-suited to automated microscopy.</p>
<p>However, ML methods developed for malaria diagnosis<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref> using Giemsa-stained blood films have so far largely failed to translate to useful deployment, for several reasons.</p>
<list list-type="simple">
<list-item>
<p>(i) The task is difficult: malaria parasites are small and closely resemble certain artifact types; field blood films are highly variable in stain color, types and numbers of artifacts, and parasite appearance; digital microscopy images vary in color, quality, and resolution; images are often full of distractor objects; and the low limits of detection required for clinical use result in low signal-to-noise ratios (e.g., one parasite per 30 large fields of view).</p>
</list-item>
<list-item>
<p>(ii) ML development has typically proceeded in a heavily ML-centric mindset, without careful attention to (or even knowledge of) the domain specifics, use cases, and clinical requirements of malaria. This yields algorithms that, almost by design, fail to meet clinical needs and cannot be built upon (see <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>).</p>
</list-item>
<list-item>
<p>(iii) ML development can only optimize what is measured, so a crucial prerequisite for successful development is a set of task-relevant metrics (<xref ref-type="bibr" rid="B20">Maier-Hein et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B34">Reinke and Tizabi, 2024</xref>). These tailored metrics have largely been lacking for malaria, for which ML development has instead been guided by generic and ill-suited ML metrics such as object-level ROC curves.</p>
</list-item>
</list>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Left: Effective ML (AI) solutions must interlock with domain requirements and will be shaped by non-ML pressures arising from the use case. Right: Solutions developed with an ML-centric perspective, neglecting the use case, will be mainly shaped by purely ML concerns and will thus fail to match clinical needs (&#x201c;interlocking&#x201d; metaphor due to Dr. Scott McClelland; jigsaw outline from <uri xlink:href="https://draradech.github.io/jigsaw/jigsaw-hex.html">https://draradech.github.io/jigsaw/jigsaw-hex.html</uri>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmala-02-1250220-g001.tif"/>
</fig>
<p>This paper seeks to accelerate the ML community&#x2019;s progress toward translatable solutions for malaria diagnosis, by describing tools and techniques which we have found to be essential for development of clinically effective ML algorithms. The intended audience are ML researchers as well as anyone evaluating the performance of ML models for malaria. It captures lessons learned by our group over a decade of applying ML to malaria diagnosis. The resulting algorithms (<xref ref-type="bibr" rid="B12">Delahunt et&#xa0;al., 2014a</xref>; <xref ref-type="bibr" rid="B23">Mehanian et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al., 2019</xref>) are, to our knowledge, the most effective and also the most extensively field tested in clinical trials (<xref ref-type="bibr" rid="B35">Torres et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B36">Vongpromek et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B15">Horning et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B9">Das et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B33">Rees-Channer et&#xa0;al., 2023</xref>) yet built for fully automated diagnosis of malaria on Giemsa-stained blood films in clinical settings (we note that in medicine the gold standard of evidence is the third-party clinical trial, not the ML-style comparison). These field trials showed that our algorithms, though state-of-the-art, still fall short of the clinical demands, and highlight the need for more robust algorithms to truly impact this category of malaria diagnosis.</p>
<p>The paper is structured as follows: Section 2 details aspects of ML work that depend on a grasp of the clinical use-case (e.g., how the disease is diagnosed in the field), lists malaria documents especially relevant to ML work, and discusses other domain knowledge resources. Section 3 first describes serious problems with commonly-used ML metrics, then describes ML metrics tailored specifically to malaria and NTDs that can be applied during development of ML algorithms to optimize and evaluate their clinical effectiveness. We focus throughout on malaria, but sometimes mention NTDs (e.g., schistosomiasis) because the same principles and methods apply.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>The clinical use case</title>
<p>To be clinically useful an ML solution must fit into a larger, ML-independent context. It must interlock with other pieces that are shaped by clinician needs, site requirements, protocols currently in use, patient needs, business environment, etc (<xref ref-type="bibr" rid="B37">Wiens et&#xa0;al., 2019</xref>). This strong constraint to mesh with non-ML considerations is often overlooked by ML practitioners, leading to algorithms that are elegant (from an ML perspective) but ill-suited for use (from a clinical perspective) (<xref ref-type="bibr" rid="B18">Koller and Bengio, 2018</xref>).</p>
<p>In particular, a clinically useful ML algorithm must fit into an existing care structure and meet or exceed existing clinical performance targets. So understanding these clinical constraints is a basic prerequisite for algorithm development. (We set aside the complex case of a disruptive technology potentially altering existing care protocols. Such cases of course require careful analysis.)</p>
<p>This section discusses some crucial points to consider, and lists resources for learning about malaria use cases.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Important domain specifics</title>
<p>Several domain-specific details are fundamental to effective algorithm development:</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Basic facts about the clinical needs</title>
<p>For example, what are the proper uses of thick vs. thin blood films for malaria?</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Performance metrics relevant in the clinic</title>
<p>Examples include patient-level sensitivity and specificity, and limit of detection (LoD). This knowledge enables ML researchers to tailor salient metrics to guide algorithm development (like those we give in Section 3), define objective functions, do internal assessment, and report algorithm results meaningfully.</p>
</sec>
<sec id="s2_1_3">
<label>2.1.3</label>
<title>Performance specifications</title>
<p>Clinicians are unwilling to reduce patient care standards, so ML models must perform at least as well as current practice to be deployable. Field performance requirements are thus vital concerns, even if a particular model iteration does not attain them (since the work can then be built upon or extended).</p>
</sec>
<sec id="s2_1_4">
<label>2.1.4</label>
<title>Domain-specific obstacles and shortcuts</title>
<p>Some difficult details need special treatment, and others allow for valuable shortcuts. For example, malaria parasites can exist at various depths of a thick blood film, so a single image plane will not capture all parasites in focus. On the plus side, the nuclei of white blood cells (WBCs) are plentiful in thick films and stain similarly to malaria parasite nuclei, so they can serve as a ready-made color reference for the rare (or absent) parasites. Shortcuts matter because generic methods applied as-is are unlikely to hit clinical performance requirements, which is a much harder task than simply outdoing another generic method in a ML-style comparison.</p>
</sec>
<sec id="s2_1_5">
<label>2.1.5</label>
<title>Structuring annotations and training sets</title>
<p>Annotations and training data are central to ML success, and must be tailored to the task. For example, malaria ring forms (the youngest parasite stage) typically have both a round nucleus and a crescent-shaped cytoplasm (examples in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>). However, after drug treatment the rings often lack visible cytoplasm, appearing in thick films as dark round dots which are very similar to a common distractor type. As a result, they have outsized impact on decision boundaries and require special care as to annotation and inclusion in training sets.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Examples of interpatient variability, thin blood films. Red arrows point to parasites, green arrow is a white blood cell. <bold>(A)</bold> Typical &#x201c;ideal&#x201d; blood film. <bold>(B)</bold> Poor quality. <bold>(C)</bold> Malaria-negative, with numerous stain and debris artifacts. By permission from <xref ref-type="bibr" rid="B11">Delahunt et al. (2019)</xref>.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmala-02-1250220-g002.tif"/>
</fig>
<p>Avenues to acquire vital domain expertise include (i) documentation and (ii) connecting with domain experts.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Documentation</title>
<p>Effective ML solutions need to design in accommodations to non-ML (e.g., clinical) constraints. Therefore, literature review to inform ML work should extend well beyond ML methods and focus on the clinical use-case itself, without an ML-centric filter. Documentation of use cases and standards of care are published by various agencies, including the World Health Organization (WHO), ministries of health, and non-government organizations (e.g., the Bill and Melinda Gates Foundation, the Global Fund, and the Worldwide Antimalarial Resistance Network).</p>
<p>Below we list some references that are especially relevant to ML researchers designing algorithms for automated malaria diagnosis using Giemsa-stained blood films.</p>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Appropriate evidence for ML</title>
<list list-type="bullet">
<list-item>
<p>The WHO has issued guidelines on how to generate meaningful evidence for ML-based medical tools (<xref ref-type="bibr" rid="B48">WHO, 2021a</xref>, especially section 1). This document is important for ML as applied to any medical use case. Crucially, evidence of algorithm performance during development must be firmly grounded in the clinical use case. This requirement underpins the metrics described below in 3.2&#x2013;3.11.</p>
</list-item>
</list>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Protocols for malaria microscopy</title>
<p>Various groups have published diagnosis protocols which detail the clinical task.</p>
<list list-type="bullet">
<list-item>
<p>WHO&#x2019;s guidelines are an essential resource (<xref ref-type="bibr" rid="B41">WHO, 2010</xref>) and (<xref ref-type="bibr" rid="B43">WHO, 2016a</xref>) (see especially SOPs 8 and 9 for diagnosis and quantitation).</p>
</list-item>
<list-item>
<p>Ministries of Health also have useful protocols, e.g., Peru (<xref ref-type="bibr" rid="B24">Ministerio de Salud, 2003</xref>) and USA (<xref ref-type="bibr" rid="B6">CDC, 2023a</xref>).</p>
</list-item>
<list-item>
<p>WWARN and the WHO have developed protocols tailored to research contexts (e.g., drug resistance sentinel sites) (<xref ref-type="bibr" rid="B44">WHO, 2016b</xref>).</p>
</list-item>
</list>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>Evaluation tests</title>
<list list-type="bullet">
<list-item>
<p>The WHO has developed a system to evaluate malaria microscopists. This uses a set of 56 blood slides with carefully specified parasitemias and species (<xref ref-type="bibr" rid="B45">WHO, 2016c</xref>, section 6). The &#x201c;WHO 56&#x201d; evaluation reflects the tasks and accuracies required in the clinic and is thus a valuable and challenging test for ML algorithms. Its difficulty gives an appreciation of the skills of human field microscopists. The defined competency levels offer clear and clinically meaningful performance targets for ML algorithms. Note that the &#x201c;WHO 56&#x201d; differs slightly from the previous version (the &#x201c;WHO 55&#x201d;) found in (<xref ref-type="bibr" rid="B40">WHO, 2009</xref>).</p>
</list-item>
<list-item>
<p>A similar but distinct evaluation set of blood slides, tailored to research rather than clinical contexts, is detailed in (<xref ref-type="bibr" rid="B44">WHO, 2016b</xref>).</p>
</list-item>
<list-item>
<p>Peru&#x2019;s quality control protocols implicitly describe performance requirements (<xref ref-type="bibr" rid="B24">Ministerio de Salud, 2003</xref>, sections 7.2 and 9.2).</p>
</list-item>
</list>
</sec>
<sec id="s2_2_4">
<label>2.2.4</label>
<title>Neglected tropical diseases</title>
<list list-type="bullet">
<list-item>
<p>The WHO has defined target product profiles, including sensitivity and specificity requirements, that are relevant to automated ML systems targeting schistosomiasis (<xref ref-type="bibr" rid="B39">WHO, 2002</xref>; <xref ref-type="bibr" rid="B49">WHO, 2021b</xref>).</p>
</list-item>
</list>
</sec>
<sec id="s2_2_5">
<label>2.2.5</label>
<title>Other performance specifications</title>
<list list-type="bullet">
<list-item>
<p>The above documents also provide detail concerning other general product requirements relevant to any ML solution that aims for translation to clinics. These issues include time-to-result, throughput, electricity/battery constraints, price, and (implicitly) computational constraints.</p>
</list-item>
</list>
</sec>
<sec id="s2_2_6">
<label>2.2.6</label>
<title>ML publications</title>
<list list-type="bullet">
<list-item>
<p>Some ML papers (e.g., <xref ref-type="bibr" rid="B15">Horning et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B28">Oyibo et&#xa0;al., 2023</xref>) cite non-ML documents relevant to use case, but this is not (yet) common practice. So ML-based literature search is insufficient.</p>
</list-item>
</list>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Domain experts</title>
<p>
<italic>Domain experts</italic> are a vital source of guidance and collaboration. They include field experts, i.e. those who work in field clinics or who do field-based research; and subject matter experts, such as WHO personnel and long-time researchers in the space (these groups overlap). The value of their experience and insight to effective algorithm development cannot be overstated.</p>
<p>As an example, our group&#x2019;s entire ML program for malaria diagnosis has depended absolutely upon expert input from a technical advisory panel, as well as on continued contacts and advice from field clinics. To the degree that our work has succeeded, this expert input has been the key ingredient (along with the closely entwined matter of data collection and curation). We would argue that ML development can only progress toward clinically useful algorithms when domain expertise is somehow integrated into the team (recent examples include <xref ref-type="bibr" rid="B52">Yang et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B22">Manescu et&#xa0;al., 2020a</xref>, <xref ref-type="bibr" rid="B21">b</xref>; <xref ref-type="bibr" rid="B17">Kassim et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B31">Poostchi et&#xa0;al., 2018a</xref>; <xref ref-type="bibr" rid="B53">Yu et&#xa0;al., 2023</xref>; and, for schistosomiasis, <xref ref-type="bibr" rid="B1">Armstrong et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B29">Oyibo et&#xa0;al., 2022</xref>, <xref ref-type="bibr" rid="B28">2023</xref>).</p>
<p>Connecting with such experts is made easier by two things. First, people (on average) love to talk about their work. Second, field experts are often (again, on average) open to engaging with ML solutions and happy to co-author serious research.</p>
<p>Sources for contacts include: (i) published work, e.g., who is leading and authoring/co-authoring relevant studies; (ii) academic institutions with concentrations of research in the space; (iii) online interest groups, e.g., on LinkedIn; and (iv) non-ML conferences, their attendees, and proceedings, e.g., the American Society of Tropical Medicine and Hygiene.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Salient metrics for ML work</title>
<p>Salient metrics are essential to ML work, both to guide development and to report results meaningfully. Unfortunately, the metrics routinely applied to ML work on malaria (e.g., object-level precision, recall, AUC, and F1 score) have disqualifying drawbacks in the malaria context.</p>
<p>A 2018 review of automated malaria detection papers (<xref ref-type="bibr" rid="B31">Poostchi et&#xa0;al., 2018b</xref>) described serious problems (which still persist): reported metrics are incomplete and not comparable between studies; metrics are object-based (not patient-based) and are thus not relevant to the clinical task; train and test sets contain objects from the same patient, which contradicts the patient-level focus; and datasets are too small. We note that in addition, incorrect assumptions are built into algorithms: for example, diagnosis on thin blood films is common in ML papers, despite being contrary to clinical practice due to practical obstacles (Long, pers. comm.; <xref ref-type="bibr" rid="B43">WHO, 2016a</xref>; though see recent work on thin film spreaders in <xref ref-type="bibr" rid="B26">Noul, 2023</xref>; <xref ref-type="bibr" rid="B27">Nowak et&#xa0;al., 2023</xref>).</p>
<p>In this section, we first (3.1) discuss some problems with commonly-used ML metrics and argue that these should not be used to report ML results for malaria.</p>
<p>We then describe in detail (3.2&#x2013;3.11) some alternative metrics which have high clinical relevance for the malaria use case. These metrics are effective tools both to guide ML development and to report meaningful ML performance results not only for malaria, but also other diseases involving parasite loads such as malaria, NTDs, or more generally any pathology where diagnosis is determined by the presence of a variable number of abnormal objects (e.g., pixels or cells in a histopathology slide).</p>
<p>A full list of mathematical notation is given in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>List of notation, using malaria as the reference context.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left" colspan="2">General terms</th>
<th valign="top" align="left" colspan="2"/>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">LoD</td>
<td valign="top" align="left">Limit of Detection</td>
<td valign="top" align="left">NTD</td>
<td valign="top" align="left">Neglected Tropical Disease</td>
</tr>
<tr>
<td valign="top" align="left">RBC</td>
<td valign="top" align="left">Red Blood Cell</td>
<td valign="top" align="left">WBC</td>
<td valign="top" align="left">White Blood Cell</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>P</italic>
</td>
<td valign="top" align="left">a parasitemia, in p/<italic>&#xb5;L</italic>
</td>
<td valign="top" align="left">#</td>
<td valign="top" align="left">&#x201d;number of&#x201d;</td>
</tr>
<tr>
<td valign="top" align="left">
<italic>&#xb5;L</italic>
</td>
<td valign="top" align="left">microliter</td>
<td valign="top" align="left">p/<italic>&#xb5;L</italic>
</td>
<td valign="top" align="left">parasites per microliter</td>
</tr>
<tr>
<td valign="top" align="left">
<italic>V</italic>
</td>
<td valign="top" align="left">estimated volume of blood examined</td>
<td valign="top" align="left">
<italic>cV</italic>
</td>
<td valign="top" align="left">clinically-relevant volume (1 <italic>&#xb5;L</italic> of blood)</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>fp</italic>
</td>
<td valign="top" align="left"># false positive objects in <italic>V</italic>
</td>
<td valign="top" align="left">
<italic>FP</italic>
</td>
<td valign="top" align="left"># false positive objects per <italic>cV</italic> for one patient</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>tp</italic>
</td>
<td valign="top" align="left"># true positive objects in <italic>V</italic>
</td>
<td valign="top" align="left">
<italic>TP</italic>
</td>
<td valign="top" align="left"># true positive objects per <italic>cV</italic> for one patient</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>n</italic>
</td>
<td valign="top" align="left"># suspected parasites in <italic>V</italic> = <italic>tp</italic> + <italic>fp</italic>
</td>
<td valign="top" align="left">
<italic>N</italic>
</td>
<td valign="top" align="left"># suspected parasites in <italic>cV</italic> = <italic>TP</italic> + <italic>FP</italic>
</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>fn</italic>
</td>
<td valign="top" align="left"># false negative objects in <italic>V</italic>
</td>
<td valign="top" align="left">
<italic>tn</italic>
</td>
<td valign="top" align="left"># true negative objects in <italic>V</italic>
</td>
</tr>
<tr>
<td valign="bottom" colspan="4" align="left" style="background-color:#7f7f7f">Terms used in metric definitions</td>
</tr>
<tr>
<td valign="top" align="left">FPR</td>
<td valign="top" align="left">False Positive Rate: <italic>FP</italic> of one patient</td>
<td valign="top" align="left">
<bold>
<italic>F</italic>
</bold>
</td>
<td valign="top" align="left">vector of patients&#x2019; FPRs</td>
</tr>
<tr>
<td valign="top" align="left">
<italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>)</td>
<td valign="top" align="left">standard deviation of <bold>
<italic>F</italic>
</bold>
</td>
<td valign="top" align="left">
<italic>&#xb5;</italic>(<bold>
<italic>F</italic>
</bold>)</td>
<td valign="top" align="left">mean of <bold>
<italic>F</italic>
</bold>
</td>
</tr>
<tr>
<td valign="top" align="left">
<italic>&#x3c3;<sub>L</sub>
</italic>(<bold>
<italic>F</italic>
</bold>)</td>
<td valign="top" align="left">lefthanded standard deviation of <bold>
<italic>F</italic>
</bold>
</td>
<td valign="top" align="left">
<italic>&#x3c3;<sub>R</sub>
</italic>(<bold>
<italic>F</italic>
</bold>)</td>
<td valign="top" align="left">righthanded standard deviation of <bold>
<italic>F</italic>
</bold>
</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>S</italic>
</td>
<td valign="top" align="left">object-level sensitivity of one patient</td>
<td valign="top" align="left">
<bold>
<italic>S</italic>
</bold>
</td>
<td valign="top" align="left">vector of patients&#x2019; <italic>S</italic>&#x2019;s</td>
</tr>
<tr>
<td valign="top" align="left">
<italic>&#x3c3;</italic>(<bold>
<italic>S</italic>
</bold>)</td>
<td valign="top" align="left">standard deviation of <italic>S</italic>
</td>
<td valign="top" align="left">
<italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>)</td>
<td valign="top" align="left">mean of <italic>S</italic>
</td>
</tr>
<tr>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im1">
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>
</td>
<td valign="top" align="left">expected FPR, e.g., <italic>&#xb5;</italic>(<bold>
<italic>F</italic>
</bold>)</td>
<td valign="top" align="left">
<inline-formula>
<mml:math display="inline" id="im2">
<mml:mover accent="true">
<mml:mi>S</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>
</td>
<td valign="top" align="left">expected sensitivity <italic>S</italic>, e.g., <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>)</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>C</italic>
</td>
<td valign="top" align="left">threshold on object classifier scores</td>
<td valign="top" align="left">
<bold>
<italic>T</italic>
</bold>
</td>
<td valign="top" align="left">threshold on # suspected parasites per <italic>cV</italic>
</td>
</tr>
<tr>
<td valign="top" align="center">
<italic>K</italic>
</td>
<td valign="top" align="left">desired patient-level specificity</td>
<td valign="top" align="left">
<italic>&#x3b1;, &#x3b2;</italic>
</td>
<td valign="top" align="left">scalars</td>
</tr>
<tr>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im3">
<mml:mover accent="true">
<mml:mi>P</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>
</td>
<td valign="top" align="left">model&#x2019;s estimate of a true <italic>P</italic>
</td>
<td valign="top" align="left">
<bold>
<italic>L</italic>
</bold>
</td>
<td valign="top" align="left">a model&#x2019;s LoD, in p/<italic>cV</italic> (i.e. p/<italic>&#xb5;L</italic>)</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s3_1">
<label>3.1</label>
<title>Problems with ROCs, AUCs, and precision</title>
<p>ML practitioners choose metrics to evaluate model performance by (i) what is customary, familiar, and convenient; (ii) what has been done by previous authors; (iii) what can generate the &#x201c;state of the art&#x201d; (SOTA) comparisons required for publication in the ML community; and (iv) what is acceptable to ML reviewers. This creates a closed loop which perpetuates the use of certain metrics without regard to their effectiveness. When entrenched metrics do not assess algorithm performance in a clinically relevant way, it blocks progress toward deployable solutions.</p>
<p>Several commonly-used ML metrics, including object-level ROC curves, AUC, object precision, and F1 score, appear frequently in the ML malaria literature. However, in the malaria context these are flawed measures of performance, for reasons given below. They also do not meet standards of evidence per (<xref ref-type="bibr" rid="B48">WHO, 2021a</xref>). They should therefore be avoided when reporting results (though they can be useful intermediate measures for internal algorithm work).</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Object-level ROC curves and AUC</title>
<p>
<italic>Object-level ROC curves</italic>, and the associated Area Under Curve (AUC), are routinely reported by ML research papers involving parasite detection. ROC curves plot sensitivity (fraction of parasites detected) vs one minus specificity (fraction of distractors misclassified as parasites). However, they have three key weaknesses in this context (except perhaps as intermediate measures for internal algorithm work).</p>
<p>First, they do not address the clinical need for patient-centric care. In particular, they ignore the crucial matter of high inter-patient variability of object-level accuracy (this variability is discussed in 3.4 and 3.3).</p>
<p>Second, real samples often have a large imbalance between distractors and positive objects, especially at parasitemias near clinical limit of detection (LoD). A common situation is a model that seeks to diagnoses malaria on thin films by labeling individual red blood cells (RBCs) as infected or not. Since 1 <italic>&#xb5;L</italic> of blood contains roughly 5 million RBCs, a parasitemia of 100 p/<italic>&#xb5;L</italic> gives 50,000 negative objects for each positive object. So a 0.999 AUC can coexist with an average of 50 False Positive objects <italic>per parasite</italic> (a very poor SNR). Since one detected parasite and one False Positive object have equal impact on diagnosis (if using the standard method described in 3.6 of exceeding a threshold count of suspected parasites in the sample), False Positive noise will swamp the diagnostic signal of detected parasites.</p>
<p>In such cases with large class imbalance (say <italic>D</italic>:1), the leftmost <inline-formula>
<mml:math display="inline" id="im1a">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>
<sup>th</sup> vertical sliver of the ROC curve, with <italic>y</italic>-axis rescaled to be full width, reflects a more meaningful (and more sobering) ROC, because this expanded sliver visually weights detected parasite (True Positive) counts and False Positive counts equally, as shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>For a 20:1 distractor-to-parasite ratio, stretching the left vertical sliver gives a more meaningful ROC curve. Diagonal red lines show operating points that give equal numbers of True Positives and False Positives.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmala-02-1250220-g003.tif"/>
</fig>
<p>Third, the object-level ROC curve depends heavily on how distractors are defined because this determines the distractor pool. For example, when using thick films to diagnose malaria, &#x201c;distractor&#x201d; can mean (i) only the most difficult objects that closely resemble parasites; or (ii) any dark blob; or even (iii) every pixel in an image. <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref> shows an example in which including only &#x201c;difficult&#x201d; distractors (top) results in a low AUC, while including additional, mostly &#x201c;easy&#x201d; distractors (bottom) gives a higher AUC with no change in actual performance as measured by the number of False Positives per detected parasite.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>For unchanged algorithm and FPR, ROCs are artificially improved by increasing the number of easy distractor objects. Left: object scores (positives are blue, negatives are red). Right: associated ROC curves.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmala-02-1250220-g004.tif"/>
</fig>
<p>More informative than the object-level ROC is the Free ROC (FROC), which plots object-level sensitivity vs. the number of False Positives per unit volume of blood (see 3.3). FROCs for object level are useful for development work: they clarify where gains can be made by favorably trading off object-level sensitivity for lower False Positive rates. When datasets lack sufficient numbers of patients, FROCs on pooled objects can provide some insight into algorithm performance, with the caveat that they ignore patient-level variability.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Patient-level ROCs</title>
<p>
<italic>Patient-level ROCs</italic> can give a useful sense of algorithm behavior near the clinical performance requirements and are well worth reporting when sufficient data exists to plot them. However, there are two caveats. First, the only salient portion of a patient-level ROC is the region near clinically relevant operating points (e.g., specificity 90%). Second, because sensitivity is parasitemia-dependent (3.4), the ROC is dependent also. Thus, a given algorithm may have much higher AUROC on a population with primarily high parasitemias than on one with lower parasitemias.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Precision</title>
<p>
<italic>Object-level Precision</italic> is the ratio of detected parasites over all detected objects, and often appears as an ML metric. This metric, as used, tends to badly underestimate the effects of parasite-to-distractor imbalances at the low LoDs required for clinical use, as follows.</p>
<p>In ML papers, precision is often calculated on datasets with the clinically unrealistic situation of roughly balanced parasite and distractor counts, either because the numbers of objects have been artificially balanced or because the positive samples had high parasitemias (i.e. many parasites per volume <italic>V</italic>). Since False Positive counts roughly scale with volume <italic>V</italic>, high parasitemia samples yield much more balanced True Positive : False Positive <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> ratios, which tend to give precisions which do not generalize to low parasitemia samples.</p>
<p>For example, a precision of 0.99 calculated on samples with <italic>P</italic> &#x2248; 10,000 p/<italic>&#xb5;L</italic> corresponds to 100 False Positives per <italic>&#xb5;L</italic> (assuming perfect sensitivity). At the required LoD of 100 p/<italic>&#xb5;L</italic>, these same 100 False Positives correspond to 100 parasites, giving precision = 0.5, a much less attractive result.</p>
<p>The related metric <italic>F1</italic>, the harmonic mean of precision and object-level sensitivity (also problematic, as noted in 3.4), is a similarly misleading metric for reporting algorithm results, and in addition has no clinical utility.</p>
<p>The rest of this section (3.2&#x2013;3.11) discusses metrics that better reflect malaria&#x2019;s clinical use case.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Patient level metrics</title>
<p>The importance of assessing algorithm performance <italic>at the patient level</italic> cannot be over-emphasized. The basic unit of clinical care is the patient<xref ref-type="fn" rid="fn2">
<sup>2</sup>
</xref>, so the most relevant metrics are defined at the patient level, not the object level. Performance assessed across pooled objects can be a useful <italic>intermediate</italic> step during ML development, but it is fundamentally unrealistic, because (i) it does not match the clinical task; (ii) it ignores interpatient variability; and (iii) it is dominated by high parasitemia samples. For example, consider four malaria-positive patients, with</p>
<p>Patient 1: 50,000 parasites/<italic>&#xb5;L</italic> (p/<italic>&#xb5;L</italic>).</p>
<p>Patients 2, 3, 4: each 300 p/<italic>&#xb5;L</italic>.</p>
<p>Suppose the algorithm detects almost all parasites in {1}, and misses all parasites in {2,3,4} (a realistic scenario due to interslide variability). Then the object-level sensitivity is 98%, while patient-level sensitivity is 25%.</p>
<p>We have found that two metrics, each defined on a per-patient basis, are particularly useful: <italic>false positive rate</italic> (FPR) and <italic>sensitivity</italic>. Each is calculated separately for each patient, using algorithm accuracy on objects within that patient&#x2019;s sample. These are covered in 3.3 and 3.4, and underpin other metrics related to specificity (3.5), LoD (3.6), and quantitation (3.9).</p>
<p>Interpatient variability (as in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>) poses great difficulty for ML, so it must be factored into algorithm evaluation. It is captured by the standard deviations of FPR and sensitivity (cf. 3.3, 3.4), to the degree that the dataset captures interpatient diversity.</p>
<p>A related issue is interclinic variability. For example, clinics can use different stain variants (e.g., Giemsa, Field, and JSB) which yield different color ranges. Even clinics with nominally identical protocols can differ substantially (see e.g., <xref ref-type="bibr" rid="B9">Das et al., 2022</xref> and a detailed example in <xref ref-type="bibr" rid="B35">Torres et&#xa0;al., 2018</xref>). Besides variations in presentation, different clinics may produce populations of samples with differently distributed FPRs and sensitivities. Implications of this for tuning algorithms are covered in 3.5.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>False positive rate</title>
<p>False Positive Rate (FPR) is the number of distractors mislabeled as parasites per clinically relevant unit of substrate, hereafter <italic>cV</italic>, e.g., 1 <italic>&#xb5;L</italic> of blood (malaria), 10 mL urine (<italic>Schistosoma haematobium</italic>), 1 gram stool (other NTDs), a specified number of cells in a histological sample, etc; but <italic>not</italic> &#x201c;per image tile&#x201d;, which generally has no clinical relevance (though image tiles can often be translated into the microscopy &#x201c;Fields of View&#x201d; used in protocols). Malaria ML papers with some FPR analysis include <xref ref-type="bibr" rid="B19">Linder et&#xa0;al. (2014)</xref>, <xref ref-type="bibr" rid="B23">Mehanian et&#xa0;al. (2017)</xref>, <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al. (2019)</xref>, and <xref ref-type="bibr" rid="B22">Manescu et&#xa0;al. (2020a)</xref>. Crucially, FPR is calculated separately for each patient. We denote the vector of FPRs for the population of patients as <italic>F</italic>.</p>
<p>FPR is <italic>not</italic> object-level specificity, which is a commonly reported but highly flawed measure in this context (see 3.1).</p>
<p>While FPR can be calculated for any sample, FPRs on positive samples may be erroneously boosted by mis- or unannotated parasites. Thus, the population&#x2019;s FPR distribution is best characterized using negative samples only.</p>
<p>Interpatient variability makes the standard deviation of FPR, <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>), a crucial performance measure. The mean FPR <italic>&#xb5;</italic>(<bold>
<italic>F</italic>
</bold>) is less relevant because it can be subtracted out, as shown in 3.6 and 3.9. However, since it tends to scale roughly with <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>), it can give a hint as to the relative magnitude of <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>) (see e.g., <xref ref-type="bibr" rid="B23">Mehanian et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al., 2019</xref>).</p>
<p>In datasets with insufficient numbers of patients, an FPR calculated over pooled objects has some value as a lower bound on <bold>
<italic>F</italic>
</bold>. In particular, it can be compared to the clinical LoD requirement. For example, a pooled-object FPR of 5,000/<italic>&#xb5;L</italic>, vs. a required 100 p/<italic>&#xb5;L</italic> LoD (malaria), is a clear sign that work is still needed. Multiple splits of a set of pooled objects does not simulate <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>), because each split will include the full patient diversity.</p>
<p>
<italic>Aside</italic>: Samples with high FPRs are sometimes criticized as being due to &#x201c;poor sample preparation&#x201d;. However, except for extreme cases this is in the eye of the beholder: human clinicians readily and successfully diagnose &#x201c;dirty&#x201d; samples on which ML algorithms fail. Thus, the need to improve sample prep is to large degree a need to accommodate ML methods&#x2019; struggles with handling highly variable sample presentations. See <xref ref-type="bibr" rid="B9">Das et&#xa0;al. (2022)</xref>, and a detailed example in <xref ref-type="bibr" rid="B35">Torres et&#xa0;al. (2018)</xref>.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Sensitivity</title>
<p>Sensitivity (aka recall) is the fraction of positive items in a set that are correctly labeled: Sensitivity = <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>tp</italic> = true positives, i.e. positive items labeled correctly, and <italic>fn</italic> = false negatives, i.e. positive items labeled as negative or missed. The &#x201c;items&#x201d; can be parasites (object-level) or malaria-positive patients (patient-level).</p>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>Pooled object sensitivity</title>
<p>Sensitivity over a pooled set of parasites from multiple patients has some value as an <italic>intermediate</italic> assessment metric during ML development (e.g., as a loss function for gradient descent training), if it is analyzed carefully to avoid problems such as imbalanced parasitemias distorting the object pool (cf. the example given in 3.2).</p>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Per-patient object sensitivity</title>
<p>A clinically realistic and useful version of object-level sensitivity measures each patient separately:</p>
<p>
<italic>Per patient object-level sensitivity S</italic> is the fraction of parasites in the examined volume <italic>V</italic> of a positive sample that are correctly labeled (e.g., by means of an object score threshold <italic>C</italic>): <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> where <italic>tp</italic> = parasites labeled correctly, and <italic>fn</italic> = parasites labeled as distractors (or missed). There is no constraint on the size of <italic>V</italic> or parasitemia, but sensitivities for patients with few parasites are less reliable (cf. the law of large numbers). Each patient&#x2019;s object-level sensitivity is calculated separately. We denote the vector of sensitivities for the (malaria-positive) population as <bold>
<italic>S</italic>
</bold>. <bold>
<italic>S</italic>
</bold> underpins metrics related to LoD (3.6) and quantitation (3.9).</p>
</sec>
<sec id="s3_4_3">
<label>3.4.3</label>
<title>Patient-level sensitivity</title>
<p>
<italic>Patient-level sensitivity</italic> is sensitivity in the usual clinical sense of the fraction of positive patients correctly diagnosed (not <italic>S</italic>). It is of course a vital metric clinically, but is complex to interpret because it depends on two things:</p>
<list list-type="simple">
<list-item>
<p>(i) The particular parasitemia distribution of the tested set: Patients with low parasitemias (close to the LoD) are harder to identify. In malaria for example (where LoD &#x2248; 100 p/<italic>&#xb5;L</italic>), if all patients have parasitemias <italic>&gt;</italic> 1000 p/<italic>&#xb5;L</italic>, 100% sensitivity is (hopefully) trivial, while if all parasitemias are under 50 p/<italic>&#xb5;L</italic>, very low sensitivity is likely.</p>
</list-item>
<list-item>
<p>(ii) The particular specificity: Sensitivity and specificity are paired and move in opposite directions, as seen in ROC curves.</p>
</list-item>
</list>
<p>Thus, reporting patient-level sensitivity is uninformative and even misleading unless one also reports (i) the parasitemia distribution, and (ii) the associated specificity on negative samples. The WHO competency levels are an important example: These levels crucially assume the parasitemia distribution of the WHO 56 diagnosis slide set, viz 20 negative slides and 20 positive slides with parasitemias between 80 and 200 p/<italic>&#xb5;L</italic> (<xref ref-type="bibr" rid="B45">WHO, 2016c</xref>). WHO competency level ratings do not apply to results on distributions with higher parasitemia samples.</p>
<p>A principled way to maximize patient-level sensitivity is given in 3.7.</p>
</sec>
<sec id="s3_4_4">
<label>3.4.4</label>
<title>Effect of species on sensitivity</title>
<p>Algorithm sensitivity results should be broken down by species&#xa0;as well as by parasitemia, because malaria species has strong impact on patient-level sensitivity. This is due to the unique synchronization and sequestration behaviors of <italic>P. falciparum</italic> (<xref ref-type="bibr" rid="B14">Garnham, 1966</xref>):</p>
<list list-type="simple">
<list-item>
<p>(i) In <italic>falciparum</italic> the large, distinctive late stage forms sequester out of the peripheral blood, leaving only the smaller ring forms that are harder to detect and disambiguate from distractor objects (especially in thick films). As a result, in our experience non-<italic>falciparum</italic> infections (i.e. <italic>vivax, ovale, malariae, knowlesi</italic>) are much easier to detect in blood films (given equal parasitemias), which allows an algorithm to have lower LoD and higher patient-level sensitivity (<xref ref-type="bibr" rid="B35">Torres et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B15">Horning et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B9">Das et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B33">Rees-Channer et&#xa0;al., 2023</xref>).</p>
</list-item>
<list-item>
<p>(ii) <italic>falciparum</italic> parasites tend to synchronize in peripheral blood, with the presenting parasites forming a narrow age distribution. This strongly impacts diagnostic methods that target the biomarker hemozoin: non-<italic>falciparum</italic> samples can be very sensitively detected due to the reliable presence of late-stage, high hemozoin parasites (<xref ref-type="bibr" rid="B2">Arndt et&#xa0;al., 2021</xref>), but even high parasitemia <italic>falciparum</italic> samples can lack detectable hemozoin due to synchronized populations of early stage ring forms (<xref ref-type="bibr" rid="B16">Jamjoom, 1988</xref>; <xref ref-type="bibr" rid="B32">Rebelo et&#xa0;al., 2011</xref>; <xref ref-type="bibr" rid="B10">Delahunt et&#xa0;al., 2014b</xref>), resulting in drastically different sensitivities by species. (Hemozoin appears to be a sensitive biomarker for <italic>falciparum</italic> in cultured blood because synchronization is absent.)</p>
</list-item>
</list>
<p>This is a high-stakes issue because <italic>falciparum</italic> is much more often fatal than non-<italic>falciparum</italic> species.</p>
</sec>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Specificity</title>
<p>Specificity is the fraction of negative items (distractor objects or patients) that are correctly diagnosed as negative:</p>
<p>Specificity = <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>tn</italic> = true negatives (negative items in <italic>V</italic> labeled correctly), and <italic>fp</italic> = false positives (negative items labeled incorrectly).</p>
<sec id="s3_5_1">
<label>3.5.1</label>
<title>Object-level specificity</title>
<p>
<italic>Object-level specificity</italic>, even if calculated for each patient separately, has little usefulness and can be highly deceptive (see 3.1).</p>
</sec>
<sec id="s3_5_2">
<label>3.5.2</label>
<title>Patient-level specificity</title>
<p>
<italic>Patient-level specificity</italic>, i.e. in the usual clinical sense, is highly salient. Clinical goals of high specificity include not overwhelming the health care system, avoiding excess treatments, and preventing misattribution. Thus, clinical use-cases generally require a high specificity (e.g., 90% for malaria diagnosis (<xref ref-type="bibr" rid="B45">WHO 2016c</xref>), 97.5% for schistosomiasis (<xref ref-type="bibr" rid="B49">WHO, 2021b</xref>)).</p>
<p>Specificity is closely tied to FPR (3.3) and can be readily tuned for an algorithm that labels objects: Suppose that objects have been detected then labeled by some method (e.g., a threshold <italic>C</italic> on object scores), that <bold>
<italic>F</italic>
</bold> (from 3.3) is gaussian, and that patient diagnosis is determined by a threshold <italic>T</italic> on the number of positively-labeled objects per <italic>cV</italic> (i.e. a standard &#x201c;detect, classify, count, then threshold&#x201d; approach). To attain a target specificity <italic>K</italic>, one can set</p>
<disp-formula id="eq1">
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo>&#xb5;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>&#x3b1;</italic> is found via the (one-sided) error function and <italic>K</italic>. Alternate formulations for the case of nongaussian <bold>
<italic>F</italic>
</bold> are given in 3.8.</p>
<p>Negative samples are easier to obtain and trivial to annotate (assuming accurate patient-level ground truth), and specificity depends only on negative samples. So <italic>T</italic> can ideally be tuned on a separate, dedicated validation set of negatives that capture a sufficient range of FPRs (both &#x201c;dirty&#x201d; and &#x201c;clean&#x201d; samples).</p>
<p>Note that different clinics can have widely different FPR distributions <bold>
<italic>F</italic>
</bold>. Because <italic>&#x3c3;</italic>(<italic>
<bold>F</bold>
</italic>) determines both specificity (<xref ref-type="disp-formula" rid="eq1">Equation 1</xref>) and LoD (3.6), different clinics may require different hyperparameters to hit the target patient specificity <italic>K</italic>, leading to different LoDs. Thus, tuning an algorithm for deployment may involve multiple validation sets of negatives (by clinic), with clinic-dependent tradeoffs between specificity and higher LoD.</p>
</sec>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Limit of detection (LoD)</title>
<p>Here, LoD roughly means the parasitemia at which the algorithm can consistently (e.g., 95% of cases) distinguish positive and negative cases. Based on the WHO evaluation criteria (<xref ref-type="bibr" rid="B45">WHO, 2016c</xref>), the required LoD for malaria microscopy is roughly 100 p/<italic>&#xb5;L</italic>, i.e. 1 parasite per 50,000 red blood cells (RBCs) or 80 white blood cells (WBCs). However, expert microscopists routinely achieve LoDs &#x2248;50 p/<italic>&#xb5;L</italic> (e.g., Vilela, pers. comm.; Bell, pers. comm.), and the lower LoD is of course clinically desirable. For helminths, LoD is implicitly 1 egg (per 10 mL urine or 1 gram stool) (<xref ref-type="bibr" rid="B39">WHO, 2002</xref>, <xref ref-type="bibr" rid="B49">2021b</xref>). Standard <italic>Loa loa</italic> diagnosis by blood microscopy has an effective LoD <italic>&gt;</italic> 200 mf/<italic>mL</italic> when <italic>V</italic> = 10<italic>&#xb5;L</italic> (<xref ref-type="bibr" rid="B25">Mischlinger et&#xa0;al., 2021</xref>).</p>
<p>LoD can be directly probed using holdout sets of low parasitemia positive samples. These are not as useful for training anyway, as they supply few parasite objects. However, this is impractical because it&#x2019;s hard to acquire enough field-prepared positive blood films near the LoD (a work-around is to ablate parasites in the image set of a sample with parasitemia above the LoD, to lower its visible parasitemia).</p>
<p>We can calculate a useful estimate of LoD from <bold>
<italic>F</italic>
</bold> and <bold>
<italic>S</italic>
</bold> as follows:</p>
<p>Denote the putative LoD as <italic>L</italic> parasites per <italic>cV</italic>, and suppose that a patient is diagnosed as &#x201c;positive&#x201d; when <italic>N</italic> &#x2265; <italic>T</italic>, where <italic>N</italic> is the number of positively-labeled objects per <italic>cV.</italic> Note that <italic>N</italic> = <italic>TP</italic> + <italic>FP</italic> in positive patients, and <italic>N</italic> = <italic>FP</italic> in negative patients, where <italic>TP</italic> and <italic>FP</italic> denote counts per <italic>cV</italic>, so <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mi>V</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> where <italic>tp</italic> is the number of parasites correctly labeled in <italic>V</italic> (similarly <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>p</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mi>V</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>).</p>
<p>&#x2022; Make <italic>T</italic> high enough to ensure to enforce 95% specificity on negative samples as described in (<xref ref-type="bibr" rid="B23">Mehanian et&#xa0;al., 2017</xref>) by setting <italic>&#x3b1;</italic> to 1.65 std devs in <xref ref-type="disp-formula" rid="eq1">Equation 1</xref>:</p>
<disp-formula id="eq2">
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo>&#xb5;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mn>1.65</mml:mn>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>&#x2022; Then for positive samples the worst case is a very &#x201c;clean&#x201d; sample with low FPR, such as the 5<sup>th</sup> percentile of samples with <italic>FP</italic> = <italic>&#xb5;</italic>(<bold>
<italic>F</italic>
</bold>) &#x2212; 1.65<italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>). In this case we must depend mostly on detected parasites to ensure <italic>N</italic> &#x2265; <italic>T</italic> for a positive diagnosis. Suppose for ease that the sample has average sensitivity = <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>). Then a sample at LoD has <italic>TP</italic> = <italic>L&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>).</p>
<p>&#x2022; To diagnose this positive sample correctly (but just barely, i.e. <italic>N</italic> = <italic>T</italic>), we need</p>
<disp-formula>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#xb5;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mo>&#xb5;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1.65</mml:mn>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mo>=</mml:mo>
<mml:mi>T</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo>&#xb5;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mn>1.65</mml:mn>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:mo>&#x21d2;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#xb5;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mn>3.3</mml:mn>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>So the estimated LoD (<italic>L</italic> per <italic>cV)</italic> has</p>
<disp-formula id="eq3">
<label>(3)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>3.3</mml:mn>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>&#x2022; Optionally, +1 can be added to the numerator (i.e. require <italic>N</italic> = <italic>T</italic> +1) to prevent unpredictable behavior should both <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>) and <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>) approach 0:</p>
<disp-formula id="eq4">
<label>(4)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>3.3</mml:mn>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In our algorithm development, we have found this estimate to be a good (slightly optimistic) proxy for actual LoD when assessing algorithms during development. In particular, it consistently tracked diagnostic accuracy on holdout sets of low parasitemia samples, i.e. lower estimated LoDs mapped to higher accuracy (sensitivity and specificity at the patient level) in holdout sets and field trials. It has the practical advantage that low parasitemia samples are unnecessary, because the vector <bold>
<italic>S</italic>
</bold> can be well characterized by high parasitemia samples. It also allows useful comparison of algorithms, as it directly addresses a key clinical requirement and is anchored to the relevant unit <italic>cV.</italic>
</p>
<p>A more nuanced (and pessimistic) proxy could account for <italic>&#x3c3;</italic>(<bold>S</bold>) by having a denominator = <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>)&#x2212;<italic>&#x3b2; &#x3c3;</italic>(<bold>
<italic>S</italic>
</bold>) for some <italic>&#x3b2;</italic>.</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Choosing operating points</title>
<p>Given a trained algorithm that uses the two hyperparameters <italic>C</italic> and <italic>T</italic>, {<italic>C, T</italic>} can be optimized in a principled way to maximize patient-level sensitivity, subject to the constraint of a fixed target specificity <italic>K</italic>:</p>
<list list-type="bullet">
<list-item>
<p>Set aside a validation set of negative samples. If there are sufficient positive samples to spare, optionally set these aside also.</p>
</list-item>
<list-item>
<p>For each <italic>C</italic>:</p>
<list list-type="simple">
<list-item>
<p>- Calculate <bold>
<italic>F</italic>
</bold> over the validation negatives, and <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>) over the validation positives if available, or (less ideal but workable) over the training set positives.</p>
</list-item>
<list-item>
<p>- Determine <italic>T</italic> = <italic>T</italic>(<italic>C, K, F</italic>) which hits the target specificity <italic>K</italic> on the validation negatives, as in 3.5.</p>
</list-item>
<list-item>
<p>- Estimate LoD as in 3.6.</p>
</list-item>
</list>
</list-item>
<list-item>
<p>Select the <italic>C</italic> with the lowest LoD.</p>
</list-item>
<list-item>
<p>Use this {<italic>C,T</italic>} pair as algorithm hyperparameters to process test sets, and report patient-level specificity and sensitivity.</p>
</list-item>
</list>
</sec>
<sec id="s3_8">
<label>3.8</label>
<title>Modified LoD and operating point formulas</title>
<p>The methods for setting <italic>T</italic> (<xref ref-type="disp-formula" rid="eq2">Equation 2</xref>) and for estimating LoD (<xref ref-type="disp-formula" rid="eq3">Equations 3</xref>, <xref ref-type="disp-formula" rid="eq4">4</xref>) both assume that the FPR vector <bold>
<italic>F</italic>
</bold> is gaussian. In our experience this is often not the case. Rather, the FPR distribution may be asymmetrical, with mostly low-FPR samples and a few high-FPR samples. This can be handled by modifying the methods in 3.5 and 3.6 as follows:</p>
<list list-type="bullet">
<list-item>
<p>For <italic>&#xb5;</italic>(<bold>
<italic>F</italic>
</bold>), use the median of <bold>
<italic>F</italic>
</bold> instead of the mean of <bold>
<italic>F</italic>
</bold>. Similarly, if the vector <bold>
<italic>S</italic>
</bold> is non-gaussian, the median can be used instead of the mean for <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>).</p>
</list-item>
<list-item>
<p>For <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>), use one-sided std devs, which can be calculated by keeping only the points to the right (or left) of the median and reflecting them across the median as centerpoint to create a symmetric distribution. This gives, for the FPR distribution above, a large right std dev <italic>&#x3c3;<sub>R</sub>
</italic>(<bold>
<italic>F</italic>
</bold>) and a small left std dev <italic>&#x3c3;<sub>L</sub>
</italic>(<bold>
<italic>F</italic>
</bold>).</p>
</list-item>
<list-item>
<p>Then the new versions of <xref ref-type="disp-formula" rid="eq1">Equations 1</xref>, <xref ref-type="disp-formula" rid="eq3">3</xref> are</p>
</list-item>
</list>
<disp-formula id="eq5">
<mml:math display="block" id="M8">
<mml:mi>T</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mi>a</mml:mi>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>R</mml:mi>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:math>
</disp-formula>
<disp-formula id="eq6">
<mml:math display="block" id="M9">
<mml:mi>L</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1.65</mml:mn>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>R</mml:mi>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:math>
</disp-formula>
<p>Two other methods of calculating <italic>T</italic> from <bold>
<italic>F</italic>
</bold> may be useful:</p>
<list list-type="order">
<list-item>
<p>Set <italic>T</italic> based on the <italic>K</italic>
<sup>th</sup> percentile of <bold>
<italic>F</italic>
</bold>.</p>
</list-item>
<list-item>
<p>Manually choose <italic>T</italic> based on a scatterplot of the <italic>FP</italic> counts in the validation negative samples.</p>
</list-item>
</list>
<p>For both these methods, the detected objects are assumed to be already classified. If a threshold <italic>C</italic> on object scores was used, then first <italic>T</italic> must be calculated for each <italic>C</italic>, before choosing the best {<italic>C,T</italic>} pair as in 3.7.</p>
<p>The manual method of choosing {<italic>C,T</italic>} takes time, but it can yield the best results in a field deployment because it is most closely tailored to the empirical FPR distribution.</p>
</sec>
<sec id="s3_9">
<label>3.9</label>
<title>Quantitation</title>
<p>Quantitation sometimes has clinical importance. For example, accurate quantitation is needed to monitor for drug-resistant malaria strains by calculating clearance curves (<xref ref-type="bibr" rid="B38">White, 2011</xref>; <xref ref-type="bibr" rid="B3">Ashley et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B46">WHO, 2016d</xref>). For helminths, quantitation targets are typically rough only (e.g., low, medium, high) (<xref ref-type="bibr" rid="B39">WHO, 2002</xref>). For <italic>Loa loa</italic>, a remarkable drug reaction necessitates accurate quantitation at certain high parasitemias only (&#x2248; 20k to 30k worms/<italic>mL</italic>) (<xref ref-type="bibr" rid="B13">Gardon et&#xa0;al., 1997</xref>; <xref ref-type="bibr" rid="B8">D&#x2019;Ambrosio et&#xa0;al., 2015</xref>).</p>
<sec id="s3_9_1">
<label>3.9.1</label>
<title>Measuring quantitation accuracy</title>
<p>Quantitation accuracy should be reported at the patient level due to high interpatient variability. For plotting quantitation error per patient, Bland-Altman plots are preferable because relative quantitation error is generally most important (<xref ref-type="bibr" rid="B44">WHO, 2016b</xref>).</p>
<p>Reporting the <italic>R</italic>
<sup>2</sup> value of a linear fit of estimated vs. true (i.e. <inline-formula>
<mml:math display="inline" id="im10">
<mml:mover accent="true">
<mml:mi>P</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> vs. <italic>P</italic>) is unsuitable when parasitemias range over orders of magnitude (common in malaria and NTDs), because effects of the <italic>L</italic>
<sub>2</sub> norm almost guarantee that high parasitemia samples will lay on the fitted line while high relative errors on low parasitemia samples will be downplayed, giving an illusion of strong fit. Fitting the log(<italic>P</italic>) rather than <italic>P</italic> values helps to reduce this illusion.</p>
</sec>
<sec id="s3_9_2">
<label>3.9.2</label>
<title>Estimating parasitemia</title>
<p>As described in <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al. (2019)</xref>, we can estimate the parasitemia <inline-formula>
<mml:math display="inline" id="im11">
<mml:mover accent="true">
<mml:mi>P</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> for a given patient by</p>
<disp-formula id="eq7">
<label>(5)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>P</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mi>V</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mover accent="true">
<mml:mi>S</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<italic>n</italic> = number of alleged parasites found in <italic>V</italic>,</p>
<p>
<inline-formula>
<mml:math display="inline" id="im12">
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> = expected FPR (e.g., <italic>&#xb5;</italic>(<bold>
<italic>F</italic>
</bold>)),</p>
<p>
<inline-formula>
<mml:math display="inline" id="im13">
<mml:mover accent="true">
<mml:mi>S</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> = expected sensitivity (e.g., <italic>&#xb5;</italic>(<bold>
<italic>S</italic>
</bold>)),</p>
<p>
<italic>cV</italic> = clinically relevant volume of substrate,</p>
<p>
<italic>V</italic> = estimate of the volume examined.</p>
<p>Three types of error affect <xref ref-type="disp-formula" rid="eq5">Equation 5</xref>: irreducible Poisson, estimates of examined volume, and counts of alleged parasites.</p>
<sec id="s3_9_2_1">
<label>3.9.2.1</label>
<title>Irreducible Poisson error</title>
<p>This is discussed below in 3.10.</p>
</sec>
<sec id="s3_9_2_2">
<label>3.9.2.2</label>
<title>Examined volume error</title>
<p>Error in estimating <italic>V</italic> impacts quantitation accuracy via the <inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mi>V</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> term of <xref ref-type="disp-formula" rid="eq5">Equation 5</xref>. For example, thick film blood volume <italic>V</italic> is typically estimated by counting WBCs (<xref ref-type="bibr" rid="B43">WHO, 2016a</xref>). Any error in the WBC count causes proportional quantitation error. This error type can be compartmentalized, for performance evaluation purposes only, as follows:</p>
<list list-type="bullet">
<list-item>
<p>Manually count WBCs on a test set to ensure oracle <italic>V</italic> estimates and use these counts to calculate <italic>V</italic>, ensuring zero error of this type.</p>
</list-item>
<list-item>
<p>Separately report the patient-level error statistics of the WBC counter.</p>
</list-item>
</list>
</sec>
<sec id="s3_9_2_3">
<label>3.9.2.3</label>
<title>Parasite counting errors</title>
<p>Errors in parasite count stem from patient-level variations in sensitivity and FPR, as follows:</p>
<list list-type="bullet">
<list-item>
<p>The number of alleged parasites per <italic>cV</italic> in the sample is <inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>f</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mi>V</mml:mi>
</mml:mfrac>
<mml:mo>=</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
</list-item>
<list-item>
<p>Let <italic>P</italic> be the true parasite count per <italic>cV</italic>. Then <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>S</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the expected number of correctly labeled true parasites per <italic>cV</italic>, and the difference between <italic>TP</italic> and <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>S</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is due to deviation of the sample&#x2019;s sensitivity from the expected <inline-formula>
<mml:math display="inline" id="im18">
<mml:mover accent="true">
<mml:mi>S</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>.</p>
</list-item>
<list-item>
<p>Similarly, the difference between <italic>FP</italic> and <inline-formula>
<mml:math display="inline" id="im19">
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>(the expected FPR) is due to the deviation of this sample&#x2019;s FPR from expected. <italic>&#x3c3;</italic>(<bold>
<italic>S</italic>
</bold>) and <italic>&#x3c3;</italic>(<bold>
<italic>F</italic>
</bold>) quantify these deviations over the population.</p>
</list-item>
<list-item>
<p>A figure of merit to assess parasite counting error, derived and discussed in <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al. (2019)</xref>, is thus</p>
</list-item>
</list>
<disp-formula id="eq8">
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>+</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>While the FPR term is usually hardest to control, it also shrinks as 1<italic>/P</italic>, so for large <italic>P</italic> the sensitivity term dominates. This effect can be leveraged by using different operating points according to whether initial estimated parasitemia is low or high, to favor FPR or sensitivity. In particular, different operating points are indicated for diagnosis (since the hard cases have low parasitemia, where FPR dominates) and for quantitation (high parasitemias, where sensitivity dominates).</p>
<p>We note that parasitemia estimates based on manual microscopy are also subject to these three error types. This complicates assessment of a model&#x2019;s quantitation accuracy against microscopy ground truth.</p>
</sec>
</sec>
</sec>
<sec id="s3_10">
<label>3.10</label>
<title>Effect of poisson statistics</title>
<p>Poisson statistics for rare events give variation in the actual number of parasites in a particular sample with volume <italic>V</italic>, given a fixed true parasitemia <italic>P</italic> over the whole sample. The variation is most visible at low parasitemias, e.g., at 100 p/<italic>&#xb5;L</italic>, where each RBC has a 1/50,000 chance of containing a parasite in thin film, or each WBC has a 1/80 chance of corresponding to a nearby parasite in thick film.</p>
<p>This variability has two main impacts:</p>
<list list-type="simple">
<list-item>
<p>(i) For diagnosis, a low LoD requires that a large volume <italic>V</italic> be examined to ensure that at least a couple true parasites are present at all. Otherwise, for a statistically predictable subset of positive patients the examined volume will contain 0 parasites, reducing patient-level sensitivity from the start. For malaria, to attain LoD of 100 p/<italic>&#xb5;L</italic> requires that at least &#x2248;0.05 <italic>&#xb5;L</italic> of blood should be examined, equivalent to 400 WBCs in thick film, or 250,000 RBCs in thin film (see <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>). The difficulty of finding this many acceptable RBCs, and the long processing time required, are two reasons why thin films are not standard protocol for manual field diagnosis; however, see progress by (<xref ref-type="bibr" rid="B26">Noul, 2023</xref>; <xref ref-type="bibr" rid="B27">Nowak et&#xa0;al., 2023</xref>). Indeed, the low sensitivity (relative to PCR) of manual microscopy at parasitemias &lt; 50 p/<italic>&#xb5;L</italic> (see, e.g., <xref ref-type="bibr" rid="B35">Torres et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B9">Das et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B33">Rees-Channer et&#xa0;al., 2023</xref>) is due largely to Poisson variability: expert microscopists can certainly recognize even a single parasite, but (following protocols) they do not examine sufficient blood to ensure that such a parasite is present when parasitemias are very low. The systems used in our group&#x2019;s studies examine &gt;0.1 <italic>&#xb5;L</italic> of thick film (<italic>&gt;</italic>800 WBCs, the red curve in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>).</p>
</list-item>
<list-item>
<p>(ii) For quantitation, a sufficiently high volume <italic>V</italic> (depending on <italic>P</italic>) must be examined to control irreducible error. For more detail and plots see S.I. of <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al. (2019)</xref>. Poisson error affects manual microscopy also, and when possible is mitigated by combining multiple manual reads (<xref ref-type="bibr" rid="B51">WWARN, 2023</xref>).</p>
</list-item>
</list>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Poisson distributions of parasite counts for various examined volumes <italic>V</italic> , assuming 100 p/<italic>&#xb5;L</italic>. Low parasitemia samples may present as negative (i.e. zero parasites) if <italic>V</italic> is too small.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmala-02-1250220-g005.tif"/>
</fig>
<p>In both cases, automated systems hold a strong advantage because they can scan higher volumes than human technicians, who often by necessity work in a high Poisson error regime (S.I. of <xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al., 2019</xref>). Manual microscopy protocols average multiple readers&#x2019; estimates (when available) to reduce quantitation error (<xref ref-type="bibr" rid="B44">WHO, 2016b</xref>; <xref ref-type="bibr" rid="B51">WWARN, 2023</xref>).</p>
<p>When reporting results on datasets of small size, authors should understand how Poisson variability limits their estimates of algorithm performance.</p>
</sec>
<sec id="s3_11">
<label>3.11</label>
<title>Malaria species identification metrics</title>
<p>Identification of malaria species is one of the three tasks assessed by the WHO 56 evaluation system (<xref ref-type="bibr" rid="B45">WHO, 2016c</xref>). Correct species ID matters clinically because (i) <italic>falciparum</italic> infections are much more likely to be fatal; and (ii) treatment plans differ by species (<xref ref-type="bibr" rid="B7">CDC, 2023b</xref>), since (for example) the hypnozoites of <italic>vivax</italic> and <italic>ovale</italic> species require special care.</p>
<p>Because not all species ID errors are equal from a clinical perspective, reported results should preferably include a confusion matrix as in (<xref ref-type="bibr" rid="B11">Delahunt et&#xa0;al., 2019</xref>).</p>
<p>Aside: In our experience, it is relatively straightforward to distinguish <italic>falciparum</italic> vs. non-<italic>falciparum</italic> on thick film alone (<xref ref-type="bibr" rid="B35">Torres et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B36">Vongpromek et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B9">Das et&#xa0;al., 2022</xref>), also (<xref ref-type="bibr" rid="B17">Kassim et&#xa0;al., 2021</xref>), and even mixed species infections that include <italic>falciparum</italic> can often be identified on thick film by comparing the ring stage and late stage parasite counts (<xref ref-type="bibr" rid="B15">Horning et&#xa0;al., 2021</xref>). However, thin films are still typically needed to distinguish between the various non-<italic>falciparum</italic> species, unless the clinical use case allows geographical priors to be leveraged. A method to distinguish non-<italic>falciparum</italic> species on thick film would yield clinical benefit by eliminating the need for thin films, due to (i) the ease of thick-only workflows (Carter, pers. comm.; Proux, pers. comm.), and (ii) thin film problems with quality (Long, pers. comm.) and difficulty of species ID at low parasitemias (Lilley, pers. comm.).</p>
<p>Staging parasites (as ring, later trophozoite, schizont, or gametocyte) is not part of the WHO evaluation, and is not generally useful clinically, except as used during species identification or when quantitating asexual forms in non-<italic>falciparum</italic> species. In <italic>falciparum</italic> (the main target of quantitation, for drug resistance studies) the difference between ring and gametocyte is glaring.</p>
<p>The methods described here were developed to address the exigencies of field-prepared blood films. They apply equally well to analysis of particular field isolates of any <italic>Plasmodium</italic> species, since the core issues (inter-sample variability, importance of FP objects, etc) apply to field isolates. The caveats connected to analysis of <italic>in vitro</italic> cultures (3.4.4) apply to field isolates as well.</p>
</sec>
</sec>
<sec id="s4" sec-type="discussion">
<label>4</label>
<title>Discussion</title>
<p>Malaria and NTDs are amenable though difficult targets for ML methods, and successful development of translatable ML solutions would yield tremendous health care benefits for currently underserved populations by enabling automated malaria diagnosis to augment the throughput capacities of hard-pressed clinicians.</p>
<p>Unfortunately communal ML progress, in which researchers build on each others&#x2019; work to reach a performance goal, is handicapped for malaria by lack of attention to clinical needs, and by widespread use of ill-suited evaluation metrics. As a result, the synergistic power of the ML community is not being applied with full force to this important task, since many papers present methods that cannot be usefully extended.</p>
<p>Individual ML research teams can radically improve the situation by grounding their ML work in an understanding of the use case, and by tailoring metrics to the clinical needs. We have described such metrics here: variation in FPR, per-patient sensitivity, LoD, patient-level sensitivity and specificity, and a figure of merit for quantitation. We have also listed some essential technical background reading from the WHO and others.</p>
<p>Peer reviewers play a special role in determining the success or failure of the communal ML effort: (i) Reviewers can assess algorithms and performance results according to whether they incorporate the requirements of the clinical use case; (ii) When authors present new metrics, well-grounded in the use-case, this can be more valuable than a comparison based on customary but inferior metrics. By recognizing when this is the case, reviewers can disrupt the cycle that perpetuates a counterproductive status quo.</p>
<p>With attention to the clinical use case and deliberate choice of metrics, the ML community can better equip itself to successfully address automated malaria and NTD diagnosis, and thus deliver concrete benefit to the populations suffering the dire effects of these illnesses.</p>
</sec>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>CD, NG, and MH researched methods. CD wrote the manuscript. NG and MH reviewed the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. Funding provided by Global Health Labs, Inc. (<ext-link ext-link-type="uri" xlink:href="http://www.ghlabs.org">www.ghlabs.org</ext-link>).</p>
</sec>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>Defined here as detecting, quantitating, and identifying the species of <italic>Plasmodium</italic> parasites in peripheral blood (<xref ref-type="bibr" rid="B5">CDC, 2023</xref>).</p>
</fn>
<fn id="fn2">
<label>2</label>
<p>We set aside population-level diagnostics such as for Vitamin A deficiency <xref ref-type="bibr" rid="B42">WHO (2011)</xref>.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Armstrong</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Coulibaly</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Essien-Baidoo</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Bogoch</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Ephraim</surname> <given-names>R. K. D.</given-names>
</name>
<name>
<surname>Fletcher</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Point-of-care sample preparation and automated quantitative detection of <italic>Schistosoma haematobium</italic> using mobile phone microscopy</article-title>. <source>Am. J. Trop. Med. Hyg.</source>  <volume>106</volume> (<issue>5</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.4269/ajtmh.21-1071</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arndt</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Koleala</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Orban</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Ibam</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Kezsmarki</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Karl</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Magneto-optical diagnosis of symptomatic malaria in Papua New Guinea</article-title>. <source>Nat. Commun.</source>  <volume>12</volume> (<issue>1</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41467-021-21110-w</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ashley</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Dhorda</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Fairhurst</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Amaratunga</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Lim</surname> <given-names>P.</given-names>
</name>
<name>
<surname>White</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Spread of artemisinin resistance in <italic>Plasmodium falciparum</italic> malaria</article-title>. <source>N. Engl. J. Med</source> <volume>371</volume> (<issue>5</issue>), <fpage>411</fpage>&#x2013;<lpage>423</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1056/NEJMoa1314981</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>BMGF</collab>
</person-group> (<year>2023</year>) <source>Neglected tropical diseases</source>. Available online at: <uri xlink:href="https://www.gatesfoundation.org/our-work/programs/global-health/neglected-tropical-diseases">https://www.gatesfoundation.org/our-work/programs/global-health/neglected-tropical-diseases</uri>.</citation>
</ref>
<ref id="B5">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>CDC</collab>
</person-group> (<year>2023</year>) <source>Malaria website</source>. Available online at: <uri xlink:href="https://www.cdc.gov/malaria/diagnosistreatment/diagnosis.html">https://www.cdc.gov/malaria/diagnosistreatment/diagnosis.html</uri>.</citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>CDC</collab>
</person-group> (<year>2023</year>a). <source>Malaria website</source> (<publisher-loc>USA</publisher-loc>: <publisher-name>Centers for Disease Control</publisher-name>). Available at: <uri xlink:href="https://www.cdc.gov/malaria/index.html">https://www.cdc.gov/malaria/index.html</uri>.</citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>CDC</collab>
</person-group> (<year>2023</year>b). <source>Algorithm for diagnosis and treatment of malaria in the United States</source> (<publisher-name>Centers for Disease Control</publisher-name>). Available at: <uri xlink:href="https://www.cdc.gov/malaria/resources/pdf/Malaria_Managment_Algorithm_202208.pdf">https://www.cdc.gov/malaria/resources/pdf/Malaria_Managment_Algorithm_202208.pdf</uri>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>D&#x2019;Ambrosio</surname> <given-names>M. V.</given-names>
</name>
<name>
<surname>Bakalar</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Bennure</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Reber</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Sakndarajah</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Fletcher</surname> <given-names>D. A.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>Point-of-care quantification of blood-borne filarial parasites with a mobile phone microscope</article-title>. <source>Sci. Trans. Med</source> <volume>7</volume> (<issue>286</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1126/scitranslmed.aaa3480</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Das</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Vongpromek</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Assawariyathipat</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Srinamon</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Kennon</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Dhorda</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Field evaluation of the diagnostic performance of EasyScan GO: a digital malaria microscopy device based on machine-learning</article-title>. <source>Malaria J.</source> <volume>21</volume>, <fpage>122</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12936-022-04146-1</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Delahunt</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Horning</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Wilson</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Proctor</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Hegg</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>b). <article-title>Limitations of haemozoin-based diagnosis of Plasmodium falciparum using dark-field microscopy</article-title>. <source>Malaria J.</source>, <fpage>393</fpage>&#x2013;<lpage>399</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1475-2875-13-147</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Delahunt</surname> <given-names>C. B.</given-names>
</name>
<name>
<surname>Jaiswal</surname> <given-names>M. S.</given-names>
</name>
<name>
<surname>Horning</surname> <given-names>M. P.</given-names>
</name>
<name>
<surname>Janko</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>C. M.</given-names>
</name>
<name>
<surname>Mehanian</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Kulhare</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Fully-automated patient-level malaria assessment on field-prepared thin blood film microscopy images</article-title>. <source>arXiv.</source> doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1908.01901</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Delahunt</surname> <given-names>C. B.</given-names>
</name>
<name>
<surname>Mehanian</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>McGuire</surname> <given-names>S. K.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Wilson</surname> <given-names>B. K.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>a). <article-title>Automated microscopy and machine learning for expert-level malaria field diagnosis</article-title>. <source>IEEE GHTC Proc</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/GHTC.2015.7344002</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gardon</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Gardon-Wendel</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Demanga-Ngangue</surname>
</name>
<name>
<surname>Demanga-Ngangue Kamgno</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chippaux</surname> <given-names>J.P.</given-names>
</name>
<name>
<surname>Boussinesq</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Serious reactions after mass treatment of onchocerciasis with ivermectin in an area endemic for loa loa infection</article-title>. <source>Lancet</source>  <volume>350</volume> (<issue>9070</issue>), <fpage>18</fpage>&#x2013;<lpage>22</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/S0140-6736(96)11094-1</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Garnham</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>1966</year>). <source>Malaria parasites and other haemosporidia</source> (<publisher-loc>Oxford, UK</publisher-loc>: <publisher-name>Blackwell Scientific Publications Ltd</publisher-name>).</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Horning</surname> <given-names>M.P.</given-names>
</name>
<name>
<surname>Delahunt</surname> <given-names>C.B.</given-names>
</name>
<name>
<surname>Bachman</surname> <given-names>C.M.</given-names>
</name>
<name>
<surname>Luchavez</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Luna</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Mehanian</surname> <given-names>C</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Performance of a fully-automated system on a WHO malaria microscopy evaluation slide set</article-title>. <source>Malaria J</source> <volume>20</volume>, <fpage>110</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12936-021-03631-3</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jamjoom</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>Patterns of pigment accumulation in <italic>Plasmodium falciparum</italic> trophozoites in peripheral blood samples</article-title>. <source>Am. J. Trop. Med. Hyg</source> <volume>39</volume> (<issue>1</issue>), <fpage>21</fpage>&#x2013;<lpage>25</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.4269/ajtmh.1988.39.21</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kassim</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Maude</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Jaeger</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Diagnosing malaria patients with <italic>Plasmodium falciparum</italic> and <italic>vivax</italic> using deep learning for thick smears</article-title>. <source>Diagnostics</source> <volume>11</volume> (<issue>11</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.3390/diagnostics11111994</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Koller</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Bengio</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>) <source>A fireside chat with Daphne Koller at ICLR</source>. Available online at: <uri xlink:href="https://www.youtube.com/watch?v=N4mdV1CIpvI">https://www.youtube.com/watch?v=N4mdV1CIpvI</uri>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Linder</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Turkki</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Walliander</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Martensson</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Diwan</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Lundin</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>A malaria diagnostic tool based on computer vision screening and visualization of <italic>Plasmodium falciparum</italic> candidate areas in digitized blood smears</article-title>. <source>PloS One</source> <volume>9</volume> (<issue>8</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0104855</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maier-Hein</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Reinke</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Godau</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Tizabi</surname> <given-names>M. D.</given-names>
</name>
<name>
<surname>Buettner</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Jager</surname> <given-names>P. F.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Metrics reloaded: Pitfalls and recommendations for image analysis validation</article-title>. <source>arXiv</source>. <uri xlink:href="https://arxiv.org/abs/2206.01653">https://arxiv.org/abs/2206.01653</uri>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2206.01653</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Manescu</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Bendkowski</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Claveau</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Elmi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Brown</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Fernandez-Reyes</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>b). <article-title>A weakly supervised deep learning approach for detecting malaria and sickle cells in blood films</article-title>. <source>Medical Image Computing and Computer Assisted Intervention &#x2013; MICCAI 2020</source>, <fpage>226</fpage>&#x2013;<lpage>235</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-59722-1_22</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manescu</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Shaw</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Elmi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Neary-Zajiczek</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Claveau</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Fernandez-Reyes</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>a). <article-title>Expert-level automated malaria diagnosis on routine blood films with deep neural networks</article-title>. <source>Am. J. Hematol.</source> <volume>95</volume> (<issue>8</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1002/ajh.25827</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mehanian</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Jaiswal</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Delahunt</surname> <given-names>C. B.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Horning</surname> <given-names>M. P.</given-names>
</name>
<name>
<surname>Bell</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Computer-automated malaria diagnosis and quantitation using convolutional neural networks</article-title>. <source>ICCV</source> pp. <fpage>116</fpage>&#x2013;<lpage>125</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCVW.2017.22</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>Ministerio de Salud</collab>
</person-group> (<year>2003</year>). <source>Manual de Procedimientos de Laboratoria Para el Diagnostico de Malaria</source> (<publisher-loc>Lima, Peru</publisher-loc>: <publisher-name>Instituto Nacional de Salud</publisher-name>).</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mischlinger</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Manego</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Mombo-Ngoma</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Ekoka Mbassi</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Hackbarth</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Ekoka Mbassi</surname> <given-names>F. A.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Diagnostic performance of capillary and venous blood samples in the detection of <italic>Loa loa</italic> and <italic>Mansonella perstans</italic> microfilaraemia using light microscopy</article-title>. <source>PloS Negl. Trop. Dis.</source> <volume>15</volume> (<issue>8</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pntd.0009623</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>Noul</collab>
</person-group> (<year>2023</year>). <source>miLab platform</source> (<publisher-loc>S. Korea</publisher-loc>: <publisher-name>Noul</publisher-name>). Available at: <uri xlink:href="https://noul.kr/en/milab">https://noul.kr/en/milab</uri>.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nowak</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Kothari</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Pannu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Algazi</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Prakash</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Inkwell: Design and validation of a low-cost open electricity-free 3d printed device for automated thin smearing of whole blood</article-title>. <source>arXiv</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2304.10200</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oyibo</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Meulah</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Bengtson</surname> <given-names>M.</given-names>
</name>
<name>
<surname>van Lieshout</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Oyibo</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Agbana</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Two-stage automated diagnosis framework for urogenital schistosomiasis in microscopy images from low-resource settings</article-title>. <source>J. Med. Imaging</source> <volume>10</volume> (<issue>4</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1117/1.JMI.10.4.044005</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oyibo</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Jujjavarapu</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Meulah</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Agbana</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Braakman</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Diehl</surname> <given-names>J. -C.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Schistoscope: An automated microscope with artificial intelligence for detection of <italic>Schistosoma haematobium</italic> eggs in resource-limited settings</article-title>. <source>Micromachines</source> <volume>13</volume> (<issue>643</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.3390/mi13050643</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Poostchi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ersoy</surname> <given-names>I.</given-names>
</name>
<name>
<surname>McMenamin</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Gordon</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Palaniappan</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Jaeger</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>a). <article-title>Malaria parasite detection and cell counting for human and mouse using thin blood smear microscopy</article-title>. <source>J. Med. Imaging</source> <volume>5</volume> (<issue>4</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1117/1.JMI.5.4.044506</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Poostchi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Silamut</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Maude</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Jaeger</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Thoma</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>b). <article-title>Image analysis and machine learning for detecting malaria</article-title>. <source>Trans. Res</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.trsl.2017.12.004</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rebelo</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Shapiro</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Amaral</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Melo-Cristino</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Hanscheid</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Haemozoin detection in infected erythrocytes for <italic>Plasmodium falciparum</italic> malaria diagnosis-prospects and limitations</article-title>. <source>Acta Tropica</source> <volume>123</volume> (<issue>1</issue>), <fpage>58</fpage>&#x2013;<lpage>61</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.actatropica.2012.03.005</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rees-Channer</surname> <given-names>R. R.</given-names>
</name>
<name>
<surname>Bachman</surname> <given-names>C. M.</given-names>
</name>
<name>
<surname>Grignard</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Gatton</surname> <given-names>M. L.</given-names>
</name>
<name>
<surname>Burkot</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Chiodini</surname> <given-names>P. L.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Evaluation of an automated microscope using machine learning for the detection of malaria in travelers returned to the UK</article-title>. <source>Front. Malaria</source> <volume>1</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fmala.2023.1148115</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Reinke</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Tizabi</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Understanding metric-related pitfalls in image analysis validation</article-title>. <source>Nat. Methods</source> <volume>21</volume>, <fpage>182</fpage>&#x2013;<lpage>194</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41592-023-02150-0</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Torres</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Bachman</surname> <given-names>C.M.</given-names>
</name>
<name>
<surname>Delahunt</surname> <given-names>C.B.</given-names>
</name>
<name>
<surname>Baldeon</surname> <given-names>J.A.</given-names>
</name>
<name>
<surname>Alava</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Bell</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). <article-title>Automated microscopy for routine malaria diagnosis: a field comparison on Giemsa-stained blood films in Peru</article-title>. <source>Malaria J</source> <volume>17</volume>, <fpage>33</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12936-018-2493-0</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vongpromek</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Proux</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Ekawati</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Archasuksan</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Bachman</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Bell</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Field evaluation of automated digital malaria microscopy: EasyScan GO</article-title>. <source>Trans. R Soc. Trop. Med. Hyg</source>. </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wiens</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Saria</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Sendak</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ghassemi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>V. X.</given-names>
</name>
<name>
<surname>Goldenberg</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Do no harm: a roadmap for responsible machine learning for health care</article-title>. <source>Nat. Med.</source> <volume>25</volume> (<issue>9</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41591-019-0548-6</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>White</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>The parasite clearance curve</article-title>. <source>Malaria J.</source> <volume>10</volume>, <fpage>278</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1475-2875-10-278</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2002</year>). <source>Prevention and control of schistosomiasis and soil-transmitted helminthiasis</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B40">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2009</year>). <source>Malaria Microscopy Quality Assurance Manual V1</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B41">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2010</year>). <source>Basic malaria microscopy. Part I. Learner&#x2019;s guide</source>. <edition>2nd ed</edition> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B42">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2011</year>). <source>Serum retinol concentrations for determining the prevalence of vitamin A deficiency in populations</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B43">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2016</year>a). <source>Microscopy examination of thick and thin blood films for identification of malaria parasites (esp SOPs 8 and 9)</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2016</year>b). <source>Microscopy for the detection, identification and quantification of malaria parasites on stained thick and thin blood films in research settings, ver 1</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B45">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2016</year>c). <source>Malaria microscopy quality assurance manual v2</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B46">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2016</year>d). <source>Malaria Microscopy Standard Operating Procedure MM-SOP-09: Malaria Parasite Counting</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B47">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2019</year>). <source>World malaria report 2019</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B48">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2021</year>a). <source>Generating evidence for artificial intelligence-based medical devices: a framework for training, validation and evaluation</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B49">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2021</year>b). <source>Diagnostic target product profiles for monitoring, evaluation and surveillance of schistosomiasis control programmes</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B50">
<citation citation-type="book">
<person-group person-group-type="author">
<collab>WHO</collab>
</person-group> (<year>2023</year>). <source>Global report on neglected tropical diseases 2023</source> (<publisher-loc>Geneva, Switzerland</publisher-loc>: <publisher-name>World Health Organization</publisher-name>).</citation>
</ref>
<ref id="B51">
<citation citation-type="web">
<person-group person-group-type="author">
<collab>WWARN</collab>
</person-group> (<year>2023</year>) <source>Obare method calculator</source>. Available online at: <uri xlink:href="https://www.wwarn.org/obare-method-calculator">https://www.wwarn.org/obare-method-calculator</uri>.</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Poostchi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Silamut</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Deep learning for smartphone-based malaria parasite detection in thick blood smears</article-title>. <source>IEEE J. BioMed. Health Inform</source> <volume>24</volume>, <fpage>1428</fpage>&#x2013;<lpage>1438</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JBHI.6221020</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Mohammed</surname> <given-names>F. O.</given-names>
</name>
<name>
<surname>Hamid</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Kassim</surname> <given-names>Y. M.</given-names>
</name>
<name>
<surname>Jaeger</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Patient-level performance evaluation of a smartphone-based malaria diagnostic application</article-title>. <source>Malaria J.</source> <volume>22</volume>, <fpage>33</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12936-023-04446-0</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>