<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Anal. Sci.</journal-id>
<journal-title>Frontiers in Analytical Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Anal. Sci.</abbrev-journal-title>
<issn pub-type="epub">2673-9283</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1105642</article-id>
<article-id pub-id-type="doi">10.3389/frans.2023.1105642</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Analytical Science</subject>
<subj-group>
<subject>Perspective</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Promoting transparency in forensic science by integrating categorical and evaluative reporting through decision theory</article-title>
<alt-title alt-title-type="left-running-head">Sigman and Williams</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frans.2023.1105642">10.3389/frans.2023.1105642</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Sigman</surname>
<given-names>Michael E.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2079664/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Williams</surname>
<given-names>Mary R.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>National Center for Forensic Science</institution>, <institution>University of Central Florida</institution>, <addr-line>Orlando</addr-line>, <addr-line>FL</addr-line>, <country>United States</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Chemistry</institution>, <institution>University of Central Florida</institution>, <addr-line>Orlando</addr-line>, <addr-line>FL</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1333844/overview">Maurice Aalders</ext-link>, University of Amsterdam, Netherlands</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1265702/overview">Luis Sarabia</ext-link>, University of Burgos, Spain</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Michael E. Sigman, <email>michael.sigman@ucf.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted toForensic Chemistry, a section of the journal Frontiers in Analytical Science</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>3</volume>
<elocation-id>1105642</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Sigman and Williams.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Sigman and Williams</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Forensic science standards often require the analyst to report in categorical terms. Categorical reporting without reference to the strength of the evidence, or the strength threshold that must be met to sustain or justify the decision, obscures the decision-making process, and allows for inconsistency and bias. Standards that promote reporting in probabilistic terms require the analyst to report the strength of the evidence without offering a conclusive interpretation of the evidence. Probabilistic reporting is often based on a likelihood ratio which depends on calibrated probabilities. While probabilistic reporting may be more objective and less open to bias than categorical reporting, the report can be difficult for a lay jury to interpret. These reporting methods may appear disparate, but the relationship between the two is easily understood and visualized by a simple decision theory construct known as the receiver operating characteristic (ROC) curve. Implementing ROC-facilitated reporting through an expanded proficiency testing regime may provide transparency in categorical reporting and potentially obviate some of the lay jury interpretation issues associated with probabilistic reporting.</p>
</abstract>
<kwd-group>
<kwd>Decision Theory</kwd>
<kwd>receiver operating characteristics</kwd>
<kwd>categorical reporting</kwd>
<kwd>evaluative reporting</kwd>
<kwd>probabilistic reporting</kwd>
<kwd>forensic science</kwd>
<kwd>transparency in reporting</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>In 2009 the National Academies of Science (NAS) issued a report entitled &#x201c;Strengthening Forensic Science in the United States: A Path Forward&#x201d;. (<xref ref-type="bibr" rid="B27">National Research Council, 2009</xref>). The main finding from the report was that &#x201c;(w)ith the exception of nuclear DNA analysis, &#x2026; no forensic method has been rigorously shown to have the capacity to consistently, and with a high degree of certainty, demonstrate a connection between evidence and a specific individual or source.&#x201d; In 2019, the Honorable Harry T. Edwards assessed progress of the forensic science community as &#x201c;still facing serious problems&#x201d; in his address to the Innocence Network Annual Conference. (<xref ref-type="bibr" rid="B10">Edwards, 2019</xref>). A significant aspect of the problem was summed up by the tautology that forensic practitioners addressing the NAS committee often didn&#x2019;t know what they didn&#x2019;t know. This condition becomes especially problematic when the expert is reporting to the court in categorical terms without reference to, and often without quantitative knowledge of, the strength of the evidence. Categorical reporting provides a multitude of terms that are not clearly defined, don&#x2019;t convey the strength of the evidence nor support more than one interpretation of the evidence (<xref ref-type="bibr" rid="B27">National Research Council, 2009</xref>; <xref ref-type="bibr" rid="B7">Cole and Biedermann, 2019</xref>). Without information regarding the strength of the evidence, the court cannot integrate the analyst&#x2019;s testimony into the overall evidence assessment.</p>
<p>The use of statistical data in forensic science can necessitate a move away from categorical statements to evaluative reporting. Under evaluative reporting the expert reports on the strength of the evidence in probabilistic terms and leaves the court to draw its own conclusions. (<xref ref-type="bibr" rid="B22">Martire et al., 2013</xref>; <xref ref-type="bibr" rid="B1">Aitken et al., 2015</xref>; <xref ref-type="bibr" rid="B6">Champod et al., 2016</xref>; <xref ref-type="bibr" rid="B3">Biedermann et al., 2017</xref>). However, the court may find probabilistic testimony difficult to interpret without guidance by the expert. The question is whether judges and jurors can interpret probabilistic reporting. (<xref ref-type="bibr" rid="B4">Brun and Teigen, 1988</xref>; <xref ref-type="bibr" rid="B9">De Keijser and Elffers, 2012</xref>; <xref ref-type="bibr" rid="B18">Friedman and Turri, 2015</xref>; <xref ref-type="bibr" rid="B31">Thompson and Newman, 2015</xref>; <xref ref-type="bibr" rid="B19">Hans and Saks, 2018</xref>; <xref ref-type="bibr" rid="B11">Eldridge, 2019</xref>; <xref ref-type="bibr" rid="B25">Melcher, 2022</xref>). A visual representation of decision theory provides a direct association between probabilistic statements, their verbal equivalents, and categorical statements. (<xref ref-type="bibr" rid="B16">Fawcett, 2004</xref>; <xref ref-type="bibr" rid="B20">Johnson, 2004</xref>; <xref ref-type="bibr" rid="B17">Fawcett, 2006</xref>).</p>
</sec>
<sec id="s2">
<title>2 Two current reporting practices</title>
<p>Reporting in categorical terms requires the analyst to make a decision regarding the interpretation of the evidence. If the question is a classification problem, the analyst must decide upon the class membership of the evidence and report their decision accordingly. If the problem is one of identification, then the binary categorical reporting must assign the questioned and controlled sample as coming from a common source. When reporting is done in categorical terms, the expert&#x2019;s opinion is often dogmatic and carries no indication of evidentiary strength or analyst uncertainty. The opinion can appear totally subjective and open to bias.</p>
<p>When reporting is done in probabilistic terms, the analyst typically reports the strength of the evidence as a likelihood ratio or the logarithm of the likelihood ratio. The likelihood ratio is typically reported in terms of two competing hypotheses. The report provides information to the court regarding the strength of the evidence relative to the two propositions, but a categorical statement is not provided. The court can then interpret the likelihood ratio, in conjunction with the prior odds of the two propositions to evaluate the posterior odds of the two propositions. This Bayesian approach can be argued to suffer from the need to estimate the prior odds and to arrive at a likelihood ratio from well-calibrated probabilities.</p>
<sec id="s2-1">
<title>2.1 Reporting examples</title>
<p>As an example of categorical reporting, under the ASTM E1618-19 standard, the analyst must report the sample as positive or negative for the presence of ignitable liquid residue. (<xref ref-type="bibr" rid="B23">Materials, 2019</xref>). In the case of comparative glass analysis, for example in a hit-and-run case, ASTM E2927-16e1 requires the analyst to report the evidence in a binary fashion as categorically representing an exclusion (different sources of the questioned and control samples) or an inclusion (same source) (<xref ref-type="bibr" rid="B2">Akmeemana et al., 2021</xref>). In both fire debris and glass analysis, recent research has focused on evaluation of the strength of evidence as a likelihood ratio that can lead to probabilistic reporting. (<xref ref-type="bibr" rid="B2">Akmeemana et al., 2021</xref>; <xref ref-type="bibr" rid="B30">Sigman et al., 2021</xref>; <xref ref-type="bibr" rid="B33">Whitehead et al., 2022</xref>). DNA evidence, which was recognized as a forensic gold-standard by the 2009 NAS report, has employed likelihood ratios to communicate the strength of the evidence for some time and advances in reporting continue, as reported in a recent review. (<xref ref-type="bibr" rid="B27">National Research Council, 2009</xref>; <xref ref-type="bibr" rid="B24">Meakin et al., 2021</xref>).</p>
</sec>
<sec id="s2-2">
<title>2.2 Errors, probabilities, and decisions</title>
<p>Various approaches have been proposed for data testing and classification, including hypothesis testing by Fisher (circa 1925) and Neyman and Pearson (circa 1928). (<xref ref-type="bibr" rid="B29">Perezgonzalez, 2015</xref>). Fisher&#x2019;s test involves establishing a null hypothesis (H<sub>0</sub>), typically that the difference in means between two populations or classes is equal to zero. Once the theoretical distribution is established for H<sub>0</sub>, the probability or <italic>p</italic>-value for any new data is calculated under H<sub>0</sub> and represents a cumulative probability of the observed result or a more extreme result. Results with a low <italic>p</italic>-value are taken as evidence against H<sub>0</sub> explaining the observed results. The Neyman-Pearson approach specifically considers an alternative hypothesis (H<sub>A</sub>) that represents a population which exists alongside a different population represented by the main hypothesis (H<sub>M</sub>). The samples in the population corresponding to H<sub>M</sub> are typically designated as class 0 and those in the adjacent population as class 1. The two populations exist alongside one another in the sense that they are each distributed over a common parameter or score, however the populations differ by some degree. The difference between populations is known as the effect size and could be as simple as the difference between the means of the two. The smaller the effect size, the more difficult it is to determine the difference between the two populations and to correctly predict a new observation&#x2019;s membership within each population (i.e., correct classification). Binary classifiers (two classes) classically attempt to minimize the expected classification error, defined as the weighted sum of type I and type II errors (defined below). (<xref ref-type="bibr" rid="B32">Tong et al., 2018</xref>). The Neyman-Pearson approach recognizes that in real-world cases these two error types may not be equally important, and it strives to limit the size of the more important (higher priority) error. Class labels can be arbitrarily switched so it is customary in the Neyman-Pearson approach to refer to the prioritized (more important) error as type I. Type I error refers to the conditional probability of mistakenly assigning a ground truth class 0 as belonging to class 1. This type I error is known as a false positive. The conditional probability of assigning a class 1 sample to class 0 is a false negative. For example, in fire debris analysis, samples that don&#x2019;t contain ignitable liquid residue would constitute class 0. The priority error (to be avoided) is to classify a class 0 sample as class 1 (i.e., containing ignitable liquid residue). The probability of committing a type I error is designated as &#x3b1; and the probability of committing a type II error is designated &#x3b2;. The goal in developing a Neyman-Pearson classification system is to keep &#x3b1; below a defined value (typically 0.05 or less) while also keeping &#x3b2; as small as possible. Data analysis methods that take into account type I and II errors are well-known and widely practiced in the analytical sciences. (<xref ref-type="bibr" rid="B28">Ortiz et al., 2010</xref>).</p>
<p>An approach to implementing a Neyman-Pearson classification without requiring assumptions regarding the probability distributions of the populations, is the &#x201c;umbrella&#x201d; method of Tong which utilizes receiver operating characteristic (ROC) curves. (<xref ref-type="bibr" rid="B32">Tong et al., 2018</xref>). The ROC method was developed by electrical engineers during World War II for the purpose of characterizing the abilities of RADAR operators to discern between targets of concern and noise that distracts from detection of the target. (<xref ref-type="bibr" rid="B5">Cal&#xec; and Longobardi, 2015</xref>). The ROC method is especially useful for binary decisions between two competing propositions. This is often the case in forensic science where the propositions of the prosecution, H<sub>p</sub>, and the defense, H<sub>d</sub>, often fall within a hierarchy of propositions. (<xref ref-type="bibr" rid="B8">Cook et al., 1998</xref>; <xref ref-type="bibr" rid="B14">Evett, 1998</xref>; <xref ref-type="bibr" rid="B12">Evett et al., 2000a</xref>). The propositions H<sub>p</sub> and H<sub>d</sub> can be defined such that they correspond to H<sub>M</sub> and H<sub>A</sub>, as discussed above. Applying the ROC method requires a ground-truth data set that must be evaluated and ranked relative to the two propositions by a score. Higher scores are associated with stronger evidence in support of a positive target detection, typically represented by H<sub>p</sub> in forensic science. The scores are sequentially treated as target-detection thresholds, such that any sample with a score equal or exceeding the threshold is assigned as positive for the target. The score-ranked data can be evaluated to determine the true positive rate (TPR) and false positive rate (FPR) of target detection at each threshold. A plot of TPR as a function of FPR, ordered by score, will produce a ROC curve. (<xref ref-type="bibr" rid="B16">Fawcett, 2004</xref>; <xref ref-type="bibr" rid="B20">Johnson, 2004</xref>; <xref ref-type="bibr" rid="B17">Fawcett, 2006</xref>; <xref ref-type="bibr" rid="B15">Fawcett and Niculescu-Mizil, 2007</xref>; <xref ref-type="bibr" rid="B5">Cal&#xec; and Longobardi, 2015</xref>). Each point on the ROC curve represents a decision threshold and projection at a right angle onto the respective axes gives the TPR and FPR statistics.</p>
</sec>
<sec id="s2-3">
<title>2.3 Relationship between categorical and probabilistic reporting</title>
<p>The ROC curve has several very useful properties. The curve is independent of the ratio of ground-truth positive and negative samples. It is independent of any parametric assumptions. The slope of a tangent to the curve at any point can be interpreted as a likelihood ratio. (<xref ref-type="bibr" rid="B5">Cal&#xec; and Longobardi, 2015</xref>). This later property of the ROC curve establishes the connection between the strengths of evidence (probabilistically reported as likelihood ratios) and a series of decision thresholds (categorically reported).</p>
<p>The ROC curve is typically plotted in a stepwise fashion; however, when the number of points is large, the curve will appear smooth, see <xref ref-type="fig" rid="F1">Figure 1</xref>. Base 10 loglikelihood ratio (LLR) values were used as scores to create the curve in <xref ref-type="fig" rid="F1">Figure 1</xref>. The scores are labeled along the curve at positions corresponding to TPR and FPR resulting from these LLR values as decision thresholds. In other words, the TPR and FPR rates are related to the strength of evidence required to support a decision in favor of H<sub>p</sub>. The two dashed lines have slopes of 10 (long dash) and 0.1 (short dash). The lines are tangents to the ROC curve at scores corresponding to their respective slopes (i.e., <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mn>10</mml:mn>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> ).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>An ROC curve (solid black line) demonstrating the relationship between evidential strength and decision thresholds. The scores that serve as decision thresholds are labeled next to their corresponding open circle symbols on the curve. The scores correspond to LLR values calculated from calibrated probabilities. Two blue dashed lines are drawn tangent to the ROC curve and the slope of the lines correspond to the likelihood ratio for the point where the tangent touches the ROC curve. The tangent line with longer dashes has a slope of 10 and corresponds to the slope of the ROC curve at a score (LLR) of 1 (see the text for additional explanation). The shorter dashed line has a slope of 0.1 and is tangent to the ROC curve at a score (LLR) of &#x2212;1. The inset diagram shows probability distributions for classes 0 and 1. The vertical &#x201c;optimal decision threshold&#x201d; corresponds to a LLR score of 1 and the shaded areas for a and b reflect the FPR and FNR (1-TPR) values where the red dashed lines intersect the corresponding axes. The inset diagram demonstrates the essential aspects of a Neyman-Pearson classifier (see text).</p>
</caption>
<graphic xlink:href="frans-03-1105642-g001.tif"/>
</fig>
<p>The straight lines drawn tangent to the ROC curve are also known as iso-performance lines. The slope of an iso-performance line is equal to the product of two ratios. The first ratio is the relative prior probabilities of H<sub>d</sub> divided by H<sub>p</sub>. The second ratio is the relative costs of a false positive assignment divided by the cost of a false negative assignment. All points falling on an iso-performance line share the same expected costs. This property allows for the selection of the optimal decision threshold along the ROC curve. An iso-performance slope is determined based on the known base rates (prior probabilities) and the acceptable cost ratio. Under this method, the optimal decision point is the ROC convex hull (CH) point that lies on the iso-performance line tangent to the curve. The ROC CH is composed of a series of straight segments connecting the outermost points on the curve.</p>
<p>The notations A&#x2013;F in <xref ref-type="fig" rid="F1">Figure 1</xref> correspond to LLR ranges for which verbal equivalents have been assigned. (<xref ref-type="bibr" rid="B14">Evett, 1998</xref>; <xref ref-type="bibr" rid="B13">Evett et al., 2000b</xref>; <xref ref-type="bibr" rid="B1">Aitken et al., 2015</xref>). New evidence (e.g., from casework) would be scored by the same method used to evaluate the samples that comprise the ROC curve. The evidential strength of new evidence is easily shown by plotting it onto the ROC curve. For example, the filled diamond plotted on the ROC curve in <xref ref-type="fig" rid="F1">Figure 1</xref> lies in the region of the curve labeled as &#x201c;B&#x201d; and provides moderate support for the prosecution&#x2019;s hypothesis, H<sub>p</sub>.</p>
<p>The ROC curve in <xref ref-type="fig" rid="F1">Figure 1</xref> is idealized and composed of a few thousand likelihood ratios calculated from calibrated probabilities to demonstrate the direct and highly visual relationship between the concepts underlying categorical and probabilistic reporting. In forensic applications, the number of data points would likely be much smaller. The scores should be numeric values that represent the degree of support for a sample belonging to the positive class. A score might not be a calibrated probability. (<xref ref-type="bibr" rid="B33">Whitehead et al., 2022</xref>). The ROC curve could more closely resemble <xref ref-type="fig" rid="F2">Figure 2</xref>. In <xref ref-type="fig" rid="F2">Figure 2</xref>, the ROC curve is shown as the solid black line and plotted in stairstep fashion. The dashed gray line connects the points on the ROC CH. The points on the ROC CH are the only points that qualify as optimal operational points for making categorical decisions. (<xref ref-type="bibr" rid="B16">Fawcett, 2004</xref>; <xref ref-type="bibr" rid="B17">2006</xref>). Following the Neyman-Pearson approach, the optimal operational threshold can be selected as a point on the ROC CH where the FPR (&#x3b1;) is less than a defined criteria (i.e., &#x3b1; &#x2264; 0.05). Each segment of the ROC CH represents an iso-performance line on the ROC plot. (<xref ref-type="bibr" rid="B17">Fawcett, 2006</xref>). The slope of each segment of the ROC CH is labeled next to the segment in <xref ref-type="fig" rid="F2">Figure 2</xref>. Note that &#x201c;Inf&#x201d; is used to note a slope approaching infinity, which is the limit of a positive change in TPR as the change in FPR approaches 0 from the positive side, as required due to the ROC curve existing in the first quadrant (FPR &#x3d; [0,1], TPR &#x3d; [0,1]). The slope of a ROC CH where the denominator was equal to 0 is mathematically undefined, and the slope of this ROC CH segment may only be considered in the limit. The probability of H<sub>p</sub> given the evidence for scores covered by the ROC CH segment is calculated by Eq. <xref ref-type="disp-formula" rid="e1">1</xref>, using the slope of the CH segment and the skew. The skew is the ratio of true-H<sub>p</sub> to true-H<sub>d</sub> case in the training set. (<xref ref-type="bibr" rid="B15">Fawcett and Niculescu-Mizil, 2007</xref>).<disp-formula id="e1">
<mml:math id="m2">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x2219;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x2219;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>An ROC curve (solid black line) comprised of a small number of loglikelihood-like scores (20 scores). The curve is plotted in a stepwise fashion and the scores are placed adjacent to open circle symbols. The dashed gray line connects the points comprising the ROC convex hull (CH). The slope of each segment of the ROC CH is plotted adjacent to the curve. The &#x201c;Inf&#x201d; label corresponds to the limiting slope of the ROC CH between scores 3 and 1. The slope of each segment of the ROC CH can be combined with the skew in the data used to generate the curve to calculate the PAV-equivalent probabilities P (H<sub>p</sub>&#x7c;E) using Eq <xref ref-type="disp-formula" rid="e1">1</xref> (see text for more details).</p>
</caption>
<graphic xlink:href="frans-03-1105642-g002.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F2">Figure 2</xref>, <inline-formula id="inf2">
<mml:math id="m3">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for new evidence with a score <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mn>3,2,1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.5</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for a score of 0, <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for a score of &#x2212;1, and <inline-formula id="inf6">
<mml:math id="m7">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>E</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for a score <inline-formula id="inf7">
<mml:math id="m8">
<mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. These values are calculated based on a skew of 1.</p>
</sec>
<sec id="s3">
<title>2.4 Implementing the methodology</title>
<p>Implementing the methodology described here requires a set of ground truth samples to evaluate following an established protocol and utilizing an accepted scoring method. Each analyst must evaluate a number of ground truth samples and use their assigned scores to generate an ROC curve. The scores should be loglikelihood-like and larger values should represent stronger support for H<sub>p</sub> (class 1). (<xref ref-type="bibr" rid="B26">Morrison, 2013</xref>). The scores can be obtained, for example, from machine learning, instrumental measurements, or subjective opinions representing an expected probability of membership in class 1. (<xref ref-type="bibr" rid="B21">J&#xf8;sang, 2016</xref>; <xref ref-type="bibr" rid="B32">Tong et al., 2018</xref>; <xref ref-type="bibr" rid="B30">Sigman et al., 2021</xref>). The number of samples to analyze should be determined within the organization and with the assistance of forensic statisticians. An example of the approach has been demonstrated in fire debris analysis with each of three analysts evaluating 20 ground truth samples each. (<xref ref-type="bibr" rid="B33">Whitehead et al., 2022</xref>). The samples must be presented to the analyst as blind or double-blind tests. This could be viewed as an extension of current proficiency exam requirements. After the development of an ROC curve by an analyst, an optimal decision threshold my be established if reporting in categorical terms is required. Casework samples must be analyzed and scored following the same protocols used to develop the ROC curve. The score obtained for the casework sample allows the determination of a PAV calibrated probability based on the covering segment of the ROCCH. Categorical reporting for the case sample would be dictated by the optimal decision threshold. The entire process is easily understood and highly visual.</p>
</sec>
</sec>
<sec id="s4">
<title>3 Discussion and conclusion</title>
<p>The relationship between decision-making and evidential strength has been demonstrated. The relationship is based on well-known and practiced engineering methods that provide a highly visual representation of the relationship. Applying these methods in forensic science could provide transparency to categorical reporting and potentially obviate some of the challenges faced by juries when trying to understand and interpret evidential strength and likelihood ratios. The ROC method provides a simple path to obtaining pooled-adjacent-violators (PAV) calibrated probabilities, which are required in forensic science. A decision threshold on an ROC curve defines the TPR and FPR rates for the method. In addition, the ROC curve provides a visualization of the trade-off between the TPR and FPR as the decision threshold is changed.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>MS: conceptualization, methodology, writing&#x2013;original draft, review and editing, supervision, project administration, funding acquisition, resources. MW: conceptualization, methodology, writing&#x2014;review and editing, supervision, funding acquisition, resources.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This research was supported by the National Center for Forensic Science, a Florida SUS recognized research center at the University of Central Florida.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Aitken</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Barrett</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Berger</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Biedermann</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Champod</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hicks</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> <source>ENFSI guideline for evaluative reporting in forensic science</source> (<year>2015</year>).</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akmeemana</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Weis</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Corzo</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ramos</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zoon</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Trejos</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> <article-title>Interpretation of chemical data from glass analysis for forensic purposes</article-title>. <source>J. Chemom.</source> (<year>2021</year>) <volume>35</volume>(<issue>1</issue>):<fpage>e3267</fpage>. <pub-id pub-id-type="doi">10.1002/cem.3267</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Biedermann</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Champod</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Willis</surname>
<given-names>S.</given-names>
</name>
</person-group> <article-title>Development of European standards for evaluative reporting in forensic science: The gap between intentions and perceptions</article-title>. <source>Int. J. Evid. Proof</source> (<year>2017</year>) <volume>21</volume>(<issue>1-2</issue>):<fpage>14</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1177/1365712716674796</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brun</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Teigen</surname>
<given-names>K. H.</given-names>
</name>
</person-group> <article-title>Verbal probabilities: Ambiguous, context-dependent, or both?</article-title> <source>Organ. Behav. Hum. Decis. Process.</source> (<year>1988</year>) <volume>41</volume>(<issue>3</issue>):<fpage>390</fpage>&#x2013;<lpage>404</lpage>. <pub-id pub-id-type="doi">10.1016/0749-5978(88)90036-2</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cal&#xec;</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Longobardi</surname>
<given-names>M.</given-names>
</name>
</person-group> <article-title>Some mathematical properties of the ROC curve and their applications</article-title>. <source>Ric. Mat.</source> (<year>2015</year>) <volume>64</volume>(<issue>2</issue>):<fpage>391</fpage>&#x2013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1007/s11587-015-0246-8</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Champod</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Biedermann</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Vuille</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Willis</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>De Kinder</surname>
<given-names>J.</given-names>
</name>
</person-group> <article-title>ENFSI guideline for evaluative reporting in forensic science: A primer for legal practitioners</article-title>. <source>Crim. Law Justice Wkly.</source> (<year>2016</year>) <volume>180</volume>(<issue>10</issue>):<fpage>189</fpage>&#x2013;<lpage>193</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cole</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Biedermann</surname>
<given-names>A.</given-names>
</name>
</person-group> <article-title>How can a forensic result Be a decision: A critical analysis of ongoing reforms of forensic reporting formats for federal examiners</article-title>. <source>Hous. L. Rev.</source> (<year>2019</year>) <volume>57</volume>:<fpage>551</fpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cook</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Evett</surname>
<given-names>I. W.</given-names>
</name>
<name>
<surname>Jackson</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Lambert</surname>
<given-names>J.</given-names>
</name>
</person-group> <article-title>A hierarchy of propositions: Deciding which level to address in casework</article-title>. <source>Sci. Justice</source> (<year>1998</year>) <volume>38</volume>(<issue>4</issue>):<fpage>231</fpage>&#x2013;<lpage>239</lpage>. <pub-id pub-id-type="doi">10.1016/s1355-0306(98)72117-3</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>De Keijser</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Elffers</surname>
<given-names>H.</given-names>
</name>
</person-group> <article-title>Understanding of forensic expert reports by judges, defense lawyers and forensic professionals</article-title>. <source>Psychol. Crime Law</source> (<year>2012</year>) <volume>18</volume>(<issue>2</issue>):<fpage>191</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1080/10683161003736744</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Edwards</surname>
<given-names>H. T.</given-names>
</name>
</person-group> <article-title>Ten years after the national academy of sciences&#x2019; landmark report on strengthening forensic science in the United States: A path forward&#x2013;where are we?</article-title> In: <source>NYU school of law, public law research paper (19-23)</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>New York University School of Law</publisher-name> (<year>2019</year>).</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eldridge</surname>
<given-names>H.</given-names>
</name>
</person-group> <article-title>Juror comprehension of forensic expert testimony: A literature review and gap analysis</article-title>. <source>Forensic Sci. Int. Synergy</source> (<year>2019</year>) <volume>1</volume>:<fpage>24</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1016/j.fsisyn.2019.03.001</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evett</surname>
<given-names>I. W.</given-names>
</name>
<name>
<surname>Jackson</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lambert</surname>
<given-names>J.</given-names>
</name>
</person-group> <article-title>More on the hierarchy of propositions: Exploring the distinction between explanations and propositions</article-title>. <source>Sci. justice J. Forensic Sci. Soc.</source> (<year>2000a</year>) <volume>40</volume>(<issue>1</issue>):<fpage>3</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1016/S1355-0306(00)71926-5</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evett</surname>
<given-names>I. W.</given-names>
</name>
<name>
<surname>Jackson</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lambert</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>McCrossan</surname>
<given-names>S.</given-names>
</name>
</person-group> <article-title>The impact of the principles of evidence interpretation on the structure and content of statements</article-title>. <source>Sci. justice J. Forensic Sci. Soc.</source> (<year>2000b</year>) <volume>40</volume>(<issue>4</issue>):<fpage>233</fpage>&#x2013;<lpage>239</lpage>. <pub-id pub-id-type="doi">10.1016/S1355-0306(00)71993-9</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evett</surname>
<given-names>I. W.</given-names>
</name>
</person-group> <article-title>Towards a uniform framework for reporting opinions in forensic science casework</article-title>. <source>Sci. Justice</source> (<year>1998</year>) <volume>3</volume>(<issue>38</issue>):<fpage>198</fpage>&#x2013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1016/s1355-0306(98)72105-7</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fawcett</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Niculescu-Mizil</surname>
<given-names>A.</given-names>
</name>
</person-group> <article-title>PAV and the ROC convex hull</article-title>. <source>Mach. Learn.</source> (<year>2007</year>) <volume>68</volume>(<issue>1</issue>):<fpage>97</fpage>&#x2013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-007-5011-0</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fawcett</surname>
<given-names>T.</given-names>
</name>
</person-group> <article-title>ROC graphs: Notes and practical considerations for researchers</article-title>. <source>Mach. Learn.</source> (<year>2004</year>) <volume>31</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fawcett</surname>
<given-names>T.</given-names>
</name>
</person-group> <article-title>An introduction to ROC analysis</article-title>. <source>Pattern Recognit. Lett.</source> (<year>2006</year>) <volume>27</volume>(<issue>8</issue>):<fpage>861</fpage>&#x2013;<lpage>874</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2005.10.010</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Friedman</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Turri</surname>
<given-names>J.</given-names>
</name>
</person-group> <article-title>Is probabilistic evidence a source of knowledge?</article-title> <source>Cognitive Sci.</source> (<year>2015</year>) <volume>39</volume>(<issue>5</issue>):<fpage>1062</fpage>&#x2013;<lpage>1080</lpage>. <pub-id pub-id-type="doi">10.1111/cogs.12182</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hans</surname>
<given-names>V. P.</given-names>
</name>
<name>
<surname>Saks</surname>
<given-names>M. J.</given-names>
</name>
</person-group> <article-title>Improving judge and jury evaluation of scientific evidence</article-title>. <source>Daedalus</source> (<year>2018</year>) <volume>147</volume>(<issue>4</issue>):<fpage>164</fpage>&#x2013;<lpage>180</lpage>. <pub-id pub-id-type="doi">10.1162/daed_a_00527</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>N. P.</given-names>
</name>
</person-group> <article-title>Advantages to transforming the receiver operating characteristic (ROC) curve into likelihood ratio co&#x2010;ordinates</article-title>. <source>Statistics Med.</source> (<year>2004</year>) <volume>23</volume>(<issue>14</issue>):<fpage>2257</fpage>&#x2013;<lpage>2266</lpage>. <pub-id pub-id-type="doi">10.1002/sim.1835</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>J&#xf8;sang</surname>
<given-names>A.</given-names>
</name>
</person-group> <source>Subjective logic</source>. <publisher-name>Springer</publisher-name> (<year>2016</year>).</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Martire</surname>
<given-names>K. A.</given-names>
</name>
<name>
<surname>Kemp</surname>
<given-names>R. I.</given-names>
</name>
<name>
<surname>Watkins</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Sayle</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Newell</surname>
<given-names>B. R.</given-names>
</name>
</person-group> <article-title>The expression and interpretation of uncertain forensic science evidence: Verbal equivalence, evidence strength, and the weak evidence effect</article-title>. <source>Law Hum. Behav.</source> (<year>2013</year>) <volume>37</volume>(<issue>3</issue>):<fpage>197</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1037/lhb0000027</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Materials</surname>
<given-names>A. S. f. T. a.</given-names>
</name>
</person-group> <source>Astm E1618&#x2013;19: Standard test method for ignitable liquid residues in extracts from fire debris samples by gas chromatography&#x2010;mass spectrometry</source>. <publisher-loc>Conshohocken, PA, USA</publisher-loc>: <publisher-name>ASTM International West</publisher-name> (<year>2019</year>).</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meakin</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Kokshoorn</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>van Oorschot</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Szkuta</surname>
<given-names>B.</given-names>
</name>
</person-group> <article-title>Evaluating forensic DNA evidence: Connecting the dots</article-title>. <source>Wiley Interdiscip. Rev. Forensic Sci.</source> (<year>2021</year>) <volume>3</volume>(<issue>4</issue>):<fpage>e1404</fpage>. <pub-id pub-id-type="doi">10.1002/wfs2.1404</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Melcher</surname>
<given-names>C. C.</given-names>
</name>
</person-group> <article-title>Uncovering the secrets of statistics as evidence in business valuations</article-title>. <source>Ct. Rev.</source> (<year>2022</year>) <volume>58</volume>:<fpage>68</fpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morrison</surname>
<given-names>G. S.</given-names>
</name>
</person-group> <article-title>Tutorial on logistic-regression calibration and fusion: Converting a score to a likelihood ratio</article-title>. <source>Aust. J. Forensic Sci.</source> (<year>2013</year>) <volume>45</volume>(<issue>2</issue>):<fpage>173</fpage>&#x2013;<lpage>197</lpage>. <pub-id pub-id-type="doi">10.1080/00450618.2012.733025</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<collab>National Research Council</collab>. <source>Strengthening forensic science in the United States: A path forward</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>National Academies Press</publisher-name> (<year>2009</year>).</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ortiz</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sarabia</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>S&#xe1;nchez</surname>
<given-names>M.</given-names>
</name>
</person-group> <article-title>Tutorial on evaluation of type I and type II errors in chemical analyses: From the analytical detection to authentication of products and process control</article-title>. <source>Anal. Chim. Acta</source> (<year>2010</year>) <volume>674</volume>(<issue>2</issue>):<fpage>123</fpage>&#x2013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1016/j.aca.2010.06.026</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Perezgonzalez</surname>
<given-names>J. D.</given-names>
</name>
</person-group> <article-title>Fisher, neyman-pearson or nhst? A tutorial for teaching data testing</article-title>. <source>Front. Psychol.</source> (<year>2015</year>) <volume>6</volume>:<fpage>223</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2015.00223</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sigman</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Thurn</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wood</surname>
<given-names>T.</given-names>
</name>
</person-group> <article-title>Validation of ground truth fire debris classification by supervised machine learning</article-title>. <source>Forensic Chem.</source> (<year>2021</year>) <volume>26</volume>:<fpage>100358</fpage>. <pub-id pub-id-type="doi">10.1016/j.forc.2021.100358</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thompson</surname>
<given-names>W. C.</given-names>
</name>
<name>
<surname>Newman</surname>
<given-names>E. J.</given-names>
</name>
</person-group> <article-title>Lay understanding of forensic statistics: Evaluation of random match probabilities, likelihood ratios, and verbal equivalents</article-title>. <source>Law Hum. Behav.</source> (<year>2015</year>) <volume>39</volume>(<issue>4</issue>):<fpage>332</fpage>&#x2013;<lpage>349</lpage>. <pub-id pub-id-type="doi">10.1037/lhb0000134</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tong</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J. J.</given-names>
</name>
</person-group> <article-title>Neyman-Pearson classification algorithms and NP receiver operating characteristics</article-title>. <source>Sci. Adv.</source> (<year>2018</year>) <volume>4</volume>(<issue>2</issue>):<fpage>eaao1659</fpage>. <pub-id pub-id-type="doi">10.1126/sciadv.aao1659</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whitehead</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Sigman</surname>
<given-names>M. E.</given-names>
</name>
</person-group> <article-title>Decision theory and linear sequential unmasking in forensic fire debris analysis: A proposed workflow</article-title>. <source>Forensic Chem.</source> (<year>2022</year>) <volume>29</volume>:<fpage>100426</fpage>. <pub-id pub-id-type="doi">10.1016/j.forc.2022.100426</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>