<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">628441</article-id>
<article-id pub-id-type="doi">10.3389/frai.2021.628441</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Interpretability Versus Accuracy: A&#x20;Comparison of Machine Learning Models Built Using Different Algorithms, Performance Measures, and Features to Predict <italic>E.&#x20;coli</italic> Levels in Agricultural Water</article-title>
<alt-title alt-title-type="left-running-head">Weller et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Predicting <italic>E.&#x20;coli</italic> Levels in Agricultural Water</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Weller</surname>
<given-names>Daniel L.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/814106/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Love</surname>
<given-names>Tanzy M. T.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wiedmann</surname>
<given-names>Martin</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/25299/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Department of Biostatistics and Computational Biology, University of Rochester, <addr-line>Rochester</addr-line>, <addr-line>NY</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Department of Food Science, Cornell University, <addr-line>Ithaca</addr-line>, <addr-line>NY</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Current Affiliation, Department of Environmental and Forest Biology, SUNY College of Environmental Science and Forestry, <addr-line>Syracuse</addr-line>, <addr-line>NY</addr-line>, <country>United&#x20;States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/718671/overview">Gregoire Mariethoz</ext-link>, University of Lausanne, Switzerland</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1159021/overview">Zeynal Topalcengiz</ext-link>, Mus Alparslan University, Turkey</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/710638/overview">Rabeah Al-Zaidy</ext-link>, King Abdullah University of Science and Technology, Saudi Arabia</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Daniel L. Weller, <email>wellerd2@gmail.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to AI in Food, Agriculture and Water, a section of the journal Frontiers in Artificial Intelligence</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>05</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>4</volume>
<elocation-id>628441</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>11</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>02</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Weller, Love and Wiedmann.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Weller, Love and Wiedmann</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Since <italic>E.&#x20;coli</italic> is considered a fecal indicator in surface water, government water quality standards and industry guidance often rely on <italic>E.&#x20;coli</italic> monitoring to identify when there is an increased risk of pathogen contamination of water used for produce production (e.g., for irrigation). However, studies have indicated that <italic>E.&#x20;coli</italic> testing can present an economic burden to growers and that time lags between sampling and obtaining results may reduce the utility of these data. Models that predict <italic>E.&#x20;coli</italic> levels in agricultural water may provide a mechanism for overcoming these obstacles. Thus, this proof-of-concept study uses previously published datasets to train, test, and compare <italic>E.&#x20;coli</italic> predictive models using multiple algorithms and performance measures. Since the collection of different feature data carries specific costs for growers, predictive performance was compared for models built using different feature types [geospatial, water quality, stream traits, and/or weather features]. Model performance was assessed against baseline regression models. Model performance varied considerably with root-mean-squared errors and Kendall&#x2019;s Tau ranging between 0.37 and 1.03, and 0.07 and 0.55, respectively. Overall, models that included turbidity, rain, and temperature outperformed all other models regardless of the algorithm used. Turbidity and weather factors were also found to drive model accuracy even when other feature types were included in the model. These findings confirm previous conclusions that machine learning models may be useful for predicting when, where, and at what level <italic>E.&#x20;coli</italic> (and associated hazards) are likely to be present in preharvest agricultural water sources. This study also identifies specific algorithm-predictor combinations that should be the foci of future efforts to develop deployable models (i.e.,&#x20;models that can be used to guide on-farm decision-making and risk mitigation). When deploying <italic>E.&#x20;coli</italic> predictive models in the field, it is important to note that past research indicates an inconsistent relationship between <italic>E.&#x20;coli</italic> levels and foodborne pathogen presence. Thus, models that predict <italic>E.&#x20;coli</italic> levels in agricultural water may be useful for assessing fecal contamination status and ensuring compliance with regulations but should not be used to assess the risk that specific pathogens of concern (e.g., <italic>Salmonella</italic>, <italic>Listeria</italic>) are present.</p>
</abstract>
<kwd-group>
<kwd>
<italic>E. coli</italic>
</kwd>
<kwd>machine learning</kwd>
<kwd>predictive model</kwd>
<kwd>food safety</kwd>
<kwd>water quality</kwd>
</kwd-group>
<contract-sponsor id="cn001">Center for Produce Safety<named-content content-type="fundref-id">10.13039/100008678</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">National Institute of Environmental Health Sciences<named-content content-type="fundref-id">10.13039/100000066</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>Following a 2018&#x20;Shiga-toxin producing <italic>Escherichia coli</italic> outbreak linked to romaine lettuce, investigators identified irrigation water contaminated by cattle feces as the probable source (<xref ref-type="bibr" rid="B9">Bottichio et&#x20;al., 2019</xref>). Such a conclusion is not uncommon, and fecal contamination of surface water has been repeatedly identified as the probable cause of enteric disease outbreaks (<xref ref-type="bibr" rid="B51">Johnson, 2006</xref>; <xref ref-type="bibr" rid="B1">Ackers et&#x20;al., 1998</xref>; <xref ref-type="bibr" rid="B92">Wachtel et&#x20;al., 2002</xref>; <xref ref-type="bibr" rid="B40">Greene et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B6">Barton Behravesh et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B32">Food and Drug Administration, 2019</xref>; <xref ref-type="bibr" rid="B33">Food and Drug Administration, 2020</xref>). As a result, non-pathogenic fecal indicator bacteria (FIBs), like <italic>E.&#x20;coli</italic>, are used to assess when and where fecal contaminants, including food and waterborne pathogens, may be present in agricultural and recreational waterways. Indeed, many countries and industry groups have established standards for agricultural and/or recreational surface water based on FIB levels; when samples are above a binary cut-off the probability of fecal contamination is deemed sufficient to require corrective action (<xref ref-type="bibr" rid="B46">Health Canada, 2012</xref>; <xref ref-type="bibr" rid="B90">US FDA, 2015</xref>; <xref ref-type="bibr" rid="B46">Health Canada, 2012</xref>; <xref ref-type="bibr" rid="B14">California Leafy Greens Marketing Agreement, 2017</xref>; <xref ref-type="bibr" rid="B29">Environmental Protection Agency, 2012</xref>; <xref ref-type="bibr" rid="B18">Corona et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B89">UK EA</xref>; <xref ref-type="bibr" rid="B30">EU Parliament, 2006</xref>; <xref ref-type="bibr" rid="B79">SA DWAF, 1996</xref>) For instance, the Australian and New&#x20;Zealand governments established trigger values for thermotolerant coliforms in water applied to food and non-food crops (<xref ref-type="bibr" rid="B3">ANZECC, 2000</xref>), while the United&#x20;States Produce Safety Rule (PSR) proposed an <italic>E. coli</italic>-based standard for surface water sources used for produce production (<xref ref-type="bibr" rid="B90">US FDA, 2015</xref>). Similarly, the California Leafy Greens Marketing Agreement requires <italic>E.&#x20;coli</italic> testing for determining the microbial quality of water used for produce production (<xref ref-type="bibr" rid="B14">California Leafy Greens Marketing Agreement, 2017</xref>). However, multiple studies have suggested that the frequency of sampling required by the PSR and similar regulations may not be sufficient to capture spatiotemporal variability in microbial water quality (<xref ref-type="bibr" rid="B27">Edge et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B60">McEgan et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B45">Havelaar et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>). Thus, supplementary or alternative approaches for monitoring surface water for potential public health hazards may be needed (<xref ref-type="bibr" rid="B27">Edge et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B60">McEgan et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B45">Havelaar et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>).</p>
<p>While an alternative to current monitoring practices is to more frequently measure FIB levels in the waterway (e.g., immediately before each irrigation event), studies that quantified costs associated with the United&#x20;States PSR found that the low-frequency testing proposed by the PSR presented a substantial economic burden to growers (<xref ref-type="bibr" rid="B15">Calvin et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B5">Astill et&#x20;al., 2018</xref>). Additional concerns about the feasibility of water testing (e.g., access/proximity to labs), and the time lag between sampling and time of water use (minimum of 24&#xa0;h) have also been raised (<xref ref-type="bibr" rid="B45">Havelaar et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B93">Wall et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>). Indeed, a study that sampled recreational waterways in Ohio over consecutive days found that a predictive model was able to better predict <italic>E.&#x20;coli</italic> levels than using <italic>E.&#x20;coli</italic> levels from samples collected on the day preceding sample collection (i.e.,&#x20;24&#xa0;h before) as the prediction; (<xref ref-type="bibr" rid="B12">Brady et&#x20;al., 2009</xref>). Predictive models may thus provide an alternative or supplementary approach to <italic>E. coli</italic>-based monitoring of agricultural and recreational surface water sources.</p>
<p>While past studies have shown that predictive models can be useful for assessing public health hazards in recreational water (<xref ref-type="bibr" rid="B72">Olyphant, 2005</xref>; <xref ref-type="bibr" rid="B49">Hou et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B11">Brady and Plona, 2009</xref>; <xref ref-type="bibr" rid="B42">Hamilton and Luffman, 2009</xref>; <xref ref-type="bibr" rid="B35">Francy et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B36">Francy et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B20">Dada and Hamilton, 2016</xref>; <xref ref-type="bibr" rid="B21">Dada, 2019</xref>; <xref ref-type="bibr" rid="B78">Rossi et&#x20;al., 2020</xref>), no models, to the author&#x2019;s knowledge, have been developed to predict <italic>E.&#x20;coli</italic> levels in surface water used for produce production (e.g., for irrigation, pesticide application, dust abatement, frost protection). Moreover, many of the recreational water quality studies only considered one algorithm during model development (e.g., (<xref ref-type="bibr" rid="B72">Olyphant, 2005</xref>; <xref ref-type="bibr" rid="B12">Brady et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B11">Brady and Plona, 2009</xref>; <xref ref-type="bibr" rid="B42">Hamilton and Luffman, 2009</xref>), including algorithms [e.g., regression, (<xref ref-type="bibr" rid="B72">Olyphant, 2005</xref>; <xref ref-type="bibr" rid="B12">Brady et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B11">Brady and Plona, 2009</xref>; <xref ref-type="bibr" rid="B42">Hamilton and Luffman, 2009</xref>)], which has more assumptions and may be less accurate than alternate algorithms (e.g., ensemble methods, support vector machines, (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>; <xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>)). As such, there is limited data on 1) how models for predicting <italic>E.&#x20;coli</italic> levels in agricultural water should be implemented and validated, or 2) how the data used to train these models should be collected (e.g., types of features to focus data collection efforts on). Addressing these knowledge gaps is key if the aim is to develop and deploy field-ready models (models that can be used to create a cost-effective tool with a GUI interface, incorporated into growers&#x2019; food safety plans, and used to guide on-farm decision-making in real-time). Thus, there is a specific need for studies that assess and compare the efficacy of models built using different algorithms and different features (e.g., weather, water quality). This latter point is particularly important since the collection of each feature type carries specific costs, including time and capital investment, worker training/expertize, and computational costs. For example, growers can often easily obtain, with no capital investment, weather data from publicly accessible stations (e.g., airport stations, AZMet [cals.arizona.edu/AZMET]), however, since these stations are unlikely to be located at a given farm, the utility of these data for training accurate predictive models need to be determined. Conversely, growers can collect physicochemical water quality data on-site provided they invest in equipment (e.g., water quality probes) and train staff to use the equipment. This proof-of-concept study aims to address these knowledge gaps and provide a framework on which future studies focused on developing field-ready models can build. Specifically, the objectives of this study were to 1) develop, assess, and compare the ability of models built using different algorithms and different combinations of feature types (e.g., geospatial, water quality, weather, and/or stream traits) to predict <italic>E.&#x20;coli</italic> levels, and 2) highlight how model interpretation is affected by the performance measure used. Since this is a proof-of-concept and not an empirical, study that used previously published data, the focus of the current paper is on identifying and comparing different algorithms, performance measures, and feature sets, and not on developing a deployable model, or characterizing relationships between <italic>E.&#x20;coli</italic> levels and features. The overarching aim of this paper is to provide a conceptual framework on which future studies can build, and to highlight what future studies should consider when selecting algorithms, performance measures, and feature&#x20;sets.</p>
<p>It is also important to remember when interpreting the findings presented here, that past research indicates an inconsistent relationship between <italic>E.&#x20;coli</italic> levels and foodborne pathogen presence (<xref ref-type="bibr" rid="B43">Harwood et&#x20;al., 2005</xref>; <xref ref-type="bibr" rid="B60">McEgan et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B73">Pachepsky et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B2">Antaki et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>). Thus, <italic>E.&#x20;coli</italic> models, like those developed here, may be useful for assessing fecal contamination status and ensure compliance with regulations but should not be used to determine if specific pathogens of concern (e.g., <italic>Salmonella</italic>, <italic>Listeria</italic>) are present. Since <italic>E.&#x20;coli</italic> is used outside food safety as an indicator of fecal contamination (e.g., for recreational water), the findings from this study may have implications for mitigating other public health concerns as&#x20;well.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>Materials and Methods</title>
<sec id="s2-1">
<title>Study Design and <italic>E.&#x20;coli</italic> Enumeration</title>
<p>Existing datasets collected in 2018 (<xref ref-type="bibr" rid="B96">Weller et&#x20;al., 2020b</xref>) and 2017 (<xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>) were used as the training and testing data, respectively, in the analyses reported here. Although the present study uses data from published empirical studies that characterized relationships between microbial water quality and environmental conditions, the study reported here is a survey focused on comparing algorithms and providing guidance for future modeling efforts.</p>
<p>Although the same sampling and laboratory protocols were used to generate both datasets, the datasets differ in the number of streams sampled (2017 &#x3d; 6 streams; 2018&#x20;&#x3d; 68 streams; <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>), and sampling frequency (2017 &#x3d; 15&#x2013;34 sampling visits per stream; 2018&#x20;&#x3d; 2&#x2013;3 visits per stream, (<xref ref-type="bibr" rid="B96">Weller et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>)). As a result, the 2017 and 2018 data represent 181 and 194 samples, respectively, (<xref ref-type="bibr" rid="B96">Weller et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>). At each sampling, a 1&#xa0;L grab sample was collected and used for <italic>E.&#x20;coli</italic> enumeration using the IDEXX Colilert-2000 test per manufacturer&#x2019;s instructions (IDEXX, Westbrook, ME). Between sample collection and enumeration (&#x3c;6&#xa0;h), samples were kept at 4&#xb0;C.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Sampling sites for each of the streams represented in the training and test data used here. Point size is proportional to <bold>(A)</bold> mean <italic>E.&#x20;coli</italic> concentration (log10 MPN/100-ml) and <bold>(B)</bold> mean turbidity levels (log10 NTUs). Municipal boundaries (yellow), and major lakes (blue) are included as references. The map depicts the Finger Lakes, Western, and Southern Tier regions of New York State, United&#x20;States.</p>
</caption>
<graphic xlink:href="frai-04-628441-g001.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>Metadata</title>
<p>Spatial data were obtained from publicly available sources and analyzed using ArcGIS version 10.2 and R version 3.5.3. Briefly, the inverse-distance weighted (IDW) proportion of cropland, developed land, forest-wetland cover, open water, and pasture land for each watershed as well as the floodplain and stream corridor upstream of each sampling site was calculated as previously described [(<xref ref-type="bibr" rid="B52">King et&#x20;al., 2005</xref>; <xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>; <xref ref-type="sec" rid="s9">Supplementary Table S1</xref>). In addition to characterizing land cover, we also determined if specific features were present in each watershed. If a feature was present, the distance to the feature closest to the sampling site, and feature density were determined (for the full list see <xref ref-type="sec" rid="s9">Supplementary Table&#x20;S1</xref>).</p>
<p>Physicochemical water quality and air temperature were measured at sample collection (<xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>). Separately, rainfall, temperature, and solar radiation data were obtained from the NEWA weather station (<ext-link ext-link-type="uri" xlink:href="http://newa.cornell.edu/">newa.cornell.edu</ext-link>) closest to each sampling site (Mean Distance &#x3d; 8.9&#xa0;km). If a station malfunctioned, data from the next nearest station were used. Average air temperature and solar radiation, and total rainfall were calculated using non-overlapping periods (e.g., 0&#x2013;1&#xa0;day before sampling, 1&#x2013;2&#xa0;days before sampling; <xref ref-type="sec" rid="s9">Supplementary Table&#x20;S1</xref>).</p>
</sec>
<sec id="s2-3">
<title>Statistical Analyses</title>
<p>All analyses were performed in R (version 3.5.3; R Core Team, Vienna, Austria) using the mlr package (<xref ref-type="bibr" rid="B8">Bischl et&#x20;al., 2016</xref>). Model training and testing were performed using the 2018 (<xref ref-type="bibr" rid="B96">Weller et&#x20;al., 2020b</xref>) and 2017 (<xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>) data, respectively. Hyperparameter tuning was performed using 3-fold cross-validation repeated 10 times. Tuning was performed to optimize root mean squared error (RMSE). After tuning, models were trained and performance assessed using RMSE, <italic>R</italic>
<sup>2</sup>, and Kendall&#x2019;s Tau (<italic>&#x3c4;</italic>). All covariates were centered and scaled before model development.</p>
<p>The algorithms used here were chosen to: 1) be comparable to algorithms used in past studies that predicted foodborne pathogen presence in farm environments [e.g., random forest, regression trees (<xref ref-type="bibr" rid="B75">Polat et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B83">Strawn et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B38">Golden et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B94">Weller et&#x20;al., 2016</xref>)], and 2) include algorithms that appear promising but have not been previously utilized for produce safety applications (e.g., extremely randomized trees, cubist). Extensive feature engineering was not done before model implementation since the aim was to 1) compare algorithm performance on the same, unaltered dataset, and 2) as an opportunity to highlight where and how (e.g., for neural nets; <xref ref-type="table" rid="T1">Table&#x20;1</xref>) feature engineering may be needed. Moreover, due to the plethora of approaches to feature selection and engineering, a separate paper focused on assessing the impact of feature selection and engineering decisions on the performance of <italic>E.&#x20;coli</italic> predictive models may be warranted. In total, 19 algorithms that fall into one of seven categories [support vector machines (SVM), cubist, decision trees, regression, neural nets, k-nearest neighbor (KNN), and forests] were used to develop the models presented here. However, a total of 26 models were developed using all predictors listed in <xref ref-type="sec" rid="s9">Supplementary Table S1</xref> (i.e.,&#x20;26 full models) since multiple variations of the SVM (4 variations), cubist (4 variations), and KNN (2 variations) were considered. While advantages and disadvantages for each algorithm are outlined briefly in <xref ref-type="table" rid="T1">Table&#x20;1</xref> and the discussion, more in-depth comparisons can be found in Kuhn and Johnson (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>).</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>List of algorithms used in the study reported here. This table was adapted from <xref ref-type="bibr" rid="B53">Kuhn and Johnson (2016)</xref> and <xref ref-type="bibr" rid="B95">Weller et al., (2020a)</xref> to i) reflect the algorithms used here, and ii) report information relevant to continuous (as opposed to categorical) data<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
<sup>,</sup>
<xref ref-type="table-fn" rid="Tfn2">
<sup>b</sup>
</xref>
</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" colspan="2" align="left">Algorithm</th>
<th rowspan="2" align="center">Package</th>
<th rowspan="2" align="center">
<italic>n</italic> &#x3c; <italic>p</italic>
</th>
<th rowspan="2" align="center">Centering and Scaling Recommended</th>
<th colspan="4" align="center">For Features, It Can Handle</th>
<th rowspan="2" align="center">Automatic Feature Selection</th>
<th rowspan="2" align="center">Interpretable</th>
</tr>
<tr>
<th align="center">Correlation</th>
<th align="center">Missingness</th>
<th align="center">Near-Zero Variance</th>
<th align="center">Noise</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="5" align="left">Tree-based Learners</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">Conditional Inference Tree</td>
<td align="left">party (<xref ref-type="bibr" rid="B48">Hothorn et al., 2006</xref>; <xref ref-type="bibr" rid="B84">Strobl et al., 2007a</xref>; <xref ref-type="bibr" rid="B86">Strobl et al., 2008</xref>; <xref ref-type="bibr" rid="B87">Strobl et al., 2009</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Evolutionary Optimal Tree</td>
<td align="left">evtree (<xref ref-type="bibr" rid="B41">Grubinger et al., 2014</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="left"/>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Regression Tree<xref ref-type="table-fn" rid="Tfn3">
<sup>c</sup>
</xref>
</td>
<td align="left">rpart (<xref ref-type="bibr" rid="B88">Therneau and Atkinson, 2019</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td colspan="5" align="left">Ensemble Learners</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">Conditional Forest</td>
<td align="left">party (<xref ref-type="bibr" rid="B48">Hothorn et al., 2006</xref>; <xref ref-type="bibr" rid="B84">Strobl et al., 2007a</xref>; <xref ref-type="bibr" rid="B86">Strobl et al., 2008</xref>; <xref ref-type="bibr" rid="B87">Strobl et al., 2009</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Extremely Randomized Trees</td>
<td align="left">extraTrees (<xref ref-type="bibr" rid="B61">Meinshausen, 2010</xref>)</td>
<td align="left"/>
<td align="left"/>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="left"/>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">Node Harvest<xref ref-type="table-fn" rid="Tfn3">
<sup>c</sup>
</xref>
</td>
<td align="left">nodeHarvest (<xref ref-type="bibr" rid="B57">Liaw et&#x20;al., 2002</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Random Forest<xref ref-type="table-fn" rid="Tfn3">
<sup>c</sup>
</xref>
</td>
<td align="left">randomForest (<xref ref-type="bibr" rid="B57">Liaw et&#x20;al., 2002</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Regularized Random Forest</td>
<td align="left">RRF (<xref ref-type="bibr" rid="B23">Deng and Runger, 2012</xref>; <xref ref-type="bibr" rid="B24">Deng and Runger, 2013</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Extreme Gradient Boosting</td>
<td align="left">xgboost (<xref ref-type="bibr" rid="B17">Chen and Guestrin, 2016</xref>; <xref ref-type="bibr" rid="B13">Brownlee, 2019</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
</tr>
<tr>
<td colspan="5" align="left">Instance-Based Learners</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">k-Nearest Neighbor</td>
<td align="left">kknn (<xref ref-type="bibr" rid="B47">Hechenbichler and Schliep, 2004</xref>)</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Weighted k-Nearest Neighbor</td>
<td align="left">kknn (<xref ref-type="bibr" rid="B47">Hechenbichler and Schliep, 2004</xref>)</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
</tr>
<tr>
<td colspan="2" align="left">Multivariate Adaptive Regression Splines</td>
<td align="left">earth (<xref ref-type="bibr" rid="B64">Milborrow, 2011</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="left"/>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
</tr>
<tr>
<td colspan="2" align="left">Neural Network<xref ref-type="table-fn" rid="Tfn4">
<sup>d</sup>
</xref>
</td>
<td align="left">nnet (<xref ref-type="bibr" rid="B91">Venable et al., 2002</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="left"/>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
</tr>
<tr>
<td colspan="2" align="left">Regression</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">Log-Linear</td>
<td align="left">stats</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Partial Least Squares</td>
<td align="left">pls (<xref ref-type="bibr" rid="B62">Mevik et al., 2019</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
<xref ref-type="table-fn" rid="Tfn5">
<sup>e</sup>
</xref>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Principal Component</td>
<td align="left">pls (<xref ref-type="bibr" rid="B62">Mevik et al., 2019</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
<xref ref-type="table-fn" rid="Tfn5">
<sup>e</sup>
</xref>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="left"/>
<td align="left"/>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="left"/>
</tr>
<tr>
<td colspan="5" align="left">Penalized Regression</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">Elastic Net</td>
<td align="left">glmnet (<xref ref-type="bibr" rid="B37">Friedman et al., 2010</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Lasso</td>
<td align="left">glmnet (<xref ref-type="bibr" rid="B37">Friedman et al., 2010</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left">Ridge</td>
<td align="left">glmnet (<xref ref-type="bibr" rid="B37">Friedman et al., 2010</xref>)</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
</tr>
<tr>
<td colspan="5" align="left">Rule-Based Algorithms</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
</tr>
<tr>
<td align="left"/>
<td align="left">Cubist</td>
<td align="left">Cubist (<xref ref-type="bibr" rid="B54">Kuhn and Quinlan, 2018</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="left"/>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
</tr>
<tr>
<td colspan="2" align="left">SVM</td>
<td align="left">e1071 (<xref ref-type="bibr" rid="B63">Meyer et al., 2020</xref>)</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>&#x2022;</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>Y</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
<td align="center">
<bold>N</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="Tfn1">
<label>
<sup>a</sup>
</label>
<p>The information reported here is based on i) <xref ref-type="bibr" rid="B53">Kuhn and Johnson, (2016</xref>), ii) the papers cited for each algorithm in the methods section, and iii) the constraints listed in the R packages below (based on the package version available in January 2020).</p>
</fn>
<fn id="Tfn2">
<label>
<sup>b</sup>
</label>
<p>Y means the algorithm meets the condition in the header. N means the algorithm does not meet this condition. &#x2022; means the algorithm is in between (e.g., random forest is not as interpretable as tree-based methods but is not a 100% black-box method like support vector machines). If the cell is blank it means there was limited information on this condition for the given algorithm.</p>
</fn>
<fn id="Tfn3">
<label>
<sup>c</sup>
</label>
<p>Preferentially selects continuous factors and categorical factors with many levels as the splitting variable resulting in variable selection bias (<xref ref-type="bibr" rid="B85">Strobl et al., 2007b</xref>; <xref ref-type="bibr" rid="B86">Strobl et al., 2008</xref>; <xref ref-type="bibr" rid="B87">Strobl et al., 2009</xref>).</p>
</fn>
<fn id="Tfn4">
<label>
<sup>d</sup>
</label>
<p>Feature selection recommended before model development.</p>
</fn>
<fn id="Tfn5">
<label>
<sup>e</sup>
</label>
<p>Centering and scaling are required but are performed as part of model fitting in the R package.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Separately from the full models, nested models were built using different feature subsets (<xref ref-type="sec" rid="s9">Supplementary Table S1</xref>). Features were divided into four categories: 1) geospatial, 2) physicochemical water quality and temperature data collected on-site, 3) all other weather data, and 4) stream traits that were observable on-site (e.g., composition of stream bottom). Nested models were built using different combinations of these four feature types, and one of nine algorithms (see <xref ref-type="sec" rid="s9">Supplementary Table S3</xref>; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> for all feature-algorithm combinations; a total of 90 nested models). The nine algorithms used to build the nested models were randomly selected from the list of 26 algorithms used to build the full models.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Graph of RMSE, which measures a model&#x2019;s ability to predict absolute <italic>E.&#x20;coli</italic> count, vs. Kendall&#x2019;s tau, which measures the model&#x2019;s ability to predict relative <italic>E.&#x20;coli</italic> concentration. The dashed line represents the RMSE for the featureless regression model; an RMSE to the right of this line indicates that the model was unable to predict absolute <italic>E.&#x20;coli</italic> counts. To facilitate readability, nested models are displayed in a separate facet from the full and log-linear models. <xref ref-type="sec" rid="s9">Supplementary Figures S3, S4</xref> display the nested model facet as a series of convex hulls to facilitate comparisons between models built using different feature types and algorithms, respectively. The top-performing models are in the top left of each&#x20;facet.</p>
</caption>
<graphic xlink:href="frai-04-628441-g002.tif"/>
</fig>
<p>Model performance was ranked using RMSE; models that tied were assigned the same rank. Two other performance measures, Kendall&#x2019;s tau (<italic>&#x3c4;</italic>) and the coefficient of determination (<italic>R</italic>
<sup>2</sup>) were also calculated. Kendall&#x2019;s tau is a rank-based measure that assesses a model&#x2019;s ability to correctly identify the relative (but not the absolute) concentration of <italic>E.&#x20;coli</italic> in novel samples (e.g., if a sample was predicted to have a high or low <italic>E.&#x20;coli</italic> concentration), while <italic>R</italic>
<sup>2</sup> assesses how much variation in <italic>E.&#x20;coli</italic> levels is predictable using the given model. Predictive performance for top-ranked models was visualized using density and split quantiles plots. An explanation of how to interpret these plots is included in the figure legends. For top-ranked models, the iml package (<xref ref-type="bibr" rid="B31">Fisher et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B65">Molnar et&#x20;al., 2018</xref>) was used to calculate permutation variable importance (PVI) and identify the features most strongly associated with accurately predicting <italic>E.&#x20;coli</italic> levels in the test and training data. Accumulated local effects plots were used to visualize the relationship between <italic>E.&#x20;coli</italic> levels, and the six factors with the highest PVI (<xref ref-type="bibr" rid="B4">Apley and Zhu, 2016</xref>).</p>
<sec id="s2-3-1">
<title>Baseline Models</title>
<p>
<xref ref-type="bibr" rid="B75">Polat et&#x20;al. (2019</xref>) developed a series of univariable models to predict <italic>Salmonella</italic> presence in Florida irrigation water. Each model was built using one of nine water quality or weather features (<xref ref-type="bibr" rid="B75">Polat et&#x20;al., 2019</xref>). Studies conducted in nonagricultural, freshwater environments (e.g., swimming beaches) that focused on developing interpretable models used similar sets of physicochemical and weather features (<xref ref-type="bibr" rid="B70">Olyphant and Whitman, 2004</xref>; <xref ref-type="bibr" rid="B34">Francy and Darner, 2006</xref>; <xref ref-type="bibr" rid="B28">Efstratiou et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B80">Shiels and Guebert, 2010</xref>; <xref ref-type="bibr" rid="B35">Francy et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B10">Bradshaw et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B21">Dada, 2019</xref>). To ensure comparability with these previous studies, and provide baseline models that could be used to gauge full and nested model performance, we developed eight log-linear and a featureless regression model. Unlike the log-linear models, the featureless model did not include any features; models outperformed by the featureless model were unable to predict <italic>E.&#x20;coli</italic> levels. Seven, separate univariable log-linear models were created using each of the following factors: air temperature at sample collection, conductivity, dissolved oxygen, pH, rainfall 0&#x2013;1&#xa0;day before sample collection, turbidity, and water temperature. An eighth model was built using air temperature, rainfall, and turbidity.</p>
</sec>
<sec id="s2-3-2">
<title>Tree-based and Forest Algorithm</title>
<p>Three tree-based algorithms were used: regression trees (CART), conditional inference trees (CTree), and evolutionary optimal trees (evTree) as described in a previous study focused on predicting pathogen presence in agricultural water (<xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>). Briefly, the three tree-based algorithms were implemented using the rpart (<xref ref-type="bibr" rid="B88">Therneau and Atkinson, 2019</xref>), party (<xref ref-type="bibr" rid="B48">Hothorn et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B85">Strobl et&#x20;al., 2007b</xref>; <xref ref-type="bibr" rid="B86">Strobl et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B87">Strobl et&#x20;al., 2009</xref>), and evtree (<xref ref-type="bibr" rid="B41">Grubinger et&#x20;al., 2014</xref>) packages, respectively. The number of splits in each tree and the min number of observations allowed in terminal nodes were tuned for each algorithm. Complexity parameters were tuned to minimize overfitting when implementing the CART and evTree algorithms, while the mincriterion parameter was set to 0.95 when implementing the CTree algorithm.</p>
<p>Six ensemble algorithms [conditional forest (condRF); extreme gradient boosting (xgBoost); node Harvest; random forest (RF); regularized random forest (RRF), and exTree] were implemented. For the three random forest algorithms, the number of factors considered for each split, and the minimum number of observations allowed in terminal nodes was tuned. To minimize overfitting the coefficient of regularization was tuned for regRF models, while the mincriterion parameter was tuned for condRF models. When implementing the xgBoost algorithm (<xref ref-type="bibr" rid="B17">Chen and Guestrin, 2016</xref>), hyperparameters were tuned that control: 1) learning rate and overfitting; 2) if splits were formed and the max. number of splits allowed; 3) number of rounds of boosting; 4) proportion of data used to build each tree; 5) number of features considered when building each tree; and 6) regularization. When implementing the node Harvest algorithm, hyperparameters were tuned that control the: 1) min number of samples to use to build each tree, and 2) max. number of splits allowed in each tree. Unlike the five other forest-based learners, the number of samples used to build each tree was not tuned when implementing the exTree algorithm since neither bagging, bootstrapping, nor boosting is performed when building exTrees (<xref ref-type="bibr" rid="B81">Simm et&#x20;al., 2014</xref>). Instead, hyperparameters were tuned that control the: 1) number of features considered when building each node; 2) the max. size of terminal nodes; and 3) the number of discretization points to select at random when defining a new node. The latter parameter highlights a key difference between the exTrees and random forest algorithms; random forests use local optimization to make the best split for a given node, which may not be globally optimal. To overcome this limitation and decrease computation time, both the variable used in new nodes, and the cutpoint used to split that variable were chosen randomly. For all ensemble methods the number of trees used was set to 20,001.</p>
</sec>
<sec id="s2-3-4">
<title>Instance-Based Algorithms</title>
<p> Two instance-based algorithms [k-nearest neighbor (kKNN) and weighted k-nearest neighbor (wKNN)] were implemented (<xref ref-type="bibr" rid="B47">Hechenbichler and Schliep, 2004</xref>). Implementation of instance-based algorithms requires tuning the number of neighbors used when predicting a novel observation. Additionally, the method for calculating distances between neighbors (Euclidean or Manhattan) was tuned when implementing the KNN algorithms. For wKNN, the weighting kernel was also tuned since several weighting approaches&#x20;exist.</p>
</sec>
<sec id="s2-3-5">
<title>Neural Nets</title>
<p>Neural networks are a non-linear regression technique and were implemented here using the mlr (<xref ref-type="bibr" rid="B8">Bischl et&#x20;al., 2016</xref>) and nnet (<xref ref-type="bibr" rid="B91">Venable et&#x20;al., 2002</xref>) packages. Unlike the other algorithms used here, neural nets cannot handle correlated or collinear predictors ((<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>); <xref ref-type="table" rid="T1">Table&#x20;1</xref>). As such, feature selection was performed before fitting the neural nets by retaining all predictors with 1) non-zero coefficients according to the full elastic net model, and 2) non-zero variable importance measures according to the full condRF model. In neural net models, the outcome is predicted using an intermediary set of unobserved variables that are linear combinations of the original predictors (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). As such, the number of intermediary variables used in the model was tuned as was the max. iterations run. Since neural net models often overfit (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>), a regularization parameter was also&#x20;tuned.</p>
</sec>
<sec id="s2-3-6">
<title>Regression and Penalized Algorithms</title>
<p>Regression models, like the baseline models developed here, are frequently used to assess associations between features and food safety outcomes [e.g., (<xref ref-type="bibr" rid="B98">Wilkes et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B7">Benjamin et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B16">Ceuppens et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B74">Pang et&#x20;al., 2017</xref>)]. However, conventional regression cannot handle correlated or collinear features or a large number of features. Various algorithms have been developed to overcome these limitations (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). We used five such algorithms here, including penalized regression, partial least squares regression, and principal component regression. Penalized regression models apply a penalty to the sum of squared estimates of errors (SSE) to control the magnitude of the parameter estimates, and account for correlation between features (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). All three penalized algorithms used here (ridge, lasso, and elastic net) were fit using the glmnet package using 10&#x20;cross-validated folds (<xref ref-type="bibr" rid="B37">Friedman et&#x20;al., 2010</xref>), which automatically tunes lambda (amount of coefficient shrinkage). For all three models, a hyperparameter was tuned that determines if the model with the min mean cross-validated error or the model within one standard error of the min. was retained. For ridge and lasso regression, alpha was set to 0 or 1, respectively, while alpha was tuned for the elastic net&#x20;model.</p>
<p>To overcome limitations associated with correlated features or having large numbers of features, principal components regression (PCR) uses a two-step approach. The dataset dimension is first reduced using principal components analysis (PCA), and then regression is performed using the principal components as features. Since PCA is performed independently of the outcome response, PCA may not produce components that explain the outcome, resulting in a poor-performing model (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). Partial least squares regression (PLS) does not suffer from this limitation. Like PCR, PLS finds underlying, linear combinations of the predictors. Unlike PCR, which selects combinations to maximally summarize the features, PLS selects combinations that maximally summarize covariance in the outcome (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). PCR and PLS models were both fit using the mlr (<xref ref-type="bibr" rid="B8">Bischl et&#x20;al., 2016</xref>) and pls (<xref ref-type="bibr" rid="B62">Mevik et&#x20;al., 2019</xref>) packages, and the number of components used was tuned. Since there are several variations of the PLS algorithm (<xref ref-type="bibr" rid="B62">Mevik et&#x20;al., 2019</xref>), the PLS algorithm used was also&#x20;tuned.</p>
</sec>
<sec id="s2-3-7">
<title>Multivariate Adaptive Regression Splines (MARS)</title>
<p>Like PLS and neural net, the MARS algorithm uses the features to create new, unobserved intermediary variables that are used to generate model predictions (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). MARS creates each new intermediary using fewer features than PLS and neural net. MARS uses a piecewise linear regression approach that allows each intermediary to model a separate part of the training data and automatically accounts for interactions (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). As in other approaches, once a full set of intermediaries has been created, pruning is performed to remove intermediaries that do not contribute to model performance (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). The MARS models created here were implemented using the mlr (<xref ref-type="bibr" rid="B8">Bischl et&#x20;al., 2016</xref>) and mda (<xref ref-type="bibr" rid="B44">Hastie and Tibshirani, 2017</xref>) packages. When fitting the MARS models the number and complexity of the intermediaries retained in the final model were&#x20;tuned.</p>
</sec>
<sec id="s2-3-8">
<title>Rule-Based Algorithms</title>
<p>Four variations of the Cubist algorithm were implemented (<xref ref-type="bibr" rid="B54">Kuhn and Quinlan, 2018</xref>). Cubist models grow a tree where each terminal node contains a separate linear regression model. Predictions are made using these terminal models but smoothed using the model immediately above the given terminal node. Ultimately, this results in a series of hierarchical paths from the top to the bottom of the tree. To prevent overfitting these paths are converted to rules, which are pruned or combined based on an adjusted error rate (<xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>). Like tree-based models, an ensemble of Cubist models can be created and the predicted <italic>E.&#x20;coli</italic> concentration from all constituent models averaged to obtain the model prediction (<xref ref-type="bibr" rid="B54">Kuhn and Quinlan, 2018</xref>). This version of the Cubist model is called boosted Cubist (BoostedCub). Separately, from BoostedCub, an instance-based Cubist can be used to create a k-nearest neighbor Cubist (kNCub (<xref ref-type="bibr" rid="B54">Kuhn and Quinlan, 2018</xref>)). kNCub works by first creating a tree, and averaging the prediction from the k-nearest training data points to predict the <italic>E.&#x20;coli</italic> concentration in a novel sample (<xref ref-type="bibr" rid="B54">Kuhn and Quinlan, 2018</xref>). The BoostedCub and kNCub can also be combined to generate a boosted, k-nearest neighbor Cubist (Boosted kNCub (<xref ref-type="bibr" rid="B54">Kuhn and Quinlan, 2018</xref>)). For all Cubist models, the number of rules included in the final model was tuned. The number of trees used was tuned for the boosted Cubist models, and the number of neighbors used was tuned for the instance-based Cubist models.</p>
</sec>
<sec id="s2-3-9">
<title>Support Vector Machines</title>
<p>Four variations of support vector machines (SVM) were implemented using the e1071 package (<xref ref-type="bibr" rid="B63">Meyer et&#x20;al., 2020</xref>). Each of the variations used a different kernel transformation. Each kernel mapped the data to higher or lower dimensional space, and the number of hyperparameters tuned reflects the dimensionality of the kernel. In order of most to least dimensionality, the kernels used were polynomial, sigmoidal, and radial; a linear kernel was also considered. Regardless of the kernel used, a penalty parameter that controls the smoothness of the hyperplane&#x2019;s decision boundary was tuned. For all SVMs built using non-linear kernels, a parameter was tuned that determines how close a sample needs to be to the hyperplane to influence it. For the sigmoid and polynomial SVMs, a parameter that allows the hyperplane to be nonsymmetrical was tuned. For the polynomial SVM, the degree of the polynomial function was&#x20;tuned.</p>
</sec>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>Results and Discussion</title>
<p>One-hundred twenty-five models were developed to predict <italic>E.&#x20;coli</italic> levels in Upstate New York streams used for agricultural purposes (e.g., produce irrigation; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>; 26 full models &#x2b;90 nested models &#x2b;9 baseline models). Full models were built using all four feature types, while nested models were built using between one and four feature types. The feature types considered were 1) geospatial, 2) physicochemical water quality and temperature data collected on-site, 3) all other weather data, and 4) stream traits observable on-site (e.g., stream bottom composition). Baseline models were either log-linear or featureless regression models.</p>
<p>The log<sub>10</sub> MPN of <italic>E.&#x20;coli</italic> per 100&#xa0;ml was similarly distributed in the training (1st quartile &#x3d; 1.95; median &#x3d; 2.33; 3rd quartile &#x3d; 2.73) and test data (1st quartile &#x3d; 1.90; median &#x3d; 2.21; 3rd quartile &#x3d; 2.54). While an advantage of this study is the use of two independently collected datasets to separately train and test the models, the size of each dataset (N &#x3d; 194 and 181 samples in the training and test data, respectively), as well as the temporal and geographic range represented (one growing season per dataset, one produce-growing region), are a limitation. However, this study&#x2019;s aim was not to develop field-ready models; instead, this study provides a conceptual framework for how field-ready models can be built once multi-region and multi-year datasets are available.</p>
<p>By using a continuous outcome, the present study complements a recent publication that focused on binary, categorical outcomes [detection/non-detection of enteric pathogens (<xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>)]. To the authors&#x2019; knowledge, this is also the first study to compare the performance of models for predicting <italic>E.&#x20;coli</italic> levels in agricultural water that were built using different feature types (i.e.,&#x20;geospatial, physicochemical water quality features, stream traits, and weather). Since the skill, capital, time, and computational power required to collect data on each feature type varies, the findings presented here will help future studies optimize data collection by focusing on key predictors (although other predictors may be important in other produce-growing regions). This in turn will help ensure that field-ready models developed as part of these future studies do not require growers to invest substantial time and money collecting multiple data types. For similar reasons (e.g., accessibility to growers, practicality for incorporating into on-farm management plans), future studies aimed at developing deployable models may want to consider the degree of feature engineering performed, however, such considerations were outside the scope of the present&#x20;study.</p>
<sec id="s3-1">
<title>Trade-Offs Between Interpretability and Accuracy Need to be Considered When Selecting the Algorithm Used for Model Development</title>
<p>Model performance varied considerably with root-mean-squared errors (RMSE), Kendall&#x2019;s Tau (<italic>&#x3c4;</italic>), and <italic>R</italic>
<sup>2</sup> ranging between 0.37 and 1.03, 0.07, and 0.55, and &#x2212;0.31 and 0.47, respectively, (<xref ref-type="sec" rid="s9">Supplementary Table S2</xref>; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>). The top-performing full models all performed comparably and were built using either boosted or bagged algorithms. In order, the top-performing models were: Boosted kNCub (RMSE &#x3d; 0.37); xgBoost (RMSE &#x3d; 0.37); condRF (RMSE &#x3d; 0.38); random forest (RMSE &#x3d; 0.38); regRF (RMSE &#x3d; 0.38); and exTree (RMSE &#x3d; 0.39; <xref ref-type="sec" rid="s9">Supplementary Table S2</xref>, <xref ref-type="fig" rid="F2">Figures 2</xref>&#x2013;<xref ref-type="fig" rid="F4">4</xref>). These full models outperformed the top-ranked nested model, which was built using the PLS algorithm and water quality, weather, and stream trait factors (RMSE &#x3d; 0.69; <xref ref-type="sec" rid="s9">Supplementary Table S2</xref>, <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>). Below the cluster of best performing models in the top-left corner of <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>, there is a second cluster of models that performed well but not as well as the top-ranked models. It is interesting to note that the best-performing model in this second cluster was also built using a boosted algorithm (Boosted Cubist). Moreover, nested models built using ensemble algorithms generally outperformed those built using a tree or instance-based algorithm (<xref ref-type="sec" rid="s9">Supplementary Figure S3</xref>). Overall, ensemble (boosted or bagged) algorithms substantially outperformed the other algorithms considered here. This is consistent with findings from a similar study (<xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>), which also found that ensemble models outperformed models built using alternative algorithms when predicting enteric pathogen presence in streams used to source irrigation water. These findings are also consistent with past studies that used ensemble methods (e.g., condRF, RF) to develop accurate models for predicting microbial contamination of recreational waters (<xref ref-type="bibr" rid="B38">Golden et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B99">Zhang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B68">Munck et&#x20;al., 2020</xref>) and agricultural environments (<xref ref-type="bibr" rid="B38">Golden et&#x20;al., 2019</xref>).</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Graph of <italic>R</italic>
<sup>2</sup>, which measures the variance in <italic>E.&#x20;coli</italic> levels that is predictable using the given model, vs. Kendall&#x2019;s tau, which measures the model&#x2019;s ability to predict relative <italic>E.&#x20;coli</italic> concentration. The dashed line represents the <italic>R</italic>
<sup>2</sup> for the featureless regression model; an <italic>R</italic>
<sup>2</sup> to the left of this line indicates that the model was unable to predict variability in <italic>E.&#x20;coli</italic> levels. To facilitate readability, nested models are displayed in a separate facet from the full and log-linear models. Better performing models are in the top right of each facet, while poor performers are in the bottom&#x20;left.</p>
</caption>
<graphic xlink:href="frai-04-628441-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Density plots and split quantiles plots showing the performance of the top-ranked full, nested, and log-linear models. The density plot shows the ability of each model to predict <italic>E.&#x20;coli</italic> counts in the test data, while the split quantiles plots show the ability of each model to predict when <italic>E.&#x20;coli</italic> levels are likely to be high or low (i.e.,&#x20;relative <italic>E.&#x20;coli</italic> concentration). The split quantiles plot is generated by sorting the test data from (i) lowest to highest predicted <italic>E.&#x20;coli</italic> concentration, and (ii) lowest to highest observed <italic>E.&#x20;coli</italic> concentration. The test data is then divided into quintiles based on the percentile the predicted value (color-coding; see legend) and observed values (<italic>x</italic>-axis goes from quintile with lowest observed <italic>E.&#x20;coli</italic> levels on the left to highest on the right) fell into. In a good model, all samples that are predicted to have a low <italic>E.&#x20;coli</italic> concentration (red) would be in the far left column, while samples that are predicted to have a high <italic>E.&#x20;coli</italic> concentration (blue) would be in the far right column.</p>
</caption>
<graphic xlink:href="frai-04-628441-g004.tif"/>
</fig>
<p>Decision trees, which were previously proposed as candidate algorithms for developing interpretable, food safety decision-support tools (<xref ref-type="bibr" rid="B59">Magee, 2019</xref>), performed poorly here and in the aforementioned enteric pathogens study (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>; <xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>). The poor performance of decision-trees is most likely due to overfitting during model training (<xref ref-type="sec" rid="s9">Supplementary Figure S1</xref>). However, the fact that the RMSE of the interpretable models (e.g., tree-based models) was generally higher than the RMSE of black-box approaches (i.e.,&#x20;less interpretable models like SVM and ensemble algorithms) (<xref ref-type="table" rid="T1">Table&#x20;1</xref>; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>), is illustrative of the trade-off between model interpretability and model performance [see (<xref ref-type="bibr" rid="B61">Meinshausen, 2010</xref>; <xref ref-type="bibr" rid="B53">Kuhn and Johnson, 2016</xref>; <xref ref-type="bibr" rid="B25">Doshi-Velez and Kim, 2017</xref>; <xref ref-type="bibr" rid="B58">Luo et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>) for more on these trade-offs]. Thus, our findings highlight the importance of weighing the need for interpretability vs. predictive accuracy before model fitting, particularly in future studies focused on developing implementable, field-ready models that growers can use for managing food safety hazards in agricultural water. When weighing these trade-offs, it is also important to consider that certain algorithms (e.g., conditional random forest) are better able to handle correlated and missing data as well as interactions between features than other algorithms (e.g., k-nearest neighbor, neural nets; <xref ref-type="table" rid="T1">Table&#x20;1</xref>). Similarly, it is important to consider whether feature selection is automatically performed as part of algorithm implementation (see <xref ref-type="table" rid="T1">Table&#x20;1</xref>). Since feature selection is performed as part of random forest implementation, they are more robust to the feature set used than neural nets or instance-based algorithms; this could explain the poor performance of the neural net, KNN, and wKNN models here. As such, random forest and similar algorithms may be able to better reflect the complexity and heterogeneity of freshwater systems particularly if feature selection will not be performed before model implementation.</p>
</sec>
<sec id="s3-2">
<title>The Measure(s) Used to Assess and Compare Model Performance Should be Determined by How the Predictive Model Will be Used, and if Actual E. coli Counts or a Relative Concentration (i.e., High Versus Low) is Needed</title>
<p>It is important to highlight that only RMSE was used in hyperparameter tuning and to identify the best performing models. While RMSE-based rankings generally matched rankings based on <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup>, some models with high RMSE (which indicates worse performance) had similar <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values to the top RMSE-ranked models (e.g., the neural net had high RMSE but <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> were similar to models with lower RMSE; <xref ref-type="sec" rid="s9">Supplementary Table S2</xref>; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>; <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>). This reflects differences in how each measure assesses performance. RMSE measures the differences between observed and predicted values, and therefore accounts for how off the prediction is from reality. As such, RMSE is a measure of how well the model can predict actual <italic>E.&#x20;coli</italic> counts. Kendall&#x2019;s <italic>&#x3c4;</italic> is a rank-based measurement that does not account for absolute differences but instead ranks the observed data in order from highest to lowest value and determines how closely the predictions from the model match this ranking (<xref ref-type="bibr" rid="B77">Rosset et&#x20;al., 2007</xref>). Thus, <italic>&#x3c4;</italic> is useful for identifying models that can predict when <italic>E.&#x20;coli</italic> levels are likely to be higher or lower (i.e.,&#x20;relative concentration) (<xref ref-type="bibr" rid="B77">Rosset et&#x20;al., 2007</xref>). The coefficient of determination (<italic>R</italic>
<sup>2</sup>) reflects the proportion of variation in the outcome that is predictable by the model. In this context, our findings suggest that the neural net model is unable to predict actual <italic>E.&#x20;coli</italic> concentrations but can correctly rank samples based on <italic>E.&#x20;coli</italic> concentration (e.g., identify when levels are likely to be elevated). As such, neural nets may be appropriate for use in applied settings where the relative concentration but not the absolute count of <italic>E.&#x20;coli</italic> is of interest (e.g., water source-specific models that are interested in deviation from baseline <italic>E.&#x20;coli</italic> levels, which could indicate a potential contamination event). Indeed, a previous study that used neural nets to predict pathogen presence in Florida irrigation water was able to achieve classification accuracies (i.e.,&#x20;classify samples as having a high or low probability of contamination) of up to 75% (<xref ref-type="bibr" rid="B75">Polat et&#x20;al., 2019</xref>). Conversely, neural nets may not be appropriate for predicting if a waterway complied with a water quality standard based on a binary <italic>E.&#x20;coli</italic> cut-off. These results illustrate the importance of carefully considering how a model will be applied (e.g., are count predictions needed or are rank predictions needed, is interpretability or predictive accuracy more important) when selecting 1) the algorithm used for model fitting, and 2) the performance measure used for model tuning and assessing model performance (e.g., RMSE, <italic>&#x3c4;</italic>,&#x20;<italic>R</italic>
<sup>2</sup>).</p>
<p>The impact of performance measure choice on model interpretation and ranking is particularly clear when we examine the log-linear models developed here. The predictive accuracy of the log-linear models varied substantially. None of the variation in <italic>E.&#x20;coli</italic> levels in the test data was predictable using the worst performing log-linear model (based on conductivity; RMSE &#x3d; 0.90; <italic>&#x3c4;</italic> &#x3d; 0.05; <italic>R</italic>
<sup>2</sup> &#x3d; 0.0), while the best-performing log-linear model (based on turbidity) was able to predict 32% of the variation in the test data (RMSE &#x3d; 0.74; <italic>&#x3c4;</italic> &#x3d; 0.44; <italic>R</italic>
<sup>2</sup> &#x3d; 0.32; <xref ref-type="sec" rid="s9">Table S2</xref>). The performance of the turbidity model developed here is comparable to turbidity-based log-linear models developed to predict <italic>E.&#x20;coli</italic> levels at Ohio swimming beaches (<italic>R</italic>
<sup>2</sup> &#x3d; 38% in (<xref ref-type="bibr" rid="B34">Francy and Darner, 2006</xref>); <italic>R</italic>
<sup>2</sup> ranged between 19 and 56% in (<xref ref-type="bibr" rid="B35">Francy et&#x20;al., 2013</xref>)). However, the RMSE of the turbidity model developed here was substantially worse than all full models except the wKNN and KNN models. Conversely, the <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values of the turbidity model were comparable to or better than 11 of the full models (<xref ref-type="sec" rid="s9">Supplementary Table S2</xref>). The only models to have substantially better <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values than the turbidity model were models built using an ensemble algorithm (e.g., Boosted kNCub, regRF, RF, xgBoost; <xref ref-type="sec" rid="s9">Supplementary Table S2</xref>). This suggests that the ability of the turbidity log-linear model to categorize the test data based on relative <italic>E.&#x20;coli</italic> concentration (e.g., into samples with high or low predicted <italic>E.&#x20;coli</italic> levels) was comparable to most full models. However, unlike the full models, the turbidity log-linear model could not predict actual <italic>E.&#x20;coli</italic> concentrations in the test data samples. In fact, the density plots in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> graphically show how the top-ranked full model was substantially better at predicting <italic>E.&#x20;coli</italic> counts compared to the top-ranked nested model and the turbidity log-linear model. Conversely, the split quantiles plots show that all three models were able to predict the relative concentration of <italic>E.&#x20;coli</italic> in the test data samples (<xref ref-type="fig" rid="F4">Figure&#x20;4</xref>). Overall, these findings reiterate the importance of determining how models will be applied in the field when designing a study. For example, if the aim is to develop an interpretable model to supplement ongoing monitoring efforts, a log-linear model based on turbidity could be useful for determining when <italic>E.&#x20;coli</italic> concentration most likely deviates from baseline levels (e.g., are expected to be higher or lower). Such a model would be most useful if a baseline level of <italic>E.&#x20;coli</italic> had been established for a given water source. However, separate models would need to be developed to establish this baseline for each water source, and the development of source-specific models could present an economic hurdle to small growers. As such, an ensemble model, like the full Boosted kNCub or XgBoost models developed here, would be more appropriate if 1) a generalized model (i.e.,&#x20;not specific to an individual water source): is needed, or 2) the model output needs to be an actual <italic>E.&#x20;coli</italic>&#x20;count.</p>
</sec>
<sec id="s3-3">
<title>Accurate Predictions for Top-Ranked Models were Driven by Turbidity and Weather</title>
<p>Among the nested models, models built using physicochemical and weather predictors consistently outperformed models built using geospatial predictors (<xref ref-type="sec" rid="s9">Supplementary Table S2</xref>; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>; <xref ref-type="sec" rid="s9">Supplementary Figure S2</xref>). Indeed, by creating a convex hull graph that groups nested models by predictor type, the substantial differences in model performance due to predictor type are evident (<xref ref-type="sec" rid="s9">Supplementary Figure S2</xref>). For example, all nested models built using only geospatial predictors or stream traits (e.g., stream bottom substrate) clustered in the bottom right of <xref ref-type="sec" rid="s9">Supplementary Figures S2</xref> (high RMSE, low <italic>&#x3c4;</italic> indicating poor performance). For eight of the nine algorithms used to build the nested models, the full models had substantially lower RMSE values than the nested models, while nested models built using physicochemical water quality and/or weather features, on average, had substantially lower RMSE values than the geospatial nested models (<xref ref-type="sec" rid="s9">Supplementary Table S2</xref>; <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>). However, since none of the nested models had an RMSE lower than the featureless regression, this indicates that all feature types (physiochemical, weather, geospatial, and stream traits) were needed to develop models that could accurately predict <italic>E.&#x20;coli</italic> counts. That being said, many of the nested models had substantially higher <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values than the featureless regression, indicating that they were able to accurately predict relative <italic>E.&#x20;coli</italic> concentration (i.e.,&#x20;if it was higher or lower). The pattern observed for RMSE holds true for the <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values, with physicochemical water quality models and weather models outperforming geospatial models. In fact, for several of the algorithms used for building the nested models, <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values for the physiochemical and/or weather models were higher than <italic>&#x3c4;</italic> and <italic>R</italic>
<sup>2</sup> values for the full models, indicating that these models were better able to predict relative <italic>E.&#x20;coli</italic> concentration than the full model. Based on permutation variable importance, the top-ranked full models&#x2019; ability to predict <italic>E.&#x20;coli</italic> levels in both the training and test data was driven by air temperature, rainfall, and turbidity (<xref ref-type="fig" rid="F5">Figure&#x20;5</xref>; <xref ref-type="sec" rid="s9">Supplementary Figures S4, S5</xref>). Similarly, the top-ranked nested models&#x2019; ability to predict <italic>E.&#x20;coli</italic> levels in the training and test data was driven by air temperature, rainfall, and turbidity, and by air temperature, solar radiation, and turbidity, respectively (<xref ref-type="fig" rid="F5">Figure&#x20;5</xref>; <xref ref-type="sec" rid="s9">Supplementary Figures S6&#x2013;S8</xref>). Overall, these findings reiterate that appropriate features to use when training models for predicting <italic>E.&#x20;coli</italic> levels in agricultural water source is dependent on how the model will be applied (i.e.,&#x20;if <italic>E.&#x20;coli</italic> counts or relative concentration is needed). However, we can also conclude that regardless of how the model will be applied, physicochemical water quality and weather factors should be included as features and that geospatial features should not be used alone for model development. However, water quality is known to vary spatially (e.g., between produce-growing regions). Since this study was only conducted in one region, separate region-specific models or a single multi-region model may be needed.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Permutation variable importance (PVI) of the 30 factors that were most strongly associated with predicting <italic>E.&#x20;coli</italic> levels in the test and training data using full, boosted, k-nearest neighbor Cubist model. The black dot shows median importance, while the line shows the upper and lower 5% and 95% quantiles of PVI values from the 150 permutations performed. Avg. Sol Rad &#x3d; Average Solar Radiation; Elev &#x3d; Elevation; FP &#x3d; Floodplain; SPDES &#x3d; Wastewater Discharge Site; Soil A&#x20;&#x3d;&#x20;Hydrologic Soil Type-A.</p>
</caption>
<graphic xlink:href="frai-04-628441-g005.tif"/>
</fig>
<p>The identification of associations between microbial water quality, and physicochemical water quality, and weather features are consistent with the scientific literature (<xref ref-type="bibr" rid="B35">Francy et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B10">Bradshaw et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B55">Lawrence, 2012</xref>; <xref ref-type="bibr" rid="B76">Rao et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B69">Nagels et&#x20;al., 2002</xref>; <xref ref-type="bibr" rid="B56">Liang et&#x20;al., 2015</xref>). More specifically, the strong association between turbidity and <italic>E.&#x20;coli</italic> levels, and rainfall and <italic>E.&#x20;coli</italic> levels has been reported by studies conducted in multiple water types (e.g., streams, canals, recreational water, irrigation water, water in cattle troughs), regions (e.g., Northeast, Southeast, Southwest), and years, indicating that these relationships are reproducible even under varying conditions and when different study designs are used (<xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>; <xref ref-type="bibr" rid="B12">Brady et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B11">Brady and Plona, 2009</xref>; <xref ref-type="bibr" rid="B34">Francy and Darner, 2006</xref>; <xref ref-type="bibr" rid="B55">Lawrence, 2012</xref>; <xref ref-type="bibr" rid="B82">Smith et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B22">Davies-Colley et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B71">Olyphant et&#x20;al., 2003</xref>; <xref ref-type="bibr" rid="B66">Money et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B19">Coulliette et&#x20;al., 2009</xref>). For example, a study that sampled the Chattahoochee River, a recreational waterway in Georgia, United&#x20;States of America, found that 78% of the variability in <italic>E.&#x20;coli</italic> levels could be explained by a model that included log<sub>10</sub> turbidity, flow event (i.e.,&#x20;base vs. stormflow), and season (<xref ref-type="bibr" rid="B55">Lawrence, 2012</xref>). In fact, the Georgia study found that for each log<sub>10</sub> increase in turbidity <italic>E.&#x20;coli</italic> levels increased by approx. 0.3 and approx. 0.8 log<sub>10</sub> MPN/100-ml under baseflow and stormflow conditions, respectively (<xref ref-type="bibr" rid="B55">Lawrence, 2012</xref>). Similarly, in the study reported here, accumulated local effects plots indicate the presence of a strong, positive association between <italic>E.&#x20;coli</italic> levels, and air temperature, rainfall, and turbidity (<xref ref-type="fig" rid="F6">Figure&#x20;6</xref>). The fact that the <italic>E. coli</italic>-rainfall and <italic>E. coli</italic>-turbidity relationships are reproducible across studies, regions, and water types makes sense when viewed through the lens of bacterial fate and transport. Both rainfall and turbidity are associated with conditions that facilitate bacterial movement into and within freshwater systems (<xref ref-type="bibr" rid="B69">Nagels et&#x20;al., 2002</xref>; <xref ref-type="bibr" rid="B67">Muirhead et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B50">Jamieson et&#x20;al.</xref>,; <xref ref-type="bibr" rid="B26">Drummond et&#x20;al., 2014</xref>). As such, it is not surprising that past studies that developed models to predict pathogen presence in agricultural water (<xref ref-type="bibr" rid="B75">Polat et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>) or <italic>E.&#x20;coli</italic> concentrations in recreational water (e.g., (<xref ref-type="bibr" rid="B34">Francy and Darner, 2006</xref>; <xref ref-type="bibr" rid="B12">Brady et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B11">Brady and Plona, 2009</xref>)), found that models built using turbidity and/or rainfall outperformed models built using other factors. For example, Polat et&#x20;al. (<xref ref-type="bibr" rid="B75">Polat et&#x20;al., 2019</xref>) found that models that included turbidity as a feature were between 6 and 15% more accurate at predicting <italic>Salmonella</italic> presence in Florida irrigation ponds than models built using other predictors. Similarly, a study that used multivariable regression to predict <italic>E.&#x20;coli</italic> level at Ohio beaches found that only rainfall and turbidity were retained in the final, best-performing model (<xref ref-type="bibr" rid="B34">Francy and Darner, 2006</xref>). Overall, the findings of this and other studies suggest that future data collection efforts (to generate data that can be used to train predictive <italic>E.&#x20;coli</italic> models) should focus on physicochemical water quality and weather as opposed to geospatial factors.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Accumulated local effects plots showing the effect of the four factors with the highest PVI when predictions were made on the training (shown in red) and test (shown in blue) data using the full boosted, k-nearest neighbor Cubist model. All predictors were centered and scaled before training each model. As a result, the units on the <italic>x</italic>-axis are the number of standard deviations above or below the mean for the given factor (e.g., in the rainfall plot 0 and two indicate the mean rainfall 0&#x2013;1&#xa0;day before sample are collection [BSC], and 2 standard deviations above this mean, respectively).</p>
</caption>
<graphic xlink:href="frai-04-628441-g006.tif"/>
</fig>
<p>While the scientific literature supports focusing future research efforts on collecting physicochemical water quality and weather data for models aimed at predicting <italic>E.&#x20;coli</italic> presence or levels in the water, this recommendation is also supported by economic and computational feasibility. It is relatively easy and inexpensive for growers to obtain and download weather data from nearby extension-run weather stations since many growers already use these websites (e.g., NEWA [<ext-link ext-link-type="uri" xlink:href="http://newa.cornell.edu/">newa.cornell.edu</ext-link>], WeatherSTEM [<ext-link ext-link-type="uri" xlink:href="http://www.weatherstem.com">www.weatherstem.com</ext-link>]) since this data is freely available. It can also be relatively inexpensive to collect turbidity and other physicochemical water quality data depending on the required precision of these measurements. Conversely, geospatial data requires either that: 1) the grower has access to software and training that allows them to extract geospatial data from government databases and calculate relevant statistics for each water source on their farm (e.g., the proportion of upstream watershed under natural cover), 2) an external group, such as consultants or universities working with industry perform this task, or 3) an external group develops a software program to perform this task. All three options would require substantial computational power, time, training, and capital.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s4">
<title>Conclusion</title>
<p>This study demonstrates that predictive models can be used to predict both relative (i.e.,&#x20;high vs. low) and absolute (i.e.,&#x20;counts) levels of <italic>E.&#x20;coli</italic> in agricultural water in New York. More specifically, the findings reported here confirm previous studies&#x2019; conclusions that machine learning models may be useful for predicting when, where, and at what level fecal contamination (and associated food safety hazards) is likely to be in agricultural water sources (<xref ref-type="bibr" rid="B75">Polat et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B95">Weller et&#x20;al., 2020a</xref>). This study also identifies specific algorithm-feature combinations (i.e.,&#x20;forest algorithms, and physicochemical water quality and weather features) that should be the foci of future efforts to develop deployable models that can guide on-farm decision-making. This study also highlights that the approach used to develop these field-ready models (i.e.,&#x20;the algorithm, performance measure, and predictors used) should be in how the model will be applied. For example, while ensemble methods can predict <italic>E.&#x20;coli</italic> counts, interpretable (i.e.,&#x20;non-black-box methods like the baseline log-linear models) cannot. Conversely, these interpretable models were able to predict when <italic>E.&#x20;coli</italic> levels are above or below a baseline. Overall, this proof-of-concept study provides foundational data that can be used to guide the design of future projects focused on developing field-ready models for predicting <italic>E.&#x20;coli</italic> levels in agricultural and possibly other (e.g., recreational) waterways. Moreover, this paper highlights that accurate models can be developed using weather (e.g., rain, temperature) and physicochemical water quality (e.g., turbidity) features. Future efforts may want to focus on models built using these (as opposed to geospatial) features. In adapting these findings to guide the development of deployable models, it is important to note that several studies suggest that <italic>E.&#x20;coli</italic> levels was an inconsistent predictor of pathogen presence in surface water (<xref ref-type="bibr" rid="B43">Harwood et&#x20;al., 2005</xref>; <xref ref-type="bibr" rid="B60">McEgan et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B73">Pachepsky et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B2">Antaki et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B97">Weller et&#x20;al., 2020c</xref>). Instead, <italic>E.&#x20;coli</italic> models, like those developed here, may be useful for assessing fecal contamination status and for ensuring compliance with regulations but should not be used to determine if specific pathogens of concern (e.g., <italic>Salmonella</italic>, <italic>Listeria</italic>) are present.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>De-identified data (e.g., excluding GPS coordinates) are available on request. Requests to access these datasets should be directed to Daniel Weller, wellerd2@gmail.com/dlw263@cornell.edu.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>DW and MW conceived of the project idea, designed the study, and wrote the grant to fund the research. DW oversaw the day-to-day aspects of data collection and led data collection and cleaning efforts. DW and TL developed the data analysis plan, which DW implemented. All authors contributed to manuscript development.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This project was funded by grants from the Center for Produce Safety under award number 2017CPS09 and the National Institute of Environmental Health Sciences of the National Institutes of Health (NIH) under award number T32ES007271. The content is solely the responsibility of the authors and does not represent the official views of the NIH, or any other United&#x20;States government agency.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<ack>
<p>We are grateful for the technical assistance of Alex Belias, Sherry Roof, Maureen Gunderson, Aziza Taylor, Kyle Markwadt, Sriya Sunil, Ahmed Gaballa, Kayla Ferris, and Julia Muuse. We would also like to thank Laura Strawn for her comments on the manuscript.</p>
</ack>
<sec id="s9">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2021.628441/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frai.2021.628441/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.docx" id="SM1" mimetype="application/docx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ackers</surname>
<given-names>M. L.</given-names>
</name>
<name>
<surname>Mahon</surname>
<given-names>B. E.</given-names>
</name>
<name>
<surname>Leahy</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Goode</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Damrow</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hayes</surname>
<given-names>P. S.</given-names>
</name>
<etal/>
</person-group> (<year>1998</year>). <article-title>An outbreak of Escherichia coli O157:H7 infections associated with leaf lettuce consumption</article-title>. <source>J.&#x20;Infect. Dis.</source> <volume>177</volume>:<fpage>1588</fpage>&#x2013;<lpage>1593</lpage>. <pub-id pub-id-type="doi">10.1086/515323</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Antaki</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Vellidis</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Harris</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Aminabadi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Levy</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Jay-Russell</surname>
<given-names>M. T.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Low concentration of <italic>Salmonella enterica</italic> and generic <italic>Escherichia coli</italic> in farm ponds and irrigation distribution systems used for mixed produce production in southern Georgia</article-title>. <source>Foodborne Pathog. Dis.</source> <volume>13</volume>, <fpage>551</fpage>&#x2013;<lpage>558</lpage>. <pub-id pub-id-type="doi">10.1089/fpd.2016.2117</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<collab>ANZECC</collab> (<year>2000</year>). <source>Guidelines for Fresh and marine water quality</source>. <edition>1st Edn</edition>. <publisher-loc>Auckland, New&#x20;Zealand, Australia and New&#x20;Zealand</publisher-loc>: <publisher-name>Australian and New&#x20;Zealand environment and conservation council, agriculture and resource management council of Australia and New&#x20;Zealand</publisher-name>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Apley</surname>
<given-names>D. W.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Visualizing the effects of predictor variables in black box supervised learning models</article-title>. <source>J.&#x20;R. Stat. Soc. Ser. B Stat. Methodol.</source> <volume>82</volume> (<issue>4</issue>), <fpage>1059</fpage>&#x2013;<lpage>1086</lpage>. <pub-id pub-id-type="doi">10.1111/rssb.12377</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Astill</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Minor</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Calvin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Thornsbury</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Before implementation of the food safety modernization act&#x2019;s produce rule: A Survey of U.S. Produce Grower</article-title>. <source>Eco. Inform. Bull.</source> <volume>194</volume> (<issue>1&#x2013;84</issue>). </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barton Behravesh</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Mody</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>Jungk</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gaul</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Redd</surname>
<given-names>J.&#x20;T.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>2008 outbreak of Salmonella Saintpaul infections associated with raw produce</article-title>. <source>N. Engl. J.&#x20;Med.</source> <volume>364</volume>:<fpage>918</fpage>&#x2013;<lpage>927</lpage>. <pub-id pub-id-type="doi">10.1056/nejmoa1005741</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Benjamin</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Jay-Russell</surname>
<given-names>M. T.</given-names>
</name>
<name>
<surname>Atwill</surname>
<given-names>E. R.</given-names>
</name>
<name>
<surname>Cooley</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Carychao</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Larsen</surname>
<given-names>R. E.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Risk factors for Escherichia coli O157 on beef cattle ranches located near a major produce production region</article-title>. <source>Epidemiol. Infect.</source> <volume>143</volume>, <fpage>81</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1017/s0950268814000521</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bischl</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kotthoff</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Schiffner</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Richter</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Studerus</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>mlr: machine learning in R</article-title>. <source>J.&#x20;Mach Learn. Res.</source> <volume>17</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bottichio</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Keaton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Thomas</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Fulton</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tiffany</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Frick</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Shiga toxin&#x2013;producing <italic>Escherichia coli</italic> infections associated with romaine lettuce&#x2014;United&#x20;States, 2018</article-title>. <source>Clin. Infect. Dis.</source> <volume>71</volume>, <fpage>e323</fpage>&#x2013;<lpage>e330</lpage>. <pub-id pub-id-type="doi">10.1093/cid/ciz1182</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bradshaw</surname>
<given-names>J.&#x20;K.</given-names>
</name>
<name>
<surname>Snyder</surname>
<given-names>B. J.</given-names>
</name>
<name>
<surname>Oladeinde</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Spidle</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Berrang</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Meinersmann</surname>
<given-names>R. J.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Characterizing relationships among fecal indicator bacteria, microbial source tracking markers, and associated waterborne pathogen occurrence in stream water and sediments in a mixed land use watershed</article-title>. <source>Water Res.</source> <volume>101</volume>, <fpage>498</fpage>&#x2013;<lpage>509</lpage>. <pub-id pub-id-type="doi">10.1016/j.watres.2016.05.014</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brady</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Plona</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Relations between environmental and water-quality variables and Escherichia coli in the cuyahoga river with emphasis on turbidity as a predictor of recreational water quality, cuyahoga valley national park, Ohio, 2008</article-title>. <publisher-loc>Columbis, OH</publisher-loc>: <publisher-name>USGS</publisher-name>, <comment>Open-File Report 2009-1192</comment>. </citation>
</ref>
<ref id="B12">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brady</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bushon</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Plona</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Predicting recreational water quality using turbidity in the cuyahoga river, cuyahoga valley national park, Ohio, 2004&#x2013;7</article-title>. <publisher-loc>Columbus, OH</publisher-loc>: <publisher-name>USGS Scientific Investigations Report</publisher-name>, <comment>2009&#x2013;5192</comment>. </citation>
</ref>
<ref id="B13">
<citation citation-type="other">
<person-group person-group-type="author">
<name>
<surname>Brownlee</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <comment>Package &#x201c;xgboost&#x201d; type package Title extreme gradient boosting</comment>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<collab>California Leafy Greens Marketing Agreement</collab> (<year>2017</year>). <source>Commodity specific food safety Guidelines for the production and harvest of lettuce and leafy greens</source>. <publisher-loc>Sacramento, CA, California, United&#x20;States</publisher-loc>: <publisher-name>California Leafy Green Handler Marketing Board</publisher-name>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Calvin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jensen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Klonsky</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Cook</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Food safety practices and costs under the California leafy greens marketing agreementt</article-title>, <comment>EIB-173</comment>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ceuppens</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Johannessen</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Allende</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tondo</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>El-Tahan</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sampers</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Risk factors for <italic>Salmonella</italic>, Shiga toxin-producing <italic>Escherichia coli</italic> and <italic>Campylobacter</italic> occurrence in primary production of leafy greens and strawberries</article-title>. <source>Ijerph</source> <volume>12</volume>, <fpage>9809</fpage>&#x2013;<lpage>9831</lpage>. <pub-id pub-id-type="doi">10.3390/ijerph120809809</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Guestrin</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>XGBoost: a scalable tree boosting system</article-title>&#x201d;. in <conf-name>Proceedings of the ACM SIGKDD international conference on knowledge discovery and data mining</conf-name>, <conf-loc>New York, New York, USA</conf-loc> (<publisher-name>Association for Computing Machinery</publisher-name>), <fpage>785</fpage>&#x2013;<lpage>794</lpage>. </citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Corona</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Las</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Belem</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ruiz</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Beltran</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Killeen</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <source>Commodity specific food safety Guidelines for the production, harvest, post-harvest, and value-added unit operations of green onions</source>. <publisher-loc>Washington, D.C., USA</publisher-loc>: <publisher-name>US Food and Drug Administration</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Coulliette</surname>
<given-names>A. D.</given-names>
</name>
<name>
<surname>Money</surname>
<given-names>E. S.</given-names>
</name>
<name>
<surname>Serre</surname>
<given-names>M. L.</given-names>
</name>
<name>
<surname>Noble</surname>
<given-names>R. T.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Space/time analysis of fecal pollution and rainfall in an eastern North Carolina estuary</article-title>. <source>Environ. Sci. Technol.</source> <volume>43</volume>, <fpage>3728</fpage>&#x2013;<lpage>3735</lpage>. <pub-id pub-id-type="doi">10.1021/es803183f</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dada</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Hamilton</surname>
<given-names>D. P.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Predictive models for determination of <italic>E.&#x20;coli</italic> concentrations at inland recreational beaches</article-title>. <source>Water Air Soil Pollut.</source> <volume>227</volume>. <pub-id pub-id-type="doi">10.1007/s11270-016-3033-6</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dada</surname>
<given-names>C. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Seeing is predicting: water clarity-based nowcast models for <italic>E.&#x20;coli</italic> prediction in surface water</article-title>. <source>Gjhs</source> <volume>11</volume>, <fpage>140</fpage>. <pub-id pub-id-type="doi">10.5539/gjhs.v11n3p140</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davies-Colley</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Valois</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Milne</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Faecal contamination and visual clarity in New&#x20;Zealand rivers: correlation of key variables affecting swimming suitability</article-title>. <source>J.&#x20;Water Health</source> <volume>16</volume>:<fpage>329</fpage>&#x2013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.2166/wh.2018.214</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Runger</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Feature selection via regularized trees</article-title>&#x201d; in <conf-name>Proceedings of the international joint conference on neural networks</conf-name> (<publisher-loc>Brisbane, Australia</publisher-loc>: <publisher-name>World Congress on Computational Intelligence</publisher-name>). <pub-id pub-id-type="doi">10.1109/IJCNN.2012.6252640</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Runger</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Gene selection with guided regularized random forest</article-title>. <source>Pattern Recognition</source> <volume>46</volume>, <fpage>3483</fpage>&#x2013;<lpage>3489</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2013.05.018</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doshi-Velez</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Towards A rigorous science of interpretable machine learning</article-title>. <comment>
<italic>arXiv</italic> [Preprint]. Available at: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://arxiv.org/abs/1702.08608">https://arxiv.org/abs/1702.08608</ext-link>.</comment>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Drummond</surname>
<given-names>J.&#x20;D.</given-names>
</name>
<name>
<surname>Davies-Colley</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Stott</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sukias</surname>
<given-names>J.&#x20;P.</given-names>
</name>
<name>
<surname>Nagels</surname>
<given-names>J.&#x20;W.</given-names>
</name>
<name>
<surname>Sharp</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Retention and remobilization dynamics of fine particles and microorganisms in pastoral streams</article-title>. <source>Water Res.</source> <volume>66</volume>, <fpage>459</fpage>&#x2013;<lpage>472</lpage>. <pub-id pub-id-type="doi">10.1016/j.watres.2014.08.025</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Edge</surname>
<given-names>T. A.</given-names>
</name>
<name>
<surname>El-Shaarawi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gannon</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Jokinen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kent</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>I. U. H.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Investigation of an <italic>Escherichia coli</italic> environmental benchmark for waterborne pathogens in agricultural watersheds in Canada</article-title>. <source>J.&#x20;Environ. Qual.</source> <volume>41</volume>, <fpage>21</fpage>. <pub-id pub-id-type="doi">10.2134/jeq2010.0253</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Efstratiou</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Mavridou</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Richardson</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Prediction of Salmonella in seawater by total and faecal coliforms and Enterococci</article-title>. <source>Mar. Pollut. Bull.</source> <volume>58</volume>, <fpage>201</fpage>&#x2013;<lpage>205</lpage>. <pub-id pub-id-type="doi">10.1016/j.marpolbul.2008.10.003</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="book">
<collab>Environmental Protection Agency</collab> (<year>2012</year>). <source>Recreational water quality criteria</source>. <publisher-loc>Washington, D.C., United&#x20;States</publisher-loc>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<collab>EU Parliament</collab> (<year>2006</year>). <article-title>Bathing water quality directive</article-title>. <source>Directive 2006/7/ECOfficial Journal of the European Union</source>. </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fisher</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rudin</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Dominici</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>All models are wrong, but many are useful: learning a variable&#x2019;s importance by studying an entire class of prediction models simultaneously</article-title>. <source>J.&#x20;Mach Learn. Res.</source> <volume>20</volume>. </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<collab>Food and Drug Administration</collab>. (<year>2019</year>). <source>Investigation summary: factors potentially contributing to the contamination of romaine lettuce implicated in the fall 2018 multi-state outbreak of <italic>E.&#x20;coli</italic> O157:H7</source>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<collab>Food and Drug Administration</collab>. (<year>2020</year>). <source>Outbreak investigation of <italic>E.&#x20;coli</italic>: Romaine from Salinas</source>, <publisher-loc>California. Washington, D.C</publisher-loc>. <comment>Available at: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.fda.gov/food/outbreaks-foodborne-illness/outbreak-investigation-e-coli-romaine-salinas-california-november-2019">https://www.fda.gov/food/outbreaks-foodborne-illness/outbreak-investigation-e-coli-romaine-salinas-california-november-2019</ext-link>.</comment>
</citation>
</ref>
<ref id="B34">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Francy</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Darner</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>Procedures for developing models to predict exceedances of recreational water-quality standards at coastal beaches</article-title>,&#x201d; in <source>Techniques and methods</source> (<publisher-loc>Reston, Virginia</publisher-loc>: <publisher-name>USGS</publisher-name>). </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Francy</surname>
<given-names>D. S.</given-names>
</name>
<name>
<surname>Stelzer</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Duris</surname>
<given-names>J.&#x20;W.</given-names>
</name>
<name>
<surname>Brady</surname>
<given-names>A. M. G.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>J.&#x20;H.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>H. E.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Predictive models for <italic>Escherichia coli</italic> concentrations at inland lake beaches and relationship of model variables to pathogen detection</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>79</volume>, <fpage>1676</fpage>&#x2013;<lpage>1688</lpage>. <pub-id pub-id-type="doi">10.1128/aem.02995-12</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Francy</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Brady</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Carvin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Corsi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fuller</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Developing and implementing predictive models for estimating recreational water quality at great lakes Beaches</article-title>. <publisher-loc>Columbus, OH</publisher-loc> <comment>Scientific investigations report</comment>. </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Friedman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hastie</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tibshirani</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Regularization paths for generalized linear models via coordinate descent</article-title>. <source>J.&#x20;Stat. Softw.</source> <volume>33</volume>, <fpage>1</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v033.i01</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Golden</surname>
<given-names>C. E.</given-names>
</name>
<name>
<surname>Rothrock</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Mishra</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Comparison between random forest and gradient boosting machine methods for predicting Listeria spp. prevalence in the environment of pastured poultry farms</article-title>. <source>Food Res. Int.</source> <volume>122</volume>, <fpage>47</fpage>&#x2013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1016/j.foodres.2019.03.062</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Greene</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Daly</surname>
<given-names>E. R.</given-names>
</name>
<name>
<surname>Talbot</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Demma</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Holzbauer</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>N. J.</given-names>
</name>
<etal/>
</person-group> (<year>2008</year>). <article-title>Recurrent multistate outbreak of Salmonella Newport associated with tomatoes from contaminated fields, 2005, 2005</article-title>. <source>Epidemiol. Infect.</source> <volume>136</volume>:<fpage>157</fpage>&#x2013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1017/s095026880700859x</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grubinger</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zeileis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pfeiffer</surname>
<given-names>K-P.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Evtree: evolutionary learning of globally optimal classification and regression trees in R</article-title>. <source>J.&#x20;Stat. Softw.</source> <volume>61</volume>, <fpage>1</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v061.i01</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hamilton</surname>
<given-names>J.&#x20;L.</given-names>
</name>
<name>
<surname>Luffman</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Precipitation, Pathogens, and turbidity trends in the little river, Tennessee</article-title>. <source>Phys. Geogr.</source> <volume>30</volume>, <fpage>236</fpage>&#x2013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.2747/0272-3646.30.3.236</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harwood</surname>
<given-names>V. J.</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>A. D.</given-names>
</name>
<name>
<surname>Scott</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Chivukula</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Lukasik</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Farrah</surname>
<given-names>S. R.</given-names>
</name>
<etal/>
</person-group> (<year>2005</year>). <article-title>Validity of the indicator organism paradigm for pathogen reduction in reclaimed water and public health protection</article-title>. <source>Aem</source> <volume>71</volume>, <fpage>3163</fpage>&#x2013;<lpage>3170</lpage>. <pub-id pub-id-type="doi">10.1128/aem.71.6.3163-3170.2005</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hastie</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tibshirani</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <source>Mda: mixture and flexible discriminant analysis</source>. <comment>R Packag version 04-10</comment>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Havelaar</surname>
<given-names>A. H.</given-names>
</name>
<name>
<surname>Vazquez</surname>
<given-names>K. M.</given-names>
</name>
<name>
<surname>Topalcengiz</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Mu&#xf1;oz-Carpena</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Danyluk</surname>
<given-names>M. D.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Evaluating the U.S. Food safety modernization act produce safety rule standard for microbial quality of agricultural water for growing produce</article-title>. <source>J.&#x20;Food Prot.</source> <volume>80</volume>, <fpage>1832</fpage>&#x2013;<lpage>1841</lpage>. <pub-id pub-id-type="doi">10.4315/0362-028x.jfp-17-122</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="book">
<collab>Health Canada</collab> (<year>2012</year>). <source>Guidelines for Canadian recreational water quality</source> <edition>3rd Edn.</edition> <publisher-loc>Ottawa, Canada</publisher-loc>.</citation>
</ref>
<ref id="B47">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hechenbichler</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Schliep</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2004</year>). <source>Weighted k-nearest-neighbor techniques and ordinal Classification Discussion paper 399, SFB 386</source>, <publisher-loc>Munich</publisher-loc>: <publisher-name>Ludwig-Maximilians University</publisher-name>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hothorn</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hornik</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zeileis</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Unbiased recursive partitioning: a conditional inference framework</article-title>. <source>J.&#x20;Comput. Graphical Stat.</source> <volume>15</volume>, <fpage>651</fpage>&#x2013;<lpage>674</lpage>. <pub-id pub-id-type="doi">10.1198/106186006x133933</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hou</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Rabinovici</surname>
<given-names>S. J.&#x20;M.</given-names>
</name>
<name>
<surname>Boehm</surname>
<given-names>A. B.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Enterococci predictions from partial least squares regression models in conjunction with a single-sample standard improve the efficacy of beach management advisories</article-title>. <source>Environ. Sci. Technol.</source> <volume>40</volume>, <fpage>1737</fpage>&#x2013;<lpage>1743</lpage>. <pub-id pub-id-type="doi">10.1021/es0515250</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="other">
<person-group person-group-type="author">
<name>
<surname>Jamieson</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Joy</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kostaschuk</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gordon</surname>
<given-names>R. J.</given-names>
</name>
</person-group> <article-title>Resuspension of sediment-associated <italic>Escherichia coli</italic> in a natural stream</article-title>. <source>J.&#x20;Environ. Qual.</source> <volume>34</volume>, <fpage>581</fpage>&#x2013;<lpage>589</lpage>. <pub-id pub-id-type="doi">10.2134/jeq2005.0581</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2006</year>). <source>The ghost map: the story of london&#x2019;s most terrifying epidemic--and how it changed science, cities, and the modern world</source>, <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Penguin Books</publisher-name>.</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>King</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Baker</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Whigham</surname>
<given-names>D. F.</given-names>
</name>
<name>
<surname>Weller</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Jordan</surname>
<given-names>T. E.</given-names>
</name>
<name>
<surname>Kazyak</surname>
<given-names>P. F.</given-names>
</name>
<etal/>
</person-group> (<year>2005</year>). <article-title>Spatial considerations for linking watershed land cover to ecological indicators in streams</article-title>. <source>Ecol. Appl.</source> <volume>15</volume>, <fpage>137</fpage>&#x2013;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1890/04-0481</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kuhn</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Applied predictive modeling</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer Nature</publisher-name>.</citation>
</ref>
<ref id="B54">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kuhn</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Quinlan</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Cubist: rule- and instance-based regression modeling</source>. <comment>R Packag version 022</comment>.</citation>
</ref>
<ref id="B55">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lawrence</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2012</year>). <source>
<italic>Escherichia coli</italic> bacteria density in relation to turbidity, streamflow characteristics, and season in the Chattahoochee River near Atlanta, Georgia, October 2000 through September 2008&#x2014;description, statistical analysis, and predictive modeling</source>. <publisher-loc>Atlanta, GA</publisher-loc>: <publisher-name>U.S. Geolo</publisher-name>. </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Goh</surname>
<given-names>S. G.</given-names>
</name>
<name>
<surname>Vergara</surname>
<given-names>G. G. R. V.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>H. M.</given-names>
</name>
<name>
<surname>Rezaeinejad</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>S. Y.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Alternative fecal indicators and their empirical relationships with enteric viruses, <italic>Salmonella enterica</italic>, and <italic>Pseudomonas aeruginosa</italic> in surface waters of a tropical urban catchment</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>81</volume>, <fpage>850</fpage>&#x2013;<lpage>860</lpage>. <pub-id pub-id-type="doi">10.1128/aem.02670-14</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liaw</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Winer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wiener</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Classification and regression by randomForest</article-title>. <source>2R News</source> <volume>2</volume> (<issue>3</issue>), <fpage>18</fpage>&#x2013;<lpage>22</lpage>. </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tseng</surname>
<given-names>H.-H.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ten Haken</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>El Naqa</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Balancing accuracy and interpretability of machine learning approaches for radiation treatment outcomes modeling</article-title>. <source>BJR&#x7c;Open</source> <volume>1</volume>, <fpage>20190021</fpage>. <pub-id pub-id-type="doi">10.1259/bjro.20190021</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Magee</surname>
<given-names>J.&#x20;F.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Decision trees: Reports from the meeting breakout groups</article-title>&#x201d; in <source>Safety and quality of water used in food production and processing attributing illness caused by Shiga toxin-producing Escherichia coli (STEC) to specific foods. microbiological risk assessment series no. 33</source> (<publisher-loc>Rome, Italy</publisher-loc>), p. <fpage>25</fpage>&#x2013;<lpage>63</lpage>. </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McEgan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mootian</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Goodridge</surname>
<given-names>L. D.</given-names>
</name>
<name>
<surname>Schaffner</surname>
<given-names>D. W.</given-names>
</name>
<name>
<surname>Danyluk</surname>
<given-names>M. D.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Predicting <italic>Salmonella</italic> populations from biological, chemical, and physical indicators in Florida surface waters</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>79</volume>, <fpage>4094</fpage>&#x2013;<lpage>4105</lpage>. <pub-id pub-id-type="doi">10.1128/aem.00777-13</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meinshausen</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Node harvest</article-title>. <source>Ann. Appl. Stat.</source> <volume>4</volume>, <fpage>2049</fpage>&#x2013;<lpage>2072</lpage>. <pub-id pub-id-type="doi">10.1214/10-aoas367</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mevik</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wehrens</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Liland</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Pls: partial least squares and principal component regression</source>. <comment>R Packag version 27-2</comment>.</citation>
</ref>
<ref id="B63">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Meyer</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Dimitriadou</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Hornik</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Weingessel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Leisch</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <source>e1071: misc Functions of the Department of Statistics</source>, <source>Probability Theory Group (Formerly: e1071)</source>. <comment>1.7-4. R package</comment>.</citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Milborrow</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Derived from mda:mars by T. Hastie and R. Tibshirani</article-title>. <comment>earth: Multivariate Adaptive Regression Spines</comment>. </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Molnar</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bischl</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Casalicchio</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Iml: an R package for interpretable machine learning</article-title>. <source>Joss</source> <volume>3</volume>, <fpage>786</fpage>. <pub-id pub-id-type="doi">10.21105/joss.00786</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Money</surname>
<given-names>E. S.</given-names>
</name>
<name>
<surname>Carter</surname>
<given-names>G. P.</given-names>
</name>
<name>
<surname>Serre</surname>
<given-names>M. L.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Modern space/time geostatistics using river distances: data integration of turbidity andE.&#x20;coliMeasurements to assess fecal contamination along the raritan river in New Jersey</article-title>. <source>Environ. Sci. Technol.</source> <volume>43</volume>, <fpage>3736</fpage>&#x2013;<lpage>3742</lpage>. <pub-id pub-id-type="doi">10.1021/es803236j</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Muirhead</surname>
<given-names>R. W.</given-names>
</name>
<name>
<surname>Davies-Colley</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Donnison</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Nagels</surname>
<given-names>J.&#x20;W.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Faecal bacteria yields in artificial flood events: quantifying in-stream stores</article-title>. <source>Water Res.</source> <volume>38</volume>, <fpage>1215</fpage>&#x2013;<lpage>1224</lpage>. <pub-id pub-id-type="doi">10.1016/j.watres.2003.12.010</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Munck</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Njage</surname>
<given-names>P. M. K.</given-names>
</name>
<name>
<surname>Leekitcharoenphon</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Litrup</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Hald</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Application of whole-genome sequences and machine learning in source attribution of Salmonella typhimurium</article-title>. <source>Risk Anal.</source> <volume>40</volume>, <fpage>1693</fpage>. <pub-id pub-id-type="doi">10.1111/risa.13510</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nagels</surname>
<given-names>J.&#x20;W.</given-names>
</name>
<name>
<surname>Davies-Colley</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Donnison</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Muirhead</surname>
<given-names>R. W.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Faecal contamination over flood events in a pastoral agricultural stream in New&#x20;Zealand</article-title>. <source>Water Sci. Technol.</source> <volume>45</volume>, <fpage>45</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.2166/wst.2002.0408</pub-id> </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olyphant</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Whitman</surname>
<given-names>R. L.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Elements of a predictive model for determining beach closures on a real time basis: the case of 63rd Street Beach Chicago</article-title>. <source>Environ. Monit. Assess.</source> <volume>98</volume>, <fpage>175</fpage>&#x2013;<lpage>190</lpage>. <pub-id pub-id-type="doi">10.1023/b:emas.0000038185.79137.b9</pub-id> </citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olyphant</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Thomas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Whitman</surname>
<given-names>R. L.</given-names>
</name>
<name>
<surname>Harper</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Characterization and statistical modeling of bacterial (<italic>Escherichia coli</italic>) outflows from watersheds that discharge into southern lake Michigan</article-title>, <source>Environ. Monit. Assess.</source> <volume>81</volume>, <fpage>289</fpage>&#x2013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.1023/A:1021345512203</pub-id> </citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olyphant</surname>
<given-names>G. A.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Statistical basis for predicting the need for bacterially induced beach closures: emergence of a paradigm?</article-title> <source>Water Res.</source> <volume>39</volume>, <fpage>4953</fpage>&#x2013;<lpage>4960</lpage>. <pub-id pub-id-type="doi">10.1016/j.watres.2005.09.031</pub-id> </citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pachepsky</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Shelton</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Dorner</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Whelan</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Can <italic>E.&#x20;coli</italic> or thermotolerant coliform concentrations predict pathogen presence or prevalence in irrigation waters?</article-title> <source>Crit. Rev. Microbiol.</source> <volume>42</volume> (<issue>3</issue>), <fpage>384</fpage>&#x2013;<lpage>393</lpage>. <pub-id pub-id-type="doi">10.3109/1040841x.2014.954524</pub-id> </citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>McEgan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mishra</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Micallef</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Pradhan</surname>
<given-names>A. K.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Identifying and modeling meteorological risk factors associated with pre-harvest contamination of <italic>Listeria</italic> species in a mixed produce and dairy farm</article-title>. <source>Food Res. Int.</source> <volume>102</volume>, <fpage>355</fpage>&#x2013;<lpage>363</lpage>. <pub-id pub-id-type="doi">10.1016/j.foodres.2017.09.029</pub-id> </citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Polat</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Topalcengiz</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Danyluk</surname>
<given-names>M. D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Prediction of <italic>Salmonella</italic> presence and absence in agricultural surface waters by artificial intelligence approaches</article-title>. <source>J.&#x20;Food Saf.</source> <volume>40</volume>, <fpage>e12733</fpage>. <pub-id pub-id-type="doi">10.1111/jfs.12733</pub-id> </citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rao</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Eisenberg</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kleinbaum</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Cevallos</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Trueba</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Levy</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Spatial variability of <italic>Escherichia coli</italic> in rivers of northern coastal Ecuador</article-title>. <source>Water</source> <volume>7</volume>, <fpage>818</fpage>&#x2013;<lpage>832</lpage>. <pub-id pub-id-type="doi">10.3390/w7020818</pub-id> </citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rosset</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Perlich</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zadrozny</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2007</year>). <source>Ranking-based evaluation of regression models</source>. <source>Knowledge Inform. Syst.</source> <volume>12</volume>, <fpage>331</fpage>&#x2013;<lpage>353</lpage>. <pub-id pub-id-type="doi">10.1109/ICDM.2005.126</pub-id> </citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rossi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wolde</surname>
<given-names>B. T.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>L. H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Prediction of recreational water safety using <italic>Escherichia coli</italic> as an indicator: case study of the Passaic and Pompton rivers, New Jersey</article-title>. <source>Sci. Total Environ.</source> <volume>714</volume>, <fpage>136814</fpage>. <pub-id pub-id-type="doi">10.1016/j.scitotenv.2020.136814</pub-id> </citation>
</ref>
<ref id="B79">
<citation citation-type="book">
<collab>SA DWAF</collab> (<year>1996</year>). <source>Water quality Guidelines</source>. <publisher-loc>South Africa</publisher-loc>.</citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shiels</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Guebert</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Implementing landscape indices to predict stream water quality in an agricultural setting: an assessment of the Lake and River Enhancement (LARE) protocol in the Mississinewa River watershed, East-Central Indiana</article-title>. <source>Ecol. Indicators</source> <volume>10</volume>, <fpage>1102</fpage>&#x2013;<lpage>1110</lpage>. <pub-id pub-id-type="doi">10.1016/j.ecolind.2010.03.007</pub-id> </citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Simm</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Magrans De Abril</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Sugiyama</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Tree-based ensemble multi-task learning method for classification and regression</article-title>. <source>IEICE Trans. Inf. Syst.</source> <volume>E97.D</volume>, <fpage>1677</fpage>&#x2013;<lpage>1681</lpage>. <pub-id pub-id-type="doi">10.1587/transinf.e97.d.1677</pub-id> </citation>
</ref>
<ref id="B82">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Smith</surname>
<given-names>R. P.</given-names>
</name>
<name>
<surname>Paiba</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Ellis-Iversen</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Short communication: turbidity as an indicator of <italic>Escherichia coli</italic> presence in water troughs on cattle farms</article-title>. <source>J.&#x20;Dairy Sci.</source> <volume>91</volume>, <fpage>2082</fpage>&#x2013;<lpage>2085</lpage>. <pub-id pub-id-type="doi">10.3168/jds.2007-0597</pub-id> </citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strawn</surname>
<given-names>L. K.</given-names>
</name>
<name>
<surname>Fortes</surname>
<given-names>E. D.</given-names>
</name>
<name>
<surname>Bihn</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Nightingale</surname>
<given-names>K. K.</given-names>
</name>
<name>
<surname>Gr&#xf6;hn</surname>
<given-names>Y. T.</given-names>
</name>
<name>
<surname>Worobo</surname>
<given-names>R. W.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Landscape and meteorological factors affecting prevalence of three food-borne pathogens in fruit and vegetable farms</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>79</volume>, <fpage>588</fpage>&#x2013;<lpage>600</lpage>. <pub-id pub-id-type="doi">10.1128/aem.02491-12</pub-id> </citation>
</ref>
<ref id="B84">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strobl</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Boulesteix</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Zeileis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hothorn</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2007a</year>). <article-title>Bias in random forest variable importance measures: illustrations, sources and a solution</article-title>. <source>BMC Bioinformatics</source> <volume>8</volume>, <fpage>25</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-8-25</pub-id> </citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strobl</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Boulesteix</surname>
<given-names>A.-L.</given-names>
</name>
<name>
<surname>Augustin</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2007b</year>). <article-title>Unbiased split selection for classification trees based on the Gini Index</article-title>. <source>Comput. Stat. Data Anal.</source> <volume>52</volume>, <fpage>483</fpage>&#x2013;<lpage>501</lpage>. <pub-id pub-id-type="doi">10.1016/j.csda.2006.12.030</pub-id> </citation>
</ref>
<ref id="B86">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strobl</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Boulesteix</surname>
<given-names>A-L.</given-names>
</name>
<name>
<surname>Kneib</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Augustin</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zeileis</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Conditional variable importance for random forests</article-title>. <source>BMC Bioinformatics</source> <volume>9</volume>, <fpage>307</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-9-307</pub-id> </citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strobl</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hothorn</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zeileis</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Party on!</article-title>. <source>R. J.</source> <volume>1</volume>, <fpage>14</fpage>&#x2013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.32614/rj-2009-013</pub-id> </citation>
</ref>
<ref id="B88">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Therneau</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Atkinson</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Rpart: recursive partitioning and regression trees. 4.1-15. R package</article-title>. </citation>
</ref>
<ref id="B89">
<citation citation-type="other">
<collab>UK EA</collab>. <source>Bathing water quality</source>.</citation>
</ref>
<ref id="B90">
<citation citation-type="book">
<collab>US FDA</collab> (<year>2015</year>). <article-title>Standards for the growing, harvesting, packing, and holding of produce for human consumption</article-title>. <publisher-loc>United&#x20;States</publisher-loc>: <publisher-name>Food Safety Modernization Act</publisher-name>. </citation>
</ref>
<ref id="B91">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Venable</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Ripley</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Venables</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Ripley</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2002</year>). <source>Modern applied statistics with S</source>, <edition>4th Edn</edition>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B92">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wachtel</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Whitehand</surname>
<given-names>L. C.</given-names>
</name>
<name>
<surname>Mandrell</surname>
<given-names>R. E.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Prevalence of <italic>Escherichia coli</italic> associated with a cabbage crop inadvertently irrigated with partially treated sewage wastewater</article-title>. <source>J.&#x20;Food Prot.</source> <volume>65</volume>, <fpage>471</fpage>&#x2013;<lpage>475</lpage>. <pub-id pub-id-type="doi">10.4315/0362-028x-65.3.471</pub-id> </citation>
</ref>
<ref id="B93">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wall</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Clements</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Fisk</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Stoeckel</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Woods</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bihn</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Meeting report: key outcomes from a collaborative summit on agricultural water standards for Fresh produce</article-title>. <source>Comprehen. Rev. Food Sci. Food Safety</source> <volume>18</volume> (<issue>3</issue>), <fpage>723</fpage>&#x2013;<lpage>737</lpage>. </citation>
</ref>
<ref id="B94">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weller</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Shiwakoti</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bergholz</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Grohn</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wiedmann</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Strawn</surname>
<given-names>L. K.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Validation of a previously developed geospatial model that predicts the prevalence of Listeria monocytogenes in New York state produce fields</article-title>. <source>Appl. Environ. Microbiol.</source> <volume>82</volume>, <fpage>797</fpage>&#x2013;<lpage>807</lpage>. <pub-id pub-id-type="doi">10.1128/aem.03088-15</pub-id> </citation>
</ref>
<ref id="B95">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weller</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Love</surname>
<given-names>T. M. T.</given-names>
</name>
<name>
<surname>Belias</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wiedmann</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020a</year>). <article-title>Predictive models may complement or provide an alternative to existing strategies for managing enteric pathogen contamination of Northeastern streams used for produce production</article-title>. <source>Front. Sustain. Food Syst.</source> <volume>4</volume>, <fpage>561517</fpage>. <pub-id pub-id-type="doi">10.3389/fsufs.2020.561517</pub-id> </citation>
</ref>
<ref id="B96">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weller</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Belias</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Roof</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wiedmann</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020b</year>). <article-title>Landscape, water quality, and weather factors associated with an increased likelihood of foodborne pathogen contamination of New York streams used to source water for produce production</article-title>. <source>Front. Sustain. Food Syst.</source> <volume>3</volume>, <fpage>124</fpage>. <pub-id pub-id-type="doi">10.3389/fsufs.2019.00124</pub-id> </citation>
</ref>
<ref id="B97">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weller</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Brassill</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Rock</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ivanek</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mudrak</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Roof</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020c</year>). <article-title>Complex interactions between weather, and microbial and physicochemical water quality impact the likelihood of detecting foodborne pathogens in agricultural water</article-title>. <source>Front. Microbiol.</source> <volume>11</volume>, <fpage>134</fpage>. <pub-id pub-id-type="doi">10.3389/fmicb.2020.00134</pub-id> </citation>
</ref>
<ref id="B98">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wilkes</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Edge</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gannon</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Jokinen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lyautey</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Medeiros</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>Seasonal relationships among indicator bacteria, pathogenic bacteria, <italic>Cryptosporidium</italic> oocysts, Giardia cysts, and hydrological indices for surface waters within an agricultural landscape</article-title>. <source>Water Res.</source> <volume>43</volume>, <fpage>2209</fpage>&#x2013;<lpage>2223</lpage>. <pub-id pub-id-type="doi">10.1016/j.watres.2009.01.033</pub-id> </citation>
</ref>
<ref id="B99">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Den Bakker</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Boxrud</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Zoonotic source attribution of <italic>Salmonella enterica</italic> serotype typhimurium using genomic surveillance data, United&#x20;States</article-title>. <source>Emerg. Infect. Dis.</source> <volume>25</volume>, <fpage>82</fpage>&#x2013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.3201/eid2501.180835</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>