<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2022.848056</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Improving Crowdsourcing-Based Image Classification Through Expanded Input Elicitation and Machine Learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Yasmin</surname> <given-names>Romena</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1481545/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Hassan</surname> <given-names>Md Mahmudulla</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1554219/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Grassel</surname> <given-names>Joshua T.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1843753/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bhogaraju</surname> <given-names>Harika</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1851360/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Escobedo</surname> <given-names>Adolfo R.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Fuentes</surname> <given-names>Olac</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Computing and Augmented Intelligence, Arizona State University</institution>, <addr-line>Tempe, AZ</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Computer Science, University of Texas at El Paso</institution>, <addr-line>El Paso, TX</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Matt Lease, University of Texas at Austin, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Valentina Poggioni, University of Perugia, Italy; Evgenia Christoforou, CYENS&#x02014;Centre of Excellence, Cyprus</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Romena Yasmin <email>ryasmin&#x00040;asu.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Machine Learning and Artificial Intelligence, a section of the journal Frontiers in Artificial Intelligence</p></fn></author-notes>
<pub-date pub-type="epub">
<day>29</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>5</volume>
<elocation-id>848056</elocation-id>
<history>
<date date-type="received">
<day>03</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Yasmin, Hassan, Grassel, Bhogaraju, Escobedo and Fuentes.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yasmin, Hassan, Grassel, Bhogaraju, Escobedo and Fuentes</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>This work investigates how different forms of input elicitation obtained from crowdsourcing can be utilized to improve the quality of inferred labels for image classification tasks, where an image must be labeled as either positive or negative depending on the presence/absence of a specified object. Five types of input elicitation methods are tested: binary classification (positive or negative); the (<italic>x, y</italic>)-coordinate of the position participants believe a target object is located; level of confidence in binary response (on a scale from 0 to 100%); what participants believe the majority of the other participants&#x00027; binary classification is; and participant&#x00027;s perceived difficulty level of the task (on a discrete scale). We design two crowdsourcing studies to test the performance of a variety of input elicitation methods and utilize data from over 300 participants. Various existing voting and machine learning (ML) methods are applied to make the best use of these inputs. In an effort to assess their performance on classification tasks of varying difficulty, a systematic synthetic image generation process is developed. Each generated image combines items from the <italic>MPEG-7 Core Experiment CE-Shape-1 Test Set</italic> into a single image using multiple parameters (e.g., density, transparency, etc.) and may or may not contain a target object. The difficulty of these images is validated by the performance of an automated image classification method. Experiment results suggest that more accurate results can be achieved with smaller training datasets when both the crowdsourced binary classification labels and the average of the self-reported confidence values in these labels are used as features for the ML classifiers. Moreover, when a relatively larger properly annotated dataset is available, in some cases augmenting these ML algorithms with the results (i.e., probability of outcome) from an automated classifier can achieve even higher performance than what can be obtained by using any one of the individual classifiers. Lastly, supplementary analysis of the collected data demonstrates that other performance metrics of interest, namely reduced false-negative rates, can be prioritized through special modifications of the proposed aggregation methods.</p></abstract>
<kwd-group>
<kwd>machine learning</kwd>
<kwd>input elicitations</kwd>
<kwd>crowdsourcing</kwd>
<kwd>human computation</kwd>
<kwd>image classification</kwd>
</kwd-group>
<contract-sponsor id="cn001">U.S. Department of Homeland Security<named-content content-type="fundref-id">10.13039/100000180</named-content></contract-sponsor>
<contract-sponsor id="cn002">National Science Foundation<named-content content-type="fundref-id">10.13039/100000001</named-content></contract-sponsor>
<counts>
<fig-count count="5"/>
<table-count count="6"/>
<equation-count count="8"/>
<ref-count count="76"/>
<page-count count="18"/>
<word-count count="14040"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>In recent years, computer vision approaches based on machine learning (ML) and, in particular, those based on deep convolutional neural networks have demonstrated significant performance improvements over conventional approaches for image classification and annotation (Krizhevsky et al., <xref ref-type="bibr" rid="B38">2012</xref>; Tan and Le, <xref ref-type="bibr" rid="B69">2019</xref>; Zhai et al., <xref ref-type="bibr" rid="B74">2021</xref>). However, these algorithms generally require a large and diverse set of annotated data to generate accurate classifications. Large amounts of annotated data are not always available, especially for tasks where producing high-quality meta-data is costly, such as image-based medical diagnosis (Cheplygina et al., <xref ref-type="bibr" rid="B6">2019</xref>), pattern recognition in geospatial remote sensing data (Rasp et al., <xref ref-type="bibr" rid="B60">2020</xref>; Stevens et al., <xref ref-type="bibr" rid="B66">2020</xref>), etc. In addition, ML algorithms are often sensitive to perturbations in the data for complex visual tasks, that to some extent are even difficult for humans, such as object detection in cluttered backgrounds and detection of adversarial examples (McDaniel et al., <xref ref-type="bibr" rid="B45">2016</xref>; Papernot et al., <xref ref-type="bibr" rid="B55">2016</xref>), due to the high dimensionality and variability of the feature space of the images.</p>
<p>Crowdsourcing has received significant attention in various domain-specific applications as a complementary approach for image classification. Its growth has been accompanied and propelled by the emergence of online crowdsourcing platforms (e.g., Amazon Mechanical Turk, Prolific), which are widely employed to recruit and compensate human participants for annotating and classifying data that are difficult for machine-only approaches. In general, crowdsourcing works by leveraging the concept of the &#x0201C;wisdom of the crowd&#x0201D; (Surowiecki, <xref ref-type="bibr" rid="B67">2005</xref>), with which the judgments or predictions of multiple participants are aggregated to sift out noise and to better approximate a ground truth (Yi et al., <xref ref-type="bibr" rid="B72">2012</xref>). Numerous studies over the last decade have established that, under the right circumstances and with the proper aggregation methods, the collective judgment of multiple non-experts is uncontroversially more accurate than those from almost any individual, including well-informed experts. This concept of using groups to make collective decisions has been successfully applied to a number of visual tasks ranging from simple classification and annotation (Russakovsky et al., <xref ref-type="bibr" rid="B61">2015</xref>) to complex real-world applications, including assessment of damages caused by natural disasters (Barrington et al., <xref ref-type="bibr" rid="B2">2012</xref>) and segmentation of biomedical images for diagnostic purposes (Gurari et al., <xref ref-type="bibr" rid="B22">2015</xref>).</p>
<p>Although ML methods have been shown to perform exceedingly well in various classification tasks, these outcomes typically depend on relatively large datasets (Hsing et al., <xref ref-type="bibr" rid="B29">2018</xref>). However, high amounts of richly annotated data are inaccessible in various situations and/or obtaining them is prohibitively costly. Yet in such situations where less data is available, ML methods provide a natural mechanism for incorporating multiple forms of crowdsourced inputs, since they are tailor-made for classification based on input features. Previous works have tended to use a single form of input (i.e., mostly binary classification labels provided by participants) as a feature for ML algorithms on visual classification tasks. However, the vast majority have overlooked other inputs that can be elicitated from the crowd. Formal studies on the merits and potential impacts of different types of elicited inputs are also lacking. This work investigates how the performance of crowdsourcing-based voting and ML methods for image classification tasks can be improved using a variety of inputs. In summary, the contributions of this work stem from the following objectives:</p>
<list list-type="bullet">
<list-item><p>Analyze the reliability and accuracy of different ML classifiers on visual screening tasks when different forms of elicited inputs are used as features.</p></list-item>
<list-item><p>Evaluate the performance of the classifiers with these additional features on both balanced and imbalanced datasets&#x02014;i.e., sets of images with equal and unequal proportions, respectively, of positive to negative images&#x02014;of varying difficulty.</p></list-item>
<list-item><p>Introduce supplementary crowdsourcing-based methods to prioritize other performance metrics of interest, namely reduced false-negative and false-positive rates.</p></list-item>
<list-item><p>Analyze the performance of the crowdsourcing-based ML classifiers when outputs of an automated classifier trained on large annotated datasets are used as an additional feature.</p></list-item>
</list>
<p>To pursue these objectives, we design a number of experiments that elicit a diversity of inputs on each classification task: binary classification (1= positive or 0= negative); the (<italic>x, y</italic>)-coordinate of the target object&#x00027;s location; level of confidence in the binary response (on a scale from 0 to 100%); guess of what the majority of participants&#x00027; binary classification is on the same task; and level of the perceived difficulty of the binary classification task (on a discrete scale). To harness the benefits of both collective human intelligence and machine intelligence, we use the elicited inputs as features for ML algorithms. The results indicate that integrating diverse forms of input elicitation, including self-reported confidence values, can improve the accuracy and efficiency of crowdsourced computation. As an additional contribution, we develop an automated image classification method based on the ResNet-50 neural network architecture (He et al., <xref ref-type="bibr" rid="B27">2015</xref>) by training it on multiple datasets of sizes ranging from 10 k to 90 k image samples. The outputs of this automated classifier are used as additional features within the crowdsourcing-based ML algorithms. These additional results demonstrate that this hybrid image classification approach can provide more accurate predictions, especially for relatively larger datasets, than what is possible by either of the two stand-alone approaches.</p>
<p>Before proceeding, it is pertinent to mention that an earlier, shorter version of this work and a subset of its results appeared in Yasmin et al. (<xref ref-type="bibr" rid="B71">2021</xref>) and were presented at the 9th AAAI Conference on Human Computation and Crowdsourcing. That earlier conference paper considered only a subset of the crowdsourcing-based ML algorithms featured herein and that smaller selection was implemented only on balanced datasets. This present work also introduces a hybrid image classification approach, and it incorporates additional descriptions, crowdsourcing experiments, and analyses.</p>
</sec>
<sec id="s2">
<title>2. Literature Review</title>
<p>In recent years, crowdsourcing has been widely applied to complete a variety of image labeling/classification tasks, from those requiring simple visual identification abilities to those that rely on domain expertise. Many studies have leveraged crowdsourcing to annotate large-scale datasets, often requiring subjective analysis such as conceptualized images (Nowak and R&#x000FC;ger, <xref ref-type="bibr" rid="B52">2010</xref>), scene-centric images (Zhou et al., <xref ref-type="bibr" rid="B75">2014</xref>), and general-purpose images from publicly available sources (Deng et al., <xref ref-type="bibr" rid="B10">2009</xref>; Everingham et al., <xref ref-type="bibr" rid="B15">2010</xref>). Crowdsourcing techniques have also been successfully tailored to many other complex visual labeling/classification contexts that require profound domain knowledge, including identifying fish and plants (He et al., <xref ref-type="bibr" rid="B26">2013</xref>; Oosterman et al., <xref ref-type="bibr" rid="B53">2014</xref>), endangered species through camera trap images (Swanson et al., <xref ref-type="bibr" rid="B68">2015</xref>), locations of targets (Salek et al., <xref ref-type="bibr" rid="B64">2013</xref>), land covers (Foody et al., <xref ref-type="bibr" rid="B16">2018</xref>), and sidewalk accessibility (Hara et al., <xref ref-type="bibr" rid="B24">2012</xref>). Due to its low cost and rapid processing capabilities, another prominent use of crowdsourcing is classification of CT images in medical applications. Such tasks have included identifying malaria-infected red blood cells (Mavandadi et al., <xref ref-type="bibr" rid="B44">2012</xref>), detecting clinical features of glaucomatous optic neuropathy (Mitry et al., <xref ref-type="bibr" rid="B48">2016</xref>), categorizing dermatological features (Cheplygina and Pluim, <xref ref-type="bibr" rid="B7">2018</xref>), labeling protein expression (Irshad et al., <xref ref-type="bibr" rid="B31">2017</xref>), and various other tasks (Nguyen et al., <xref ref-type="bibr" rid="B51">2012</xref>; Mitry et al., <xref ref-type="bibr" rid="B47">2013</xref>).</p>
<p>Despite its effectiveness at processing high work volumes, numerous technical challenges need to be addressed to maximize the benefits of the crowdsourcing paradigm. One such technical challenge involves deploying effective mechanisms for judgment/estimation aggregation, that is, the combining or fusing of multiple sources of potentially conflicting information into a single representative judgment. Since the quality of the predictions is highly dependent on the method employed to consolidate the crowdsourced inputs (Mao et al., <xref ref-type="bibr" rid="B42">2013</xref>), a vast number of works have focused on developing effective algorithms to tackle this task. Computational social choice is a field dedicated to the rigorous analysis and design of such data aggregation mechanisms (Brandt et al., <xref ref-type="bibr" rid="B4">2016</xref>). Researchers in this field have studied the properties of various voting rules, which have been applied extensively to develop better classification algorithms. The most commonly used method across various types of tasks is Majority Voting (MV) (Hastie and Kameda, <xref ref-type="bibr" rid="B25">2005</xref>). MV attains high accuracy on simple idealized tasks, but its performance tends to degrade on those that require more expertise. One related shortcoming is that MV usually elicits and utilizes only one input from each participant&#x02014;typically a binary response in crowdsourcing. Relying on a single form of input elicitation may decrease the quality of the collective judgment due to cognitive biases such as anchoring, bandwagon effect, decoy effect, etc. (Eickhoff, <xref ref-type="bibr" rid="B12">2018</xref>). Studies have also found that the choice of input modality, for example, using rankings or ratings to specify a subjective response, can play a significant role in the accuracy of group decisions (Escobedo et al., <xref ref-type="bibr" rid="B13">2022</xref>) and predictions (Rankin and Grube, <xref ref-type="bibr" rid="B59">1980</xref>). These difficulties in data collection and aggregation mechanisms become even more prominent when the task at hand is complex (e.g., see Yoo et al., <xref ref-type="bibr" rid="B73">2020</xref>). Researchers have suggested many potential ways of mitigating these limitations. One promising direction is the collection of richer data, i.e., using multiple forms of input elicitation. As a parallel line of inquiry, previous works suggest that specialized aggregation methods for integrating this data should be considered for making good use of these different pieces of information (Kemmer et al., <xref ref-type="bibr" rid="B34">2020</xref>).</p>
<p>A logical enhancement of MV for the harder tasks is to elicit the participant&#x00027;s level of confidence (as a proxy of expertise) and to integrate these inputs within the aggregation mechanism. In the context of group decision-making, Grofman et al. (<xref ref-type="bibr" rid="B21">1983</xref>) suggested weighing each individual&#x00027;s inputs based on self-reported confidence of their respective responses, in accordance with the belief that individuals can estimate reliably the accuracy of their own judgments (Griffin and Tversky, <xref ref-type="bibr" rid="B20">1992</xref>). More recently, Hamada et al. (<xref ref-type="bibr" rid="B23">2020</xref>) designed a wisdom of the crowds study that asked a set of participants to rank and rate 15 items they would need for survival and used weighted confidence values to aggregate their inputs. The results were sensitive to the size of the group (i.e., number of participants); when the group was small (fewer than 10 participants), the confidence values reportedly had little impact on the results. In a more realistic application, Saha Roy et al. (<xref ref-type="bibr" rid="B63">2021</xref>) used binary classification and stated confidence in these inputs to locate target objects in natural scene images. Their study showed that using the weighted average of confidence values improved collective judgment. It is important to remark that these and the vast majority of related studies incorporate the self-reported confidence inputs at face value. The Slating algorithm developed by Koriat (<xref ref-type="bibr" rid="B37">2012</xref>) represents a different approach that determines the response according to the most confident participant. For additional uses of confidence values to make decisions, we refer the reader to Mannes et al. (<xref ref-type="bibr" rid="B41">2014</xref>) and Litvinova et al. (<xref ref-type="bibr" rid="B40">2020</xref>).</p>
<p>Although subjective confidence values can be a valid predictor of accuracy in some cases (Matoulkova, <xref ref-type="bibr" rid="B43">2017</xref>; G&#x000F6;rzen et al., <xref ref-type="bibr" rid="B19">2019</xref>), in many others they may degrade performance owing to cognitive biases that prevent a realistic assessment of one&#x00027;s abilities (Saab et al., <xref ref-type="bibr" rid="B62">2019</xref>). Another natural approach is to weigh responses based on some form of worker reliability. Khattak and Salleb-Aouissi (<xref ref-type="bibr" rid="B35">2011</xref>) used trapping questions with expert-annotated labels to estimate the expertise level of workers. For domain-specific tasks where the majority can be systematically biased, Prelec et al. (<xref ref-type="bibr" rid="B57">2017</xref>) introduced the Surprisingly Popular Voting method, which elicits two responses from participants: their own answer and what they think the majority of other participants&#x00027; answer is. It then selects the answer that is &#x0201C;more popular than people predict.&#x0201D; Other aggregation approaches include reference-based scoring models (Xu and Bailey, <xref ref-type="bibr" rid="B70">2012</xref>) and probabilistic inference-based iterative models (Ipeirotis et al., <xref ref-type="bibr" rid="B30">2010</xref>; Karger et al., <xref ref-type="bibr" rid="B33">2011</xref>).</p>
<p>In addition to crowdsourcing-based methods, automated image classification has become popular due to the breakthrough performances achieved by deep neural networks. Krizhevsky et al. (<xref ref-type="bibr" rid="B38">2012</xref>) used a convolutional neural network called AlexNet on a large dataset for the first time and achieved significant performance in image classification tasks compared to other contemporary methods. Since then, hundreds of studies have further improved classification capabilities, and a few have shown human-level performance when trained on large, noise-free datasets (Assiri, <xref ref-type="bibr" rid="B1">2020</xref>; Dai et al., <xref ref-type="bibr" rid="B9">2021</xref>). However, as the size and/or quality of training datasets decreases, the performance of these networks quickly degrades (Dodge and Karam, <xref ref-type="bibr" rid="B11">2017</xref>; Geirhos et al., <xref ref-type="bibr" rid="B17">2017</xref>).</p>
<p>A two-way relationship between AI and crowdsourcing can help compensate for some of the disadvantages associated with the two separate decision-making approaches. Human-elicited inputs interact with machine learning for a variety of reasons, but most are in service of the latter. A wider variety of ML models use human judgment to improve the accuracy and diversity in training data sets. For example, Chang et al. (<xref ref-type="bibr" rid="B5">2017</xref>) uses crowdsourcing to label images of cats and dogs since, unlike machines, humans can recognize these animals in many different contexts such as cartoons and advertisements. Human-elicited inputs are given more importance in specialized fields like law and medicine. For example, a study conducted by Gennatas et al. (<xref ref-type="bibr" rid="B18">2020</xref>) uses clinicians&#x00027; inputs to improve ML training datasets and as a feedback mechanism using what is aptly termed &#x0201C;Expert-augmented machine learning.&#x0201D; In a similarly promising direction, Hekler et al. (<xref ref-type="bibr" rid="B28">2019</xref>) uses a combination of responses from a user study and a convolutional neural network to classify images with skin cancer; the overall accuracy of their hybrid system was higher than both components in isolation.</p>
<p>Unlike human-AI interaction, human-AI collaboration is an emerging focus that can lead to the formulation of more efficient and inclusive solutions. Mora et al. (<xref ref-type="bibr" rid="B49">2020</xref>) designed an augmented reality shopping assistant that guides human clothing choices based on social media presence, historical purchase history, etc. As part of this focus, human-in-the-loop applications seek a more balanced integration of the abilities of humans and machines by sequentially alternating a feedback loop between them. For example, Koh et al. (<xref ref-type="bibr" rid="B36">2017</xref>) conducted a study where a field operator wearing smart glasses uses an artificial intelligence agent for remote assistance for hardware assembly tasks. Yet, few studies seek to combine human judgments and ML outputs to form a collective decision. Developing such equitable human-AI collaboration methods could be particularly beneficial in situations where the transparency, interpretability, and overall reliability of AI-aided decisions are of paramount concern.</p>
</sec>
<sec id="s3">
<title>3. Crowdsourcing-Based ML Classification</title>
<p>This section introduces different forms of input elicitations and describes how they can be utilized within a crowdsourcing-based ML classifier. Consider the image label aggregation problem where a set of images <italic>I</italic> are to be labeled by a set of participants <italic>P</italic>; without loss of generality, assume each image and participant has a unique identifier, that is, <italic>I</italic> &#x0003D; {<italic>i</italic><sub>1</sub>, <italic>i</italic><sub>2</sub>, &#x02026;, <italic>i</italic><sub><italic>n</italic></sub>} and <italic>P</italic> &#x0003D; {<italic>p</italic><sub>1</sub>, <italic>p</italic><sub>2</sub>, &#x02026;, <italic>p</italic><sub><italic>m</italic></sub>}, where <italic>n</italic> and <italic>m</italic> represent the total number of images and participants, respectively. For each image <italic>i</italic><sub><italic>k</italic></sub> &#x02208; <italic>I</italic>, the objective is to infer the binary ground truth label <italic>y</italic><sub><italic>k</italic></sub> &#x02208; {0, 1}, where <italic>y</italic><sub><italic>k</italic></sub> &#x0003D; 1 if the specified target object is present in the image (i.e., positive image) and <italic>y</italic><sub><italic>k</italic></sub> &#x0003D; 0 otherwise (i.e., negative image). Since in these experiments each worker may label only a subset of the images, let <italic>P</italic>(<italic>i</italic><sub><italic>k</italic></sub>) &#x02286; <italic>P</italic> be the set of participants who complete the labeling task of image <italic>i</italic><sub><italic>k</italic></sub> &#x02208; <italic>I</italic>. In contrast to most crowdsourced labeling tasks where only a single label estimate is elicited per classification task, in the featured experiments each participant is asked to provide multiple inputs from the following five options. The first input is their binary response <inline-formula><mml:math id="M1"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> (i.e., classification label) indicating the presence/absence of the target object in image <italic>i</italic><sub><italic>k</italic></sub>. The second input is a coordinate-pair <inline-formula><mml:math id="M2"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> indicating the location of the target object (elicited only when <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>). The third input is a numeric value <inline-formula><mml:math id="M4"><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>100</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> indicating the degree of confidence in the binary response <inline-formula><mml:math id="M5"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. The fourth input is another binary choice <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> indicating what <italic>p</italic><sub><italic>j</italic></sub> estimates the binary response assigned by the majority of participants to <italic>i</italic><sub><italic>k</italic></sub> is; this input is referred to in this study as the Guess of Majority Elicitation (GME). The fifth input is a discrete rating <inline-formula><mml:math id="M7"><mml:msubsup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, whose values are mapped from four linguistic responses&#x02014;1: &#x0201C;not at all difficult,&#x0201D; 2: &#x0201C;somewhat difficult,&#x0201D; 3: &#x0201C;very difficult,&#x0201D; and 4: &#x0201C;extremely difficult&#x0201D;&#x02014;indicating, in increasing order, the perceived difficulty of task <italic>i</italic><sub><italic>k</italic></sub>.</p>
<p>Before proceeding, it is worth motivating the use of participant confidence values in the proposed methods. Previous research has found that participants can accurately assess their individual confidence in their independently formed decisions (e.g., see Meyen et al., <xref ref-type="bibr" rid="B46">2021</xref>). However, a pertinent concern regarding these confidence values is that, even if some participants are accurate in judging their performance at certain times, humans are generally prone to metacognitive biases, i.e., overconfidence or underconfidence in their actual abilities (Oyama et al., <xref ref-type="bibr" rid="B54">2013</xref>). Hence, self-reported confidence should not be taken at face value, and specific confidence values should not be assumed to convey the same meaning across different individuals. In an attempt to mitigate such biases, the confidence values, <inline-formula><mml:math id="M8"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> provided by participant <italic>p</italic><sub><italic>j</italic></sub> &#x02208; <italic>P</italic> are rescaled linearly between 0 and 100, with the lowest confidence value expressed by <italic>p</italic><sub><italic>j</italic></sub> being mapped to 0 and the greatest to 100. Letting <italic>I</italic><sup><italic>j</italic></sup> &#x02286; <italic>I</italic> be the set of images for which <italic>p</italic><sub><italic>j</italic></sub> provides a label, the confidence of participant <italic>p</italic><sub><italic>j</italic></sub> at classifying image <italic>i</italic><sub><italic>k</italic></sub> is rescaled as</p>
<disp-formula id="E1"><mml:math id="M9"><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mn>100</mml:mn><mml:mo>.</mml:mo></mml:math></disp-formula>
<p>The remainder of this section describes how the collected input elicitations are used as features in ML classifiers to generate predictions.</p>
<sec>
<title>3.1. Features for Crowdsourcing-Based ML Methods</title>
<p>A total of seven features were extracted from the five inputs elicitations discussed in the beginning of this section for use with the ML classifiers; these features are described in the ensuing paragraphs.</p>
<list list-type="bullet">
<list-item><p><bold>Binary Choice Elicitation</bold>: For each image <italic>i</italic><sub><italic>k</italic></sub> &#x02208; <italic>I</italic>, the binary choice elicitation values are divided into two sets: one containing the participants with response <inline-formula><mml:math id="M10"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> and the other containing participants with response <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. The number of participants in each set can be used as an input feature within a ML classifier. However, since the number of participants can vary from image to image in practical settings, it is more prudent to use the relative size of the sets. Note that these relative sizes are complements of each other, that is, the fraction of participants who chose <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> as their binary choice label can be determined by subtracting from 1.0 the fraction of participants who chose <inline-formula><mml:math id="M13"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. Therefore, to remove redundancy and co-linearity within the features, only one of these values is used as an input and is given as</p>
<p><disp-formula id="E2"><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>where <inline-formula><mml:math id="M15"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is the fraction of participants who specify that the target object is present in image <italic>i</italic><sub><italic>k</italic></sub>.</p></list-item>
<list-item><p><bold>Spatial Elicitation</bold>: A clustering-based approach is implemented to identify participants whose location coordinates <inline-formula><mml:math id="M16"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>&#x02014;elicited only when they specify that the target object is present&#x02014;are close to each other. For each image <italic>i</italic><sub><italic>k</italic></sub> &#x02208; <italic>I</italic>, participants with binary choice label <inline-formula><mml:math id="M17"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> are divided into multiple clusters using the Density Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm (Ester et al., <xref ref-type="bibr" rid="B14">1996</xref>). The reasons for choosing this algorithm are twofold. First, DBSCAN is able to identify groups of points that are close to each other but form arbitrary shapes; since the target images have varying shapes and sizes, this is what one would expect to see in a single image if all collected data points were overlaid onto it. Second, this clustering algorithm can easily mark as outliers/noise the points that are in low density areas, i.e., coordinate points that have significant distance from each other. After clustering, the fraction of participants belonging to the largest cluster is used as an input feature within the ML classifiers. For image <italic>i</italic><sub><italic>k</italic></sub>, this input feature can be expressed as</p>
<p><disp-formula id="E3"><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">SE</mml:mtext></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>where <italic>n</italic><sub><italic>r</italic></sub> is the number of participants in cluster <italic>r</italic> and <italic>R</italic><sub><italic>k</italic></sub> is the set of clusters identified by DBSCAN for image <italic>i</italic><sub><italic>k</italic></sub>.</p></list-item>
<list-item><p><bold>Confidence Elicitation</bold>: Although previous works have explored using confidence scores to improve annotation quality of crowdsourced data (Ipeirotis et al., <xref ref-type="bibr" rid="B30">2010</xref>), very few have incorporated this input within a machine learning model. The confidence values are divided into two sets based on <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and the respective averages are used as additional features for the ML classifier. For image <italic>i</italic><sub><italic>k</italic></sub> &#x02208; <italic>I</italic>, these two input features can be expressed as</p>
<p><disp-formula id="E4"><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">conf</mml:mtext><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>;</mml:mo><mml:mtext class="textrm" mathvariant="normal">and</mml:mtext></mml:math></disp-formula></p>
<p><disp-formula id="E5"><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">conf</mml:mtext><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Here, the confidence values are rescaled linearly between 0 and 100 before incorporating them as the features.</p></list-item>
<list-item><p><bold>Guess of Majority Elicitation</bold>: Similar to BCE, GME is converted into a single feature based on the number of participants whose <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> response value is 1 and is written as</p>
<p><disp-formula id="E6"><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">GME,</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p></list-item>
<list-item><p><bold>Perceived Difficulty Elicitation :</bold> Previous research has shown that a task&#x00027;s perceived difficulty level can be used to some extent to improve the quality of annotation. In most cases, the difficulty level is set based on inputs from experts, that is, participants with specialized knowledge with respect to the task at hand (Khattak and Salleb-Aouissi, <xref ref-type="bibr" rid="B35">2011</xref>), or it is estimated from the classification labels collected from participants (Karger et al., <xref ref-type="bibr" rid="B33">2011</xref>). Unlike these works, the featured experiments gather the perceived difficulty of each task directly from each participant to evaluate the reliability of this information and its potential use within ML classifiers. For each image <italic>i</italic><sub><italic>k</italic></sub> &#x02208; <italic>I</italic>, the difficulty elicitation values <inline-formula><mml:math id="M24"><mml:msubsup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are divided into two sets: one for the participants with response <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, and the other for the remaining participants with response <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. The average values from each set are then used as additional features for the ML classifier; these two input features can be expressed as</p>
<p><disp-formula id="E7"><mml:math id="M27"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">PDE,</mml:mtext><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>;</mml:mo><mml:mtext class="textrm" mathvariant="normal">and</mml:mtext></mml:math></disp-formula></p>
<p><disp-formula id="E8"><mml:math id="M28"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textit" mathvariant="italic">PDE,</mml:mtext><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p></list-item>
</list>
</sec>
</sec>
<sec id="s4">
<title>4. Experiment Design</title>
<p>Prior to introducing the components of the experiment design, we describe the <italic>MPEG-7 Core Experiment CE-Shape-1 Test Set</italic> (Jeannin and Bober, <xref ref-type="bibr" rid="B32">1999</xref>; Ralph, <xref ref-type="bibr" rid="B58">1999</xref>), which is the source data from which the featured crowdsourcing activities are constructed. The dataset is composed of black and white images of a diverse set of shapes and objects including animals, geometric shapes, common household objects, etc. In total, the dataset consists of 1, 200 objects/shapes (referred to here as <italic>templates</italic>) divided into 60 object/shape classes, with each class containing 20 members. <xref ref-type="fig" rid="F1">Figure 1</xref> provides representative templates from some of these classes.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Object/shape templates from the MPEG-7 core experiment CE-shape-1 test set.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-848056-g0001.tif"/>
</fig>
<p>The images used in the crowdsourcing experiment are constructed by instantiating and placing multiple MPEG-7 Core Experiment CE-Shape-1 Test Set templates onto a single image frame. The instantiation of the image template is specified with six adjustable parameters: density, scale, color, transparency, rotation, and target object. See <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref> for a detailed description of these parameters.</p>
<sec>
<title>4.1. Description of Activities</title>
<p>For the crowdsourcing activities, we designed two studies, each of which elicits multiple forms of input from participants to complete a number of image classification tasks. A user interface was designed and implemented to perform the two studies, which differ based on the subsets of input elicitations tested and the class balance ratios of the image datasets (more details are provided later in this subsection). The interfaces were developed in HTML and Javascript and then deployed using Amazon Mechanical Turk (MTurk). Participants were first briefed about the nature of the study and shown a short walk-through video explaining the interface. Afterwards, participants proceeded to the image classification tasks, which were shown in a randomized order. After completing an experiment, participants were disallowed to participate in further experiments. <xref ref-type="fig" rid="F2">Figures 2</xref>, <xref ref-type="fig" rid="F3">3</xref> provide examples of the user interfaces, both of which instituted a 60 s time limit to view each image before it was hidden. If the participant completed the input elicitations before the time limit, they were allowed to proceed to the next image; on the other hand, if the time limit was reached, the image was hidden from view but participants could take as much time as they needed to finish providing their inputs. The time limit was imposed to ensure the scalable implementation of a high number of tasks. In particular, the goal is to develop activities that can capture enough quality inputs from participants while mitigating potential cognitive fatigue. In preliminary experiments, we found that participants rarely exceeded 45 s. In the featured studies (to be described in the next two paragraphs), the full 60 s were utilized in only 7% of the tasks, with an average time of around 27 s. The number of tasks given to the participants varied by experiment and ranged from 16 to 40 images (see <xref ref-type="table" rid="T1">Table 1</xref> for details). We deemed this number of tasks to be reasonable and not cognitively burdensome to participants based on findings of prior studies with shared characteristics. For instance, Zhou et al. (<xref ref-type="bibr" rid="B76">2018</xref>) performed a visual identification crowdsourcing study where participants were assigned up to 80 tasks, each of which took a median time of 29.4 s to complete. The authors found that accuracy decreased negligibly for this workload (i.e., twice as large as in the featured studies).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Image classification task UI for balanced dataset&#x02014;image contains bat (lower right).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-848056-g0002.tif"/>
</fig>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Image classification task UI for imbalanced dataset&#x02014;image contains bat (center left).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-848056-g0003.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of experiment image parameters.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left" colspan="2"><bold>Exp</bold>.</th>
<th valign="top" align="center"><bold>Images</bold></th>
<th valign="top" align="center"><bold>Density</bold></th>
<th valign="top" align="center"><bold>Scale</bold></th>
<th valign="top" align="left"><bold>Color</bold></th>
<th valign="top" align="left"><bold>Transparency</bold></th>
<th valign="top" align="left"><bold>Target</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left" rowspan="4">SetA</td>
<td valign="top" align="center">&#x00023;1</td>
<td valign="middle" align="center" rowspan="4">16</td>
<td valign="middle" align="center" rowspan="4">{100, 120,<break/> 140, 160}</td>
<td valign="middle" align="center" rowspan="4">{<italic>T</italic>(0.2 &#x000B1; 0.12), .., <italic>T</italic>(0.65 &#x000B1; 0.12)}</td>
<td valign="middle" align="left" rowspan="4">Discrete: {4}</td>
<td valign="middle" align="left" rowspan="4"><italic>U</italic>(100, 200)</td>
<td valign="top" align="left">Bat</td>
</tr>
<tr>
<td valign="top" align="center">&#x00023;2</td>
<td valign="top" align="left">Butterfly</td>
</tr>
<tr>
<td valign="top" align="center">&#x00023;3</td>
<td valign="top" align="left">Apple</td>
</tr>
<tr>
<td valign="top" align="center">&#x00023;4</td>
<td valign="top" align="left">Stingray</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="middle" align="left" rowspan="3">SetB</td>
<td valign="top" align="center">&#x00023;5</td>
<td valign="middle" align="center" rowspan="3">24</td>
<td valign="top" align="center">{80}</td>
<td valign="top" align="center">{<italic>T</italic>(0.2 &#x000B1; 0.05), <italic>T</italic>(0.3 &#x000B1; 0.05)}</td>
<td valign="top" align="left">Discrete: {1,&#x02026;,6}</td>
<td valign="middle" align="left" rowspan="3"><italic>U</italic>(140, 170)</td>
<td valign="top" align="left">Bat</td>
</tr>
<tr>
<td valign="top" align="center">&#x00023;6</td>
<td valign="top" align="center">{80,100,120}</td>
<td valign="top" align="center">{<italic>T</italic>(0.2 &#x000B1; 0.05), .., <italic>T</italic>(0.4 &#x000B1; 0.05)}</td>
<td valign="top" align="left"><italic>U</italic>(10, 255) for R,G,&#x00026; B</td>
<td valign="top" align="left">Turtle</td>
</tr>
<tr>
<td valign="top" align="center">&#x00023;7</td>
<td valign="top" align="center">{100, 150}</td>
<td valign="top" align="center">{<italic>T</italic>(0.2 &#x000B1; 0.05), <italic>T</italic>(0.3 &#x000B1; 0.05)}</td>
<td/>
<td valign="top" align="left">Various-7</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="middle" align="left" rowspan="3" style="border-bottom: thin solid #000000;">SetC</td>
<td valign="top" align="center">&#x00023;8</td>
<td valign="middle" align="center" rowspan="3">40</td>
<td valign="middle" align="center" rowspan="3">{90, 100, 115, 150}</td>
<td valign="middle" align="center" rowspan="3">{<italic>T</italic>(0.25, 0.35, 0.40)}</td>
<td valign="middle" align="left" rowspan="3">Discrete: {4}</td>
<td valign="middle" align="left" rowspan="3"><italic>U</italic>(150, 200)</td>
<td valign="middle" align="left" rowspan="3">Bat</td>
</tr>
<tr>
<td valign="top" align="center">&#x00023;9</td>
</tr>
<tr>
<td valign="top" align="center" style="border-bottom: thin solid #000000;">&#x00023;10</td>
</tr>
<tr>
<td valign="middle" align="left" rowspan="3">SetD</td>
<td valign="top" align="center">&#x00023;11</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x00023;12</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="center">&#x00023;13</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<p>In the first study, seven experiments were completed and grouped into two sets: Experiment Set A (four experiments) and Experiment Set B (three experiments). Each experiment used a balanced set of images, with half containing the target template (i.e., positive images); target objects were chosen so as to avoid confusion with other template classes. See <xref ref-type="table" rid="T1">Table 1</xref> for image generation parameters, and see <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref> for additional related details. The parameter ranges selected for Experiment Set A were designed to keep the difficulty of the classification tasks relatively moderate. On the other hand, a more complex set of parameters was selected for Experiment Set B to expand the range of difficulty. These differences are reflected in the individual performance achieved in these two experiment sets, measured by the respective average number of correct classifications obtained by participants. For Experiment Set A, individual performance averages ranged between 59 and 77% for each of the four experiments, whereas for Experiment Set B, they were between 54 and 82% for each of the three experiments.</p>
<p>In the second study, six experiments were conducted. These experiments were also grouped into two sets: Experiment Set C (three experiments) and Experiment Set D (three experiments). Each consisted of image sets with an imbalanced ratio of positive-to-negative images. Experiment Set C had a 20-80 balance, meaning that 20% of the images were positive, and 80% were negative; Experiment Set D had a 10&#x02013;90 balance. The results of Experiment Sets A and B revealed that <italic>scale</italic> and <italic>density</italic> are the only factors that had a statistically significant impact on individual performance. Based on this insight, we constructed a simple linear regression model with these two parameters as the predictors and <italic>proportion of correct participants</italic> as the responses; the model is very significant (<italic>p</italic> &#x0003C; 0.001), and its adjusted R-squared value is 0.65. The model was used to generate image sets with an approximated difficulty level by modifying the scale and density parameters accordingly. It should be noted that the true difficulty of each image varies based on the random generation process. The model was implemented to design experiments consisting of classification tasks of reasonable difficulty&#x02014;that is, neither trivial nor impossible to complete. Images of four levels of difficulty were generated for Experiment Sets C and D. At each difficulty level, the density was varied while keeping the other parameters consistent across images. This resulted in images that appear similar, but with different amounts of &#x0201C;clutter&#x0201D;. The four difficulties generated were categorized as &#x0201C;very difficult,&#x0201D; &#x0201C;difficult,&#x0201D; &#x0201C;average,&#x0201D; and &#x0201C;easy.&#x0201D; See <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref> for details and sample images of each difficulty. Experiment Sets C and D use an even split of each difficulty (i.e., 25% of generated images from each level). For the three respective experiments, individual average accuracy values ranged between 65 and 73% for Set C and between 58 and 72% for Set D.</p>
<p><xref ref-type="fig" rid="F2">Figures 2</xref>, <xref ref-type="fig" rid="F3">3</xref> show the user interface presented to participants in the first and second study, respectively. For each classification task (i.e., image) in the first study, participants were asked to provide a binary response indicating whether or not a target object is present. If they responded affirmatively, they were then prompted to locate the target object by clicking on it. Then, participants were asked to rate their confidence in their binary response on a scale from 0 to 100%. Finally, participants were asked to guess the binary response of the majority of participants. The second study asked participants similar questions as the first study. For each classification task, participants were also asked to provide a binary responses indicating whether or not a target object is present and their level of confidence in this response. If they responded affirmatively, however, they were then prompted to locate the target object by drawing a bounding box around it; the centroid of the bounding box was used as the (<italic>x, y</italic>)-coordinate gathered from this elicitation. In replacement to the last question of the first study, participants were asked to rate the difficulty of the specific image being classified based on a discrete scale. The rating choices provided were &#x0201C;not difficult at all,&#x0201D; &#x0201C;somewhat difficult,&#x0201D; &#x0201C;very difficult,&#x0201D; and &#x0201C;extremely difficult.&#x0201D; These labels were mapped to 1, 2, 3, and 4, respectively, for use in the aggregation algorithms.</p>
</sec>
<sec>
<title>4.2. Participant Demographics and Filtering of Insincere Participants</title>
<p>A total of 356 participants were recruited and compensated for their participation using Amazon MTurk. Participants in Experiment Set A were paid $1.25, those in Experiment Set B were paid $2.00, and those in Experiment Sets C and D were paid $3.75. The differences in compensation can be attributed to the number of questions and the difficulty of image classification tasks of the respective experiment sets. Participants were made aware of the compensation amount before beginning the study. Payment was based only on completion and not on performance. Before proceeding, it is necessary to delve further into the quality of the participants recruited <italic>via</italic> the MTurk platform, and the quality of data they provide. Because of the endemic presence in most crowdsourcing platforms of annotators who do not demonstrate an earnest effort (Christoforou et al., <xref ref-type="bibr" rid="B8">2021</xref>), some criteria should be defined to detect such insincere participants and filter out low-quality inputs. This work defined two criteria for characterizing (and filtering out) an annotator as insincere:</p>
<list list-type="bullet">
<list-item><p><bold>Criterion 1:</bold> The participant answered over 75% of the questions in no more than 10 s per question.</p></list-item>
<list-item><p><bold>Criterion 2:</bold> The participant&#x00027;s binary responses were exclusively 0 or exclusively 1 over the entire question set.</p></list-item>
</list>
<p>Criteria 1 was imposed based on the following reasoning. In general, classification of negative images takes longer than classification of positive images. Even if it is assumed that participants can spot the positive images immediately (i.e., within 10 s), it should take more than 10 s to reply to the negative images that are of moderate to high difficulty. Because each Experiment Set in this study contained at least 50% negative images (Experiment Sets C and D contain a higher percentage) and only a small minority were of low difficulty, a conservative estimate that participants should take longer than 10 s to answer at least 25% of the images was set (i.e., to be more lenient toward the participants). Further analysis of the behavior of the participants in relation to the task completion times supporting this observation has been added to <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref>.</p>
<p>From the initial 356 participants, 50 participants were removed from the four experiment sets using the above criteria. Among them, 15 fell under criterion 1 and the rest under criterion 2. As expected, filtering out these data provided less noisy inputs to the crowdsourcing-based aggregation methods. From the remaining 306 participants, 276 completed the demographics survey. Their reported ages ranged from 21 to 71 years old, with a mean and median of 36 and 33, respectively. 156 participants reported their gender as male, 120 as female, and 0 as other. In terms of reported education level, 23 participants finished a high school/GED, 17 some college, 16 a 2-year degree, 148 a 4-year degree, 70 a master&#x00027;s degree, 1 a professional degree, and 2 a doctoral degree.</p>
</sec>
<sec>
<title>4.3. Distribution of Crowdsourced Data</title>
<p>Before proceeding to the computational results, it is pertinent to analyze the data collected from the crowdsourcing experiments. First, let us analyze the relationship between the perceived difficulty levels reported by the participants (i.e., input feature PDE) and the difficulty levels utilized in the proposed image generation algorithm (see Section 4.1 for details). The average difficulty values reported by participants for images categorized by the algorithm as &#x0201C;very difficult,&#x0201D; &#x0201C;difficult,&#x0201D; &#x0201C;average,&#x0201D; and &#x0201C;easy&#x0201D; were 2.89, 2.73, 2.62, and 2.03, respectively. This evinces a clear correlation, with the &#x0201C;very difficult&#x0201D; images having the highest average perceived difficulty values and the rest reflecting a decreasing order of difficulty, which supports the ability of the image generation method used in this study to control the classification task difficulty, according to the four above-mentioned categories.</p>
<p>Next, let us analyze the correctness of the binary response values collected from the participants. <xref ref-type="fig" rid="F4">Figure 4</xref> shows the percentage of participants who answered each question accurately; question numbers have been reordered for each of the four datasets by increasing participant accuracy. The positive and negative images for the balanced and imbalanced datasets are presented in separate graphs. The plots show that, for the balanced datasets (Experiments Sets A and B), the accuracy on the positive images is significantly lower than on the negative images. Moreover, in Experiment Set B, nearly half of the positive images have accuracy values below 0.4, whereas in Experiment Set A most images have values above 0.4. This is a good indication of the higher difficulty level of Experiment Set B. For the imbalanced datasets, in both Experiments Sets C and D, nearly all negative images have accuracy values above 0.4. In Experiment Set C, there is an almost even distribution of the positive images above and below 0.6, whereas in Experiment Set D nearly 60% of the positive images have accuracy values below 0.5, indicating that Experiment Set D was comparatively more difficult.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Distribution of Binary Classification results from crowdsoured data. <bold>(A)</bold> Balanced dataset. <bold>(B)</bold> Imbalanced dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-848056-g0004.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<title>5. Computational Results</title>
<p>This section compares the performance of the voting and crowdsourcing-based ML methods presented in Section 3 on both balanced and imbalanced datasets. As a baseline of comparison for the proposed crowdsourcing-based ML methods, three traditional voting methods are used: Majority Voting (MV), Confidence Weighted Majority Voting (CWMV), and Surprisingly Popular Voting (SPV). The details of these methods can be found in <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref>. For the ML methods, four binary classification approaches were selected: K-Nearest Neighbor (KNN), Logistic Regression (LR), Random Forest Classifier (RF), and Linear Support Vector Machines (SVM-Linear). These were selected as reasonable representatives of commonly available methods. The ML classifiers were trained and evaluated using built-in functions of the Python <italic>scikit-learn</italic> library (Pedregosa et al., <xref ref-type="bibr" rid="B56">2011</xref>). The hyper-parameters were optimized on a linear grid search with a nested 5-fold cross-validation strategy. However, due to the small size of the datasets, a Leave-One-Out (LOO) cross-validation strategy was used to train and evaluate the classifiers.</p>
<p>In the DBSCAN clustering approach used for extracting the Spatial Elicitation (SE), the maximum distance between two data points in the cluster (&#x003F5;) and the minimum data points required to form a cluster (<italic>MinPts</italic>) was set to 50 and 3, respectively. The former was set based on the size of the target objects used relative to the size of the image frame (1, 080 &#x000D7; 1, 080); the latter was set to ensure a sufficiently low probability of forming a cluster with random inputs. To obtain a rough estimate of this probability, consider the case where three participants with binary response <inline-formula><mml:math id="M29"><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> randomly select their location coordinates on an image with area <italic>A</italic>. The probability of two points having a maximum distance of <italic>r</italic> (i.e., falling within a circle with radius <italic>r</italic>) is &#x003C0;<italic>r</italic><sup>2</sup>/<italic>A</italic> and, therefore, the probability of the three points being identified as a cluster by DBSCAN is 2(&#x003C0;<italic>r</italic><sup>2</sup>/<italic>A</italic>)2. Setting <italic>r</italic> &#x0003D; &#x003F5; &#x0003D; 50 and <italic>A</italic> &#x0003D; 1, 080 &#x000D7; 1, 080 for our experiment, this probability value becomes 0.01, which is sufficiently small and justifies the use of the selected parameters.</p>
<sec>
<title>5.1. Performance of Aggregation Methods on Balanced Datasets</title>
<p>This section compares the performance of the voting and ML methods on balanced datasets (Experiment Sets A and B). The initial study elicits four out of the five inputs listed in Section 3.1: BCE, GME, CE, and SE. The results are summarized in <xref ref-type="table" rid="T2">Tables 2</xref> and <xref ref-type="table" rid="T3">3</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Performance analysis of voting methods for balanced dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>MV</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>CWMV</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>SPV</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td valign="top" align="center"><bold>Acc</bold>.</td>
<td valign="top" align="center"><bold>FNR</bold></td>
<td valign="top" align="center"><bold>Acc</bold>.</td>
<td valign="top" align="center"><bold>FNR</bold></td>
<td valign="top" align="center"><bold>Acc</bold>.</td>
<td valign="top" align="center"><bold>FNR</bold></td>
</tr>
<tr>
<td valign="top" align="left">Experiment Set A</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.94</td>
</tr>
<tr>
<td valign="top" align="left">Experiment Set B</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.92</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Performance analysis of crowdsourcing based ML methods for balanced dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Input</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>KNN</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>LR</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>RF</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>SVM-Linear</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Elicitations</bold></th>
<th valign="top" align="center"><bold>Acc</bold>.</th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>Acc</bold>.</th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>Acc</bold>.</th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>Acc</bold>.</th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center" colspan="13"><bold>Experiment Set A</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center"><bold>0.95</bold></td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center"><bold>0.19</bold></td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center">0.90</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center"><bold>0.93</bold></td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.91</td>
</tr>
<tr>
<td valign="top" align="left">BCE-SE</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left">BCE-GME</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center"><bold>0.19</bold></td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center">0.91</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center"><bold>0.16</bold></td>
<td valign="top" align="center">0.91</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-GME</td>
<td valign="top" align="center">0.8</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.90</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE-GME</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.89</td>
</tr>
<tr>
<td valign="top" align="center" colspan="13"><bold>Experiment Set B</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center"><bold>0.22</bold></td>
<td valign="top" align="center"><bold>0.82</bold></td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">BCE-SE</td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center"><bold>0.19</bold></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.80</td>
</tr>
<tr>
<td valign="top" align="left">BCE-GME</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center"><bold>0.88</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center"><bold>0.82</bold></td>
<td valign="top" align="center"><bold>0.22</bold></td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-GME</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center"><bold>0.76</bold></td>
<td valign="top" align="center"><bold>0.22</bold></td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE-GME</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.83</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold values denote best performance among the different input elicitation combinations for each Crowdsourcing-based ML method</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>The performance of the ML methods is quantified <italic>via</italic> three performance metrics: accuracy (Acc.), false-negative rate (FNR), and area under the ROC curve (AUC). For the voting methods, only the first two of these metrics are reported. For each of the ML classifiers, the best accuracy, FNR, and AUC values among the different input elicitation combinations are marked in bold. Before proceeding, it is worthwhile to mention two additional points regarding the values presented in the table. First, each row in <xref ref-type="table" rid="T3">Table 3</xref> represents a different combination of inputs used as features for the ML classifiers. For example, BCE-CE indicates that both binary and confidence elicitation inputs (i.e., <inline-formula><mml:math id="M30"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textit" mathvariant="italic">conf</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textit" mathvariant="italic">conf</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) were used as features for the ML classifiers, whereas BCE-CE-SE-GME indicates that all four input elicitations (i.e., <inline-formula><mml:math id="M32"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textit" mathvariant="italic">conf</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textit" mathvariant="italic">conf</mml:mtext></mml:mstyle><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textit" mathvariant="italic">SE</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M33"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext class="textit" mathvariant="italic">GME,</mml:mtext></mml:mstyle><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) of Experiment Set A and B were used as the respective input features. Second, when calculating the accuracy and FNR values of the voting methods, images with undecided outcomes (i.e., ties) are considered as a third separate label.</p>
<p>Let us first discuss the performance of the aggregation models in terms of accuracy. For Experiment Sets A and B, the average accuracy value of MV was stable at around 72%. The CWMV method performed significantly better than MV, achieving an average accuracy value of around 77%. SPV was the worst performer across the board, with an average accuracy value of &#x0003C;50% (i.e., worse than a purely random classifier). This low performance can be largely attributed to the excessive number of tied labels generated compared to the other methods. In SPV, 18 out of the 136 instances were classified as tied (i.e., participants were undecided regarding the guess of the majority&#x00027;s estimate). By comparison, there were only three tied instances with MV and none with CWMV.</p>
<p>The results of the ML classifiers in Experiment Set A were relatively consistent in terms of both accuracy and AUC values for all seven combinations of the input elicitations. The classifiers performed particularly well, attaining accuracy values above 83% for all combinations; this can be partly explained by the fact that the images in this experiment set were generated using parameter ranges that were more consistent and less variable in difficulty. In Experiment Set B, the ML classifiers reached higher accuracy and AUC values under certain combinations of the input elicitations. For RF, LR, and KNN, a noticeable increase in AUC values (from 76 to 85%) results when using the BCE-CE combination compared to the standalone BCE input; the accuracy values in these cases either increased or stayed the same. Altogether, these results suggest that integrating CE into an ML classifier can help attain more accurate predictions when the sample size is small and the difficulty level of the images is more varied. Furthermore, they show that the ML classifiers outperformed the voting methods, with the LR classifier achieving the highest values in terms of both accuracy and AUC scores.</p>
<p>Another performance metric of interest is FNR, which denotes the fraction of images the methods label as 0 (i.e., negative) when their true label is 1 (i.e., positive). A high FNR may be concerning in many critical engineering and medical applications where a false-negative may be more detrimental than a false-positive since the latter can be easily verified in subsequent steps. For example, FNR has significant importance in detecting lung cancer from chest X-rays. If the model falsely classifies an X-ray as negative, the patient may not receive needed medical care in a timely fashion. Returning to <xref ref-type="table" rid="T2">Table 2</xref>, the FNRs of the three voting methods are high across the board, with SPV again having the worst performance. The high FNRs of MV and CWMV can be attributed to the fact that people tend to label the image as negative whenever they fail to find the target object and that these methods are unable to extract additional useful information from the responses.</p>
<p>In Experiment Set A, the accuracy values are the highest for the BCE-CE combination, whereas the FNR values are the lowest for the single BCE input. On the other hand, in Experiment Set B, although the accuracy values are the same for both input combinations, FNR values decrease for the BCE-CE combination. Moreover, for SVM, the reduction in FNR values is significant for Experiment Set B (from 42 to 25%) for the BCE-CE combination. This outcome reiterates the advantages of integrating CE into ML classifiers for more complex datasets.</p>
</sec>
<sec>
<title>5.2. Performance of Aggregation Methods on Imbalanced Datasets</title>
<p>This section compares the performance of the voting and crowdsourcing-based ML methods on imbalanced datasets (Experiment Sets C and D). Similar to the balanced datasets, a total of four input elicitations are utilized. However, for this study, the GME input is replaced by the PDE input (i.e., a rating value to assess the difficulty of the classification task), which is explained as follows. Recall from the discussion of Section 5.1 that none of the ML classifiers obtained a performance improvement when using the GME input relative to the other elicitation combinations. Moreover, the only method that utilizes the GME elicitation, SPV, was the worst-performing among the three voting methods. The inability of the GME input to provide any additional information during the classification process prompted its removal from subsequent studies. Due to this modification, only two voting methods (MV and CWMV) are explored for the imbalanced datasets.</p>
<p>When the dataset is balanced, accuracy by itself is a good indicator of the model&#x00027;s performance. However, when the dataset is imbalanced, accuracy can often be misleading as it provides an overly optimistic estimation of the classifier&#x00027;s performance on the majority class (&#x0201C;0&#x0201D; in this experiment). In such cases, a more accurate evaluation metric is the <italic>F</italic><sub>1</sub>-score (Sokolova et al., <xref ref-type="bibr" rid="B65">2006</xref>), defined as the harmonic mean of the precision and recall values and can be expressed as, <italic>F</italic><sub>1</sub>-<inline-formula><mml:math id="M34"><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">score</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">Precision</mml:mtext></mml:mstyle><mml:mo>&#x000D7;</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">Recall</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">Precision</mml:mtext></mml:mstyle><mml:mo>&#x0002B;</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">Recall</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">TP</mml:mtext></mml:mstyle><mml:mo>/</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">TP</mml:mtext></mml:mstyle><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">FP</mml:mtext></mml:mstyle><mml:mo>&#x0002B;</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">FN</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></inline-formula> where, TP, FP, and FN refers to the number of true-positives (images the methods label as 1 when their true label is 1), false-positives (images the methods label as 1 when their true label is 0), and false-negatives (images the methods label as 0 when their true label is 1), respectively. Since both Experiment Sets C and D are highly imbalanced (with an average of 15% of their images belonging to the positive class) the <italic>F</italic><sub>1</sub>-score is reported instead of accuracy to better estimate the performance of the classifiers.</p>
<p>The overall results for the voting and machine learning methods are summarized in <xref ref-type="table" rid="T4">Tables 4</xref>, <xref ref-type="table" rid="T5">5</xref>, respectively. The performance of the ML methods is quantified <italic>via</italic> three performance metrics: <italic>F</italic><sub>1</sub>-score, FNR, and AUC; for the voting methods, only the first two of these metrics are reported. Let us first discuss the performance of the aggregation methods in terms of <italic>F</italic><sub>1</sub>-score. For Experiment Sets C and D, MV and CWMV have comparable scores, with both having the same value in the first set and MV outperforming CWMV by a slight margin in the second set. Moving on to the ML methods, for Experiment Set C, the ML classifiers displayed comparable <italic>F</italic><sub>1</sub>-scores for combinations BCE-CE, BCE-SE, BCE-CE-SE, and BCE-CE-SE-PDE. In addition, all four of these input combinations performed better than the standalone BCE input. The RF and KNN classifiers achieved the highest values with the combination BCE-CE-SE-PDE. In contrast, the LR and SVM classifiers achieved the highest values with the BCE-SE combination. Overall, the LR classifier achieved the best performance for this set with inputs BCE-SE. In Experiment Set D, the results followed a different pattern. In this case, the classifiers achieved the same or higher values when the BCE-CE combination was used compared to the BCE-SE or BCE-CE-SE combinations, indicating that the SE input does not provide any additional information for this experiment set. Because this dataset is highly skewed toward the negative class (10&#x02013;90 balance), we conjecture that participants may have become demotivated to closely inspect difficult images from the positive class. Whatever the cause, smaller clusters were obtained from these images, reducing the effectiveness of the SE input in many cases. In Experiment Set D, the highest performance was achieved by the SVM classifier for the BCE-CE input. These results once again indicate that, even though the self-reported confidence values are not particularly helpful when used within the traditional voting methods context (Li and Varshney, <xref ref-type="bibr" rid="B39">2017</xref>; Saab et al., <xref ref-type="bibr" rid="B62">2019</xref>)&#x02014;as can also be seen by the performance of the CWMV algorithm in this study&#x02014;incorporating them into an ML classifier can help attain better performance, specifically higher <italic>F</italic><sub>1</sub>-scores for highly imbalanced datasets.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Performance analysis of voting methods for imbalanced dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>MV</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>CWMV</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Experiment Set C</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.25</td>
</tr>
<tr>
<td valign="top" align="left">Experiment Set D</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.50</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Performance analysis of crowdsourcing based ML methods for imbalanced datasets.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Input</bold> <break/> <bold>Elicitations</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>KNN</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>LR</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>RF</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>SVM-Linear</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold><italic>AUC</italic></bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center" colspan="13"><bold>Experiment Set C</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.90</td>
</tr>
<tr>
<td valign="top" align="left">BCE-SE</td>
<td valign="top" align="center"><bold>0.81</bold></td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center"><bold>0.95</bold></td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">BCE-PDE</td>
<td valign="top" align="center"><bold>0.81</bold></td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.9</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center"><bold>0.90</bold></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.88</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-PDE</td>
<td valign="top" align="center"><bold>0.81</bold></td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.90</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE-PDE</td>
<td valign="top" align="center"><bold>0.81</bold></td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center"><bold>0.81</bold></td>
<td valign="top" align="center"><bold>0.29</bold></td>
<td valign="top" align="center"><bold>0.90</bold></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="center" colspan="13"><bold>Experiment Set D</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center"><bold>0.58</bold></td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center"><bold>0.33</bold></td>
<td valign="top" align="center"><bold>0.87</bold></td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center"><bold>0.42</bold></td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE</td>
<td valign="top" align="center"><bold>0.59</bold></td>
<td valign="top" align="center"><bold>0.58</bold></td>
<td valign="top" align="center"><bold>0.76</bold></td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center"><bold>0.63</bold></td>
<td valign="top" align="center"><bold>0.50</bold></td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center"><bold>0.67</bold></td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center"><bold>0.87</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-SE</td>
<td valign="top" align="center"><bold>0.59</bold></td>
<td valign="top" align="center"><bold>0.58</bold></td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.8</td>
</tr>
<tr>
<td valign="top" align="left">BCE-PDE</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center"><bold>0.57</bold></td>
<td valign="top" align="center"><bold>0.33</bold></td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center"><bold>0.42</bold></td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE</td>
<td valign="top" align="center"><bold>0.59</bold></td>
<td valign="top" align="center"><bold>0.58</bold></td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center"><bold>0.87</bold></td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center"><bold>0.67</bold></td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.84</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-PDE</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center"><bold>0.87</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE-PDE</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center"><bold>0.58</bold></td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center"><bold>0.67</bold></td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.85</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold values denote best performance among the different input elicitation combinations for each Crowdsourcing-based ML method</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>In terms of FNR, the performance of the CWMV method was markedly better than the MV method for both Experiment Sets. The assigned labels for the positive images in Experiment Set D for the two voting methods are almost identical, with the exception of a single image which the latter labeled as a tie (i.e., undecided), contributing to the decrease in performance. Note that none of the images in Experiment Set C was labeled as a tie by either of the voting methods. Among the ML methods, LR significantly outperformed all of the other classifiers for Experiment Set D. Although in Experiment Set C the FNR for the BCE-SE combination (25%) was lower than for the BCE combination (33%), in Experiment Set D a significant increase (33&#x02013;50%) can be seen between these two combinations. Overall, ML classifiers outperformed MV; however, CWMV showed comparable performance for both experiment sets. Note that a distinctive advantage of CWMV over the ML methods is that it does not require training data.</p>
</sec>
<sec>
<title>5.3. Changing the Threshold of Positive Classification</title>
<p>This section examines how voting methods can be modified to emphasize other important metrics of image classification. In particular, it seeks to prioritize reduced false-negative rates, which are relevant in various critical applications. The FNRs can be reduced by lowering the threshold at which a positive classification is returned by a classification method (i.e., changing the tipping point for returning a positive collective response). However, care must be exercised when lowering the threshold since this implicitly increases false-positive rates (FPRs), which can also be problematic.</p>
<p>By default, the threshold at which voting methods return a positive response is fixed; for example, MV requires more than 50% of positive responses to return the positive class. <xref ref-type="fig" rid="F5">Figure 5</xref> illustrates the impacts of adjusting the thresholds for the voting methods as well as for the ML methods; the figure separates FNRs from FPRs for each method. Using MV as an example, decreasing the threshold from 0.5 to 0.3 results in relatively small increases to the FPR and larger decreases to the FNR; further decreases cause a disproportionate increase to FPRs. Hence, these inflection points can help guide how the thresholds can be set for each voting method to prioritize FNR. A similar observation can be made about the FNRs of the ML methods (except for LR) for the imbalanced datasets. However, this does not hold for the ML methods for the balanced datasets&#x02014;for example, reducing the threshold to 0.3 causes a significant increase in FPRs compared to the decrease in FNRs. This suggests that caution must be exercised when changing the threshold of positive classification of ML classifiers.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Change in FNR/FPR of different aggregation methods under varying thresholds. <bold>(A)</bold> Experiment Set A, <bold>(B)</bold> Experiment Set B, <bold>(C)</bold> Experiment Set C, <bold>(D)</bold> Experiment Set D.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-848056-g0005.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<title>6. Enhancement of Crowdsourcing-Based ML Methods With an Automated Classifier</title>
<p>In order to assess the difficulty of the image classification problem presented to participants and to evaluate the potential of hybrid human-ML approaches, we developed a deep learning image classification approach that leverages large training datasets. Our classifier is based on ResNet-50, a popular variant of ResNet architecture (He et al., <xref ref-type="bibr" rid="B27">2015</xref>), which has shown very good performance on multiple image classification tasks. It has been extensively used by the computer vision research community and adopted as a baseline architecture in many studies done over the last few years (Bello et al., <xref ref-type="bibr" rid="B3">2021</xref>).</p>
<p>For training the classifier, we generated a balanced dataset of 100 k samples, with 10 k samples set aside as the validation set and the rest used as the training set. The images are representative of an even mixture of the difficulty classes used to generate Experiment Sets C and D. We trained and evaluated the performance of the network using training set sizes ranging from 10 k samples to 90 k samples, increasing the training set size by 10 k every iteration, totaling nine different training sessions. Each training session was started from the previous session&#x00027;s best-performing checkpoint of the network and the corresponding optimization state and continued for 35 epochs. See <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref> for a complete description of the ResNet classifier used as well as a detailed analysis of its performance.</p>
<p>We emphasize that this work does not aim to advance the state-of-the-art results for automated image classification. Instead, the focus of the automated classification method is to explore the benefits and limitations of a hybrid method introduced herein that integrates the outputs of a well-known deep neural network into the crowdsourcing-based classification methods. In particular, the proposed method uses the output of the automated classifier as an additional feature of the featured ML methods. <xref ref-type="table" rid="T6">Table 6</xref> summarizes the results for the small imbalanced test sets used in Experiment Sets C and D as the training set grows larger. Due to the imbalanced nature of these test sets, this table and the rest of the analysis focus on <italic>F</italic><sub>1</sub>-score, false-negative rate (FNR), and area under the ROC curve (AUC). Before proceeding, it is worthwhile to mention two additional points regarding the values presented in the table. First, the input elicitation RC represents the probability value of positive classification obtained from the automated classifier when used as a feature. For example, BCE-RC indicates that both the binary elicitation inputs and the probability scores from the ResNet-50 were used as features for the ML classifiers. Second, the Combined Set C&#x00026;D is created by merging the data from Experiment Sets C and D, thereby effectively doubling the size of the training set relative to the individual experiment sets.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Performance analysis of Crowdsourcing-based ML methods with expanded inputs from ResNet-50.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Input</bold><break/> <bold>Elicitations</bold></th>
<th valign="top" align="center"><bold>Size of</bold> <break/> <bold>dataset</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>ResNet50</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>KNN</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>LR</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>RF</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>SVM-Linear</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub></bold></th>
<th valign="top" align="center"><bold>FNR</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="17"><bold>Experiment Set C</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-SE-PDE<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">10k</td>
<td valign="top" align="center">0.36</td>
<td valign="top" align="center">0.21</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center"><bold>0.82</bold></td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center"><bold>0.93</bold></td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.91</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">30k</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center"><bold>0.93</bold></td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">50k</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center"><bold>0.90</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.21</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.98</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center"><bold>0.90</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.21</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.97</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">70k</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center"><bold>0.04</bold></td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center"><bold>0.93</bold></td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center"><bold>0.94</bold></td>
<td valign="top" align="center"><bold>0.04</bold></td>
<td valign="top" align="center">0.99</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center"><bold>0.04</bold></td>
<td valign="top" align="center"><bold>1.00</bold></td>
<td valign="top" align="center"><bold>0.93</bold></td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center"><bold>0.94</bold></td>
<td valign="top" align="center"><bold>0.04</bold></td>
<td valign="top" align="center">0.99</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">90k</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.00</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.98</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.98</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.98</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.99</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.98</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.9</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.98</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.98</bold></td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.99</td>
</tr>
<tr>
<td valign="top" align="left" colspan="17"><bold>Experiment Set D</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.87</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">10k</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.11</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.78</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">30k</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center"><bold>0.33</bold></td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.8</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.87</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">50k</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center"><bold>0.80</bold></td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.96</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center"><bold>0.80</bold></td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.91</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">70k</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">90k</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.95</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.92</td>
</tr>
<tr>
<td valign="top" align="left" colspan="17"><bold>Combined Set C&#x00026;D</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.9</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.9</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">10k</td>
<td valign="top" align="center">0.27</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.5</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center"><bold>0.91</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.25</td>
<td valign="top" align="center">0.9</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center"><bold>0.91</bold></td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">30k</td>
<td valign="top" align="center">0.63</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center"><bold>0.25</bold></td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center"><bold>0.31</bold></td>
<td valign="top" align="center"><bold>0.94</bold></td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center"><bold>0.91</bold></td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center"><bold>0.28</bold></td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center"><bold>0.93</bold></td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">50k</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.87</bold></td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.11</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center">0.14</td>
<td valign="top" align="center">0.97</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center">0.14</td>
<td valign="top" align="center">0.98</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">70k</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.11</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center"><bold>0.06</bold></td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.93</bold></td>
<td valign="top" align="center"><bold>0.06</bold></td>
<td valign="top" align="center">0.97</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.93</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center"><bold>0.06</bold></td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center"><bold>0.08</bold></td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.92</bold></td>
<td valign="top" align="center"><bold>0.06</bold></td>
<td valign="top" align="center">0.97</td>
</tr>
<tr>
<td valign="top" align="left">RC</td>
<td valign="top" align="center">90k</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">BCE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.98</td>
</tr>
<tr>
<td valign="top" align="left">BCE-CE-RC</td>
<td/>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center"><bold>0.97</bold></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.98</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>&#x0002A;Denotes the input combinations that achieved the best performance among the Crowdsourcing-based ML methods. Bold values denote cases where hybrid method outperforms both the Resnet-50 classifier and the Crowdsourcing-based ML methods</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="T6">Table 6</xref> marks in bold those cases in which the performance of the hybrid method according to a given metric is better than both the completely automated approach (ResNet-50) and the results achieved by the crowdsourcing-based ML methods (according to the best input combination). As expected, when the ResNet-50 performance is poor, using its output as a feature hurts the overall results. Conversely, when the ResNet-50 performance is near perfect, it is difficult to improve upon its performance by adding information obtained from the crowd. However, apart from those extremes, exploiting the output of the ResNet-50 is beneficial in most cases, particularly regarding <italic>F</italic><sub>1</sub>-score and AUC.</p>
<p>The proposed hybrid methods, which use the results from the automated classifier as an additional input feature for the crowdsourcing-based ML methods, exhibited a robust performance. They attained maximum <italic>F</italic><sub>1</sub>-scores of 0.98, 0.96 0.97 and minimum FNRs of 0.04, 0.08, 0.06 for Experiment Set C, D, and Combined Set C&#x00026;D, respectively, all of which represent significant improvements over what crowdsourcing-based methods achieved on a standalone basis. While these top results were associated with the automated classifier training set of 90k samples, impressive results were obtained using smaller datasets for Combined Set C&#x00026;D, compared to Experiment Set C and D separately. As an example, incorporating the output of the automated classifier trained on 50k samples with the crowdsourcing-based methods for Combined Set C&#x00026;D improved the <italic>F</italic><sub>1</sub>-score significantly (see <xref ref-type="table" rid="T5">Tables 5</xref>, <xref ref-type="table" rid="T6">6</xref>). However, the hybrid approach did not show better results for Experiment Sets C and D separately over the same training set size in some cases. This can be explained by the fact that Experiment Sets C and D have fewer data points than Combined Set C&#x00026;D. This attests that, while crowdsourcing-based methods supplemented with the outputs of the automated classifier perform very well on small datasets, too few data points can negatively affect the performance of the hybrid approach.</p>
</sec>
<sec sec-type="discussion" id="s7">
<title>7. Discussion</title>
<p>This section highlights key observations related to the research questions, along with the limitations of the study. The experiment results demonstrate that supplementing binary choice elicitation with other forms of inputs can generate better classifiers. When the training sets is small, incorporating binary labels along with confidence values regarding these responses within any of the four ML classifiers tested in this work generated more dependable results for datasets of varying levels of difficulty. These diverse inputs also helped improve other performance metrics such as AUC values, which measure an ML model&#x00027;s capability to distinguishing between labels. While voting methods had a rather poor performance with respect to FNRs, a simple parametric modification (i.e., changing the threshold value) was shown to significantly reduce these values with comparatively small increases to FPRs. When the training sets is larger, integrating the inputs from the automated classifier with the crowdsourcing-based ML methods decreased FNRs even further. Those methods achieved near-perfect FNRs thanks to a large dataset that was used to train the automated classifier. The <italic>F</italic><sub>1</sub>-score was also improved significantly through this hybrid approach. Although smaller training sets of 50k samples slightly reduced the performance of the automated classifier, the numbers were still better than those obtained by standalone crowdsourcing-based methods. Altogether, the results demonstrate that including diverse inputs as features within an ML classifier, it is possible to obtain better classifications at a relatively low cost.</p>
<p>The methodology for aggregating crowd information to improve image classification outcomes presented in this paper could have wide-ranging applications. Through suitable adaptations and enhancements, it could be applied for various types of real-world screening tasks, such as inspecting luggage at travel checkpoints (e.g., airports, metro), X-ray imaging for medical diagnosis, online image labeling, AI model training using CAPTCHAs, etc. Moreover, the image classification problem featured herein is a special case of the overall participant information aggregation problem; therefore, the findings in this paper could be extended to various other classification applications that utilize the wisdom of the crowd concept.</p>
<p>The presented studies admittedly have some limitations. For starters, the approach used to filter &#x0201C;insincere participants" was relatively simple. To obtain a better quality dataset, future studies will seek to deploy more sophisticated quality control techniques for filtering out unreliable or poor quality participants, e.g., using Honeypot questions (Mortensen et al., <xref ref-type="bibr" rid="B50">2017</xref>). A second limitation is that the synthetic images generated for this work have certain characteristics that may overly benefit automated classification methods but may not generalize to various real-world situations. It is possible, for example, that the images might have tiny consistent details that are not visible to human eyes due to the nature of the image generation process. In that case, the automated classification method had an unfair advantage of exploiting those details to improve performance effectively. Future studies will assess the featured methods on more realistic datasets drawn from other practical contexts.</p>
</sec>
<sec sec-type="conclusions" id="s8">
<title>8. Conclusion</title>
<p>Although crowdsourcing methods have been productive in image classification, they do not tap into the full potential of the wisdom of the crowd in one important respect. These methods have largely overlooked the fact that difficult tasks can be amplified to elicit and integrate multiple inputs from each participant; an easy-to-implement option, for example, is eliciting the level of confidence in one&#x00027;s binary response. This paper investigates how different types of information can be utilized with machine learning to enhance the capabilities of crowdsourcing-based classification. It makes four main contributions. First, it introduces a systematic synthetic image generation process that can be used to create image classification tasks of varying difficulty. Second, it demonstrates that while reported confidence in one&#x00027;s response does not significantly raise the performance of voting methods, this intuitive form of input can enhance the performance of machine learning methods, particularly when smaller training datasets are available. Third, it explains how aggregation methods can be adapted to prioritize other metrics of interest of image classification (e.g., reduced false-negative rates). Fourth, it demonstrates that under the right circumstances, automated classifiers can significantly improve classification performance when integrated with crowdsourcing-based methods.</p>
<p>The code used to generate the synthetic images can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/O-ARE/2D-Image-Generation-HCOMP">https://github.com/O-ARE/2D-Image-Generation-HCOMP</ext-link>. In addition, the code used to train and evaluate the automated classifier can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/O-ARE/2d-image-classification">https://github.com/O-ARE/2d-image-classification</ext-link>.</p>
</sec>
<sec sec-type="data-availability" id="s9">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article can be made available by the authors upon request.</p>
</sec>
<sec id="s10">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by Tiffany Dunning, IRB Coordinator, Arizona State University. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s11">
<title>Author Contributions</title>
<p>RY, AE, OF, MH, and JG contributed to the conception and design of the study. HB organized the database. RY, JG, and MH performed the computational analysis. RY wrote the first draft of the manuscript. MH, JG, and HB wrote sections of the manuscript. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s12">
<title>Funding</title>
<p>This material was based upon work supported by the U.S. Department of Homeland Security under Grant Award Number 17STQAC00001-05-00, which all authors gratefully acknowledge. The lead PI of the project (AE) and two of the students (JG and HB) also gratefully acknowledge support from the National Science Foundation under Award Number 1850355. An earlier, shorter version of this paper and a smaller subset of the results featured herein have been published in Yasmin et al. (<xref ref-type="bibr" rid="B71">2021</xref>) and presented in the 9th AAAI Conference on Human Computation and Crowdsourcing.</p>
</sec>
<sec id="s13">
<title>Author Disclaimer</title>
<p>The views and conclusions contained in this document are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of the U.S. Department of Homeland Security or of the National Science Foundation.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s14">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack>
<p>The authors thank all participants in this study, which received institutional IRB approval prior to deployment.</p>
</ack>
<sec sec-type="supplementary-material" id="s15">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2022.848056/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frai.2022.848056/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Assiri</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Stochastic optimization of plain convolutional neural networks with simple methods</article-title>. <source>arXiv [Preprint] arXiv:2001.08856</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2001.08856</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barrington</surname> <given-names>L.</given-names></name> <name><surname>Ghosh</surname> <given-names>S.</given-names></name> <name><surname>Greene</surname> <given-names>M.</given-names></name> <name><surname>Har-Noy</surname> <given-names>S.</given-names></name> <name><surname>Berger</surname> <given-names>J.</given-names></name> <name><surname>Gill</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Crowdsourcing earthquake damage assessment using remote sensing imagery</article-title>. <source>Ann. Geophys</source>. <fpage>54</fpage>. <pub-id pub-id-type="doi">10.4401/ag-5324</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bello</surname> <given-names>I.</given-names></name> <name><surname>Fedus</surname> <given-names>W.</given-names></name> <name><surname>Du</surname> <given-names>X.</given-names></name> <name><surname>Cubuk</surname> <given-names>E. D.</given-names></name> <name><surname>Srinivas</surname> <given-names>A.</given-names></name> <name><surname>Lin</surname> <given-names>T.-Y.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Revisiting resnets: improved training and scaling strategies</article-title>. <source>arXiv [Preprint] arXiv:2103.07579</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2103.07579</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brandt</surname> <given-names>F.</given-names></name> <name><surname>Conitzer</surname> <given-names>V.</given-names></name> <name><surname>Endriss</surname> <given-names>U.</given-names></name> <name><surname>Lang</surname> <given-names>J.</given-names></name> <name><surname>Procaccia</surname> <given-names>A. D.</given-names></name></person-group> (<year>2016</year>). <source>Handbook of Computational Social Choice</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/CBO9781107446984</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>J. C.</given-names></name> <name><surname>Amershi</surname> <given-names>S.</given-names></name> <name><surname>Kamar</surname> <given-names>E.</given-names></name></person-group> (<year>2017</year>). <article-title>Revolt: collaborative crowdsourcing for labeling machine learning datasets</article-title>, in <source>Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>2334</fpage>&#x02013;<lpage>2346</lpage>. <pub-id pub-id-type="doi">10.1145/3025453.3026044</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheplygina</surname> <given-names>V.</given-names></name> <name><surname>de Bruijne</surname> <given-names>M.</given-names></name> <name><surname>Pluim</surname> <given-names>J. P.</given-names></name></person-group> (<year>2019</year>). <article-title>Not-so-supervised: a survey of semi-supervised, multi-instance, and transfer learning in medical image analysis</article-title>. <source>Med. Image Anal</source>. <volume>54</volume>, <fpage>280</fpage>&#x02013;<lpage>296</lpage>. <pub-id pub-id-type="doi">10.1016/j.media.2019.03.009</pub-id><pub-id pub-id-type="pmid">30959445</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cheplygina</surname> <given-names>V.</given-names></name> <name><surname>Pluim</surname> <given-names>J. P.</given-names></name></person-group> (<year>2018</year>). <article-title>Crowd disagreement about medical images is informative</article-title>, in <source>Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis</source>, eds <person-group person-group-type="editor"><name><surname>Stoyanov</surname> <given-names>D.</given-names></name> <name><surname>Taylor</surname> <given-names>Z.</given-names></name> <name><surname>Balocco</surname> <given-names>S.</given-names></name> <name><surname>Sznitman</surname> <given-names>R.</given-names></name> <name><surname>Martel</surname> <given-names>A.</given-names></name> <name><surname>Maier-Hein</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<publisher-loc>Granada</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>105</fpage>&#x02013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01364-6_12</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Christoforou</surname> <given-names>E.</given-names></name> <name><surname>Fern&#x000E1;ndez Anta</surname> <given-names>A.</given-names></name> <name><surname>S&#x000E1;nchez</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>An experimental characterization of workers&#x00027; behavior and accuracy in crowdsourced tasks</article-title>. <source>PLoS ONE</source> <volume>16</volume>:<fpage>e0252604</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0252604</pub-id><pub-id pub-id-type="pmid">34133447</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dai</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name> <name><surname>Tan</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Coatnet: Marrying convolution and attention for all data sizes</article-title>. <source>ArXiv, abs/2106.04803</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2106.04803</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.-J.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>Imagenet: A large-scale hierarchical image database</article-title>, in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>), <fpage>248</fpage>&#x02013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206848</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dodge</surname> <given-names>S.</given-names></name> <name><surname>Karam</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>A study and comparison of human and deep learning recognition performance under visual distortions</article-title>, in <source>2017 26th International Conference on Computer Communication and Networks (ICCCN)</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1109/ICCCN.2017.8038465</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eickhoff</surname> <given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>Cognitive biases in crowdsourcing</article-title>, in <source>Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>162</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1145/3159652.3159654</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Escobedo</surname> <given-names>A. R.</given-names></name> <name><surname>Moreno-Centeno</surname> <given-names>E.</given-names></name> <name><surname>Yasmin</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>An axiomatic distance methodology for aggregating multimodal evaluations</article-title>. <source>Inform. Sci</source>. <volume>590</volume>, <fpage>322</fpage>&#x02013;<lpage>345</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2021.12.124</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ester</surname> <given-names>M.</given-names></name> <name><surname>Kriegel</surname> <given-names>H.-P.</given-names></name> <name><surname>Sander</surname> <given-names>J.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>1996</year>). <article-title>A density-based algorithm for discovering clusters in large spatial databases with noise</article-title>, in <source>KDD, Vol. 96</source> (<publisher-loc>Portland, OR</publisher-loc>), <fpage>226</fpage>&#x02013;<lpage>231</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Everingham</surname> <given-names>M.</given-names></name> <name><surname>Van Gool</surname> <given-names>L.</given-names></name> <name><surname>Williams</surname> <given-names>C. K.</given-names></name> <name><surname>Winn</surname> <given-names>J.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>The pascal visual object classes (VOC) challenge</article-title>. <source>Int. J. Comput. Vis</source>. <volume>88</volume>, <fpage>303</fpage>&#x02013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-009-0275-4</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Foody</surname> <given-names>G.</given-names></name> <name><surname>See</surname> <given-names>L.</given-names></name> <name><surname>Fritz</surname> <given-names>S.</given-names></name> <name><surname>Moorthy</surname> <given-names>I.</given-names></name> <name><surname>Perger</surname> <given-names>C.</given-names></name> <name><surname>Schill</surname> <given-names>C.</given-names></name> <name><surname>Boyd</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>Increasing the accuracy of crowdsourced information on land cover <italic>via</italic> a voting procedure weighted by information inferred from the contributed data</article-title>. <source>ISPRS Int. J. Geo-Inform</source>. <volume>7</volume>:<fpage>80</fpage>. <pub-id pub-id-type="doi">10.3390/ijgi7030080</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Geirhos</surname> <given-names>R.</given-names></name> <name><surname>Janssen</surname> <given-names>D. H.</given-names></name> <name><surname>Sch&#x000FC;tt</surname> <given-names>H. H.</given-names></name> <name><surname>Rauber</surname> <given-names>J.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name> <name><surname>Wichmann</surname> <given-names>F. A.</given-names></name></person-group> (<year>2017</year>). <article-title>Comparing deep neural networks against humans: object recognition when the signal gets weaker</article-title>. <source>arXiv [Preprint] arXiv:1706.06969</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1706.06969</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gennatas</surname> <given-names>E. D.</given-names></name> <name><surname>Friedman</surname> <given-names>J. H.</given-names></name> <name><surname>Ungar</surname> <given-names>L. H.</given-names></name> <name><surname>Pirracchio</surname> <given-names>R.</given-names></name> <name><surname>Eaton</surname> <given-names>E.</given-names></name> <name><surname>Reichmann</surname> <given-names>L. G.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Expert-augmented machine learning</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>117</volume>, <fpage>4571</fpage>&#x02013;<lpage>4577</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1906831117</pub-id><pub-id pub-id-type="pmid">32071251</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>G&#x000F6;rzen</surname> <given-names>T.</given-names></name> <name><surname>Laux</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2019</year>). <source>Extracting the Wisdom From the Crowd: A Comparison of Approaches to Aggregating Collective Intelligence</source>. Technical Report, <publisher-name>Paderborn University, Faculty of Business Administration and Economics</publisher-name>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Griffin</surname> <given-names>D.</given-names></name> <name><surname>Tversky</surname> <given-names>A.</given-names></name></person-group> (<year>1992</year>). <article-title>The weighing of evidence and the determinants of confidence</article-title>. <source>Cogn. Psychol</source>. <volume>24</volume>, <fpage>411</fpage>&#x02013;<lpage>435</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0285(92)90013-R</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grofman</surname> <given-names>B.</given-names></name> <name><surname>Owen</surname> <given-names>G.</given-names></name> <name><surname>Feld</surname> <given-names>S. L.</given-names></name></person-group> (<year>1983</year>). <article-title>Thirteen theorems in search of the truth</article-title>. <source>Theory Decis</source>. <volume>15</volume>, <fpage>261</fpage>&#x02013;<lpage>278</lpage>. <pub-id pub-id-type="doi">10.1007/BF00125672</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gurari</surname> <given-names>D.</given-names></name> <name><surname>Theriault</surname> <given-names>D.</given-names></name> <name><surname>Sameki</surname> <given-names>M.</given-names></name> <name><surname>Isenberg</surname> <given-names>B.</given-names></name> <name><surname>Pham</surname> <given-names>T. A.</given-names></name> <name><surname>Purwada</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>How to collect segmentations for biomedical images? A benchmark evaluating the performance of experts, crowdsourced non-experts, and algorithms</article-title>, in <source>2015 IEEE Winter Conference on Applications of Computer Vision</source> (<publisher-loc>Austin, TX</publisher-loc>), <fpage>1169</fpage>&#x02013;<lpage>1176</lpage>. <pub-id pub-id-type="doi">10.1109/WACV.2015.160</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hamada</surname> <given-names>D.</given-names></name> <name><surname>Nakayama</surname> <given-names>M.</given-names></name> <name><surname>Saiki</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Wisdom of crowds and collective decision-making in a survival situation with complex information integration</article-title>. <source>Cogn. Res</source>. <volume>5</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1186/s41235-020-00248-z</pub-id><pub-id pub-id-type="pmid">33057843</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hara</surname> <given-names>K.</given-names></name> <name><surname>Le</surname> <given-names>V.</given-names></name> <name><surname>Froehlich</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>A feasibility study of crowdsourcing and google street view to determine sidewalk accessibility</article-title>, in <source>Proceedings of the 14th International ACM SIGACCESS Conference on Computers and Accessibility</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>273</fpage>&#x02013;<lpage>274</lpage>. <pub-id pub-id-type="doi">10.1145/2384916.2384989</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>R.</given-names></name> <name><surname>Kameda</surname> <given-names>T.</given-names></name></person-group> (<year>2005</year>). <article-title>The robust beauty of majority rules in group decisions</article-title>. <source>Psychol. Rev</source>. <volume>112</volume>:<fpage>494</fpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.112.2.494</pub-id><pub-id pub-id-type="pmid">15783295</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>van Ossenbruggen</surname> <given-names>J.</given-names></name> <name><surname>de Vries</surname> <given-names>A. P.</given-names></name></person-group> (<year>2013</year>). <article-title>Do you need experts in the crowd? A case study in image annotation for marine biology</article-title>, in <source>Proceedings of the 10th Conference on Open Research Areas in Information Retrieval</source> (<publisher-loc>Lisbon</publisher-loc>), <fpage>57</fpage>&#x02013;<lpage>60</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep residual learning for image recognition</article-title>. <source>arXiv [Preprint] arXiv:1512.03385</source>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id><pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hekler</surname> <given-names>A.</given-names></name> <name><surname>Utikal</surname> <given-names>J. S.</given-names></name> <name><surname>Enk</surname> <given-names>A. H.</given-names></name> <name><surname>Hauschild</surname> <given-names>A.</given-names></name> <name><surname>Weichenthal</surname> <given-names>M.</given-names></name> <name><surname>Maron</surname> <given-names>R. C.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Superior skin cancer classification by the combination of human and artificial intelligence</article-title>. <source>Eur. J. Cancer</source> <volume>120</volume>, <fpage>114</fpage>&#x02013;<lpage>121</lpage>.<pub-id pub-id-type="pmid">31518967</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hsing</surname> <given-names>P.-Y.</given-names></name> <name><surname>Bradley</surname> <given-names>S.</given-names></name> <name><surname>Kent</surname> <given-names>V. T.</given-names></name> <name><surname>Hill</surname> <given-names>R. A.</given-names></name> <name><surname>Smith</surname> <given-names>G. C.</given-names></name> <name><surname>Whittingham</surname> <given-names>M. J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Economical crowdsourcing for camera trap image classification</article-title>. <source>Remote Sens. Ecol. Conserv</source>. <volume>4</volume>, <fpage>361</fpage>&#x02013;<lpage>374</lpage>. <pub-id pub-id-type="doi">10.1002/rse2.84</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ipeirotis</surname> <given-names>P. G.</given-names></name> <name><surname>Provost</surname> <given-names>F.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Quality management on amazon mechanical Turk</article-title>, in <source>Proceedings of the ACM SIGKDD Workshop on Human Computation</source> (<publisher-loc>Washington, DC</publisher-loc>), <fpage>64</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1145/1837885.1837906</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Irshad</surname> <given-names>H.</given-names></name> <name><surname>Oh</surname> <given-names>E.-Y.</given-names></name> <name><surname>Schmolze</surname> <given-names>D.</given-names></name> <name><surname>Quintana</surname> <given-names>L. M.</given-names></name> <name><surname>Collins</surname> <given-names>L.</given-names></name> <name><surname>Tamimi</surname> <given-names>R. M.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Crowdsourcing scoring of immunohistochemistry images: evaluating performance of the crowd and an automated computational method</article-title>. <source>Sci. Rep</source>. <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1038/srep43286</pub-id><pub-id pub-id-type="pmid">28230179</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jeannin</surname> <given-names>S.</given-names></name> <name><surname>Bober</surname> <given-names>M.</given-names></name></person-group> (<year>1999</year>). <source>Description of Core Experiments for MPEG-7 Motion/Shape</source>. MPEG-7, ISO/IEC/JTC1/SC29/WG11/MPEG99 N, 2690.</citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Karger</surname> <given-names>D. R.</given-names></name> <name><surname>Oh</surname> <given-names>S.</given-names></name> <name><surname>Shah</surname> <given-names>D.</given-names></name></person-group> (<year>2011</year>). <article-title>Iterative learning for reliable crowdsourcing systems</article-title>, in <source>Neural Information Processing Systems</source> (<publisher-loc>Granada</publisher-loc>).</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kemmer</surname> <given-names>R.</given-names></name> <name><surname>Yoo</surname> <given-names>Y.</given-names></name> <name><surname>Escobedo</surname> <given-names>A.</given-names></name> <name><surname>Maciejewski</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Enhancing collective estimates by aggregating cardinal and ordinal inputs</article-title>, in <source>Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 8</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>73</fpage>&#x02013;<lpage>82</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Khattak</surname> <given-names>F. K.</given-names></name> <name><surname>Salleb-Aouissi</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Quality control of crowd labeling through expert evaluation</article-title>, in <source>Proceedings of the NIPS 2nd Workshop on Computational Social Science and the Wisdom of Crowds, Vol. 2</source> (<publisher-loc>Sierra Nevada</publisher-loc>), <fpage>5</fpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Koh</surname> <given-names>W. L.</given-names></name> <name><surname>Kaliappan</surname> <given-names>J.</given-names></name> <name><surname>Rice</surname> <given-names>M.</given-names></name> <name><surname>Ma</surname> <given-names>K.-T.</given-names></name> <name><surname>Tay</surname> <given-names>H. H.</given-names></name> <name><surname>Tan</surname> <given-names>W. P.</given-names></name></person-group> (<year>2017</year>). <article-title>Preliminary investigation of augmented intelligence for remote assistance using a wearable display</article-title>, in <source>TENCON 2017-2017 IEEE Region 10 Conference</source> (<publisher-loc>Penang</publisher-loc>), <fpage>2093</fpage>&#x02013;<lpage>2098</lpage>. <pub-id pub-id-type="doi">10.1109/TENCON.2017.8228206</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koriat</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>The self-consistency model of subjective confidence</article-title>. <source>Psychol. Rev</source>. <volume>119</volume>:<fpage>80</fpage>. <pub-id pub-id-type="doi">10.1037/a0025648</pub-id><pub-id pub-id-type="pmid">22022833</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>, in <source>Advances in Neural Information Processing Systems, Vol. 25</source>, eds <person-group person-group-type="editor"><name><surname>Pereira</surname> <given-names>F.</given-names></name> <name><surname>Burges</surname> <given-names>C. J. C.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<publisher-loc>Penang</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>).</citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Varshney</surname> <given-names>P. K.</given-names></name></person-group> (<year>2017</year>). <article-title>Does confidence reporting from the crowd benefit crowdsourcing performance?</article-title> in <source>Proceedings of the 2nd International Workshop on Social Sensing</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>49</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1145/3055601.3055607</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Litvinova</surname> <given-names>A.</given-names></name> <name><surname>Herzog</surname> <given-names>S. M.</given-names></name> <name><surname>Kall</surname> <given-names>A. A.</given-names></name> <name><surname>Pleskac</surname> <given-names>T. J.</given-names></name> <name><surname>Hertwig</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>How the &#x0201C;wisdom of the inner crowd&#x0201D; can boost accuracy of confidence judgments</article-title>. <source>Decision</source> <volume>7</volume>:<fpage>183</fpage>. <pub-id pub-id-type="doi">10.1037/dec0000119</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mannes</surname> <given-names>A. E.</given-names></name> <name><surname>Soll</surname> <given-names>J. B.</given-names></name> <name><surname>Larrick</surname> <given-names>R. P.</given-names></name></person-group> (<year>2014</year>). <article-title>The wisdom of select crowds</article-title>. <source>J. Pers. Soc. Psychol</source>. <volume>107</volume>:<fpage>276</fpage>. <pub-id pub-id-type="doi">10.1037/a0036677</pub-id><pub-id pub-id-type="pmid">25090129</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mao</surname> <given-names>A.</given-names></name> <name><surname>Procaccia</surname> <given-names>A. D.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>Better human computation through principled voting</article-title>, in <source>AAAI</source> (<publisher-loc>Bellevue, WA</publisher-loc>).</citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Matoulkova</surname> <given-names>B. K.</given-names></name></person-group> (<year>2017</year>). <source>Wisdom of the crowd: comparison of the CWM, simple average and surprisingly popular answer method</source> (Master&#x00027;s thesis). <publisher-name>Erasmus University Rotterdam</publisher-name>, <publisher-loc>Rotterdam, Netherlands</publisher-loc>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mavandadi</surname> <given-names>S.</given-names></name> <name><surname>Dimitrov</surname> <given-names>S.</given-names></name> <name><surname>Feng</surname> <given-names>S.</given-names></name> <name><surname>Yu</surname> <given-names>F.</given-names></name> <name><surname>Sikora</surname> <given-names>U.</given-names></name> <name><surname>Yaglidere</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Distributed medical image analysis and diagnosis through crowd-sourced games: a malaria case study</article-title>. <source>PLoS ONE</source> <volume>7</volume>:<fpage>e37245</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0037245</pub-id><pub-id pub-id-type="pmid">22606353</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McDaniel</surname> <given-names>P.</given-names></name> <name><surname>Papernot</surname> <given-names>N.</given-names></name> <name><surname>Celik</surname> <given-names>Z. B.</given-names></name></person-group> (<year>2016</year>). <article-title>Machine learning in adversarial settings</article-title>. <source>IEEE Secur. Privacy</source> <volume>14</volume>, <fpage>68</fpage>&#x02013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.1109/MSP.2016.51</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meyen</surname> <given-names>S.</given-names></name> <name><surname>Sigg</surname> <given-names>D. M.</given-names></name> <name><surname>von Luxburg</surname> <given-names>U.</given-names></name> <name><surname>Franz</surname> <given-names>V. H.</given-names></name></person-group> (<year>2021</year>). <article-title>Group decisions based on confidence weighted majority voting</article-title>. <source>Cogn. Res</source>. <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1186/s41235-021-00279-0</pub-id><pub-id pub-id-type="pmid">33721120</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mitry</surname> <given-names>D.</given-names></name> <name><surname>Peto</surname> <given-names>T.</given-names></name> <name><surname>Hayat</surname> <given-names>S.</given-names></name> <name><surname>Morgan</surname> <given-names>J. E.</given-names></name> <name><surname>Khaw</surname> <given-names>K.-T.</given-names></name> <name><surname>Foster</surname> <given-names>P. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Crowdsourcing as a novel technique for retinal fundus photography classification: analysis of images in the epic norfolk cohort on behalf of the ukbiobank eye and vision consortium</article-title>. <source>PLoS ONE</source> <volume>8</volume>:<fpage>e71154</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0071154</pub-id><pub-id pub-id-type="pmid">23990935</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mitry</surname> <given-names>D.</given-names></name> <name><surname>Zutis</surname> <given-names>K.</given-names></name> <name><surname>Dhillon</surname> <given-names>B.</given-names></name> <name><surname>Peto</surname> <given-names>T.</given-names></name> <name><surname>Hayat</surname> <given-names>S.</given-names></name> <name><surname>Khaw</surname> <given-names>K.-T.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>The accuracy and reliability of crowdsource annotations of digital retinal images</article-title>. <source>Transl. Vis. Sci. Technol</source>. <volume>5</volume>:<fpage>6</fpage>. <pub-id pub-id-type="doi">10.1167/tvst.5.5.6</pub-id><pub-id pub-id-type="pmid">27668130</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mora</surname> <given-names>D.</given-names></name> <name><surname>Zimmermann</surname> <given-names>R.</given-names></name> <name><surname>Cirqueira</surname> <given-names>D.</given-names></name> <name><surname>Bezbradica</surname> <given-names>M.</given-names></name> <name><surname>Helfert</surname> <given-names>M.</given-names></name> <name><surname>Auinger</surname> <given-names>A.</given-names></name> <name><surname>Werth</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Who wants to use an augmented reality shopping assistant application?&#x0201D;</article-title> in <source>Proceedings of the 4th International Conference on Computer-Human Interaction Research and Applications - WUDESHI-DR</source> (<publisher-loc>SciTePress</publisher-loc>), <fpage>309</fpage>&#x02013;<lpage>318</lpage>. <pub-id pub-id-type="doi">10.5220/0010214503090318</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mortensen</surname> <given-names>M. L.</given-names></name> <name><surname>Adam</surname> <given-names>G. P.</given-names></name> <name><surname>Trikalinos</surname> <given-names>T. A.</given-names></name> <name><surname>Kraska</surname> <given-names>T.</given-names></name> <name><surname>Wallace</surname> <given-names>B. C.</given-names></name></person-group> (<year>2017</year>). <article-title>An exploration of crowdsourcing citation screening for systematic reviews</article-title>. <source>Res. Synthes. Methods</source> <volume>8</volume>, <fpage>366</fpage>&#x02013;<lpage>386</lpage>. <pub-id pub-id-type="doi">10.1002/jrsm.1252</pub-id><pub-id pub-id-type="pmid">28677322</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nguyen</surname> <given-names>T. B.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Anugu</surname> <given-names>V.</given-names></name> <name><surname>Rose</surname> <given-names>N.</given-names></name> <name><surname>McKenna</surname> <given-names>M.</given-names></name> <name><surname>Petrick</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Distributed human intelligence for colonic polyp classification in computer-aided detection for CT colonography</article-title>. <source>Radiology</source> <volume>262</volume>, <fpage>824</fpage>&#x02013;<lpage>833</lpage>. <pub-id pub-id-type="doi">10.1148/radiol.11110938</pub-id><pub-id pub-id-type="pmid">22274839</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nowak</surname> <given-names>S.</given-names></name> <name><surname>R&#x000FC;ger</surname> <given-names>S.</given-names></name></person-group> (<year>2010</year>). <article-title>How reliable are annotations <italic>via</italic> crowdsourcing: a study about inter-annotator agreement for multi-label image annotation</article-title>, in <source>Proceedings of the International Conference on Multimedia Information Retrieval</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>557</fpage>&#x02013;<lpage>566</lpage>. <pub-id pub-id-type="doi">10.1145/1743384.1743478</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oosterman</surname> <given-names>J.</given-names></name> <name><surname>Nottamkandath</surname> <given-names>A.</given-names></name> <name><surname>Dijkshoorn</surname> <given-names>C.</given-names></name> <name><surname>Bozzon</surname> <given-names>A.</given-names></name> <name><surname>Houben</surname> <given-names>G.-J.</given-names></name> <name><surname>Aroyo</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Crowdsourcing knowledge-intensive tasks in cultural heritage</article-title>, in <source>Proceedings of the 2014 ACM Conference on Web Science</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>267</fpage>&#x02013;<lpage>268</lpage>. <pub-id pub-id-type="doi">10.1145/2615569.2615644</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oyama</surname> <given-names>S.</given-names></name> <name><surname>Baba</surname> <given-names>Y.</given-names></name> <name><surname>Sakurai</surname> <given-names>Y.</given-names></name> <name><surname>Kashima</surname> <given-names>H.</given-names></name></person-group> (<year>2013</year>). <article-title>Accurate integration of crowdsourced labels using workers&#x00027; self-reported confidence scores</article-title>, in <source>Twenty-Third International Joint Conference on Artificial Intelligence</source> (<publisher-loc>Beijing</publisher-loc>).</citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Papernot</surname> <given-names>N.</given-names></name> <name><surname>McDaniel</surname> <given-names>P.</given-names></name> <name><surname>Jha</surname> <given-names>S.</given-names></name> <name><surname>Fredrikson</surname> <given-names>M.</given-names></name> <name><surname>Celik</surname> <given-names>Z. B.</given-names></name> <name><surname>Swami</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>The limitations of deep learning in adversarial settings</article-title>, in <source>2016 IEEE European Symposium on Security and Privacy (EuroS&#x00026;P)</source> (<publisher-loc>Saarbr&#x000FC;cken</publisher-loc>), <fpage>372</fpage>&#x02013;<lpage>387</lpage>. <pub-id pub-id-type="doi">10.1109/EuroSP.2016.36</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedregosa</surname> <given-names>F.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Gramfort</surname> <given-names>A.</given-names></name> <name><surname>Michel</surname> <given-names>V.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Grisel</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Scikit-learn: machine learning in Python</article-title>. <source>J. Mach. Learn. Res</source>. <volume>12</volume>, <fpage>2825</fpage>&#x02013;<lpage>2830</lpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1201.0490</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prelec</surname> <given-names>D.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name> <name><surname>McCoy</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>A solution to the single-question crowd wisdom problem</article-title>. <source>Nature</source> <volume>541</volume>:<fpage>532</fpage>. <pub-id pub-id-type="doi">10.1038/nature21054</pub-id><pub-id pub-id-type="pmid">28128245</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Ralph</surname> <given-names>R</given-names></name></person-group>. <article-title>Mpeg-7 core experiment ce-shape-1 test set</article-title>. (<year>1999</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://dabi.temple.edu/external/shape/MPEG7/dataset.html">https://dabi.temple.edu/external/shape/MPEG7/dataset.html</ext-link></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rankin</surname> <given-names>W. L.</given-names></name> <name><surname>Grube</surname> <given-names>J. W.</given-names></name></person-group> (<year>1980</year>). <article-title>A comparison of ranking and rating procedures for value system measurement</article-title>. <source>Eur. J. Soc. Psychol</source>. <volume>10</volume>, <fpage>233</fpage>&#x02013;<lpage>246</lpage>. <pub-id pub-id-type="doi">10.1002/ejsp.2420100303</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rasp</surname> <given-names>S.</given-names></name> <name><surname>Schulz</surname> <given-names>H.</given-names></name> <name><surname>Bony</surname> <given-names>S.</given-names></name> <name><surname>Stevens</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Combining crowdsourcing and deep learning to explore the mesoscale organization of shallow convection</article-title>. <source>Bull. Am. Meteorol. Soc</source>. <volume>101</volume>, <fpage>E1980</fpage>&#x02013;<lpage>E1995</lpage>. <pub-id pub-id-type="doi">10.1175/BAMS-D-19-0324.1</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russakovsky</surname> <given-names>O.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Krause</surname> <given-names>J.</given-names></name> <name><surname>Satheesh</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Imagenet large scale visual recognition challenge</article-title>. <source>Int. J. Comput. Vis</source>. <volume>115</volume>, <fpage>211</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-015-0816-y</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saab</surname> <given-names>F.</given-names></name> <name><surname>Elhajj</surname> <given-names>I. H.</given-names></name> <name><surname>Kayssi</surname> <given-names>A.</given-names></name> <name><surname>Chehab</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Modelling cognitive bias in crowdsourcing systems</article-title>. <source>Cogn. Syst. Res</source>. <volume>58</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.cogsys.2019.04.004</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saha Roy</surname> <given-names>T.</given-names></name> <name><surname>Mazumder</surname> <given-names>S.</given-names></name> <name><surname>Das</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>Wisdom of crowds benefits perceptual decision making across difficulty levels</article-title>. <source>Sci. Rep</source>. <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-80500-0</pub-id><pub-id pub-id-type="pmid">33436921</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Salek</surname> <given-names>M.</given-names></name> <name><surname>Bachrach</surname> <given-names>Y.</given-names></name> <name><surname>Key</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>Hotspotting-a probabilistic graphical model for image object localization through crowdsourcing</article-title>, in <source>Twenty-Seventh AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Bellevue, WA</publisher-loc>).</citation>
</ref>
<ref id="B65">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sokolova</surname> <given-names>M.</given-names></name> <name><surname>Japkowicz</surname> <given-names>N.</given-names></name> <name><surname>Szpakowicz</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>Beyond accuracy, f-score and roc: a family of discriminant measures for performance evaluation</article-title>, in <source>Australasian Joint Conference on Artificial Intelligence</source> (<publisher-loc>Hobart, TAS</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>1015</fpage>&#x02013;<lpage>1021</lpage>. <pub-id pub-id-type="doi">10.1007/11941439_114</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stevens</surname> <given-names>B.</given-names></name> <name><surname>Bony</surname> <given-names>S.</given-names></name> <name><surname>Brogniez</surname> <given-names>H.</given-names></name> <name><surname>Hentgen</surname> <given-names>L.</given-names></name> <name><surname>Hohenegger</surname> <given-names>C.</given-names></name> <name><surname>Kiemle</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Sugar, gravel, fish and flowers: mesoscale cloud patterns in the trade winds</article-title>. <source>Q. J. R. Meteorol. Soc</source>. <volume>146</volume>, <fpage>141</fpage>&#x02013;<lpage>152</lpage>. <pub-id pub-id-type="doi">10.1002/qj.3662</pub-id></citation>
</ref>
<ref id="B67">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Surowiecki</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <source>The Wisdom of Crowds</source>. <publisher-loc>Anchor</publisher-loc>.</citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swanson</surname> <given-names>A.</given-names></name> <name><surname>Kosmala</surname> <given-names>M.</given-names></name> <name><surname>Lintott</surname> <given-names>C.</given-names></name> <name><surname>Simpson</surname> <given-names>R.</given-names></name> <name><surname>Smith</surname> <given-names>A.</given-names></name> <name><surname>Packer</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Snapshot serengeti, high-frequency annotated camera trap images of 40 mammalian species in an African savanna</article-title>. <source>Sci. Data</source> <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1038/sdata.2015.26</pub-id><pub-id pub-id-type="pmid">26097743</pub-id></citation></ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>M.</given-names></name> <name><surname>Le</surname> <given-names>Q.</given-names></name></person-group> (<year>2019</year>). <article-title>EfficientNet: rethinking model scaling for convolutional neural networks</article-title>, in <source>International Conference on Machine Learning</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>6105</fpage>&#x02013;<lpage>6114</lpage>.</citation>
</ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>A.</given-names></name> <name><surname>Bailey</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <article-title>A reference-based scoring model for increasing the findability of promising ideas in innovation pipelines</article-title>, in <source>Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>1183</fpage>&#x02013;<lpage>1186</lpage>. <pub-id pub-id-type="doi">10.1145/2145204.2145380</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yasmin</surname> <given-names>R.</given-names></name> <name><surname>Grassel</surname> <given-names>J. T.</given-names></name> <name><surname>Hassan</surname> <given-names>M. M.</given-names></name> <name><surname>Fuentes</surname> <given-names>O.</given-names></name> <name><surname>Escobedo</surname> <given-names>A. R.</given-names></name></person-group> (<year>2021</year>). <article-title>Enhancing image classification capabilities of crowdsourcing-based methods through expanded input elicitation</article-title>, in <source>Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 9</source> (<publisher-loc>Virtual</publisher-loc>), <fpage>166</fpage>&#x02013;<lpage>178</lpage>.</citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yi</surname> <given-names>S. K. M.</given-names></name> <name><surname>Steyvers</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>M. D.</given-names></name> <name><surname>Dry</surname> <given-names>M. J.</given-names></name></person-group> (<year>2012</year>). <article-title>The wisdom of the crowd in combinatorial problems</article-title>. <source>Cogn. Sci</source>. <volume>36</volume>, <fpage>452</fpage>&#x02013;<lpage>470</lpage>. <pub-id pub-id-type="doi">10.1111/j.1551-6709.2011.01223.x</pub-id><pub-id pub-id-type="pmid">22268680</pub-id></citation></ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoo</surname> <given-names>Y.</given-names></name> <name><surname>Escobedo</surname> <given-names>A.</given-names></name> <name><surname>Skolfield</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>A new correlation coefficient for comparing and aggregating non-strict and incomplete rankings</article-title>. <source>Eur. J. Oper. Res</source>. <volume>285</volume>, <fpage>1025</fpage>&#x02013;<lpage>1041</lpage>. <pub-id pub-id-type="doi">10.1016/j.ejor.2020.02.027</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhai</surname> <given-names>X.</given-names></name> <name><surname>Kolesnikov</surname> <given-names>A.</given-names></name> <name><surname>Houlsby</surname> <given-names>N.</given-names></name> <name><surname>Beyer</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Scaling vision transformers</article-title>. <source>arXiv [Preprint] arXiv:2106.04560</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2106.04560</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Lapedriza</surname> <given-names>A.</given-names></name> <name><surname>Xiao</surname> <given-names>J.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Learning deep features for scene recognition using places database</article-title>. <source>Adv. Neural Inform. Process. Syst</source>. <volume>27</volume>, <fpage>487</fpage>&#x02013;<lpage>495</lpage>. <pub-id pub-id-type="doi">10.1101/265918</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>N.</given-names></name> <name><surname>Siegel</surname> <given-names>Z. D.</given-names></name> <name><surname>Zarecor</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>N.</given-names></name> <name><surname>Campbell</surname> <given-names>D. A.</given-names></name> <name><surname>Andorf</surname> <given-names>C. M.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Crowdsourcing image analysis for plant phenomics to generate ground truth data for machine learning</article-title>. <source>PLoS Comput. Biol</source>. <volume>14</volume>:<fpage>e1006337</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1006337</pub-id><pub-id pub-id-type="pmid">30059508</pub-id></citation></ref>
</ref-list>
</back>
</article>