<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Radiol.</journal-id>
<journal-title>Frontiers in Radiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Radiol.</abbrev-journal-title>
<issn pub-type="epub">2673-8740</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fradi.2023.1088068</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Radiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Localization supervision of chest x-ray classifiers using label-specific eye-tracking annotation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes"><name><surname>Bigolin Lanfredi</surname><given-names>Ricardo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref><uri xlink:href="https://loop.frontiersin.org/people/2039991/overview" /></contrib>
<contrib contrib-type="author"><name><surname>Schroeder</surname><given-names>Joyce D.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/2336909/overview"/></contrib>
<contrib contrib-type="author"><name><surname>Tasdizen</surname><given-names>Tolga</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/28741/overview"/></contrib>
</contrib-group>
<aff id="aff1"><label><sup>1</sup></label><addr-line>Scientific Computing and Imaging Institute</addr-line>, <institution>University of Utah</institution>, <addr-line>Salt Lake City, UT</addr-line>, <country>United States</country></aff>
<aff id="aff2"><label><sup>2</sup></label><addr-line>Department of Radiology and Imaging Sciences</addr-line>, <institution>University of Utah</institution>, <addr-line>Salt Lake City, UT</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p><bold>Edited by:</bold> Henning M&#x00FC;ller, University of Applied Sciences Western Switzerland (HES-SO Valais), Switzerland</p></fn>
<fn fn-type="edited-by"><p><bold>Reviewed by:</bold> Yashin Dicente Cid, Roche Diagnostics S.L., Spain, Alexandros Karargyris, H&#x00F4;pitaux Universitaires de Strasbourg, France</p></fn>
<corresp id="cor1"><label>&#x002A;</label><bold>Correspondence:</bold> Ricardo Bigolin Lanfredi <email>ricbl@sci.utah.edu</email></corresp>
</author-notes>
<pub-date pub-type="epub"><day>22</day><month>06</month><year>2023</year></pub-date>
<pub-date pub-type="collection"><year>2023</year></pub-date>
<volume>3</volume><elocation-id>1088068</elocation-id>
<history>
<date date-type="received"><day>03</day><month>11</month><year>2022</year></date>
<date date-type="accepted"><day>05</day><month>06</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Bigolin Lanfredi, Schroeder and Tasdizen.</copyright-statement>
<copyright-year>2023</copyright-year><copyright-holder>Bigolin Lanfredi, Schroeder and Tasdizen</copyright-holder><license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Convolutional neural networks (CNNs) have been successfully applied to chest x-ray (CXR) images. Moreover, annotated bounding boxes have been shown to improve the interpretability of a CNN in terms of localizing abnormalities. However, only a few relatively small CXR datasets containing bounding boxes are available, and collecting them is very costly. Opportunely, eye-tracking (ET) data can be collected during the clinical workflow of a radiologist. We use ET data recorded from radiologists while dictating CXR reports to train CNNs. We extract snippets from the ET data by associating them with the dictation of keywords and use them to supervise the localization of specific abnormalities. We show that this method can improve a model&#x2019;s interpretability without impacting its image-level classification.</p>
</abstract>
<kwd-group>
<kwd>eye tracking</kwd>
<kwd>chest x-ray (CXR)</kwd>
<kwd>interpretability</kwd>
<kwd>annotation</kwd>
<kwd>localization</kwd>
<kwd>gaze</kwd>
</kwd-group>
<contract-num rid="cn001">R21EB028367</contract-num>
<contract-sponsor id="cn001">National Institute of Biomedical Imaging and Bioengineering<named-content content-type="fundref-id">10.13039/100000070</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="7"/><equation-count count="122"/><ref-count count="31"/><page-count count="0"/><word-count count="0"/></counts><custom-meta-wrap><custom-meta><meta-name>section-at-acceptance</meta-name><meta-value>Artificial Intelligence in Radiology</meta-value></custom-meta></custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro"><label>1.</label><title>Introduction</title>
<p>Along with the success of deep learning methods in medical image analysis, interpretability methods have been used to validate that models are working as expected (<xref ref-type="bibr" rid="B1">1</xref>). The interpretability of deep learning models has also been listed as a critical research priority for artificial intelligence in medical imaging (<xref ref-type="bibr" rid="B2">2</xref>). The employment of bounding box annotations during training has been shown to improve a model&#x2019;s ability to highlight abnormalities and, consequently, their interpretability (<xref ref-type="bibr" rid="B3">3</xref>). However, bounding boxes for medical images are costly to acquire since they require expert annotation, whereas image-level labels can readily be extracted from radiology reports. This fact is exemplified by the relatively small size of bounding box datasets for chest x-rays (CXRs) (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B5">5</xref>) when compared to the size of CXR datasets with image-level labels. Eye-tracking (ET) data, on the other hand, contain implicit information about the location of labels, and its collection may be easier to scale up than bounding boxes if the acquisition of gaze from radiologists is implemented in clinical practice.</p>
<p>CXRs are the most common medical imaging exam in the United States (<xref ref-type="bibr" rid="B6">6</xref>). This type of imaging has also had much attention from deep learning practitioners, with successful technical results (<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B7">7</xref>). Despite their universality, reading a CXR is considered one of the hardest interpretations performed by radiologists (<xref ref-type="bibr" rid="B8">8</xref>), with high inter-rater variability in reported abnormalities (<xref ref-type="bibr" rid="B9">9</xref>, <xref ref-type="bibr" rid="B10">10</xref>). Moreover, abnormalities in CXRs can appear in a vast diversity of locations, including the lungs, mediastinum, pleural space, vessels, airways, and ribs (<xref ref-type="bibr" rid="B11">11</xref>). They can also be described as hundreds of different findings (<xref ref-type="bibr" rid="B11">11</xref>). The diverse aspect of the report has been simplified for use in deep learning applications, where, in several cases, a simplified subset of the most common labels has been automatically extracted from reports for use in a multi-label formulation (<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B12">12</xref>, <xref ref-type="bibr" rid="B13">13</xref>). Since radiologists pay close attention to several areas when dictating a CXR report, scanning almost the whole image for signs of several abnormalities, the ET data accumulated during the full report dictation might highlight several areas with no evidence of abnormality. Therefore, the use of the temporal aspect of the report, by processing the ET data with the dictation-transcription timestamps, may achieve a more precise localization for specific abnormalities.</p>
<p>We propose to use ET data with timestamped dictations of radiology reports to identify when the presence of specific abnormalities was dictated, identify the times when radiologists would have visually attended to such abnormality, and extract the associated gaze locations. The extracted information can be used as label-specific annotation for supervising models to highlight abnormalities spatially. The localization supervision is performed using a combination of a multiple instance learning loss over the last spatial layer of a convolutional neural network (CNN) (<xref ref-type="bibr" rid="B3">3</xref>) and a multi-task learning loss (<xref ref-type="bibr" rid="B14">14</xref>), adding an output representing label-specific ET maps. To complement the annotations of an ET dataset, we employ a large dataset of CXRs with image-level labels in a weak supervision formulation. We evaluate the classification performance and the ability to localize abnormalities of a model trained with data annotated by ET. This model is compared against baselines using no annotated data and hand-annotated data. We show that using ET data during training improves the localization performance of generated interpretable heatmaps without compromising area under the receiver operating characteristic curve (AUC) classification scores and that this type of data might have value in replacing hand-labeling, depending on the costs and benefits of each type of data collection. In addition, to the best of our knowledge, this study offers the first estimation of how the value of ET data compares to the value of hand-annotated localization data when a very large dataset with image-level labels is available for weak supervision. The main contributions of this paper are:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>developing knowledge of the complexities of the use of eye-tracking data for annotation in radiology, including the proposed method of considering the lag between a radiologist&#x2019;s gaze and dictation for accumulating ET data specific for each abnormality label; and</p></list-item>
<list-item><label>&#x2022;</label><p>informing about expected relative value between ET data and manual annotations for the decision of starting future more comprehensive eye-tracking studies.</p></list-item>
</list></p>
</sec>
<sec id="s2"><label>2.</label><title>Materials and methods</title>
<sec id="s2a"><label>2.1.</label><title>Extracting localization information from eye-tracking data</title>
<p>We designed a pipeline to extract disease locations from ET data. This pipeline has two main parts: extracting label mentions in reports and generating an ET heatmap for a given detected label. A representation of the pipeline is shown in <xref ref-type="fig" rid="F1">Figure&#x00A0;1</xref>. The pipeline requires the ET dataset to contain timestamps, transcriptions of report dictations, and fixations, i.e., locations in the image where radiologists stabilized their gaze for some time.</p>
<fig id="F1" position="float"><label>Figure 1</label>
<caption><p>Diagram of the use of ET data from a radiologist to train a CNN for improved localization. (<bold>A</bold>) The ET heatmap from the dictation of the full report over its corresponding CXR. (<bold>B</bold>) Label-specific ET heatmap for the label <italic>Opacity</italic>. The keywords associated with this label, found by the adapted CheXpert labeler, are highlighted in orange. The listed sentence represents the timestamps from which fixations were extracted for generating the label-specific heatmap. (<bold>C</bold>) A representation of the employed loss function, which compares the extracted heatmap against an encoded spatial vector and a decoded version of it.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fradi-03-1088068-g001.tif"/>
</fig>
<sec id="s2a1"><label>2.1.1.</label><title>Label mention extraction</title>
<p>To extract labels from reports, we adopted a modified version of the CheXpert labeler (<xref ref-type="bibr" rid="B12">12</xref>), which uses a set of hand-crafted rules to detect label mentions and negation structures.</p>
</sec>
<sec id="s2a2"><label>2.1.2.</label><title>Fixation association with each mention</title>
<p>From observing the gaze of radiologists on a few examples, we noticed patterns that seemed to be in common for all radiologists:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>for the first moments after being shown a CXR, radiologists looked all over the image without dictating anything;</p></list-item>
<list-item><label>&#x2022;</label><p>when dictating, radiologists usually looked at regions corresponding to the content of the current sentence or the following sentence (when near the end of the dictation of the current sentence).</p></list-item>
</list>From these two observations, we decided to generate ET heatmaps for detected labels from the fixations of the sentences where the label was mentioned, the previous sentence, and the pause between sentences. We accumulated fixations within a limit of 1.5 s previous to the start of the mentioning sentence and up to the last mention in the mentioning sentence. An illustration of the method for choosing which fixations were included in the heatmaps is given in <xref ref-type="fig" rid="F2">Figure&#x00A0;2</xref>. An example with a case from the ET data we used are given in <xref ref-type="fig" rid="F1">Figures&#x00A0;1A</xref>,<xref ref-type="fig" rid="F1">B</xref>. From applying this extraction method, abnormality mentions within the same report sentence are associated with the same localization heatmap.</p>
<fig id="F2" position="float"><label>Figure 2</label>
<caption><p>Multiple instance learning technique. Method of accumulation of fixations for generating heatmaps for each sentence that mentions at least one of the abnormality labels. The chosen fixations could be between the start of the previous sentence and the last mention of the current sentence or between 1.5 s before the start of the current sentence and the last mention of the current sentence, whichever has the shortest duration.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fradi-03-1088068-g002.tif"/>
</fig>
</sec>
<sec id="s2a3"><label>2.1.3.</label><title>Heatmap generation</title>
<p>Heatmaps were generated by placing Gaussians over each fixation location with a standard deviation of one degree of visual angle, following Le Meur et al. (<xref ref-type="bibr" rid="B15">15</xref>). Fixations had the amplitude of their Gaussians weighted by their duration. The heatmap for each detected mention of a label was normalized to have a maximum value of 1. Multiple mentions of a label for the same CXR were aggregated with a maximum function.</p>
</sec>
</sec>
<sec id="s2b"><label>2.2.</label><title>Multiple instance learning</title>
<p>We used the multiple instance learning loss term from Li et al. (<xref ref-type="bibr" rid="B3">3</xref>) to train an encoder with the extracted ET heatmaps. The encoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM1"><mml:mi>E</mml:mi></mml:math></inline-formula>, as represented in <xref ref-type="fig" rid="F1">Figure&#x00A0;1C</xref>, was built to output a grid of cells, where each cell represented a multi-label classifier for label presence in the homologous region of the image. The image-level prediction <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM2"><mml:msub><mml:mi>C</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for label <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM3"><mml:mi>k</mml:mi></mml:math></inline-formula> was formulated as<disp-formula id="disp-formula1"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM1"><mml:msub><mml:mi>C</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:munder><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM4"><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the logit output for class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM5"><mml:mi>k</mml:mi></mml:math></inline-formula> and grid cell <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM6"><mml:mi>j</mml:mi></mml:math></inline-formula> for image <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM7"><mml:mi>x</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM8"><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula> is the set of all grid cells for image <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM9"><mml:mi>x</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM10"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x22C5;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the sigmoid function. <xref ref-type="disp-formula" rid="disp-formula1">Equation (1)</xref> is a soft version of the Boolean <italic>OR</italic> function, assigning a positive image-level label when at least one of the grid cells was found to contain that class.</p>
<p>During training, the loss function depended on the presence of a localization annotation. For images annotated with localization (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM11"><mml:mi>A</mml:mi></mml:math></inline-formula>), grid cells were trained to match a resized version of the annotation, as shown in <xref ref-type="fig" rid="F1">Figure&#x00A0;1C</xref>. We used the loss<disp-formula id="disp-formula2"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM2"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:munder><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mi mathvariant="italic">j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="italic">B</mml:mi><mml:mrow><mml:mi mathvariant="italic">k</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi mathvariant="italic">j</mml:mi><mml:mi mathvariant="italic">k</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:munder><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mi mathvariant="italic">j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="italic">&#x0393;</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:msub><mml:mo mathvariant="italic">&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="italic">B</mml:mi><mml:mrow><mml:mi mathvariant="italic">k</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi mathvariant="italic">j</mml:mi><mml:mi mathvariant="italic">k</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM12"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the set of grid cells labeled as containing evidence of disease <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM13"><mml:mi>k</mml:mi></mml:math></inline-formula> for image <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM14"><mml:mi>x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM15"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the output of the loss function for annotated images of class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM16"><mml:mi>k</mml:mi></mml:math></inline-formula>.</p>
<p>For images that did not contain localization annotations (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM17"><mml:mi>U</mml:mi></mml:math></inline-formula>), the loss depended on the image-level label. For positive images, at least one grid cell should be positive. We used the loss<disp-formula id="disp-formula3"><label>(3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM3"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="italic">C</mml:mi><mml:mi mathvariant="italic">k</mml:mi></mml:msub><mml:mo mathvariant="italic" stretchy="false">(</mml:mo><mml:mi mathvariant="italic">x</mml:mi><mml:mo mathvariant="italic" stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM18"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the output of the loss function for unannotated images labeled as positive for class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM19"><mml:mi>k</mml:mi></mml:math></inline-formula>. For negative images, all grid cells should be negative. We used<disp-formula id="disp-formula4"><label>(4)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM4"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mo>&#x2212;</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:munder><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mi mathvariant="italic">j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="italic">&#x0393;</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mi mathvariant="italic">j</mml:mi><mml:mi mathvariant="italic">k</mml:mi><mml:mi mathvariant="italic">x</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM20"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mo>&#x2212;</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the output of the loss function for unannotated images labeled as negative for class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM21"><mml:mi>k</mml:mi></mml:math></inline-formula>. The multiple instance learning loss term <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM22"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> was then formulated as<disp-formula id="disp-formula5"><label>(5)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM5"><mml:mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM23"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a hyperparameter controlling the relative importance of annotated images during training, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM24"><mml:mi>X</mml:mi></mml:math></inline-formula> is the set of all images, annotated and unannotated, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM25"><mml:mi>K</mml:mi></mml:math></inline-formula> is the set of all classes, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM26"><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM27"><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM28"><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mo>&#x2212;</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are the output of indicator functions that were 1 when the image <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM29"><mml:mi>x</mml:mi></mml:math></inline-formula> was annotated, positive for class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM30"><mml:mi>k</mml:mi></mml:math></inline-formula> (unannotated) and negative for class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM31"><mml:mi>k</mml:mi></mml:math></inline-formula> (unannotated), respectively.</p>
<sec id="s2b1"><label>2.2.1.</label><title>Avoiding numerical underflow and balanced range normalization</title>
<p>To avoid numerical underflow and have a more uniform range for the output of models, Li et al. (<xref ref-type="bibr" rid="B3">3</xref>) suggested normalizing the factors of the products in <xref ref-type="disp-formula" rid="disp-formula1">Equations (1)</xref> to <xref ref-type="disp-formula" rid="disp-formula5">(5)</xref>, i.e., <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM32"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM33"><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, to the range <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM34"><mml:mo stretchy="false">[</mml:mo><mml:mn>0.98</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. After running tests, we achieved better results by balancing this normalization, changing the range of each product factor to <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM35"><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mn>0.0056738</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM36"><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is the number of factors being multiplied. This range allows all products to have a similar expected range and keeps the same [0.98,1] range when <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM37"><mml:msub><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>256</mml:mn></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="s2c"><label>2.3.</label><title>Multi-task learning</title>
<p>Inspired by a work by Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>), we added another loss term to our method. With the intuition of giving more supervision to the representations calculated by encoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM38"><mml:mi>E</mml:mi></mml:math></inline-formula> and, consequently, improving its representations, we added the task of predicting a high-resolution ET map for each label, performed with the help of decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM39"><mml:mi>D</mml:mi></mml:math></inline-formula>, as shown in <xref ref-type="fig" rid="F1">Figure&#x00A0;1C</xref>. From testing the network&#x2019;s performance, we modified the method in that class outputs were calculated according to <xref ref-type="disp-formula" rid="disp-formula1">Equation (1)</xref> instead of adding fully connected layers as suggested by Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>). Another difference in our method was that our decoder output had one channel for each of the ten labels in our classification task. The output of decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM40"><mml:mi>D</mml:mi></mml:math></inline-formula> could then provide estimations for the localization of abnormalities. In other words, the output of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM41"><mml:mi>D</mml:mi></mml:math></inline-formula> is an interpretability output: an alternative to the spatial activations or other interpretability methods, such as GradCAM (<xref ref-type="bibr" rid="B16">16</xref>). Decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM42"><mml:mi>D</mml:mi></mml:math></inline-formula> had an architecture with three blocks, each composed of a sequence of a bilinear upsampling layer, a convolution layer, and a batch normalization layer. The loss <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM43"><mml:msub><mml:mi>L</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> added for this task is the pixel-level cross-entropy between the output of decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM44"><mml:mi>D</mml:mi></mml:math></inline-formula> and the label-specific ET map in the same resolution, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM45"><mml:mn>256</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>256</mml:mn></mml:math></inline-formula>, formulated as<disp-formula id="disp-formula6"><label>(6)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM6"><mml:msub><mml:mi>L</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">[</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>A</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">[</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mtext>log</mml:mtext><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mo>&#x2209;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mtext>log</mml:mtext><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">]</mml:mo></mml:mrow><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM46"><mml:msub><mml:mrow><mml:mn mathvariant="double-struck">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:msup><mml:mi>A</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the output of an indicator function that was 1 when the image <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM47"><mml:mi>x</mml:mi></mml:math></inline-formula> was annotated and positive for class <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM48"><mml:mi>k</mml:mi></mml:math></inline-formula>. This loss was used to train both decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM49"><mml:mi>D</mml:mi></mml:math></inline-formula> and encoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM50"><mml:mi>E</mml:mi></mml:math></inline-formula> and, as shown in <xref ref-type="disp-formula" rid="disp-formula6">Equation (6)</xref>, was applied only for channels corresponding to positive ground-truth labels.</p>
</sec>
<sec id="s2d"><label>2.4.</label><title>Multi-resolution architecture</title>
<p>As shown in <xref ref-type="fig" rid="F3">Figure&#x00A0;3</xref>, for our encoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM51"><mml:mi>E</mml:mi></mml:math></inline-formula> we adapted the Resnet-50 (<xref ref-type="bibr" rid="B17">17</xref>) architecture by replacing its average pooling and last linear layer with two convolutional layers separated by batch normalization (<xref ref-type="bibr" rid="B18">18</xref>) and ReLU activation (CNN Block 5 from <xref ref-type="fig" rid="F3">Figure&#x00A0;3</xref>). To improve the results for labels with small findings in the original image, we modified the network such that spatial maps with <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM52"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> resolution were used as inputs to CNN Block 5.</p>
<fig id="F3" position="float"><label>Figure 3</label>
<caption><p>Encoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM53"><mml:mi>E</mml:mi></mml:math></inline-formula> as a modified Resnet-50 architecture to include the multi-resolution branches.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fradi-03-1088068-g003.tif"/>
</fig>
</sec>
<sec id="s2e"><label>2.5.</label><title>Loss function</title>
<p>The final loss function <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM54"><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x22C5;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, to be minimized while training, is given by<disp-formula id="disp-formula7"><label>(7)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="DM7"><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mi>I</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM55"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a hyperparameter controlling the relative importance of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM56"><mml:msub><mml:mi>L</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The classification output of our model only influences the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM57"><mml:msub><mml:mi>L</mml:mi><mml:mi>I</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> term.</p>
</sec>
<sec id="s2f"><label>2.6.</label><title>Datasets</title>
<p>We used two datasets in our study. The REFLACX dataset (<xref ref-type="bibr" rid="B19">19</xref>&#x2013;<xref ref-type="bibr" rid="B21">21</xref>) provides ET data and reports from five radiologists for CXRs from the MIMIC-CXR-JPG dataset (<xref ref-type="bibr" rid="B21">21</xref>&#x2013;<xref ref-type="bibr" rid="B23">23</xref>). Additionally, the REFLACX dataset contains image-level labels and radiologist-drawn abnormality ellipses, which can be used to validate the locations highlighted by our tested models. Except for the experiments described in Sections <xref ref-type="sec" rid="s2g">2.7</xref> and <xref ref-type="sec" rid="s2h">2.8</xref>, we used examples from Phase 3 from the REFLACX dataset. The MIMIC-CXR-JPG dataset, which contains patients who visited the emergency department of the Beth Israel Deaconess Medical Center between 2011 and 2016, was also utilized for its unannotated CXRs and image-level labels. Images from the MIMIC-CXR-JPG dataset were filtered using the same criteria as the REFLACX dataset: only labeled frontal images from studies with a single frontal image were considered. The test sets for both datasets were kept the same. A few subjects from the training set of the REFLACX dataset were assigned to its validation set so that around 10&#x0025; of the REFLACX dataset was part of the validation set. The same subjects were also assigned to the validation set of the MIMIC-CXR-JPG dataset. The train, validation, and test splits had, respectively, 1,724, 277, and 506 images for the annotated set and 187,519, 4,275, and 2,701 images for the unannotated set. The use of both datasets did not require ethics approval because they are publicly available de-identified datasets. All ellipses and ET data we used originated from the REFLACX dataset and the split sizes of that dataset reflect the number of such annotations that we had available.</p>
<p>The sets of labels from the annotated and unannotated datasets were different. We decided to use the ten labels listed in <xref ref-type="table" rid="T1">Table&#x00A0;1</xref>. We provide, in <xref ref-type="table" rid="T2">Table&#x00A0;2</xref>, a list of the labels from each dataset that were considered equivalent to each of the ten labels used in this study and, in <xref ref-type="table" rid="T3">Table&#x00A0;3</xref>, the number of examples of each label present in the datasets.</p>
<table-wrap id="T1" position="float"><label>Table 1</label>
<caption><p>Per-label AUC metric on the test set for the two baselines and our method.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Label</th>
<th valign="top" align="center"><italic>Unannotated</italic></th>
<th valign="top" align="center"><italic>Ellipse</italic></th>
<th valign="top" align="center"><italic>ET model</italic> (ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">AMC</td>
<td valign="top" align="center">0.658 [0.654, 0.662]</td>
<td valign="top" align="center">0.665 [0.660, 0.670]</td>
<td valign="top" align="center">0.660 [0.655, 0.665]</td>
</tr>
<tr>
<td valign="top" align="left">Atelectasis</td>
<td valign="top" align="center">0.749 [0.744, 0.754]</td>
<td valign="top" align="center">0.751 [0.748, 0.753]</td>
<td valign="top" align="center">0.748 [0.746, 0.751]</td>
</tr>
<tr>
<td valign="top" align="left">ECS</td>
<td valign="top" align="center">0.768 [0.764, 0.772]</td>
<td valign="top" align="center">0.768 [0.765, 0.771]</td>
<td valign="top" align="center">0.765 [0.760, 0.771]</td>
</tr>
<tr>
<td valign="top" align="left">Consolidation</td>
<td valign="top" align="center">0.704 [0.697, 0.711]</td>
<td valign="top" align="center">0.709 [0.703, 0.716]</td>
<td valign="top" align="center">0.710 [0.706, 0.713]</td>
</tr>
<tr>
<td valign="top" align="left">Edema</td>
<td valign="top" align="center">0.839 [0.837, 0.841]</td>
<td valign="top" align="center">0.835 [0.833, 0.837]</td>
<td valign="top" align="center">0.838 [0.836, 0.840]</td>
</tr>
<tr>
<td valign="top" align="left">Fracture</td>
<td valign="top" align="center">0.714 [0.690, 0.737]</td>
<td valign="top" align="center">0.710 [0.696, 0.724]</td>
<td valign="top" align="center">0.714 [0.701, 0.727]</td>
</tr>
<tr>
<td valign="top" align="left">Lung Lesion</td>
<td valign="top" align="center">0.760 [0.747, 0.774]</td>
<td valign="top" align="center">0.751 [0.740, 0.763]</td>
<td valign="top" align="center">0.745 [0.738, 0.752]</td>
</tr>
<tr>
<td valign="top" align="left">Opacity</td>
<td valign="top" align="center">0.784 [0.782, 0.786]</td>
<td valign="top" align="center">0.785 [0.783, 0.788]</td>
<td valign="top" align="center">0.783 [0.779, 0.787]</td>
</tr>
<tr>
<td valign="top" align="left">Pleural abnormality</td>
<td valign="top" align="center">0.868 [0.865, 0.870]</td>
<td valign="top" align="center">0.869 [0.867, 0.871]</td>
<td valign="top" align="center">0.868 [0.865, 0.871]</td>
</tr>
<tr>
<td valign="top" align="left">Pneumothorax</td>
<td valign="top" align="center">0.825 [0.811, 0.838]</td>
<td valign="top" align="center">0.810 [0.799, 0.820]</td>
<td valign="top" align="center">0.819 [0.810, 0.827]</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T2" position="float"><label>Table 2</label>
<caption><p>List of labels that were grouped to form the labels from the analysis presented in this paper.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Labels used in our models</th>
<th valign="top" align="left">Labels from public datasets (REFLACX and MIMIC-CXR)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Abnormal Mediastinal Contour (AMC)</td>
<td valign="top" align="left">abnormal mediastinal contour (AMC), enlarged cardiomediastinum</td>
</tr>
<tr>
<td valign="top" align="left">Atelectasis</td>
<td valign="top" align="left">atelectasis</td>
</tr>
<tr>
<td valign="top" align="left">Enlarged Cardiac Silhouette (ECS)</td>
<td valign="top" align="left">enlarged cardiac silhouette (ECS), cardiomegaly</td>
</tr>
<tr>
<td valign="top" align="left">Consolidation</td>
<td valign="top" align="left">consolidation</td>
</tr>
<tr>
<td valign="top" align="left">Edema</td>
<td valign="top" align="left">pulmonary edema, edema</td>
</tr>
<tr>
<td valign="top" align="left">Fracture</td>
<td valign="top" align="left">fracture, acute fracture</td>
</tr>
<tr>
<td valign="top" align="left">Lung Lesion</td>
<td valign="top" align="left">lung nodule or mass, lung lesion</td>
</tr>
<tr>
<td valign="top" align="left">Opacity</td>
<td valign="top" align="left">pulmonary edema, edema, lung nodule or mass, atelectasis, consolidation, groundglass opacity, interstitial lung disease, pneumonia, lung opacity</td>
</tr>
<tr>
<td valign="top" align="left">Pleural Abnormality</td>
<td valign="top" align="left">pleural abnormality, pleural other, pleural effusion</td>
</tr>
<tr>
<td valign="top" align="left">Pneumothorax</td>
<td valign="top" align="left">pneumothorax</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T3" position="float"><label>Table 3</label>
<caption><p>Number of positive examples for each of the splits for both employed datasets: REFLACX (R) and MIMIC-CXR-JPG (M).</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Label</th>
<th valign="top" align="center">Train R</th>
<th valign="top" align="center">Val R</th>
<th valign="top" align="center">Test R</th>
<th valign="top" align="center">Train M</th>
<th valign="top" align="center">Val M</th>
<th valign="top" align="center">Test M</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Abnormal Mediastinal Contour</td>
<td valign="top" align="center">59</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">13,551</td>
<td valign="top" align="center">351</td>
<td valign="top" align="center">289</td>
</tr>
<tr>
<td valign="top" align="left">Atelectasis</td>
<td valign="top" align="center">503</td>
<td valign="top" align="center">69</td>
<td valign="top" align="center">189</td>
<td valign="top" align="center">46,981</td>
<td valign="top" align="center">1,196</td>
<td valign="top" align="center">754</td>
</tr>
<tr>
<td valign="top" align="left">Enlarged Cardiac Silhouette</td>
<td valign="top" align="center">340</td>
<td valign="top" align="center">55</td>
<td valign="top" align="center">173</td>
<td valign="top" align="center">41,681</td>
<td valign="top" align="center">1,138</td>
<td valign="top" align="center">821</td>
</tr>
<tr>
<td valign="top" align="left">Consolidation</td>
<td valign="top" align="center">543</td>
<td valign="top" align="center">95</td>
<td valign="top" align="center">196</td>
<td valign="top" align="center">12,311</td>
<td valign="top" align="center">352</td>
<td valign="top" align="center">253</td>
</tr>
<tr>
<td valign="top" align="left">Edema</td>
<td valign="top" align="center">222</td>
<td valign="top" align="center">39</td>
<td valign="top" align="center">124</td>
<td valign="top" align="center">33,383</td>
<td valign="top" align="center">971</td>
<td valign="top" align="center">841</td>
</tr>
<tr>
<td valign="top" align="left">Fracture</td>
<td valign="top" align="center">34</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">27</td>
<td valign="top" align="center">4,028</td>
<td valign="top" align="center">62</td>
<td valign="top" align="center">71</td>
</tr>
<tr>
<td valign="top" align="left">Lung Lesion</td>
<td valign="top" align="center">77</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">41</td>
<td valign="top" align="center">6,015</td>
<td valign="top" align="center">114</td>
<td valign="top" align="center">106</td>
</tr>
<tr>
<td valign="top" align="left">Opacity</td>
<td valign="top" align="center">865</td>
<td valign="top" align="center">131</td>
<td valign="top" align="center">337</td>
<td valign="top" align="center">97,640</td>
<td valign="top" align="center">2,560</td>
<td valign="top" align="center">1,861</td>
</tr>
<tr>
<td valign="top" align="left">Pleural Abnormality</td>
<td valign="top" align="center">486</td>
<td valign="top" align="center">69</td>
<td valign="top" align="center">205</td>
<td valign="top" align="center">50,914</td>
<td valign="top" align="center">1,426</td>
<td valign="top" align="center">1,036</td>
</tr>
<tr>
<td valign="top" align="left">Pneumothorax</td>
<td valign="top" align="center">48</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">9,653</td>
<td valign="top" align="center">258</td>
<td valign="top" align="center">96</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2g"><label>2.7.</label><title>Labeler</title>
<p>The set of labels from the REFLACX dataset is slightly different from the ones provided by the CheXpert labeler. With the help of a cardiothoracic subspecialty-trained radiologist, we modified the labeler to output a new set of labels. Modifications were also made to improve the identification of the already present labels after observing common mistakes on a separate validation set composed of 20&#x0025; of Phase 1 and Phase 2 from the REFLACX dataset. We adjusted rules for negation finding and added/adapted expressions to match and unmatch labels.<xref ref-type="fn" rid="FN0001"><sup>1</sup></xref></p>
</sec>
<sec id="s2h"><label>2.8.</label><title>Location extraction</title>
<p>We tested several methods for extracting label-specific localization of abnormalities from the eye-tracking data. All methods involved the accumulation of fixations into heatmaps, with different starting and ending accumulation times, after extracting the label&#x2019;s mention time from the dictation. For a first stage of validation, the starting times we considered were:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>MAX(Start of mention sentence - TIME, Start of the previous sentence),</p></list-item>
<list-item><label>&#x2022;</label><p>MAX(First mention in the sentence - TIME, Start of the previous sentence),</p></list-item>
<list-item><label>&#x2022;</label><p>MAX(End of mention sentence - TIME, Start of the previous sentence),</p></list-item>
<list-item><label>&#x2022;</label><p>start of first report sentence,</p></list-item>
<list-item><label>&#x2022;</label><p>start of the previous sentence,</p></list-item>
<list-item><label>&#x2022;</label><p>end of the previous sentence,</p></list-item>
<list-item><label>&#x2022;</label><p>start of mention sentence,</p></list-item>
<list-item><label>&#x2022;</label><p>start of the recording of data for that CXR,</p></list-item>
</list>where TIME is a time delay assuming the values of 2.5 s, 5.0 s, and 7.5 s. The end times we considered were:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>start of mention sentence,</p></list-item>
<list-item><label>&#x2022;</label><p>end of mention sentence,</p></list-item>
<list-item><label>&#x2022;</label><p>end of the first mention,</p></list-item>
<list-item><label>&#x2022;</label><p>end of the last mention.</p></list-item>
</list>We tested all combinations between starting times and end times with a duration of 0 s or more. We compared the extracted heatmaps with the validation hand-annotated ellipses using the IoU metric with a validated threshold. After this validation, we finetuned, as a second stage of validation, the time delay by testing more times (0.5 s, 0.75 s, 1 s, 1.25 s, 1.5 s, 1.75 s, 2 s, 2.5 s, 3 s, 3.5 s, 4 s, 4.5 s, 5 s).</p>
</sec>
<sec id="s2i"><label>2.9.</label><title>Validation and evaluation</title>
<p>For our experiments,<xref ref-type="fn" rid="FN0002"><sup>2</sup></xref> we used PyTorch 1.10.2 (<xref ref-type="bibr" rid="B24">24</xref>). Hyperparameters commonly used for CXR classifiers were employed during training and were not tuned for any tested method. Models were trained for 60 epochs with the AMSGrad Adam optimizer (<xref ref-type="bibr" rid="B25">25</xref>) using a learning rate equal to 0.001 and weight decay of 0.00001. A batch size of 20 images was chosen for the use of GPUs with 16GB of memory or more. Images were resized such that their longest dimension had 512 pixels, whereas the other dimension was padded with black pixels to reach a length of 512 pixels. For training, images were augmented with rotation up to 45 degrees, translation up to 15&#x0025;, and scaling up to 15&#x0025;. The grid supervised by loss <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM58"><mml:msub><mml:mi>L</mml:mi><mml:mi>I</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> had 1,024 cells (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM59"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>). We used the max-pooling operation to convert the ET heatmap annotations to the same dimension. We thresholded the ET heatmaps at 0.15. This number was chosen after visual analysis of their histograms of intensities. We used <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM60"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>A</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM61"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>300</mml:mn></mml:math></inline-formula> after validation of AUC and IoU values for our proposed method considering the following values: 0.3, 1, 3, 10, 30, and 100 for <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM62"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:math></inline-formula> and 0.1, 0.3, 1, 3, 10, 30, 100, 300, 1,000, 3,000 for <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM63"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>T</mml:mi></mml:msub></mml:math></inline-formula>. We trained models with five different seeds and report their average results and 95&#x0025; confidence intervals. Experiments were run in internal servers containing Nvidia GPUs (TITAN RTX 24GB, RTX A6000 48GB, Tesla V100-SXM2 16 GB). Each training run took approximately two to three days in one GPU.</p>
<p>As baselines, we evaluated a model trained without the annotated data (<italic>Unannotated</italic>) and a model trained with data annotated by the drawn ground truth (GT) truth ellipses (<italic>Ellipse</italic>). The ellipses were represented by binary heatmaps and were processed in the same way as the ET heatmaps. The loss function, CNN architecture, and training hyperparameters were the same for all methods. We did not include a cross-entropy classification loss baseline because it achieved lower scores than the presented methods. The best epoch for each method was chosen using the average AUC on the validation set. The best model heatmap threshold for each method and label was calculated using the average validation intersection over union (IoU) over the five seeds, considering the full range of thresholds.</p>
<p>We evaluated our model (<italic>ET model</italic>) and the two baselines by calculating the test AUC for image-level labels of the MIMIC-CXR-JPG dataset and the test IoU for localization of abnormalities, compared against the drawn GT ellipses. IoU was calculated individually per positive label for all images with a positive label. We calculated three heatmaps for each label and input image: the output of decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM64"><mml:mi>D</mml:mi></mml:math></inline-formula>, the spatial activations, and the output of the GradCAM method (<xref ref-type="bibr" rid="B16">16</xref>). We tested which heatmap had the best IoU validation results for each of the three reported training methods. We report results for the <italic>Ellipse</italic> and <italic>ET model</italic> using the output of decoder <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM65"><mml:mi>D</mml:mi></mml:math></inline-formula> and for the <italic>Unannotated</italic> model using the spatial activations. The heatmaps for each method were upscaled to the resolution of the GT ellipses using nearest-neighbor interpolation.</p>
</sec>
</sec>
<sec id="s3" sec-type="results"><label>3.</label><title>Results</title>
<sec id="s3a"><label>3.1.</label><title>Labeler</title>
<p>Labeler quality estimations, after modifications, are shown in <xref ref-type="table" rid="T4">Table&#x00A0;4</xref> and were calculated with the rest of the data from Phase 2, representing 80&#x0025; of the cases. Results were variable depending on the label, and misdetections should be expected when using this labeler.</p>
<table-wrap id="T4" position="float"><label>Table 4</label>
<caption><p>Results of the label detection with a modified version of the CheXpert labeler (<xref ref-type="bibr" rid="B12">12</xref>).</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="left"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Label</th>
<th valign="top" align="center">Recall</th>
<th valign="top" align="center">Precision</th>
<th valign="top" align="center">Label</th>
<th valign="top" align="center">Recall</th>
<th valign="top" align="center">Precision</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">AMC</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="left">Interstitial Lung Disease</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.27</td>
</tr>
<tr>
<td valign="top" align="left">Acute Fracture</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="left">Lung Nodule or Mass</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.50</td>
</tr>
<tr>
<td valign="top" align="left">Atelectasis</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="left">Pleural Abnormality</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.98</td>
</tr>
<tr>
<td valign="top" align="left">Consolidation</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="left">Pneumothorax</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">1.00</td>
</tr>
<tr>
<td valign="top" align="left">ECS</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="left">Pulmonary Edema</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">Groundglass Opacity</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="left"/>
<td valign="top" align="center"/>
<td valign="top" align="center"/>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn1"><p>AMC stands for <italic>Abnormal Mediastinal Contour</italic> and ECS for <italic>Enlarged Cardiac Silhouette</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3b"><label>3.2.</label><title>Location extraction</title>
<p>For the first validation, the highest-scoring accumulated heatmaps used a starting time of MAX(Start of mention sentence <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM66"><mml:mo>&#x2212;</mml:mo><mml:mn>2.5</mml:mn></mml:math></inline-formula> s, Start of the previous sentence) and an ending at the end of the last mention present in the sentence. For the second stage of the location extraction validation, when we tested delay times in a higher resolution, the time with the best IoU was 1.5 s with an IoU of 0.233, justifying our approach as presented in Section <xref ref-type="sec" rid="s2a">2.1</xref>.</p>
</sec>
<sec id="s3c"><label>3.3.</label><title>Comparison with baselines</title>
<p>Results, averaged over all labels, are presented in <xref ref-type="table" rid="T5">Table&#x00A0;5</xref>. The average AUC for the <italic>ET model</italic> was not significantly different from the baselines. Regarding localization, the IoU values showed that training with the ET data was significantly better than training without annotated data and worse than training with the hand-labeled localization ellipses. Results for AUC and IoU of individual labels are presented in <xref ref-type="table" rid="T1">Tables&#x00A0;1</xref>, <xref ref-type="table" rid="T6">6</xref>. AUC was stable among all methods for almost all labels. Successful and unsuccessful heatmaps generated by our trained models are shown in <xref ref-type="fig" rid="F4">Figure&#x00A0;4</xref>.</p>
<fig id="F4" position="float"><label>Figure 4</label>
<caption><p>Localization output for the models for random test CXRs.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fradi-03-1088068-g004.tif"/>
</fig>
<table-wrap id="T5" position="float"><label>Table 5</label>
<caption><p>Results on the test set comparing the <italic>ET model</italic> with the two baselines.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Method</th>
<th valign="top" align="center"><italic>Unannotated</italic></th>
<th valign="top" align="center"><italic>Ellipse</italic></th>
<th valign="top" align="center"><italic>ET model</italic> (ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">AUC</td>
<td valign="top" align="center">0.767 [0.763, 0.771]</td>
<td valign="top" align="center">0.765 [0.763, 0.768]</td>
<td valign="top" align="center">0.765 [0.763, 0.767]</td>
</tr>
<tr>
<td valign="top" align="left">IoU</td>
<td valign="top" align="center">0.201 [0.198, 0.204]</td>
<td valign="top" align="center">0.335 [0.330, 0.339]</td>
<td valign="top" align="center">0.256 [0.253, 0.260]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn2"><p>AUC and IoU were averaged over the scores of all labels.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="T6" position="float"><label>Table 6</label>
<caption><p>Per-label IoU metric on the test set for the two baselines and our method.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Label</th>
<th valign="top" align="center"><italic>Unannotated</italic></th>
<th valign="top" align="center"><italic>Ellipse</italic></th>
<th valign="top" align="center"><italic>ET model</italic> (ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">AMC</td>
<td valign="top" align="center">0.111 [0.100, 0.122]</td>
<td valign="top" align="center">0.297 [0.264, 0.331]</td>
<td valign="top" align="center">0.214 [0.187, 0.242]</td>
</tr>
<tr>
<td valign="top" align="left">Atelectasis</td>
<td valign="top" align="center">0.245 [0.241, 0.250]</td>
<td valign="top" align="center">0.385 [0.378, 0.392]</td>
<td valign="top" align="center">0.335 [0.320, 0.350]</td>
</tr>
<tr>
<td valign="top" align="left">ECS</td>
<td valign="top" align="center">0.386 [0.346, 0.427]</td>
<td valign="top" align="center">0.747 [0.743, 0.751]</td>
<td valign="top" align="center">0.379 [0.357, 0.401]</td>
</tr>
<tr>
<td valign="top" align="left">Consolidation</td>
<td valign="top" align="center">0.242 [0.235, 0.249]</td>
<td valign="top" align="center">0.380 [0.374, 0.385]</td>
<td valign="top" align="center">0.324 [0.314, 0.334]</td>
</tr>
<tr>
<td valign="top" align="left">Edema</td>
<td valign="top" align="center">0.314 [0.299, 0.330]</td>
<td valign="top" align="center">0.466 [0.460, 0.472]</td>
<td valign="top" align="center">0.401 [0.396, 0.406]</td>
</tr>
<tr>
<td valign="top" align="left">Fracture</td>
<td valign="top" align="center">0.012 [0.006, 0.018]</td>
<td valign="top" align="center">0.004 [0.004, 0.004]</td>
<td valign="top" align="center">0.007 [0.006, 0.008]</td>
</tr>
<tr>
<td valign="top" align="left">Lung Lesion</td>
<td valign="top" align="center">0.113 [0.104, 0.123]</td>
<td valign="top" align="center">0.203 [0.193, 0.214]</td>
<td valign="top" align="center">0.213 [0.194, 0.231]</td>
</tr>
<tr>
<td valign="top" align="left">Opacity</td>
<td valign="top" align="center">0.260 [0.255, 0.264]</td>
<td valign="top" align="center">0.387 [0.382, 0.391]</td>
<td valign="top" align="center">0.341 [0.336, 0.347]</td>
</tr>
<tr>
<td valign="top" align="left">Pleural abnormality</td>
<td valign="top" align="center">0.210 [0.202, 0.218]</td>
<td valign="top" align="center">0.297 [0.283, 0.311]</td>
<td valign="top" align="center">0.246 [0.240, 0.251]</td>
</tr>
<tr>
<td valign="top" align="left">Pneumothorax</td>
<td valign="top" align="center">0.114 [0.103, 0.125]</td>
<td valign="top" align="center">0.180 [0.173, 0.187]</td>
<td valign="top" align="center">0.103 [0.089, 0.118]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>When training the <italic>Ellipse</italic> model with only 15&#x0025; of the annotated dataset, we achieved an IoU of .257 [.248,.266]. Therefore, given the IoU provided for the <italic>ET model</italic> in <xref ref-type="table" rid="T5">Table&#x00A0;5</xref>, the best estimation for the value of ET data is 15&#x0025;, i.e., around one-seventh, of the value of the hand-annotated data.</p>
</sec>
<sec id="s3d"><label>3.4.</label><title>Ablation study</title>
<p>We present in <xref ref-type="table" rid="T7">Table&#x00A0;7</xref> the results of an ablation study with each modification added to the original method from Li et al. (<xref ref-type="bibr" rid="B3">3</xref>), which is presented in Section <xref ref-type="sec" rid="s2b">2.2</xref>, including the label specific heatmaps from Section <xref ref-type="sec" rid="s2a">2.1</xref>, the balanced range normalization from Sections <xref ref-type="sec" rid="s2b1">2.2.1</xref>, the multi-resolution architecture from Section <xref ref-type="sec" rid="s2d">2.4</xref>, and the multi-task learning from Section <xref ref-type="sec" rid="s2c">2.3</xref>. <xref ref-type="table" rid="T7">Table&#x00A0;7</xref> presents the elements added to the model in chronological order of addition to our project in its first five rows. We show an advantage from each modification for all methods. We also added a row with the removal of only the label-specific heatmaps to show that they had a big impact on the final IoU. Its removal caused a decrease of around 0.062 (24.2&#x0025;) on the IoU, reaching an IoU similar to the <italic>Unannotated</italic> model.</p>
<table-wrap id="T7" position="float"><label>Table 7</label>
<caption><p>IoU metric on the test set indicating the advantage of using each of the modifications to the multiple instance learning (MIL) method proposed by Li et al. (<xref ref-type="bibr" rid="B3">3</xref>).</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">MIL</th>
<th valign="top" align="center">LSH</th>
<th valign="top" align="center">BRN</th>
<th valign="top" align="center">MRA</th>
<th valign="top" align="center">MTL</th>
<th valign="top" align="center">IoU for <italic>ET model</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">&#x2713;</td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM67"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM68"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM69"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM70"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center">.165 [.162, .168]</td>
</tr>
<tr>
<td valign="top" align="left">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM71"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM72"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM73"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center">.180 [.176, .185]</td>
</tr>
<tr>
<td valign="top" align="left">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM74"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM75"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center">.200 [.197, .203]</td>
</tr>
<tr>
<td valign="top" align="left">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM76"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center">.218 [.216, .221]</td>
</tr>
<tr>
<td valign="top" align="left">&#x2713;</td>
<td valign="top" align="center"><inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="IM77"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">.194 [.192, .197]</td>
</tr>
<tr>
<td valign="top" align="left">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="top" align="center">.256 [.252, .260]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn3"><p>We tested different combinations of the following methods: label-specific heatmaps (LSH) (our contribution), balanced range normalization (BRN) (our contribution), multi-resolution architecture (MRA) (our contribution), and multi-task learning (MTL) (<xref ref-type="bibr" rid="B14">14</xref>).</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Furthermore, we evaluated the impact of the quality of the label mention extraction, performed using the modified version of the CheXpert labeler. To check how IoU scores change with respect to the quality of the labeler, we intentionally lowered the quality of Pleural Abnormality label extraction, for which both recall and precision were high in <xref ref-type="table" rid="T4">Table&#x00A0;4</xref>. We randomly removed the mention of approximately 45&#x0025; of the cases with mentions of Pleural Abnormality and randomly assigned a mention of another label as a Pleural Abnormality mention in approximately 5&#x0025; of the cases where there was no mention of Pleural Abnormality. With these changes, we estimate that the recall for mentions of Pleural Abnormality is 0.51, and the precision is 0.86. We trained a model with these modified mentions for Pleural Abnormality and the original mentions for the other nine abnormality labels. We achieved an IoU of 0.227 [0.205,0.249], likely to lie between the <italic>Unannotated</italic> (0.210 [0.202, 0.218]) and <italic>ET model</italic> (0.246 [0.240, 0.251]), shown in <xref ref-type="table" rid="T6">Table&#x00A0;6</xref>. There was no impact on the Pleural Abnormality AUC (0.869 [0.868, 0.870]).</p>
</sec>
</sec>
<sec id="s4" sec-type="discussion"><label>4.</label><title>Discussion</title>
<p>Other studies have applied ET data for localizing abnormalities and improving the localization of models. Stember et al. (<xref ref-type="bibr" rid="B26">26</xref>) showed that radiologists looked at the location of a label when indicating the presence of tumors in MRIs. However, we use a less restrictive and more challenging scenario, with freeform reports and multiple types of abnormalities reported. Saab et al. (<xref ref-type="bibr" rid="B27">27</xref>) showed that, when gaze data are aggregated in hand-crafted features and used in a multi-task setup, the GradCAM heatmap of a model overlaps more often with the location of pneumothoraces. Li et al. (<xref ref-type="bibr" rid="B28">28</xref>) developed an attention-guided network for glaucoma diagnosis split into three sequential stages: a prediction of an attention map supervised by an ET heatmap, an intermediary classification network trained to refine the attention map through guided backpropagation, and a final classification network. Wang et al. (<xref ref-type="bibr" rid="B29">29</xref>) performed osteoarthritis grading by enforcing the class activation map (CAM) heatmap to be similar to the ET map, allowing for uncertainty in the ET map. We tested adding this method to our loss, but no improvement was seen. All of these methods used ET datasets where radiologists focused on a single task/abnormality, making their ET data intrinsically label-specific and their setup distant from clinical practice. The method we propose uses dictations to identify moments when the radiologist looks at evidence of multiple abnormalities, allowing for the use of ET data collected during clinical report dictation. We used a multi-task formulation as part of our loss following Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>). However, even though they used a dataset where radiologists looked for several types of abnormalities, they showed only the impact of a single ET heatmap for all labels. Our study focuses on more complex uses of the ET data, with the generation of label-specific heatmaps.</p>
<p>Contrary to other works, such as Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>) and Li et al. (<xref ref-type="bibr" rid="B3">3</xref>), we achieved no improvements in classification performance in our setup when applying a variety of localization losses to our model. One of the reasons we might achieve different levels of improvement from Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>) is that we use a much larger dataset for training the model, weakly including most of the MIMIC-CXR-JPG dataset. The use of abundant unannotated data might reduce the impact of the annotated data on the final model. However, improvements in the ability to localize the abnormalities were still shown in our experiments.</p>
<p>We tried to apply the same method to train localization with the 1,064 images from the dataset shared by Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>). However, the performance was similar to the <italic>Unannotated</italic> model. There are several reasons why the method might not be generalizable to the other dataset, including:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>It is possible that the method choices were only well-adapted to some of the five radiologists in the REFLACX dataset and not generic for all radiologists, including the single one in the other dataset.</p></list-item>
<list-item><label>&#x2022;</label><p>Some of the method choices might have only been well-adapted to the characteristics of the ET data of the REFLACX dataset. The other dataset, for example, has data collected at 60 Hz instead of 1,000 Hz and has two cases with the first fixations happening after the first mention of an abnormality.</p></list-item>
<list-item><label>&#x2022;</label><p>The size of the other dataset might be too small, or the distribution of cases might be different from the validation set of the REFLACX dataset, which we used to evaluate performance. For example, the dataset shared by Karargyris et al. (<xref ref-type="bibr" rid="B14">14</xref>) includes only posterior anterior (PA) CXRs and had only one case where the labeler identified a mention of Pneumothorax.</p></list-item>
</list>The method we proposed for producing label-specific localization annotations from ET data and for training models to produce heatmaps that match the annotation improved the interpretability of deep learning models for CXRs, as measured by comparing produced heatmaps against hand-annotated localization of abnormalities. From the IoU achieved by training the model with 15&#x0025; of the available bounding boxes, we showed that, in our setup, around seven CXRs with ET data provide the same level of efficacy in localization supervision as one CXR with expert-annotated ellipses, showing the possible value of using this type of data for scaling up annotations. This relative value might be used in calculations involving the costs and benefits of each data collection method when deciding on how to get annotations. As shown by our ablation study, the use of label-specific annotations was essential to the added value of using the ET data. We also showed in the ablation study that the performance of the label extraction algorithm has a corresponding impact on the improvements over IoU. These results show that there is an opportunity for improvement of the IoU results for labels that had a low recall and/or precision on <xref ref-type="table" rid="T4">Table&#x00A0;4</xref>. Recent advances in large language models (<xref ref-type="bibr" rid="B30">30</xref>) suggest that these models could significantly improve the accuracy of label extraction in the future.</p>
<p>The ET data were relatively noisy, and the achieved IoU for our label-specific training heatmaps was relatively low, limiting the achieved IoU for our proposed method of using ET data for training. In future work, we will investigate other methods of extracting the localization information to reduce the noise in the data, including methods of unsupervised alignment.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability"><title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://www.physionet.org/content/reflacx-xray-localization/1.0.0/">https://www.physionet.org/content/reflacx-xray-localization/1.0.0/</ext-link> <ext-link ext-link-type="uri" xlink:href="https://www.physionet.org/content/mimic-cxr-jpg/2.0.0/">https://www.physionet.org/content/mimic-cxr-jpg/2.0.0/</ext-link>.</p>
</sec>
<sec id="s6" sec-type="ethics-statement"><title>Ethics statement</title>
<p>Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. Written informed consent for participation was not required for this study in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec id="s7" sec-type="author-contributions"><title>Author contributions</title>
<p>RBL wrote the manuscript, coded and conducted the experiments, and ran the analyses. JDS participated in the study design and provided feedback for the manuscript. TT is the project PI, coordinating the study design, leading discussions about the project, and editing the manuscript. All authors reviewed the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8" sec-type="funding-information"><title>Funding</title>
<p>This research was funded by the National Institute of Biomedical Imaging and Bioengineering of the National Institutes of Health under Award Number R21EB028367.</p>
</sec>
<ack><title>Acknowledgments</title>
<p>Yichu Zhou participated in the modification and evaluation of the CheXpert labeler. This article has appeared in a preprint (<xref ref-type="bibr" rid="B31">31</xref>).</p>
</ack>
<sec id="s9" sec-type="COI-statement"><title>Conflict of interest</title>
<p>The author TT declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer"><title>Publisher&#x0027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="FN0001"><p><sup>1</sup>The final set of rules can be found in our code repository at <ext-link ext-link-type="uri" xlink:href="https://github.com/ricbl/eye-tracking-localization">https://github.com/ricbl/eye-tracking-localization</ext-link>.</p></fn>
<fn id="FN0002"><p><sup>2</sup>The code for our experiments can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/ricbl/eye-tracking-localization">https://github.com/ricbl/eye-tracking-localization</ext-link>.</p></fn>
</fn-group>
<ref-list><title>References</title>
<ref id="B1"><label>1.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Rajpurkar</surname><given-names>P</given-names></name><name><surname>Irvin</surname><given-names>J</given-names></name><name><surname>Zhu</surname><given-names>K</given-names></name><name><surname>Yang</surname><given-names>B</given-names></name><name><surname>Mehta</surname><given-names>H</given-names></name><name><surname>Duan</surname><given-names>T</given-names></name></person-group>, et al. <comment>Chexnet: Radiologist-level pneumonia detection on chest x-rays withdeep learning (2017)</comment>.</citation></ref>
<ref id="B2"><label>2.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Langlotz</surname><given-names>CP</given-names></name><name><surname>Allen</surname><given-names>B</given-names></name><name><surname>Erickson</surname><given-names>BJ</given-names></name><name><surname>Kalpathy-Cramer</surname><given-names>J</given-names></name><name><surname>Bigelow</surname><given-names>K</given-names></name><name><surname>Cook</surname><given-names>TS</given-names></name><etal/></person-group> <article-title>A road map for foundational research on artificial intelligence in medical imaging: from the 2018 NIH/RSNA/ACR/the academy workshop</article-title>. <source>Radiology</source>. (<year>2019</year>) <volume>291</volume>:<fpage>781</fpage>&#x2013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1148/radiol.2019190613</pub-id> <pub-id pub-id-type="pmid">30990384</pub-id></citation></ref>
<ref id="B3"><label>3.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Li</surname><given-names>Z</given-names></name><name><surname>Wang</surname><given-names>C</given-names></name><name><surname>Han</surname><given-names>M</given-names></name><name><surname>Xue</surname><given-names>Y</given-names></name><name><surname>Wei</surname><given-names>W</given-names></name><name><surname>Li</surname><given-names>L</given-names></name></person-group>, et al. <comment>Thoracic disease identification, localization with limited supervision. In: <italic>2018 IEEE Conference on Computer Vision, Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18&#x2013;22, 2018</italic>. IEEEComputer Society (2018). p. 8290&#x2013;8299</comment>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00865</pub-id></citation></ref>
<ref id="B4"><label>4.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Nguyen</surname><given-names>HQ</given-names></name><name><surname>Lam</surname><given-names>K</given-names></name><name><surname>Le</surname><given-names>LT</given-names></name><name><surname>Pham</surname><given-names>HH</given-names></name><name><surname>Tran</surname><given-names>DQ</given-names></name><name><surname>Nguyen</surname><given-names>DB</given-names></name></person-group>, et al. <comment>Vindr-cxr: Anopen dataset of chest x-rays with radiologist&#x2019;s annotations (2021)</comment>.</citation></ref>
<ref id="B5"><label>5.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wang</surname><given-names>X</given-names></name><name><surname>Peng</surname><given-names>Y</given-names></name><name><surname>Lu</surname><given-names>L</given-names></name><name><surname>Lu</surname><given-names>Z</given-names></name><name><surname>Bagheri</surname><given-names>M</given-names></name><name><surname>Summers</surname><given-names>RM</given-names></name></person-group>. <comment>Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In: <italic>2017 IEEE Conference on Computer Vision and Pattern Recognition,CVPR 2017, Honolulu, HI, USA, July 21&#x2013;26, 2017</italic>. IEEE Computer Society (2017). p. 3462&#x2013;3471</comment>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.369</pub-id></citation></ref>
<ref id="B6"><label>6.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mettler</surname><given-names>FA</given-names></name><name><surname>Bhargavan</surname><given-names>M</given-names></name><name><surname>Faulkner</surname><given-names>K</given-names></name><name><surname>Gilley</surname><given-names>DB</given-names></name><name><surname>Gray</surname><given-names>JE</given-names></name><name><surname>Ibbott</surname><given-names>GS</given-names></name></person-group>, et al. <article-title>Radiologic and nuclear medicine studies in the united states and worldwide: frequency, radiation dose, and comparison with other radiation sources&#x2014;1950&#x2013;2007</article-title>. <source>Radiology</source>. (<year>2009</year>) <volume>253</volume>:<fpage>520</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1148/radiol.2532082010</pub-id>. <comment>PMID: 19789227</comment><pub-id pub-id-type="pmid">19789227</pub-id></citation></ref>
<ref id="B7"><label>7.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lakhani</surname><given-names>P</given-names></name><name><surname>Sundaram</surname><given-names>B</given-names></name></person-group>. <article-title>Deep learning at chest radiography: automated classification of pulmonary tuberculosis by using convolutional neural networks</article-title>. <source>Radiology</source>. (<year>2017</year>) <volume>284</volume>:<fpage>574</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1148/radiol.2017162326</pub-id><pub-id pub-id-type="pmid">28436741</pub-id></citation></ref>
<ref id="B8"><label>8.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nyboe</surname><given-names>J</given-names></name></person-group>. <article-title>Evaluation of efficiency in interpretation of chest x-ray films</article-title>. <source>Bull World Health Organ</source>. (<year>1966</year>) <volume>35</volume>:<fpage>535</fpage>&#x2013;<lpage>45</lpage>. <pub-id pub-id-type="pmid">5297553</pub-id>.<pub-id pub-id-type="pmid">5297553</pub-id></citation></ref>
<ref id="B9"><label>9.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balabanova</surname><given-names>Y</given-names></name><name><surname>Coker</surname><given-names>R</given-names></name><name><surname>Fedorin</surname><given-names>I</given-names></name><name><surname>Zakharova</surname><given-names>S</given-names></name><name><surname>Plavinskij</surname><given-names>S</given-names></name><name><surname>Krukov</surname><given-names>N</given-names></name></person-group>, et al. <article-title>Variability in interpretation of chest radiographs among Russian clinicians, implications for screening programmes: observational study</article-title>. <source>BMJ</source>. (<year>2005</year>) <volume>331</volume>:<fpage>379</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1136/bmj.331.7513.379</pub-id><pub-id pub-id-type="pmid">16096305</pub-id></citation></ref>
<ref id="B10"><label>10.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quekel</surname><given-names>LG</given-names></name><name><surname>Kessels</surname><given-names>AG</given-names></name><name><surname>Goei</surname><given-names>R</given-names></name><name><surname>van Engelshoven</surname><given-names>JM</given-names></name></person-group>. <article-title>Detection of lung cancer on the chest radiograph: a study on observer performance</article-title>. <source>Eur J Radiol</source>. (<year>2001</year>) <volume>39</volume>:<fpage>111</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1016/S0720-048X(01)00301-1</pub-id><pub-id pub-id-type="pmid">11522420</pub-id></citation></ref>
<ref id="B11"><label>11.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bustos</surname><given-names>A</given-names></name><name><surname>Pertusa</surname><given-names>A</given-names></name><name><surname>Salinas</surname><given-names>JM</given-names></name></person-group>. <article-title>Padchest: A large chest x-ray image dataset with multi-label annotated reports</article-title>. <source>Med Image Anal</source>. (<year>2020</year>) <volume>66</volume>:<fpage>101797</fpage>. <pub-id pub-id-type="doi">10.1016/j.media.2020.101797</pub-id><pub-id pub-id-type="pmid">32877839</pub-id></citation></ref>
<ref id="B12"><label>12.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Irvin</surname><given-names>J</given-names></name><name><surname>Rajpurkar</surname><given-names>P</given-names></name><name><surname>Ko</surname><given-names>M</given-names></name><name><surname>Yu</surname><given-names>Y</given-names></name><name><surname>Ciurea-Ilcus</surname><given-names>S</given-names></name><name><surname>Chute</surname><given-names>C</given-names></name></person-group>, et al. <comment>CheXpert: A large chest radiograph dataset with uncertainty labels, expert comparison. In: <italic>The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January27&#x2013;February 1, 2019</italic>. AAAI Press (2019). p. 590&#x2013;597</comment>. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.3301590</pub-id></citation></ref>
<ref id="B13"><label>13.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname><given-names>AEW</given-names></name><name><surname>Pollard</surname><given-names>TJ</given-names></name><name><surname>Berkowitz</surname><given-names>SJ</given-names></name><name><surname>Greenbaum</surname><given-names>NR</given-names></name><name><surname>Lungren</surname><given-names>MP</given-names></name><name><surname>Deng</surname><given-names>CY</given-names></name><etal/></person-group> <article-title>MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports</article-title>. <source>Sci Data</source>. (<year>2019</year>) <volume>6</volume>:<fpage>317</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-019-0322-0</pub-id><pub-id pub-id-type="pmid">31831740</pub-id></citation></ref>
<ref id="B14"><label>14.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Karargyris</surname><given-names>A</given-names></name><name><surname>Kashyap</surname><given-names>S</given-names></name><name><surname>Lourentzou</surname><given-names>I</given-names></name><name><surname>Wu</surname><given-names>JT</given-names></name><name><surname>Sharma</surname><given-names>A</given-names></name><name><surname>Tong</surname><given-names>M</given-names></name></person-group>, et al. <article-title>Creation and validation of a chest x-ray dataset with eye-tracking and report dictation for AI development</article-title>. <source>Sci Data</source>. (<year>2021</year>) <volume>8</volume>:<fpage>92</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-021-00863-5</pub-id> <pub-id pub-id-type="pmid">33767191</pub-id></citation></ref>
<ref id="B15"><label>15.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le Meur</surname><given-names>O</given-names></name><name><surname>Baccino</surname><given-names>T</given-names></name></person-group>. <article-title>Methods for comparing scanpaths, saliency maps: strengths, weaknesses</article-title>. <source>Behav Res Methods</source>. (<year>2013</year>) <volume>45</volume>(<issue>1</issue>):<fpage>251</fpage>&#x2013;<lpage>66</lpage>. <comment>10.3758/s13428-012-0226-9</comment><pub-id pub-id-type="pmid">22773434</pub-id></citation></ref>
<ref id="B16"><label>16.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Selvaraju</surname><given-names>RR</given-names></name><name><surname>Cogswell</surname><given-names>M</given-names></name><name><surname>Das</surname><given-names>A</given-names></name><name><surname>Vedantam</surname><given-names>R</given-names></name><name><surname>Parikh</surname><given-names>D</given-names></name><name><surname>Batra</surname><given-names>D</given-names></name></person-group>. <comment>Grad-CAM: Visual explanations from deep networks via gradient-based localization. In: <italic>IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22&#x2013;29, 2017</italic>. IEEE Computer Society (2017). p. 618&#x2013;626</comment>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.74</pub-id></citation></ref>
<ref id="B17"><label>17.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>He</surname><given-names>K</given-names></name><name><surname>Zhang</surname><given-names>X</given-names></name><name><surname>Ren</surname><given-names>S</given-names></name><name><surname>Sun</surname><given-names>J</given-names></name></person-group>. <comment>Deep residual learning for image recognition. In: <italic>2016 IEEE Conference on Computer Vision, Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27&#x2013;30, 2016</italic>. IEEE Computer Society (2016). p. 770&#x2013;778</comment>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id></citation></ref>
<ref id="B18"><label>18.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Ioffe</surname><given-names>S</given-names></name><name><surname>Szegedy</surname><given-names>C</given-names></name></person-group>. <comment>Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Bach FR, Blei DM, editors. <italic>Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6&#x2013;11 July 2015, JMLR Workshop, Conference Proceedings</italic>. (JMLR.org) (2015). vol. 37, p. 448&#x2013;456</comment></citation></ref>
<ref id="B19"><label>19.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Bigolin Lanfredi</surname><given-names>R</given-names></name><name><surname>Zhang</surname><given-names>M</given-names></name><name><surname>Auffermann</surname><given-names>W</given-names></name><name><surname>Chan</surname><given-names>J</given-names></name><name><surname>Duong</surname><given-names>P</given-names></name><name><surname>Srikumar</surname><given-names>V</given-names></name></person-group>, et al. <comment>REFLACX: Reports and eye-tracking data for localization of abnormalities in chest x-rays (2021)</comment>. <comment>doi:10.13026/E0DJ-8498</comment></citation></ref>
<ref id="B20"><label>20.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bigolin Lanfredi</surname><given-names>R</given-names></name><name><surname>Zhang</surname><given-names>M</given-names></name><name><surname>Auffermann</surname><given-names>WF</given-names></name><name><surname>Chan</surname><given-names>J</given-names></name><name><surname>Duong</surname><given-names>PT</given-names></name><name><surname>Srikumar</surname><given-names>V</given-names></name></person-group>, et al. <article-title>Reflacx, a dataset of reports and eye-tracking data for localization of abnormalities in chest x-rays</article-title>. <source>Sci Data</source>. (<year>2022</year>) <volume>9</volume>:<fpage>350</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-022-01441-z</pub-id><pub-id pub-id-type="pmid">35717401</pub-id></citation></ref>
<ref id="B21"><label>21.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberger</surname><given-names>AL</given-names></name><name><surname>Amaral</surname><given-names>LAN</given-names></name><name><surname>Glass</surname><given-names>L</given-names></name><name><surname>Hausdorff</surname><given-names>JM</given-names></name><name><surname>Ivanov</surname><given-names>PC</given-names></name><name><surname>Mark</surname><given-names>RG</given-names></name></person-group>, et al. <article-title>PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals</article-title>. <source>Circulation</source>. (<year>2000</year>) <volume>101</volume>:<fpage>e215</fpage>&#x2013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1161/01.CIR.101.23.e215</pub-id><pub-id pub-id-type="pmid">10851218</pub-id></citation></ref>
<ref id="B22"><label>22.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Johnson</surname><given-names>A</given-names></name><name><surname>Lungren</surname><given-names>M</given-names></name><name><surname>Peng</surname><given-names>Y</given-names></name><name><surname>Lu</surname><given-names>Z</given-names></name><name><surname>Mark</surname><given-names>R</given-names></name><name><surname>Berkowitz</surname><given-names>S</given-names></name></person-group>, et al. <comment>MIMIC-CXR-JPG- chest radiographs with structured labels (version 2.0.0) (2019)</comment>. <comment>doi:10.13026/8360-t248</comment></citation></ref>
<ref id="B23"><label>23.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Johnson</surname><given-names>AEW</given-names></name><name><surname>Pollard</surname><given-names>TJ</given-names></name><name><surname>Berkowitz</surname><given-names>SJ</given-names></name><name><surname>Greenbaum</surname><given-names>NR</given-names></name><name><surname>Lungren</surname><given-names>MP</given-names></name><name><surname>Deng</surname><given-names>C</given-names></name></person-group>, et al. <comment>MIMIC-CXR-JPG: A large publicly available database of labeled chest radiographs. CoRR[Preprint] abs/1901.07042 (2019)</comment>.</citation></ref>
<ref id="B24"><label>24.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Paszke</surname><given-names>A</given-names></name><name><surname>Gross</surname><given-names>S</given-names></name><name><surname>Massa</surname><given-names>F</given-names></name><name><surname>Lerer</surname><given-names>A</given-names></name><name><surname>Bradbury</surname><given-names>J</given-names></name><name><surname>Chanan</surname><given-names>G</given-names></name></person-group>, et al. <comment>PyTorch: An imperative style, high-performance deep learning library. In: Wallach HM, Larochelle H, Beygelzimer A, d&#x2019;Alch&#x00E9;-Buc F, Fox EB, Garnett R, editors, <italic>Advances in Neural Information ProcessingSystems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019,December 8&#x2013;14, 2019, Vancouver, BC, Canada</italic> (2019). p. 8024&#x2013;8035</comment>.</citation></ref>
<ref id="B25"><label>25.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Reddi</surname><given-names>SJ</given-names></name><name><surname>Kale</surname><given-names>S</given-names></name><name><surname>Kumar</surname><given-names>S</given-names></name></person-group>. <comment>On the convergence of adam and beyond. In: <italic>International Conference on Learning Representations</italic> (2018)</comment>.</citation></ref>
<ref id="B26"><label>26.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stember</surname><given-names>JN</given-names></name><name><surname>Celik</surname><given-names>H</given-names></name><name><surname>Gutman</surname><given-names>D</given-names></name><name><surname>Swinburne</surname><given-names>N</given-names></name><name><surname>Young</surname><given-names>R</given-names></name><name><surname>Eskreis-Winkler</surname><given-names>S</given-names></name></person-group>, et al. <article-title>Integrating eye tracking and speech recognition accurately annotates MR brain images for deep learning: proof of principle</article-title>. <source>Radiol Artificial Intell</source>. (<year>2021</year>) <volume>3</volume>:<fpage>e200047</fpage>. <pub-id pub-id-type="doi">10.1148/ryai.2020200047</pub-id></citation></ref>
<ref id="B27"><label>27.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Saab</surname><given-names>K</given-names></name><name><surname>Hooper</surname><given-names>SM</given-names></name><name><surname>Sohoni</surname><given-names>NS</given-names></name><name><surname>Parmar</surname><given-names>J</given-names></name><name><surname>Pogatchnik</surname><given-names>B</given-names></name><name><surname>Wu</surname><given-names>S</given-names></name></person-group>, et al. <comment>Observational supervision for medical image classification using gaze data. In: <italic>Medical Image Computingand Computer Assisted Intervention - MICCAI 2021</italic>. Springer International Publishing (2021). p. 603&#x2013;614</comment>. <pub-id pub-id-type="doi">10.1007/978-3-030-87196-3-56</pub-id></citation></ref>
<ref id="B28"><label>28.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Li</surname><given-names>L</given-names></name><name><surname>Xu</surname><given-names>M</given-names></name><name><surname>Wang</surname><given-names>X</given-names></name><name><surname>Jiang</surname><given-names>L</given-names></name><name><surname>Liu</surname><given-names>H</given-names></name></person-group>. <comment>Attention based glaucoma detection: a large-scale database and CNN model. In: <italic>IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16&#x2013;20, 2019</italic>. Computer Vision Foundation/IEEE (2019). p. 10571&#x2013;10580</comment>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01082</pub-id></citation></ref>
<ref id="B29"><label>29.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname><given-names>S</given-names></name><name><surname>Ouyang</surname><given-names>X</given-names></name><name><surname>Liu</surname><given-names>T</given-names></name><name><surname>Wang</surname><given-names>Q</given-names></name><name><surname>Shen</surname><given-names>D</given-names></name></person-group>. <article-title>Follow my eye: using gaze to supervise computer-aided diagnosis</article-title>. <source>IEEE Trans Med Imaging</source>. (<year>2022</year>) <volume>41</volume>:<fpage>1688</fpage>&#x2013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2022.3146973</pub-id><pub-id pub-id-type="pmid">35085074</pub-id></citation></ref>
<ref id="B30"><label>30.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Agrawal</surname><given-names>M</given-names></name><name><surname>Hegselmann</surname><given-names>S</given-names></name><name><surname>Lang</surname><given-names>H</given-names></name><name><surname>Kim</surname><given-names>Y</given-names></name><name><surname>Sontag</surname><given-names>DA</given-names></name></person-group>. <comment>Large language models are few-shot clinical information extractors. In: Goldberg Y, Kozareva Z, Zhang Y, editors, <italic>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, AbuDhabi, United Arab Emirates, December 7&#x2013;11, 2022</italic>. Association for Computational Linguistics (2022). p. 1998&#x2013;2022</comment>.</citation></ref>
<ref id="B31"><label>31.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Bigolin Lanfredi</surname><given-names>R</given-names></name><name><surname>Schroeder</surname><given-names>JD</given-names></name><name><surname>Tasdizen</surname><given-names>T</given-names></name></person-group>. <comment>Localization supervision of chest x-ray classifiers using label-specific eye-tracking annotation. CoRR [Preprint] abs/2207.09771 (2022)</comment>. <pub-id pub-id-type="doi">10.48550/arXiv.2207.09771</pub-id></citation></ref></ref-list>
</back>
</article>