<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2021.750639</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Comparing Object Recognition in Humans and Deep Convolutional Neural Networks&#x2014;An Eye Tracking Study</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>van Dyck</surname> <given-names>Leonard Elia</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1423972/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kwitt</surname> <given-names>Roland</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1479961/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Denzler</surname> <given-names>Sebastian Jochen</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1489590/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Gruber</surname> <given-names>Walter Roland</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1488206/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Psychology, University of Salzburg</institution>, <addr-line>Salzburg</addr-line>, <country>Austria</country></aff>
<aff id="aff2"><sup>2</sup><institution>Center for Cognitive Neuroscience, University of Salzburg</institution>, <addr-line>Salzburg</addr-line>, <country>Austria</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Computer Science, University of Salzburg</institution>, <addr-line>Salzburg</addr-line>, <country>Austria</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Britt Anderson, University of Waterloo, Canada</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Kohitij Kar, Massachusetts Institute of Technology, United States; Mohammad Ebrahimpour, University of San Francisco, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Leonard Elia van Dyck, <email>leonard.vandyck@plus.ac.at</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Perception Science, a section of the journal Frontiers in Neuroscience</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>750639</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>16</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2021 van Dyck, Kwitt, Denzler and Gruber.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>van Dyck, Kwitt, Denzler and Gruber</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Deep convolutional neural networks (DCNNs) and the ventral visual pathway share vast architectural and functional similarities in visual challenges such as object recognition. Recent insights have demonstrated that both hierarchical cascades can be compared in terms of both exerted behavior and underlying activation. However, these approaches ignore key differences in spatial priorities of information processing. In this proof-of-concept study, we demonstrate a comparison of human observers (<italic>N</italic> = 45) and three feedforward DCNNs through eye tracking and saliency maps. The results reveal fundamentally different resolutions in both visualization methods that need to be considered for an insightful comparison. Moreover, we provide evidence that a DCNN with biologically plausible receptive field sizes called <italic>vNet</italic> reveals higher agreement with human viewing behavior as contrasted with a standard ResNet architecture. We find that image-specific factors such as category, animacy, arousal, and valence have a direct link to the agreement of spatial object recognition priorities in humans and DCNNs, while other measures such as difficulty and general image properties do not. With this approach, we try to open up new perspectives at the intersection of biological and computer vision research.</p>
</abstract>
<kwd-group>
<kwd>seeing</kwd>
<kwd>vision</kwd>
<kwd>object recognition</kwd>
<kwd>brain</kwd>
<kwd>deep neural network</kwd>
<kwd>eye tracking</kwd>
<kwd>saliency map</kwd>
</kwd-group>
<counts>
<fig-count count="10"/>
<table-count count="1"/>
<equation-count count="1"/>
<ref-count count="61"/>
<page-count count="15"/>
<word-count count="10277"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="S1">
<title>Introduction</title>
<p>In the last few years, advances in deep learning have turned rather simple convolutional neural networks, once developed to simulate the complex nature of biological vision, into sophisticated objects of investigation themselves. Especially the increasing synergy between neural and computer sciences has facilitated this interdisciplinary progress with the aim to enable machines to see and to further the understanding of visual perception in living organisms along the way. Computer vision has given rise to deep convolutional neural networks (DCNNs) that exceed human benchmark performance in key challenges of visual perception (<xref ref-type="bibr" rid="B32">Krizhevsky et al., 2012</xref>; <xref ref-type="bibr" rid="B25">He et al., 2015</xref>). Among them, the fundamental ability of core object recognition, which allows humans to identify an enormous number of objects despite their substantial variations in appearance (<xref ref-type="bibr" rid="B14">DiCarlo et al., 2012</xref>) and thus to classify visual inputs into meaningful categories based on previously acquired knowledge (<xref ref-type="bibr" rid="B7">Cadieu et al., 2014</xref>).</p>
<p>In the brain, this information processing task is solved particularly by the ventral visual pathway (<xref ref-type="bibr" rid="B28">Ishai et al., 1999</xref>), which passes information through a hierarchical cascade of retinal ganglion cells (RGC), lateral geniculate nucleus (LGN), visual cortex areas (V1, V2, and V4), and inferior temporal cortex (ITC) (<xref ref-type="bibr" rid="B55">Tanaka, 1996</xref>; <xref ref-type="bibr" rid="B45">Riesenhuber and Poggio, 1999</xref>; <xref ref-type="bibr" rid="B46">Rolls, 2000</xref>; <xref ref-type="bibr" rid="B14">DiCarlo et al., 2012</xref>). This organization shares vast similarities with the purely feedforward architectures of DCNNs in a way that visual information can pass through by means of a single end-to-end sweep (see <xref ref-type="fig" rid="F1">Figure 1</xref>). While in most cases this processing mechanism seems to suffice for so called <italic>early solved</italic> natural images, a substantial body of literature proposes that especially <italic>late-solved</italic> challenge images benefit from recurrent processing through neural interconnections and loops (<xref ref-type="bibr" rid="B34">Lamme and Roelfsema, 2000</xref>; <xref ref-type="bibr" rid="B31">Kar et al., 2019</xref>; <xref ref-type="bibr" rid="B30">Kar and DiCarlo, 2020</xref>). Moreover, electrophysiological findings therefore suggest that recurrence may set in increasingly after around 150 ms to stimulus onset (<xref ref-type="bibr" rid="B13">DiCarlo and Cox, 2007</xref>; <xref ref-type="bibr" rid="B9">Cichy et al., 2014</xref>; <xref ref-type="bibr" rid="B10">Contini et al., 2017</xref>; <xref ref-type="bibr" rid="B56">Tang et al., 2018</xref>; <xref ref-type="bibr" rid="B44">Rajaei et al., 2019</xref>; <xref ref-type="bibr" rid="B52">Seijdel et al., 2020</xref>). Interestingly, DCNNs seem to face difficulties in recognizing exactly these late-solved (<xref ref-type="bibr" rid="B31">Kar et al., 2019</xref>) and manipulated challenge images (<xref ref-type="bibr" rid="B15">Dodge and Karam, 2017</xref>; <xref ref-type="bibr" rid="B20">Geirhos et al., 2017</xref>, <xref ref-type="bibr" rid="B21">2018b</xref>; <xref ref-type="bibr" rid="B59">van Dyck and Gruber, 2020</xref>), which may require additional recurrent processing.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Object recognition in the human brain and DCNNs. <bold>(A)</bold> In the brain, visual information enters via the retina before it passes through the ventral visual pathway, consisting of the visual cortex areas (V1, V2, and V4) and inferior temporal cortex (ITC). After a first feedforward sweep of information (&#x223C;150 ms), recurrent processes reconnecting higher to lower areas of this hierarchical cascade become activated and allow more in-depth visual processing. <bold>(B)</bold> ResNet18 (<xref ref-type="bibr" rid="B26">He et al., 2016</xref>) is a standard DCNN with 5 convolutional layers. <bold>(C)</bold> vNet (<xref ref-type="bibr" rid="B38">Mehrer et al., 2021</xref>) is a novel DCNN with 10 convolutional layers and modified effective kernel sizes, which simulate the progressively increasing <italic>receptive field size</italic> (RFS) in the ventral visual pathway. <bold>(D)</bold> Schematic overview of convolutions with a constant and increasing RFS. An increasing RFS raises the number of pixels represented within individual neurons of the later layers. Below individual layer names, the first number represents the feature map size while the second and third indicate the obtained output size.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g001.tif"/>
</fig>
<p>Naturally, images of objects, regardless of whether represented on a biological retina or in a computer matrix, are highly complex. Therefore, in the brain, object recognition is influenced by a number of bottom-up (<xref ref-type="bibr" rid="B51">Rutishauser et al., 2004</xref>) and top-down processes (<xref ref-type="bibr" rid="B2">Bar, 2003</xref>; <xref ref-type="bibr" rid="B3">Bar et al., 2006</xref>). In simplified terms, however, it is thought to solve this challenge by transferring given two-dimensional information into an invariant three-dimensional representation, encoded in an even higher-dimensional neuronal space (<xref ref-type="bibr" rid="B14">DiCarlo et al., 2012</xref>). As pointed out by <xref ref-type="bibr" rid="B37">Marr (1982)</xref>, the two-dimensional image allows first and foremost the extraction of shape features such as edges and regions, the precursors of the so-called <italic>primal sketch</italic> (<xref ref-type="bibr" rid="B37">Marr, 1982</xref>, p. 37). Then, when textures and shades start to enrich the outline, a <italic>2.5D sketch</italic> (<xref ref-type="bibr" rid="B37">Marr, 1982</xref>, p. 37) emerges. Finally, in combination with previously acquired knowledge, the representation of an invariant <italic>3D model</italic> can be inferred (<xref ref-type="bibr" rid="B37">Marr, 1982</xref>, p. 37). These processing steps are reflected not only by the architecture of the ventral visual pathway but can also be found in its artificial replica. While in DCNNs, earlier layers are mostly sensitive to specific configurations of edges, blobs, and colors (e.g., the edge of an orange stripe), the following convolutions start to combine them into texture-like feature groups (e.g., an orange-white striped pattern), until later layers assemble whole object parts (e.g., the fin of an anemone fish) and eventually infer the object class label (e.g., an anemone fish). Interestingly, recent literature suggests that DCNNs trained on the ImageNet dataset (<xref ref-type="bibr" rid="B12">Deng et al., 2009</xref>; <xref ref-type="bibr" rid="B49">Russakovsky et al., 2015</xref>) are strongly biased toward texture and use it more frequently rather than shape information to classify images (<xref ref-type="bibr" rid="B22">Geirhos et al., 2018a</xref>). This preference seems to contradict findings in humans that clearly identify shape as the single most important cue for object recognition (<xref ref-type="bibr" rid="B35">Landau et al., 1988</xref>). In addition, several studies have examined the effect of context in human object recognition (<xref ref-type="bibr" rid="B41">Oliva and Torralba, 2007</xref>; <xref ref-type="bibr" rid="B23">Greene and Oliva, 2009</xref>). Contextual cues are representational regularities such as spatial layout (e.g., houses are usually located on the ground), inter-object dependencies (e.g., houses are usually attached to a street), point of view (e.g., houses are usually looked at from a specific perspective), and other summary statistics (i.e., entropy and power spectral density). Likewise, these same cues are certainly not meaningless to DCNNs. In fact, mainly work around removing or manipulating one of these hints such as the object&#x2019;s regular pose (<xref ref-type="bibr" rid="B1">Alcorn et al., 2019</xref>) or background (<xref ref-type="bibr" rid="B4">Beery et al., 2018</xref>) have demonstrated that indeed contextual learning, which is common practice in learning systems, can be found here as well (<xref ref-type="bibr" rid="B19">Geirhos et al., 2020</xref>).</p>
<p>Nevertheless, a key difference between humans and DCNNs lies within the modulating effects of appraisal. While millions of years of evolution have tuned <italic>in vivo</italic> object recognizers such as humans to seek and avoid different kinds of stimuli (e.g., find nourishment and avoid dangers quickly), this subjective experience of for example arousal and valence is lacking completely in <italic>in silico</italic> models. Furthermore, the neuroscientific literature contains many examples suggesting that particularly the automatic detection of fear-relevant and threatful objects is solved by an even faster subcortical route, which skips parts of the ventral visual pathway through shortcuts to the amygdala (<xref ref-type="bibr" rid="B40">&#x00D6;hman, 2005</xref>; <xref ref-type="bibr" rid="B42">Pessoa and Adolphs, 2010</xref>). This for example might enable threat-superiority effects in terms of reaction times and reduced position effects in a visual search paradigm (<xref ref-type="bibr" rid="B5">Blanchette, 2006</xref>). While the plausibility of this so-called <italic>low road</italic> is highly discussed (<xref ref-type="bibr" rid="B8">Cauchoix and Crouzet, 2013</xref>), the differences in performance are well-documented.</p>
<p>As an interim summary, it can be noted that the ventral visual pathway and DCNNs suggest conceptual overlaps but also substantial differences in many regards. Hence, more recent evidence from <xref ref-type="bibr" rid="B38">Mehrer et al. (2021)</xref> further highlights the importance of biological plausibility for the fit between brain and DCNN activity, as their novel architecture called <italic>vNet</italic> simulates the progressively increasing foveal <italic>receptive field size</italic> (hereafter abbreviated as RFS) along the ventral visual pathway (<xref ref-type="bibr" rid="B60">Wandell and Winawer, 2015</xref>; <xref ref-type="bibr" rid="B24">Grill-Spector et al., 2017</xref>). However, as their analyses point out, this modification does not lead to higher congruence with for example fMRI activity of human observers when compared to a standard architecture such as AlexNet (<xref ref-type="bibr" rid="B32">Krizhevsky et al., 2012</xref>). This raises many questions about the definite impact of this RFS modification. Therefore, as vNet&#x2019;s hierarchical organization is designed to resemble that of the ventral visual pathway more accurately as compared to a standard DCNN, here we hypothesize that the major advantage of vNet may not be visible within more brain-like activations, as also not found by <xref ref-type="bibr" rid="B38">Mehrer et al. (2021)</xref>, but rather more similar spatial priorities of information processing compared through eye tracking and GradCAM. Consequently, following an important distinction in human-machine comparisons by <xref ref-type="bibr" rid="B17">Firestone (2020)</xref>, we believe that this resemblance in underlying <italic>competence</italic> should lead to higher similarity in object recognition behavior and further observable <italic>performance</italic>. As in this specific human-machine comparison rather divergent architectures are compared, we believe that a RFS modification can result in more similar spatial priorities of information processing and likewise object recognition behavior, without immediately suggesting a higher match between activity patterns of neural components and individual DCNN layers.</p>
<p>As DCNNs are getting more and more complex, several <italic>attribution</italic>-tools have been developed to understand (to some extent) their classifications. At first glance, these new visualization algorithms, also called <italic>saliency maps</italic>, resemble well-established methods in cognitive neuroscience. In eye tracking, the execution of a visual task is analyzed by mapping gaze behavior onto specific regions of interest, which receive special cognitive or computational priorities during information processing. While eye tracking and saliency maps share this concept, the methodological way this is achieved seems fundamentally different. In eye tracking measurements, a human observer is presented with a stimulus, which is only presented for a limited time, while viewing behavior is being recorded. Saliency maps such as <italic>Gradient-weighted Class Activation Mapping</italic>, also known as <italic>GradCAM</italic> (<xref ref-type="bibr" rid="B53">Selvaraju et al., 2017</xref>), extract class activations within a specific layer of the DCNN to explain obtained predictions. Therefore, the algorithm uses the gradient of the loss function to compute a weight for every feature map. The weighted sum of these activations is class-discriminative and hence allows the localization and visualization of all relevant regions that contributed to the probability of a given class. However, DCNNs do not operate on a meaningful time scale (other than related to computational power) when trying to classify images. Despite the conceptual similarity between eye tracking and saliency maps, only a few attempts have been made to draw this comparison of black boxes (<xref ref-type="bibr" rid="B16">Ebrahimpour et al., 2019</xref>). Importantly, a major challenge in this field is to encourage and conduct fair human-machine comparisons, as it is only possible to infer similarities and differences if there are no fundamental constraints within the comparison itself (<xref ref-type="bibr" rid="B17">Firestone, 2020</xref>; <xref ref-type="bibr" rid="B18">Funke et al., 2020</xref>). In this study, we compare human eye tracking to DCNN saliency maps in an approximately species-fair object recognition task and examine a wide range of possible factors influencing similarity measures.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2.SS1">
<title>General Procedure</title>
<p>In order to test our hypotheses, we designed a fair human-machine comparison that allowed us to investigate behavioral, eye tracking, and physiological data from human observers performing a laboratory experiment, as well as predictions and activations of three DCNN architectures on identical visual stimuli and under roughly similar conditions. The main task in this experiment was to categorize briefly shown images based on a forced-choice format of 12 basic-level categories, namely <italic>human</italic>, <italic>dog</italic>, <italic>cat</italic>, <italic>bird</italic>, <italic>fish</italic>, <italic>snake</italic>, <italic>car</italic>, <italic>train</italic>, <italic>house</italic>, <italic>bed</italic>, <italic>flower</italic>, <italic>ball</italic> (see section Methods). Basic-level categories (e.g., <italic>dog</italic>) were chosen over detailed category concepts which are more often used in the field of computer vision (e.g., <italic>border collie</italic>), as they are more naturally utilized in human object classification (<xref ref-type="bibr" rid="B47">Rosch, 1999</xref>).</p>
</sec>
<sec id="S2.SS2">
<title>Human Observers&#x2014;Eye Tracking Experiment</title>
<p>A total of 45 valid participants (28 female, 17 male) with an age between 18 and 31 years (M = 22.64, <italic>SD</italic> = 2.57) were tested in the eye tracking experiment. Participants were required to have normal or corrected-to-normal vision without problems of color perception and other eye diseases. Two participants with contact lenses had to be excluded from further analyses due to insufficient eye tracking precision. The experimental procedure was admitted by the University of Salzburg ethics committee, in line with the declaration of Helsinki and agreed to by participants via written consent before the experiment. Psychology students received accredited participation hours for taking part in the experiment.</p>
<p>The experiment consisted of two main tasks (see <xref ref-type="fig" rid="F2">Figure 2</xref>). Participants had to concentrate on a fixation cross for 500 ms until an image appeared at center 12.93 degrees of visual angle away on the left or right side. If, as in one condition, the image was presented for a short duration of 150 ms, a visual backward mask (1/f noise) followed for the same duration, and the participant had to classify the presented object based on a forced-choice format of 12 basic-level categories by clicking on the respective class symbol. If, as in the other condition, the image was presented for a long duration of 3000 ms, participants had to rate it afterward in its arousal (1 = <italic>Very low</italic>, 4 = <italic>Neither low nor high</italic>, 7 = <italic>Very high</italic>) and valence (1 = <italic>Very negative</italic>, 4 = <italic>Neutral</italic>, 7 = <italic>Very positive</italic>) on a scale from 1 to 7. As both classification and rating tasks were balanced out in occurrence and previously pseudo-randomized, it was impossible for the observers to differentiate between the two conditions before the short presentation time was exceeded. Based on this central assumption, both conditions should not vary in viewing behavior. Participants were familiarized with the experimental procedure during training trials which were excluded from further analyses. The whole experiment consisted of 420 test trials (210 per condition), took about 1 h to complete, and was divided into three blocks with resting breaks in between. In this way, a single participant classified one half of the entire dataset and rated the other half. To obtain categorization and rating results for all images, there were two versions of the experiment with interchanged conditions for both halves.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Object recognition paradigm. <bold>(A)</bold> During categorization trials, human observers had to focus on a fixation cross (500 ms, including a 150 ms fixation control) before an image was presented on either the left or the right side (150 ms), followed by a visual backward mask (150 ms, 1/f noise), and a forced-choice categorization (self-paced, max. 7,500 ms). <bold>(B)</bold> During rating trials and after the fixation cross, an image was presented again on the left or the right side (3,000 ms), followed by an arousal rating and valence rating (both self-paced, max. 7,500 ms each). As conditions were pseudo-randomized, human observers were not able to anticipate, whether an image needed to be categorized or rated until the initial 150 ms had elapsed. This allows the assumption that in the first 150 ms the conditions should not vary in viewing behavior.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g002.tif"/>
</fig>
<p>The participants performed the experiment in a laboratory room, where they were seated in front of a screen (1,920 &#x00D7; 1,080 pixels, 50 Hz) and had to place their head into a chin rest located at a distance of 60 cm. The right eye was tracked and recorded with an EyeLink 1000 (SR Research Ltd., Mississauga, ON, Canada) desktop mount, at a sampling rate of 1,000 Hz. For presentational purposes, original images were scaled by factor two onscreen (448 &#x00D7; 448) but stayed unchanged in image resolution (224 &#x00D7; 224). This way, the presented images had 11.52 degrees of visual angle in size. Recorded eye tracking data were preprocessed using DataViewer (Version 4.2, SR Research Ltd., Mississauga, ON, Canada) and analyzed after the participants gaze crossed an invisible boundary framing the entire image. In this way, participants were able to process appearing images already peripherally for the first couple of milliseconds to allow a meaningfully programmed first fixation (see <xref ref-type="fig" rid="F3">Figure 3</xref>). Fixations were compiled for the first 150 ms within the image. Here, x- and y-coordinates were downscaled again from expanded presentation size (448 &#x00D7; 448) to original image size (224 &#x00D7; 224). Average heatmaps were computed from individual sampling points of either the first 150 ms (= feedforward) or the entire presentation time after 150 ms (= recurrent) within the image using in-house built MATLAB scripts. The obtained heatmaps for all participants were averaged per image, Gaussian filtered with a standard deviation of 15 pixels, and normalized.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Dispersion of eye tracking fixations along x- and y-coordinates across categories. Root Mean Squared Errors (RMSEs) were computed for individual images and indicate the magnitude of deviation among the fixations of individual participants on a single image. Generally, the fixations seemed to be rather precise on an image-by-image level with an average dispersion of below 30 pixels across all participants. This suggests that meaningful features were targeted and that the centroids, which were used for further analyses, can be regarded as characteristic for the human observer sample. The dispersion along <bold>(A)</bold> x-coordinates was slightly higher as compared to <bold>(B)</bold> y-coordinates, which is thought to reflect the reported central fixation and saccadic motor biases. The dispersion of an image can also be increased by the presence of multiple meaningful features, without any loss of precision. However, this problem should occur rarely, as most of the images showed only one dominant object.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g003.tif"/>
</fig>
<p>Additionally, before the eye tracking experiment started, participants were wired with ECG and SCL electrodes. Physiological measurements were collected using a Varioport biosignal recorder (Becker Meditec, Karlsruhe, Germany) and analyzed with ANSLAB (<xref ref-type="bibr" rid="B6">Blechert et al., 2016</xref>) (Version 2.51, Salzburg, Austria). Here, the relative change in mean activity from a baseline of 1000 ms before the image presentation to the time window of 3000ms during the image presentation was used.</p>
</sec>
<sec id="S2.SS3">
<title>Deep Convolutional Neural Network Training and GradCAM Saliency Maps</title>
<p>Our implementation (in PyTorch) of vNet follows the original GroupNorm variant of vNet from <xref ref-type="bibr" rid="B38">Mehrer et al. (2021)</xref>, the ResNet18 architecture follows the original proposal from <xref ref-type="bibr" rid="B26">He et al. (2016)</xref>. For training both network types, we use the (publicly available) Ecoset <italic>training</italic> set, constrained to the 12 categories described in section Methods below. As the number of images differs significantly across categories, we artificially balance the corpus by drawing (uniformly at random) <italic>N</italic> images per category, where <italic>N</italic> corresponds to the number of images in the smallest category (i.e., <italic>fish</italic>). During training, all (&#x223C;15 k) images are first resized to a spatial resolution of 256&#x00D7; 256, then cropped to the center square of size 224 &#x00D7; 224 and eventually normalized (by subtracting the channel-wise mean and dividing by the channel-wise standard deviation, computed from the training corpus). Note that no data augmentation is applied across all experiments. We minimize the cross-entropy loss using stochastic gradient descent (SGD) with momentum (0.9) and weight decay (1e&#x2013;4) under a cosine learning rate schedule, starting at an initial learning rate of 0.01. We train for 80 epochs using a batch size of 128. Results for the fine-tuned ResNet18 are obtained by replacing the final linear classifier of an ImageNet-trained ResNet18, freezing all earlier layers, and fine-tuning for 50 epochs (in the same setup as described before). When referring to <italic>early</italic>, <italic>middle</italic>, and <italic>late</italic> layers in the manuscript, we refer to GradCAM outputs generated from activations after the 2nd, 6th, and 9th layer for vNet, and activations after the 1st, 3rd, and 4th ResNet18 block (as all ResNet architectures total four blocks). Eventually, if the aim is to investigate the similarity in spatial priorities of information processing that underlie object recognition behavior with the demonstrated approach, the output layer of a DCNN would suffice for the comparison. In this proof-of-concept study, however, we also include earlier layers for sanity checks and further model comparisons.</p>
</sec>
<sec id="S2.SS4">
<title>Dataset</title>
<p>Images were part of 12 basic-level categories from the ecologically motivated Ecoset dataset, which was created by <xref ref-type="bibr" rid="B38">Mehrer et al. (2021)</xref> in order to better capture the organization of human-relevant categories. The dataset consisted of 6 animate (namely <italic>human</italic>, <italic>dog</italic>, <italic>cat</italic>, <italic>bird</italic>, <italic>fish</italic>, and <italic>snake</italic>) and 6 inanimate categories (namely <italic>car</italic>, <italic>train</italic>, <italic>house</italic>, <italic>bed</italic>, <italic>flower</italic>, and <italic>ball</italic>) with 30 images per category. During preprocessing, images were randomly drawn from the test set, cropped toward the biggest possible central square, and resized to 224 &#x00D7; 224 pixels. All images were visually checked and excluded if multiple categories (e.g., <italic>human</italic> and <italic>dog</italic>), overlayed text, or image effects (e.g., grayscale images) were visible or the object was fully removed during preprocessing steps. In categories where less than 30 images from the test set remained, Ecoset images were complemented with ImageNet examples (<italic>n</italic> = 28, across 4 categories). Additionally, in 3 animate (namely <italic>human</italic>, <italic>dog</italic>, and <italic>snake</italic>) and 3 inanimate categories (namely <italic>car</italic>, <italic>house</italic>, <italic>flower</italic>), 10 images per category of the respective objects were added from the Open Affective Standardized Image Set (OASIS) (<xref ref-type="bibr" rid="B33">Kurdi et al., 2017</xref>). Here, arousal and valence ratings from a large number of participants (<italic>N</italic> = 822) were already available and increased the variability while serving as a sanity check for own ratings. In total, the dataset consisted of 420 test images.</p>
</sec>
</sec>
<sec sec-type="results" id="S3">
<title>Results</title>
<sec id="S3.SS1">
<title>Performance</title>
<p>The first set of analyses investigated object recognition performance in human observers and DCNNs. Therefore, human predictions obtained during categorization trials (see section Methods) were compared against model predictions. Generally, as the categorization data were not normally distributed, non-parametric tests were applied to compare the human observer sample against fixed-accuracy values of individual DCNNs. On average, human observers reached a recognition accuracy of 89.96%. One-sample Wilcoxon tests indicated that human observers were significantly outperformed by fine-tuned ResNet18 with 95.48% [<italic>V</italic> = 0, CI = (89.52, 91.43), <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.85] but significantly more correct than both <italic>trained-from-scratch</italic> ResNet18 [<italic>V</italic> = 1,035, CI = (89.52, 91.43), <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.85] and vNet [<italic>V</italic> = 1,034, CI = (89.52, 91.43), <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.85] with accuracies of 70.48 and 76.43%, respectively. The results endorse both sides of the literature by demonstrating that especially DCNNs trained on large datasets can exceed human benchmark performance (<xref ref-type="bibr" rid="B32">Krizhevsky et al., 2012</xref>; <xref ref-type="bibr" rid="B54">Szegedy et al., 2015</xref>; <xref ref-type="bibr" rid="B26">He et al., 2016</xref>; <xref ref-type="bibr" rid="B27">Huang et al., 2017</xref>), but also simultaneously reminds of possible limits due to the amount of provided training data. Nevertheless, following analyses focus predominantly on the two equally trained DCNNs, as they are more suitable for a fair comparison of architectures.</p>
<p>Remarkably, vNet outperformed ResNet18 throughout all grouping variables (animacy, arousal, valence, and category) and generally seemed to be closer to the human benchmark level of accuracy (see <xref ref-type="fig" rid="F4">Figure 4</xref>). However, across categories, applied Kruskal Wallis tests revealed that both ResNet18 [X<sup>2</sup>(29) = 60.18, <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.14] and vNet [X<sup>2</sup>(33) = 58.65, <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.11] performed significantly dissimilar to human observers and each other [X<sup>2</sup>(33) = 62.76, <italic>p</italic> = 0.001, <italic>r</italic> = 0.61]. Moreover, based on previous findings, we hypothesized that human observers should be significantly better at recognizing animate compared to inanimate objects (<xref ref-type="bibr" rid="B39">New et al., 2007</xref>). However, Wilcoxon rank sum tests indicated the existence of this effect but in the opposite direction, as human observers were significantly more accurate in recognizing inanimate (Median = 93.27) compared to animate objects (Median = 88.68; W = 345.5, <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.55). Additionally, ResNet18 and vNet mirrored this behavior with a substantial increase in accuracy from animate (ResNet18: 61.90%/vNet: 68.10%) to inanimate objects (ResNet18: 79.05%/vNet: 84.76%).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Object recognition accuracy of the human observer sample, fine-tuned ResNet18 (marked as FT), and both trained-from-scratch ResNet18 and vNet. Remarkably, vNet is more accurate then ResNet18 across all grouping variables. <bold>(A)</bold> Human observers were outperformed by fine-tuned ResNet18 but more accurate than both ResNet18 and vNet. <bold>(B)</bold> Human observers were better at recognizing inanimate compared to animate objects. This relationship held for DCNNs as well. <bold>(C)</bold> The effect of arousal on human observers indicated more inaccurate recognition with increasing arousal. <bold>(D)</bold> The effect of valence on human observers indicated more accurate recognition with increasing valence. <bold>(E)</bold> Human observers showed a small effect of category, while both trained from scratch DCNNs faced large variability between individual categories. Confidence intervals for the human observer sample were estimated with Hodges-Lehmann procedure on a significance level of <italic>p</italic> = 0.05.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g004.tif"/>
</fig>
<p>Furthermore, low, medium, and high arousal and valence groups, obtained by tertile splits of human observer ratings (33rd and 66th percentile, see section Methods), revealed a significant, negative Spearman rank correlation between human accuracy and arousal (rho = &#x2212;0.22, <italic>p</italic> &#x003C; 0.001). Interestingly, this relationship did not disappear when partial correlations controlling for animacy were computed (rho = &#x2212;0.20, <italic>p</italic> = 0.001). Subsequent Kruskal Wallis tests hinted at an effect of arousal [X<sup>2</sup>(2) = 38.51, <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.50] and Bonferroni-corrected, pairwise Wilcoxon tests revealed a significant decrease of accuracy between all three increasing levels of arousal. A significant positive correlation (rho = 0.34, <italic>p</italic> &#x003C; 0.001), partial correlation controlling for animacy (rho = 0.37, <italic>p</italic> &#x003C; 0.001), and an effect [X<sup>2</sup>(2) = 36.90, <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.48] with a significant increase of accuracy between low and medium levels in Bonferroni-corrected <italic>post-hoc</italic> tests were found for valence. Similarly, the accuracies of ResNet18 and vNet seemed to follow both effects. Images that were assessed as more calm and positive by human observers during rating trials lead to better recognition performance in human observers and DCNNs.</p>
<p>It is fundamental to note that only the performance of vNet exhibited a significant, positive Spearman rank correlation with human observer performance (see <xref ref-type="table" rid="T1">Table 1</xref>; rho = 0.14, <italic>p</italic> = 0.009), which may indicate a better fit to human categorization behavior. Contrary to our expectations, image properties (namely entropy, shape, texture, and power spectral peak-to-mean ratio) seemed to be rather unrelated to performance. Yet, as expected, the parameters were highly correlated among each other and demonstrated the statistical regularities of complex natural images.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Spearman rank correlation between human observer accuracy, DCNN accuracy, and image properties.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td/>
<td valign="top" align="center">Human Acc.</td>
<td valign="top" align="center">ResNet18 (FT) Acc.</td>
<td valign="top" align="center">ResNet18 Acc.</td>
<td valign="top" align="center">vNet Acc.</td>
<td valign="top" align="center">Entropy</td>
<td valign="top" align="center">Shape</td>
<td valign="top" align="center">Texture</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Human Acc.</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">ResNet18 (FT) Acc.</td>
<td valign="top" align="center">0.071</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">ResNet18 Acc.</td>
<td valign="top" align="center">0.030</td>
<td valign="top" align="center">0.236<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">vNet Acc.</td>
<td valign="top" align="center">0.127<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;</xref></td>
<td valign="top" align="center">0.203<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td valign="top" align="center">0.551<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Entropy</td>
<td valign="top" align="center">0.041</td>
<td valign="top" align="center">&#x2013;0.071</td>
<td valign="top" align="center">&#x2013;0.053</td>
<td valign="top" align="center">0.027</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Shape</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">&#x2212;0.096<xref ref-type="table-fn" rid="tfn1">&#x002A;</xref></td>
<td valign="top" align="center">&#x2013;0.025</td>
<td valign="top" align="center">&#x2013;0.019</td>
<td valign="top" align="center">0.299<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Texture</td>
<td valign="top" align="center">0.036</td>
<td valign="top" align="center">&#x2013;0.056</td>
<td valign="top" align="center">&#x2013;0.044</td>
<td valign="top" align="center">0.000</td>
<td valign="top" align="center">0.432<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td valign="top" align="center">0.763<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">Peak-to-Mean</td>
<td valign="top" align="center">0.015</td>
<td valign="top" align="center">0.058</td>
<td valign="top" align="center">0.072</td>
<td valign="top" align="center">0.048</td>
<td valign="top" align="center">&#x2212;0.271<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td valign="top" align="center">&#x2212;0.591<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
<td valign="top" align="center">&#x2212;0.874<xref ref-type="table-fn" rid="tfn1">&#x002A;&#x002A;&#x002A;</xref></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="tfn1"><p><italic>Acc. stands for Accuracy. Peak-to-Mean stands for Power Spectral Peak-to-Mean Ratio. Computed correlation used spearman-method with listwise-deletion. &#x002A;p &#x003C; 0.05, &#x002A;&#x002A;p &#x003C; 0.01, &#x002A;&#x002A;&#x002A;p &#x003C; 0.001.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>In order to shine more light on the classification errors made by human observers and DCNNs, which underlie the reported performances, categorization patterns were investigated (see <xref ref-type="fig" rid="F5">Figure 5</xref>). Interestingly, as human observers seemed to have difficulties with relatively common classes (such as dog and cat or fish and snake), their classification behavior suggests that these conceptually similar classes could have lead to confusions. It should also be taken into account that other top-down and bottom-up influencing factors such as contextual cues may be especially similar between these classes. Moreover, both trained-from-scratch ResNet18 and vNet were found to misclassify images in similar ways.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Categorization matrices reveal the classification patterns of true vs. predicted labels that underly raw object recognition performances and thereby help to understand especially classification errors. The off-diagonal misclassifications show that human observers seemed to have problems with conceptually similar classes (such as dog and cat or fish and snake). Interestingly, especially both trained-from-scratch ResNet18 and vNet made similar mistakes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g005.tif"/>
</fig>
</sec>
<sec id="S3.SS2">
<title>Fixations&#x2014;Feedforward vs. Feedforward Processing</title>
<p>In an attempt to identify priorities during feedforward information processing in human observers and DCNNs, we inspected human fixations and global maxima of saliency maps. Fixations were compiled for the first 150 ms, the theoretical time of a feedforward pass, and assigned to 1 out of 16 equally sized target blocks. Similarly, for GradCAM, the centroid of the single highest scoring patch was defined as the global maximum and used as an equivalent with the respective target block. As displayed in <xref ref-type="fig" rid="F6">Figure 6</xref>, we found that human fixations were subject to a <italic>central fixation bias</italic> (<xref ref-type="bibr" rid="B57">Tatler, 2007</xref>; <xref ref-type="bibr" rid="B48">Rothkegel et al., 2017</xref>), as most individual and almost all average fixations were located within the center blocks. Furthermore, a <italic>saccadic motor bias</italic> was visible, as average fixations of images presented on the left side were predominantly located on the right side of the image and vice versa. This pattern is thought to reflect the preference of the saccadic system for smaller amplitude eye movements over larger ones (<xref ref-type="bibr" rid="B58">Tatler et al., 2006</xref>). In most cases, however, meaningful fixations on object features could be clearly identified on the individual level.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>Uncorrected spatial priorities during feedforward object recognition in human observers and DCNNs. <bold>(A)</bold> Human fixations were affected by a central fixation and saccadic motor bias, while being normally distributed on a continuous scale. <bold>(B,C)</bold> ResNet18 and vNet GradCAM maxima displayed a near uniform distribution in early and middle layers with a 7 &#x00D7; 7 and 4&#x00D7; 4 grid in the late layers. Generally, maxima followed a discrete segmentation that was identical to the output size of the respective layer.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g006.tif"/>
</fig>
<p>Although we hypothesized GradCAM maxima to differ substantially from human fixations, the results of ResNet18 and vNet across early, middle, and late layers proposed fundamentally diverging distributions deeper down the architectures. While human fixations were found to be normally distributed on a continuous scale of coordinates, GradCAM maxima, especially in middle and late layers, were located on discrete grids of different sizes. As these grids of possible maxima (ResNet18 Late = 7 &#x00D7; 7 and vNet Late = 4 &#x00D7; 4) were identical with the output sizes of the respective layers (see <xref ref-type="fig" rid="F1">Figures 1B,C</xref>), we attributed this behavior to both the agglomerative nature of convolutions and the resulting technical aspects of how the GradCAM algorithm extracts class activations from layers.</p>
<p>To our knowledge, this characteristic of DCNN attribution methods has not been explicitly considered during previous human-machine comparisons so far. Since these findings restrict further comparisons of Euclidean distance measures between human fixations and GradCAM maxima, we proceeded by investigating only the specific target blocks, in which the respective fixations or maxima fell. As displayed in <xref ref-type="fig" rid="F7">Figure 7</xref>, this analysis promoted more similar object recognition priorities between human observers and vNet.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Corrected spatial priorities during feedforward object recognition in human observers and DCNNs&#x2019; late layers. As the smallest resolution of output sizes allowed this resolution, all fixations and maxima were assigned to respective target blocks accordingly. <bold>(A)</bold> The target block percentages implied more similar spatial priorities between human observers and vNet, as mostly center blocks were targeted. In contrast, both ResNet18 models focused more on marginal target blocks. <bold>(B)</bold> Match between human and DCNN late layer target blocks across different grouping variables. vNet matched human target blocks more frequently. <bold>(C&#x2013;E)</bold> DCNNs had a higher agreement on images of animate objects, with higher arousal ratings, and higher valence ratings. <bold>(F)</bold> vNet obtained especially high agreement on specific categories such as <italic>human</italic> and <italic>dog</italic>, while fine-tuned ResNet18 seemed to systematically choose different target block in <italic>house</italic> images. <bold>(G)</bold> Surprisingly, GradCAM maxima matched most frequently with human fixations in early, decreased in middle, and increased again in late layers. The dotted line represents chance level at 6.25%.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g007.tif"/>
</fig>
<p>In principle, the agreement between human observers and DCNNs in target blocks (see <xref ref-type="fig" rid="F7">Figure 7</xref>) was rather low. The results indicated persisting differences after correcting for the discovered mismatches in resolution, as vNet coincided with human target blocks in 17.14% compared to ResNet18 with 11.67%. Generally, the agreement seemed to be higher for animate objects compared to inanimate objects and vNet prioritized especially more human target blocks on images of specific animate objects (especially <italic>human</italic> and <italic>dog</italic>). These findings further strengthened our confidence that vNet compared to ResNet18 indeed utilizes spatial priorities that are more similar to those of human observers during feedforward object recognition. These insights offer compelling evidence for a more human-like vNet and generally the impact of a RFS modification in terms of fit to human eye tracking data.</p>
</sec>
<sec id="S3.SS3">
<title>Heatmaps&#x2014;Recurrent vs. Feedforward Processing</title>
<p>Further analyses were conducted to compare human observers and DCNNs based on eye tracking and saliency heatmaps. Here, the focus was shifted away from feedforward mechanisms, as additionally recurrent processes with human eye movements after 150 ms were investigated. We hypothesized that differences, which had already existed during early processing (i.e., effects of animacy, arousal, and valence), should be amplified in the brain due to mostly top-down processes setting in during this time window. It is important to note that this analysis is rather unfair in its nature, as it compares feedforward processing in DCNNs with additional recurrent processing in humans. However, in the light of this knowledge, it is entirely possible to test further hypotheses. Generally speaking, the obtained heatmaps illustrated the expected idiosyncrasies. With a few exceptions, human observers fixated the specific object within the first milliseconds and later shifted their attention to more relevant features (such as faces or arousing image parts). In contrast, DCNNs displayed their hierarchical organization with activation of especially specific shape and texture features in early, feature groups in middle, and whole objects in late layers (see <xref ref-type="fig" rid="F8">Figure 8</xref>).</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption><p>Examples of human observer heatmaps and DCNN GradCAMs of different layers. Images with <bold>(A)</bold> the highest and <bold>(B)</bold> the lowest correlation between human observers and ResNet18. Images with <bold>(C)</bold> the highest and <bold>(D)</bold> the lowest correlation between human observers and vNet. Images highlighted in green were categorized correctly while images highlighted in red were categorized incorrectly.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g008.tif"/>
</fig>
<p>Mean absolute error (MAE), defined as the individual deviation of a GradCAM from its human heatmap equivalent, was computed for all individual images. On average, MAEs of 0.29 for ResNet18, and 0.22 for vNet were found. These results fit well with previous findings by <xref ref-type="bibr" rid="B16">Ebrahimpour et al. (2019)</xref>, who reported values of around 0.40 for a scene viewing task, and our previous outcomes promoting vNet as a better model for human eye tracking heatmaps. On top of that, MAE did not seem to be associated with general performance, as the fine-tuned ResNet18, which significantly outperformed human observers, reached only 0.32.</p>
<disp-formula id="S3.E1"><label>(1)</label><mml:math id="M1" display="block"><mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>W</mml:mi><mml:mpadded width="+5pt"><mml:mi>x</mml:mi></mml:mpadded><mml:mi>H</mml:mi></mml:mrow></mml:mfrac><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>W</mml:mi></mml:munderover><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>H</mml:mi></mml:munderover><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p><bold>Equation 1</bold> mean absolute error (MAE) where W and H are the width and height of the original image in pixels (here 224 &#x00D7; 224), while E and S are the eye tracking heatmap and saliency map of the same image.</p>
<p>In order to link these results to the aforementioned control vs. challenge distinction by <xref ref-type="bibr" rid="B31">Kar et al. (2019)</xref>, we treated images which were categorized correctly by human observers (avg. accuracy = 100%) and DCNNs as control images (<italic>n</italic> = 145) and images which were categorized correctly by human observers (avg. accuracy = 100%) but categorized incorrectly by ResNet18 and vNet as challenge images (<italic>n</italic> = 60). Analyses across individual layers showed that the reported difference in MAE between both DCNNs seemed to emerge in late layers (see <xref ref-type="fig" rid="F9">Figure 9</xref>). These findings suggest that control and challenge images do not seem to be treated differently by the visual system in terms of their spatial priorities of information processing. Taken together with the findings of <xref ref-type="bibr" rid="B31">Kar et al. (2019)</xref> this means that the match between eye tracking and GradCAM data was not influenced by this distinction based on their difficulty.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption><p>Mean absolute error (MAE) of ResNet18 and vNet from human observer heatmaps for individual images. Control images were categorized correctly by human observers, ResNet18, and vNet, while challenge images were categorized correctly by human observers but categorized incorrectly by ResNet18 and vNet. <bold>(A,B)</bold> MAEs seemed to be lowest in early layers. Here, the funnel shaped distribution hinted toward a lower boundary which may be a consequence of the discrete scale of GradCAMs. <bold>(C)</bold> As the distribution spread apart in later layers, the clearly smaller MAE of vNet became visible as more images could be found below the diagonal break-even line.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g009.tif"/>
</fig>
</sec>
<sec id="S3.SS4">
<title>Arousal and Valence</title>
<p>On average, human observers rated images with median scores of 4.04 in arousal and 4.27 in valence (see <xref ref-type="fig" rid="F10">Figure 10</xref>). In terms of arousal, Kruskal Wallis tests suggested that animate objects received significantly higher scores (Median = 4.41) compared to inanimate objects (Median = 3.59; W = 37971, <italic>p</italic> &#x003C; 0.001, <italic>r</italic> = 0.28), whereas for valence, no significant difference was discovered (W = 22555, <italic>p</italic> = 0.685, <italic>r</italic> = 0.02). Here, a substantial disagreement is evident, as mean heart rate [arousal: X<sup>2</sup>(2) = 0.62, <italic>p</italic> = 0.732, <italic>r</italic> = 0.02/valence: X<sup>2</sup>(2) = 5.40, <italic>p</italic> = 0.067, <italic>r</italic> = 0.05] and skin conductance response [arousal: M = X; X<sup>2</sup>(2) = 4.26, <italic>p</italic> = 0.119, <italic>r</italic> = 0.04/valence: M = X; X<sup>2</sup>(2) = 0.53, <italic>p</italic> = 0.768, <italic>r</italic> = 0.03] did not differ substantially between images of low, medium, and high arousal and valence ratings. However, available ratings from the OASIS dataset (arousal: Median = 4.06/valence: Median = 3.92) were more or less consistent with our ratings.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption><p>Human observers&#x2019; arousal and valence ratings across categories. <bold>(A)</bold> Arousal ratings showed that animate categories were perceived as more arousing when compared to inanimate categories. <bold>(B)</bold> Valence ratings suggest category-specific effects (i.e., <italic>snake</italic>), as with a few exceptions most images were rated as rather positive. The dotted line represents neutral scores.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-750639-g010.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="S4">
<title>Discussion</title>
<p>In this proof-of-concept study, we investigated the similarity of information processing in human observers and feedforward DCNN models during object recognition. For this purpose, human eye tracking heatmaps were compared to saliency maps of GradCAM, a customary attribution technique. Most importantly, during this endeavor, we found that GradCAM outputs, unlike eye tracking heatmaps, are produced on a discrete scale. While this is clear given the construction of DCNNs, to the best of our knowledge, this phenomenon has neither been regarded in previous studies nor stated explicitly in the literature of the field. As a natural consequence, this finding constrains several established results, as for example by <xref ref-type="bibr" rid="B16">Ebrahimpour et al. (2019)</xref>, and also our own heatmap comparisons, as different resolutions underlying these visualizations pose an evident challenge. Therefore, it seems necessary to draw attention to this fact, as it might endanger the fairness of human-machine comparisons (<xref ref-type="bibr" rid="B17">Firestone, 2020</xref>; <xref ref-type="bibr" rid="B18">Funke et al., 2020</xref>). We even believe to find more evidence for this problem as displayed in <xref ref-type="fig" rid="F9">Figure 9</xref>, where especially in earlier layers MAE values seemed to hit a lower boundary which was possibly due to fundamental mismatches in resolution and therefore impossible to undercut. Nevertheless, final comparisons of fixations and maxima were not influenced by this effect, as we controlled for spatial inaccuracy by proceeding with analyses on the target block level of the lowest maxima grid.</p>
<p>Generally, our results corroborate the assumption that the novel vNet architecture by <xref ref-type="bibr" rid="B38">Mehrer et al. (2021)</xref> captures human object recognition behavior more accurate compared to standard DCNNs commonly used throughout computer vision problems. Interestingly, this seemed to be the case on a performance level, as only vNet&#x2019;s performance was significantly correlated to human performance, and on a functional level of both before 150 ms, as it matched target blocks of human fixations spatially more consistent, and after 150 ms, where it yielded a lower MAE in late layers. On top of that, we argue that this higher similarity was not a side effect of higher performance compared to the equally trained ResNet18, as the fine-tuned ResNet18 model even outperformed human observers significantly and yet agreed the least with human fixations during feedforward processing. As already stated in previous literature (i.e., <xref ref-type="bibr" rid="B15">Dodge and Karam, 2017</xref>), the covariation in performance of human observers and DCNNs was found to be rather small. Nevertheless, these findings should be examined in the light of their relative impact, as vNet&#x2019;s correlation with human accuracy was not only of statistically significant importance, but also substantially higher. However, the reported effects should be treated with caution, as human-machine comparisons are prone to a wide range of confounding factors, either related to human cognition, such as the reported top-down and bottom-up processes, or related to deep learning problems, such as training settings and learning algorithms. We are aware of the fact that the number of parameters of both ResNet18 models and vNet is of necessity substantially different. In our view, these results emphasize the validity of a RFS modification as a method for designing more human-like models in computer vision. Moreover, as demonstrated by <xref ref-type="bibr" rid="B36">Luo et al. (2016)</xref>, effective RFS follows a Gaussian distribution and become heavily increased by deep learning techniques such as subsampling (in most cases average or max pooling) and dilated convolutions, which are all commonly used in current architectures. Hence, as in contrast to DCNN saliency maps, human viewing behavior is highly focal, it would be interesting to see if the match of spatial priorities can be further increased in models that lack these computations. At the same time, future studies should target this topic by comparing DCNNs that only differ in RFS along their hierarchical architecture.</p>
<p>Moreover, we were able to identify control and challenge images based on the notion of <xref ref-type="bibr" rid="B31">Kar et al. (2019)</xref>. Our analyses suggested no effect of image difficulty on the similarity between eye tracking heatmaps and saliency maps. While the authors&#x2019; original findings on neural activity show that control and challenge images are processed differently especially during late time periods (&#x003E;150 ms), their results also propose that the two image groups share a similar early response and may not be treated differently by the visual system via the retina at all. Our reported null result regarding the viewing behavior may complement this line of argument well, as the authors even mentioned that on visual inspection no specific image properties differed between the groups. Unfortunately, in this case, the allocation to early- and late-solved images through neural recordings was not possible in the experimental setup at hand.</p>
<p>Surprisingly, in contradiction with earlier findings in humans (<xref ref-type="bibr" rid="B39">New et al., 2007</xref>), our results indicated that both human observers and DCNNs were more accurate in recognizing inanimate compared animate objects. This advantage for inanimate objects has been also reported on the level of basic categories by other studies before (<xref ref-type="bibr" rid="B43">Pra&#x00DF; et al., 2013</xref>). Meanwhile, we discovered substantially more agreement between human and DCNN target blocks for animate objects. In turn, this suggests that both human observers and DCNNs were prioritizing more similar features during object recognition of animate objects. Therefore, we believe that this effect may be due to high efficient face processing mechanisms (<xref ref-type="bibr" rid="B11">Crouzet et al., 2010</xref>), which could have led to more similarity in specific <italic>face-heavy</italic> categories (such as <italic>human</italic> and <italic>dog</italic>). Furthermore, as animacy could not fully explain arousal and valence effects on a behavioral and functional level of DCNNs, we argue that these effects could be a consequence of naturally learned optimization effects in visual systems, which have also been reported for image memorability judgments that automatically develop in DCNNs trained on object recognition and even predict variation in neural spiking activity (<xref ref-type="bibr" rid="B29">Jaegle et al., 2019</xref>; <xref ref-type="bibr" rid="B50">Rust and Mehrpour, 2020</xref>). This interpretation, however, needs to be treated with caution, as arousal and valence scores were obtained by human ratings which again underlie a wide range of effects such as for example acquired knowledge, attention, and memorability.</p>
<p>To summarize, in this paper we outline a novel concept of comparing human and computer vision during object recognition. In theory, this approach seems suitable for evaluating similarities and differences in priorities of information processing and may help to further pinpoint the specific impact of model adjustments toward more biological plausibility. We demonstrate this by showing that a RFS modification, which agrees conceptually with the ventral visual pathway, increases the model fit to human viewing behavior. Practically, we believe that our method will be improved by including different attribution techniques such as <italic>Occlusion Sensitivity</italic> (<xref ref-type="bibr" rid="B61">Zeiler and Fergus, 2014</xref>), which estimates class activations through a combination of occlusions and classifications, and thereby allow more adequate and comparable resolutions in the future. Furthermore, we hope that our idea will open up new perspectives on comparative vision at the intersection between biological and computer vision research.</p>
</sec>
<sec sec-type="data-availability" id="S5">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="S6">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by the University of Salzburg Ethics Committee. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="S7">
<title>Author Contributions</title>
<p>LD, SD, and WG contributed to the conception and design of the study. LD programmed the experiment and wrote the first draft of the manuscript. LD and SD collected the eye tracking data. RK implemented the neural network architectures and contributed their data and wrote sections of the manuscript. LD and WG analyzed and interpreted the data. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="S8">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="S9">
<title>Funding</title>
<p>An Open Access Publication Fee was granted by the Center for Cognitive Neuroscience, University of Salzburg.</p>
</sec>
<ack>
<p>We thank Michael Christian Leitner and Stefan Hawelka for their support regarding the eye tracking measurements, as well as Frank Wilhelm and Michael Liedlgruber for sharing their experience regarding the emotional stimuli and physiological measurements.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alcorn</surname> <given-names>M. A.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Gong</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Mai</surname> <given-names>L.</given-names></name> <name><surname>Ku</surname> <given-names>W.-S.</given-names></name><etal/></person-group> (<year>2019</year>). &#x201C;<article-title>Strike (with) a pose: neural networks are easily fooled by strange poses of familiar objects</article-title>,&#x201D; in <source><italic>Paper presented at the Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Long Beach, CA</publisher-loc>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00498</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bar</surname> <given-names>M.</given-names></name></person-group> (<year>2003</year>). <article-title>A cortical mechanism for triggering top-down facilitation in visual object recognition.</article-title> <source><italic>J. Cogn. Neurosci.</italic></source> <volume>15</volume> <fpage>600</fpage>&#x2013;<lpage>609</lpage>. <pub-id pub-id-type="doi">10.1162/089892903321662976</pub-id> <pub-id pub-id-type="pmid">12803970</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bar</surname> <given-names>M.</given-names></name> <name><surname>Kassam</surname> <given-names>K. S.</given-names></name> <name><surname>Ghuman</surname> <given-names>A. S.</given-names></name> <name><surname>Boshyan</surname> <given-names>J.</given-names></name> <name><surname>Schmid</surname> <given-names>A. M.</given-names></name> <name><surname>Dale</surname> <given-names>A. M.</given-names></name><etal/></person-group> (<year>2006</year>). <article-title>Top-down facilitation of visual recognition.</article-title> <source><italic>Proc. Natl Acad. Sci. U.S.A.</italic></source> <volume>103</volume>:<issue>449</issue>. <pub-id pub-id-type="doi">10.1073/pnas.0507062103</pub-id> <pub-id pub-id-type="pmid">16407167</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beery</surname> <given-names>S.</given-names></name> <name><surname>Van Horn</surname> <given-names>G.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Recognition in Terra Incognita</article-title>,&#x201D; in <source><italic>Paper presented at the Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>, <publisher-loc>Munich</publisher-loc>. <pub-id pub-id-type="doi">10.1007/978-3-030-01270-0_28</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blanchette</surname> <given-names>I.</given-names></name></person-group> (<year>2006</year>). <article-title>Snakes, spiders, guns, and syringes: how specific are evolutionary constraints on the detection of threatening stimuli?</article-title> <source><italic>Q. J. Exp. Psychol.</italic></source> <volume>59</volume> <fpage>1484</fpage>&#x2013;<lpage>1504</lpage>. <pub-id pub-id-type="doi">10.1080/02724980543000204</pub-id> <pub-id pub-id-type="pmid">16846972</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blechert</surname> <given-names>J.</given-names></name> <name><surname>Peyk</surname> <given-names>P.</given-names></name> <name><surname>Liedlgruber</surname> <given-names>M.</given-names></name> <name><surname>Wilhelm</surname> <given-names>F. H.</given-names></name></person-group> (<year>2016</year>). <article-title>ANSLAB: integrated multichannel peripheral biosignal processing in psychophysiological science.</article-title> <source><italic>Behav. Res. Methods</italic></source> <volume>48</volume> <fpage>1528</fpage>&#x2013;<lpage>1545</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-015-0665-1</pub-id> <pub-id pub-id-type="pmid">26511369</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cadieu</surname> <given-names>C. F.</given-names></name> <name><surname>Hong</surname> <given-names>H.</given-names></name> <name><surname>Yamins</surname> <given-names>D. L. K.</given-names></name> <name><surname>Pinto</surname> <given-names>N.</given-names></name> <name><surname>Ardila</surname> <given-names>D.</given-names></name> <name><surname>Solomon</surname> <given-names>E. A.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Deep neural networks rival the representation of primate IT cortex for core visual object recognition.</article-title> <source><italic>PLoS Comput. Biol.</italic></source> <volume>10</volume>:<issue>e1003963</issue>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003963</pub-id> <pub-id pub-id-type="pmid">25521294</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cauchoix</surname> <given-names>M.</given-names></name> <name><surname>Crouzet</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>How plausible is a subcortical account of rapid visual recognition?</article-title> <source><italic>Front. Hum. Neurosci.</italic></source> <volume>7</volume>:<issue>39</issue>. <pub-id pub-id-type="doi">10.3389/fnhum.2013.00039</pub-id> <pub-id pub-id-type="pmid">23450981</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cichy</surname> <given-names>R. M.</given-names></name> <name><surname>Pantazis</surname> <given-names>D.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Resolving human object recognition in space and time.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>17</volume> <fpage>455</fpage>&#x2013;<lpage>462</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3635</pub-id> <pub-id pub-id-type="pmid">24464044</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Contini</surname> <given-names>E. W.</given-names></name> <name><surname>Wardle</surname> <given-names>S. G.</given-names></name> <name><surname>Carlson</surname> <given-names>T. A.</given-names></name></person-group> (<year>2017</year>). <article-title>Decoding the time-course of object recognition in the human brain: from visual features to categorical decisions.</article-title> <source><italic>Neuropsychologia</italic></source> <volume>105</volume> <fpage>165</fpage>&#x2013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuropsychologia.2017.02.013</pub-id> <pub-id pub-id-type="pmid">28215698</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crouzet</surname> <given-names>S. M.</given-names></name> <name><surname>Kirchner</surname> <given-names>H.</given-names></name> <name><surname>Thorpe</surname> <given-names>S. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Fast saccades toward faces: face detection in just 100 ms.</article-title> <source><italic>J. Vis.</italic></source> <volume>10</volume> <fpage>16</fpage>&#x2013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1167/10.4.16</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Kai</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>F.-F.</given-names></name></person-group> (<year>2009</year>). &#x201C;<article-title>ImageNet: a large-scale hierarchical image database</article-title>,&#x201D; in <source><italic>Paper Presented at the 2009 IEEE Conference on Computer Vision and Pattern Recognition</italic></source>, <publisher-loc>Miami, FL</publisher-loc>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206848</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name> <name><surname>Cox</surname> <given-names>D. D.</given-names></name></person-group> (<year>2007</year>). <article-title>Untangling invariant object recognition.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>11</volume> <fpage>333</fpage>&#x2013;<lpage>341</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2007.06.010</pub-id> <pub-id pub-id-type="pmid">17631409</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name> <name><surname>Zoccolan</surname> <given-names>D.</given-names></name> <name><surname>Rust</surname></name> <name><surname>Nicole</surname> <given-names>C.</given-names></name></person-group> (<year>2012</year>). <article-title>How does the brain solve visual object recognition?</article-title> <source><italic>Neuron</italic></source> <volume>73</volume> <fpage>415</fpage>&#x2013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.01.010</pub-id> <pub-id pub-id-type="pmid">22325196</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dodge</surname> <given-names>S.</given-names></name> <name><surname>Karam</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>A study and comparison of human and deep learning recognition performance under visual distortions</article-title>,&#x201D; in <source><italic>Paper Presented at the 26th International Conference on Computer Communication and Networks (ICCCN)</italic></source>, <publisher-loc>Vancouver, BC</publisher-loc>.</citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ebrahimpour</surname> <given-names>M. K.</given-names></name> <name><surname>Falandays</surname> <given-names>J. B.</given-names></name> <name><surname>Spevack</surname> <given-names>S.</given-names></name> <name><surname>Noelle</surname> <given-names>D. C.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Do humans look where deep convolutional neural networks &#x201C;attend&#x201D;?</article-title>,&#x201D; in <source><italic>Paper Presented at the Advances in Visual Computing</italic></source>, <publisher-loc>Cham</publisher-loc>. <pub-id pub-id-type="doi">10.1007/978-3-030-33723-0_5</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Firestone</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>Performance vs. competence in human&#x2013;machine comparisons.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>117</volume> <issue>26562</issue>. <pub-id pub-id-type="doi">10.1073/pnas.1905334117</pub-id> <pub-id pub-id-type="pmid">33051296</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funke</surname> <given-names>C. M.</given-names></name> <name><surname>Borowski</surname> <given-names>J.</given-names></name> <name><surname>Stosio</surname> <given-names>K.</given-names></name> <name><surname>Brendel</surname> <given-names>W.</given-names></name> <name><surname>Wallis</surname> <given-names>T. S.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>The notorious difficulty of comparing human and machine perception.</article-title> <source><italic>arXiv</italic> [Preprint]</source> <comment>arXiv: 2004.09406</comment>, <pub-id pub-id-type="doi">10.32470/CCN.2019.1295-0</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Geirhos</surname> <given-names>R.</given-names></name> <name><surname>Jacobsen</surname> <given-names>J.-H.</given-names></name> <name><surname>Michaelis</surname> <given-names>C.</given-names></name> <name><surname>Zemel</surname> <given-names>R.</given-names></name> <name><surname>Brendel</surname> <given-names>W.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Shortcut learning in deep neural networks.</article-title> <source><italic>arXiv</italic> [Preprint]</source> <comment>arXiv:1312.6199</comment>, <pub-id pub-id-type="doi">10.1038/s42256-020-00257-z</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Geirhos</surname> <given-names>R.</given-names></name> <name><surname>Janssen</surname> <given-names>D. H.</given-names></name> <name><surname>Sch&#x00FC;tt</surname> <given-names>H. H.</given-names></name> <name><surname>Rauber</surname> <given-names>J.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name> <name><surname>Wichmann</surname> <given-names>F. A.</given-names></name></person-group> (<year>2017</year>). <article-title>Comparing deep neural networks against humans: object recognition when the signal gets weaker.</article-title> <source><italic>arXiv</italic> [Preprint]</source> <comment>arXiv:1706.06 969</comment>,</citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Geirhos</surname> <given-names>R.</given-names></name> <name><surname>Temme</surname> <given-names>C. R.</given-names></name> <name><surname>Rauber</surname> <given-names>J.</given-names></name> <name><surname>Sch&#x00FC;tt</surname> <given-names>H. H.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name> <name><surname>Wichmann</surname> <given-names>F. A.</given-names></name></person-group> (<year>2018b</year>). &#x201C;<article-title>Generalisation in humans and deep neural networks</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the 32nd Conference on Neural Information Processing Systems (NIPS)</italic></source>, <publisher-loc>Montr&#x00E9;al, QC</publisher-loc>.</citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Geirhos</surname> <given-names>R.</given-names></name> <name><surname>Rubisch</surname> <given-names>P.</given-names></name> <name><surname>Michaelis</surname> <given-names>C.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name> <name><surname>Wichmann</surname> <given-names>F. A.</given-names></name> <name><surname>Brendel</surname> <given-names>W.</given-names></name></person-group> (<year>2018a</year>). <article-title>ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.</article-title> <source><italic>arXiv</italic> [Preprint]</source> <comment>arXiv:1811.12231</comment>,</citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Greene</surname> <given-names>M. R.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Recognition of natural scenes from global properties: seeing the forest without representing the trees.</article-title> <source><italic>Cogn. Psychol.</italic></source> <volume>58</volume> <fpage>137</fpage>&#x2013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1016/j.cogpsych.2008.06.001</pub-id> <pub-id pub-id-type="pmid">18762289</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grill-Spector</surname> <given-names>K.</given-names></name> <name><surname>Weiner</surname> <given-names>K. S.</given-names></name> <name><surname>Kay</surname> <given-names>K.</given-names></name> <name><surname>Gomez</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>The functional neuroanatomy of human face perception.</article-title> <source><italic>Annu. Rev. Vis. Sci.</italic></source> <volume>3</volume> <fpage>167</fpage>&#x2013;<lpage>196</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-vision-102016-061214</pub-id> <pub-id pub-id-type="pmid">28715955</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Delving deep into rectifiers: Surpassing human-level performance on imagenet classification</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the IEEE International Conference on Computer Vision (ICCV)</italic></source>, <publisher-loc>Santiago</publisher-loc>. <pub-id pub-id-type="doi">10.1109/ICCV.2015.123</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Deep residual learning for image recognition</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Las Vegas, NV</publisher-loc>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Van Der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Densely connected convolutional networks</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</italic></source>, <publisher-loc>Honolulu, HI</publisher-loc>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.243</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ishai</surname> <given-names>A.</given-names></name> <name><surname>Ungerleider</surname> <given-names>L. G.</given-names></name> <name><surname>Martin</surname> <given-names>A.</given-names></name> <name><surname>Schouten</surname> <given-names>J. L.</given-names></name> <name><surname>Haxby</surname> <given-names>J. V.</given-names></name></person-group> (<year>1999</year>). <article-title>Distributed representation of objects in the human ventral visual pathway.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>96</volume>:<issue>9379</issue>. <pub-id pub-id-type="doi">10.1073/pnas.96.16.9379</pub-id> <pub-id pub-id-type="pmid">10430951</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaegle</surname> <given-names>A.</given-names></name> <name><surname>Mehrpour</surname> <given-names>V.</given-names></name> <name><surname>Mohsenzadeh</surname> <given-names>Y.</given-names></name> <name><surname>Meyer</surname> <given-names>T.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Rust</surname> <given-names>N.</given-names></name></person-group> (<year>2019</year>). <article-title>Population response magnitude variation in inferotemporal cortex predicts image memorability.</article-title> <source><italic>eLife</italic></source> <volume>8</volume>:<issue>e47596</issue>. <pub-id pub-id-type="doi">10.7554/eLife.47596</pub-id> <pub-id pub-id-type="pmid">31464687</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kar</surname> <given-names>K.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2020</year>). <article-title>Fast recurrent processing via ventral prefrontal cortex is needed by the primate ventral stream for robust core visual object recognition.</article-title> <source><italic>bioRxiv</italic>[Preprint]</source> <pub-id pub-id-type="doi">10.1101/2020.05.10.086959</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kar</surname> <given-names>K.</given-names></name> <name><surname>Kubilius</surname> <given-names>J.</given-names></name> <name><surname>Schmidt</surname> <given-names>K.</given-names></name> <name><surname>Issa</surname> <given-names>E. B.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2019</year>). <article-title>Evidence that recurrent circuits are critical to the ventral stream&#x2019;s execution of core object recognition behavior.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>22</volume> <fpage>974</fpage>&#x2013;<lpage>983</lpage>. <pub-id pub-id-type="doi">10.1038/s41593-019-0392-5</pub-id> <pub-id pub-id-type="pmid">31036945</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). &#x201C;<article-title>Imagenet classification with deep convolutional neural networks</article-title>,&#x201D; in <source><italic>Paper Presented at the Advances in Neural Information Processing Systems</italic></source>, <publisher-loc>Lake Tahoe, NV</publisher-loc>.</citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kurdi</surname> <given-names>B.</given-names></name> <name><surname>Lozano</surname> <given-names>S.</given-names></name> <name><surname>Banaji</surname> <given-names>M. R.</given-names></name></person-group> (<year>2017</year>). <article-title>Introducing the open affective standardized image set (OASIS).</article-title> <source><italic>Behav. Res. Methods</italic></source> <volume>49</volume> <fpage>457</fpage>&#x2013;<lpage>470</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-016-0715-3</pub-id> <pub-id pub-id-type="pmid">26907748</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lamme</surname> <given-names>V. A.</given-names></name> <name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name></person-group> (<year>2000</year>). <article-title>The distinct modes of vision offered by feedforward and recurrent processing.</article-title> <source><italic>Trends Neurosci.</italic></source> <volume>23</volume> <fpage>571</fpage>&#x2013;<lpage>579</lpage>. <pub-id pub-id-type="doi">10.1016/s0166-2236(00)01657-x</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Landau</surname> <given-names>B.</given-names></name> <name><surname>Smith</surname> <given-names>L. B.</given-names></name> <name><surname>Jones</surname> <given-names>S. S.</given-names></name></person-group> (<year>1988</year>). <article-title>The importance of shape in early lexical learning.</article-title> <source><italic>Cogn. Dev.</italic></source> <volume>3</volume> <fpage>299</fpage>&#x2013;<lpage>321</lpage>. <pub-id pub-id-type="doi">10.1016/0885-2014(88)90014-7</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Urtasun</surname> <given-names>R.</given-names></name> <name><surname>Zemel</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Understanding the effective receptive field in deep convolutional neural networks</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the 30th International Conference on Neural Information Processing Systems</italic></source>, <publisher-loc>Barcelona</publisher-loc>.</citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name></person-group> (<year>1982</year>). <source><italic>Vision: A Computational Investigation Into the Human Representation and Processing of Visual Information.</italic></source> <publisher-loc>San Francisco, CA</publisher-loc>: <publisher-name>W. H. Freeman</publisher-name>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehrer</surname> <given-names>J.</given-names></name> <name><surname>Spoerer</surname> <given-names>C. J.</given-names></name> <name><surname>Jones</surname> <given-names>E. C.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Kietzmann</surname> <given-names>T. C.</given-names></name></person-group> (<year>2021</year>). <article-title>An ecologically motivated image dataset for deep learning yields better models of human vision.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>118</volume>:<issue>e2011417118</issue>. <pub-id pub-id-type="doi">10.1073/pnas.2011417118</pub-id> <pub-id pub-id-type="pmid">33593900</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>New</surname> <given-names>J.</given-names></name> <name><surname>Cosmides</surname> <given-names>L.</given-names></name> <name><surname>Tooby</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Category-specific attention for animals reflects ancestral priorities, not expertise.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>104</volume>:<issue>16598</issue>. <pub-id pub-id-type="doi">10.1073/pnas.0703913104</pub-id> <pub-id pub-id-type="pmid">17909181</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x00D6;hman</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>The role of the amygdala in human fear: automatic detection of threat.</article-title> <source><italic>Psychoneuroendocrinology</italic></source> <volume>30</volume> <fpage>953</fpage>&#x2013;<lpage>958</lpage>. <pub-id pub-id-type="doi">10.1016/j.psyneuen.2005.03.019</pub-id> <pub-id pub-id-type="pmid">15963650</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2007</year>). <article-title>The role of context in object recognition.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>11</volume> <fpage>520</fpage>&#x2013;<lpage>527</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2007.09.009</pub-id> <pub-id pub-id-type="pmid">18024143</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pessoa</surname> <given-names>L.</given-names></name> <name><surname>Adolphs</surname> <given-names>R.</given-names></name></person-group> (<year>2010</year>). <article-title>Emotion processing and the amygdala: from a &#x2018;low road&#x2019; to &#x2018;many roads&#x2019; of evaluating biological significance.</article-title> <source><italic>Nat. Rev. Neurosci.</italic></source> <volume>11</volume> <fpage>773</fpage>&#x2013;<lpage>782</lpage>. <pub-id pub-id-type="doi">10.1038/nrn2920</pub-id> <pub-id pub-id-type="pmid">20959860</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pra&#x00DF;</surname> <given-names>M.</given-names></name> <name><surname>Grimsen</surname> <given-names>C.</given-names></name> <name><surname>K&#x00F6;nig</surname> <given-names>M.</given-names></name> <name><surname>Fahle</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Ultra rapid object categorization: effects of level, animacy and context.</article-title> <source><italic>PLoS One</italic></source> <volume>8</volume>:<issue>e68051</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0068051</pub-id> <pub-id pub-id-type="pmid">23840810</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rajaei</surname> <given-names>K.</given-names></name> <name><surname>Mohsenzadeh</surname> <given-names>Y.</given-names></name> <name><surname>Ebrahimpour</surname> <given-names>R.</given-names></name> <name><surname>Khaligh-Razavi</surname> <given-names>S.-M.</given-names></name></person-group> (<year>2019</year>). <article-title>Beyond core object recognition: recurrent processes account for object recognition under occlusion.</article-title> <source><italic>PLoS Comput. Biol.</italic></source> <volume>15</volume>:<issue>e1007001</issue>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1007001</pub-id> <pub-id pub-id-type="pmid">31091234</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riesenhuber</surname> <given-names>M.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>1999</year>). <article-title>Hierarchical models of object recognition in cortex.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>2</volume> <fpage>1019</fpage>&#x2013;<lpage>1025</lpage>. <pub-id pub-id-type="doi">10.1038/14819</pub-id> <pub-id pub-id-type="pmid">10526343</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rolls</surname> <given-names>E. T.</given-names></name></person-group> (<year>2000</year>). <article-title>Functions of the primate temporal lobe cortical visual areas in invariant visual object and face recognition.</article-title> <source><italic>Neuron</italic></source> <volume>27</volume> <fpage>205</fpage>&#x2013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1016/s0896-6273(00)00030-1</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosch</surname> <given-names>E.</given-names></name></person-group> (<year>1999</year>). <article-title>Principles of categorization.</article-title> <source><italic>Concepts</italic></source> <volume>189</volume> <fpage>312</fpage>&#x2013;<lpage>322</lpage>. <pub-id pub-id-type="doi">10.1016/B978-1-4832-1446-7.50028-5</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rothkegel</surname> <given-names>L. O. M.</given-names></name> <name><surname>Trukenbrod</surname> <given-names>H. A.</given-names></name> <name><surname>Sch&#x00FC;tt</surname> <given-names>H. H.</given-names></name> <name><surname>Wichmann</surname> <given-names>F. A.</given-names></name> <name><surname>Engbert</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Temporal evolution of the central fixation bias in scene viewing.</article-title> <source><italic>J. Vis.</italic></source> <volume>17</volume>:<issue>3</issue>. <pub-id pub-id-type="doi">10.1167/17.13.3</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russakovsky</surname> <given-names>O.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Krause</surname> <given-names>J.</given-names></name> <name><surname>Satheesh</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>ImageNet large scale visual recognition challenge.</article-title> <source><italic>Int. J. Comput. Vis.</italic></source> <volume>115</volume> <fpage>211</fpage>&#x2013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-015-0816-y</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rust</surname> <given-names>N. C.</given-names></name> <name><surname>Mehrpour</surname> <given-names>V.</given-names></name></person-group> (<year>2020</year>). <article-title>Understanding image memorability.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>24</volume> <fpage>557</fpage>&#x2013;<lpage>568</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2020.04.001</pub-id> <pub-id pub-id-type="pmid">32386889</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rutishauser</surname> <given-names>U.</given-names></name> <name><surname>Walther</surname> <given-names>D.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2004</year>). &#x201C;<article-title>Is bottom-up attention useful for object recognition?</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004</italic></source>, <publisher-loc>Washington, DC</publisher-loc>.</citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seijdel</surname> <given-names>N.</given-names></name> <name><surname>Loke</surname> <given-names>J.</given-names></name> <name><surname>van de Klundert</surname> <given-names>R.</given-names></name> <name><surname>van der Meer</surname> <given-names>M.</given-names></name> <name><surname>Quispel</surname> <given-names>E.</given-names></name> <name><surname>van Gaal</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>On the necessity of recurrent processing during object recognition: it depends on the need for scene segmentation.</article-title> <source><italic>bioRxiv</italic> [Preprint]</source> <pub-id pub-id-type="doi">10.1101/2020.11.11.377655</pub-id> <comment>bioRxiv: 2020.2011.2011.37 7655</comment>,</citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Selvaraju</surname> <given-names>R. R.</given-names></name> <name><surname>Cogswell</surname> <given-names>M.</given-names></name> <name><surname>Das</surname> <given-names>A.</given-names></name> <name><surname>Vedantam</surname> <given-names>R.</given-names></name> <name><surname>Parikh</surname> <given-names>D.</given-names></name> <name><surname>Batra</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Grad-cam: Visual explanations from deep networks via gradient-based localization</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the IEEE International Conference on Computer Vision</italic></source>, <publisher-loc>Venice</publisher-loc>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.74</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Sermanet</surname> <given-names>P.</given-names></name> <name><surname>Reed</surname> <given-names>S.</given-names></name> <name><surname>Anguelov</surname> <given-names>D.</given-names></name><etal/></person-group> (<year>2015</year>). &#x201C;<article-title>Going deeper with convolutions</article-title>,&#x201D; in <source><italic>Paper Presented at the Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Boston, MA</publisher-loc>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298594</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanaka</surname> <given-names>K.</given-names></name></person-group> (<year>1996</year>). <article-title>Inferotemporal cortex and object vision.</article-title> <source><italic>Annu. Rev. Neurosci.</italic></source> <volume>19</volume> <fpage>109</fpage>&#x2013;<lpage>139</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.ne.19.030196.000545</pub-id> <pub-id pub-id-type="pmid">8833438</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>H.</given-names></name> <name><surname>Schrimpf</surname> <given-names>M.</given-names></name> <name><surname>Lotter</surname> <given-names>W.</given-names></name> <name><surname>Moerman</surname> <given-names>C.</given-names></name> <name><surname>Paredes</surname> <given-names>A.</given-names></name> <name><surname>Ortega Caro</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2018</year>). <article-title>Recurrent computations for visual pattern completion.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>115</volume> <fpage>8835</fpage>&#x2013;<lpage>8840</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1719397115</pub-id> <pub-id pub-id-type="pmid">30104363</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tatler</surname> <given-names>B. W.</given-names></name></person-group> (<year>2007</year>). <article-title>The central fixation bias in scene viewing: selecting an optimal viewing position independently of motor biases and image feature distributions.</article-title> <source><italic>J. Vis.</italic></source> <volume>7</volume> <fpage>4.1</fpage>&#x2013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1167/7.14.4</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tatler</surname> <given-names>B. W.</given-names></name> <name><surname>Baddeley</surname> <given-names>R. J.</given-names></name> <name><surname>Vincent</surname> <given-names>B. T.</given-names></name></person-group> (<year>2006</year>). <article-title>The long and the short of it: spatial statistics at fixation vary with saccade amplitude and task.</article-title> <source><italic>Vis. Res.</italic></source> <volume>46</volume> <fpage>1857</fpage>&#x2013;<lpage>1862</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2005.12.005</pub-id> <pub-id pub-id-type="pmid">16469349</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Dyck</surname> <given-names>L. E.</given-names></name> <name><surname>Gruber</surname> <given-names>W. R.</given-names></name></person-group> (<year>2020</year>). <article-title>Seeing eye-to-eye? A comparison of object recognition performance in humans and deep convolutional neural networks under image manipulation.</article-title> <source><italic>arXiv</italic> [Preprint]</source> <comment>arXiv: 2007.06294.</comment>,</citation></ref>
<ref id="B60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wandell</surname> <given-names>B. A.</given-names></name> <name><surname>Winawer</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Computational neuroimaging and population receptive fields.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>19</volume> <fpage>349</fpage>&#x2013;<lpage>357</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2015.03.009</pub-id> <pub-id pub-id-type="pmid">25850730</pub-id></citation></ref>
<ref id="B61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeiler</surname> <given-names>M. D.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). &#x201C;<article-title>Visualizing and understanding convolutional networks</article-title>,&#x201D; in <source><italic>Paper presented at the European Conference on Computer Vision</italic></source>, (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>). <pub-id pub-id-type="doi">10.1007/978-3-319-10590-1_53</pub-id></citation></ref>
</ref-list>
</back>
</article>
