<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2022.760085</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The Data Efficiency of Deep Learning Is Degraded by Unnecessary Input Dimensions</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>D&#x00027;Amario</surname> <given-names>Vanessa</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1439483/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Srivastava</surname> <given-names>Sanjana</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Sasaki</surname> <given-names>Tomotake</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1441087/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Boix</surname> <given-names>Xavier</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1439137/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology</institution>, <addr-line>Cambridge, MA</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Center for Brains, Minds and Machines</institution>, <addr-line>Cambridge, MA</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Computer Science, Stanford University</institution>, <addr-line>Stanford, CA</addr-line>, <country>United States</country></aff>
<aff id="aff4"><sup>4</sup><institution>Artificial Intelligence Laboratory, Fujitsu Limited</institution>, <addr-line>Kawasaki</addr-line>, <country>Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Omri Barak, Technion Israel Institute of Technology, Israel</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Wei Lin, Fudan University, China; Mariofanna Milanova, University of Arkansas at Little Rock, United States; Jeremy Bernstein, California Institute of Technology, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Vanessa D&#x00027;Amario <email>vanessad&#x00040;mit.edu</email></corresp>
<corresp id="c002">Xavier Boix <email>xboix&#x00040;mit.edu</email></corresp>
<fn fn-type="present-address" id="fn002"><p>&#x02020;Present address: Vanessa D&#x00027;Amario, Fujitsu Research of America, Inc., Sunnyvale, CA, United States</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>31</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>760085</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>01</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 D&#x00027;Amario, Srivastava, Sasaki and Boix.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>D&#x00027;Amario, Srivastava, Sasaki and Boix</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Biological learning systems are outstanding in their ability to learn from limited training data compared to the most successful learning machines, <italic>i.e.</italic>, Deep Neural Networks (DNNs). What are the key aspects that underlie this data efficiency gap is an unresolved question at the core of biological and artificial intelligence. We hypothesize that one important aspect is that biological systems rely on mechanisms such as foveations in order to reduce unnecessary input dimensions for the task at hand, <italic>e.g.</italic>, background in object recognition, while state-of-the-art DNNs do not. Datasets to train DNNs often contain such unnecessary input dimensions, and these lead to more trainable parameters. Yet, it is not clear whether this affects the DNNs&#x00027; data efficiency because DNNs are robust to increasing the number of parameters in the hidden layers, and it is uncertain whether this holds true for the input layer. In this paper, we investigate the impact of unnecessary input dimensions on the DNNs data efficiency, namely, the amount of examples needed to achieve certain generalization performance. Our results show that unnecessary input dimensions that are task-unrelated substantially degrade data efficiency. This highlights the need for mechanisms that remove task-unrelated dimensions, such as foveation for image classification, in order to enable data efficiency gains.</p></abstract>
<kwd-group>
<kwd>data efficiency</kwd>
<kwd>overparameterization</kwd>
<kwd>object recognition</kwd>
<kwd>object background</kwd>
<kwd>unnecessary input dimensions</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="32"/>
<page-count count="8"/>
<word-count count="5432"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The success of Deep Neural Networks (DNNs) contrasts with the still distant goal of learning with few training examples as in biological systems, <italic>i.e.</italic>, in a data efficient manner (Hassabis et al., <xref ref-type="bibr" rid="B14">2017</xref>). Understanding the principles that underlie such differential is a question at the core of both artificial and biological intelligence. In this paper, we introduce the hypothesis that an important aspect for data efficiency is that biological systems rely on mechanisms such as foveations in order to reduce unnecessary input dimensions, <italic>e.g.</italic>, background in object recognition, while state-of-the-art DNNs do not.</p>
<p>DNNs are usually trained on high dimensional datasets (<italic>e.g.</italic>, images and text), and many input dimensions of the DNN may be unnecessary to predict the ground-truth label as they are unrelated and/or redundant to the task at hand. Machine learning theory for linear and kernel methods predicts that unnecessary input dimensions may degrade the DNN&#x00027;s data efficiency (Hastie et al., <xref ref-type="bibr" rid="B15">2009</xref>), as the classifier may overfit to the unnecessary input dimensions if not enough training examples are provided to learn to discard them.</p>
<p>However, DNNs have challenged classic machine learning measures of complexity (<italic>e.g.</italic>, VC dimensions, Rademacher complexity) as they can achieve high test accuracy despite having a number of trainable parameters much larger than the number of training examples, <italic>i.e.</italic>, DNNs are overparameterized (Zhang et al., <xref ref-type="bibr" rid="B29">2017</xref>; Nakkiran et al., <xref ref-type="bibr" rid="B22">2020</xref>). Since unnecessary input dimensions lead to more overparameterization, it is unclear in what way DNNs suffer from unnecessary input dimensions and whether more data is needed to learn to discard them.</p>
<p>To foreshadow the results, we find that the DNNs&#x00027; data efficiency depends on whether the unnecessary dimensions are <italic>task-unrelated</italic> or <italic>task-related</italic> (redundant with respect to other input dimensions). Namely, increasing the number of <italic>task-unrelated</italic> dimensions leads to a substantial drop of data efficiency, while increasing the number of <italic>task-related</italic> dimensions that are linear combinations of other <italic>task-related</italic> dimensions, helps to alleviate the negative impact of the <italic>task-unrelated</italic> dimensions. These results suggest that mechanisms to discard unnecessary input dimensions, such as foveations for object recognition, are necessary to enable data efficiency gains.</p>
</sec>
<sec id="s2">
<title>2. Related Works</title>
<p>We now relate our work with the effect of background on the generalization abilities of DNNs in object recognition, and also with the DNNs generalization abilities depending on the number of parameters of the network.</p>
<sec>
<title>2.1. Object&#x00027;s Background and DNN Generalization</title>
<p>The data collection process is often biased (Torralba and Efros, <xref ref-type="bibr" rid="B26">2011</xref>). One of the most prominent factors of such dataset bias is the background, such that some aspects of the background systematically co-occur with certain objects, <italic>e.g.</italic>, airplanes may tend to always appear in the sky. This co-occurrence is a confounding factor for the network, and the network may learn to associate the background with the object, <italic>e.g.</italic>, the sky may be regarded as part of the airplane. Previous works have shown that DNNs for image recognition fail to classify objects in novel and uncommon backgrounds (Choi et al., <xref ref-type="bibr" rid="B9">2012</xref>; Volokitin et al., <xref ref-type="bibr" rid="B27">2017</xref>; Beery et al., <xref ref-type="bibr" rid="B5">2018</xref>; Tian et al., <xref ref-type="bibr" rid="B25">2018</xref>). Remarkably, popular object recognition datasets are biased to such an extent that DNNs can predict the object category even when the objects are removed from the image (Zhu et al., <xref ref-type="bibr" rid="B32">2017</xref>; Tian et al., <xref ref-type="bibr" rid="B25">2018</xref>; Xiao et al., <xref ref-type="bibr" rid="B28">2021</xref>). Barbu et al. (<xref ref-type="bibr" rid="B4">2019</xref>) introduced a new benchmark which addresses the biased co-occurrence of objects and background, among other types of bias. DNNs exhibit large performance drops in this benchmark compared to ImageNet (Deng et al., <xref ref-type="bibr" rid="B12">2009</xref>). Recently, Borji has shown that a large portion of the performance drop comes from the bias in object&#x00027;s background, as classifying the object in isolation substantially alleviates the performance drop (Borji, <xref ref-type="bibr" rid="B8">2021</xref>).</p>
<p>In contrast to previous works, we analyse the impact of the object&#x00027;s background to the DNN&#x00027;s generalization performance when the dataset is unbiased, <italic>i.e.</italic>, there is no significant correlation between the objects and backgrounds and the statistics of the object&#x00027;s background are the same between training and testing times. To the best of our knowledge, our work is the first to investigate the effects of object&#x00027;s background on DNNs when these are unbiased. We show that just the presence of background, even if it is unbiased, can degrade the data efficiency of the DNN.</p>
</sec>
<sec>
<title>2.2. Overparameterization and Data Dimensionality</title>
<p>A remarkable characteristic of DNNs is that the test error follows a double-descend when the DNN&#x00027;s width is increased by adding more hidden units. Thus, the test error decreases as the network&#x00027;s width is increased in both the underparameterized and overparameterized regimes, except in a critical region between these two where a substantial error increase can take place (Belkin et al., <xref ref-type="bibr" rid="B6">2019</xref>; Advani et al., <xref ref-type="bibr" rid="B1">2020</xref>; Nakkiran et al., <xref ref-type="bibr" rid="B22">2020</xref>). The overparameterized regime has received a lot of attention because DNNs with many more parameters than training examples can achieve high test accuracy, and a theoretical understanding of this phenomenon is an active area of research. Robustness to overparametrization relates to unnecessary input dimensions because unnecessary input dimensions also increase the number of parameters of the network, albeit in the input layer rather than in the intermediate layers. As we show in the sequel, increasing the number of unnecessary input dimensions can have the opposite effect of increasing the number of hidden units in the test error.</p>
<p>A theoretical understanding of this phenomenon using mathematical tools is an open question. The PAC Bayes theory appears as a promising approach to describe the generalization capacity of DNNs [<italic>e.g.</italic>, (Dziugaite and Roy, <xref ref-type="bibr" rid="B13">2017</xref>; De Palma et al., <xref ref-type="bibr" rid="B11">2019</xref>; Bernstein and Yue, <xref ref-type="bibr" rid="B7">2021</xref>)]. While these theoretical results provide insights about the trends of the behaviour of the DNN, an empirical, quantitative assessment of the effect of unnecessary dimensions to the DNN&#x00027;s data efficiency is missing. Our analysis derives from theoretical insights of the exact solution of a linear network in a regression task. In this way, we can relate and compare empirical results for DNNs with cases that are well understood theoretically.</p>
<p>Another strand of research relates the structure of the dataset with the generalization ability of the network. Several works in statistical learning theory for kernel machines relate the spectrum of the dataset with the generalization performance (Zhang, <xref ref-type="bibr" rid="B30">2005</xref>). For neural networks, Ansuini et al. (<xref ref-type="bibr" rid="B3">2019</xref>); Recanatesi et al. (<xref ref-type="bibr" rid="B23">2019</xref>) define the intrinsic dimensionality based on the dimension of the data manifold. These works analyze how the network reduces the intrinsic dimension across layers. Yet, these metrics based on manifolds do not provide insights about how specific aspects of the dataset, <italic>e.g.</italic>, unnecessary dimensions, contribute to the intrinsic dimensionality.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Unnecessary Input Dimensions and Data Efficiency</title>
<p>We aim at analyzing the effect of unnecessary input dimensions on the data efficiency of DNNs. Let <bold>x</bold> be a vector representing a data sample, and let <bold>y</bold> be the ground-truth label of <bold>x</bold>. We define <italic>f</italic>(<bold>x</bold>) &#x0003D; <bold>y</bold> as the target function of the learning problem. Also, we use [<bold>x</bold>; <bold>u</bold>] to denote the data sample <bold>x</bold> with unnecessary input dimensions appended to it. The unnecessary dimensions do not affect the target function of the learning problem, <italic>i.e.</italic>, <italic>g</italic>([<bold>x</bold>; <bold>u</bold>]) &#x0003D; <italic>f</italic>(<bold>x</bold>) &#x0003D; <bold>y</bold>, where <italic>g</italic> is the target function of the learning problem with unnecessary input dimensions. Each sample can have a different set of dimensions that are unnecessary, <italic>e.g.</italic>, one sample could be [<bold>x</bold><sub>1</sub>; <bold>u</bold><sub>1</sub>] and another be [<bold>u</bold><sub>2</sub>; <bold>x</bold><sub>2</sub>]. Note that this variability is present in object recognition because the dimensions representing the object&#x00027;s background are unnecessary and vary across data samples, as the object can be in different image locations.</p>
<p>We define two types of unnecessary input dimensions: <italic>task-unrelated</italic> and <italic>task-related</italic>. Unnecessary input dimensions are <italic>task-unrelated</italic> when they are independent of <bold>x</bold>, <italic>i.e.</italic>, they can not be predicted from <bold>x</bold>, as in unbiased object&#x00027;s background. Otherwise, the unnecessary dimensions are <italic>task-related</italic>, which are equivalent to redundant dimensions. An example that leads to more <italic>task-related</italic> unnecessary dimensions is upscaling the image.</p>
<p>To study the effect of unnecessary input dimensions, we measure the test accuracy of DNNs trained with different amounts of unnecessary input dimensions and training examples. Given a DNN architecture and a dataset with a fixed amount of unnecessary dimensions, we define the <italic>data efficiency</italic> of the DNN as the Area Under the Test Curve (AUTC) for the DNN trained with different number of training examples. The curve is monotonically increasing, as more training examples lead to higher test accuracy, and the AUTC measures the area under it. We normalize the AUTC to be between 0 and 1, where 1 is the maximum achievable, and it corresponds to 100% test accuracy for all number of training examples. In the experiments where the number of training examples spans several orders of magnitude, we calculate the AUTC by converting the number of training examples in logarithmic scale, such that all orders of magnitude are equally taken into account.</p>
</sec>
<sec id="s4">
<title>4. Datasets and Networks</title>
<p>We now introduce the datasets and networks we use in the experiments (refer to <xref ref-type="supplementary-material" rid="SM1">Appendices 1</xref>, <xref ref-type="supplementary-material" rid="SM1">2</xref> for additional details).</p>
<sec>
<title>4.1. Linearly Separable Dataset</title>
<p>We use a linearly separable dataset for binary classification, as it facilitates relating results of classic machine learning and DNNs. We generate a binary classification dataset of 30 input dimensions, which follow a Gaussian distribution with (&#x003BC; &#x0003D; 0, &#x003C3; &#x0003D; 1). The ground-truth label is the output of a linear classifier, such that the dataset is linearly separable with a hyperplane randomly chosen. Unnecessary input dimensions are appended to the data samples. <italic>Task-unrelated</italic> dimensions follow a Gaussian distribution with (&#x003BC; &#x0003D; 0, &#x003C3; &#x0003D; 0.1). The <italic>task-related</italic> dimensions are linear combinations of the dimensions of the original dataset samples.</p>
<p>We evaluate the following linear and Multi-Layer Perceptron (MLP) networks: linear network trained with square loss (pseudo-inverse solution), MLP with linear activation functions trained with either square loss or cross entropy loss, and MLP with ReLU trained with cross entropy loss.</p>
</sec>
<sec>
<title>4.2. Non-linearly Separable Dataset With Different Noise Distributions</title>
<p>To further evaluate the generality of results on data distributions that are not linearly separable, we use a mixture of Gaussians to generate non-linearly separable datasets for binary classification. Each class consists of three multivariate Gaussians of dimensions <italic>p</italic> &#x0003D; 30. We generate a sample by randomly selecting with the same probability one of the three distribution. To give a more comprehensive evaluation on the effect of different types of noise, we generate unnecessary dimensions using Gaussian distributions with different variance, and we also evaluate two other noise distributions, namely, Gaussian noise with &#x003A3;<sub><italic>ii</italic></sub> &#x0003D; 1, &#x02200;<italic>i</italic>, with &#x003A3;<sub><italic>ij</italic></sub> &#x0003D; 0.5, &#x02200;<italic>i</italic>&#x02260;<italic>j</italic>, and salt and pepper noise, where each vector component can assume value (0, or <italic>u</italic>), based on a Bernoullian distribution on {&#x02212;1, 1}.</p>
<p>We consider the MLP with ReLU and soft-max with cross-entropy loss because among the different variants it is the only well suited to fit non-linearly separable data.</p>
</sec>
<sec>
<title>4.3. Object Recognition Datasets</title>
<p>We evaluate object recognition datasets based on extensions of the MNIST dataset (LeCun et al., <xref ref-type="bibr" rid="B20">1998</xref>) and the Stanford Dogs dataset (Khosla et al., <xref ref-type="bibr" rid="B18">2011</xref>).</p>
<p><bold>Synthetic and Natural MNIST</bold>. We generate two datasets based on MNIST: the synthetic MNIST and the natural MNIST, which have synthetic and natural background, respectively. In both datasets, the MNIST digit is always at the center of the image and normalized between 0 and 1.</p>
<p>In the synthetic MNIST dataset, the <italic>task-unrelated</italic> dimensions are sampled from a Gaussian distribution with (&#x003BC; &#x0003D; 0, &#x003C3; &#x0003D; 0.2) and the <italic>task-related</italic> dimensions are the result of upscaling the MNIST digit. We also combine <italic>task-related</italic> and <italic>unrelated</italic> dimensions by fixing the size of the image and changing the ratio of <italic>task-related</italic> and <italic>unrelated</italic> dimensions by upscaling the MNIST digit.</p>
<p>In the natural MNIST dataset, the background is taken from the Places dataset (Zhou et al., <xref ref-type="bibr" rid="B31">2014</xref>), as in Volokitin et al. (<xref ref-type="bibr" rid="B27">2017</xref>). The size of the image is constant across experiments (256 &#x000D7; 256 pixels), and the size of the MNIST digits determines the amount of <italic>task-related</italic> and <italic>unrelated</italic> dimensions.</p>
<p>We use the MLP with ReLU and cross entropy loss, and also Convolutional Neural Networks (CNNs). The architecture of the CNN consists of three convolutional layers each with max-pooling, followed by two fully connected layers. Since the receptive field size of the CNN neurons may have an impact on the data efficiency, we evaluate different receptive field sizes. We use a factor <italic>r</italic> to scale the receptive field size, such that the convolution filter size is (<italic>r</italic>&#x000B7;3) &#x000D7; (<italic>r</italic>&#x000B7;3) and the pooling region size is (<italic>r</italic>&#x000B7;2) &#x000D7; (<italic>r</italic>&#x000B7;2). We experiment by either fixing <italic>r</italic> to a constant value or adapting <italic>r</italic> to the scale of the MNIST digit, such that the receptive fields of the neurons capture the same object region independently of the scale of the digit.</p>
<p><bold>Stanford Dogs</bold>. Recall our analysis focuses on unnecessary input dimensions that are unbiased. We use the Stanford Dogs dataset (Khosla et al., <xref ref-type="bibr" rid="B18">2011</xref>) as it is reasonable to assume that the bias between breeds of dogs and background is negligible. This dataset contains natural images (227 &#x000D7; 227 pixels) of dogs at different image positions. The amount of <italic>task-unrelated</italic> dimensions is determined by the dog size, which is different for each image. To evaluate the effect of unnecessary input dimensions, we introduce the following five versions of the dataset. Case 1 corresponds to the original image. In case 2, we multiply by zero the pixels of the background, which reduces the variability of the <italic>task-unrelated</italic> dimensions. In case 3, the dog is centered in the image. In case 4, we fix the ratio of <italic>task-related/unrelated</italic> dimensions by centering the dog and scaling it to half of the image size. In case 5, we remove the background by cropping and scaling the dog.</p>
<p>We use a ResNet-18 (He et al., <xref ref-type="bibr" rid="B17">2016</xref>), following the standard pre-processing of the image used in ImageNet.</p>
</sec>
</sec>
<sec sec-type="results" id="s5">
<title>5. Results</title>
<p>In this section, we report results, first on the linearly separable datasets and then, on the object recognition datasets.</p>
<sec>
<title>5.1. Linearly Separable Dataset</title>
<p><xref ref-type="fig" rid="F1">Figure 1A</xref> shows the test accuracy of the pseudo-inverse solution for different number of training examples and unnecessary input dimensions. <xref ref-type="fig" rid="F1">Figure 1B</xref> reports the data efficiency of all networks tested for different number of unnecessary input dimensions. Recall that the data efficiency is measured with the AUTC and summarizes the test accuracy as a function of the amount of training examples, <italic>e.g.</italic>, for the pseudo-inverse solution, the curves in <xref ref-type="fig" rid="F1">Figure 1A</xref> are summarized by the AUTC in <xref ref-type="fig" rid="F1">Figure 1B</xref>. We observe that increasing the amount of <italic>task-unrelated</italic> dimensions harms the data efficiency, <italic>i.e.</italic>, the AUTC drops. Also, the <italic>task-related</italic> dimensions alone do not harm data efficiency, and they alleviate the effect of the <italic>task-unrelated</italic> dimensions.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><italic>Data Efficiency of Linear and Fully Connected Networks Trained on Dataset For Binary Classification</italic>. Data efficiency for different amount of unnecessary input dimensions. Error bars indicate standard deviation across experiment repetitions. <bold>(A)</bold> Test accuracy of the pseudo-inverse solution as function of different amount of training examples. Each curve corresponds to a different amount of additional <italic>task-unrelated</italic> (left plot), <italic>task-related</italic> (middle plot) and <italic>task-related/unrelated</italic> dimensions, as reported in the legend. <bold>(B)</bold> We report the Area Under the Test Curve (AUTC) of the accuracy for different amount of training examples. We indicate with the gradient bar on the <italic>x</italic>-axis the amount of additional <italic>task-unrelated</italic> dimensions (left and right plot), and additional <italic>task-related</italic> dimensions, in the middle plot. The legend indicate the different networks trained and tested on the dataset. <bold>(C)</bold> AUTC values for an MLP with ReLU activation and cross entropy loss, for different types of <italic>task-unrelated</italic> dimensions: Gaussian independent components with increasing variance, Gaussian with non diagonal covariance, and salt and pepper noise. The gradient on the <italic>x</italic>-axis indicates the number of <italic>task-unrelated</italic> dimensions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-760085-g0001.tif"/>
</fig>
<p>These results clarify the difference between robustness to overparameterization in intermediate layers and unnecessary input dimensions. Note that the effect on the test accuracy of increasing the number of hidden units is the opposite of increasing the number of <italic>task-unrelated</italic> input dimensions, <italic>i.e.</italic>, DNNs are not robust to all kinds of overparameterization.</p>
<p>Analytical results for linear regression using the square loss predicts an analogous effect of <italic>task-unrelated</italic> dimensions on the solution. For sake of clarity, we retrace these results in the <xref ref-type="supplementary-material" rid="SM1">Appendix 3.1</xref>, where we outline the effect of additional <italic>task-unrelated</italic> dimensions that are Gaussian-distributed on the pseudo-inverse solution. There, we show that <italic>task-unrelated</italic> dimensions lead to the pseudo-inverse, Tihkonov-regularized solution calculated in the dataset without <italic>task-unrelated</italic> dimensions. Since in this case the regularization can not be tuned or switched off as it is fixed by the number of <italic>task-unrelated</italic> dimensions, it is likely to harm the test accuracy, as we have observed.</p>
<p>The regularizer is beneficial in some specific cases. Following from regularization theory (Hastie et al., <xref ref-type="bibr" rid="B15">2009</xref>), <xref ref-type="supplementary-material" rid="SM1">Appendix 3.2</xref> highlights a noisy regression problem in which certain amounts of <italic>task-unrelated</italic> dimensions help to improve generalization. In a classification problem, Tiknhonov regularization may also be beneficial in some cases. This can be seen in <xref ref-type="fig" rid="F1">Figure 1A</xref>, where we observe that for a given number of training examples, increasing the number of <italic>task-unrelated</italic> dimensions improves the test accuracy in some cases. This specific trend relates to the aforementioned double descend of DNNs (Belkin et al., <xref ref-type="bibr" rid="B6">2019</xref>; Advani et al., <xref ref-type="bibr" rid="B1">2020</xref>). As shown in Nakkiran et al. (<xref ref-type="bibr" rid="B22">2020</xref>), the location of the critical region is affected by the number of training examples and the complexity of the model. Here, the complexity of the model is affected by the number of <italic>task-unrelated</italic> dimensions due to its regularization effect.</p>
</sec>
<sec>
<title>5.2. Non-linearly Separable Datasets With Different Distributions of <italic>Task-Unrelated</italic> Dimensions</title>
<p>In <xref ref-type="fig" rid="F1">Figure 1C</xref>, we show the data efficiency of MLP with ReLU trained with cross entropy loss on non-linearly separable datasets. On the left of the quadrant, we report different distributions of the <italic>task-unrelated</italic> dimensions: Gaussian noise with different <italic><bold>&#x003C3;</bold></italic> (corresponding to multiplicative factor applied on the identical covariance matrix), Gaussian noise with non-diagonal covariance matrix and salt and pepper noise. The amount of <italic>task-unrelated</italic> dimensions is reported through the colored bar (indicated with a gradient). We observe that, similarly to previous results in the linearly separable dataset (<xref ref-type="fig" rid="F1">Figure 1B</xref>), <italic>task-unrelated</italic> dimensions harm data efficiency. Also, as expected, data efficiency deteriorates as the variance of Gaussian noise increases. The combination of <italic>task-related/unrelated</italic> dimensions alleviates the detrimental effect of <italic>task-unrelated</italic> dimensions. These empirical results on MLPs show a similar trend to the one predicted for linear networks, with exception of a less pronounced effect of the double descent behavior.</p>
</sec>
<sec>
<title>5.3. Object Recognition Datasets</title>
<p><xref ref-type="fig" rid="F2">Figure 2A</xref> shows the log-AUTC for the MLP and the CNN for different amount of unnecessary dimensions (an increasing amount as we move left to right), for the synthetic MNIST dataset. In <xref ref-type="supplementary-material" rid="SM1">Appendix 4.1</xref>, we report the test accuracy for different number of unnecessary dimensions, which further strengthens the results of <xref ref-type="fig" rid="F1">Figure 1</xref>. Conclusions are consistent with the previous results in the linearly separable dataset. Also, we observe that CNNs are overall much more data efficient than MLPs, which is expected because of their more adequate inductive bias given by the weight sharing of the convolutions.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><italic>Data Efficiency in Object Recognition Datasets</italic>. Results of DNNs&#x00027; data efficiency for different amount of unnecessary input dimensions, tasks and networks. Error bars indicate standard deviation across experiment repetitions. <bold>(A)</bold> log-AUTC for CNNs and MLPs trained on synthetic MNIST, for different number of unnecessary dimensions. <bold>(B)</bold> Left plot: log-AUTC for networks trained on Natural MNIST for larger amount of training examples; Right: test accuracy on the smallest training set. <bold>(C)</bold> AUTC on the Stanford Dogs dataset for the five cases shown on the left of the panel.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-16-760085-g0002.tif"/>
</fig>
<p><xref ref-type="fig" rid="F2">Figure 2B</xref> shows results in natural MNIST dataset for different ratios of <italic>task-related/unrelated</italic> dimensions. The plots compare CNNs with different receptive field sizes, represented by the factor <italic>r</italic> (see Section 4.3). Since the CNN achieves high accuracy with few examples, the mean and standard deviation of the log-AUTC (left plot) hardly show any variation when computed on more than 20 training examples per class. Yet, the gap of the testing accuracy is considerable for 20 training examples per class (right plot). These results confirm that <italic>task-unrelated</italic> dimensions degrade data efficiency independently of the receptive field sizes (see <xref ref-type="supplementary-material" rid="SM1">Appendix 4.2</xref> for additional results further supporting these conclusions).</p>
<p><xref ref-type="fig" rid="F2">Figure 2C</xref> shows results on the Stanford Dogs dataset, namely the AUTC score across the five cases of unnecessary dimensions that we evaluate. This dataset serves to assess a more realistic scenario, where the objects can appear at different positions and scales. We observe that the <italic>task-unrelated</italic> dimensions, which come from the background, harm the data efficiency (cases 1 to 4 versus case 5). Putting to zero the unnecessary dimensions improves the data efficiency of models trained on the original dataset (cases 2 to 4 vs. case 1). This is because the <italic>task-unrelated</italic> dimensions become redundant as they all take the same value in all images. We also observe that removing the variability of the position and scale of the object hardly affects the data efficiency (case 2 to 4). Thus, learning to discard the background requires more training examples than learning to handle the variability in scale and position of the object.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusions</title>
<p>We have analyzed the effect of unnecessary input dimensions (<italic>e.g.</italic>, object&#x00027;s background). We found that <italic>task-unrelated</italic> dimensions harm the data efficiency, while increasing the number of <italic>task-related</italic> dimensions that are linear combinations of other <italic>task-related</italic> dimensions help to alleviate the negative effect of <italic>task-unrelated</italic> dimensions. These results demonstrate that the robustness of DNNs to overparameterization is limited, as increasing the number of <italic>task-unrelated</italic> input dimensions is a form of overparameterization that degrades the accuracy. Also, our results add to the growing body of works in object recognition that shows that bias in the object&#x00027;s background can undermine the reliability of DNNs. Here we have shown that the problem runs far deeper, as the object&#x00027;s background negatively affects the network even when there is no bias.</p>
<p>Taken together, these results suggest that data efficiency gains could be enabled by mechanisms that remove <italic>task-unrelated</italic> dimensions, such as foveation for image classification (Luo et al., <xref ref-type="bibr" rid="B21">2016</xref>; Akbas and Eckstein, <xref ref-type="bibr" rid="B2">2017</xref>), or also by adapting to DNNs regularization techniques that encourage predictions from a sparse subset of input dimensions (<italic>e.g.</italic>, &#x02113;<sub>1</sub> regularization for linear regression Hastie et al., <xref ref-type="bibr" rid="B16">2019</xref>). Also, our results can be extended to other domains, such as natural language processing and clinical tasks, as the effect of unnecessary dimensions may have been investigated, <italic>e.g.</italic>, (Laksana et al., <xref ref-type="bibr" rid="B19">2020</xref>), but their effects in the data efficiency remain largely unexplored.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: Stanford Dogs dataset: <ext-link ext-link-type="uri" xlink:href="https://www.tensorflow.org/datasets/catalog/stanford_dogs">https://www.tensorflow.org/datasets/catalog/stanford_dogs</ext-link>; MNIST dataset: <ext-link ext-link-type="uri" xlink:href="https://www.tensorflow.org/datasets/catalog/mnist">https://www.tensorflow.org/datasets/catalog/mnist</ext-link>; PLACES dataset: <ext-link ext-link-type="uri" xlink:href="http://places.csail.mit.edu/">http://places.csail.mit.edu/</ext-link>; Synthetic datasets can be generated using the following code: <ext-link ext-link-type="uri" xlink:href="https://github.com/vanessadamario/data_efficiency/blob/main/synthetic_framework/main.py">https://github.com/vanessadamario/data_efficiency/blob/main/synthetic_framework/main.py</ext-link>; The code supporting the conclusions of this article is publicly accessible in the following github repository: <ext-link ext-link-type="uri" xlink:href="https://github.com/vanessadamario/data_efficiency.git">https://github.com/vanessadamario/data_efficiency.git</ext-link>.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>VD&#x00027;A implemented the experiments and carried out the analysis, with contributions of SS and XB. VD&#x00027;A and XB conceived the experiments with contributions of SS and TS. VD&#x00027;A and XB wrote the manuscript with contributions of TS. XB and TS supervised the study. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>This work has been supported by the Center for Brains, Minds, and Machines (funded by NSF STC award CCF-1231216), XB by the R01EY020517 grant from the National Eye Institute (NIH) and XB and VD&#x00027;A by Fujitsu Laboratories Ltd. (Contract No. 40008819) and the MIT-Sensetime Alliance on Artificial Intelligence.</p>

</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>This study received funding from Fujitsu Laboratories Ltd. The funder through TS had the following involvement with the study: conception of the experiment, writing of this article, and supervision of the study. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<ack><p>We would like to thank Pawan Sinha and Tomaso Poggio for useful discussions and insightful advice provided during this project.</p>
</ack>
<sec sec-type="supplementary-material" id="s11">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fncom.2022.760085/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fncom.2022.760085/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Presentation_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Advani</surname> <given-names>M. S.</given-names></name> <name><surname>Saxe</surname> <given-names>A. M.</given-names></name> <name><surname>Sompolinsky</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>High-dimensional dynamics of generalization error in neural networks</article-title>. <source>Neural Netw.</source> <volume>132</volume>:<fpage>428</fpage>&#x02013;<lpage>446</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2020.08.022</pub-id><pub-id pub-id-type="pmid">33022471</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><collab>Akbas E. and Eckstein, M. P..</collab></person-group> (<year>2017</year>). <article-title>Object detection through search with a foveated visual system</article-title>. <source>PLOS Comput. Biolo.</source> <volume>13</volume>:<fpage>e1005743</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1005743</pub-id><pub-id pub-id-type="pmid">28991906</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ansuini</surname> <given-names>A.</given-names></name> <name><surname>Laio</surname> <given-names>A.</given-names></name> <name><surname>Macke</surname> <given-names>J. H.</given-names></name> <name><surname>Zoccolan</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Intrinsic dimension of data representations in deep neural networks,</article-title> in <source>Advances in Neural Information Processing Systems (NeurIPS)</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>6109</fpage>&#x02013;<lpage>6119</lpage>.<pub-id pub-id-type="pmid">29126070</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barbu</surname> <given-names>A.</given-names></name> <name><surname>Mayo</surname> <given-names>D.</given-names></name> <name><surname>Alverio</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Gutfreund</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>ObjectNet: a large-scale bias-controlled dataset for pushing the limits of object recognition models,</article-title> in <source>Advances in Neural Information Processing Systems (NeurIPS)</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>9448</fpage>&#x02013;<lpage>9458</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Beery</surname> <given-names>S.</given-names></name> <name><surname>Van Horn</surname> <given-names>G.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Recognition in terra incognita,</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>456</fpage>&#x02013;<lpage>473</lpage>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Belkin</surname> <given-names>M.</given-names></name> <name><surname>Hsu</surname> <given-names>D.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <name><surname>Mandal</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Reconciling modern machine-learning practice and the classical bias&#x02013;variance trade-off</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>116</volume>, <fpage>15849</fpage>&#x02013;<lpage>15854</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1903070116</pub-id><pub-id pub-id-type="pmid">31341078</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bernstein</surname> <given-names>J.</given-names></name> <name><surname>Yue</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>On the implicit biases of architecture &#x00026; gradient descent</article-title>. <source>arXiv preprint</source> arXiv:2110.04274.</citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A..</given-names></name></person-group> (<year>2021</year>). <article-title>Contemplating real-world object classification,</article-title> in <source>Proceedings of the International Conference on Learning Representations (ICLR)</source>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Choi</surname> <given-names>M. J.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Willsky</surname> <given-names>A. S.</given-names></name></person-group> (<year>2012</year>). <article-title>Context models and out-of-context objects</article-title>. <source>Pattern Recogn. Lett.</source> <volume>33</volume>, <fpage>853</fpage>&#x02013;<lpage>862</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2011.12.004</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Daubechies</surname> <given-names>I..</given-names></name></person-group> (<year>1992</year>). <source>Ten Lectures on Wavelets</source>. <publisher-loc>Philadelphia, PA</publisher-loc>: <publisher-name>SIAM</publisher-name>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>De Palma</surname> <given-names>G.</given-names></name> <name><surname>Kiani</surname> <given-names>B. T.</given-names></name> <name><surname>Lloyd</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Random deep neural networks are biased towards simple functions,</article-title> in <source>Advances in Neural Information Processing Systems (NeurIPS)</source> (<publisher-loc>Vancouver, BC</publisher-loc>).</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.-J.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>ImageNet: a large-scale hierarchical image database,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Miami, FL</publisher-loc>), <fpage>248</fpage>&#x02013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dziugaite</surname> <given-names>G. K.</given-names></name> <name><surname>Roy</surname> <given-names>D. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data,</article-title> in <source>Proceedings of the Conference on Uncertainty in Artificial Intelligence (UAI)</source> (<publisher-loc>Sydney, NSW</publisher-loc>).</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassabis</surname> <given-names>D.</given-names></name> <name><surname>Kumaran</surname> <given-names>D.</given-names></name> <name><surname>Summerfield</surname> <given-names>C.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Neuroscience-inspired artificial intelligence</article-title>. <source>Neuron</source> <volume>95</volume>, <fpage>245</fpage>&#x02013;<lpage>258</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2017.06.011</pub-id><pub-id pub-id-type="pmid">28728020</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name> <name><surname>Friedman</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <source>The Elements of Statistical Learning: Data Mining, Inference, and Prediction</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer Science &#x00026; Business Media</publisher-name>.<pub-id pub-id-type="pmid">19443179</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name> <name><surname>Wainwright</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <source>Statistical Learning With Sparsity: the Lasso and Generalizations</source>. <publisher-loc>London</publisher-loc>: <publisher-name>CRC</publisher-name>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>.<pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Jayadevaprakash</surname> <given-names>N.</given-names></name> <name><surname>Yao</surname> <given-names>B.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2011</year>). <article-title>Novel dataset for fine-grained image categorization,</article-title> in <source>First Workshop on Fine-Grained Visual Categorization, IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://vision.stanford.edu/aditya86/ImageNetDogs/">http://vision.stanford.edu/aditya86/ImageNetDogs/</ext-link> (accessed November 15, 2021).</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laksana</surname> <given-names>E.</given-names></name> <name><surname>Aczon</surname> <given-names>M.</given-names></name> <name><surname>Ho</surname> <given-names>L.</given-names></name> <name><surname>Carlin</surname> <given-names>C.</given-names></name> <name><surname>Ledbetter</surname> <given-names>D.</given-names></name> <name><surname>Wetzel</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>The impact of extraneous features on the performance of recurrent neural network models in clinical tasks</article-title>. <source>J. Biomed. Inf.</source> <volume>102</volume>:<fpage>103351</fpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2019.103351</pub-id><pub-id pub-id-type="pmid">31870949</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Haffner</surname> <given-names>P.</given-names></name></person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x02013;<lpage>2324</lpage>. <pub-id pub-id-type="doi">10.1109/5.726791</pub-id> Available online at: <ext-link ext-link-type="uri" xlink:href="http://yann.lecun.com/exdb/mnist/">http://yann.lecun.com/exdb/mnist/</ext-link> (accessed: November 15, 2021).<pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Boix</surname> <given-names>X.</given-names></name> <name><surname>Roig</surname> <given-names>G.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name></person-group> (<year>2016</year>). <source>Foveation-Based Mechanisms Alleviate Adversarial examples. Technical Report CBMM Memo No. 44. Center for Brains, Minds and Machines</source>. Available Online at: <ext-link ext-link-type="uri" xlink:href="https://cbmm.mit.edu/sites/default/files/publications/cbmm_memo_044.pdf">https://cbmm.mit.edu/sites/default/files/publications/cbmm_memo_044.pdf</ext-link></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nakkiran</surname> <given-names>P.</given-names></name> <name><surname>Kaplun</surname> <given-names>G.</given-names></name> <name><surname>Bansal</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>T.</given-names></name> <name><surname>Barak</surname> <given-names>B.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep double descent: where bigger models and more data hurt,</article-title> in <source>Proceedings of the International Conference on Learning Representations (ICLR)</source>. (Addis Ababa).</citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Recanatesi</surname> <given-names>S.</given-names></name> <name><surname>Farrell</surname> <given-names>M.</given-names></name> <name><surname>Advani</surname> <given-names>M.</given-names></name> <name><surname>Moore</surname> <given-names>T.</given-names></name> <name><surname>Lajoie</surname> <given-names>G.</given-names></name> <name><surname>Shea-Brown</surname> <given-names>E.</given-names></name></person-group> (<year>2019</year>). <article-title>Dimensionality compression and expansion in deep neural networks</article-title>. <source>arXiv preprint</source> arXiv:1906.00443.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sch&#x000F6;lkopf</surname> <given-names>B.</given-names></name> <name><surname>Herbrich</surname> <given-names>R.</given-names></name> <name><surname>Smola</surname> <given-names>A. J.</given-names></name></person-group> (<year>2001</year>). <article-title>A generalized representer theorem,</article-title> in <source>Proceedings of the International Conference on Computational Learning Theory (COLT)</source> (<publisher-loc>Amsterdam</publisher-loc>), <fpage>416</fpage>&#x02013;<lpage>426</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>M.</given-names></name> <name><surname>Yi</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>Eliminating background-bias for robust person re-identification,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>5794</fpage>&#x02013;<lpage>5803</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Efros</surname> <given-names>A. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Unbiased look at dataset bias,</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Colorado Springs, CO</publisher-loc>), <fpage>1521</fpage>&#x02013;<lpage>1528</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Volokitin</surname> <given-names>A.</given-names></name> <name><surname>Roig</surname> <given-names>G.</given-names></name> <name><surname>Poggio</surname> <given-names>T. A.</given-names></name></person-group> (<year>2017</year>). <article-title>Do deep neural networks suffer from crowding?</article-title> in <source>Advances in Neural Information Processing Systems (NIPS)</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>5628</fpage>&#x02013;<lpage>5638</lpage>.</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>K. Y.</given-names></name> <name><surname>Engstrom</surname> <given-names>L.</given-names></name> <name><surname>Ilyas</surname> <given-names>A.</given-names></name> <name><surname>Madry</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Noise or signal: the role of image backgrounds in object recognition,</article-title> in <source>Proceedings of the International Conference on Learning Representations (ICLR)</source>.</citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Bengio</surname> <given-names>S.</given-names></name> <name><surname>Hardt</surname> <given-names>M.</given-names></name> <name><surname>Recht</surname> <given-names>B.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name></person-group> (<year>2017</year>). <article-title>Understanding deep learning requires rethinking generalization,</article-title> in <source>Proceedings of the International Conference on Learning Representations (ICLR)</source> (<publisher-loc>Toulon</publisher-loc>).</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>T..</given-names></name></person-group> (<year>2005</year>). <article-title>Learning bounds for kernel regression using effective data dimensionality</article-title>. <source>Neural Comput.</source> <volume>17</volume>, <fpage>2077</fpage>&#x02013;<lpage>2098</lpage>. <pub-id pub-id-type="doi">10.1162/0899766054323008</pub-id><pub-id pub-id-type="pmid">15992491</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Lapedriza</surname> <given-names>A.</given-names></name> <name><surname>Xiao</surname> <given-names>J.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Learning deep features for scene recognition using Places database,</article-title> in <source>Advances in Neural Information Processing Systems (NIPS)</source> (Montreal, QC), <fpage>487</fpage>&#x02013;<lpage>495</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://places.csail.mit.edu/">http://places.csail.mit.edu/</ext-link> (accessed November 15, 2021).</citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Xie</surname> <given-names>L.</given-names></name> <name><surname>Yuille</surname> <given-names>A. L.</given-names></name></person-group> (<year>2017</year>). <article-title>Object recognition with and without objects,</article-title> in <source>Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)</source> (<publisher-loc>Melbourne, QC</publisher-loc>), <fpage>3609</fpage>&#x02013;<lpage>3615</lpage>.</citation>
</ref>
</ref-list> 
</back>
</article> 