<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3-mathml3.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" dtd-version="1.3" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title-group>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2025.1665874</article-id>
<article-version article-version-type="Version of Record" vocab="NISO-RP-8-2008"/>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Original Research</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Turing Test for artificial nets devoted to vision</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" equal-contrib="yes">
<name><surname>Vila-Tom&#x000E1;s</surname> <given-names>Jorge</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2287672"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Software" vocab-term-identifier="https://credit.niso.org/contributor-roles/software/">Software</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
</contrib>
<contrib contrib-type="author" equal-contrib="yes">
<name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>Pablo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2361442"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Software" vocab-term-identifier="https://credit.niso.org/contributor-roles/software/">Software</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Qiang</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Software" vocab-term-identifier="https://credit.niso.org/contributor-roles/software/">Software</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
<uri xlink:href="https://loop.frontiersin.org/people/1101416"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Laparra</surname> <given-names>Valero</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Funding acquisition" vocab-term-identifier="https://credit.niso.org/contributor-roles/funding-acquisition/">Funding acquisition</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
<uri xlink:href="https://loop.frontiersin.org/people/187819"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Malo</surname> <given-names>Jes&#x000FA;s</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/136519"/>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Funding acquisition" vocab-term-identifier="https://credit.niso.org/contributor-roles/funding-acquisition/">Funding acquisition</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
</contrib>
</contrib-group>
<aff id="aff1"><label>1</label><institution>Image Processing Lab, Universitat de Val&#x000E8;ncia</institution>, <city>Valencia</city>, <country country="es">Spain</country></aff>
<aff id="aff2"><label>2</label><institution>TReNDS, Georgia State, Georgia Tech, and Emory</institution>, <city>Atlanta, GA</city>, <country country="us">United States</country></aff>
<author-notes>
<corresp id="c001"><label>&#x0002A;</label>Correspondence: Jes&#x000FA;s Malo, <email xlink:href="mailto:jesus.malo@uv.es">jesus.malo@uv.es</email></corresp>
<fn fn-type="equal" id="fn001"><label>&#x02020;</label><p>These authors have contributed equally to this work</p></fn></author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-01-05">
<day>05</day>
<month>01</month>
<year>2026</year>
</pub-date>
<pub-date publication-format="electronic" date-type="collection">
<year>2025</year>
</pub-date>
<volume>8</volume>
<elocation-id>1665874</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>07</month>
<year>2025</year>
</date>
<date date-type="rev-recd">
<day>23</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>17</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2026 Vila-Tom&#x000E1;s, Hern&#x000E1;ndez-C&#x000E1;mara, Li, Laparra and Malo.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>Vila-Tom&#x000E1;s, Hern&#x000E1;ndez-C&#x000E1;mara, Li, Laparra and Malo</copyright-holder>
<license>
<ali:license_ref start_date="2026-01-05">https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</license-p>
</license>
</permissions>
<abstract>
<p>In this work<xref ref-type="fn" rid="fn0003"><sup>1</sup></xref> we argue that, despite recent claims about successful modeling of the visual brain using deep nets, the problem is far from being solved, particularly for low-level vision. Open issues include <italic>where should we read from in ANNs to check behavior? What should be the read-out? Is this ad-hoc read-out considered part of the brain model or not?</italic> In order to understand vision-ANNs, <italic>should we use artificial psychophysics or artificial physiology?</italic> Anyhow, <italic>should artificial tests literally match the experiments done with humans?</italic> These questions suggest a clear need for biologically sensible tests for deep models of the visual brain, and more generally, to understand ANNs devoted to generic vision tasks. Following our use of low-level facts from <italic>Vision Science</italic> in Image Processing, we present a low-level dataset compiling the basic spatio-chromatic properties that describe the adaptive bottleneck of the retina-V1 pathway and are not currently available in popular databases such as BrainScore. We propose its use for qualitative and quantitative model evaluation. As an illustration of the proposed methods, we check the behavior of three recent models with similar deep architectures: (1) A parametric model tuned via the psychophysical method of Maximum Differentiation [Malo &#x00026; Simoncelli SPIE 15, Martinez et al. PLOS 18, Martinez et al. Front. Neurosci. 19], (2) A non-parametric model (the <italic>PerceptNet</italic>) tuned to maximize the correlation with humans on subjective image distortions [Hepburn et al. IEEE ICIP 20], and (3) A model with the same encoder as the <italic>PerceptNet</italic>, but tuned for image segmentation [Hernandez-Camara et al. Patt.Recogn.Lett. 23, Hernandez-Camara et al. Neurocomp. 25]. Results on the proposed 10 compelling psycho/physio visual properties show that the first (parametric) model is the one with behavior closest to humans.</p></abstract>
<kwd-group>
<kwd>evaluation of AI models</kwd>
<kwd>neural networks for vision</kwd>
<kwd>human vision</kwd>
<kwd>Turing Test</kwd>
<kwd>low-level visual psychophysics</kwd>
<kwd>linear &#x0002B; non-linear cascade</kwd>
<kwd>image quality</kwd>
<kwd>image segmentation</kwd>
</kwd-group>
<funding-group>
<award-group id="gs1">
<funding-source id="sp1">
<institution-wrap>
<institution>Ministerio de Ciencia e Innovaci&#x000F3;n</institution>
<institution-id institution-id-type="doi" vocab="open-funder-registry" vocab-identifier="10.13039/open_funder_registry">10.13039/501100004837</institution-id>
</institution-wrap>
</funding-source>
</award-group>
<award-group id="gs2">
<funding-source id="sp2">
<institution-wrap>
<institution>Generalitat Valenciana</institution>
<institution-id institution-id-type="doi" vocab="open-funder-registry" vocab-identifier="10.13039/open_funder_registry">10.13039/501100003359</institution-id>
</institution-wrap>
</funding-source>
</award-group>
<award-group id="gs3">
<funding-source id="sp3">
<institution-wrap>
<institution>Fundaci&#x000F3;n BBVA</institution>
<institution-id institution-id-type="doi" vocab="open-funder-registry" vocab-identifier="10.13039/open_funder_registry">10.13039/100007406</institution-id>
</institution-wrap>
</funding-source>
</award-group>
<funding-statement>The author(s) declared that financial support was received for this work and/or its publication. The <italic>invited talk</italic> at the <italic>Artificial Intelligence Evaluation Workshop 2022</italic> was funded by the University of Bristol. The computational work was partially funded by MCIN/AEI/FEDER/UE under Grants PID2020-118071GB-I00 and PID2023-152133NB-I00, by Spanish MIU under Grant FPU21/02256, and by Generalitat Valenciana under Projects GV/2021/074, CIPROM/2021/056, and by the grant BBVA Foundations of Science program: Maths, Stats, Comp. Sci. and AI (VIS4NN). Some computer resources were provided by Artemisa, funded by the EU ERDF through the Instituto de F&#x000ED;sica Corpuscular, IFIC (CSIC-UV). The audio draft of this work (<italic>the talk</italic>) was recorded at El Saler beach (Valencia) and then transcription was done at the <italic>Lisboa</italic> restaurant (Valencia): its staff was particularly helpful at writing time during <italic>Fallas</italic> 2025.</funding-statement>
</funding-group>
<counts>
<fig-count count="17"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="145"/>
<page-count count="25"/>
<word-count count="18411"/>
</counts>
<custom-meta-group>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Machine Learning and Artificial Intelligence</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<label>1</label>
<title>Introduction</title>
<sec>
<label>1.1</label>
<title>Prologue</title>
<p>This work reproduces our <italic>talk</italic> (otherwise unpublished in print) at the AI Evaluation Workshop in June 2022 at the AI Dept. of the University of Bristol organized by Prof. Raul Santos of the Eng. Maths Dept. of UoB (<xref ref-type="bibr" rid="B96">Malo et al., 2022</xref>). That <italic>talk</italic> proposed an original methodology (with experimental results) to evaluate deep nets devoted to vision tasks and was the seed of our current (as of 2025) work with Prof. Jeff Bowers of the Psychol. Dept. of UoB, as a low-level complement to his (high-level) proposals in <xref ref-type="bibr" rid="B12">Bowers et al. (2023</xref>) and <xref ref-type="bibr" rid="B10">Biscione et al. (2024</xref>). Journal publication of this 2022 <italic>talk</italic> is pertinent for a wider audience because this approach, based on low-level visual psychophysics, is still unusual in the AI and machine learning communities, despite some researchers are independently proposing very similar evaluations quite recently (<xref ref-type="bibr" rid="B16">Cai et al., 2025</xref>; <xref ref-type="bibr" rid="B39">Hammou et al., 2025</xref>). As shown below, our proposed evaluation program includes facts that go beyond the luminance, color, and contrast masking properties considered in <xref ref-type="bibr" rid="B16">Cai et al. (2025</xref>) and <xref ref-type="bibr" rid="B39">Hammou et al. (2025</xref>). The work of Rafal Mantiuk&#x00027;s lab shares the same spirit and focus on low-level psychophysics, but his focus on <italic>quantitative comparison</italic> is in contrast with our proposal, which, while including quantitative comparison, also stresses the <italic>qualitative understanding</italic> of the response curves. In that way, AI researchers can spot major conceptual errors in deep models easily. Moreover, as explained below, the selected visual stimuli<xref ref-type="fn" rid="fn0004"><sup>2</sup></xref> (and associated psychophysical properties) allow us to intuitively infer modifications in the architectures in order to correct the detected errors.</p>
</sec>
<sec>
<label>1.2</label>
<title>Motivation: is that model really human-like?</title>
<p>The motivation for our proposal starts by reviewing the claims about how deep learning models are the ultimate tool to model the visual brain, as recalled in <xref ref-type="bibr" rid="B12">Bowers et al. (2023</xref>). Claims cited by Bowers et al. include <xref ref-type="bibr" rid="B57">Kubilius et al. (2019</xref>), <xref ref-type="bibr" rid="B103">Mehrer et al. (2021</xref>), <xref ref-type="bibr" rid="B145">Zhuang et al. (2021</xref>), <xref ref-type="bibr" rid="B127">Storrs et al. (2021</xref>), <xref ref-type="bibr" rid="B113">Rajalingham et al. (2018</xref>) and <xref ref-type="bibr" rid="B74">Macpherson et al. (2021</xref>), and other examples in the same vein include <xref ref-type="bibr" rid="B14">Cadena et al. (2019</xref>) and <xref ref-type="bibr" rid="B13">Burg et al. (2021</xref>). A skeptical tone about claims (<xref ref-type="bibr" rid="B12">Bowers et al., 2023</xref>) is a good practice in science<xref ref-type="fn" rid="fn0005"><sup>3</sup></xref>. Two examples of this skepticism regarding the eventual plausibility of models include major scientists such as <italic>Tomaso Poggio</italic> and <italic>Horace Barlow</italic>. In the 70s, Marr and Poggio proposed a taxonomy of the approaches to the vision problem: their famous <italic>separate abstraction levels</italic>, namely, computational, algorithmic, and implementation (<xref ref-type="bibr" rid="B98">Marr and Poggio, 1977</xref>; <xref ref-type="bibr" rid="B97">Marr, 1978</xref>). However, 42 years later, in view of the current tools to optimize models, Poggio himself questioned the separability of these levels (<xref ref-type="bibr" rid="B112">Poggio, 2021</xref>). This taxonomy has been inspiring for decades, but now it is under debate (<xref ref-type="bibr" rid="B67">Lengyel, 2024</xref>; <xref ref-type="bibr" rid="B111">Pillow, 2024</xref>; <xref ref-type="bibr" rid="B89">Malo and Hern&#x000E1;ndez-C&#x000E1;mara, 2024</xref>; <xref ref-type="bibr" rid="B46">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025c</xref>). For example, work on color illusions (<xref ref-type="bibr" rid="B33">Gomez-Villa et al., 2020b</xref>), on CSFs in autoencoders (<xref ref-type="bibr" rid="B68">Li et al. 2022</xref>), and on subjective distances between images in ANNs (<xref ref-type="bibr" rid="B46">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025c</xref>; <xref ref-type="bibr" rid="B41">Hepburn et al., 2022</xref>) stress the relation between the computational and the algorithmic levels, thus questioning previous (purely computational) explanations that disregard architecture (<xref ref-type="bibr" rid="B85">Malo and Guti&#x000E9;rrez, 2006</xref>; <xref ref-type="bibr" rid="B59">Laparra et al., 2012</xref>; <xref ref-type="bibr" rid="B61">Laparra and Malo, 2015</xref>). In a similar vein, Horace Barlow, 50 years after his inspiring <italic>Efficient Coding Hypothesis</italic> (<xref ref-type="bibr" rid="B4">Barlow, 1959</xref>, <xref ref-type="bibr" rid="B5">1961</xref>), questioned his own purely infomax approach (<xref ref-type="bibr" rid="B6">Barlow, 2001</xref>) <xref ref-type="fn" rid="fn0006"><sup>4</sup></xref>.</p>
<p>That skepticism is the core of the spirit in <xref ref-type="bibr" rid="B12">Bowers et al. (2023</xref>), and also the motivation of this work, which has two key ideas:</p>
<list list-type="bullet">
<list-item><p>The use of AI techniques (e.g., deep learning) to understand the visual brain may not be as easy as people thought back in 2022, and even now. More explanatory tests are required.</p></list-item>
<list-item><p>Our specific proposal here is a Turing-like test (<xref ref-type="bibr" rid="B132">Turing, 1950</xref>) based on 10 properties of low-level human vision (our <italic>Decalogue</italic>) to check if a certain artificial model behaves as the (low-level) human visual brain.</p></list-item>
</list>
</sec>
<sec>
<label>1.3</label>
<title>Structure of the paper</title>
<p>Section 2 states that the question <italic>Are the models sensible from the point of view of low-level physiology and psychophysics?</italic> remains open from the perspective of modeling and evaluation. In Section 3, we propose our contribution: an easy-to-use test (consisting of online available visual stimuli) and associated responses for qualitative and quantitative evaluation of deep learning vision models. These stimuli visually illustrate low-level phenomena described by classical <italic>Vision Science</italic>. In Section 4, we illustrate the proposed method through the qualitative and quantitative evaluation of three recent models: (1) a classically formulated, not end-to-end optimized model with a functional form derived from classical vision science literature, where the specific values of its parameters have been psychophysically measured (<xref ref-type="bibr" rid="B95">Malo and Simoncelli, 2015</xref>; <xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>, <xref ref-type="bibr" rid="B99">2019</xref>; <xref ref-type="bibr" rid="B81">Malo et al., 2024</xref>). (2) A network with a bio-inspired architecture but with free parameters end-to-end optimized to reproduce subjective image quality, the <italic>PerceptNet</italic> (<xref ref-type="bibr" rid="B40">Hepburn et al., 2020</xref>). It resembles AlexNet and VGG, but it was specifically designed to accommodate the known aspects of the retina&#x02013;cortex visual pathway using a constrained version of divisive normalization (<xref ref-type="bibr" rid="B3">Ball&#x000E9; et al., 2017</xref>). And (3) a model with the same encoder as the <italic>PerceptNet</italic>, but augmented with a decoder, and both (encoder and decoder) are trained for image segmentation (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>; <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>), which is also a biologically plausible task. Section 5 discusses what can be learned from the proposed test and shows an example of how a model can be fixed. Note that even in the engineering case where one does not necessarily need the networks to resemble humans, one would always want them to have good adaptation properties to achieve good generalization, and potential failures in this regard become clearly evident through the proposed tests. Finally, Section 6 concludes the paper. <xref ref-type="supplementary-material" rid="SM1">Supplementary materials</xref> include (a) the ground-truth values of the response curves in the proposed tests, (b) in-depth results of the CSFs of the models using six different read-out strategies at all the layers of the considered networks, and (c) a critical discussion of the aggregation of quantitative quality descriptors used in BrainScore (<xref ref-type="bibr" rid="B119">Schrimpf et al., 2018</xref>).</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Open issues in modeling vision</title>
<p>As pointed out in <xref ref-type="bibr" rid="B131">Torralba et al. (2024</xref>), the basic question, as in human vision, is how to deal with deep models which are hardly explainable black boxes once trained.</p>
<sec>
<label>2.1</label>
<title>Uncertain computational goal</title>
<p>First, the more general open issue is the discussion on the <italic>computational goal</italic> that eventually explains the organization and behavior of visual systems. Consider architectures/tasks such as the ones presented in <xref ref-type="fig" rid="F1">Figure 1</xref>. These tasks are related to low-, mid-, and high-level tasks arguably implemented by biological vision. In biology, enhancement of the blurry and noisy signal in the retina has been proposed as an explanation of the LGN, as pursuing this goal may reproduce some of its spatio-chromatic (<xref ref-type="bibr" rid="B2">Atick et al., 1992</xref>; <xref ref-type="bibr" rid="B68">Li et al., 2022</xref>) and purely chromatic (<xref ref-type="bibr" rid="B33">Gomez-Villa et al., 2020b</xref>) features. Another example is the compression, possibly, happening in part at the LGN bottleneck and at the feature selection after V1. Bandwidth limitation, dimensionality reduction, and attention focus are sensible goals in this regard (<xref ref-type="bibr" rid="B51">Karklin and Simoncelli, 2011</xref>; <xref ref-type="bibr" rid="B70">Lindsey et al., 2019</xref>; <xref ref-type="bibr" rid="B144">Zhaoping, 2014</xref>). A number of compression algorithms [for images (<xref ref-type="bibr" rid="B135">Wallace, 1991</xref>; <xref ref-type="bibr" rid="B93">Malo et al., 1995</xref>, <xref ref-type="bibr" rid="B82">2000a</xref>; <xref ref-type="bibr" rid="B129">Taubman and Marcellin, 2013</xref>; <xref ref-type="bibr" rid="B79">Malo et al., 2006</xref>; <xref ref-type="bibr" rid="B3">Ball&#x000E9; et al., 2017</xref>) and video <xref ref-type="bibr" rid="B64">Le Gall, 1992</xref>; <xref ref-type="bibr" rid="B83">Malo et al., 2000b</xref>,<xref ref-type="bibr" rid="B87">c</xref>, <xref ref-type="bibr" rid="B86">2001</xref>] have been based on human vision models. Segmentation is arguably another (mid-level) task that has to be done by biological vision, and biological non-linearities have been shown to improve segmentation in images (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>; <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>) and video (<xref ref-type="bibr" rid="B87">Malo et al., 2000c</xref>, <xref ref-type="bibr" rid="B86">2001</xref>). Arguably, segmentation is implemented in the <italic>where</italic> channel from the lower-level primitives extracted in V1 (<xref ref-type="bibr" rid="B35">Goodale et al., 1991</xref>; <xref ref-type="bibr" rid="B105">Milner and Goodale, 1992</xref>). Higher-level tasks such as classification are supposed to happen in the <italic>what</italic> channel (<xref ref-type="bibr" rid="B71">Logothetis and Sheinberg, 1996</xref>; <xref ref-type="bibr" rid="B54">Kreiman et al., 2000</xref>). Similarly, in standard models such as the one depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>, biological non-linearities have been shown to have a significant role in classification (<xref ref-type="bibr" rid="B21">Coen-Cagli and Schwartz, 2013</xref>; <xref ref-type="bibr" rid="B104">Miller et al., 2022</xref>).</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>Image denoising, image compression, image segmentation, and image classification architectures with (eventually) biological correlates in the LGN, the V1, and beyond. However, it is not obvious how these tasks may be combined to explain biological vision. Images reproduced with permission from: M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, &#x0201C;The Cityscapes Dataset for Semantic Urban Scene Understanding,&#x0201D; in <italic>Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, 2016.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0001.tif">
<alt-text>Diagram illustrating a neural network process for image analysis with four stages. A noisy street image labeled &#x00022;Noisy input (Retina)&#x00022; is processed through Denoising (LGN), Compression (LGN-Cortex), Segmentation (where), and Classification (what). Each stage shows respective transformations, ending with a clear identification of &#x00022;2 cars&#x00022;.</alt-text>
</graphic>
</fig>
</sec>
<sec>
<label>2.2</label>
<title>Uncertain read-out mechanisms</title>
<p>As stated in the introduction, in the age of automatic differentiation where the classical Marr-Poggio levels are not that separated, the <italic>computational goal</italic> is not the only open issue. For instance, in order to check if a (mathematical) model is biologically sensible, where should we read the signals from? The read-out mechanism is also important. Note that the fact that a certain layer has the necessary information in order to solve a task (read-out in <italic>any complicated</italic> way, e.g., a highly specialized dense network) is not enough to say that this layer represents the way the visual brain works: the necessary information is already present in the retina (if read in a proper way) and, of course, the retina is not a good model for the rest of the visual brain. This problem is illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig position="float" id="F2">
<label>Figure 2</label>
<caption><p>Given a deep model successfully trained for some visual task, the read-out location and read-out mechanism (or decoder) are important to assess its biological plausibility.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0002.tif">
<alt-text>Flowchart illustrating an encoding and decoding process with color-coded circles and rectangles. Left side shows encoding with red, yellow, and blue circles, indicating different features. Right side shows decoding with purple and blue circles. Green and turquoise labels suggest &#x00022;All features?&#x00022; and &#x00022;One feature?&#x00022;. Arrows connect elements, implying a sequence from encoding to decoding.</alt-text>
</graphic>
</fig>
<p>In the case of doing <italic>artificial physiology</italic>, i.e., reading the signals from certain neurons or layers, or <italic>artificial psychophysics</italic>, i.e., trying to make decisions from the responses of the network to decide if a certain stimulus is visible or not, one should propose a <italic>read-out mechanism</italic> to summarize the responses into a decision variable (see <xref ref-type="fig" rid="F3">Figure 3</xref>). The selection of the <italic>read-out mechanism</italic> is not trivial. In fact, the quality of the read-out information may strongly depend on the complexity of this (arbitrarily selected) mechanism. As a result, one may not be able to tell if the model itself is good, or if the good behavior has to be attributed to a clever read-out that is not part of the model. Examples include the use of classifiers at certain locations of the network to make a decision on visibility, as in (<xref ref-type="bibr" rid="B21">Coen-Cagli and Schwartz, 2013</xref>; <xref ref-type="bibr" rid="B1">Akbarinia et al., 2023</xref>), or without classifiers relying on the model output (<xref ref-type="bibr" rid="B43">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025a</xref>); or the (more classical) use of Euclidean distances between stimuli to tell if they are discriminable (<xref ref-type="bibr" rid="B130">Teo and Heeger, 1994</xref>; <xref ref-type="bibr" rid="B68">Li et al., 2022</xref>; <xref ref-type="bibr" rid="B46">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025c</xref>). These (arbitrary) decisions definitely affect the characterization of the system, e.g., its frequency response (<xref ref-type="bibr" rid="B68">Li et al., 2022</xref>; <xref ref-type="bibr" rid="B1">Akbarinia et al., 2023</xref>). For example, linear or non-linear classifiers effectively apply different (non-Euclidean) distance metrics (<xref ref-type="bibr" rid="B26">Duda and Hart, 1973</xref>) and, hence, they should lead to different decisions.</p>
<fig position="float" id="F3">
<label>Figure 3</label>
<caption><p>In artificial physiology <bold>(Left)</bold> and in artificial psychophysics <bold>(Right)</bold>, the arbitrary decoder to read out model activations is critical.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0003.tif">
<alt-text>Diagram illustrating the process of decoding neural signals. On the left, physiology involves a brain generating a signal from a visual stimulus. The center section shows an artificial physiology model predicting neural signals decoded by a system. On the right, psychophysics depicts a person viewing a screen and processing visual information, with the thought bubble &#x00022;I see it.&#x00022;</alt-text>
</graphic>
</fig>
<p>Another (more particular) discussion is the debate on the summation, which is classical in vision science (<xref ref-type="bibr" rid="B36">Graham, 1989</xref>): for instance, which Minkowski exponent is more physiologically plausible? Note that using different norms and summation schemes definitely leads to different results (<xref ref-type="bibr" rid="B62">Laparra et al., 2010</xref>). A final (also non-obvious) way of assessing stimuli in the network is measuring differences in the statistical properties of the response (<xref ref-type="bibr" rid="B136">Wang et al., 2004</xref>; <xref ref-type="bibr" rid="B25">Ding et al., 2022</xref>) or measuring information flow along the network (<xref ref-type="bibr" rid="B125">Sheikh et al., 2005</xref>; <xref ref-type="bibr" rid="B124">Sheikh and Bovik, 2006</xref>; <xref ref-type="bibr" rid="B76">Malo, 2020</xref>; <xref ref-type="bibr" rid="B90">Malo et al., 2021</xref>; <xref ref-type="bibr" rid="B69">Li et al., 2024</xref>). These options require making non-trivial decisions such as which statistical descriptors make sense (<xref ref-type="bibr" rid="B30">Gatys et al., 2016</xref>; <xref ref-type="bibr" rid="B25">Ding et al., 2022</xref>), or how to set the level of noise in the network (<xref ref-type="bibr" rid="B125">Sheikh et al., 2005</xref>; <xref ref-type="bibr" rid="B124">Sheikh and Bovik, 2006</xref>; <xref ref-type="bibr" rid="B76">Malo, 2020</xref>). In this regard, models can be improved either by changing the architecture and the measures of information (<xref ref-type="bibr" rid="B90">Malo et al., 2021</xref>; <xref ref-type="bibr" rid="B60">Laparra et al., 2025</xref>), or by better estimations of the internal noise (<xref ref-type="bibr" rid="B80">Malo et al., 2025</xref>).</p>
</sec>
<sec>
<label>2.3</label>
<title>Uncertain experimental setting</title>
<p>And finally, the third open issue is the way of doing the evaluation: <italic>the experiment implementation matters</italic>. In particular, <italic>should we use artificial physiology or artificial psychophysics?</italic> Current techniques by the machine learning community to visualize the behavior of the networks (<xref ref-type="bibr" rid="B75">Mahendran and Vedaldi, 2016</xref>; <xref ref-type="bibr" rid="B73">Luo et al., 2016</xref>) are based on classical single-cell recordings, such as the very concept of <italic>receptive field</italic> (<xref ref-type="bibr" rid="B47">Hubel et al., 1959</xref>; <xref ref-type="bibr" rid="B48">Hubel and Wiesel, 1961</xref>; <xref ref-type="bibr" rid="B114">Ringach, 2002</xref>), and the identification of sensitive neurons by looking at the stimulus that maximizes the neuron response, which is a common practice in visual neuroscience (<xref ref-type="bibr" rid="B128">Tailby et al., 2008</xref>). However, there are more sophisticated techniques such as <italic>reverse correlation</italic> which are used both in physiology (<xref ref-type="bibr" rid="B115">Ringach and Shapley, 2004</xref>) and in psychophysics (<xref ref-type="bibr" rid="B27">Eckstein and Ahumada, 2002</xref>), and these are not yet widely used in machine learning. Regarding the experimental setting, <italic>should one go for a literal reproduction of the experiments with humans, or should one try an idealized version of the experiment?</italic> This open question can be illustrated by the example in <xref ref-type="fig" rid="F4">Figure 4</xref> on the spectral sensitivity of a network.</p>
<fig position="float" id="F4">
<label>Figure 4</label>
<caption><p>In measuring the spectral sensitivity of certain elements of a network, one may try a <italic>literal</italic> reproduction of human psychophysics <bold>(Left)</bold> or an <italic>idealized</italic> experiment <bold>(Right)</bold>. The <italic>literal reproduction</italic> could be done through matching experiments (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>): finding the ratio of energies necessary to match the response to quasi-spectral stimuli of different wavelengths. The <italic>idealized</italic> version of the experiment could be based on measuring the increment of response (distance) due to equal energy quasi-monochromatic stimuli with regard to a common reference.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0004.tif">
<alt-text>Comparison of literal and idealized methods for matching brightness and distance of equienergetic stimuli. The left panel shows graphs and color scales for literal matching with variable energy. The right panel depicts idealized matching using distance functions. Both panels include spectral graphs with corresponding color-coded bars and analytical plots.</alt-text>
</graphic>
</fig>
<p>This is a non-trivial question because, for instance, some techniques to assess visual illusions in a model involves the inversion of the inner representation (<xref ref-type="bibr" rid="B110">Otazu et al., 2010</xref>; <xref ref-type="bibr" rid="B33">Gomez-Villa et al., 2020b</xref>), which <italic>does not happen in the human brain</italic>, while others, similarly to human psychophysics (<xref ref-type="bibr" rid="B137">Ware and Cowan, 1982</xref>), are based on <italic>matching</italic> the response at the inner representation (<xref ref-type="bibr" rid="B33">Gomez-Villa et al., 2020b</xref>, <xref ref-type="bibr" rid="B34">2025</xref>). As stated above, this has implications for deciding at which layer one should impose the matching (or where to read from).</p>
</sec>
<sec>
<label>2.4</label>
<title>Better evaluation techniques are needed</title>
<p>All these non-trivial decisions (despite that they all belong to low-level characterizations of the visual system) clearly point out the need for better methodologies for model evaluation in order to assess how close different models may be to the visual brain. These better methods should easily show the impact of the open issues mentioned above.</p>
<p>In this context, our proposal here is simple: <italic>just provide the code to generate a set of well-selected stimuli that illustrate a number of classical low-level visual psychophysics facts and have them prepared as inputs to evaluate image-computable models</italic>. The first version of such a <italic>low-level Turing Test</italic> (back in 2022) included stimuli for 10 wellknown behaviors (our <italic>Decalogue</italic>). That Decalogue is being extended to 20 properties in our on-going (2025) collaboration with Prof. Bowers (<xref ref-type="bibr" rid="B78">Malo and Bowers, 2024</xref>).</p>
<p>The selected stimuli here (which include color and texture) are behind the current understanding of early vision as a set of linear-non-linear layers (<xref ref-type="bibr" rid="B118">Rust and Movshon, 2005</xref>; <xref ref-type="bibr" rid="B120">Sch&#x000FC;tt and Wichmann, 2017</xref>; <xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>, <xref ref-type="bibr" rid="B99">2019</xref>; <xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al., 2020</xref>; <xref ref-type="bibr" rid="B81">Malo et al., 2024</xref>; <xref ref-type="bibr" rid="B8">Bertalm&#x000ED;o et al., 2024</xref>). Our proposal follows the tradition of previous (too simple) low-level datasets such as the OSA ModelFest initiative (<xref ref-type="bibr" rid="B20">Carney et al., 1999</xref>), but low-level psychophysics has not been extensively included in the (today&#x00027;s popular) BrainScore (<xref ref-type="bibr" rid="B119">Schrimpf et al., 2018</xref>), nor in the high-level criticisms made by <xref ref-type="bibr" rid="B12">Bowers et al. (2023</xref>) and <xref ref-type="bibr" rid="B10">Biscione et al. (2024</xref>).</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Our proposal: a low-level vision Turing Test for deep-nets</title>
<sec>
<label>3.1</label>
<title>The Decalogue: facts and foundations</title>
<p>The set of facts and associated stimuli included in our proposal is summarized in <xref ref-type="table" rid="T1">Table 1</xref>. Among the rich literature on low-level visual psychophysics, the selection of those specific properties is grounded in two main reasons.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Properties of human vision (and associated stimuli) of our Decalogue that are behind the current understanding of the information bottleneck happening between the retina and the V1 cortex.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th/>
<th valign="top" align="left"><bold>Facts / properties</bold></th>
<th valign="top" align="left"><bold>Stimuli</bold></th>
<th valign="top" align="left"><bold>Modality</bold></th>
<th valign="top" align="left"><bold>Response</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="left">Spectral sensitivities (achromatic and opponent)</td>
<td valign="top" align="left">Quasi-spectral</td>
<td valign="top" align="left">Color</td>
<td valign="top" align="left">Linear</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="left">Brightness &#x00026; color response saturation</td>
<td valign="top" align="left">Color calibrated</td>
<td valign="top" align="left">Color</td>
<td valign="top" align="left">Non-linear</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">Achromatic contrast sensitivity (bandwidth)</td>
<td valign="top" align="left">Achrom. Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Linear</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="left">Chromatic contrast sensitivity (bandwidth)</td>
<td valign="top" align="left">Chrom. Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Linear</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">Spatio-chromatic receptive fields</td>
<td valign="top" align="left">Deltas / noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Linear</td>
</tr>
<tr>
<td valign="top" align="left">6</td>
<td valign="top" align="left">Non-linear contrast response: saturation</td>
<td valign="top" align="left">Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Non-linear</td>
</tr>
<tr>
<td valign="top" align="left">7</td>
<td valign="top" align="left">Non-linear contrast response: freq. dependent</td>
<td valign="top" align="left">Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Non-linear</td>
</tr>
<tr>
<td valign="top" align="left">8</td>
<td valign="top" align="left">Context effects: energy</td>
<td valign="top" align="left">Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Non-linear</td>
</tr>
<tr>
<td valign="top" align="left">9</td>
<td valign="top" align="left">Context effects: frequency</td>
<td valign="top" align="left">Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Non-linear</td>
</tr>
<tr>
<td valign="top" align="left">10</td>
<td valign="top" align="left">Context effects: Orientation</td>
<td valign="top" align="left">Gabors/noise</td>
<td valign="top" align="left">Texture</td>
<td valign="top" align="left">Non-linear</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>They include color, texture, and motion processing abilities of human early vision. In the original literature, the stimuli were specifically designed to probe the linear or the non-linear behavior of the system.</p>
</table-wrap-foot>
</table-wrap>
<p><bold>First</bold>, they describe the visual information adaptively captured (and discarded) by the front end of human vision. On the one hand, linear sensitivities describe the spectral, chromatic, and spatio-temporal bandwidth and relative weight given by the system to the frequency components of the input stimuli. This linear description in terms of sensitivity filters is the first-order approximation to the visual bottleneck. More interestingly, this bottleneck is adaptive: in classical models of vision science, extra non-linear mechanisms are proposed between the linear filters to account for the adaptive responses to the specific eigen-stimuli of the linear filters. The stimuli in the tests we compile here were specifically designed to probe those linear and non-linear mechanisms of human vision. The power and relevance of the selected stimuli for a complete characterization of the low-level bottleneck of image-computable models is suggested by the fact that, for decades, the straightforward use of these facts (with minor or no optimization at all) led to competitive image (<xref ref-type="bibr" rid="B135">Wallace, 1991</xref>; <xref ref-type="bibr" rid="B93">Malo et al., 1995</xref>, <xref ref-type="bibr" rid="B82">2000a</xref>; <xref ref-type="bibr" rid="B129">Taubman and Marcellin, 2013</xref>; <xref ref-type="bibr" rid="B79">Malo et al., 2006</xref>) and video coding algorithms (<xref ref-type="bibr" rid="B64">Le Gall, 1992</xref>; <xref ref-type="bibr" rid="B83">Malo et al., 2000b</xref>,<xref ref-type="bibr" rid="B87">c</xref>, <xref ref-type="bibr" rid="B86">2001</xref>) and distortion metrics (<xref ref-type="bibr" rid="B23">Daly, 1993</xref>; <xref ref-type="bibr" rid="B108">Null, 1993</xref>; <xref ref-type="bibr" rid="B130">Teo and Heeger, 1994</xref>; <xref ref-type="bibr" rid="B94">Malo et al., 1997</xref>; <xref ref-type="bibr" rid="B141">Watson and Malo, 2002</xref>; <xref ref-type="bibr" rid="B62">Laparra et al., 2010</xref>) equipped with color constancy and contrast adaptation (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>; <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>; <xref ref-type="bibr" rid="B42">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2024</xref>). Checking if the response of a network is human-like for those stimuli would imply that the bottleneck of the network would have <italic>statistically</italic> good adaptive behavior (<xref ref-type="bibr" rid="B109">Olshausen and Field, 1996</xref>; <xref ref-type="bibr" rid="B121">Schwartz and Simoncelli, 2001</xref>; <xref ref-type="bibr" rid="B85">Malo and Guti&#x000E9;rrez, 2006</xref>; <xref ref-type="bibr" rid="B91">Malo and Laparra, 2010</xref>; <xref ref-type="bibr" rid="B59">Laparra et al., 2012</xref>; <xref ref-type="bibr" rid="B61">Laparra and Malo, 2015</xref>; <xref ref-type="bibr" rid="B32">Gomez-Villa et al., 2020a</xref>; <xref ref-type="bibr" rid="B76">Malo, 2020</xref>, <xref ref-type="bibr" rid="B77">2022</xref>). This adaptivity could be useful for the mentioned applications and also for domain adaptation (<xref ref-type="bibr" rid="B122">Sengar et al., 2025</xref>).</p>
<p><bold>Second</bold>, effects elicited by the selected stimuli are visually compelling and hence, the user of the test can check (by the eye) if the model under consideration behaves like humans or not. On the one hand, sensitivity surfaces to simple (isolated) stimuli are standardized and ready for direct quantitative comparison (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>; <xref ref-type="bibr" rid="B49">Hurvich and Jameson, 1957</xref>; <xref ref-type="bibr" rid="B53">Krauskopf and Gegenfurtner, 1992</xref>; <xref ref-type="bibr" rid="B17">Campbell and Robson, 1968</xref>; <xref ref-type="bibr" rid="B107">Mullen, 1985</xref>; <xref ref-type="bibr" rid="B31">Georgeson and Sullivan, 1975</xref>; <xref ref-type="bibr" rid="B22">Daly, 1990</xref>; <xref ref-type="bibr" rid="B94">Malo et al., 1997</xref>; <xref ref-type="bibr" rid="B52">Kelly, 1979</xref>; <xref ref-type="bibr" rid="B24">D&#x000ED;ez-Ajenjo et al., 2011</xref>). On the other hand, as illustrated below, non-linear responses when using stimuli in a context (under adaptation) have specific qualitative behaviors that are easy to see (<xref ref-type="bibr" rid="B29">Foley, 1994</xref>; <xref ref-type="bibr" rid="B140">Watson and Solomon, 1997</xref>; <xref ref-type="bibr" rid="B102">Martinez-Uriegas, 1997</xref>). In this way, that eventual model deviations from human-like behavior are easy to detect. Moreover, Section 3.3 proposes ways to summarize these visual behaviors in a numerical score. The proposed tests do not give definitive answers to the points raised in Section 2, but they are useful to stress the impact of those issues in easy-to-view ways and rule out models accordingly.</p>
</sec>
<sec>
<label>3.2</label>
<title>The Decalogue: specific examples</title>
<p>In this section, we show four examples of the proposed Decalogue with series of calibrated stimuli (from the colorimetric and the spatial perspectives) that illustrate the non-linear response of humans to (i) luminance in different backgrounds leading to different perceptions of <italic>brightness</italic>, (ii) deviations in opponent color directions under different induction conditions leading to different perceptions of <italic>hue</italic> and <italic>saturation</italic>, (iii) texture masking due to the energy of the background, and (iv) texture masking due to the similarity between the features of the background and the test.</p>
<p>It is important to note that the properties illustrated here (properties 2, 8, 10) are examples of curves that are not standardized, as opposed to other facts in the proposed Decalogue (properties 1, 3, 4, 6, 7), in which strict comparisons (RMSE or Pearson correlation) are possible. The fact that, even in these non-standardized examples, the qualitative behavior is so compelling implies that checking the order of the curves using rank correlations is useful to quantitatively describe the alignment between artificial models and humans.</p>
<sec>
<label>3.2.1</label>
<title>Luminance and brightness</title>
<p>The first set of stimuli refers to a series of luminance-calibrated achromatic samples that illustrate the perception of brightness in backgrounds of different luminance. They illustrate the Weber law (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>; <xref ref-type="bibr" rid="B28">Fairchild, 2013</xref>) and the crispening effect (<xref ref-type="bibr" rid="B142">Whittle, 1992</xref>), i.e., the achromatic part of Property 2 in <xref ref-type="table" rid="T1">Table 1</xref>. These effects have been related to the statistics of natural images (<xref ref-type="bibr" rid="B63">Laughlin, 1981</xref>; <xref ref-type="bibr" rid="B59">Laparra et al., 2012</xref>) and with sophisticated models of retinal adaptation (<xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al., 2020</xref>).</p>
<p><xref ref-type="fig" rid="F5">Figure 5</xref> shows a series of these stimuli in the (linearly spaced) range of luminance [0.5, 120] <italic>cd</italic>/<italic>m</italic><sup>2</sup> on (linearly spaced) backgrounds of luminance in the range [1, 160] <italic>cd</italic>/<italic>m</italic><sup>2</sup>. These stimuli are easily generated in digital levels (i.e., ready to feed conventional artificial models) with the code provided in this work<xref ref-type="fn" rid="fn0007"><sup>5</sup></xref>, which makes use of the calibration of the Matlab toolbox Colorlab (<xref ref-type="bibr" rid="B92">Malo and Luque, 2002</xref>) on a standard computer screen.</p>
<fig position="float" id="F5">
<label>Figure 5</label>
<caption><p>Series of stimuli eliciting non-linear brightness perception. Luminance-calibrated linearly spaced tests in different luminance-calibrated backgrounds. This illustrates the Weber law (<xref ref-type="bibr" rid="B28">Fairchild, 2013</xref>) as well as Whittle&#x00027;s crispening effect (<xref ref-type="bibr" rid="B142">Whittle, 1992</xref>), as summarized in <xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al. (2020</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0005.tif">
<alt-text>Five rows of small rectangles on the left vary in luminance, each highlighted at levels of 1, 40.5, 60.5, 120.5, and 160 cd/m2. A graph on the right shows luminance response curves, with colored lines corresponding to the rectangles. Circles highlight specific points on each curve.</alt-text>
</graphic>
</fig>
<p>Let&#x00027;s describe the perceived brightness of the stimuli in this test.</p>
<p><bold>First</bold>, the series of stimuli in the darkest background clearly shows the saturation non-linearity of the brightness vs luminance curve: note that the jumps in perceived brightness for the low-luminance tests are distinctly bigger than the equivalent jumps for the same increments in luminance at the high-luminance end. In the axis of perceived brightness, the above implies that the response (blue curve) has a large slope (high sensitivity) at the low-luminance end and a saturation of such response (lower sensitivity) at the high-luminance end. That makes the <italic>qualitative</italic> saturating blue curve of brightness vs luminance.</p>
<p><bold>Second</bold>, when one increases the luminance of the background (e.g., from 1 <italic>cd</italic>/<italic>m</italic><sup>2</sup> to 40 <italic>cd</italic>/<italic>m</italic><sup>2</sup>), the brightness of the (same) samples is lower than in the previous series, so the <italic>qualitative</italic> brightness response to this second series of stimuli is below the previous one (as depicted by the <italic>qualitative</italic> black curve).</p>
<p><bold>Third</bold>, by looking at the stimuli highlighted in gray in the 1 <italic>cd</italic>/<italic>m</italic><sup>2</sup> and the 40 <italic>cd</italic>/<italic>m</italic><sup>2</sup> backgrounds, it is obvious that in the brighter background, the stimuli with equivalent brightness are shifted to the right on the scale of luminance, which means that the response in black (for the stimuli in the brighter background) is shifted right-down with regard to the curve in blue (for the stimuli in the darker background). Moreover, this means that the black curve has a sigmoidal shape as it should start from zero brightness. Similar visual reasoning implies that this shift progressively increases as one increases the luminance of the background, as <italic>qualitatively</italic> illustrated by the samples highlighted in gray along the diagonal of the panel with the stimuli (left), leading to the shift in the curves (right).</p>
<p><bold>Fourth</bold>, the (same) stimuli in the brightest background elicit a brightness response with a substantially different shape: the sigmoid has substantially shifted to the right (red curve), and, all in all, one can see a smooth transition of the sigmoidal response curves from the blue curve to the red curve. The crispening effect (increased sensitivity around backgrounds of similar luminance) is illustrated by the shift to the right of the points of maximum slope in the response curves.</p>
<p>Finally, <bold>fifth</bold>, the decreasing brightness of the samples of the same luminance in backgrounds of progressively greater luminance (as illustrated by the samples highlighted in green) illustrates brightness induction (<xref ref-type="bibr" rid="B28">Fairchild, 2013</xref>).</p>
<p>Of course, the <italic>qualitative</italic> visual observations made here do not try to substitute for the rich <italic>quantitative</italic> literature in which these responses are determined by accurate psychophysics (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>; <xref ref-type="bibr" rid="B28">Fairchild, 2013</xref>). However, (1) the phenomena are compelling enough that one can see the qualitative trends of the curves by eye, and, as seen in the numerical experiments below, (2) these trends (visible in ready-to-use digital images) are enough to spot divergences with human behavior in certain artificial models or discriminate between models in terms of their similarity to human behavior, which is the ultimate goal of the tests presented here.</p></sec>
<sec>
<label>3.2.2</label>
<title>Non-linear response to saturation and color adaptation</title>
<p>Responses to constant deviations from white in the red-green and yellow-blue directions of the Jameson &#x00026; Hurvich color space (<xref ref-type="bibr" rid="B49">Hurvich and Jameson, 1957</xref>; <xref ref-type="bibr" rid="B134">Vila-Tom&#x000E1;s et al., 2023</xref>) with equiluminant stimuli describe the non-linear perception of hue and saturation, as pointed out in <xref ref-type="bibr" rid="B53">Krauskopf and Gegenfurtner (1992</xref>) and <xref ref-type="bibr" rid="B116">Romero et al. (1993</xref>) in similar opponent spaces, i.e., the chromatic version of property 2 in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> shows colorimetrically calibrated stimuli with such deviations (in the range [-20, 20] of the linear RG and YB tristimulus values of the Jameson and Hurvich space) in different backgrounds, which are easy to generate and modify by using the code provided in this work<xref ref-type="fn" rid="fn0008"><sup>6</sup></xref>. As in the previous test, let&#x00027;s describe the perceived hue and saturation of the stimuli to infer the qualitative shape of the responses.</p>
<fig position="float" id="F6">
<label>Figure 6</label>
<caption><p>Series of non-linear perceived saturation (or response of the opponent chromatic channels) vs. linearly spaced increments in colorimetrically calibrated color opponent directions in different chromatic contexts. This illustrates the non-linear effects pointed out in <xref ref-type="bibr" rid="B53">Krauskopf and Gegenfurtner (1992</xref>) and <xref ref-type="bibr" rid="B116">Romero et al. (1993</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0006.tif">
<alt-text>Two side-by-side graphs show color perception data. The left graph depicts RG (red-green) contrasts with circular markers over horizontal bars transitioning from green to red. The right graph depicts YB (yellow-blue) contrasts with circular markers over horizontal bars transitioning from blue to orange. Both graphs have vertical scales from &#x02013;21.5 to 21.5 and horizontal scales from &#x02013;20 to 20, showing varied patterns.</alt-text>
</graphic>
</fig>
<p><bold>First</bold>, take the stimuli in gray backgrounds and note that the jumps in perceived hue are bigger around the central (achromatic) stimuli than at the extremes with more saturated stimuli (either red, green, yellow, or blue): Judge the jumps in saturation close to the achromatic stimulus and at the extremes of the chromatic axes. Similarly to the responses for brightness, these differences imply a sigmoidal response to saturation when the stimuli linearly depart from white in constant steps: see the qualitative responses in gray for both the red-green and the yellow-blue directions. <bold>Second</bold>, these sigmoidal responses shift to the right or to the left, as can be seen from the shift of the stimuli that are perceived as achromatic in the different backgrounds (e.g., see the stimuli highlighted in gray). Note that a stimulus is seen as achromatic when the response of the mechanism tuned to red-green or yellow-blue is zero. See the corresponding shifts in the zero crossings of the sigmoids (also highlighted in gray). Finally, <bold>third</bold>, the shift of the responses is bigger as the saturation of the background is increased.</p>
<p>Again, the goal of this test is not to substitute the original accurate psychophysics done on humans (<xref ref-type="bibr" rid="B53">Krauskopf and Gegenfurtner, 1992</xref>; <xref ref-type="bibr" rid="B116">Romero et al., 1993</xref>) to point out these phenomena. On the contrary, they just represent an easy way to get digital images that can be used to test artificial models and check if their responses qualitatively behave like humans.</p></sec>
<sec>
<label>3.2.3</label>
<title>Texture masking 1 (energy): non-linear adaptive contrast response</title>
<p>The same kind of qualitative derivation of human-like responses can be applied to the perceived contrast of textured patterns with calibrated frequency content and controlled luminance. The test presented here illustrates the fact that perceived contrast non-linearly depends on linearly increasing Michelson contrast (<xref ref-type="bibr" rid="B66">Legge and Foley, 1980</xref>; <xref ref-type="bibr" rid="B65">Legge, 1981</xref>) and this response decreases with (is masked by) the energy of a background of similar texture (<xref ref-type="bibr" rid="B29">Foley, 1994</xref>; <xref ref-type="bibr" rid="B140">Watson and Solomon, 1997</xref>). This corresponds to property 8 in <xref ref-type="table" rid="T1">Table 1</xref>. The stimuli presented in the following example can be reproduced and modified both in frequency orientation, average luminance, and contrast with the code provided<xref ref-type="fn" rid="fn0009"><sup>7</sup></xref>.</p>
<p><xref ref-type="fig" rid="F7">Figure 7</xref> shows Gaussian-windowed test noise patches of 4 cycles/degree (cpd) in images subtending 1 degree with an average luminance of 50 <italic>cd</italic>/<italic>m</italic><sup>2</sup> and linearly spaced RMSE contrasts (from left to right) in the range [0, 0.3]. The different rows show the same tests on different backgrounds of noise of 4 cpd with linearly spaced RMSE contrast in the range [0, 0.25].</p>
<fig position="float" id="F7">
<label>Figure 7</label>
<caption><p>Series of non-linearly perceived contrast (or response of the mechanisms tuned to certain texture) vs. linearly spaced increments in contrast-calibrated test of controlled spatial frequency in backgrounds of different (controlled) energy. This illustrates the non-linear effects pointed out in <xref ref-type="bibr" rid="B31">Georgeson and Sullivan (1975</xref>), <xref ref-type="bibr" rid="B66">Legge and Foley (1980</xref>), <xref ref-type="bibr" rid="B65">Legge (1981</xref>), and <xref ref-type="bibr" rid="B22">Daly (1990</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0007.tif">
<alt-text>Grid of wavy lines with varying densities on the left, labeled with numerical values. Circles overlay certain patterns in blue and yellow. On the right, a graph plots RMSE contrast against an axis, with colored lines corresponding to circled patterns.</alt-text>
</graphic>
</fig>
<p>Similarly to the previous cases, let us describe the perceived contrast along the two dimensions of the panel: test: left to right, and background: top to bottom. Again, the qualitative shape of the responses will be determined by the perceived jumps of contrast of the tests (from left to right) and by their variation as one increases the energy of the background (from top to bottom).</p>
<p><bold>First</bold>, for the zero-contrast background (first, top row), the jumps in perceived contrast in the low-contrast end (left) are bigger than the jumps in perceived contrast in the high-contrast end (right). See the differences in perceived contrast in the tests highlighted in blue. This implies a saturating contrast response curve (as in the previous examples), i.e., the blue curve.</p>
<p><bold>Second</bold>, as the contrast of the background is increased (see stimuli highlighted in orange), the perceived contrast of the test is reduced. This implies that subsequent curves (black and lighter shades of gray) are below the initial blue curve.</p>
<p>Finally, <bold>third</bold>, in order to perceive the tests with equivalent contrasts in backgrounds of progressively greater energy, the necessary contrast of the test increases; this means that the sigmoidal curves shift to the right.</p>
<p>As in the previous examples, the qualitative behavior illustrated by this series of digital images generated by our code should give (in artificial models) corresponding saturating curves with smooth variation from the blue (zero-contrast background) condition to the red (high-contrast background) condition, and hence the lower response curve.</p></sec>
<sec>
<label>3.2.4</label>
<title>Texture masking 2 (features): interaction between orientations</title>
<p>Reduction of sensitivity (the so-called masking) also happens when a certain test is presented on top of a background that shares some feature with the test (<xref ref-type="bibr" rid="B117">Ross et al., 1991</xref>; <xref ref-type="bibr" rid="B29">Foley, 1994</xref>; <xref ref-type="bibr" rid="B140">Watson and Solomon, 1997</xref>), i.e., properties 9 and 10 in <xref ref-type="table" rid="T1">Table 1</xref>. The next example, <xref ref-type="fig" rid="F8">Figure 8</xref>, refers to the specific case of interaction between orientations of test and background. It can be reproduced and modified in terms of frequency, orientation, contrast, and average luminance with the code provided<xref ref-type="fn" rid="fn0010"><sup>8</sup></xref>.</p>
<fig position="float" id="F8">
<label>Figure 8</label>
<caption><p>Series of non-linearly perceived contrast (or response of the mechanisms tuned to certain texture) vs. linearly spaced increments in contrast-calibrated tests of controlled spatial frequency in backgrounds of different orientations. This illustrates the non-linear effects pointed out in <xref ref-type="bibr" rid="B29">Foley (1994</xref>) and <xref ref-type="bibr" rid="B140">Watson and Solomon (1997</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0008.tif">
<alt-text>Grid of small images with varying orientations and contrasts is shown on the left. Circles in different colors highlight specific images. A line graph on the right shows RMSE contrast plotted against a variable, with lines corresponding to the highlighted images. Each line matches the color of the circles in the grid.</alt-text>
</graphic>
</fig>
<p><xref ref-type="fig" rid="F8">Figure 8</xref> shows 6 cpd horizontal Gabor patches with an average luminance of 50 <italic>cd</italic>/<italic>m</italic><sup>2</sup> and RMSE contrast increasing linearly from left to right in the range [0, 0.3]. These Gabor patches are shown on top of band-pass noise of contrast 0.2, with the same frequency, but different orientation. The numbers in the different rows show the angular difference between tests and backgrounds. The figure shows some compelling facts that lead to clear qualitative trends in the response curves.</p>
<p><bold>First</bold>, the test is better seen (has bigger visibility or perceived contrast) when the background is orthogonal to the test (in the first row). In that row, the different jumps in visibility in the low-contrast and high-contrast ends (tests highlighted in cyan) indicate a saturating response, as in the previous examples (response curve in blue).</p>
<p><bold>Second</bold>, the necessary contrast to detect the test smoothly increases as the difference in orientation between test and background decreases: see that the tests highlighted in blue, black, shades of gray, and red approximately have the same visibility over the different backgrounds with angular differences in the range [90,0] deg. The trend is similar for negative angular differences. This implies a smooth variation (decrease) of the response curves in terms of the difference between test and background.</p>
<p><bold>Third</bold>, the biggest masking is obtained when test and background are aligned (the red curve is clearly the lowest response curve). This and the previous fact imply that the general trend is this smooth transition of the non-linear curves from the blue to the red.</p>
</sec>
</sec>
<sec>
<label>3.3</label>
<title>Proposed methodology</title>
<p>Our proposal is simple: use the code provided here<xref ref-type="fn" rid="fn0011"><sup>9</sup></xref> to generate the stimuli (digital images well-calibrated in luminance, color, and spatio-temporal frequency) that illustrate the 10 compelling properties listed in <xref ref-type="table" rid="T1">Table 1</xref> describing the adaptive information bottleneck of low-level human vision. The resulting digital images are organized in series that correspond to progressive stimulation of a vision system in particular ways. The interesting point is that this set of controlled stimulation conditions leads to intuitive responses (as shown above), or even to standardized sensitivity curves or surfaces that are also provided with the code see Section 1 of <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
<p>Once stimuli are generated, they are used to feed any artificial image-computable model. Then, depending on the model, the user decides where to read from the network under consideration and the read-out mechanism to get <italic>visibility</italic> values to generate artificial series of response curves.</p>
<p>The resulting curves can be qualitatively assessed by checking the correspondence of their shape with our own visual experience as in the examples above. However, in order to simplify the (blind-quantitative) use by the non-expert, we propose two different correlation measures for the cases where experimental ground truth is (or is not) available:</p>
<list list-type="bullet">
<list-item><p><bold>Pearson Correlation:</bold> in the case of properties where clear ground truth is available, we propose to quantify the alignment using Pearson correlation between the ground truth and the predictions of the models. Pearson correlation is insensitive to a global (arbitrary) scale of the response, and this is generally acceptable because a single global scale is not important. In order to reduce the relevance of this arbitrary scale, we propose to evaluate a single Pearson correlation measure for groups of comparable curves (for instance, response curves for different frequencies but the same contrast stimulation&#x02013;e.g., in Props. 6&#x02013;8-, or curves with known scaling between the achromatic and chromatic response&#x02013;e.g., in spectral sensitivities in Prop. 1, CSFs in Props. 3 and 4, or scale of achromatic and chromatic curves in Props. 6 and 7). In those cases, by evaluating the Pearson correlation for groups of curves, it captures if they are scaled with the proper relative size (which should be reproduced by the models). Using the ground truth curves shown in <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref> (and associated code), the Pearson correlation can be computed for about 70% of the proposed tests (Props. 1, 3&#x02013;7, the cases with no adaptation in Prop. 2, and also the no-masking curves in Props. 8&#x02013;10).</p></list-item>
<list-item><p><bold>Rank correlation:</bold> In the cases of brightness and chromatic adaptation (Prop. 2) and cases where there is spatial masking with the same stimulus (Prop. 8) or with different stimuli (Props. 9 and 10), direct visualization of the tests shows the trend for the variation of the curves (mainly shifting and attenuation of the curves when increasing the strength of the adaptor&#x02013;either luminance or saturation of the background or illuminant or contrast of the masker). The specialized literature describes these trends, but in too sparse (or not comparable) situations that are not properly captured by the (ready for ANNs) stimuli in the database. In this situation, Pearson correlation is not applicable. However, the relative order (or rank) between the curves for different strengths of the adaptor is known. Reproduction of this known rank of curves at several locations of the abscissas is possible using rank correlation (e.g., Spearman or Kendall correlations)to obtain a quantitative descriptor per Property. Although necessary in cases where no experimental data is available, this kind of descriptor of &#x0201C;ranking of curves&#x0201D; can also be used in cases with ground truth because Pearson correlation may miss (or not clearly capture) the rank between the curves, which usually is a qualitative trend that should be reproduced by the models.</p></list-item>
</list>
<p>Choosing an appropriate layer to read from and a read-out mechanism can be a crucial step because, as it happens in the human-visual system, certain behaviors are expected to be happening at different points of the visual processing pipeline. BrainScore (<xref ref-type="bibr" rid="B119">Schrimpf et al., 2018</xref>) proposes using an independent test set to evaluate how each layer of an artificial model matches to certain areas of the brain. Then, it&#x00027;s only a matter of reading from the layer that has a better match with the behavior we want to measure. As a different example, within this work we employ models that have been designed to accommodate certain parts of the human visual system at certain layers, so choosing the appropriate layer to read from is straightforward. We also provide a set of control experiments (<xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>; Section 2) that showcase how does the read-out mechanism affect the obtained results so that the interested reader can make an informed choice.</p>
<p>In the case of properties 2, 8, 9 and 10 these curves have to be compared with the kind of qualitative curves described above, which given the clarity of the selected stimuli can be drawn by simple visual observation of the stimuli as described above.</p>
<p>In the case of property 5 [existence of center-surround and Gabor-like receptive fields tuned to achromatic, red-green and yellow-blue patterns (<xref ref-type="bibr" rid="B123">Shapley and Hawken, 2011</xref>)], the more straightforward method is checking their presence by reading the response to deltas from single neurons or from the Jacobian of the network at that layer (<xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>; <xref ref-type="bibr" rid="B33">Gomez-Villa et al., 2020b</xref>; <xref ref-type="bibr" rid="B68">Li et al., 2022</xref>). Other indirect methods could be (1) using reverse correlation feeding the network with controlled noise [also generable using Vistalab (<xref ref-type="bibr" rid="B84">Malo and Gutierrez, 2002</xref>) following the appropriate literature (<xref ref-type="bibr" rid="B27">Eckstein and Ahumada, 2002</xref>)], or (2) using artificial psychophysics based on adaptation [e.g., the Blakemore and Campbell experiment (<xref ref-type="bibr" rid="B11">Blakemore and Campbell, 1969</xref>)]. However, this very last method to measure proeprty 5 relies on fulfillment of adaptation proeprties 6&#x02013;10, which may not hold in non-human networks.</p>
<p>In the above (non-standardized) cases the general trends of the curves can be qualitatively assessed in detail: general shape of the curves, the blue response and the red curve being the biggest and the lowest, respectively, and the transition from one to the other. Note that user of the provided code can change the parameters of the stimuli and infer new curves by applying a similar visual analysis. For the receptive fields they can be analyzed using shape parameters in the spatial or the Fourier domain as classically done in visual neuroscience (<xref ref-type="bibr" rid="B114">Ringach, 2002</xref>; <xref ref-type="bibr" rid="B115">Ringach and Shapley, 2004</xref>; <xref ref-type="bibr" rid="B101">Martinez et al., 2017</xref>; <xref ref-type="bibr" rid="B72">Loxley, 2017</xref>) and the same for the chromatic tuning in standard color spaces (<xref ref-type="bibr" rid="B128">Tailby et al., 2008</xref>; <xref ref-type="bibr" rid="B38">Gutmann et al., 2014</xref>; <xref ref-type="bibr" rid="B33">Gomez-Villa et al., 2020b</xref>).</p>
<p>Finally in the case of sensitivity curves or surfaces which are standardized or available in the code (proeprties 1, 3, 4, 6 and 7) the visibility values obtained from the models can be numerically compared with the provided ground truth.</p>
<p>The above quantitative descriptors clearly display some limitations. First, they obviously overlook relevant qualitative properties of the results. For example: (a) Correlations do not capture the change of slope happening in crispening and chromatic adaptation in Prop. 2. This change of slopes is relevant to decide between models (<xref ref-type="bibr" rid="B59">Laparra et al., 2012</xref>; <xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al., 2020</xref>). More complicated descriptors of the qualitative shape of these specific curves could be developed, but they are out of the scope of this work. (b) Assessment of receptive fields in certain layers is done by the eye in many influential works (<xref ref-type="bibr" rid="B109">Olshausen and Field, 1996</xref>; <xref ref-type="bibr" rid="B7">Bell and Sejnowski, 1997</xref>; <xref ref-type="bibr" rid="B50">Hyvarinen et al., 2009</xref>; <xref ref-type="bibr" rid="B55">Krizhevsky et al., 2012a</xref>) that report the emergence of Gabor-like receptive fields and their similarity with biological receptive fields in V1 just by inspection of the filters. Analysis of the spatio-frequency properties and chromatic properties of the receptive fields, as in <xref ref-type="bibr" rid="B38">Gutmann et al. (2014</xref>); <xref ref-type="bibr" rid="B101">Martinez et al. (2017</xref>); <xref ref-type="bibr" rid="B72">Loxley (2017</xref>); <xref ref-type="bibr" rid="B33">Gomez-Villa et al. (2020b</xref>) comparing with (<xref ref-type="bibr" rid="B114">Ringach, 2002</xref>) and <xref ref-type="bibr" rid="B128">Tailby et al. (2008</xref>), is also possible but more unusual.</p>
<p>Second, the derivation of a single score per model as the average of scores over experiments [as the compromise solution taken in BrainScore (<xref ref-type="bibr" rid="B119">Schrimpf et al., 2018</xref>)] is arguable. The selection of properties to average over is always somewhat arbitrary, and alternative selections lead to global scores with different variances (and discriminative power). <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref> illustrates this limitation by showing bootstrap examples in the set of quantitative results shown below.</p>
<p>This qualitative/quantitative methodology is summarized in <xref ref-type="fig" rid="F9">Figure 9</xref> and applied in the next experimental section for three illustrative networks.</p>
<fig position="float" id="F9">
<label>Figure 9</label>
<caption><p>The proposed method: feed the model with a series of images, compute responses (using the simplest possible read-out mechanism), and make quantitative comparisons with standard sensitivity surfaces or qualitative comparisons checking the non-linearity using different adaptation conditions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0009.tif">
<alt-text>Diagram depicting a process involving image analysis and comparison through a neural network. Images labeled &#x00022;Images to feed the network&#x00022; are processed by &#x00022;YOUR NET&#x00022;, producing a graph. Outputs are compared for quality scores (red, yellow, green icons) and quantitative scores, including Pearson correlation with ground truth and Kendall order of curves. Charts represent contrast versus response.</alt-text>
</graphic>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments: analysis of three illustrative deep models</title>
<sec>
<label>4.1</label>
<title>Networks and experimental setting</title>
<p>In our experiments, we check the behavior on the proposed <italic>Decalogue</italic> of three recent networks of similar architecture:</p>
<list list-type="order">
<list-item><p>A parametric vision model, the <italic><bold>BioMultiLayer</bold></italic> network (<xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>), which consists of a cascade of four linear&#x0002B;non-linear stages that account for (1) color opponency and adaptation, (2) contrast computation, (3) contrast sensitivities and energy masking, and (4) wavelet analysis and cross-masking between textures. The linear parts of all the stages were not optimized, but they were directly inspired by classical psychophysical or physiological literature. The non-linear parts were implemented via divisive normalization (<xref ref-type="bibr" rid="B18">Carandini and Heeger, 1994</xref>; <xref ref-type="bibr" rid="B79">Malo et al., 2006</xref>; <xref ref-type="bibr" rid="B62">Laparra et al., 2010</xref>; <xref ref-type="bibr" rid="B81">Malo et al., 2024</xref>; <xref ref-type="bibr" rid="B19">Carandini and Heeger, 2012</xref>). The non-linearities of the 2nd and 3rd stages of the model were tuned via the psychophysical method of maximum differentiation in <xref ref-type="bibr" rid="B95">Malo and Simoncelli (2015</xref>). The non-linear parts of the 1st and 4th stages were tuned to reproduce subjective opinions on distortion and contrast masking facts (<xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>, <xref ref-type="bibr" rid="B99">2019</xref>). The statistical properties of the model and its relations with recurrent models were studied in <xref ref-type="bibr" rid="B32">Gomez-Villa et al. (2020a</xref>) and <xref ref-type="bibr" rid="B81">Malo et al. (2024</xref>), respectively.</p></list-item>
<list-item><p>A non-parametric model to predict subjective image quality, the <italic><bold>PerceptNet</bold></italic> (<xref ref-type="bibr" rid="B40">Hepburn et al., 2020</xref>), which starts with a non-linear front-end at the retina followed by a cascade of three linear&#x0002B;non-linear stages. The architecture was intended to accommodate similar vision facts that motivated the <italic>BioMultiLayer</italic>. The <italic>PerceptNet</italic> architecture is similar to AlexNet (<xref ref-type="bibr" rid="B56">Krizhevsky et al., 2012b</xref>) but its non-linearities were formulated using an end-to-end optimizable divisive normalization (<xref ref-type="bibr" rid="B58">Laparra et al., 2017</xref>; <xref ref-type="bibr" rid="B3">Ball&#x000E9; et al., 2017</xref>). Both the linear and the non-linear parts of <italic>PerceptNet</italic> were end-to-end tuned to maximize the correlation with humans on subjective image distortions (<xref ref-type="bibr" rid="B40">Hepburn et al., 2020</xref>). Non-parametric layers of <italic>PerceptNet</italic> are not easy to interpret, as pointed out recently (<xref ref-type="bibr" rid="B133">Vila-Tom&#x000E1;s et al., 2025</xref>).</p></list-item>
<list-item><p>An image segmentation model, the <italic><bold>Bio U-Net</bold></italic> (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>), with the same style encoder as the non-parametric <italic>PerceptNet</italic> (a cascade of linear &#x0002B; divisive normalization stages), but augmented with a decoder that recovers the original dimension of the input signal and predicts a class per pixel for semantic segmentation. The encoder and the decoder were tuned to optimize segmentation in different databases. The benefits of the biologically inspired non-linearities of this model for segmentation have been further studied thereafter (<xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>).</p></list-item>
</list>
<p>We assumed a visual field of 2 degrees with a sampling frequency of 64 cycles/deg, i.e., we fed the models with 128 &#x000D7; 128 images. We measured the responses of the models to specific tests through the Euclidean departure between the response to test&#x0002B;background with regard to the response to the isolated background.</p>
</sec>
<sec>
<label>4.2</label>
<title>Results I: qualitative analysis</title>
<p>In this section, we qualitatively assess the results by checking the shape and relative scales of the curves obtained by the models in relation to the observations made in Section 3.2. In the next section, we apply the quantification of the alignment proposed in Section 3.3.</p>
<sec>
<label>4.2.1</label>
<title>Spectral sensitivities and color responses (properties 1 and 2)</title>
<p><xref ref-type="fig" rid="F10">Figure 10</xref>-top shows the response of the models to quasi-monochromatic stimuli<xref ref-type="fn" rid="fn0012"><sup>10</sup></xref> to get the spectral sensitivity of the neurons (property 1). In order to point out the relevance of the layer from which responses are measured, in the case of the <italic>BioMultiLayer</italic> network, we consider direct read-out of the response (with sign) in the first linear layer (subplots A and B) and in the last non-linear layer (subplots C and D). In this network, the first linear layer has achromatic and opponent channels defined by construction, so the <italic>V</italic><sub>&#x003BB;</sub> (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>) (subplot A) and the opponent curves of Jameson &#x00026; Hurvich (<xref ref-type="bibr" rid="B49">Hurvich and Jameson, 1957</xref>) (subplot B) are trivially obtained. Interestingly, the spectral sensitivities at the last non-linear layer are wide-band positive in the first channel of the network and opponent in the other two channels, but their shapes are substantially modified with regard to the human-like behavior at the first layer. These differences and the uneven relative scaling between the achromatic and chromatic responses justify the qualitative scores given in each case. We can conclude that spectral sensitivity in this model is human-like at the front end but degrades throughout the network. In other words, as suggested in Section 2.2, read-out location matters, and a certain kind of information should be extracted from a specific place in the model.</p>
<fig position="float" id="F10">
<label>Figure 10</label>
<caption><p>Spectral sensitivities of the considered models <bold>(top)</bold> and corresponding responses to luminance and linear deviations from white in the cardinal red-green and yellow-blue directions <bold>(bottom)</bold>. Subplots <bold>(A)</bold>, <bold>(C)</bold>, <bold>(E)</bold>, and <bold>(G)</bold> display the achromatic spectral sensitivity for the models, while subplots <bold>(B)</bold>, <bold>(D)</bold>, <bold>(F)</bold>, and <bold>(H)</bold> display the chromatic sensitivity obtained through hue cancellation (<xref ref-type="bibr" rid="B134">Vila-Tom&#x000E1;s et al., 2023</xref>). Knowledge of the standard spectral sensitivity, the CIE <italic>V</italic><sub>&#x003BB;</sub> curve (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>), or the standard spectral sensitivity of the opponent channels (<xref ref-type="bibr" rid="B49">Hurvich and Jameson, 1957</xref>) indicates which model is correct in Prop. 1. Subplots <bold>(I)</bold>, <bold>(L)</bold>, <bold>(O)</bold>, and <bold>(R)</bold> show the brightness responses to luminance. Subplots <bold>(J)</bold>, <bold>(K)</bold>, <bold>(M)</bold>, <bold>(N)</bold>, <bold>(P)</bold>, <bold>(Q)</bold>, <bold>(S)</bold>, and <bold>(T)</bold> show the responses to deviations from white. In this case, the stimuli proposed here (<xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>) and the associated human behavior described above indicate the correct trends in the responses of Prop. 2.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0010.tif">
<alt-text>Chart comparing three models: BioMultiLayer, PerceptNet, and Bio-UNet across two properties. Property 1 features graphs labeled A-H, illustrating visibility and spectral data with different models&#x00027; predictions. Property 2 includes graphs I-T, showing brightness versus luminance and linear or nonlinear relationships. Models are evaluated with symbols indicating correctness, approximation, or error, denoted by checkmarks, tilde marks, and crosses, respectively.</alt-text>
</graphic>
</fig>
<p>The <italic>PerceptNet</italic> has a color space change after the retinal non-linearity. We measure the spectral sensitivity at that point because the design assumed that achromatic and chromatic channels could emerge there. Results show that the first channel displays an all-positive but bimodal response (subplot E), and for the other two channels, one of them certainly displays opponent-like responses, but the other is basically insensitive (subplot F).</p>
<p>The very same location of the encoder of the segmentation <italic>Bio-U-Net</italic> has very different sensitivities despite it having the same architecture as the <italic>PerceptNet</italic> up to that layer. The sensitivity of the (supposedly) achromatic channel is very noisy, and the other two channels are clearly non-human (subplots G and H).</p>
<p>On the other hand, <xref ref-type="fig" rid="F10">Figure 10</xref>-bottom checks property 2 by showing the responses to (i) luminance and to deviations from white in the (ii) red-green and (iii) yellow-blue directions (left, center, and right, respectively). In the achromatic case, tests in the range [0.5, 120] <italic>cd</italic>/<italic>m</italic><sup>2</sup> are shown on top of backgrounds of different luminance in the range [1, 160] <italic>cd</italic>/<italic>m</italic><sup>2</sup>. The response curves in different backgrounds are depicted in blue, black, and progressively lighter shades of gray until red, as in <xref ref-type="fig" rid="F5">Figure 5</xref>. In the chromatic cases, responses are computed with tests on an achromatic background (black curve) and on backgrounds of progressively saturated color (reddish and greenish curves, and bluish and yellowish curves as in <xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<p>For the <italic>BioMultiLayer</italic> model, we have such responses for two different layers: first (I, J, K) and fourth (L, M, and N). The achromatic response of the first layer is certainly non-linear for the darkest background, and the response gets attenuated when the luminance of the background is increased (see the transition from curves in blue to red in subplot I). However, these responses do not reproduce the crispening (sigmoids shifting to high luminance), and responses for high luminance backgrounds are too linear. As a result, the achromatic behavior of this layer has been qualified as non-human. The chromatic responses display sigmoidal shapes, and they shift in the right directions under different backgrounds (subplots J and K). However, the non-linearities for the chromatic backgrounds are very smooth compared to the sharpness of the non-linearity for the achromatic background. As a result, the human similarity of chromatic behaviors has been qualified as intermediate. In contrast, the achromatic response of the fourth layer (subplot L) does reproduce the non-linear behavior and crispening, so it has been qualified as more human-like than the achromatic response of the first layer. Shifts of the chromatic non-linearities are stronger depending on the background, but the non-linearities in achromatic backgrounds (black curves in subplots M and N) are still too sharp. Therefore, the score remains the same.</p>
<p>The <italic>PerceptNet</italic> model displays non-linear behavior and crispening in the responses to the achromatic series (subplot O). However, note how the curves corresponding to light backgrounds exceed the response on dark backgrounds, so human similarity has been qualified as intermediate. The responses to red-green series in <italic>PerceptNet</italic> shift in the right directions on different backgrounds, but they are too linear (and hence wrong) in subplot P. In contrast, the blue-yellow responses (subplot Q) display a rather human behavior.</p>
<p>Finally, the <italic>Bio-U-Net</italic> shows a clearly non-human achromatic response: note the noise and wrong order in the curves with no trace of crispening (subplot R). In contrast, the responses to the chromatic series display the expected sigmoidal shape with the shift in the proper directions for the different chromatic backgrounds (subplots S and T). Noisy and unstable responses are what determined the intermediate score.</p></sec>
<sec>
<label>4.2.2</label>
<title>Achromatic and chromatic contrast sensitivities and receptive fields (props. 3, 4, and 5)</title>
<p>The top row of <xref ref-type="fig" rid="F11">Figure 11</xref> shows the achromatic Contrast Sensitivity Function (property 3, black curve) and the red-green and yellow-blue Contrast Sensitivity Functions (property 4, red and blue curves, respectively). These CSFs have been computed from the responses to noise patterns of controlled spatial frequency and the same low contrast (<italic>C</italic><sub>RMSE</sub> &#x0003D; 0.05) for every frequency. Patterns were generated in the corresponding color channel of the Jameson &#x00026; Hurvich color space (<xref ref-type="bibr" rid="B49">Hurvich and Jameson, 1957</xref>) that isolates luminance, red&#x02013;green, and yellow&#x02013;blue components. We consider the responses at the last layer of the networks, and we plot the Euclidean distance between the responses for each pattern and for a flat image of the same average color.</p>
<fig position="float" id="F11">
<label>Figure 11</label>
<caption><p>Achromatic and chromatic CSFs of the considered models [top subplots <bold>(A&#x02013;C)</bold>]. Subplots <bold>(D&#x02013;F)</bold> show the receptive fields tuned to 3, 6, and 12 cpd (different line styles) computed for the different models using the Blakemore &#x00026; Campbell CSF adaptation method (<xref ref-type="bibr" rid="B11">Blakemore and Campbell, 1969</xref>). Subplots <bold>(G&#x02013;I)</bold> show the receptive fields obtained using delta stimuli (<xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>) at early layers of the models and subplots <bold>(J&#x02013;L)</bold> at late layers.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0011.tif">
<alt-text>Comparison of three models: BioMultiLayer, PerceptNet, and Bio-UNet across properties. Panels A, D, G, and J show BioMultiLayer graphs and visual data with green check marks. Panels B, E, H, and K display PerceptNet&#x00027;s results with mixed evaluations. Panels C, F, I, and L present Bio-UNet&#x00027;s results with mostly red crosses. Graphs detail sensitivity and frequency, while visual data display model responses.</alt-text>
</graphic>
</fig>
<p>The CSFs of the <italic>BioMultiLayer</italic> model (subplot A) strongly resemble the human CSFs (<xref ref-type="bibr" rid="B17">Campbell and Robson, 1968</xref>; <xref ref-type="bibr" rid="B107">Mullen, 1985</xref>): the achromatic response is band-pass with peak sensitivity around 4 cpd and high cut-off frequency (above 32 cpd), and the chromatic responses are lower and basically low-pass with cut-off frequencies of about 15 cpd.</p>
<p>The achromatic CSF of the <italic>PerceptNet</italic> is also band-pass (black curve in subplot B), but the chromatic CSFs are far from human because their shape is also band-pass and the responses to modulations in the YB direction are much bigger than the responses to equivalent achromatic modulations.</p>
<p>Finally, the <italic>Bio-U-Net</italic> (subplot C) displays a strongly non-human behavior: see the non-plausible high-pass behavior of the responses to achromatic gratings, and the bigger responses to chromatic gratings in the mid-frequency range.</p>
<p><xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref> makes a systematic study of the impact on the CSFs of different read-out locations along the three models and six different read-out strategies, thus illustrating the points made in Section 2.2.</p>
<p>The receptive fields of the models (property 5) have been proved in two ways: (1) a <italic>psychophysical</italic> method based on the Blakemore &#x00026; Campbell experiment (<xref ref-type="bibr" rid="B11">Blakemore and Campbell, 1969</xref>), which relies on the attenuation of the CSF under adaptation for different frequencies, and (2) a <italic>physiological</italic> method based on recording the response to deltas in the luminance, red&#x02013;green, and yellow&#x02013;blue channels (<xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>). This is another example of two different experimental settings (psychophysical and physiological) to measure the alignment mentioned in Section 2.2.</p>
<p>In the <italic>BioMultiLayer</italic> model, the attenuation of the achromatic CSF when the gratings are shown on top of backgrounds of specific frequencies (subplot D) reveals the existence of narrow-band sensors with bandwidth that increases with frequency, which is consistent with human behavior (<xref ref-type="bibr" rid="B11">Blakemore and Campbell, 1969</xref>; <xref ref-type="bibr" rid="B126">Simoncelli and Adelson, 1990</xref>). This comes from the fact that the linear part of the 4th layer of this model is made of wavelet kernels, and their response is non-linearly attenuated by the activity of neighbor sensors tuned to the same feature through divisive normalization.</p>
<p>On the other hand, when checking the shape of the receptive fields using delta functions, one gets two biologically plausible results: (a) in the 3rd layer of the <italic>BioMultiLayer</italic> network, receptive fields are center-surround patterns in the achromatic, red&#x02013;green, and yellow&#x02013;blue directions (subplot G), and (b) in the fourth layer, one gets local frequency filters with different orientations and scales (subplot J) as happens in biological vision at LGN (<xref ref-type="bibr" rid="B15">Cai et al., 1997</xref>; <xref ref-type="bibr" rid="B123">Shapley and Hawken, 2011</xref>) and V1 (<xref ref-type="bibr" rid="B47">Hubel et al., 1959</xref>; <xref ref-type="bibr" rid="B138">Watson, 1983</xref>).</p>
<p>For <italic>PerceptNet</italic>, results are quite different: first, the Blakemore and Campbell experiment shows non-human wide-band mechanisms (subplot E). This is not only due to the non-human nature of the CSFs, it also means that any frequency leads to attenuation of the responses to patterns of any frequency. This departure from human behavior is also visible when getting the receptive fields from the last layer of the network using deltas: one gets oriented filters, but all with the same size and with low-frequency blobs. Moreover, chromatic information is spread along all the filters (subplot K), as opposed to what happens in the early layers, where one gets achromatic, red&#x02013;green, and yellow&#x02013;blue responses (subplot H).</p>
<p>Finally, in the <italic>Bio-U-Net</italic> model, the Blakemore and Campbell experiment also leads to non-human ratios of the CSFs (subplot F). The receptive fields obtained from deltas in the first layers lead to center-surround blobs but not in definite chromatic directions (subplot I). In the central layers of the encoder, one gets larger receptive fields which have no clear spatial oscillations nor preferred chromatic directions (subplot L).</p></sec>
<sec>
<label>4.2.3</label>
<title>Contrast saturation, dependence on frequency (properties 6 and 7)</title>
<p>The top row of <xref ref-type="fig" rid="F12">Figure 12</xref> shows the visibility response (Euclidean difference of response with respect to the response of a uniform gray image) for achromatic patterns of different frequencies and for red-green and yellow-blue patterns of different frequencies, all seen in isolation (properties 6 and 7). Line styles for the different frequencies in the achromatic and chromatic cases are different according to the different order expected for band-pass and low-pass systems. In every case, a match with human behavior would be illustrated by having the blue curve at the top and the red curve at the bottom, with a smooth transition from black to light gray in between.</p>
<fig position="float" id="F12">
<label>Figure 12</label>
<caption><p>Contrast responses of the considered models in different masking conditions and for achromatic and chromatic textures of different frequencies. The color code (indicated in the subplots corresponding to Perceptnet, but applicable to the equivalent curves of the other models) has been designed so that human response curves would be in the blue-red order, as in <xref ref-type="fig" rid="F7">Figures 7</xref>, <xref ref-type="fig" rid="F8">8</xref>. In this way, it is obvious which model reproduces human behavior better.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0012.tif">
<alt-text>Chart comparing visibility across three models: BioMultiLayer, PerceptNet, and Bio-UNet. The charts display visibility versus contrast in achromatic, red-green, and yellow-blue categories. Indicators show performance across properties and frequencies (3 and 12 cycles per degree). Results vary, with some configurations excelling while others underperform. Each model&#x00027;s effectiveness is marked with check or cross symbols for different conditions.</alt-text>
</graphic>
</fig>
<p>The <italic>BioMultiLayer</italic> model leads to saturating responses with larger intensities in the achromatic case (left) than in the chromatic cases (subplots A and B-C), as in humans (<xref ref-type="bibr" rid="B140">Watson and Solomon, 1997</xref>; <xref ref-type="bibr" rid="B102">Martinez-Uriegas, 1997</xref>). The achromatic response to mid-frequency (3 cpd, in blue) is clearly bigger than the response to the other frequencies, which is smoothly reduced for higher frequencies (from 6 to 24 cpd) and also attenuated for 1.5cpd. On the other hand, the chromatic responses are basically ordered according to frequency in a low-pass fashion. All these trends are in good agreement with human behavior.</p>
<p>The achromatic responses of the <italic>PerceptNet</italic>, though band-pass, exhibit quite a linear, non-saturating or even expanding behavior (subplot D). Moreover, these achromatic responses are not bigger than the response to chromatic patterns, particularly the yellow-blue (subplot F), which is contrary to human perception.</p>
<p>The responses for the chromatic patterns in the <italic>Bio-U-Net</italic> model exhibit human-like saturation, and they are in the right frequency order (subplots H and I), but they are larger than the responses for achromatic patterns (subplot G), which is contrary to human perception.</p></sec>
<sec>
<label>4.2.4</label>
<title>Energy masking and feature masking (properties 8&#x02013;10)</title>
<p>Each panel of the second row in <xref ref-type="fig" rid="F12">Figure 12</xref> shows the responses to a 3 cpd achromatic pattern (left) and a 12 cpd achromatic pattern (right) seen on top of a mask (noise of the same frequency and orientation) with progressively larger RMSE contrast (in the range [0,0.3]) leading to different response curves in different colors (from blue to red), thus checking the effect of the energy of the background (prop. 8). The color code has been selected so that the no-mask case is depicted in blue (less attenuated in humans), and colors from black to light gray and red are taken for progressively bigger contrasts of the mask.</p>
<p>The responses of the <italic>BioMultiLayer</italic> in <xref ref-type="fig" rid="F12">Figure 12</xref> (subplots J, K) progressively attenuate as the energy of the background is increased, in line with the reduction in visibility of the test shown in each column of <xref ref-type="fig" rid="F7">Figure 7</xref>. And this happens both for low and high frequency, with bigger responses for the mid-frequency. Therefore, the behavior is qualitatively human. The <italic>PerceptNet</italic> displays a completely non-human behavior: for the 3 cpd tests, progressively larger masks induce enhancement of the expansive (non-saturating) response, and the responses for the high-frequency patterns are larger, linear, and do not show significant variation with the mask. Finally, the <italic>Bio-U-Net</italic> model does display human-like attenuation of the response to 3 cpd patterns (subplot N). However, the responses to 12 cpd patterns (subplot O) are not human-like because of their (large) size, expansive shape, and increase with the energy of the mask.</p>
<p>The panels of the third row of <xref ref-type="fig" rid="F12">Figure 12</xref> show the responses for an achromatic test of 3 cpd (left) and 12 cpd (right) seen on top of backgrounds of different frequencies (and 0.2 contrast) compared to the no-mask condition, i.e., it checks the frequency cross-masking (property 9). The color code has been selected so that the no-mask case is depicted in blue (less attenuated in humans), and colors from black to light gray and red are taken for progressively closer frequencies in mask and test, which lead to increased attenuation of response in humans.</p>
<p>The response of the <italic>BioMultiLayer</italic> model is bigger in the no-mask condition, displays substantial attenuation when the background shares the same frequency as the test (red curves in subplots P and Q), and responses are bigger for 3 cpd than for 12 cpd. In each case, the optimal frequency is not the one that leads to the biggest attenuation, but it is close to it. The <italic>PerceptNet</italic> responses are not human because for the low frequency, subplot R, responses are not saturating regardless of the mask, and the responses for high frequency are larger, linear, and the presence of backgrounds leads to larger responses (subplot S). The <italic>Bio-U-Net</italic> does not show human-like trends because in the case that displays a saturating response, the presence of a background leads to responses larger than in the no-mask case (black curve in subplot T). The behavior in subplot U is non-human for the same reasons stated in subplots N and O.</p>
<p>Finally, the last row of <xref ref-type="fig" rid="F12">Figure 12</xref> shows the responses for low- and high-frequency achromatic patterns (left and right, respectively) seen on top of backgrounds of the same frequency but different orientations; i.e., it checks the orientation cross-masking (property 10). Again, the color code has been chosen so that in a human, the blue curve would be at the top and the red would be at the bottom, as in <xref ref-type="fig" rid="F8">Figure 8</xref>.</p>
<p>For this last example, the <italic>BioMultiLayer</italic> model gets bigger attenuation for the background of the same orientation, particularly for high frequency (see the red curves), and the other orientations lead to responses that are between the no-mask condition (in blue) and the same-orientation background (in red) in subplots V and W. The other models give clearly non-human results because (on top of the arguments used in previous cases) stimulation on backgrounds of the same orientation (red curves) does not lead to the expected attenuation, and bigger attenuation is obtained for backgrounds that are almost orthogonal to the test, which is not what humans experience in <xref ref-type="fig" rid="F8">Figure 8</xref>.</p></sec>
<sec>
<label>4.2.5</label>
<title>Summary of qualitative results</title>
<p>The qualitative evaluation of the considered models over the proposed tests is summarized in <xref ref-type="table" rid="T2">Table 2</xref>. From this table, there is a clear ranking of the alignment between the models and humans. It is not surprising that the parametric model (the <italic>BioMultiLayer</italic>) has greater alignment in the linear parts (props. 1 and 5) since sensitivities and center-surround and Gabor receptive fields were parametrically built into that model.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Summary of qualitative results, which for these models, is enough to clearly identify the best alignment.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th/>
<th valign="top" align="left"><bold>Facts / properties</bold></th>
<th valign="top" align="center"><bold>BioMultiLayer</bold></th>
<th valign="top" align="center"><bold>PerceptNet</bold></th>
<th valign="top" align="center"><bold>Bio-UNet</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="left">Spectral sensitivities (achromatic and opponent)</td>
<td valign="top" align="center"><bold>&#x0007E;</bold>&#x02713;</td>
<td valign="top" align="center"> <bold>&#x000D7;&#x0007E;</bold></td>
<td valign="top" align="center">&#x000D7;&#x000D7;</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="left">Brightness &#x00026; Color response saturation</td>
<td valign="top" align="center">&#x02713;<bold>&#x0007E;</bold></td>
<td valign="top" align="center"><bold>&#x0007E;</bold><bold>&#x0007E;</bold></td>
<td valign="top" align="center"> <bold>&#x000D7;&#x0007E;</bold></td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">Achromatic contrast sensitivity (bandwidth)</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="center">&#x000D7;</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="left">Chromatic Contrast Sensitivity (Bandwidth)</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="center"><bold>&#x0007E;</bold></td>
<td valign="top" align="center">&#x000D7;</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">Spatio-chromatic receptive fields</td>
<td valign="top" align="center">&#x02713;&#x02713;&#x02713;</td>
<td valign="top" align="center"><bold>&#x000D7;&#x0007E;</bold><bold>&#x0007E;</bold></td>
<td valign="top" align="center"><bold>&#x000D7;&#x000D7;&#x0007E;</bold></td>
</tr>
<tr>
<td valign="top" align="left">6</td>
<td valign="top" align="left">Non-linear contrast response: saturation</td>
<td valign="top" align="center">&#x02713; &#x02713;</td>
<td valign="top" align="center">&#x000D7;&#x000D7;</td>
<td valign="top" align="center">&#x000D7;&#x02713;</td>
</tr>
<tr>
<td valign="top" align="left">7</td>
<td valign="top" align="left">Non-linear contrast response: frequency order</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="center">&#x000D7;</td>
<td valign="top" align="center"><bold>&#x0007E;</bold></td>
</tr>
<tr>
<td valign="top" align="left">8</td>
<td valign="top" align="left">Context effects: energy</td>
<td valign="top" align="center">&#x02713;</td>
<td valign="top" align="center">&#x000D7;</td>
<td valign="top" align="center"><bold>&#x0007E;</bold></td>
</tr>
<tr>
<td valign="top" align="left">9</td>
<td valign="top" align="left">Context effects: frequency</td>
<td valign="top" align="center"><bold>&#x0007E;</bold></td>
<td valign="top" align="center">&#x000D7;</td>
<td valign="top" align="center">&#x000D7;</td>
</tr>
<tr>
<td valign="top" align="left">10</td>
<td valign="top" align="left">Context effects: orientation</td>
<td valign="top" align="center"><bold>&#x0007E;</bold></td>
<td valign="top" align="center">&#x000D7;</td>
<td valign="top" align="center">&#x000D7;</td>
</tr></tbody>
</table>
</table-wrap>
<p>More interestingly, the band-pass behavior of the sensors emerged from modifications in the CSFs in our simulation of the Blakemore and Campbell experiment. It is also interesting that the close reproduction of the band-pass and low-pass behavior and the relative scaling of the CSFs obtained from responses to sinusoids (an original check done here) was not built in. This indicates that the (non-trivial) gain of the center-surround cells and the Gabor cells was properly adjusted through the indirect psychophysical experiments done to set their parameters. As a result, the relative order of the (saturated) frequency responses (prop. 7) is also ok, both for achromatic and chromatic textures. The saturation of the responses to Gabor stimuli in isolation (prop. 6) is better reproduced in the parametric model than in the Bio-UNet. The difference between them is more evident when one digs deeper using props. 8-10 because they need proper interaction between texture sensors, and this was only easy to do in a parametric model such as the <italic>BioMultiLayer</italic>.</p>
<p>However, note that the reproduction of the interaction between features (both in color, property 2, and in texture, properties 9 and 10) is not properly reproduced, not even in the <italic>BioMultiLayer</italic>, pointing out that more work is needed to adjust its parameters, as discussed below.</p>
<p>According to the proposed test, the other two models (the non-parametric <italic>PerceptNet</italic> and the <italic>Bio-U-Net</italic>) are <italic>less human</italic>, in that order of alignment. This also makes sense because the <italic>PerceptNet</italic> was tuned to reproduce low-level human opinion on distortion, while the <italic>Bio-U-Net</italic> was just tuned to reproduce a specific mid-level vision goal such as image segmentation. In the discussion, we elaborate more on the combination of goals that may explain the organization of the visual system.</p>
<p>In any case, we see that even with this qualitative application of the proposed test (again, quantitative comparisons could be done with properties 1, 3, 4, and 6, even for moving patterns), a significant ranking is possible, and, as discussed below, the qualitative behaviors, when they are properly understood, suggest significant changes in the architectures and training of the models. Quantitative automation of the optimization should be iteratively done by alternating goals of different natures, as suggested in (<xref ref-type="bibr" rid="B99">Martinez et al. 2019</xref>): optimize for conventional goals and then fine-tune to reproduce the effects pointed out by the test proposed here (or the other way around).</p>
</sec>
</sec>
<sec>
<label>4.3</label>
<title>Results II: quantitative analysis</title>
<p>The table in <xref ref-type="fig" rid="F13">Figure 13</xref> shows the quantitative description of the alignment with humans of each model for each property using the method proposed in Section 3.3 that considers RMSE alignment up to a scale factor using Pearson correlation and preservation of the order between curves using a rank correlation (in this case, Kendall correlation). All the scores are in a comparable [&#x02212;1, 1] range, where bigger value means higher alignment. The table shows separate descriptors for read-outs done at different layers (<italic>L</italic><sub>1</sub> or <italic>L</italic><sub>4</sub>), for the order of the curves for properties measured at different chromatic channels (<italic>achromatic, red&#x02013;green, and yellow&#x02013;blue</italic>), or the order of the curves for different frequencies (<italic>Low f</italic> and <italic>High f</italic> ). Maximum correlations for each property are highlighted in green.</p>
<fig position="float" id="F13">
<label>Figure 13</label>
<caption><p>Alignment with human behavior measured in two ways: alignment between ground truth and prediction (Pearson correlation, <italic>left panel</italic>) and preservation of order between curves (rank Kendall correlation, <italic>right panel</italic>). Props. 8&#x02013;10 share the same &#x003C1;<sub><italic>p</italic></sub> values here because they share the same (known) no-masking curves. In props. 8&#x02013;10, the interesting masking behavior is described by the order of the curves, quantified by &#x003C1;<sub><italic>k</italic></sub>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0013.tif">
<alt-text>Six graphs compare responses under different divisive normalization conditions. The top row shows Yellow-Blue responses, while the bottom row shows Red-Green responses. Left column: Modified layer 1 with sharper normalization (worse), shows lower correlation values. Middle column: Original layer 1 with too nonlinear normalization (not right), moderate correlation values. Right column: Modified layer 1 with smoother normalization (better), highest correlation values. The graphs display correlations and stimulus responses for linear tristimulus inputs.</alt-text>
</graphic>
</fig>
<p>The plain average of the descriptors, as suggested in BrainScore (<xref ref-type="bibr" rid="B119">Schrimpf et al., 2018</xref>), leads to the following model ranking: <italic><bold>BioMultiLayer</bold></italic> : <bold>0.68</bold>&#x000B1;0.14, <italic><bold>PerceptNet</bold></italic> : <bold>0.26</bold>&#x000B1;0.26. <italic><bold>Bio-UNet</bold></italic> : <bold>0.20</bold>&#x000B1;0.34, and Two-sample Kolmogorov-Smirnov tests (<xref ref-type="bibr" rid="B106">Monahan, 2011</xref>) show that the <italic>BioMultiLayer</italic> is significantly more aligned with humans than the other two (with <italic>p</italic> &#x0003D; 8&#x000B7;10<sup>&#x02212;3</sup> and <italic>p</italic> &#x0003D; 7&#x000B7;10<sup>&#x02212;3</sup>, respectively), while the other two models are not significantly different (<italic>p</italic> &#x0003D; 0.26). This global conclusion is consistent with the qualitative analysis shown above.</p>
<p>The trend of individual scores in the table of quantitative descriptors agrees with the qualitative analysis shown in <xref ref-type="table" rid="T2">Table 2</xref>. Nevertheless, as anticipated in Section 3.3, the numerical descriptors may miss certain qualitative differences. For example, the rank correlation for the chromatic case of Property 2 in <italic>PerceptNet</italic> states that the alignment is good for both the RG and the YB cases. However, inspection of plots P and Q in <xref ref-type="fig" rid="F10">Figure 10</xref>, compared with the expected qualitative behavior in <xref ref-type="fig" rid="F6">Figure 6</xref>, shows that while the order in the RG curves of plot (<xref ref-type="fig" rid="F10">Figure 10</xref>). P is correct, the shape of the curves is clearly not sigmoidal (i.e., wrong).</p>
<p>On the other hand, plain aggregation of correlation results over many (arbitrarily selected) phenomena, as done here following BrainScore (<xref ref-type="bibr" rid="B119">Schrimpf et al., 2018</xref>), has an obvious impact on the power of the descriptor. Bootstrap experiments on the results of the quantitative table illustrate this fact in <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
<p>In summary, the goal of the proposed &#x0201C;visual&#x0201D; Turing Tests is to allow one to &#x0201C;experience&#x0201D; the shape of the response curves by looking at the test stimuli generated by the software; however, the suggested quantification also allows a blind assessment of the alignment between networks and humans. Nevertheless, like any cost function, the value of the specific aggregated quantitative descriptor has to be taken with caution, and it is essential to always check that it properly captures the relevant qualitative (visual) behavior.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion: what can be learned from the proposed methodology?</title>
<p>In this section, we discuss the benefits of the proposed <italic>Decalogue</italic> for generic artificial models. Benefits go beyond the evaluation of the human nature of models: even if we don&#x00027;t need a certain model to be similar to humans, the behaviors described by the human-like curves elicited by the stimuli in the <italic>Decalogue</italic> imply human-like bottlenecks and adaptation properties that one would like in efficient and robust artificial vision systems. Similarly, we also discuss the benefits of the architectures from classical vision science models that reproduce such behaviors.</p>
<sec>
<label>5.1</label>
<title>(Non-human) curves suggest changes in the models</title>
<p>Failures to achieve the expected result in the proposed Test may suggest changes in the parameters of the models or even in their architecture. This is easy to see in the models considered here because they are relatively simple and interpretable, but this may also be the case in more recent (more complex) models.</p>
<p>First, consider, for instance, the non-human (not-right) behavior of the <italic>BioMultiLayer</italic> in the perception of color in the central panel of <xref ref-type="fig" rid="F14">Figure 14</xref>. In that case, it displays too non-linear (too-sharp) responses to YB and RG stimuli in the no-adaptation case (curves in gray). As this model is made of interpretable divisive normalization layers<xref ref-type="fn" rid="fn0013"><sup>11</sup></xref>, the strength of the non-linearity can be modulated by the relative weight of the pool in the denominator. Following that intuition, we made two modifications to the original configuration of the first layer of the <italic>BioMultiLayer</italic>: we decreased and increased the parameter <italic>b</italic><sub><italic>i</italic></sub> that controls the relative weight of the pool in the denominator. By doing so, and re-computing the responses to the color stimuli in the test) we get the behaviors shown in the left and right panels of <xref ref-type="fig" rid="F14">Figure 14</xref>: the interventions either alleviate the problem (right panel) or make it worse (left panel). This can be seen both in the qualitative shape of the curves and in the quantitative scores. The same rationale can be applied to other failures (e.g., too sharp/smooth responses in <xref ref-type="fig" rid="F12">Figure 12</xref>, subplot P for textures) as suggested in <xref ref-type="bibr" rid="B99">Martinez et al. (2019</xref>).</p>
<fig position="float" id="F14">
<label>Figure 14</label>
<caption><p>Fixing errors in a model (BioMultiLayer) from the results of the test. <italic>Central Column</italic>: displays the non-linear YB and RG responses of the first layer of the original BioMultiLayer (as depicted in <xref ref-type="fig" rid="F10">Figures 10J</xref>, <xref ref-type="fig" rid="F10">K</xref>). <bold>(Left)</bold> and <bold>(Right)</bold> panels show the effect of different interventions in the parameters of the model (see text).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0014.tif">
<alt-text>Three graphs illustrate saturation responses for achromatic, red-green, and yellow-blue contrasts, each showing a solid and dashed line. Below, four graphs show pixel value outputs across four DN layers with black, gray, and white backgrounds. The graphs depict varying responses and contrast effects in visual processing.</alt-text>
</graphic>
</fig>
<p>When measuring the response of conventional networks using the spatially and chromatically calibrated stimuli proposed here, one can get human-like behaviors such as the ones shown in Section 3. For instance, shallow autoencoders optimized for image deblurring and denoising display human-like saturation when responding to achromatic and chromatic gratings of controlled spatial frequency: see <xref ref-type="fig" rid="F15">Figure 15</xref> (top), reproduced from <xref ref-type="bibr" rid="B68">Li et al. (2022</xref>). In this case, the slope of the response of these autoencoders (their sensitivity) is bigger for achromatic gratings than for red-green and yellow-blue gratings (Properties 3 and 4), and it reduces with the contrast of the gratings, just as in humans (Property 6).</p>
<fig position="float" id="F15">
<label>Figure 15</label>
<caption><p><bold>(Top)</bold> Human-like saturation behavior in contrast response (Property 6) happening in generic shallow autoencoders (<xref ref-type="bibr" rid="B68">Li et al., 2022</xref>). <bold>(Bottom)</bold> Human-like behavior obtained in image segmentation U-Nets when they are equipped with bio-inspired divisive normalization to improve their performance (<xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0015.tif">
<alt-text>Three sections show different visual analyses. The top left illustrates color temperature effects on an image at 3000 K (orange tint), equienergetic (neutral), and 10000 K (blue tint). Below these are line graphs comparing RMSE values for Red &#x0002B; Degrad, Degrad, and Blue &#x0002B; Degrad, with RMSE values of 29.3, 28.4, and 27.3, respectively. The right side compares images of a street with no fog and high fog, accompanied by input and output visual data representations.</alt-text>
</graphic>
</fig>
<p>As shown in the example above, this adaptive behavior and its benefits can be enforced in conventional networks by changing their architecture, for instance, by including divisive normalization or alternative bio-inspired non-linear layers. Examples include benefits in autoencoding and compression (<xref ref-type="bibr" rid="B79">Malo et al., 2006</xref>; <xref ref-type="bibr" rid="B3">Ball&#x000E9; et al., 2017</xref>), denoising and enhancement (<xref ref-type="bibr" rid="B37">Guti&#x000E9;rrez et al., 2006</xref>; <xref ref-type="bibr" rid="B58">Laparra et al., 2017</xref>), segmentation (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>; <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>), classification (<xref ref-type="bibr" rid="B21">Coen-Cagli and Schwartz, 2013</xref>; <xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al., 2020</xref>; <xref ref-type="bibr" rid="B104">Miller et al., 2022</xref>), or robustness to adversarial attacks with few layers given the strong non-linearity due to this kind of biological computation (<xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al., 2020</xref>). Moreover, the inclusion of these non-linearities, if done parametrically, (e.g., by using parametric expressions) reduces the training time and increases generalization because of the drastic reduction in the number of parameters of the network (<xref ref-type="bibr" rid="B133">Vila-Tom&#x000E1;s et al., 2025</xref>).</p>
<p>Of course, for more complex (non-interpretable) models, the interventions may not be that simple. However, enforcing the behaviors shown here, for instance by imposing certain band-pass behaviors (similar to properties 3 and 4) through regularization, will change the models and may improve their results. For example, complex vision-language models with human-like CSFs also have human-like responses to some adversarial attacks (<xref ref-type="bibr" rid="B43">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025a</xref>). Testing vision-language models with the stimuli in the proposed test is particularly interesting because one can interact with the model verbally (as done in humans) and ask if there are visible patterns in the images. In that way, sensitivities can be derived via psychometric functions (<xref ref-type="bibr" rid="B43">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025a</xref>) as opposed to using (arguable) distortion metrics in the internal representations of the model [as done, for instance, in <xref ref-type="bibr" rid="B68">Li et al. (2022</xref>) and <xref ref-type="bibr" rid="B1">Akbarinia et al. (2023</xref>) and illustrated in Section 2 of <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>].</p>
</sec>
<sec>
<label>5.2</label>
<title>Changes in the optimization goal or data statistics to achieve human-like adaptation</title>
<p>Misalignment with human behavior in proposed tests may also suggest changes in the optimization goal and in the statistics of the training data.</p>
<p>For instance, it is known that information maximization arguments lead to the emergence of Gabor-like receptive fields tuned to achromatic and opponent-chromatic directions (<xref ref-type="bibr" rid="B50">Hyvarinen et al., 2009</xref>; <xref ref-type="bibr" rid="B38">Gutmann et al., 2014</xref>). However, that sensible goal can be complemented with denoising-deblurring tasks so that center-surround cells and proper contrast sensitivity do emerge (<xref ref-type="bibr" rid="B2">Atick et al., 1992</xref>; <xref ref-type="bibr" rid="B51">Karklin and Simoncelli, 2011</xref>; <xref ref-type="bibr" rid="B70">Lindsey et al., 2019</xref>; <xref ref-type="bibr" rid="B68">Li et al., 2022</xref>). Moreover, if the contrast non-linearities do not emerge, they may be enforced by the segmentation goal in the encoder, as in (<xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al. 2025b</xref>; see <xref ref-type="fig" rid="F15">Figure 15</xref>, bottom), which shows curves where excitation is moderated by the presence of active neighbors. Regarding the poor emergence of plausible receptive fields in the considered non-parametric models (PerceptNet and BioUNet), this <italic>error</italic> makes sense in the context of the recently proposed <italic>feature-spreading</italic> problem (<xref ref-type="bibr" rid="B133">Vila-Tom&#x000E1;s et al., 2025</xref>): if the goal is not demanding enough (as is usual in conventional goals), the features spread along all layers of the net in a way that the weak goal(s) is (are) fulfilled, but the layers remain biologically non-plausible.</p>
<p>Regarding suggestions on the training data, the behavior of the achromatic and chromatic CSFs proposed here was checked in <xref ref-type="bibr" rid="B68">Li et al. (2022</xref>) in scenes with well-controlled illumination. The behavior found in the autoencoder CSFs in those cases resembles Von Kries adaptation, as anticipated in (<xref ref-type="bibr" rid="B38">Gutmann et al. 2014</xref>). <xref ref-type="fig" rid="F16">Figure 16</xref> (left) shows that under low-temperature (reddish) illumination, the red-tuned channel is relatively attenuated with regard to the blue-tuned channel, and the other way around under high-temperature (bluish) illumination, as would happen using a Von Kries computation (<xref ref-type="bibr" rid="B28">Fairchild, 2013</xref>) or imposing the shifts in the response curves (<xref ref-type="bibr" rid="B53">Krauskopf and Gegenfurtner, 1992</xref>) shown in the Decalogue.</p>
<fig position="float" id="F16">
<label>Figure 16</label>
<caption><p><bold>(Left)</bold> Adaptation in the CSFs in autoencoders obtained from training in the proper (colormetrically calibrated) environments for color adaptation (<xref ref-type="bibr" rid="B68">Li et al., 2022</xref>). <bold>(Right)</bold> Improved contrast perception by using bio-inspired divisive normalization (with adaptive contrast responses such as the ones described in our proposal) in the model (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>). Images reproduced with permission from: M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, &#x0201C;The Cityscapes Dataset for Semantic Urban Scene Understanding,&#x0201D; in <italic>Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic>, 2016.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0016.tif">
<alt-text>Top section features two graphs: a grayscale heat map on the left labeled \(P(w_{i} | w^{&#x0002A;}_{i})\) with axes \(w_{i}\) and \(w^{&#x0002A;}_{i}\), and another on the right labeled \(P(r_{i}^{(1.7)} | r_{i}^{(1.7)})\) with a mathematical equation in between. The bottom section shows three images of a building facade with different reconstruction methods: (a) \(min_{u(x)} ||x &#x02013; \hat{x}||_{2}\), (b) \(min_{u(x)} NLDP(x, \hat{x})\), and (c) \(min_{u(x)} 1 &#x02013; MS-SSIM(x, \hat{x})\).</alt-text>
</graphic>
</fig>
<p>Finally, the emergence of the contrast-dependent non-linearities of property 8 (or the ability for contrast enhancement) may be enforced by including low-contrast images in the training of regular networks. For example, <xref ref-type="fig" rid="F16">Figure 16</xref> (right) shows that a network using Div. Norm. leads to contrast enhancement despite it not being trained on high-fog images, as anticipated by the contrast-dependent non-linearities shown in <xref ref-type="fig" rid="F15">Figure 15</xref> (bottom). However, in regular UNets (that do not include Div. Norm.), this behavior can be obtained (to a lesser degree) by including high-fog images in the training (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>). In fact, the behavior in the BioUNet that includes Div. Norm. is not completely human (plot 12.N is ok but plots 12.O or 12.U are not). This may be in part due to a lack of constraints in the Div. Norm. [free kernels in <xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al. (2023</xref>) and <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al. (2025b</xref>) as opposed to more sensible parametric kernels in <xref ref-type="bibr" rid="B100">Martinez et al. (2018</xref>, <xref ref-type="bibr" rid="B99">2019</xref>)], but also because its behavior was obtained by training only with good-quality (clear day) images.</p>
<p>Of course, the limitations imposed by low-dynamic range and 8-bit quantized images imply that the artificial systems face scenes with limited variability, and hence, they may not develop certain non-linearities that are more pronounced (or conspicuous) in humans. In fact, including biologically inspired non-linear layers to cope with such variability (as suggested in Section 5.1) is an easy way to improve the efficiency of gamut mapping techniques (<xref ref-type="bibr" rid="B58">Laparra et al., 2017</xref>) and the robustness of networks devoted to vision (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>; <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>; <xref ref-type="bibr" rid="B9">Bertalm&#x000ED;o et al., 2020</xref>).</p>
</sec>
<sec>
<label>5.3</label>
<title>Human-like curves imply better priors for natural image statistics</title>
<p>Two examples may illustrate how the non-linear responses to Gabor stimuli shown in textured contexts as presented in the proposed Decalogue capture the statistics of natural images: the described non-linear behaviors are a robust prior that may benefit anynetwork intended to work in vision.</p>
<p>First, in <xref ref-type="fig" rid="F17">Figure 17</xref> (top) we show that the energy of neighbor Gabor-like coefficients is correlated [bow-tie conditional probabilities of Gabor coefficients in natural images, as reported in <xref ref-type="bibr" rid="B121">Schwartz and Simoncelli (2001</xref>)], but the non-linear responses in textured backgrounds make the resulting coefficients independent (<xref ref-type="bibr" rid="B91">Malo and Laparra, 2010</xref>). In this regard, interactions between coefficients (e.g., as in Div. Norm.) are strictly required to remove redundancy because it is known that mutual information (redundancy) is invariant under point-wise transforms (<xref ref-type="bibr" rid="B60">Laparra et al., 2025</xref>).</p>
<fig position="float" id="F17">
<label>Figure 17</label>
<caption><p><bold>(Top)</bold> Image PDF factorization from the contrast non-linearites (div. norm.) illustrated in <xref ref-type="fig" rid="F7">Figures 7</xref>, <xref ref-type="fig" rid="F8">8</xref>, as shown in <xref ref-type="bibr" rid="B91">Malo and Laparra (2010</xref>). <bold>(Bottom)</bold> autoencoders trained with distortion metrics based on the contrast non-linearities described in our proposal capture natural image statistics despite being trained with few samples (<xref ref-type="bibr" rid="B41">Hepburn et al., 2022</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-08-1665874-g0017.tif">
<alt-text>Diagram of a transformation process with two grayscale plots and a mathematical formula. The first plot shows a gradient radiating from the center. The second plot, following a red arrow through the formula, displays horizontal banding. Below are three labeled images: (a)darkened image of a building, (b) clearer building facade, (c) complete facade with brighter details.</alt-text>
</graphic>
</fig>
<p>Second, non-Euclidean metrics based on the non-linear responses to the stimuli presented here [e.g., metrics like those reported in (<xref ref-type="bibr" rid="B62">Laparra et al. 2010</xref>, <xref ref-type="bibr" rid="B58">2017</xref>); <xref ref-type="bibr" rid="B99">Martinez et al. (2019</xref>); <xref ref-type="bibr" rid="B81">Malo et al. (2024</xref>)] represent a robust prior of the PDF of natural images as illustrated by the fact shown in <xref ref-type="fig" rid="F17">Figure 17</xref> (bottom): in autoencoders with access to very few samples, the use of this kind of perceptual metric in the loss function, makes the reconstruction of images much more robust than those using (naive) Euclidean metrics because the perceptual metric is already capturing the statistics of natural images, although samples are missing (<xref ref-type="bibr" rid="B41">Hepburn et al., 2022</xref>).</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Final remarks</title>
<p>First, we noted that there are many open problems when we evaluate the human nature of artificial networks: there is a non-trivial relationship between the training environment, the task, and the architecture (<xref ref-type="bibr" rid="B112">Poggio, 2021</xref>; <xref ref-type="bibr" rid="B89">Malo and Hern&#x000E1;ndez-C&#x000E1;mara, 2024</xref>; <xref ref-type="bibr" rid="B46">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025c</xref>). That complexity implies it is difficult to choose the layer(s) to measure from and the read-out mechanism to check the human nature of the model responses. These problems point out the need for new tests of human alignment that are independent of the training data and goal.</p>
<p>This motivates our proposal: a set of stimuli, <italic>a Decalogue</italic>, based on classical low-level vision science. The stimuli and associated human responses describe the adaptive information bottleneck in the retina-V1 pathway. On the one hand, some of the sensitivity surfaces are standardized (<xref ref-type="bibr" rid="B143">Wyszecki and Stiles, 2000</xref>; <xref ref-type="bibr" rid="B139">Watson and Ramirez, 2000</xref>; <xref ref-type="bibr" rid="B141">Watson and Malo, 2002</xref>; <xref ref-type="bibr" rid="B107">Mullen, 1985</xref>; <xref ref-type="bibr" rid="B22">Daly, 1990</xref>; <xref ref-type="bibr" rid="B94">Malo et al., 1997</xref>; <xref ref-type="bibr" rid="B52">Kelly, 1979</xref>), or data is readily available (<xref ref-type="bibr" rid="B11">Blakemore and Campbell, 1969</xref>; <xref ref-type="bibr" rid="B24">D&#x000ED;ez-Ajenjo et al., 2011</xref>) and allow quantitative comparison [namely properties 1, 3, 4, 5, 6], as done in <xref ref-type="bibr" rid="B134">Vila-Tom&#x000E1;s et al. (2023</xref>), <xref ref-type="bibr" rid="B68">Li et al. (2022</xref>), <xref ref-type="bibr" rid="B1">Akbarinia et al. (2023</xref>), <xref ref-type="bibr" rid="B39">Hammou et al. (2025</xref>), and <xref ref-type="bibr" rid="B16">Cai et al. (2025</xref>). On the other hand, we showed that the responses involving tests in different illuminations or textured backgrounds (namely properties 2, 7&#x02013;10) have clear qualitative trends that allow a quantitative assessment of the rank of the curves. As a consequence, we proposed a numerical description of the alignment combining Pearson and Kendall correlations.</p>
<p>This qualitative/quantitative analysis of the responses was applied to evaluate and rank three illustrative models: (1) a parametric one based on physiology, classical psychophysics, and maximum differentiation measurements (<xref ref-type="bibr" rid="B100">Martinez et al., 2018</xref>, <xref ref-type="bibr" rid="B99">2019</xref>; <xref ref-type="bibr" rid="B32">Gomez-Villa et al., 2020a</xref>; <xref ref-type="bibr" rid="B81">Malo et al., 2024</xref>), (2) a non-parametric model, the <italic>PerceptNet</italic> (<xref ref-type="bibr" rid="B40">Hepburn et al., 2020</xref>), that includes trainable divisive normalization to reproduce human opinion on subjective image quality, and (3) a U-net with the same encoder as the <italic>PerceptNet</italic> but trained for image segmentation (<xref ref-type="bibr" rid="B45">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2023</xref>; <xref ref-type="bibr" rid="B44">Hern&#x000E1;ndez-C&#x000E1;mara et al., 2025b</xref>). Experiments show that the proposed tests illustrate in easy-to-see ways the impact of the read-out location and strategy. Moreover, the quantitative and qualitative results are consistent and successfully rank the models according to their different origins: the two models with less alignment have been trained for tasks that are not enough to fully explain human behavior or are too flexible so that they easily develop non-human behavior.</p>
<p>Finally, in the discussion, we have seen that the proposed test can be useful to modify the architecture of the networks, both in their linear and non-linear parts. The test is useful to question the tasks or restrictions that are used in training (e.g., infoMax, noise, compression bottlenecks, classification, and segmentation). It is also useful to question the data used in the training, either in their generality or balance. Moreover, we discussed how the use of human behaviors represented by the data in the proposed test gives rise to priors related to the statistics of natural images.</p>
<p>In summary, we argue that the analysis of any kind of network, not only those that are specifically dedicated to modeling human vision, but any devoted to vision, can benefit, in great measure, from seeing how they respond to the proposed test.</p></sec>
</body>
<back>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/<xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>JV-T: Methodology, Software, Validation, Writing &#x02013; review &#x00026; editing. PH-C: Methodology, Software, Validation, Writing &#x02013; review &#x00026; editing. QL: Methodology, Software, Validation, Writing &#x02013; review &#x00026; editing. VL: Conceptualization, Funding acquisition, Writing &#x02013; review &#x00026; editing. JM: Conceptualization, Funding acquisition, Methodology, Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
<p>The author JM declared that they were an editorial board member of Frontiers at the time of submission. This had no impact on the peer review process and the final decision.</p>
</sec>
<sec sec-type="ai-statement" id="s10">
<title>Generative AI statement</title>
<p>The author(s) declared that generative AI was not used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p></sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s12">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2025.1665874/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frai.2025.1665874/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Presentation_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/></sec>
<ref-list>
<title>References</title>
<ref id="B1">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Akbarinia</surname> <given-names>A.</given-names></name> <name><surname>Morgenstern</surname> <given-names>Y.</given-names></name> <name><surname>Gegenfurtner</surname> <given-names>K.</given-names></name></person-group> (<year>2023</year>). <article-title>Contrast sensitivity function in deep networks</article-title>. <source>Neural Netw</source>. <volume>164</volume>:<fpage>228</fpage>&#x02013;<lpage>244</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neunet.2023.04.032</pub-id><pub-id pub-id-type="pmid">37156217</pub-id></mixed-citation>
</ref>
<ref id="B2">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Atick</surname> <given-names>J. J.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Redlich</surname> <given-names>A. N.</given-names></name></person-group> (<year>1992</year>). <article-title>Understanding retinal color coding from first principles</article-title>. <source>Neural Comput</source>. <volume>4</volume>, <fpage>559</fpage>&#x02013;<lpage>572</lpage>. doi: <pub-id pub-id-type="doi">10.1162/neco.1992.4.4.559</pub-id></mixed-citation>
</ref>
<ref id="B3">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ball&#x000E9;</surname> <given-names>J.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E.</given-names></name></person-group> (<year>2017</year>). <article-title>End-to-end optimized image compression</article-title>. <source>ICLR ArXiV:1611.01704</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1611.01704</pub-id></mixed-citation>
</ref>
<ref id="B4">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Barlow</surname> <given-names>H.</given-names></name></person-group> (<year>1959</year>). <article-title>&#x0201C;Sensory mechanisms, the reduction of redundancy, and intelligence,&#x0201D;</article-title> in <source>Proc. of the Nat. Phys. Lab. Symposium on the Mechanization of Thought Process</source> (<publisher-loc>London</publisher-loc>: <publisher-name>HM Stationery Office</publisher-name>), <fpage>535</fpage>&#x02013;<lpage>539</lpage>.</mixed-citation>
</ref>
<ref id="B5">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Barlow</surname> <given-names>H.</given-names></name></person-group> (<year>1961</year>). <article-title>&#x0201C;Possible principles underlying the transformation of sensory messages,&#x0201D;</article-title> in <source>Sensory Communication</source>, ed. W. Rosenblith (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>217</fpage>&#x02013;<lpage>234</lpage>.</mixed-citation>
</ref>
<ref id="B6">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Barlow</surname> <given-names>H.</given-names></name></person-group> (<year>2001</year>). <article-title>Redundancy reduction revisited</article-title>. <source>Network: Comp. Neur. Syst</source>. <volume>12</volume>, <fpage>241</fpage>&#x02013;<lpage>253</lpage>. doi: <pub-id pub-id-type="doi">10.1088/0954-898X/12/3/301</pub-id><pub-id pub-id-type="pmid">11563528</pub-id></mixed-citation>
</ref>
<ref id="B7">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bell</surname> <given-names>A.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T.</given-names></name></person-group> (<year>1997</year>). <article-title>The &#x02018;independent components&#x00027; of natural scenes are edge filters</article-title>. <source>Vis. Res</source>. <volume>37</volume>, <fpage>3327</fpage>&#x02013;<lpage>3338</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0042-6989(97)00121-1</pub-id><pub-id pub-id-type="pmid">9425547</pub-id></mixed-citation>
</ref>
<ref id="B8">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Dur&#x000E1;n-Vizca&#x000ED;no</surname> <given-names>A.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Wichmann</surname> <given-names>F.</given-names></name></person-group> (<year>2024</year>). <article-title>Plaid masking explained with input-dependent dendritic nonlinearities</article-title>. <source>Sci. Rep</source>. <volume>14</volume>:<fpage>24856</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-024-75471-5</pub-id><pub-id pub-id-type="pmid">39438555</pub-id></mixed-citation>
</ref>
<ref id="B9">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Gomez-Villa</surname> <given-names>A.</given-names></name> <name><surname>Mart&#x000ED;n</surname> <given-names>A.</given-names></name> <name><surname>Vazquez</surname> <given-names>J.</given-names></name> <name><surname>Kane</surname> <given-names>D.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Evidence for the intrinsically nonlinear nature of receptive fields in vision</article-title>. <source>Sci. Rep</source>. <volume>10</volume>:<fpage>16277</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s41598-020-73113-0</pub-id><pub-id pub-id-type="pmid">33004868</pub-id></mixed-citation>
</ref>
<ref id="B10">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Biscione</surname> <given-names>V.</given-names></name> <name><surname>Yin</surname> <given-names>D.</given-names></name> <name><surname>Malhotra</surname> <given-names>G.</given-names></name> <name><surname>Dujmovic</surname> <given-names>M.</given-names></name> <name><surname>Montero</surname> <given-names>M. L.</given-names></name> <name><surname>Puebla</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>MindSet: Vision: a toolbox for testing DNNs on key psychological experiments</article-title>. <source>arXiv preprint arXiv:2404.05290</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2404.05290</pub-id></mixed-citation>
</ref>
<ref id="B11">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Blakemore</surname> <given-names>C.</given-names></name> <name><surname>Campbell</surname> <given-names>F.</given-names></name></person-group> (<year>1969</year>). <article-title>On the existence of neurons selectivity sensitive to the orientation and size of retinal images</article-title>. <source>J. Physiol</source>. <volume>203</volume>, <fpage>237</fpage>&#x02013;<lpage>260</lpage>. doi: <pub-id pub-id-type="doi">10.1113/jphysiol.1969.sp008862</pub-id></mixed-citation>
</ref>
<ref id="B12">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bowers</surname> <given-names>J. S.</given-names></name> <name><surname>Malhotra</surname> <given-names>G.</given-names></name> <name><surname>Dujmovi&#x00107;</surname> <given-names>M.</given-names></name> <name><surname>Montero</surname> <given-names>M. L.</given-names></name> <name><surname>Tsvetkov</surname> <given-names>C.</given-names></name> <name><surname>Biscione</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Deep problems with neural network models of human vision</article-title>. <source>Behav. Brain Sci</source>. <volume>46</volume>:<fpage>e385</fpage>. doi: <pub-id pub-id-type="doi">10.1017/S0140525X23002777</pub-id><pub-id pub-id-type="pmid">36453586</pub-id></mixed-citation>
</ref>
<ref id="B13">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Burg</surname> <given-names>M.</given-names></name> <name><surname>Cadena</surname> <given-names>S.</given-names></name> <name><surname>Denfield</surname> <given-names>G. H.</given-names></name> <name><surname>Walker</surname> <given-names>E.</given-names></name> <name><surname>Tolias</surname> <given-names>A.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Learning divisive normalization in primary visual cortex</article-title>. <source>PLoS Comput. Biol</source>. <volume>16</volume>:<fpage>e1009028</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pcbi.1009028</pub-id><pub-id pub-id-type="pmid">34097695</pub-id></mixed-citation>
</ref>
<ref id="B14">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Cadena</surname> <given-names>S.</given-names></name> <name><surname>Denfield</surname> <given-names>G.</given-names></name> <name><surname>Walker</surname> <given-names>E.</given-names></name> <name><surname>Gatys</surname> <given-names>L.</given-names></name> <name><surname>Tolias</surname> <given-names>A.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Deep convolutional models improve predictions of macaque v1 responses to natural images</article-title>. <source>PLoS Comput. Biol</source>. <volume>15</volume>:<fpage>e1006897</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pcbi.1006897</pub-id><pub-id pub-id-type="pmid">31013278</pub-id></mixed-citation>
</ref>
<ref id="B15">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>D.</given-names></name> <name><surname>DeAngelis</surname> <given-names>G.</given-names></name> <name><surname>Freeman</surname> <given-names>R.</given-names></name></person-group> (<year>1997</year>). <article-title>Spatiotemporal receptive field organization in the lateral geniculate nucleus of cats and kittens</article-title>. <source>J. Neurophysiol</source>. <volume>78</volume>, <fpage>1045</fpage>&#x02013;<lpage>1061</lpage>. doi: <pub-id pub-id-type="doi">10.1152/jn.1997.78.2.1045</pub-id><pub-id pub-id-type="pmid">9307134</pub-id></mixed-citation>
</ref>
<ref id="B16">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>Y.</given-names></name> <name><surname>Yin</surname> <given-names>F.</given-names></name> <name><surname>Hammou</surname> <given-names>D.</given-names></name> <name><surname>Mantiuk</surname> <given-names>R.</given-names></name></person-group> (<year>2025</year>). Do computer vision foundation models learn the low-level characteristics of the human visual system? <italic>CVPR ArXiV: 2502.20256</italic>. doi: <pub-id pub-id-type="doi">10.1109/CVPR52734.2025.01866</pub-id></mixed-citation>
</ref>
<ref id="B17">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>F.</given-names></name> <name><surname>Robson</surname> <given-names>J.</given-names></name></person-group> (<year>1968</year>). <article-title>Application of Fourier analysis to the visibility of gratings</article-title>. <source>J. Physiol</source>. <volume>197</volume>, <fpage>551</fpage>&#x02013;<lpage>566</lpage>. doi: <pub-id pub-id-type="doi">10.1113/jphysiol.1968.sp008574</pub-id><pub-id pub-id-type="pmid">5666169</pub-id></mixed-citation>
</ref>
<ref id="B18">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Carandini</surname> <given-names>M.</given-names></name> <name><surname>Heeger</surname> <given-names>D.</given-names></name></person-group> (<year>1994</year>). <article-title>Summation and division by neurons in visual cortex</article-title>. <source>Science</source> <volume>264</volume>, <fpage>1333</fpage>&#x02013;<lpage>1336</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.8191289</pub-id></mixed-citation>
</ref>
<ref id="B19">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Carandini</surname> <given-names>M.</given-names></name> <name><surname>Heeger</surname> <given-names>D. J.</given-names></name></person-group> (<year>2012</year>). <article-title>Normalization as a canonical neural computation</article-title>. <source>Nat. Rev. Neurosci</source>. <volume>13</volume>, <fpage>51</fpage>&#x02013;<lpage>62</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nrn3136</pub-id><pub-id pub-id-type="pmid">22108672</pub-id></mixed-citation>
</ref>
<ref id="B20">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Carney</surname> <given-names>T.</given-names></name> <name><surname>Klein</surname> <given-names>S.</given-names></name> <name><surname>Tyler</surname> <given-names>C.</given-names></name> <name><surname>Silverstein</surname> <given-names>A.</given-names></name> <name><surname>Beutter</surname> <given-names>B.</given-names></name> <name><surname>Levi</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>1999</year>). <article-title>Development of an image/threshold database for designing and testing human vision models</article-title>. <source>Hum. Vis. Electr. Imaging IV</source> <volume>3644</volume>, <fpage>542</fpage>&#x02013;<lpage>551</lpage>. doi: <pub-id pub-id-type="doi">10.1117/12.348473</pub-id></mixed-citation>
</ref>
<ref id="B21">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Coen-Cagli</surname> <given-names>R.</given-names></name> <name><surname>Schwartz</surname> <given-names>O.</given-names></name></person-group> (<year>2013</year>). <article-title>The impact on midlevel vision of statistically optimal divisive normalization in v1</article-title>. <source>J. Vis</source>. <volume>13</volume>, <fpage>13</fpage>&#x02013;<lpage>13</lpage>. doi: <pub-id pub-id-type="doi">10.1167/13.8.13</pub-id><pub-id pub-id-type="pmid">23857950</pub-id></mixed-citation>
</ref>
<ref id="B22">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Daly</surname> <given-names>S.</given-names></name></person-group> (<year>1990</year>). <article-title>Application of a noise-adaptive Contrast Sensitivity Function to image data compression</article-title>. <source>Opt. Eng</source>. <volume>29</volume>, <fpage>977</fpage>&#x02013;<lpage>987</lpage>. doi: <pub-id pub-id-type="doi">10.1117/12.55666</pub-id></mixed-citation>
</ref>
<ref id="B23">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Daly</surname> <given-names>S.</given-names></name></person-group> (<year>1993</year>). <article-title>&#x0201C;Visible differences predictor: an algorithm for the assessment of image fidelity,&#x0201D;</article-title> in <source>Digital Images and Human Vision</source>, ed. A. Watson (Cambridge, MA: MIT Press), <fpage>179</fpage>&#x02013;<lpage>206</lpage>. doi: <pub-id pub-id-type="doi">10.1117/12.135952</pub-id></mixed-citation>
</ref>
<ref id="B24">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>D&#x000ED;ez-Ajenjo</surname> <given-names>M.</given-names></name> <name><surname>Capilla</surname> <given-names>P.</given-names></name> <name><surname>Luque</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Red-green vs. blue-yellow spatio-temporal contrast sensitivity across the visual field</article-title>. <source>J. Mod. Opt</source>. <volume>58</volume>, <fpage>1736</fpage>&#x02013;<lpage>1748</lpage>. doi: <pub-id pub-id-type="doi">10.1080/09500340.2011.606374</pub-id></mixed-citation>
</ref>
<ref id="B25">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ding</surname> <given-names>K.</given-names></name> <name><surname>Ma</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E. P.</given-names></name></person-group> (<year>2022</year>). <article-title>Image quality assessment: Unifying structure and texture similarity</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>44</volume>, <fpage>2567</fpage>&#x02013;<lpage>2581</lpage>. doi: <pub-id pub-id-type="doi">10.1109/tpami.2020.3045810</pub-id><pub-id pub-id-type="pmid">33338012</pub-id></mixed-citation>
</ref>
<ref id="B26">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Duda</surname> <given-names>R.</given-names></name> <name><surname>Hart</surname> <given-names>P.</given-names></name></person-group> (<year>1973</year>). <source>Pattern Classification and Scene Analysis</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>John Wiley and Sons</publisher-name>.</mixed-citation>
</ref>
<ref id="B27">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Eckstein</surname> <given-names>M.</given-names></name> <name><surname>Ahumada</surname> <given-names>A.</given-names></name></person-group> (<year>2002</year>). <article-title>Classification images: a tool to analyze visual strategies</article-title>. <source>J. Vis</source>. 2. doi: <pub-id pub-id-type="doi">10.1167/2.1.i</pub-id><pub-id pub-id-type="pmid">12678601</pub-id></mixed-citation>
</ref>
<ref id="B28">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fairchild</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <source>Color Appearance Models. The Wiley-IS &#x00026; T Series in Imaging Science and Technology</source>. Wiley. doi: <pub-id pub-id-type="doi">10.1002/9781118653128</pub-id></mixed-citation>
</ref>
<ref id="B29">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Foley</surname> <given-names>J. M.</given-names></name></person-group> (<year>1994</year>). <article-title>Human luminance pattern-vision mechanisms: masking experiments require a new model</article-title>. <source>J. Opt. Soc. Am. A</source> <volume>11</volume>, <fpage>1710</fpage>&#x02013;<lpage>1719</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSAA.11.001710</pub-id><pub-id pub-id-type="pmid">8046537</pub-id></mixed-citation>
</ref>
<ref id="B30">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Gatys</surname> <given-names>L. A.</given-names></name> <name><surname>Ecker</surname> <given-names>A. S.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Image style transfer using convolutional neural networks,&#x0201D;</article-title> in <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2414</fpage>&#x02013;<lpage>2423</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2016.265</pub-id></mixed-citation>
</ref>
<ref id="B31">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Georgeson</surname> <given-names>M.</given-names></name> <name><surname>Sullivan</surname> <given-names>G.</given-names></name></person-group> (<year>1975</year>). <article-title>Contrast constancy: deblurring in human vision by spatial frequency channels</article-title>. <source>J. Physiol</source>. <volume>252</volume>, <fpage>627</fpage>&#x02013;<lpage>656</lpage>. doi: <pub-id pub-id-type="doi">10.1113/jphysiol.1975.sp011162</pub-id><pub-id pub-id-type="pmid">1206570</pub-id></mixed-citation>
</ref>
<ref id="B32">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gomez-Villa</surname> <given-names>A.</given-names></name> <name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2020a</year>). <article-title>Visual information flow in Wilson&#x02013;Cowan networks</article-title>. <source>J. Neurophysiol</source>. <volume>123</volume>, <fpage>2249</fpage>&#x02013;<lpage>2268</lpage>. doi: <pub-id pub-id-type="doi">10.1152/jn.00487.2019</pub-id><pub-id pub-id-type="pmid">32159407</pub-id></mixed-citation>
</ref>
<ref id="B33">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gomez-Villa</surname> <given-names>A.</given-names></name> <name><surname>Martin</surname> <given-names>A.</given-names></name> <name><surname>Vazquez</surname> <given-names>J.</given-names></name> <name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2020b</year>). <article-title>Color illusions also deceive CNNs for low-level vision tasks: analysis and implications</article-title>. <source>Vision Res</source>. <volume>176</volume>, <fpage>156</fpage>&#x02013;<lpage>174</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.visres.2020.07.010</pub-id><pub-id pub-id-type="pmid">32896717</pub-id></mixed-citation>
</ref>
<ref id="B34">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gomez-Villa</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Parraga</surname> <given-names>C.</given-names></name> <name><surname>Twardowski</surname> <given-names>B.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Vazquez-Corral</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>The art of deception: color visual illusions and diffusion models</article-title>. <source>IEEE Comp. Vis. Patt. Recogn</source>. (<italic>CVPR)</italic> doi: <pub-id pub-id-type="doi">10.1109/CVPR52734.2025.01737</pub-id></mixed-citation>
</ref>
<ref id="B35">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Goodale</surname> <given-names>M.</given-names></name> <name><surname>Milner</surname> <given-names>A.</given-names></name> <name><surname>Jakobson</surname> <given-names>L.</given-names></name> <name><surname>Carey</surname> <given-names>D.</given-names></name></person-group> (<year>1991</year>). <article-title>A neurological dissociation between perceiving objects and grasping them</article-title>. <source>Nature</source> <volume>349</volume>, <fpage>154</fpage>&#x02013;<lpage>156</lpage>. doi: <pub-id pub-id-type="doi">10.1038/349154a0</pub-id><pub-id pub-id-type="pmid">1986306</pub-id></mixed-citation>
</ref>
<ref id="B36">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Graham</surname> <given-names>N. V. S.</given-names></name></person-group> (<year>1989</year>). <source>Visual Pattern Analyzers</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press, xvi 646</publisher-name>. doi: <pub-id pub-id-type="doi">10.1093/acprof:oso/9780195051544.001.0001</pub-id></mixed-citation>
</ref>
<ref id="B37">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Guti&#x000E9;rrez</surname> <given-names>J.</given-names></name> <name><surname>Ferri</surname> <given-names>F.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Regularization operators for natural images based on nonlinear perception models</article-title>. <source>IEEE Tr. Im. Proc</source>. <volume>15</volume>, <fpage>189</fpage>&#x02013;<lpage>200</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2005.860345</pub-id><pub-id pub-id-type="pmid">16435549</pub-id></mixed-citation>
</ref>
<ref id="B38">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gutmann</surname> <given-names>M. U.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Hyv&#x000E4;rinen</surname> <given-names>A.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Spatio-chromatic adaptation via higher-order canonical correlation analysis of natural images</article-title>. <source>PLoS ONE</source> <volume>9</volume>:<fpage>e86481</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0086481</pub-id><pub-id pub-id-type="pmid">24533049</pub-id></mixed-citation>
</ref>
<ref id="B39">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hammou</surname> <given-names>D.</given-names></name> <name><surname>Cai</surname> <given-names>Y.</given-names></name> <name><surname>Madhusudanarao</surname> <given-names>P.</given-names></name> <name><surname>Bampis</surname> <given-names>C. G.</given-names></name> <name><surname>Mantiuk</surname> <given-names>R. K.</given-names></name></person-group> (<year>2025</year>). Do image and video quality metrics model low-level human vision? <italic>ArXiV: 2503.16264</italic>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2503.16264</pub-id></mixed-citation>
</ref>
<ref id="B40">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Hepburn</surname> <given-names>A.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>McConville</surname> <given-names>R.</given-names></name> <name><surname>Santos-Rodriguez</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Perceptnet: a human visual system inspired neural network for estimating perceptual distance,&#x0201D;</article-title> in <source>2020 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Abu Dhabi</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>121</fpage>&#x02013;<lpage>125</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICIP40778.2020.9190691</pub-id></mixed-citation>
</ref>
<ref id="B41">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hepburn</surname> <given-names>A.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Santos-Rodriguez</surname> <given-names>R.</given-names></name> <name><surname>Ball&#x000E9;</surname> <given-names>J.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;On the relation between statistical learning and perceptual distances,&#x0201D;</article-title> in <source>International Conference on Learning Representations (ICLR)</source>.</mixed-citation>
</ref>
<ref id="B42">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Daud&#x000E9;n</surname> <given-names>P.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Alignment of color discrimination in humans and image segmentation networks</article-title>. <source>Front. Psychol</source>. <volume>15</volume>:<fpage>1415958</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2024.1415958</pub-id><pub-id pub-id-type="pmid">39507086</pub-id></mixed-citation>
</ref>
<ref id="B43">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Gomez-Villa</surname> <given-names>A.</given-names></name> <name><surname>Ja&#x000E9;n-Lorites</surname> <given-names>J. M.</given-names></name> <name><surname>Vila-Tom&#x000E1;s</surname> <given-names>J.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name></person-group> (<year>2025a</year>). <article-title>Contrast sensitivity function of multimodal vision-language models</article-title>. <source>arXiv preprint arXiv:2508.10367</source>.</mixed-citation>
</ref>
<ref id="B44">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Vila-Tom&#x000E1;s</surname> <given-names>J.</given-names></name> <name><surname>Dauden-Oliver</surname> <given-names>P.</given-names></name> <name><surname>Alabau-Bosque</surname> <given-names>N.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2025b</year>). <article-title>Why divisive normalization works in image segmentation?</article-title> <source>Neurocomputing</source> <volume>649</volume>:<fpage>130569</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2025.130569</pub-id></mixed-citation>
</ref>
<ref id="B45">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Vila-Tom&#x000E1;s</surname> <given-names>J.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Neural networks with divisive normalization for image segmentation</article-title>. <source>Patt. Rec. Lett</source>. <volume>173</volume>, <fpage>64</fpage>&#x02013;<lpage>71</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.patrec.2023.07.017</pub-id></mixed-citation>
</ref>
<ref id="B46">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Vila-Tomas</surname> <given-names>J.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2025c</year>). <article-title>Dissecting the effectiveness of deep features as metric of perceptual image quality</article-title>. <source>Neural Netw</source>. <volume>185</volume>:<fpage>107189</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neunet.2025.107189</pub-id><pub-id pub-id-type="pmid">39874824</pub-id></mixed-citation>
</ref>
<ref id="B47">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hubel</surname> <given-names>D. H.</given-names></name> <name><surname>Wiesel</surname> <given-names>T. N.</given-names></name></person-group> (<year>1959</year>). <article-title>Receptive fields of single neurones in the cat&#x00027;s striate cortex</article-title>. <source>J. Physiol</source>. <volume>148</volume>, <fpage>574</fpage>&#x02013;<lpage>591</lpage>. doi: <pub-id pub-id-type="doi">10.1113/jphysiol.1959.sp006308</pub-id><pub-id pub-id-type="pmid">14403679</pub-id></mixed-citation>
</ref>
<ref id="B48">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hubel</surname> <given-names>D. H.</given-names></name> <name><surname>Wiesel</surname> <given-names>T. N.</given-names></name></person-group> (<year>1961</year>). <article-title>Integrative action in the cat&#x00027;s lateral geniculate body</article-title>. <source>J. Physiol</source>. <volume>155</volume>:<fpage>385</fpage>. doi: <pub-id pub-id-type="doi">10.1113/jphysiol.1961.sp006635</pub-id><pub-id pub-id-type="pmid">13716436</pub-id></mixed-citation>
</ref>
<ref id="B49">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hurvich</surname> <given-names>L. M.</given-names></name> <name><surname>Jameson</surname> <given-names>D.</given-names></name></person-group> (<year>1957</year>). <article-title>An opponent-process theory of color vision</article-title>. <source>Psychol. Rev</source>. <volume>64</volume>(<issue>Part 1</issue>), <fpage>384</fpage>&#x02013;<lpage>404</lpage>. doi: <pub-id pub-id-type="doi">10.1037/h0041403</pub-id><pub-id pub-id-type="pmid">13505974</pub-id></mixed-citation>
</ref>
<ref id="B50">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hyvarinen</surname> <given-names>A.</given-names></name> <name><surname>Hurri</surname> <given-names>J.</given-names></name> <name><surname>Hoyer</surname> <given-names>P.</given-names></name></person-group> (<year>2009</year>). <source>Natural Image Statistics: A Probabilistic Approach to Early Computational Vision</source>. London: Springer. doi: <pub-id pub-id-type="doi">10.1007/978-1-84882-491-1</pub-id></mixed-citation>
</ref>
<ref id="B51">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Karklin</surname> <given-names>Y.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E.</given-names></name></person-group> (<year>2011</year>). <article-title>Efficient coding of natural images with a population of noisy linear-nonlinear neurons</article-title>. <source>NIPS</source> 24. <pub-id pub-id-type="pmid">26273180</pub-id></mixed-citation>
</ref>
<ref id="B52">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kelly</surname> <given-names>D. H.</given-names></name></person-group> (<year>1979</year>). <article-title>Motion and vision. ii. stabilized spatio-temporal threshold surface</article-title>. <source>J. Opt. Soc. Am</source>. <volume>69</volume>, <fpage>1340</fpage>&#x02013;<lpage>1349</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSA.69.001340</pub-id><pub-id pub-id-type="pmid">521853</pub-id></mixed-citation>
</ref>
<ref id="B53">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Krauskopf</surname> <given-names>J.</given-names></name> <name><surname>Gegenfurtner</surname> <given-names>K.</given-names></name></person-group> (<year>1992</year>). <article-title>Color discrimination and adaptation</article-title>. <source>Vis. Res</source>. <volume>32</volume>, <fpage>2165</fpage>&#x02013;<lpage>2175</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0042-6989(92)90077-V</pub-id><pub-id pub-id-type="pmid">1304093</pub-id></mixed-citation>
</ref>
<ref id="B54">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kreiman</surname> <given-names>G.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name> <name><surname>Fried</surname> <given-names>I.</given-names></name></person-group> (<year>2000</year>). <article-title>Category specific visual responses of single neurons in the human medial temporal lobe</article-title>. <source>Nat. Neurosci</source>. <volume>3</volume>, <fpage>946</fpage>&#x02013;<lpage>953</lpage>. doi: <pub-id pub-id-type="doi">10.1038/78868</pub-id><pub-id pub-id-type="pmid">10966627</pub-id></mixed-citation>
</ref>
<ref id="B55">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2012a</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>NIPS</source> 25.</mixed-citation>
</ref>
<ref id="B56">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012b</year>). <article-title>&#x0201C;Imagenet classification with deep convolutional neural networks,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, vol 25, eds. F. Pereira, C. Burges, L. Bottou, and K. Weinberger (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>).</mixed-citation>
</ref>
<ref id="B57">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kubilius</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Brain-like object recognition with high-performing shallow recurrent anns</article-title>. <source>ICLR, Arxiv: 1909.06161</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1909.06161</pub-id></mixed-citation>
</ref>
<ref id="B58">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Berardino</surname> <given-names>A.</given-names></name> <name><surname>Ball&#x000E9;</surname> <given-names>J.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E. P.</given-names></name></person-group> (<year>2017</year>). <article-title>Perceptually optimized image rendering</article-title>. <source>J. Opt. Soc. Am. A</source> <volume>34</volume>, <fpage>1511</fpage>&#x02013;<lpage>1525</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSAA.34.001511</pub-id><pub-id pub-id-type="pmid">29036154</pub-id></mixed-citation>
</ref>
<ref id="B59">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Jim&#x000E9;nez</surname> <given-names>S.</given-names></name> <name><surname>Camps-Valls</surname> <given-names>G.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Nonlinearities and adaptation of color vision from Sequential Principal Curves Analysis</article-title>. <source>Neural Comp</source>. <volume>24</volume>, <fpage>2751</fpage>&#x02013;<lpage>2788</lpage>. doi: <pub-id pub-id-type="doi">10.1162/NECO_a_00342</pub-id><pub-id pub-id-type="pmid">22845821</pub-id></mixed-citation>
</ref>
<ref id="B60">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Johnson</surname> <given-names>J.</given-names></name> <name><surname>Camps-Valls</surname> <given-names>G.</given-names></name> <name><surname>Santos-Rodriguez</surname> <given-names>R.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2025</year>). <article-title>Estimating information theoretic measures via multidimensional gaussianization</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell</source>. <volume>47</volume>, <fpage>1293</fpage>&#x02013;<lpage>1308</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2024.3495827</pub-id><pub-id pub-id-type="pmid">39527441</pub-id></mixed-citation>
</ref>
<ref id="B61">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Visual aftereffects and sensory nonlinearities from a single statistical framework</article-title>. <source>Front. Hum. Neurosci</source>. <volume>9</volume>:<fpage>557</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fnhum.2015.00557</pub-id><pub-id pub-id-type="pmid">26528165</pub-id></mixed-citation>
</ref>
<ref id="B62">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Mu noz Mar&#x000ED;</surname> <given-names>J.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Divisive normalization image quality metric revisited</article-title>. <source>JOSA A</source> <volume>27</volume>, <fpage>852</fpage>&#x02013;<lpage>864</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSAA.27.000852</pub-id><pub-id pub-id-type="pmid">20360827</pub-id></mixed-citation>
</ref>
<ref id="B63">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laughlin</surname> <given-names>S.</given-names></name></person-group> (<year>1981</year>). <article-title>A simple coding procedure enhances a neuron&#x00027;s information capacity</article-title>. <source>Zeitschrift F&#x000FC;r Naturforschung C</source> <volume>36</volume>, <fpage>910</fpage>&#x02013;<lpage>912</lpage>. doi: <pub-id pub-id-type="doi">10.1515/znc-1981-9-1040</pub-id><pub-id pub-id-type="pmid">7303823</pub-id></mixed-citation>
</ref>
<ref id="B64">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Le Gall</surname> <given-names>D. J.</given-names></name></person-group> (<year>1992</year>). <article-title>The MPEG video compression algorithm</article-title>. <source>Signal Process.: Image Commun</source>. <volume>4</volume>, <fpage>129</fpage>&#x02013;<lpage>140</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0923-5965(92)90019-C</pub-id></mixed-citation>
</ref>
<ref id="B65">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Legge</surname> <given-names>G.</given-names></name></person-group> (<year>1981</year>). <article-title>A power law for contrast discrimination</article-title>. <source>Vis. Res</source>. <volume>18</volume>, <fpage>68</fpage>&#x02013;<lpage>91</lpage>. <pub-id pub-id-type="pmid">7269325</pub-id></mixed-citation>
</ref>
<ref id="B66">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Legge</surname> <given-names>G.</given-names></name> <name><surname>Foley</surname> <given-names>J.</given-names></name></person-group> (<year>1980</year>). <article-title>Contrast masking in human vision</article-title>. <source>J. Opt. Soc. Am</source>. <volume>70</volume>, <fpage>1458</fpage>&#x02013;<lpage>1471</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSA.70.001458</pub-id><pub-id pub-id-type="pmid">7463185</pub-id></mixed-citation>
</ref>
<ref id="B67">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Lengyel</surname> <given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>Marr&#x00027;s three levels of analysis are useful as a framework for neuroscience</article-title>. <source>J. Physiol</source>. <volume>602</volume>, <fpage>1911</fpage>&#x02013;<lpage>1914</lpage>. doi: <pub-id pub-id-type="doi">10.1113/JP279549</pub-id><pub-id pub-id-type="pmid">38628044</pub-id></mixed-citation>
</ref>
<ref id="B68">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Gomez-Villa</surname> <given-names>A.</given-names></name> <name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Contrast sensitivity functions in autoencoders</article-title>. <source>J. Vis</source>. 22. doi: <pub-id pub-id-type="doi">10.1167/jov.22.6.8</pub-id><pub-id pub-id-type="pmid">35587354</pub-id></mixed-citation>
</ref>
<ref id="B69">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Steeg</surname> <given-names>G. V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Functional connectivity via total correlation: analytical results in visual areas</article-title>. <source>Neurocomputing</source> <volume>571</volume>:<fpage>127143</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2023.127143</pub-id></mixed-citation>
</ref>
<ref id="B70">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Lindsey</surname> <given-names>J.</given-names></name> <name><surname>Ocko</surname> <given-names>S.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name> <name><surname>Deny</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <source>The Effects of Neural Resource Constraints on Early Visual Representations</source>. ICLR.</mixed-citation>
</ref>
<ref id="B71">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Logothetis</surname> <given-names>N.</given-names></name> <name><surname>Sheinberg</surname> <given-names>D.</given-names></name></person-group> (<year>1996</year>). <article-title>Visual object recognition</article-title>. <source>Ann. Rev. Neurosci</source>. <volume>19</volume>, <fpage>577</fpage>&#x02013;<lpage>621</lpage>. doi: <pub-id pub-id-type="doi">10.1146/annurev.ne.19.030196.003045</pub-id></mixed-citation>
</ref>
<ref id="B72">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Loxley</surname> <given-names>P. N.</given-names></name></person-group> (<year>2017</year>). <article-title>The two-dimensional gabor function adapted to natural image statistics: a model of simple-cell receptive fields and sparse structure in images</article-title>. <source>Neural Comput</source>. <volume>29</volume>, <fpage>2769</fpage>&#x02013;<lpage>2799</lpage>. doi: <pub-id pub-id-type="doi">10.1162/neco_a_00997</pub-id><pub-id pub-id-type="pmid">28777727</pub-id></mixed-citation>
</ref>
<ref id="B73">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Urtasun</surname> <given-names>R.</given-names></name> <name><surname>Zemel</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>Understanding the effective receptive field in deep convolutional neural networks</article-title>. <source>arXiv [Preprint]</source>. arXiv:1701.04128. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1701.04128</pub-id></mixed-citation>
</ref>
<ref id="B74">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Macpherson</surname> <given-names>T.</given-names></name> <name><surname>Churchland</surname> <given-names>A.</given-names></name> <name><surname>Sejnowski</surname> <given-names>T.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J.</given-names></name> <name><surname>Kamitani</surname> <given-names>Y.</given-names></name> <name><surname>Takahashi</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Natural and artificial intelligence: a brief introduction to the interplay between ai and neuroscience research</article-title>. <source>Neural Netw</source>. <volume>144</volume>, <fpage>603</fpage>&#x02013;<lpage>613</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neunet.2021.09.018</pub-id><pub-id pub-id-type="pmid">34649035</pub-id></mixed-citation>
</ref>
<ref id="B75">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mahendran</surname> <given-names>A.</given-names></name> <name><surname>Vedaldi</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Visualizing deep convolutional neural networks using natural pre-images</article-title>. <source>Int. J. Comput. Vis</source>. <volume>120</volume>, <fpage>233</fpage>&#x02013;<lpage>255</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11263-016-0911-8</pub-id></mixed-citation>
</ref>
<ref id="B76">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Spatio-chromatic information available from different neural layers via gaussianization</article-title>. <source>J. Math. Neurosci</source>. <volume>10</volume>:<fpage>18</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s13408-020-00095-8</pub-id><pub-id pub-id-type="pmid">33175257</pub-id></mixed-citation>
</ref>
<ref id="B77">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Information flow in biological networks for color vision</article-title>. <source>Entropy</source> <volume>24</volume>:<fpage>1442</fpage>. doi: <pub-id pub-id-type="doi">10.3390/e24101442</pub-id><pub-id pub-id-type="pmid">37420462</pub-id></mixed-citation>
</ref>
<ref id="B78">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Bowers</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <source>The Low-Level Mindset: Compelling Low-Level Visual Psychophysics to Evaluate Image Computable Vision Models</source>. Invited talk, Psychol. Dept. University of Bristol.</mixed-citation>
</ref>
<ref id="B79">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Epifanio</surname> <given-names>I.</given-names></name> <name><surname>Navarro</surname> <given-names>R.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E.</given-names></name></person-group> (<year>2006</year>). <article-title>Non-linear image representation for efficient perceptual coding</article-title>. <source>IEEE Trans. Image Process</source>. <volume>15</volume>, <fpage>68</fpage>&#x02013;<lpage>80</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2005.860325</pub-id></mixed-citation>
</ref>
<ref id="B80">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Esteve-Taboada</surname> <given-names>J.</given-names></name> <name><surname>Aguilar</surname> <given-names>G.</given-names></name> <name><surname>Maertens</surname> <given-names>M.</given-names></name> <name><surname>Wichmann</surname> <given-names>F.</given-names></name></person-group> (<year>2025</year>). <article-title>Estimating the contribution of early and late noise in vision from psychophysical data</article-title>. <source>J. Vis</source>. <volume>25</volume>, <fpage>12</fpage>&#x02013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1167/jov.25.1.12</pub-id><pub-id pub-id-type="pmid">39804617</pub-id></mixed-citation>
</ref>
<ref id="B81">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Esteve-Taboada</surname> <given-names>J.</given-names></name> <name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name></person-group> (<year>2024</year>). <article-title>Cortical divisive normalization from wilson-cowan neural dynamics</article-title>. <source>J. Nonlinear Sci</source>. <volume>34</volume>:<fpage>35</fpage>. doi: <pub-id pub-id-type="doi">10.1007/s00332-023-10009-z</pub-id></mixed-citation>
</ref>
<ref id="B82">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Ferri</surname> <given-names>F.</given-names></name> <name><surname>Albert</surname> <given-names>J. J.</given-names></name> <name><surname>Soret Artigas</surname> <given-names>J.</given-names></name></person-group> (<year>2000a</year>). <article-title>The role of perceptual contrast non-linearities in image transform coding</article-title>. <source>Image Vis. Comput</source>. <volume>18</volume>, <fpage>233</fpage>&#x02013;<lpage>246</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0262-8856(99)00010-4</pub-id></mixed-citation>
</ref>
<ref id="B83">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Ferri</surname> <given-names>F.</given-names></name> <name><surname>Gutierrez</surname> <given-names>J.</given-names></name> <name><surname>Epifanio</surname> <given-names>I.</given-names></name></person-group> (<year>2000b</year>). <article-title>Importance of quantizer design compared to optimal multigrid motion estimation in video coding</article-title>. <source>Electron. Lett</source>. <volume>36</volume>, <fpage>507</fpage>&#x02013;<lpage>509</lpage>. doi: <pub-id pub-id-type="doi">10.1049/el:20000645</pub-id></mixed-citation>
</ref>
<ref id="B84">
<mixed-citation publication-type="web"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Gutierrez</surname> <given-names>J.</given-names></name></person-group> (<year>2002</year>). <source>VistaLab: The Matlab Toolbox for Linear Spatio-Temporal Vision Models</source>. Univ. Valencia. Available online at: <ext-link ext-link-type="uri" xlink:href="https://isp.uv.es/code/vision_and_color/colorlab/vistalab/">https://isp.uv.es/code/vision_and_color/colorlab/vistalab/</ext-link> (Accessed December 11, 2025).</mixed-citation>
</ref>
<ref id="B85">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Guti&#x000E9;rrez</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>V1 non-linear properties emerge from local-to-global non-linear ICA</article-title>. <source>Netw.: Comput. Neural Syst</source>. <volume>17</volume>, <fpage>85</fpage>&#x02013;<lpage>102</lpage>. doi: <pub-id pub-id-type="doi">10.1080/09548980500439602</pub-id><pub-id pub-id-type="pmid">16613796</pub-id></mixed-citation>
</ref>
<ref id="B86">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Gutierrez</surname> <given-names>J.</given-names></name> <name><surname>Epifanio</surname> <given-names>I.</given-names></name> <name><surname>Ferri</surname> <given-names>F. J.</given-names></name> <name><surname>Artigas</surname> <given-names>J. M.</given-names></name></person-group> (<year>2001</year>). <article-title>Perceptual feed-back in multigrid motion estimation using an improved DCT quantization</article-title>. <source>IEEE Trans. Image Process</source>. <volume>10</volume>, <fpage>1411</fpage>&#x02013;<lpage>1427</lpage>. doi: <pub-id pub-id-type="doi">10.1109/83.951528</pub-id></mixed-citation>
</ref>
<ref id="B87">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Gutierrez</surname> <given-names>J.</given-names></name> <name><surname>Epifanio</surname> <given-names>I.</given-names></name> <name><surname>Ferri</surname> <given-names>F.</given-names></name></person-group> (<year>2000c</year>). <article-title>Perceptually weighted optical flow for motion-based segmentation in MPEG-4 paradigm</article-title>. <source>Electron. Lett</source>. <volume>36</volume>, <fpage>1693</fpage>&#x02013;<lpage>1694</lpage>. doi: <pub-id pub-id-type="doi">10.1049/el:20001222</pub-id></mixed-citation>
</ref>
<ref id="B88">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Guti&#x000E9;rrez</surname> <given-names>J.</given-names></name> <name><surname>Rovira</surname> <given-names>J.</given-names></name></person-group> (<year>2004</year>). <article-title>&#x0201C;Perturbation analysis of the changes in V1 receptive fields due to context,&#x0201D;</article-title> in <source>Gordon Research Conference: Sensory Coding and the Natural Environment</source> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Gordon Research Conference</publisher-name>).</mixed-citation>
</ref>
<ref id="B89">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name></person-group> (<year>2024</year>). <article-title>A separate theory-on-top level may be inspiring, but it is neither separate nor enough</article-title>. <source>J. Physiol</source>. <volume>602</volume>, <fpage>1918</fpage>&#x02013;<lpage>1918</lpage>.</mixed-citation>
</ref>
<ref id="B90">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Kheravdar</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name></person-group> (<year>2021</year>). <article-title>Visual information fidelity with better vision models and better mutual information estimates</article-title>. <source>J. Vis</source>. <volume>21</volume>:<fpage>2351</fpage>. doi: <pub-id pub-id-type="doi">10.1167/jov.21.9.2351</pub-id></mixed-citation>
</ref>
<ref id="B91">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name></person-group> (<year>2010</year>). <article-title>Psychophysically tuned divisive normalization approximately factorizes the pdf of natural images</article-title>. <source>Neural Comput</source>. <volume>22</volume>, <fpage>3179</fpage>&#x02013;<lpage>3206</lpage>. doi: <pub-id pub-id-type="doi">10.1162/NECO_a_00046</pub-id><pub-id pub-id-type="pmid">20858127</pub-id></mixed-citation>
</ref>
<ref id="B92">
<mixed-citation publication-type="web"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Luque</surname> <given-names>M.</given-names></name></person-group> (<year>2002</year>). <source>ColorLab: A Matlab Toolbox for Color Science and Calibrated Color Image Processing</source>. Univ. Valencia. Available online at: <ext-link ext-link-type="uri" xlink:href="https://isp.uv.es/code/vision_and_color/colorlab/content/">https://isp.uv.es/code/vision_and_color/colorlab/content/</ext-link> (Accessed December 11, 2025).</mixed-citation>
</ref>
<ref id="B93">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Pons</surname> <given-names>A.</given-names></name> <name><surname>Artigas</surname> <given-names>J.</given-names></name></person-group> (<year>1995</year>). <article-title>Bit allocation algorithm for codebook design in vector quantization fully based on human visual system non-linearities for suprathreshold contrasts</article-title>. <source>Electron. Lett</source>. <volume>31</volume>, <fpage>1229</fpage>&#x02013;<lpage>1231</lpage>. doi: <pub-id pub-id-type="doi">10.1049/el:19950863</pub-id></mixed-citation>
</ref>
<ref id="B94">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Pons</surname> <given-names>A.</given-names></name> <name><surname>Artigas</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Subjective image fidelity metric based on bit allocation of the human visual system in the dct domain</article-title>. <source>Image Vis. Comput</source>. <volume>15</volume>, <fpage>535</fpage>&#x02013;<lpage>548</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0262-8856(96)00004-2</pub-id></mixed-citation>
</ref>
<ref id="B95">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E. P.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Geometrical and statistical properties of vision models obtained via maximum differentiation,&#x0201D;</article-title> in <source>Human Vision and Electronic Imaging XX, volume 9394 of SPIE</source>, eds. B. Rogowitz, T. Pappas, and H. de Ridder (<publisher-loc>San Francisco, CA</publisher-loc>: <publisher-name>SPIE</publisher-name>), 93940L. doi: <pub-id pub-id-type="doi">10.1117/12.2085653</pub-id></mixed-citation>
</ref>
<ref id="B96">
<mixed-citation publication-type="web"><person-group person-group-type="author"><name><surname>Malo</surname> <given-names>J.</given-names></name> <name><surname>Vila-Tom&#x000E1;s</surname> <given-names>J.</given-names></name> <name><surname>Hernandez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;A turing test for artificial nets devoted to vision,&#x0201D;</article-title> in <source>Invited talk at the Artif. Intell. Evaluation Workshop, AI Dept. University of Bristol, June 2022</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://isp.uv.es/docs/talk_AI_Bristol_Malo_et_al_2022.pdf">http://isp.uv.es/docs/talk_AI_Bristol_Malo_et_al_2022.pdf</ext-link> (Accessed December 11, 2025).</mixed-citation>
</ref>
<ref id="B97">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name></person-group> (<year>1978</year>). <source>Vision: A Computational Investigation into the Human Representation and Processing of Visual Information</source>. New York, NY: W.H. Freeman and Co.</mixed-citation>
</ref>
<ref id="B98">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Marr</surname> <given-names>D.</given-names></name> <name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>1977</year>). <article-title>From understanding computation to understanding neural circuitry</article-title>. <source>Neurosci. Res. Prog. Bull</source>. <volume>15</volume>, <fpage>470</fpage>&#x02013;<lpage>488</lpage>.</mixed-citation>
</ref>
<ref id="B99">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Martinez</surname> <given-names>M.</given-names></name> <name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>In praise of artifice reloaded: caution with natural image databases in modeling vision</article-title>. <source>Front. Neurosci</source>. <volume>13</volume>:<fpage>8</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fnins.2019.00008</pub-id><pub-id pub-id-type="pmid">30894796</pub-id></mixed-citation>
</ref>
<ref id="B100">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Martinez</surname> <given-names>M.</given-names></name> <name><surname>Cyriac</surname> <given-names>P.</given-names></name> <name><surname>Batard</surname> <given-names>T.</given-names></name> <name><surname>Bertalm&#x000ED;o</surname> <given-names>M.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Derivatives and inverse of cascaded linear&#x0002B;nonlinear neural models</article-title>. <source>PLoS ONE</source> <volume>13</volume>:<fpage>e0201326</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0201326</pub-id><pub-id pub-id-type="pmid">30321175</pub-id></mixed-citation>
</ref>
<ref id="B101">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Martinez</surname> <given-names>M.</given-names></name> <name><surname>Martinez</surname> <given-names>L.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Topographic independent component analysis reveals random scrambling of orientation in visual space</article-title>. <source>PLoS ONE</source> <volume>12</volume>:<fpage>e0178345</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0178345</pub-id><pub-id pub-id-type="pmid">28640816</pub-id></mixed-citation>
</ref>
<ref id="B102">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Martinez-Uriegas</surname> <given-names>E.</given-names></name></person-group> (<year>1997</year>). <article-title>&#x0201C;Color detection and color contrast discrimination thresholds,&#x0201D;</article-title> in <source>Proc. OSA Meeting</source>, 81.</mixed-citation>
</ref>
<ref id="B103">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mehrer</surname> <given-names>J.</given-names></name> <name><surname>Spoerer</surname> <given-names>C. J.</given-names></name> <name><surname>Jones</surname> <given-names>E. C.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Kietzmann</surname> <given-names>T. C.</given-names></name></person-group> (<year>2021</year>). <article-title>An ecologically motivated image dataset for deep learning yields better models of human vision</article-title>. <source>Proc. Nat. Acad. Sci. U. S. A</source>. <volume>118</volume>:<fpage>e2011417118</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.2011417118</pub-id><pub-id pub-id-type="pmid">33593900</pub-id></mixed-citation>
</ref>
<ref id="B104">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>M.</given-names></name> <name><surname>Chung</surname> <given-names>S.</given-names></name> <name><surname>Miller</surname> <given-names>K. D.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Divisive feature normalization improves image recognition performance in alexnet,&#x0201D;</article-title> in <source>Int. Conf. Learn. Repres</source>. (ICLR).</mixed-citation>
</ref>
<ref id="B105">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Milner</surname> <given-names>A. D.</given-names></name> <name><surname>Goodale</surname> <given-names>M. A.</given-names></name></person-group> (<year>1992</year>). <article-title>Separate visual pathways for perception and action</article-title>. <source>Trends Neurosci</source>. <volume>15</volume>, <fpage>20</fpage>&#x02013;<lpage>25</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0166-2236(92)90344-8</pub-id><pub-id pub-id-type="pmid">1374953</pub-id></mixed-citation>
</ref>
<ref id="B106">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Monahan</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <source>Numerical Methods of Statistics. Cambridge Series in Statistical and Probabilistic Mathematics</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. doi: <pub-id pub-id-type="doi">10.1017/CBO9780511977176</pub-id></mixed-citation>
</ref>
<ref id="B107">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mullen</surname> <given-names>K. T.</given-names></name></person-group> (<year>1985</year>). <article-title>The CSF of human colour vision to red-green and yellow-blue chromatic gratings</article-title>. <source>J. Physiol</source>. <volume>359</volume>, <fpage>381</fpage>&#x02013;<lpage>400</lpage>. doi: <pub-id pub-id-type="doi">10.1113/jphysiol.1985.sp015591</pub-id></mixed-citation>
</ref>
<ref id="B108">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Null</surname> <given-names>C. H.</given-names></name></person-group> (<year>1993</year>). <source>Digital Images and Human Vision</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</mixed-citation>
</ref>
<ref id="B109">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Olshausen</surname> <given-names>B.</given-names></name> <name><surname>Field</surname> <given-names>D.</given-names></name></person-group> (<year>1996</year>). <article-title>Emergence of simple-cell receptive field properties by learning a sparse code for natural images</article-title>. <source>Nature</source> <volume>281</volume>, <fpage>607</fpage>&#x02013;<lpage>609</lpage>. doi: <pub-id pub-id-type="doi">10.1038/381607a0</pub-id><pub-id pub-id-type="pmid">8637596</pub-id></mixed-citation>
</ref>
<ref id="B110">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Otazu</surname> <given-names>X.</given-names></name> <name><surname>Parraga</surname> <given-names>C.</given-names></name> <name><surname>Vanrell</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Toward a unified chromatic induction model</article-title>. <source>J. Vis</source>. <volume>10</volume>:<fpage>5</fpage>. doi: <pub-id pub-id-type="doi">10.1167/10.12.5</pub-id><pub-id pub-id-type="pmid">21047737</pub-id></mixed-citation>
</ref>
<ref id="B111">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Pillow</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Cross talk opposing view: Marr&#x00027;s three levels of analysis are not useful as a framework for neuroscience</article-title>. <source>J. Physiol</source>. <volume>602</volume>, <fpage>1915</fpage>&#x02013;<lpage>1917</lpage>. doi: <pub-id pub-id-type="doi">10.1113/JP279550</pub-id><pub-id pub-id-type="pmid">38628062</pub-id></mixed-citation>
</ref>
<ref id="B112">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Poggio</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <source>From Marr&#x00027;s Vision to the Problem of Human Intelligence</source>. MIT-CBMM Memos.</mixed-citation>
</ref>
<ref id="B113">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Rajalingham</surname> <given-names>R.</given-names></name> <name><surname>Issa</surname> <given-names>E.</given-names></name> <name><surname>Bashivan</surname> <given-names>P.</given-names></name> <name><surname>Kar</surname> <given-names>K.</given-names></name> <name><surname>Schmidt</surname> <given-names>K.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Large-scale, high-resolution comparison of the core visual object recognition behavior of humans, monkeys, and state-of-the-art deep artificial neural networks</article-title>. <source>J. Neurosci</source>. <volume>38</volume>, <fpage>7255</fpage>&#x02013;<lpage>7269</lpage>. doi: <pub-id pub-id-type="doi">10.1523/JNEUROSCI.0388-18.2018</pub-id><pub-id pub-id-type="pmid">30006365</pub-id></mixed-citation>
</ref>
<ref id="B114">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ringach</surname> <given-names>D.</given-names></name></person-group> (<year>2002</year>). <article-title>Spatial structure and symmetry of simple-cell receptive fields in macaque primary visual cortex</article-title>. <source>J. Neurophysiol</source>. <volume>88</volume>, <fpage>455</fpage>&#x02013;<lpage>463</lpage>. doi: <pub-id pub-id-type="doi">10.1152/jn.2002.88.1.455</pub-id><pub-id pub-id-type="pmid">12091567</pub-id></mixed-citation>
</ref>
<ref id="B115">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ringach</surname> <given-names>D.</given-names></name> <name><surname>Shapley</surname> <given-names>R.</given-names></name></person-group> (<year>2004</year>). <article-title>Reverse correlation in neurophysiology</article-title>. <source>Cognit. Sci</source>. <volume>28</volume>, <fpage>147</fpage>&#x02013;<lpage>166</lpage>. doi: <pub-id pub-id-type="doi">10.1207/s15516709cog2802_2</pub-id></mixed-citation>
</ref>
<ref id="B116">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Romero</surname> <given-names>J.</given-names></name> <name><surname>Garc&#x000ED;a</surname> <given-names>J. A.</given-names></name> <name><surname>del Barco</surname> <given-names>L. J.</given-names></name> <name><surname>Hita</surname> <given-names>E.</given-names></name></person-group> (<year>1993</year>). <article-title>Evaluation of color-discrimination ellipsoids in two-color spaces</article-title>. <source>J. Opt. Soc. Am. A</source> <volume>10</volume>, <fpage>827</fpage>&#x02013;<lpage>837</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSAA.10.000827</pub-id></mixed-citation>
</ref>
<ref id="B117">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ross</surname> <given-names>J.</given-names></name> <name><surname>Speed</surname> <given-names>H. D.</given-names></name> <name><surname>Campbell</surname> <given-names>F. W.</given-names></name></person-group> (<year>1991</year>). <article-title>Contrast adaptation and contrast masking in human vision</article-title>. <source>Proc. R. Soc. Lond. Ser. B: Biol. Sci</source>. <volume>246</volume>, <fpage>61</fpage>&#x02013;<lpage>70</lpage>. doi: <pub-id pub-id-type="doi">10.1098/rspb.1991.0125</pub-id><pub-id pub-id-type="pmid">1684669</pub-id></mixed-citation>
</ref>
<ref id="B118">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Rust</surname> <given-names>N.</given-names></name> <name><surname>Movshon</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>In praise of artifice</article-title>. <source>Nature Neurosci</source>. <volume>8</volume>, <fpage>1647</fpage>&#x02013;<lpage>1650</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nn1606</pub-id><pub-id pub-id-type="pmid">16306892</pub-id></mixed-citation>
</ref>
<ref id="B119">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schrimpf</surname> <given-names>M.</given-names></name> <name><surname>Kubilius</surname> <given-names>J.</given-names></name> <name><surname>Hong</surname> <given-names>H.</given-names></name> <name><surname>Majaj</surname> <given-names>N.</given-names></name> <name><surname>Rajalingham</surname> <given-names>R.</given-names></name> <name><surname>Issa</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2018</year>). Brain-score: which artificial neural network for object recognition is most brain-like? <italic>bioRxiv</italic>. doi: <pub-id pub-id-type="doi">10.1101/407007</pub-id></mixed-citation>
</ref>
<ref id="B120">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sch&#x000FC;tt</surname> <given-names>H.</given-names></name> <name><surname>Wichmann</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>An image-computable psychophysical spatial vision model</article-title>. <source>J. Vis</source>. <volume>17</volume>:<fpage>12</fpage>. doi: <pub-id pub-id-type="doi">10.1167/17.12.12</pub-id><pub-id pub-id-type="pmid">29053781</pub-id></mixed-citation>
</ref>
<ref id="B121">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schwartz</surname> <given-names>O.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E.</given-names></name></person-group> (<year>2001</year>). <article-title>Natural signal statistics and sensory gain control</article-title>. <source>Nat. Neurosci</source>. <volume>4</volume>, <fpage>819</fpage>&#x02013;<lpage>825</lpage>. doi: <pub-id pub-id-type="doi">10.1038/90526</pub-id><pub-id pub-id-type="pmid">11477428</pub-id></mixed-citation>
</ref>
<ref id="B122">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sengar</surname> <given-names>S. S.</given-names></name> <name><surname>Hasan</surname> <given-names>A. B.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2025</year>). <article-title>Generative artificial intelligence: a systematic review and applications</article-title>. <source>Multimedia Tools Applic</source>. <volume>84</volume>, <fpage>23661</fpage>&#x02013;<lpage>23700</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11042-024-20016-1</pub-id></mixed-citation>
</ref>
<ref id="B123">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Shapley</surname> <given-names>R.</given-names></name> <name><surname>Hawken</surname> <given-names>M. J.</given-names></name></person-group> (<year>2011</year>). <article-title>Color in the cortex: single- and double-opponent cells</article-title>. <source>Vis. Res</source>. <volume>51</volume>, <fpage>701</fpage>&#x02013;<lpage>717</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.visres.2011.02.012</pub-id><pub-id pub-id-type="pmid">21333672</pub-id></mixed-citation>
</ref>
<ref id="B124">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sheikh</surname> <given-names>H.</given-names></name> <name><surname>Bovik</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Image information and visual quality</article-title>. <source>IEEE Trans. Image Process</source>. <volume>15</volume>, <fpage>430</fpage>&#x02013;<lpage>444</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2005.859378</pub-id><pub-id pub-id-type="pmid">16479813</pub-id></mixed-citation>
</ref>
<ref id="B125">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sheikh</surname> <given-names>H.</given-names></name> <name><surname>Bovik</surname> <given-names>A.</given-names></name> <name><surname>de Veciana</surname> <given-names>G.</given-names></name></person-group> (<year>2005</year>). <article-title>An information fidelity criterion for image quality assessment using natural scene statistics</article-title>. <source>IEEE Trans. Image Process</source>. <volume>14</volume>, <fpage>2117</fpage>&#x02013;<lpage>2128</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2005.859389</pub-id><pub-id pub-id-type="pmid">16370464</pub-id></mixed-citation>
</ref>
<ref id="B126">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Simoncelli</surname> <given-names>E.</given-names></name> <name><surname>Adelson</surname> <given-names>E.</given-names></name></person-group> (<year>1990</year>). <source>Subband Image Coding, chapter Subband Transforms</source>. <publisher-loc>Norwell, MA</publisher-loc>: <publisher-name>Kluwer Academic Publishers</publisher-name>, <fpage>143</fpage>&#x02013;<lpage>192</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-1-4757-2119-5_4</pub-id></mixed-citation>
</ref>
<ref id="B127">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Storrs</surname> <given-names>K.</given-names></name> <name><surname>Kietzmann</surname> <given-names>T.</given-names></name> <name><surname>Walther</surname> <given-names>A.</given-names></name> <name><surname>Mehrer</surname> <given-names>J.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2021</year>). <article-title>Diverse deep neural networks all predict human inferior temporal cortex well, after training and fitting</article-title>. <source>J. Cogn. Neurosci</source>. <volume>33</volume>, <fpage>2044</fpage>&#x02013;<lpage>2064</lpage>. doi: <pub-id pub-id-type="doi">10.1162/jocn_a_01755</pub-id><pub-id pub-id-type="pmid">34272948</pub-id></mixed-citation>
</ref>
<ref id="B128">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tailby</surname> <given-names>C.</given-names></name> <name><surname>Solomon</surname> <given-names>S.</given-names></name> <name><surname>Dhruv</surname> <given-names>N.</given-names></name> <name><surname>Lennie</surname> <given-names>P.</given-names></name></person-group> (<year>2008</year>). <article-title>Habituation reveals fundamental chromatic mechanisms in striate cortex of macaque</article-title>. <source>J. Neurosci</source>. <volume>28</volume>, <fpage>1131</fpage>&#x02013;<lpage>1139</lpage>. doi: <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4682-07.2008</pub-id><pub-id pub-id-type="pmid">18234891</pub-id></mixed-citation>
</ref>
<ref id="B129">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Taubman</surname> <given-names>D.</given-names></name> <name><surname>Marcellin</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <source>JPEG2000 Image Compression Fundamentals, Standards and Practice</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer Publishing Company, Incorporated</publisher-name>.</mixed-citation>
</ref>
<ref id="B130">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Teo</surname> <given-names>P.</given-names></name> <name><surname>Heeger</surname> <given-names>D.</given-names></name></person-group> (<year>1994</year>). <article-title>Perceptual image distortion</article-title>. <source>Proc. SPIE</source> <volume>2179</volume>, <fpage>127</fpage>&#x02013;<lpage>141</lpage>.</mixed-citation>
</ref>
<ref id="B131">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Isola</surname> <given-names>P.</given-names></name> <name><surname>Freeman</surname> <given-names>W. T.</given-names></name></person-group> (<year>2024</year>). <source>Foundations of Computer Vision</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</mixed-citation>
</ref>
<ref id="B132">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Turing</surname> <given-names>A. M.</given-names></name></person-group> (<year>1950</year>). <article-title>Computing machinery and intelligence</article-title>. <source>Mind</source> <volume>LIX</volume>, <fpage>433</fpage>&#x02013;<lpage>460</lpage>. doi: <pub-id pub-id-type="doi">10.1093/mind/LIX.236.433</pub-id></mixed-citation>
</ref>
<ref id="B133">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Vila-Tom&#x000E1;s</surname> <given-names>J.</given-names></name> <name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Laparra</surname> <given-names>V.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2025</year>). <article-title>Parametric perceptnet: A bio-inspired deep-net trained for image quality assessment</article-title>. <source>ArXiV 2412.03210</source>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.2412.03210</pub-id></mixed-citation>
</ref>
<ref id="B134">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Vila-Tom&#x000E1;s</surname> <given-names>J.</given-names></name> <name><surname>Hern&#x000E1;ndez-C&#x000E1;mara</surname> <given-names>P.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Artificial psychophysics questions classical hue cancellation experiments</article-title>. <source>Front. Neurosci</source>. <volume>17</volume>:<fpage>1208882</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fnins.2023.1208882</pub-id><pub-id pub-id-type="pmid">37483357</pub-id></mixed-citation>
</ref>
<ref id="B135">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wallace</surname> <given-names>G. K.</given-names></name></person-group> (<year>1991</year>). <article-title>The JPEG still picture compression standard</article-title>. <source>Commun. ACM</source> <volume>34</volume>, <fpage>30</fpage>&#x02013;<lpage>44</lpage>. doi: <pub-id pub-id-type="doi">10.1145/103085.103089</pub-id></mixed-citation>
</ref>
<ref id="B136">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Bovik</surname> <given-names>A. C.</given-names></name> <name><surname>Sheikh</surname> <given-names>H. R.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E. P.</given-names></name></person-group> (<year>2004</year>). <article-title>Perceptual image quality assessment: from error visibility to structural similarity</article-title>. <source>IEEE Trans Image Process</source>. <volume>13</volume>, <fpage>600</fpage>&#x02013;<lpage>612</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2003.819861</pub-id></mixed-citation>
</ref>
<ref id="B137">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ware</surname> <given-names>C.</given-names></name> <name><surname>Cowan</surname> <given-names>W.</given-names></name></person-group> (<year>1982</year>). <article-title>Changes in perceived color due to chromatic interactions</article-title>. <source>Vision Res</source>. <volume>22</volume>, <fpage>1353</fpage>&#x02013;<lpage>1362</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0042-6989(82)90225-5</pub-id><pub-id pub-id-type="pmid">7157673</pub-id></mixed-citation>
</ref>
<ref id="B138">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Watson</surname> <given-names>A.</given-names></name></person-group> (<year>1983</year>). <article-title>&#x0201C;Detection and recognition of simple spatial forms,&#x0201D;</article-title> in <source>Physical and Biological Processing of Images</source>, vol. 11, eds. O. Braddick and A. Sleigh (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>100</fpage>&#x02013;<lpage>114</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-642-68888-1_8</pub-id></mixed-citation>
</ref>
<ref id="B139">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Watson</surname> <given-names>A.</given-names></name> <name><surname>Ramirez</surname> <given-names>C.</given-names></name></person-group> (<year>2000</year>). <article-title>A standard observer for spatial vision</article-title>. <source>Investig. Opht. Vis. Sci</source>. <volume>41</volume>:<fpage>S713</fpage>.</mixed-citation>
</ref>
<ref id="B140">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Watson</surname> <given-names>A.</given-names></name> <name><surname>Solomon</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>A model of visual contrast gain control and pattern masking</article-title>. <source>JOSA A</source> <volume>14</volume>, <fpage>2379</fpage>&#x02013;<lpage>2391</lpage>. doi: <pub-id pub-id-type="doi">10.1364/JOSAA.14.002379</pub-id></mixed-citation>
</ref>
<ref id="B141">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Watson</surname> <given-names>A. B.</given-names></name> <name><surname>Malo</surname> <given-names>J.</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;Video quality measures based on the standard spatial observer,&#x0201D;</article-title> in <source>Proceedings. International Conference on Image Processing, vol. 3, III</source> (<publisher-loc>Rochester, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>). doi: <pub-id pub-id-type="doi">10.1109/ICIP.2002.1038898</pub-id></mixed-citation>
</ref>
<ref id="B142">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Whittle</surname> <given-names>P.</given-names></name></person-group> (<year>1992</year>). <article-title>Brightness, discriminability and the &#x0201C;crispening effect&#x0201D;</article-title>. <source>Vis. Res</source>. <volume>32</volume>, <fpage>1493</fpage>&#x02013;<lpage>1507</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0042-6989(92)90205-W</pub-id><pub-id pub-id-type="pmid">1455722</pub-id></mixed-citation>
</ref>
<ref id="B143">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Wyszecki</surname> <given-names>G.</given-names></name> <name><surname>Stiles</surname> <given-names>W.</given-names></name></person-group> (<year>2000</year>). <source>Color Science: Concepts and Methods, Quantitative Data and Formulae</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x00026; Sons</publisher-name>.</mixed-citation>
</ref>
<ref id="B144">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zhaoping</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <source>Understanding Vision: Theory, Models, and Data</source>. Oxford: Oxford University Press. doi: <pub-id pub-id-type="doi">10.1093/acprof:oso/9780199564668.001.0001</pub-id></mixed-citation>
</ref>
<ref id="B145">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zhuang</surname> <given-names>C.</given-names></name> <name><surname>Yan</surname> <given-names>S.</given-names></name> <name><surname>Nayebi</surname> <given-names>A.</given-names></name> <name><surname>Schrimpf</surname> <given-names>M.</given-names></name> <name><surname>Frank</surname> <given-names>M.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Unsupervised neural network models of the ventral visual stream</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A</source>. <volume>118</volume>:<fpage>e2014196118</fpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.2014196118</pub-id><pub-id pub-id-type="pmid">33431673</pub-id></mixed-citation>
</ref>
</ref-list>
<fn-group>
<fn fn-type="custom" custom-type="edited-by" id="fn0001">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2812493/overview">Sandhya Avasthi</ext-link>, ABES Engineering College, India</p>
</fn>
<fn fn-type="custom" custom-type="reviewed-by" id="fn0002">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/76072/overview">Taiki Fukiage</ext-link>, NTT Communication Science Laboratories, Japan</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1376257/overview">Sandeep Singh Sengar</ext-link>, Cardiff Metropolitan University, United Kingdom</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2964913/overview">S. Sivachitralakshmi</ext-link>, SRM Institute of Science and Technology, India</p>
</fn>
</fn-group>
<fn-group>
<fn id="fn0003"><label>1</label><p>Concept and results first presented at the <italic>AI Evaluation Workshop</italic> at the University of Bristol, June 2022.</p></fn>
<fn id="fn0004"><label>2</label><p>All made online available here since 2022: <ext-link ext-link-type="uri" xlink:href="http://isp.uv.es/docs/TuringTestVision.zip">http://isp.uv.es/docs/TuringTestVision.zip</ext-link>.</p></fn>
<fn id="fn0005"><label>3</label><p>In the <italic>talk</italic> (<xref ref-type="bibr" rid="B96">Malo et al., 2022</xref>), we mentioned the humorous comment of Dr. Paninski at NYU back in 2001 after he carefully listened to the details of our brand new model: <italic>yes, yes, that is nice, but the brain doesn&#x00027;t work like that, does it?</italic>.</p></fn>
<fn id="fn0006"><label>4</label><p>In the <italic>talk</italic> <xref ref-type="bibr" rid="B96">Malo et al. (2022</xref>), we guessed that, following that skepticism, Barlow questioned our preliminary work on the use of Principal Curves to explain color and texture non-linearities of human vision purely based on image data, back in 2004 (<xref ref-type="bibr" rid="B88">Malo et al., 2004</xref>): <italic>yes, that is interesting, but the visual brain may not work like that</italic>.</p></fn>
<fn id="fn0007"><label>5</label><p>See the script <monospace>StimuliColorNonLinearities.m</monospace>.</p></fn>
<fn id="fn0008"><label>6</label><p>See the script <monospace>StimuliColorNonLinearities.m</monospace> of this work which also uses the Toolbox Colorlab (<xref ref-type="bibr" rid="B92">Malo and Luque, 2002</xref>) for calibration.</p></fn>
<fn id="fn0009"><label>7</label><p>See the script <monospace>StimuliMaskEnergy.m</monospace> which makes extensive use of the Toolbox Vistalab (<xref ref-type="bibr" rid="B84">Malo and Gutierrez, 2002</xref>).</p></fn>
<fn id="fn0010"><label>8</label><p>See the script <monospace>StimuliMaskOrient.m</monospace>, which makes extensive use of the Toolbox Vistalab (<xref ref-type="bibr" rid="B84">Malo and Gutierrez, 2002</xref>).</p></fn>
<fn id="fn0011"><label>9</label><p>The Decalogue Toolbox is available here: <ext-link ext-link-type="uri" xlink:href="http://isp.uv.es/docs/TuringTestVision.zip">http://isp.uv.es/docs/TuringTestVision.zip</ext-link>.</p></fn>
<fn id="fn0012"><label>10</label><p>Spectrally narrow Gaussians (5 nm width) of constant energy centered on different wavelengths along the visible spectrum on top of a low energy flat spectrum, as in <xref ref-type="fig" rid="F4">Figure 4</xref>. In this way all the stimuli can be faithfully represented in digital values.</p></fn>
<fn id="fn0013"><label>11</label><p>The divisive normalization transform (<xref ref-type="bibr" rid="B18">Carandini and Heeger, 1994</xref>, <xref ref-type="bibr" rid="B19">2012</xref>) is a function <italic>y</italic> &#x0003D; <italic>f</italic>(<italic>x</italic>) with this generic sigmoidal form: <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula> in which each output <italic>y</italic><sub><italic>i</italic></sub> is driven by the corresponding input <italic>x</italic><sub><italic>i</italic></sub>, <italic>normalized</italic> by a pool of the activity of the neighbor responses. It is a sort of local batch normalization (<xref ref-type="bibr" rid="B55">Krizhevsky et al., 2012a</xref>).</p></fn>
</fn-group>
</back>
</article>