<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Hum. Neurosci.</journal-id>
<journal-title>Frontiers in Human Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Hum. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5161</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnhum.2022.862588</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Human Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>An improved saliency model of visual attention dependent on image content</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Novin</surname> <given-names>Shabnam</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/228500/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Fallah</surname> <given-names>Ali</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1651651/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Rashidi</surname> <given-names>Saeid</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1651059/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Daliri</surname> <given-names>Mohammad Reza</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/7246/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Faculty of Biomedical Engineering, Amirkabir University of Technology (AUT)</institution>, <addr-line>Tehran</addr-line>, <country>Iran</country></aff>
<aff id="aff2"><sup>2</sup><institution>Faculty of Medical Sciences and Technologies, Science and Research Branch, Islamic Azad University</institution>, <addr-line>Tehran</addr-line>, <country>Iran</country></aff>
<aff id="aff3"><sup>3</sup><institution>Neuroscience and Neuroengineering Research Laboratory, Biomedical Engineering Department, School of Electrical Engineering, Iran University of Science and Technology</institution>, <addr-line>Tehran</addr-line>, <country>Iran</country></aff>
<aff id="aff4"><sup>4</sup><institution>School of Cognitive Sciences (SCS), Institute for Research in Fundamental Sciences (IPM)</institution>, <addr-line>Tehran</addr-line>, <country>Iran</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Elias Ebrahimzadeh, Institute for Research in Fundamental Sciences (IPM), Iran</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Alexander Toet, Netherlands Organisation for Applied Scientific Research, Netherlands; Anna Belardinelli, Honda Research Institute Europe GmbH, Germany</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Shabnam Novin <email>novin.sh.novin&#x00040;gmail.com</email></corresp>
<corresp id="c002">Ali Fallah <email>afallah&#x00040;aut.ac.ir</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Cognitive Neuroscience, a section of the journal Frontiers in Human Neuroscience</p></fn></author-notes>
<pub-date pub-type="epub">
<day>28</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>862588</elocation-id>
<history>
<date date-type="received">
<day>26</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>11</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Novin, Fallah, Rashidi and Daliri.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Novin, Fallah, Rashidi and Daliri</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Many visual attention models have been presented to obtain the saliency of a scene, i.e., the visually significant parts of a scene. However, some mechanisms are still not taken into account in these models, and the models do not fit the human data accurately. These mechanisms include which visual features are informative enough to be incorporated into the model, how the conspicuity of different features and scales of an image may integrate to obtain the saliency map of the image, and how the structure of an image affects the strategy of our attention system. We integrate such mechanisms in the presented model more efficiently compared to previous models. First, besides low-level features commonly employed in state-of-the-art models, we also apply medium-level features as the combination of orientations and colors based on the visual system behavior. Second, we use a variable number of center-surround difference maps instead of the fixed number used in the other models, suggesting that human visual attention operates differently for diverse images with different structures. Third, we integrate the information of different scales and different features based on their weighted sum, defining the weights according to each component&#x00027;s contribution, and presenting both the local and global saliency of the image. To test the model&#x00027;s performance in fitting human data, we compared it to other models using the CAT2000 dataset and the Area Under Curve (AUC) metric. Our results show that the model has high performance compared to the other models (AUC = 0.79 and sAUC = 0.58) and suggest that the proposed mechanisms can be applied to the existing models to improve them.</p></abstract>
<kwd-group>
<kwd>visual attention</kwd>
<kwd>saliency model</kwd>
<kwd>medium-level features</kwd>
<kwd>center-surround difference map</kwd>
<kwd>eye fixations</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="22"/>
<ref-count count="66"/>
<page-count count="20"/>
<word-count count="11351"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>The brain uses attention mechanisms to process a massive amount of visual information, selecting the significant items while ignoring the unimportant ones (Carrasco, <xref ref-type="bibr" rid="B16">2011</xref>). Human visual attention acts based on two mechanisms of bottom-up and top-down interacting through various brain areas (Doricchi et al., <xref ref-type="bibr" rid="B22">2010</xref>; Dehghani et al., <xref ref-type="bibr" rid="B20">2020</xref>). Bottom-up attention functions in an involuntary and fast way based on the psychophysical properties of the scene, while top-down attention functions in a voluntary and slow way based on pre-defined goals (Connor et al., <xref ref-type="bibr" rid="B18">2004</xref>).</p>
<p>Many computational models have been presented for visual attention. Few have modeled top-down visual attention (Deco and Zihl, <xref ref-type="bibr" rid="B19">2001</xref>; Denison et al., <xref ref-type="bibr" rid="B21">2021</xref>; Novin et al., <xref ref-type="bibr" rid="B48">2021</xref>), while most of them have modeled bottom-up visual attention, as its mechanisms are better understood (Zhang et al., <xref ref-type="bibr" rid="B63">2008</xref>). Visual attention models that aim to obtain the saliency of a scene are called saliency models. Saliency refers to how some parts of a scene stand out compared to their surroundings in our visual perception (Borji and Itti, <xref ref-type="bibr" rid="B9">2013</xref>). Saliency models are typically evaluated based on how well they can predict human attentional gaze while viewing different images (Krasovskaya and MacInnes, <xref ref-type="bibr" rid="B36">2019</xref>; Ullah et al., <xref ref-type="bibr" rid="B55">2020</xref>).</p>
<p>Many different approaches have been proposed for saliency detection. Some of them operate based on objects&#x00027; information in the image (Liu et al., <xref ref-type="bibr" rid="B41">2010</xref>). The Bayesian-based models combine the prior constraints with scene information in a probabilistic way to find the rare features (Zhang et al., <xref ref-type="bibr" rid="B63">2008</xref>). Some models determine the most informative parts of the scene as salient (Bruce and Tsotsos, <xref ref-type="bibr" rid="B13">2005</xref>). The frequency-based models determine the image regions with unique frequencies as salient (Hou and Zhang, <xref ref-type="bibr" rid="B28">2007</xref>). The other category is learning-based models that train the model with human data (Judd et al., <xref ref-type="bibr" rid="B34">2009</xref>).</p>
<p>Recently, many saliency models have focused on using learning algorithms. Although learning algorithms have significantly improved the performance of saliency models (Kummerer et al., <xref ref-type="bibr" rid="B40">2017</xref>), they operate as a black box rather than including particular mechanisms to address the behavior of visual attention. In addition, they usually require many data to train the model. Here, we present a saliency model to investigate some of the mechanisms of bottom-up attention behavior. The missing understanding of the mechanisms of bottom-up attention and decreasing focus on investigating these mechanisms motivated us to do this study.</p>
<p>There are still open problems in the bottom-up aspect of the saliency models, addressed in recent studies (Zhang and Sclaroff, <xref ref-type="bibr" rid="B62">2016</xref>; Ayoub et al., <xref ref-type="bibr" rid="B4">2018</xref>; Wang et al., <xref ref-type="bibr" rid="B58">2018</xref>; Molin et al., <xref ref-type="bibr" rid="B45">2021</xref>). These problems include determining the appropriate features to apply (Kummerer et al., <xref ref-type="bibr" rid="B40">2017</xref>; Narayanaswamy et al., <xref ref-type="bibr" rid="B47">2020</xref>), establishing center-surround differences in the image (Zhang and Sclaroff, <xref ref-type="bibr" rid="B62">2016</xref>; Ayoub et al., <xref ref-type="bibr" rid="B4">2018</xref>), and integrating multiple features and scales (Jian et al., <xref ref-type="bibr" rid="B33">2019</xref>; Narayanaswamy et al., <xref ref-type="bibr" rid="B47">2020</xref>) to obtain salient areas in the image.</p>
<p>Here, we make some improvements based on the existing knowledge about the visual system. (1) We investigate some missing features to be considered. (2) How can we integrate the image information better for creating the saliency map. (3) How can we improve the mechanism of calculating center-surround differences in an image by adapting it to the image structure. In the following, we review previous studies in line with our focus on these improvements.</p>
<p>The rest of our paper is structured as follows. In Section Related study, we review the related studies and end with the contributions of our model. Afterward, in Section Methods, we explain our model. In Section Results, we show the results and compare the model&#x00027;s performance with other models. In Section Discussion, we discuss the results. Finally, in Section Conclusion, we make conclusions.</p>
</sec>
<sec id="s2">
<title>Related study</title>
<sec>
<title>Itti&#x00027;s base model of visual attention</title>
<p>The base model of saliency was presented by Itti et al. (<xref ref-type="bibr" rid="B31">1998</xref>), commonly known as the Itti model. It was known as a pioneer and benchmark model because it is based mainly on the behavior of the human early visual system. The Itti model predicts the saliency map by applying the center-surround (C&#x02013;S) mechanism to three features of intensity, color, and orientation that are known to be critical low-level features in bottom-up attention (Wolfe and Horowitz, <xref ref-type="bibr" rid="B60">2004</xref>; Frintrop et al., <xref ref-type="bibr" rid="B25">2015</xref>). The C&#x02013;S difference maps between different scales are calculated for each feature map to simulate visual system behavior, which is attracted by areas that are more distinct from their surroundings (Casagrande and Norton, <xref ref-type="bibr" rid="B17">1991</xref>). Then, the C&#x02013;S difference maps are combined across various scales and features to obtain the saliency map, which is a topographic representation of saliency for each pixel in an image.</p>
<p>Later, many models made improvements to different steps of the base model of Itti (Borji and Itti, <xref ref-type="bibr" rid="B9">2013</xref>). In the following, we review some of these studies.</p>
</sec>
<sec>
<title>Decomposing images into multi scales using wavelets</title>
<p>In most models, the image is decomposed to multiple scales using the classic Gaussian filter, while here, we apply wavelet-based decomposition. Recently, using the wavelet transform (WT) in saliency models has been shown to be beneficial (Murray et al., <xref ref-type="bibr" rid="B46">2011</xref>; Imamoglu et al., <xref ref-type="bibr" rid="B29">2013</xref>; Ma et al., <xref ref-type="bibr" rid="B42">2015</xref>). WT has the advantage of simultaneously providing spatial and frequency information at each image scale (Murray et al., <xref ref-type="bibr" rid="B46">2011</xref>). Also, it extracts oriented details of the image in horizontal, vertical, and diagonal dimensions at each scale (Imamoglu et al., <xref ref-type="bibr" rid="B29">2013</xref>). WT is a powerful tool for spatial-frequency (Antonini et al., <xref ref-type="bibr" rid="B3">1992</xref>) and time-frequency analysis (Sadjadi et al., <xref ref-type="bibr" rid="B51">2021</xref>). In a spatial-frequency analysis, WT decomposes the image into multiple levels by iteratively performing horizontal and, subsequently, vertical sub-sampling on it through a set of filters (Antonini et al., <xref ref-type="bibr" rid="B3">1992</xref>).</p>
<p>Murray et al. (<xref ref-type="bibr" rid="B46">2011</xref>) model the center-surround effect based on contrast energy ratios at the central and surrounding regions using WT. Then they weigh the scales of the wavelet pyramid by a contrast sensitivity function (CSF). Finally, they obtain the saliency map by combining the inverse WT of the weighted maps. In Murray et al. (<xref ref-type="bibr" rid="B46">2011</xref>), the computation of the saliency map is mainly based on local contrasts in the image. Imamoglu et al. (<xref ref-type="bibr" rid="B29">2013</xref>) obtained the feature maps by applying WT until the coarsest possible level. Their model obtains the general saliency map by modulating the locations&#x00027; local saliency with their global saliency. Abkenar and Ahmad (<xref ref-type="bibr" rid="B1">2016</xref>) proposed a saliency model according to the wavelet coefficients calculated for superpixels to make the model applicable to more complex images.</p>
</sec>
<sec>
<title>Using multi-scale features to create C&#x02013;S difference maps</title>
<p>The experiment of Bonnar et al. (<xref ref-type="bibr" rid="B5">2002</xref>) indicates that the perceived information of an image can be present at different scales. Hence, integration of the information of different scales is needed to obtain the final saliency. The question is how many different scales of C&#x02013;S difference maps we need to integrate. Previous models have used different numbers, and no evaluation has been done on what can be a proper number to choose. The base model of Itti et al. (<xref ref-type="bibr" rid="B31">1998</xref>) employs 6 C&#x02013;S difference maps created for low-level features. Zhao and Koch (<xref ref-type="bibr" rid="B65">2011</xref>) use a similar model to Itti&#x00027;s, adding the face feature. In Zhang et al. (<xref ref-type="bibr" rid="B63">2008</xref>), Goferman et al. (<xref ref-type="bibr" rid="B27">2012</xref>), and Ma et al. (<xref ref-type="bibr" rid="B43">2013</xref>), four scales of feature maps are used. The authors Kruthiventi et al. (<xref ref-type="bibr" rid="B38">2017</xref>) and Qi et al. (<xref ref-type="bibr" rid="B50">2019</xref>) proposed a multi-scale convolutional neural network (CNN), where each CNN is trained to obtain the salient locations at a particular scale. Vig et al. (<xref ref-type="bibr" rid="B56">2014</xref>) combine various models to take advantage of each one. The authors discuss that one of the factors that makes the models different is the scale on which they perform.</p>
<p>In contrast to the common strategy of using a fixed number of scales, we discuss that images with different structures may require a different number of scales to present the saliency of objects in the image.</p>
</sec>
<sec>
<title>Integration of information to create the saliency map</title>
<p>Conventionally, in many models, the saliency map is obtained by linearly combining different scales and features (like Itti et al., <xref ref-type="bibr" rid="B31">1998</xref>; Borji, <xref ref-type="bibr" rid="B6">2012</xref>; Goferman et al., <xref ref-type="bibr" rid="B27">2012</xref>; Imamoglu et al., <xref ref-type="bibr" rid="B29">2013</xref>; Wei and Luo, <xref ref-type="bibr" rid="B59">2015</xref>; Zeng et al., <xref ref-type="bibr" rid="B61">2015</xref>). The objects produce different amounts of saliency at different scales, depending on their size, details, etc. Similarly, different features are not equally relevant to describe an object&#x00027;s saliency. Therefore, a linear combination of different scales and features may not produce results fitting to human data.</p>
<p>Some models made improvements by applying different methods than linear combinations. Itti and Koch (<xref ref-type="bibr" rid="B30">2001</xref>) compared four different methods for normalizing the feature maps based on their distributions. Murray et al. (<xref ref-type="bibr" rid="B46">2011</xref>) weigh the scales using the contrast sensitivity function proposed by Otazu et al. (<xref ref-type="bibr" rid="B49">2010</xref>) and fitting it to psychophysical data. They obtain the final saliency map by combining the Euclidean norm of the maps of different channels. Narayanaswamy et al. (<xref ref-type="bibr" rid="B47">2020</xref>) prioritize the feature maps at multiple levels based on the 2D entropy of the maps. Then they calculate the model score for several channel combinations to find the informative channels. Zhao and Koch (<xref ref-type="bibr" rid="B65">2011</xref>) set the weights of feature maps based on learning different datasets using the least square method. Borji et al. (<xref ref-type="bibr" rid="B8">2011</xref>) use an evolutionary optimization method to set the weights of the scales and feature maps so that they lead to maximum scores and minimum processing costs. Singh et al. (<xref ref-type="bibr" rid="B52">2020</xref>), in their optimization method, define an objective to increase the activity in the saliency map at the location of a salient object and decrease the activity in the background.</p>
<p>Here, we make some improvements to previous models to consider some missing mechanisms. Our contributions are listed below:</p>
<list list-type="simple">
<list-item><p>- We propose that in addition to the commonly used low-level and high-level features, the medium-level features based on the combination of orientations and colors play a role in bottom-up attention.</p></list-item>
<list-item><p>- We apply a weighting method for across-scale and across-feature integration, presenting the image&#x00027;s local and global saliency. Furthermore, we compare the weighting method for integrating features&#x00027; conspicuity maps with the method of calculating their maximum.</p></list-item>
<list-item><p>- We propose using a variable number of center-surround difference maps depending on the structure of the images. This is an important part of our contributions.</p></list-item>
</list>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>Methods</title>
<p>The block diagram of the proposed model is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The input of the model is an RGB image. The model consists of five layers, where each layer&#x00027;s output is shown for a sample image in the figure. In the first layer, the visual features of the image are extracted. In the second layer, the scale pyramids are acquired for each feature using wavelet transform. In the third layer, the difference between the high and low-resolution pyramid levels is calculated for each feature in the specified levels to make the center-surround difference maps. In the fourth layer, C&#x02013;S difference maps are integrated for each feature to create the feature&#x00027;s conspicuity map. Finally, in the fifth layer, the conspicuity maps of different features are combined to make the final saliency map. In the following, the computations within the layers are described in detail.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The proposed model structure. The outputs of each layer are shown for a sample image. RG and BY denote red-green and blue-yellow channels, respectively. For more details, refer to the main text.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0001.tif"/>
</fig>
<sec>
<title>Layer I&#x02013;Extracting the features</title>
<p>First, the input RGB images are converted to CIELab color space. Compared to RGB used in Itti&#x00027;s base model and many other models, CIELab is a perceptually uniform color space (Frintrop, <xref ref-type="bibr" rid="B23">2006</xref>; Borji, <xref ref-type="bibr" rid="B6">2012</xref>). The CIELab color space consists of L, a, and b channels, representing luminance, the green-red opponent colors, and the blue-yellow opponent colors, similar to human color perception (Frintrop, <xref ref-type="bibr" rid="B23">2006</xref>; Borji, <xref ref-type="bibr" rid="B6">2012</xref>; Ma et al., <xref ref-type="bibr" rid="B43">2013</xref>). The intensity, color, and orientation features are extracted from the CIELab image.</p>
<p>The intensity feature is computed by Equation (1), where R, G, and B, respectively, stand for the red, green, and blue channels of the RGB image.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>2989</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>R</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5870</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>G</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>1140</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>B</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Since a and b channels in CIElab space are based on the opponent color model of human visual cells (Frintrop, <xref ref-type="bibr" rid="B23">2006</xref>), we use a and b channels, respectively, as green-red and blue-yellow color opponency features to account for this behavior of visual cells.</p>
<p>A set of 8 orientation features is obtained by convolving the intensity map in (1) with a set of eight Gabor filters at the wavelength of 10 and orientations of <inline-formula><mml:math id="M2"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>4</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>5</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>6</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>7</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Previous models have used various low-level and high-level features, and it is still under debate how much and in which image areas each of these two kinds of features contribute to predicting the human gaze data (Kummerer et al., <xref ref-type="bibr" rid="B40">2017</xref>). Here, in addition to the low-level features mentioned above, we consider the role of medium-level features. The neurophysiological findings of Ts&#x00027;o and Gilbert (<xref ref-type="bibr" rid="B54">1988</xref>) and the study of Koene and Zhaoping (<xref ref-type="bibr" rid="B35">2007</xref>) provide evidence for the existence of primary visual cells driven by saliency according to the conjunction of color and orientation. Based on these studies, we suggest that medium-level features based on the combination of orientations and colors also play a role in bottom-up visual attention. A set of eight medium-level features is obtained by convolving the red-green channel with a set of eight Gabor filters at the wavelength of 10 and orientations of <inline-formula><mml:math id="M3"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>4</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>5</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>6</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>7</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. Similarly, a set of eight medium-level features is obtained by convolving the blue-yellow channel with the same Gabor filters.</p>
</sec>
<sec>
<title>Layer II&#x02013;Creating feature pyramids</title>
<p>In most models, the pyramids are built using Gaussian decomposition. Here, according to the advantages described in Section Decomposing images into multi scales using wavelets, wavelet decomposition is applied. The wavelet function is selected based on its properties: orthogonality, symmetry, compact support, and regularity. The Daubechies and Symlet wavelets are more frequently used in saliency studies (Jian et al., <xref ref-type="bibr" rid="B32">2015</xref>; Zhu et al., <xref ref-type="bibr" rid="B66">2019</xref>). Here, we choose the Symlet wavelet having more symmetry, which avoids phase distortion. We use Symlet of order 4 (sym4), which has also been used in some other studies (Zhang et al., <xref ref-type="bibr" rid="B64">2011</xref>; Ghasemi et al., <xref ref-type="bibr" rid="B26">2013</xref>). Using a sym4 wavelet, a pyramid of eight scales is obtained for each feature.</p>
</sec>
<sec>
<title>Layer III&#x02014;Creating center-surround difference maps</title>
<p>In order to obtain areas with high contrast compared to their surroundings, the difference between different scales of the image is calculated. For each feature pyramid, we calculate the difference between the scales of low numbers and the scales of two and three higher numbers. The lower scales with a high resolution and the higher scales with a low resolution can be considered respectively as the center and surround areas to compute the center-surround difference maps for each feature, as described in Equation (2).</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0229D;</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0229D;</mml:mo><mml:mi>R</mml:mi><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0229D;</mml:mo><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0229D;</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0229D;</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>Y</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>Y</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0229D;</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>Y</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>c</italic> refers to the center part represented by lower scales and <italic>s</italic> refers to the surrounding region represented by higher scales. &#x003B8; denotes the orientation of the Gabor function. The symbol &#x0229D; denotes the C&#x02013;S difference operator, where the high-scale image is interpolated to the size of the low-scale image, and then the two images are subtracted.</p>
<p>Most saliency models use a fixed number of C&#x02013;S difference maps. Here, we discuss that the human visual system does not attend to the contrasts in the same way for all images, containing different amounts of detailed and coarse content. Hence, different numbers of contrast maps are required for different images to model visual attention. Several factors, such as the image&#x00027;s crowdedness, the size of the objects, and the variety of each feature, may determine the amount of detailed and coarse content in an image that can affect the required number of contrast maps. We investigated our hypothesis by applying 4, 6, and 10 contrast maps for each feature. The numbers 4, 6, and 10 are chosen exemplary to refer to, respectively, a low, medium, and a high number of contrast maps. We will compare all three groups of results with the human data in the results section. For the final results, we will compute the model&#x00027;s performance based on the maximum value between the results for using 4, 6, and 10 contrast maps. The values of <italic>c</italic> and <italic>s</italic> in Equation (2), denoting the scale number, are defined as described in equations (3&#x02013;6).</p>
<disp-formula id="E5"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">Using 4 contrast maps</mml:mtext><mml:mo>:</mml:mo><mml:mi>c</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E6"><label>(4)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">Using 6 contrast maps</mml:mtext><mml:mo>:</mml:mo><mml:mi>c</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E7"><label>(5)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">Using 10 contrast maps</mml:mtext><mml:mo>:</mml:mo><mml:mi>c</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>5</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E8"><label>(6)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">The s values are defined as&#x000A0;</mml:mtext><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B4;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B4;</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the case of using <italic>c&#x003F5;</italic> {1, 2}, we would obtain 4 contrast maps for each feature. Thus, it would yield 27 &#x000D7; 4 = 108 contrast maps in total. In the same way, in the case of using <italic>c&#x003F5;</italic> {1, 2, 3}, we would obtain 27 &#x000D7; 6 = 162 contrast maps in total, and in the case of using <italic>c&#x003F5;</italic> {1, 2, 3, 4, 5}, we would obtain 27 &#x000D7; 10 = 270 contrast maps in total. The results for a sample input image are shown in the third layer of the model in <xref ref-type="fig" rid="F1">Figure 1</xref> by applying 6 contrast maps as an example.</p>
<p>In <xref ref-type="table" rid="T1">Table 1</xref>, the calculations in equations 2&#x02013;6 are detailed to show how the difference between scale numbers is calculated in the three approaches of using 4, 6, and 10 contrast maps.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>A detailed description of the calculations in equations 2&#x02013;6 for obtaining C&#x02013;S difference maps for each feature.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Approach</bold></th>
<th valign="top" align="left"><bold>Difference between scale numbers c and s, calculated for obtaining C&#x02013;S maps</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Using 4 contrast maps</td>
<td valign="top" align="left">1&#x0229D;3, 1&#x0229D;4, 2&#x0229D;4, 2&#x0229D;5</td>
</tr>
<tr>
<td valign="top" align="left">Using 6 contrast maps</td>
<td valign="top" align="left">1&#x0229D;3, 1&#x0229D;4, 2&#x0229D;4, 2&#x0229D;5, 3&#x0229D;5, 3&#x0229D;6</td>
</tr>
<tr>
<td valign="top" align="left">Using 10 contrast maps</td>
<td valign="top" align="left">1&#x0229D;3, 1&#x0229D;4, 2&#x0229D;4, 2&#x0229D;5, 3&#x0229D;5, 3&#x0229D;6, 4&#x0229D;6, 4&#x0229D;7, 5&#x0229D;7, 5&#x0229D;8</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The symbol &#x0229D; denotes the center-surround difference operator. In each row, the numbers written before and after &#x0229D; refer, respectively, to parameters c and s.</p>
</table-wrap-foot>
</table-wrap>
<p>As it is seen in equations (3&#x02013;6) and <xref ref-type="table" rid="T1">Table 1</xref>, using each number of C&#x02013;S maps yields the contrast between specific scales. The lower scales (close to scale 1) contain the fine (detailed) content of the image, and the higher scales (close to scale 8) contain the coarse (rough) content of the image. As a result, we can say that using 4 C&#x02013;S maps yields the contrasts calculated as very detailed scales (1, 2) minus detailed scales (3) and also minus middle-coarse (4, 5) scales. Using 6 C&#x02013;S maps yields the same contrasts as 4 C&#x02013;S maps. In addition, it yields the contrasts computed as detailed scales (3) minus middle-coarse (5, 6) scales. Using 10 C&#x02013;S maps yields the same contrasts as 6 C&#x02013;S maps. In addition, it yields the contrasts computed as middle-coarse scales (4, 5) minus coarse scales (7, 8).</p>
</sec>
<sec>
<title>Layer IV&#x02013;Obtaining conspicuity maps</title>
<p>In this step, the contrast maps of each feature are normalized between [0, 1], and then they are combined using a weighted summation to construct the conspicuity map for the related feature. Following the behavior of the visual system, we use the contrast sensitivity function to calculate the weight of different spatial information. The Contrast sensitivity function shows how much the human visual system is sensitive to contrast changes in a scene in different spatial frequencies. We define the weight of each contrast map dependent on the importance of its spatial frequency information. We use the widely accepted CSF model proposed by Mannos and Sakrison (<xref ref-type="bibr" rid="B44">1974</xref>), described by Equation (7), which is also used in some other saliency models (Buzatu, <xref ref-type="bibr" rid="B14">2012</xref>; Wang et al., <xref ref-type="bibr" rid="B58">2018</xref>); however, they do not apply it for weighting contrast maps. The normalized contrast maps are transformed to the frequency domain by Fourier transform (Brigham and Morrow, <xref ref-type="bibr" rid="B12">1967</xref>). Then, the contrast sensitivity function <italic>C</italic>(<italic>f</italic>) is calculated by Equation (7).</p>
<disp-formula id="E9"><label>(7)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>0499</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>2964</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>f</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>114</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where f is the spatial frequency.</p>
<p>We obtain two conspicuity maps for each feature based on its contrast maps&#x00027; global and local weighting. For the global weighting, the weight of the normalized contrast map number <italic>i</italic> [termed as &#x003C9;<sub><italic>i</italic></sub> in Equation (8)] is calculated by averaging <italic>C</italic><sub><italic>i</italic></sub>(<italic>f</italic>) over the map. Then, the weighted and normalized sum of the contrast maps for a specific feature <italic>j</italic> is calculated to get the conspicuity map for feature <italic>j</italic> [termed as <italic>Global conspicuity</italic><sub><italic>j</italic></sub> in Equation (9)].</p>
<disp-formula id="E10"><label>(8)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E11"><label>(9)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>G</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>i</italic> refers to the index of the contrast map for a specific feature, <italic>j</italic> refers to the index of the feature, and <italic>N</italic> = 4, 6, and 10, respectively, for the case of using 4, 6, and 10 contrast maps.</p>
<disp-formula id="E12"><mml:math id="M14"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>j</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>Y</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x003B8;</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>4</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>5</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>6</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>7</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>For the local weighting, the <italic>C</italic>(<italic>f</italic>) matrix is multiplied pixel-wise in the associated normalized contrast map. Then the weighted and normalized sum of the contrast maps for each feature is calculated to get its conspicuity map [termed as <italic>Local conspicuity</italic><sub><italic>j</italic></sub> in Equation (10)].</p>
<disp-formula id="E13"><label>(10)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The global and local conspicuity maps of each feature are combined according to Equation (11) to yield its final conspicuity map. The weights &#x003B1; = 0.95 and &#x003B2; = 0.05 in Equation (11) are set and optimized by trial and error to better fit the results to the human data in MIT dataset provided in Judd et al. (<xref ref-type="bibr" rid="B34">2009</xref>). The conspicuity maps for a sample input image are shown in the fourth layer of the model in <xref ref-type="fig" rid="F1">Figure 1</xref> for the case of applying six contrast maps as an example.</p>
<disp-formula id="E14"><label>(11)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>Layer V&#x02013;Obtaining saliency map</title>
<p>The saliency map represents the saliency degree at every location of the image as the level of brightness in a grayscale image. The final saliency map is obtained by combining the conspicuity maps of different features. Here, we compare four different methods for integrating various features to find the best method among them. These four integration methods that are described below were developed based on a preliminary study testing 12 different variations of the presented integration methods.</p>
<sec>
<title>Four investigated methods for integrating features</title>
<sec>
<title>First method</title>
<p>In the first method, we obtain the saliency values by calculating the maximum value among the conspicuity maps of different features at each pixel using Equation (12). The reasoning is that we use the values that have more potential to make high conspicuity at each location. This strategy would be a more pixel-wise and local strategy to predict human attentional focus.</p>
<disp-formula id="E16"><label>(12)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E17"><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>j</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>Y</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x003B8;</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>4</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>5</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>6</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>7</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In Equation (12), the conspicuity maps of orientation features of <italic>O</italic><sup><italic>I</italic></sup>, <italic>O</italic><sup><italic>RG</italic></sup>, and <italic>O</italic><sup><italic>BY</italic></sup> are obtained by adding the conspicuity maps of the associated feature for 8 orientations and then scale normalizing between [0, 1].</p>
</sec>
<sec>
<title>Second method</title>
<p>In the second method, we combine the conspicuity maps using a weighted summation. The weight of each conspicuity map is defined based on the difference between the global maximum and average of local maxima in the related map as below. This method is used in a similar way in some other studies (Itti and Koch, <xref ref-type="bibr" rid="B30">2001</xref>; Frintrop et al., <xref ref-type="bibr" rid="B24">2007</xref>).</p>
<p>First, all conspicuity maps are normalized between [0, 1]. Second, we find the local maximum areas in each conspicuity map <italic>j</italic>. Then we calculate the global maximum value (<italic>Max</italic><sub><italic>j</italic></sub>) and the average value (<italic>mean</italic><sub><italic>j</italic></sub>) among the local maxima except <italic>Max</italic><sub><italic>j</italic></sub>. The weights of the conspicuity maps are calculated by Equation (13).</p>
<disp-formula id="E18"><label>(13)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>w</italic><sub><italic>j</italic></sub> is the weight of the conspicuity map of feature <italic>j</italic>, and</p>
<disp-formula id="E19"><mml:math id="M21"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>j</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mi>Y</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>G</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>Y</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x003B8;</mml:mi><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>4</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>5</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>6</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>7</mml:mn><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The amount of difference between the global maximum and averaged local maxima in Equation (13) represents the amount of contrast that the related conspicuity map can make to draw attention.</p>
<p>Third, the final saliency map, <italic>Sal map</italic>, is obtained by Equation (14) as the weighted sum of the conspicuity maps of various features.</p>
<disp-formula id="E20"><label>(14)</label><mml:math id="M22"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>Third method</title>
<p>In the third method, we obtain the weights directly by calculating the difference between the global maximum and the local maxima, in contrast to the averaged local maxima used in the second method. Thus, the weights in the third method would be as matrices and multiplied point-wise in the conspicuity maps. The weights in the second method rely more on the global saliency that each conspicuity map can produce, and the weights in the third method rely more on the local saliency of each conspicuity map.</p>
<p>In the third method, we calculate the values of the conspicuity map <italic>j</italic> at the location of local maxima to produce the matrix <italic>LocalMax</italic><sub><italic>j</italic></sub>. The weights of the conspicuity maps are calculated by Equation (15).</p>
<disp-formula id="E21"><label>(15)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>W</italic><sub><italic>j</italic></sub> is the weight of the conspicuity map of feature <italic>j</italic>.</p>
<p>The final saliency map, <italic>Sal map</italic>, is obtained by Equation (16) as the weighted sum of the conspicuity maps of different features.</p>
<disp-formula id="E22"><label>(16)</label><mml:math id="M24"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>Fourth method</title>
<p>According to comparing the results of the second and third methods, both methods had almost the same overall performance for the MIT dataset. However, for different images, one of the two methods performed slightly better than the other. This indicates that depending on the structure of the image, the global or local weighting of conspicuity maps may be more efficient. Based on this, in the fourth method, we add linearly the saliency maps obtained by means of the second and third methods to incorporate the properties of both methods. The results showed a little better performance for the fourth method compared to the second and third methods.</p>
<p>We use the fourth method as the final approach for combining different features based on weighted summation. The saliency map for a sample input image is shown in the last layer of the model in <xref ref-type="fig" rid="F1">Figure 1</xref> based on the fourth method for obtaining the saliency map.</p>
<p>Furthermore, in the results section, we will compare the results of the first method, which is based on maximizing the conspicuity maps, with the fourth method, which is based on the weighted summation of the maps. The comparison results show that the weighting method leads to better results.</p>
</sec>
</sec>
</sec>
<sec>
<title>Different versions of the model</title>
<p>We define three versions of our model to investigate the effect of each improvement that we made.</p>
<p><bold>Proposed model 1</bold> refers to the version in that we used a fixed number (6) of C&#x02013;S difference maps for each feature, and we did not apply the medium-level features. The model was chosen to contain six C&#x02013;S maps to be in line with the base model of Itti et al. (<xref ref-type="bibr" rid="B31">1998</xref>). In this model version, we investigate the effect of applying our approaches for integrating the information of different scales (described in Section Layer IV&#x02013;obtaining conspicuity maps) and features (described in Section Layer V&#x02013;obtaining saliency map). This model version is considered as a baseline to be compared with proposed model 2.</p>
<p><bold>Proposed model 2</bold> refers to the version in which we extended proposed model 1 by applying the medium-level features (described in Section Layer I&#x02013;Extracting the features) to investigate the effect of applying them. This model version is considered as a baseline to be compared with proposed model 3.</p>
<p><bold>Proposed model 3</bold> refers to the full version of the model in that we extended proposed model 2 by applying different numbers of 4, 6, and 10 for the number of C&#x02013;S difference maps (described in Section Layer III&#x02013;creating center-surround difference maps), and we calculated the maximum score among the scores of using 4, 6, and 10 maps. In this model version, we investigate the effect of applying variable numbers for C&#x02013;S difference maps.</p>
</sec>
<sec>
<title>Dataset overview and evaluation metrics</title>
<p>We evaluated our model using the CAT2000 dataset (Borji and Itti, <xref ref-type="bibr" rid="B10">2015</xref>; Bylinskii et al., <xref ref-type="bibr" rid="B15">2019</xref>), which includes 24 subjects&#x00027; eye tracking data on 2,000 images classified into 20 different categories of 100 images. The images and fixation maps in the dataset have a size of 1,080 &#x000D7; 1,920. We resize the images to 450 &#x000D7; 800. The CAT2000 dataset contains images of different types, including art, cartoons, black white, indoor, outdoor, low resolution, noisy, objects, and outdoor natural. This large variety provides a proper way to validate the model&#x00027;s ability to predict human data related to images with different structures.</p>
<p>The resulting saliency maps were resized to the size of data fixation maps (1,080 &#x000D7; 1,920). The model&#x00027;s performance in predicting human fixations was evaluated based on AUC (Area Under Curve) metric (Bylinskii et al., <xref ref-type="bibr" rid="B15">2019</xref>), which shows the area under ROC (Receiver Operating Characteristic) curve. The ROC curve plots the true positive rate vs. the false positive rate based on comparing the model&#x00027;s fixations to the dataset fixations. A higher AUC denotes a higher performance. According to previous studies that reviewed the models using different evaluation metrics, AUC is the most common metric (Kummerer et al., <xref ref-type="bibr" rid="B39">2018</xref>; Bylinskii et al., <xref ref-type="bibr" rid="B15">2019</xref>). Our goal is to investigate the effect of improvements we made on the model. For this purpose, we found the location-based metric of AUC sufficient to evaluate the results considering that we will compare each improved model version with its lower model version based on this metric. In addition, to compare our model to the other models, we also used the shuffled AUC (sAUC) metric (Bylinskii et al., <xref ref-type="bibr" rid="B15">2019</xref>). The sAUC is similar to AUC with the difference that for AUC, the negative set is selected uniformly at random from the fixation map of the image, while, for sAUC, the negative set is selected from fixation maps of the other images sampled from the dataset. The sAUC is designed to penalize the models that explicitly apply center bias (Zhang et al., <xref ref-type="bibr" rid="B63">2008</xref>). Since some of the to-be-compared models apply center bias to their results, we used sAUC to compare the models better. We used the AUC-Borji and sAUC algorithms presented in Borji et al. (<xref ref-type="bibr" rid="B11">2013</xref>).</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>Results</title>
<sec>
<title>Visualizing the results of the model on sample images</title>
<p>In <xref ref-type="fig" rid="F2">Figure 2</xref>, the saliency map results and the AUC scores of the model for sample images from various categories are shown to show the model&#x00027;s performance for different image structures and conditions. The results are obtained using the fourth method of combining conspicuity maps described in Section Layer V&#x02013;obtaining saliency map.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Saliency map results and the AUC scores of the model for sample images of the CAT2000 dataset from various categories of Action, Art, Black White, Low Resolution, and Noisy.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0002.tif"/>
</fig>
</sec>
<sec>
<title>Comparing the methods for integrating conspicuity maps</title>
<p>Before comparing the model results to the other models, we first compare the first and fourth methods of obtaining the saliency map, described in Section Layer V&#x02013;obtaining saliency map as two different investigated approaches for integrating conspicuity maps. The first method is based on calculating the maximum of conspicuity maps, and the fourth method is based on the weighted summation of the maps. The mean AUC score calculated among all the images of the dataset was higher for the fourth method (AUC = 0.75) compared to the first method (AUC = 0.73). However, for some images, like the example ones shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, the first method performed better, as is seen in the saliency maps and the AUC scores. This shows that, although weighting the maps generally performs better; however, in some images, the conspicuity maps of particular features may dominate the other features, and thus, calculating the maximum of conspicuity maps can make better results for these images.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Showing some out of the common images where the first method of obtaining the saliency map is better than the fourth method described in Section Layer V-obtaining saliency map because the conspicuity maps of particular features dominate the other features. The results are shown for the case of applying six center-surround difference maps for each feature.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0003.tif"/>
</fig>
<p>For example, in <xref ref-type="fig" rid="F3">Figure 3</xref>, in the image of the first row, the edges stand out more than the other features, and therefore, the first method of obtaining the saliency map has resulted in a higher AUC score than the fourth method. In the image of the second row, in specific areas of the image, the intensity or the red-green feature is very bold compared to the other features, and since the first method calculates the maximum of maps over each area, it has gained a higher AUC score. Similarly, in the image of the third row, in the specific areas of the image, the red-green and edge features dominate the other features. In the image of the fourth row, the image contains specific areas of uniform structure, wherein each area, one of the features is bolder compared to other features, making the first method perform better than the fourth method. In the image of the fifth row, the image belongs to the Pattern category, and the pattern used has specific features like edges to be bolder than the other features.</p>
<p>For the rest of the results in this section, the results of the fourth method will be used.</p>
</sec>
<sec>
<title>Comparing different versions of the model to the other models</title>
<p>Many saliency models have been proposed to predict human fixations (Borji and Itti, <xref ref-type="bibr" rid="B9">2013</xref>; Borji et al., <xref ref-type="bibr" rid="B11">2013</xref>; Borji, <xref ref-type="bibr" rid="B7">2019</xref>). We compare our model to several models that used the CAT2000 dataset like ours, and their results are publicly available (<ext-link ext-link-type="uri" xlink:href="http://saliency.mit.edu/results_cat2000.html">http://saliency.mit.edu/results_cat2000.html</ext-link>). We use the metric of AUC-Borji and sAUC (Borji et al., <xref ref-type="bibr" rid="B11">2013</xref>) to compare our model to the models shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, with AUC scores ranging from low to highest reported scores. We use the IttiKoch model (Walther and Koch, <xref ref-type="bibr" rid="B57">2006</xref>) as an extension of the base model of Itti et al. (<xref ref-type="bibr" rid="B31">1998</xref>), the Achanta model (Achanta et al., <xref ref-type="bibr" rid="B2">2009</xref>) as one of the most cited models in the frequency domain, and the SUN saliency model (Zhang et al., <xref ref-type="bibr" rid="B63">2008</xref>) and the Fast and Efficient Saliency (FES) model (Tavakoli et al., <xref ref-type="bibr" rid="B53">2011</xref>) as two of the Bayesian-based models. We use the Murray model (Murray et al., <xref ref-type="bibr" rid="B46">2011</xref>) that uses wavelet transform to generate scales like our model. We also use learning-based models, including MSI-Net (Kroner et al., <xref ref-type="bibr" rid="B37">2020</xref>) and the Judd model (Judd et al., <xref ref-type="bibr" rid="B34">2009</xref>). The models of Judd et al. (<xref ref-type="bibr" rid="B34">2009</xref>) and Vig et al. (<xref ref-type="bibr" rid="B56">2014</xref>), and Zhang and Sclaroff (<xref ref-type="bibr" rid="B62">2016</xref>) have the highest AUC score reported on the mentioned website for evaluation on the CAT2000 dataset based on the AUC-Borji metric. Their high AUC is due to setting their parameters according to fixations on trained images. <xref ref-type="fig" rid="F4">Figure 4</xref> shows our model&#x00027;s mean scores among 20 categories in the dataset, compared to the mentioned models.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Comparing the performance of our model to several models evaluated on the CAT2000 dataset based on <bold>(A)</bold> mean AUC score and <bold>(B)</bold> mean sAUC score. For more details, refer to the main text.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0004.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F4">Figure 4</xref>, the scores of the three versions of the model described in Section Different versions of the model are compared to other models. In <xref ref-type="table" rid="T2">Table 2</xref>, the specifications and the scores of the compared models and our model are described. <xref ref-type="fig" rid="F4">Figure 4</xref> and <xref ref-type="table" rid="T2">Table 2</xref> include the models that require no learning as well as the learning-based ones. Our model belongs to non-learning-based models.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Specifications of the models that were compared to our model in <xref ref-type="fig" rid="F4">Figure 4</xref>, including the non-learning-based models and learning-based ones.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Model name and references</bold></th>
<th valign="top" align="left"><bold>Features</bold></th>
<th valign="top" align="left"><bold>Number of scales</bold></th>
<th valign="top" align="left"><bold>Method of integrating scales</bold></th>
<th valign="top" align="left"><bold>Method of integrating features</bold></th>
<th valign="top" align="left"><bold>Learning data</bold></th>
<th valign="top" align="left"><bold>AUC</bold></th>
<th valign="top" align="left"><bold>sAUC</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">IttiKoch (Walther and Koch, <xref ref-type="bibr" rid="B57">2006</xref>)</td>
<td valign="top" align="left">Low-level: intensity, color, orientations</td>
<td valign="top" align="left">6</td>
<td valign="top" align="left">Linear summation</td>
<td valign="top" align="left">Linear summation, and then estimating the proto-object region based on the salient locations</td>
<td valign="top" align="left">No</td>
<td valign="top" align="left">0.53</td>
<td valign="top" align="left">0.52</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">SUN saliency (Zhang et al., <xref ref-type="bibr" rid="B63">2008</xref>)</td>
<td valign="top" align="left">Low-level: intensity, color</td>
<td valign="top" align="left">4</td>
<td valign="top" align="left">Data-Driven Bayesian approach</td>
<td valign="top" align="left">Linear summation</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">0.69</td>
<td valign="top" align="left">0.57</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Achanta model (Achanta et al., <xref ref-type="bibr" rid="B2">2009</xref>)</td>
<td valign="top" align="left">Low-level: color, luminance</td>
<td valign="top" align="left">The model is frequency-tuned</td>
<td valign="top" align="left">&#x02013;</td>
<td valign="top" align="left">The difference between arithmetic mean pixel value and Gaussian blurred image is calculated for various features</td>
<td valign="top" align="left">No</td>
<td valign="top" align="left">0.55</td>
<td valign="top" align="left">0.52</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Judd model (Judd et al., <xref ref-type="bibr" rid="B34">2009</xref>)</td>
<td valign="top" align="left">Low-level: intensity, color, orientations Mid-level: horizon line High-level: persons, faces</td>
<td valign="top" align="left">3</td>
<td valign="top" align="left">Data-Based learning method</td>
<td valign="top" align="left">Data-Based learning method</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">0.84</td>
<td valign="top" align="left">0.56</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Murray model (Murray et al., <xref ref-type="bibr" rid="B46">2011</xref>)</td>
<td valign="top" align="left">Low-level: intensity, color</td>
<td valign="top" align="left">Largest dimension of image</td>
<td valign="top" align="left">&#x02013;</td>
<td valign="top" align="left">Euclidean norm of saliency maps of different channels, where each map is calculated based on weighted coefficients of WT of the image. The weights are defined based on CSF</td>
<td valign="top" align="left">Yes (for setting parameters of the weights defined by contrast</td>
<td valign="top" align="left">0.70</td>
<td valign="top" align="left">0.59</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">FES (Tavakoli et al., <xref ref-type="bibr" rid="B53">2011</xref>)</td>
<td valign="top" align="left">Low-level: CIELab values</td>
<td valign="top" align="left">3</td>
<td valign="top" align="left">Linear summation</td>
<td valign="top" align="left">Bayesian approach</td>
<td valign="top" align="left">Yes (for approximating probability values in Bayesian approach)</td>
<td valign="top" align="left">0.76</td>
<td valign="top" align="left">0.54</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">MSI-Net (Kroner et al., <xref ref-type="bibr" rid="B37">2020</xref>)</td>
<td valign="top" align="left">High-level: image&#x00027;s semantic information</td>
<td valign="top" align="left">The scales are obtained through three convolutional layers</td>
<td valign="top" align="left">Encoder-Decoder approach</td>
<td valign="top" align="left">The feature maps are combined through a convolutional neural network (CNN)</td>
<td valign="top" align="left">Yes</td>
<td valign="top" align="left">0.82</td>
<td valign="top" align="left">0.59</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Proposed model 1</td>
<td valign="top" align="left">Low-level: intensity, color, orientations</td>
<td valign="top" align="left">6</td>
<td valign="top" align="left">Weighted summation. The weights are defined based on CSF, calculated locally and globally</td>
<td valign="top" align="left">Weighted summation. The weights are defined based on the difference between the global maximum and local maxima, calculated locally and globally</td>
<td valign="top" align="left">No</td>
<td valign="top" align="left">0.73</td>
<td valign="top" align="left">0.54</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Proposed model 2</td>
<td valign="top" align="left">Low-level: intensity, color, orientations Medium-level: combination of colors and orientations</td>
<td valign="top" align="left">6</td>
<td valign="top" align="left">Same as proposed model 1</td>
<td valign="top" align="left">Same as proposed model 1</td>
<td valign="top" align="left">No</td>
<td valign="top" align="left">0.75</td>
<td valign="top" align="left">0.56</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Proposed model 3 (full version of our model)</td>
<td valign="top" align="left">Same as proposed model 2</td>
<td valign="top" align="left">Variable (4, 6, and 10)</td>
<td valign="top" align="left">Same as proposed model 1 and 2</td>
<td valign="top" align="left">Same as proposed model 1 and 2</td>
<td valign="top" align="left">No</td>
<td valign="top" align="left">0.79</td>
<td valign="top" align="left">0.58</td>
</tr></tbody>
</table>
</table-wrap>
<p><bold>Proposed model 1</bold> refers to the version in that we used a fixed number (6) of C&#x02013;S difference maps, and we did not apply the medium-level features. The AUC results in <xref ref-type="fig" rid="F4">Figure 4</xref> show that proposed model 1 has high performance (AUC = 0.73) in its related category of non-learning-based models. This high performance indicates that the algorithms used for integrating the information of different scales (described in Section Layer IV&#x02013;obtaining conspicuity maps) and different features (described in Section Layer V&#x02013;obtaining saliency map) were efficient in making acceptable results to fit human data.</p>
<p>Proposed model 1 is considered as a baseline to compare to <bold>proposed model 2</bold>, in which we add medium-level features (described in Section Layer I&#x02013;extracting the features). The AUC score shown in <xref ref-type="fig" rid="F4">Figure 4</xref> improved by 0.02 (AUC = 0.75) compared to the proposed model 1. The AUC increase suggests that the medium-level features play a role in bottom-up visual attention.</p>
<p>Proposed model 2 is considered as a baseline to compare to <bold>proposed model 3</bold>, in which we add the approach of using variable numbers of 4, 6, and 10 C&#x02013;S difference maps. We calculate the maximum AUC and sAUC score among the AUC and sAUC scores using 4, 6, and 10 C&#x02013;S difference maps (described in Section Layer III&#x02013;creating center-surround difference maps). As is seen in <xref ref-type="fig" rid="F4">Figure 4</xref>, applying variable numbers for C&#x02013;S difference maps has a considerable effect on improving the AUC results of the model for 0.04 (AUC = 0.79). This suggests that the human visual system may apply different strategies for contrast maps for different images depending on their contents. AUC increase points to better fitting to human gaze data, and gaze is an indicator of human visual attention. In other words, the results indicate a better fit for human attentional behavior.</p>
</sec>
<sec>
<title>The results of using a variable number of contrast maps</title>
<p>To make our proposal about using a variable number of C&#x02013;S difference maps more visible, in <xref ref-type="fig" rid="F5">Figure 5</xref>, we compare the model results between three cases of using 4, 6, and 10 C&#x02013;S difference maps for some sample images from the dataset. Referring to the description in Section Layer III&#x02013;creating center-surround difference maps, if an image has mainly detailed content, using 4 C&#x02013;S maps would probably lead to better performance than 6 or 10 C&#x02013;S maps. If an image has detailed and middle-coarse content, using 6 C&#x02013;S maps would be better, and if an image has detailed, middle-coarse and coarse content, using 10 C&#x02013;S maps would be better. We may say that the more big structures an image has, the more coarse content it would have, and a higher number of C&#x02013;S maps would probably lead to better performance. The edge density over the images may give us a rough estimation of the detailed and coarse content of the images. <xref ref-type="fig" rid="F6">Figure 6</xref> shows the results for edge detection of sample images in <xref ref-type="fig" rid="F5">Figure 5</xref> as examples for the best number of 4, 6, and 10 for C&#x02013;S difference maps. However, finding a direct relation between the best number of C&#x02013;S maps and the image content is not straightforward because the image content cannot be described by a single factor; but rather by several factors such as edge densities, frequency content, histogram of intensities, the number and the size of objects in the image, and image texture.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparing the AUC results of the model between three different cases of using 4, 6, and 10 C&#x02013;S difference maps for some sample images from the dataset. The results show that images with different contents require a different number of C&#x02013;S difference maps to result in high performance for the model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The results for edge detection of sample images of <xref ref-type="fig" rid="F5">Figure 5</xref>. The images are shown as examples for the best number of 4, 6, and 10 for C&#x02013;S difference maps. The edge detection shows a rough estimation of the detailed and coarse content of the images. The edges are obtained for the intensity channel of the images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0006.tif"/>
</fig>
<p>Based on the results in <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref>, we propose that finding a mechanism to apply in the models to get information about the image content and producing the C&#x02013;S difference maps based on the image content can make the models more efficient. In connection with the factors mentioned above to describe the image content, some approaches like frequency analysis, calculating image crowdedness, object detection, and edge detection can be investigated. Another approach that may be useful is a hierarchical segmentation of the image to determine the number of big and small structures in the image.</p>
<p>In <xref ref-type="fig" rid="F7">Figure 7</xref>, the AUC results of the model for the 20 categories of the dataset are shown for the three cases of applying 4, 6, and 10 C&#x02013;S difference maps. The AUC score is calculated in each category by averaging the AUC scores for the 100 images of the associated category for each case of using 4, 6, and 10 maps. The model has the highest performance for the Sketch category, probably because of the simpler structure of the images in this category, and it has the lowest performance for the Satellite category, probably because of the very low resolution and bad quality of the images in this category. We use approximation information of the wavelet transform of the image, and not the detailed information of WT, while due to the very low quality of the images in the satellite category, more detailed information of the image is required to make a high performance by the model.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>The mean AUC scores of the model for the 20 categories of the CAT2000 dataset for the three cases of applying 4, 6, and 10 C&#x02013;S difference maps in the model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnhum-16-862588-g0007.tif"/>
</fig>
<p>In each category, dependent on the overall content of the images in that category, the mean AUC score for that category may be higher for the case of applying 4, 6, or 10 C&#x02013;S difference maps. As discussed in the results of <xref ref-type="fig" rid="F5">Figure 5</xref>, relating the best number of C&#x02013;S maps and the image content requires deeper investigation. However, the results of some categories may be interpreted based on their image types. For example, in the LowResolution category, because of the images&#x00027; low resolution, the images&#x00027; detailed content is blurred; thus, the salient areas may be obtained mostly based on the coarse content of the images. Therefore, using 10 C&#x02013;S difference maps would lead to better performance for most of the images in this category, and a clearer difference is visible between the AUC results of using 10 C&#x02013;S maps compared to 4 and 6 C&#x02013;S maps. Similarly, in the category Noisy, because of the presence of noise in the images, and in the category Satellite, because of the images&#x00027; low quality, the detailed content of the images cannot be detected easily. Therefore, similar to the category LowResolution, using 10 C&#x02013;S maps would lead to better performance in these two categories.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>Discussion</title>
<p>Recently, the focus of saliency models on investigating bottom-up attention has decreased, while there are still open questions about applying bottom-up mechanisms. These questions include what features should be applied, how to obtain the contrasting areas, and how to integrate the information of different scales and features of the image. Here we present a bottom-up saliency model to address these questions and improve previous models. We present our improvements in different versions of the model as proposed models 1, 2, and 3 and compare their performances to show the effect of each improvement separately.</p>
<p>First, in proposed model 1, we integrate the information of different scales (described in Section Layer IV&#x02013;obtaining conspicuity maps) and features (described in Section Layer V&#x02013;obtaining saliency map) based on their weighted sum. The weight of contrast maps of different scales for a specific feature depends on the importance of the spatial frequency information of that map, which we calculate it using the contrast sensitivity function. This function of the visual system is not considered in so many models. The authors (Murray et al., <xref ref-type="bibr" rid="B46">2011</xref>; Buzatu, <xref ref-type="bibr" rid="B14">2012</xref>; Wang et al., <xref ref-type="bibr" rid="B58">2018</xref>) use this function in their model; however, they do not apply it for weighting different scales of the feature maps. In addition, we utilize both global and local weighting on the contrast maps of the scales.</p>
<p>The weights of the conspicuity maps of different features are defined based on a preliminary study testing 12 different variations of the integration methods and applying four different methods among them. Finally, we choose the method with the best performance among these four methods. The weights are defined based on the difference between the global maximum and local maxima in the related conspicuity map. This weighting method for features is applied in a similar way in some other models (Itti and Koch, <xref ref-type="bibr" rid="B30">2001</xref>; Frintrop et al., <xref ref-type="bibr" rid="B24">2007</xref>). In contrast to them, we calculate the weights of conspicuity maps to present both the local and global contrast of an area. Comparing our model 1 to other models (<xref ref-type="fig" rid="F4">Figure 4</xref>, <xref ref-type="table" rid="T2">Table 2</xref>) based on AUC and sAUC metrics suggests that the integration mechanisms applied in the model perform better than other non-learning-based models to make the model fit human data.</p>
<p>Furthermore, we compare our weighting method for integrating the conspicuity maps of different features with the method of calculating the maximum of conspicuity maps. The comparison results show that overall, the mean AUC score for the dataset images is higher for the weighting method than for calculating the maximum of conspicuity maps. However, as shown in the results section for some sample images (<xref ref-type="fig" rid="F3">Figure 3</xref>), in some images, the maximum method performs better than the weighting method. This shows that, although the weighting method leads to better overall performance, in some images, particular features may be too salient compared to other features in various areas of the image, and this overcoming saliency causes calculating the maximum of conspicuity maps leads to better results.</p>
<p>Second, in proposed model 2, we extend proposed model 1 so that, in addition to the low-level features commonly used in other models, we apply medium-level features based on the combination of color features with orientations (described in Section Layer I&#x02013;extracting the features). The increased performance of proposed model 2 compared to proposed model 1 (<xref ref-type="fig" rid="F4">Figure 4</xref>) suggests that the medium-level features also play a role in bottom-up visual attention behavior.</p>
<p>Third, in proposed model 3, we extend proposed model 2 so that we apply a variable number of C&#x02013;S difference maps instead of a fixed number as common in other models (described in Section Layer III&#x02013;creating center-surround difference maps). We discuss that human visual attention does not act the same for different images with different contents. To investigate this idea, we implemented the model using 4, 6, and 10 contrast maps for each feature. The chosen numbers refer, respectively, to the low, medium, and a high number of contrast maps. We calculated the AUC score by computing the maximum score among the scores of using 4, 6, and 10 C&#x02013;S difference maps. The considerably increased performance (AUC by 0.04 and sAUC by 0.02) of proposed model 3 compared to proposed model 2 in <xref ref-type="fig" rid="F4">Figure 4</xref> suggests that our proposal about applying a variable number of contrast maps can make the model more fitting to human data. This suggests that human visual attention may not act the same for different types of images, and the proper number of contrast maps for each image depends on the amount of detailed and coarse content in the image, as is shown for some sample images of the dataset in <xref ref-type="fig" rid="F5">Figure 5</xref>. In addition, comparing the results of applying different numbers of contrast maps for all 20 categories of the CAT2000 dataset (<xref ref-type="fig" rid="F7">Figure 7</xref>) confirmed our proposal. The center-surround difference mechanism is the base and essential mechanism of visual attention, and this improvement was the most important improvement made in our model.</p>
<p>We propose to apply a mechanism in the model to get information about the image content and adapt the number of C&#x02013;S difference maps to the image content, as discussed in Section The results of using a variable number of contrast maps. This adaptation mechanism can be used in the saliency models to improve their performances. Furthermore, the adaptation mechanism would minimize the additional computations due to applying a variable number of C&#x02013;S maps. The detailed and coarse content of the images may be roughly estimated by the images&#x00027; edge map, as shown in <xref ref-type="fig" rid="F6">Figure 6</xref> for some sample images. However, further investigation is required to find the factors that can be used to describe image content and be relative to the best number of C&#x02013;S maps. The factors such as edge densities, the ratio of low and high frequencies in the image, histogram of intensities, image crowdedness, and image texture could be investigated. Moreover, hierarchical segmentation can be applied to the image to extract the amount of big and small structures to indicate the amount of detailed and coarse content in the image. The analysis of image content could be improved even further to make it more adaptive in the way that a different number of C&#x02013;S maps is applied to distinct image regions with different structures. Finding an analyzing method to adapt the number of C&#x02013;S maps based on the image content is the scope of future study.</p>
<p>We validated our model&#x00027;s performance to describe human data by applying it to the CAT2000 dataset and measuring its performance based on the AUC and sAUC metrics. Comparing the results of our model to the other models that used the same dataset (<xref ref-type="fig" rid="F4">Figure 4</xref>) shows that our model has a high performance in the category of non-learning-based models, with the AUC of 0.73, 0.75, and 0.79, and sAUC of 0.54, 0.56, 0.58, respectively, for model versions 1, 2, and full version 3. The high performances of saliency models have been reported mostly for learning-based methods. The high performance of these learning-based models is related to the fact that they predict the fixations on the images by setting their parameters according to fixations on trained images. These learning-based models look promising; however, they usually require large human datasets to perform well. Also, learning-based methods may not be perfect for showing attention mechanisms. Usually, they show us what would be attended by the visual system but not how or why it may be attended. The advantage of our model is that it shows high performance according to AUC and sAUC metrics based on the applied mechanisms, and it does not have the complexity of learning-based methods that need to learn human data. We suggest that our proposed mechanisms can be added to the learning-based models to make them achieve higher performance and fit better with human data.</p>
</sec>
<sec sec-type="conclusions" id="s6">
<title>Conclusion</title>
<p>We proposed a saliency model that addresses some mechanisms of visual attention behavior that other models do not take into account. We defined three versions of the model to investigate the effect of each improvement separately. As a first improvement, in proposed model 1, we applied a weighted summation method for integrating the information of different scales and different features, defining the weights according to the contribution of each component, and presenting both the local and global saliency of the image. The model&#x00027;s high performance compared to the other models indicates that the integration methods were efficient. Furthermore, we compared the weighted summation method for combining the conspicuity maps of different features with the method of calculating the maximum of conspicuity maps. The comparison showed that although the weighted summation method leads by average to better performance, however, in some images where particular features dominate the other features, calculating the maximum of conspicuity maps can make better results. Second, in proposed model 2, in addition to the common low-level features, we added the medium-level features to proposed model 1. The increased AUC and sAUC of proposed model 2 compared to model 1 suggests that medium-level features may play a role in the behavior of visual attention. Third, and most importantly, in proposed model 3, instead of the common approach of using a fixed number of C&#x02013;S difference maps, we added the mechanism of variable numbers of these maps to proposed model 2, proposing that human visual attention performs differently for different images. The increased AUC and sAUC of proposed model 3 compared to model 2 confirms our proposal about the center-surround mechanism. The mechanism of a variable number of C&#x02013;S difference maps can improve further to make it adaptive to the image content.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>SN contributed to the conceptualization of the study, methodology, writing the first draft of the manuscript, and visualization. AF contributed to the supervision of the study. SR and MD contributed to manuscript revision. All authors approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>This research has been supported by the Cognitive Sciences and Technologies Council of Iran.</p>
</sec>
<ack><p>The authors would like to thank Frederik Beuth for his helpful discussions and comments.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abkenar</surname> <given-names>M. R.</given-names></name> <name><surname>Ahmad</surname> <given-names>M. O.</given-names></name></person-group> (<year>2016</year>). <article-title>Superpixel-based salient region detection using the wavelet transform</article-title>, in <source>2016 IEEE International Symposium on Circuits and Systems (ISCAS)</source> (<publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2719</fpage>&#x02013;<lpage>2722</lpage>. <pub-id pub-id-type="doi">10.1109/ISCAS.2016.7539154</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Achanta</surname> <given-names>R.</given-names></name> <name><surname>Hemami</surname> <given-names>S.</given-names></name> <name><surname>Estrada</surname> <given-names>F.</given-names></name> <name><surname>Susstrunk</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Frequency-tuned salient region detection</article-title>, in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1597</fpage>&#x02013;<lpage>1604</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206596</pub-id><pub-id pub-id-type="pmid">35957389</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Antonini</surname> <given-names>M.</given-names></name> <name><surname>Barlaud</surname> <given-names>M.</given-names></name> <name><surname>Mathieu</surname> <given-names>P.</given-names></name> <name><surname>Daubechies</surname> <given-names>I.</given-names></name></person-group> (<year>1992</year>). <article-title>Image coding using wavelet transform</article-title>. <source>IEEE Trans. Image Proc.</source> <volume>1</volume>, <fpage>205</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1109/83.136597</pub-id><pub-id pub-id-type="pmid">18296155</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ayoub</surname> <given-names>N.</given-names></name> <name><surname>Gao</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Tobji</surname> <given-names>R.</given-names></name> <name><surname>Yao</surname> <given-names>N.</given-names></name></person-group> (<year>2018</year>). <article-title>Visual saliency detection based on color frequency features under Bayesian framework</article-title>. <source>KSII Trans. Internet Inform. Syst.</source> <volume>12</volume>, <fpage>676</fpage>&#x02013;<lpage>692</lpage>. <pub-id pub-id-type="doi">10.3837/tiis.2018.02.008</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bonnar</surname> <given-names>L.</given-names></name> <name><surname>Gosselin</surname> <given-names>F.</given-names></name> <name><surname>Schyns</surname> <given-names>P. G.</given-names></name></person-group> (<year>2002</year>). <article-title>Understanding Dali&#x00027;s Slave market with the disappearing bust of voltaire: a case study in the scale information driving perception</article-title>. <source>Perception</source> <volume>31</volume>, <fpage>683</fpage>&#x02013;<lpage>691</lpage>. <pub-id pub-id-type="doi">10.1068/p3276</pub-id><pub-id pub-id-type="pmid">12092795</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Exploiting local and global patch rarities for saliency detection</article-title>, in <source>Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Providence, RI</publisher-loc>), <fpage>478</fpage>&#x02013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2012.6247711</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Saliency prediction in the deep learning era: successes and limitations</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>43</volume>, <fpage>679</fpage>&#x02013;<lpage>700</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2019.2935715</pub-id><pub-id pub-id-type="pmid">31425064</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A.</given-names></name> <name><surname>Ahmadabadi</surname> <given-names>M. N.</given-names></name> <name><surname>Araabi</surname> <given-names>B. N.</given-names></name></person-group> (<year>2011</year>). <article-title>Cost-sensitive learning of top-down modulation for attentional control</article-title>. <source>Mach. Vis. Applic.</source> <volume>22</volume>, <fpage>61</fpage>&#x02013;<lpage>76</lpage>. <pub-id pub-id-type="doi">10.1007/s00138-009-0192-0</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A.</given-names></name> <name><surname>Itti</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>State-of-the-Art in visual attention modeling</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>1</volume>, <fpage>185</fpage>&#x02013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2012.89</pub-id><pub-id pub-id-type="pmid">22487985</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A.</given-names></name> <name><surname>Itti</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>CAT2000: a large scale fixation dataset for boosting saliency research</article-title>. <source>rXiv Preprint</source> arXiv:1505.03581. Retrieved from <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1505.03581">http://arxiv.org/abs/1505.03581</ext-link></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Borji</surname> <given-names>A.</given-names></name> <name><surname>Tavakoli</surname> <given-names>H. R.</given-names></name> <name><surname>Sihite</surname> <given-names>D. N.</given-names></name> <name><surname>Itti</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>Analysis of scores, datasets, and models in visual saliency prediction</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Sydney, NSW</publisher-loc>), <fpage>921</fpage>&#x02013;<lpage>928</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2013.118</pub-id><pub-id pub-id-type="pmid">22868572</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brigham</surname> <given-names>E. O.</given-names></name> <name><surname>Morrow</surname> <given-names>R. E.</given-names></name></person-group> (<year>1967</year>). <article-title>The fast Fourier transform</article-title>. <source>IEEE Spectr.</source> <volume>4</volume>, <fpage>63</fpage>&#x02013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1109/MSPEC.1967.5217220</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bruce</surname> <given-names>N. D. B.</given-names></name> <name><surname>Tsotsos</surname> <given-names>J. K.</given-names></name></person-group> (<year>2005</year>). <article-title>Saliency based on information maximization</article-title>, in <source>Proceedings of the 18th International Conference on Neural Information Processing Systems</source>, eds <person-group person-group-type="editor"><name><surname>Weiss</surname> <given-names>Y.</given-names></name> <name><surname>Scholkopf</surname> <given-names>B.</given-names></name> <name><surname>Platt</surname> <given-names>J. C.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>155</fpage>&#x02013;<lpage>162</lpage>. <pub-id pub-id-type="doi">10.5555/2976248.2976268</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buzatu</surname> <given-names>O. L.</given-names></name></person-group> (<year>2012</year>). <article-title>Human visual perception concepts as mechanisms for saliency detection</article-title>. <source>Acta Tech. Napocensis</source> <volume>53</volume>, <fpage>25</fpage>&#x02013;<lpage>30</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bylinskii</surname> <given-names>Z.</given-names></name> <name><surname>Judd</surname> <given-names>T.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Durand</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <article-title>What do different evaluation metrics tell us about saliency models?</article-title> <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>41</volume>, <fpage>740</fpage>&#x02013;<lpage>757</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2018.2815601</pub-id><pub-id pub-id-type="pmid">29993800</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carrasco</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Visual attention: the past 25 years</article-title>. <source>Vis. Res.</source> <volume>51</volume>, <fpage>1484</fpage>&#x02013;<lpage>1525</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2011.04.012</pub-id><pub-id pub-id-type="pmid">21549742</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Casagrande</surname> <given-names>V. A.</given-names></name> <name><surname>Norton</surname> <given-names>T. T.</given-names></name></person-group> (<year>1991</year>). <article-title>The neural basis of vision function: vision and visual dysfunction BT&#x02014;the neural basis of vision function: vision and visual dysfunction</article-title>. <source>Neural Basis Vis. Funct. Vis. Vis. Dysfunct.</source> <volume>4</volume>, <fpage>41</fpage>&#x02013;<lpage>84</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Connor</surname> <given-names>C. E.</given-names></name> <name><surname>Egeth</surname> <given-names>H. E.</given-names></name> <name><surname>Yantis</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>Visual attention: bottom-up versus top-down</article-title>. <source>Curr. Biol.</source> <volume>14</volume>, <fpage>R850</fpage>&#x02013;<lpage>R852</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2004.09.041</pub-id><pub-id pub-id-type="pmid">15458666</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deco</surname> <given-names>G.</given-names></name> <name><surname>Zihl</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>Top-down selective visual attention: a neurodynamical approach</article-title>. <source>Vis. Cogn.</source> <volume>8</volume>, <fpage>118</fpage>&#x02013;<lpage>139</lpage>. <pub-id pub-id-type="doi">10.1080/13506280042000054</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dehghani</surname> <given-names>A.</given-names></name> <name><surname>Soltanian-Zadeh</surname> <given-names>H.</given-names></name> <name><surname>Hossein-Zadeh</surname> <given-names>G.-A.</given-names></name></person-group> (<year>2020</year>). <article-title>Global data-driven analysis of brain connectivity during emotion regulation by electroencephalography neurofeedback</article-title>. <source>Brain Connect.</source> <volume>10</volume>, <fpage>302</fpage>&#x02013;<lpage>315</lpage>. <pub-id pub-id-type="doi">10.1089/brain.2019.0734</pub-id><pub-id pub-id-type="pmid">32458692</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Denison</surname> <given-names>R. N.</given-names></name> <name><surname>Carrasco</surname> <given-names>M.</given-names></name> <name><surname>Heeger</surname> <given-names>D. J.</given-names></name></person-group> (<year>2021</year>). <article-title>A dynamic normalization model of temporal attention</article-title>. <source>Nat. Hum. Behav</source>. <volume>5</volume>, <fpage>1674</fpage>&#x02013;<lpage>1685</lpage>. <pub-id pub-id-type="doi">10.1038/s41562-021-01129-1</pub-id><pub-id pub-id-type="pmid">34140658</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doricchi</surname> <given-names>F.</given-names></name> <name><surname>Macci</surname> <given-names>E.</given-names></name> <name><surname>Silvetti</surname> <given-names>M.</given-names></name> <name><surname>Macaluso</surname> <given-names>E.</given-names></name></person-group> (<year>2010</year>). <article-title>Neural correlates of the spatial and expectancy components of endogenous and stimulus-driven orienting of attention in the Posner task</article-title>. <source>Cereb. Cortex</source> <volume>20</volume>, <fpage>1574</fpage>&#x02013;<lpage>1585</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhp215</pub-id><pub-id pub-id-type="pmid">19846472</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frintrop</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>VOCUS: a visual attention system for object detection and goal-directed search</article-title>. <source>Lecture Notes Artif. Intell.</source> <volume>3899</volume>, <fpage>1</fpage>&#x02013;<lpage>197</lpage>. <pub-id pub-id-type="doi">10.1007/11682110</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Frintrop</surname> <given-names>S.</given-names></name> <name><surname>Klodt</surname> <given-names>M.</given-names></name> <name><surname>Rome</surname> <given-names>E.</given-names></name></person-group> (<year>2007</year>). <article-title>A real-time visual attention system using integral images</article-title>. <source>Science</source>. <volume>11</volume>, <fpage>191</fpage>&#x02013;<lpage>193</lpage>. Retrieved from <ext-link ext-link-type="uri" xlink:href="http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.91.2449&#x00026;amp;rep=rep1&#x00026;amp;type=pdf">http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.91.2449&#x00026;amp;rep=rep1&#x00026;amp;type=pdf</ext-link></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Frintrop</surname> <given-names>S.</given-names></name> <name><surname>Werner</surname> <given-names>T.</given-names></name> <name><surname>Martin Garcia</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Traditional saliency reloaded: a good old model in new shape</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>82</fpage>&#x02013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298603</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ghasemi</surname> <given-names>J. B.</given-names></name> <name><surname>Heidari</surname> <given-names>Z.</given-names></name> <name><surname>Jabbari</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Toward a continuous wavelet transform-based search method for feature selection for classification of spectroscopic data</article-title>. <source>Chemometr. Intell. Lab. Syst.</source> <volume>127</volume>, <fpage>185</fpage>&#x02013;<lpage>194</lpage>. <pub-id pub-id-type="doi">10.1016/j.chemolab.2013.06.008</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goferman</surname> <given-names>S.</given-names></name> <name><surname>Zelnik-Manor</surname> <given-names>L.</given-names></name> <name><surname>Tal</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Context-Aware saliency detection</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>34</volume>, <fpage>1915</fpage>&#x02013;<lpage>1926</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2011.272</pub-id><pub-id pub-id-type="pmid">22201056</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2007</year>). <article-title>Saliency detection: a spectral residual approach</article-title>, in <source>2007 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Minneapolis, MN</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2007.383267</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Imamoglu</surname> <given-names>N.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Fang</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>A saliency detection model using low-level features based on wavelet transform</article-title>. <source>IEEE Trans. Multimedia</source> <volume>15</volume>, <fpage>96</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2012.2225034</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Itti</surname> <given-names>L.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>2001</year>). <article-title>Feature combination strategies for saliency-based visual attention systems</article-title>. <source>J. Electro. Imaging</source> <volume>10</volume>, <fpage>161</fpage>&#x02013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1117/1.1333677</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Itti</surname> <given-names>L.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name> <name><surname>Niebur</surname> <given-names>E.</given-names></name></person-group> (<year>1998</year>). <article-title>A model of saliency-based visual attention for rapid scene analysis</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>20</volume>, <fpage>1254</fpage>&#x02013;<lpage>1259</lpage>. <pub-id pub-id-type="doi">10.1109/34.730558</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jian</surname> <given-names>M.</given-names></name> <name><surname>Lam</surname> <given-names>K.-M.</given-names></name> <name><surname>Dong</surname> <given-names>J.</given-names></name> <name><surname>Shen</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Visual-patch-attention-aware saliency detection</article-title>. <source>IEEE Trans. Cybernet.</source> <volume>8</volume>, <fpage>1575</fpage>&#x02013;<lpage>1586</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2014.2356200</pub-id><pub-id pub-id-type="pmid">25291809</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jian</surname> <given-names>M.</given-names></name> <name><surname>Zhou</surname> <given-names>Q.</given-names></name> <name><surname>Cui</surname> <given-names>C.</given-names></name> <name><surname>Nie</surname> <given-names>X.</given-names></name> <name><surname>Luo</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Assessment of feature fusion strategies in visual attention mechanism for saliency detection</article-title>. <source>Patt. Recogn. Lett.</source> <volume>127</volume>, <fpage>37</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2018.08.022</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Judd</surname> <given-names>T.</given-names></name> <name><surname>Ehinger</surname> <given-names>K.</given-names></name> <name><surname>Durand</surname> <given-names>F.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Learning to predict where humans look</article-title>, in <source>2009 IEEE 12th International Conference on Computer Vision</source> (<publisher-loc>Kyoto</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2106</fpage>&#x02013;<lpage>2113</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2009.5459462</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koene</surname> <given-names>A. R.</given-names></name> <name><surname>Zhaoping</surname> <given-names>L.</given-names></name></person-group> (<year>2007</year>). <article-title>Feature-specific interactions in salience from combined feature contrasts: evidence for a bottom-up saliency map in V1</article-title>. <source>J. Vis.</source> <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1167/7.7.6</pub-id><pub-id pub-id-type="pmid">17685802</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krasovskaya</surname> <given-names>S.</given-names></name> <name><surname>MacInnes</surname> <given-names>W. J.</given-names></name></person-group> (<year>2019</year>). <article-title>Salience models: a computational cognitive neuroscience review</article-title>. <source>Vision</source> <volume>3</volume>, <fpage>56</fpage>. <pub-id pub-id-type="doi">10.3390/vision3040056</pub-id><pub-id pub-id-type="pmid">31735857</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kroner</surname> <given-names>A.</given-names></name> <name><surname>Senden</surname> <given-names>M.</given-names></name> <name><surname>Driessens</surname> <given-names>K.</given-names></name> <name><surname>Goebel</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Contextual encoder&#x02013;decoder network for visual saliency prediction</article-title>. <source>Neural Netw.</source> <volume>129</volume>, <fpage>261</fpage>&#x02013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2020.05.004</pub-id><pub-id pub-id-type="pmid">32563023</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kruthiventi</surname> <given-names>S. S. S.</given-names></name> <name><surname>Ayush</surname> <given-names>K.</given-names></name> <name><surname>Babu</surname> <given-names>R. V.</given-names></name></person-group> (<year>2017</year>). <article-title>Deepfix: a fully convolutional neural network for predicting human eye fixations</article-title>. <source>IEEE Trans. Image Process.</source> <volume>26</volume>, <fpage>4446</fpage>&#x02013;<lpage>4456</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2710620</pub-id><pub-id pub-id-type="pmid">28692956</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kummerer</surname> <given-names>M.</given-names></name> <name><surname>Wallis</surname> <given-names>T. S. A.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Saliency benchmarking made easy: separating models, maps and metrics</article-title>, in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>770</fpage>&#x02013;<lpage>787</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01270-0_47</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kummerer</surname> <given-names>M.</given-names></name> <name><surname>Wallis</surname> <given-names>T. S. A.</given-names></name> <name><surname>Gatys</surname> <given-names>L. A.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Understanding low-and high-level contributions to fixation prediction</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>4789</fpage>&#x02013;<lpage>4798</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.513</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Yuan</surname> <given-names>Z.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Zheng</surname> <given-names>N.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Learning to detect a salient object</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>33</volume>, <fpage>353</fpage>&#x02013;<lpage>367</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2010.70</pub-id><pub-id pub-id-type="pmid">21193811</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Lam</surname> <given-names>K.-M.</given-names></name> <name><surname>Zhong</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>Efficient saliency analysis based on wavelet transform and entropy theory</article-title>. <source>J. Vis. Commun. Image Represent.</source> <volume>30</volume>, <fpage>201</fpage>&#x02013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2015.04.008</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Lam</surname> <given-names>K. M.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>Saliency analysis based on multi-scale wavelet decomposition</article-title>, in <source>2013 16th International IEEE Conference on Intelligent Transportation Systems: Intelligent Transportation Systems for All Modes, ITSC</source> (<publisher-loc>The Hague</publisher-loc>), <fpage>1977</fpage>&#x02013;<lpage>1980</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mannos</surname> <given-names>J.</given-names></name> <name><surname>Sakrison</surname> <given-names>D.</given-names></name></person-group> (<year>1974</year>). <article-title>The effects of a visual fidelity criterion of the encoding of images</article-title>. <source>IEEE Trans. Inform. Theory</source> <volume>20</volume>, <fpage>525</fpage>&#x02013;<lpage>536</lpage>. <pub-id pub-id-type="doi">10.1109/TIT.1974.1055250</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Molin</surname> <given-names>J. L.</given-names></name> <name><surname>Thakur</surname> <given-names>C. S.</given-names></name> <name><surname>Niebur</surname> <given-names>E.</given-names></name> <name><surname>Etienne-Cummings</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>A neuromorphic proto-object based dynamic visual saliency model with a hybrid FPGA implementation</article-title>. <source>IEEE Trans. Biomed. Circ. Syst.</source> <volume>15</volume>, <fpage>580</fpage>&#x02013;<lpage>594</lpage>. <pub-id pub-id-type="doi">10.1109/TBCAS.2021.3089622</pub-id><pub-id pub-id-type="pmid">34133287</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Murray</surname> <given-names>N.</given-names></name> <name><surname>Vanrell</surname> <given-names>M.</given-names></name> <name><surname>Otazu</surname> <given-names>X.</given-names></name> <name><surname>Parraga</surname> <given-names>C. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Saliency estimation using a non-parametric low-level vision model</article-title>, in <source>Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Colorado Springs, CO</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>433</fpage>&#x02013;<lpage>440</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2011.5995506</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Narayanaswamy</surname> <given-names>M.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Fung</surname> <given-names>W. K.</given-names></name> <name><surname>Fough</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>A low-complexity wavelet-based visual saliency model to predict fixations</article-title>, in <source>2020 27th IEEE International Conference on Electronics, Circuits and Systems (ICECS)</source> (<publisher-loc>Glasgow</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1109/ICECS49266.2020.9294905</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Novin</surname> <given-names>S.</given-names></name> <name><surname>Fallah</surname> <given-names>A.</given-names></name> <name><surname>Rashidi</surname> <given-names>S.</given-names></name> <name><surname>Beuth</surname> <given-names>F.</given-names></name> <name><surname>Hamker</surname> <given-names>F. H.</given-names></name></person-group> (<year>2021</year>). <article-title>A neuro-computational model of visual attention with multiple attentional control sets</article-title>. <source>Vis. Res.</source> <volume>189</volume>, <fpage>104</fpage>&#x02013;<lpage>118</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2021.08.009</pub-id><pub-id pub-id-type="pmid">34749237</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Otazu</surname> <given-names>X.</given-names></name> <name><surname>Parraga</surname> <given-names>C. A.</given-names></name> <name><surname>Vanrell</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Toward a unified chromatic induction model</article-title>. <source>J. Vis.</source> <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1167/10.12.5</pub-id><pub-id pub-id-type="pmid">21047737</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qi</surname> <given-names>F.</given-names></name> <name><surname>Lin</surname> <given-names>C.</given-names></name> <name><surname>Shi</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>A convolutional encoder-decoder network with skip connections for saliency prediction</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>60428</fpage>&#x02013;<lpage>60438</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2915630</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sadjadi</surname> <given-names>S. M.</given-names></name> <name><surname>Ebrahimzadeh</surname> <given-names>E.</given-names></name> <name><surname>Shams</surname> <given-names>M.</given-names></name> <name><surname>Seraji</surname> <given-names>M.</given-names></name> <name><surname>Soltanian-Zadeh</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Localization of epileptic foci based on simultaneous EEG&#x02013;fMRI data</article-title>. <source>Front. Neurol.</source> <volume>12</volume>, <fpage>645594</fpage>. <pub-id pub-id-type="doi">10.3389/fneur.2021.645594</pub-id><pub-id pub-id-type="pmid">33986718</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>N.</given-names></name> <name><surname>Mishra</surname> <given-names>K. K.</given-names></name> <name><surname>Bhatia</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>SEAM-an improved environmental adaptation method with real parameter coding for salient object detection</article-title>. <source>Multimedia Tools Applic.</source> <volume>79</volume>, <fpage>12995</fpage>&#x02013;<lpage>13010</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-020-08678-z</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tavakoli</surname> <given-names>H. R.</given-names></name> <name><surname>Rahtu</surname> <given-names>E.</given-names></name> <name><surname>Heikkil&#x000E4;</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>Fast and efficient saliency detection using sparse sampling and kernel density estimation</article-title>, in <source>Proceedings of the 17th Scandinavian Conference on Image Analysis</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>666</fpage>&#x02013;<lpage>675</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-21227-7_62</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ts&#x00027;o</surname> <given-names>D. Y.</given-names></name> <name><surname>Gilbert</surname> <given-names>C. D.</given-names></name></person-group> (<year>1988</year>). <article-title>The organization of chromatic and spatial interactions in the primate striate cortex</article-title>. <source>J. Neurosci.</source> <volume>8</volume>, <fpage>1712</fpage>&#x02013;<lpage>1727</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.08-05-01712.1988</pub-id><pub-id pub-id-type="pmid">3367218</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ullah</surname> <given-names>I.</given-names></name> <name><surname>Jian</surname> <given-names>M.</given-names></name> <name><surname>Hussain</surname> <given-names>S.</given-names></name> <name><surname>Guo</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A brief survey of visual saliency detection</article-title>. <source>Multimedia Tools Applic.</source> <volume>79</volume>, <fpage>34605</fpage>&#x02013;<lpage>34645</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-020-08849-y</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vig</surname> <given-names>E.</given-names></name> <name><surname>Dorr</surname> <given-names>M.</given-names></name> <name><surname>Cox</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>Large-scale optimization of hierarchical features for saliency prediction in natural images</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Columbus, OH</publisher-loc>), <fpage>2798</fpage>&#x02013;<lpage>2805</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2014.358</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Walther</surname> <given-names>D.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>2006</year>). <article-title>Modeling attention to salient proto-objects</article-title>. <source>Neural Netw.</source> <volume>19</volume>, <fpage>1395</fpage>&#x02013;<lpage>1407</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2006.10.001</pub-id><pub-id pub-id-type="pmid">17098563</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Wan</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Visual saliency based just noticeable difference estimation in DWT domain</article-title>. <source>Information</source> <volume>9</volume>, <fpage>178</fpage>. <pub-id pub-id-type="doi">10.3390/info9070178</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>L.</given-names></name> <name><surname>Luo</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>A biologically inspired computational approach to model top-down and bottom-up visual attention</article-title>. <source>Optik</source> <volume>126</volume>, <fpage>522</fpage>&#x02013;<lpage>529</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijleo.2015.01.004</pub-id><pub-id pub-id-type="pmid">21931178</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wolfe</surname> <given-names>J. M.</given-names></name> <name><surname>Horowitz</surname> <given-names>T. S.</given-names></name></person-group> (<year>2004</year>). <article-title>What attributes guide the deployment of visual attention and how do they do it?</article-title> <source>Nat. Rev. Neurosci.</source> <volume>5</volume>, <fpage>495</fpage>&#x02013;<lpage>501</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1411</pub-id><pub-id pub-id-type="pmid">15152199</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>W.</given-names></name> <name><surname>Yang</surname> <given-names>M.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name> <name><surname>Al-Kabbany</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>An improved saliency detection using wavelet transform</article-title>, in <source>2015 IEEE International Conference on Communication Software and Networks (ICCSN)</source> (<publisher-loc>Chengdu</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>345</fpage>&#x02013;<lpage>351</lpage>. <pub-id pub-id-type="doi">10.1109/ICCSN.2015.7296181</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Sclaroff</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Exploiting surroundedness for saliency detection: a boolean map approach</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>38</volume>, <fpage>889</fpage>&#x02013;<lpage>902</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2473844</pub-id><pub-id pub-id-type="pmid">26336114</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Tong</surname> <given-names>M. H.</given-names></name> <name><surname>Marks</surname> <given-names>T. K.</given-names></name> <name><surname>Shan</surname> <given-names>H.</given-names></name> <name><surname>Cottrell</surname> <given-names>G. W.</given-names></name></person-group> (<year>2008</year>). <article-title>SUN: a Bayesian framework for saliency using natural statistics</article-title>. <source>J. Vis.</source> <volume>8</volume>, <fpage>1</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1167/8.7.32</pub-id><pub-id pub-id-type="pmid">19146264</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>Infrared small target detection based on morphology and wavelet transform</article-title>, in <source>2011 2nd International Conference on Artificial Intelligence, Management Science and Electronic Commerce (AIMSEC)</source> (<publisher-loc>Deng Feng</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4033</fpage>&#x02013;<lpage>4036</lpage>.</citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Q.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>Learning a saliency map using fixated locations in natural scenes</article-title>. <source>J. Vis.</source> <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1167/11.3.9</pub-id><pub-id pub-id-type="pmid">21393388</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name> <name><surname>Mu</surname> <given-names>N.</given-names></name></person-group> (<year>2019</year>). <article-title>Saliency detection based on the combination of high-level knowledge and low-level cues in foggy images</article-title>. <source>Entropy</source> <volume>21</volume>, <fpage>374</fpage>. <pub-id pub-id-type="doi">10.3390/e21040374</pub-id><pub-id pub-id-type="pmid">33267088</pub-id></citation></ref>
</ref-list>
</back>
</article>