<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2017.00058</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Joint Learning of Binocularly Driven Saccades and Vergence by Active Efficient Coding</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhu</surname> <given-names>Qingpeng</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<uri xlink:href="http://frontiersin.org/people/u/285802"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Triesch</surname> <given-names>Jochen</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/1023"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Shi</surname> <given-names>Bertram E.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/12769"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology</institution>, <addr-line>Hong Kong</addr-line>, <country>Hong Kong</country></aff>
<aff id="aff2"><sup>2</sup><institution>Frankfurt Institute for Advanced Studies</institution>, <addr-line>Frankfurt am Main</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jeffrey L. Krichmar, University of California, Irvine, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Michael Beyeler, University of Washington, United States; Jaerock Kwon, Kettering University, United States</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Qingpeng Zhu, <email>qzhuad&#x00040;ust.hk</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>03</day>
<month>11</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>11</volume>
<elocation-id>58</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>07</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>10</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Zhu, Triesch and Shi.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Zhu, Triesch and Shi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>This paper investigates two types of eye movements: vergence and saccades. Vergence eye movements are responsible for bringing the images of the two eyes into correspondence, whereas saccades drive gaze to interesting regions in the scene. Control of both vergence and saccades develops during early infancy. To date, these two types of eye movements have been studied separately. Here, we propose a computational model of an active vision system that integrates these two types of eye movements. We hypothesize that incorporating a saccade strategy driven by bottom-up attention will benefit the development of vergence control. The integrated system is based on the active efficient coding framework, which describes the joint development of sensory-processing and eye movement control to jointly optimize the coding efficiency of the sensory system. In the integrated system, we propose a binocular saliency model to drive saccades based on learned binocular feature extractors, which simultaneously encode both depth and texture information. Saliency in our model also depends on the current fixation point. This extends prior work, which focused on monocular images and saliency measures that are independent of the current fixation. Our results show that the proposed saliency-driven saccades lead to better vergence performance and faster learning in the overall system than random saccades. Faster learning is significant because it indicates that the system actively selects inputs for the most effective learning. This work suggests that saliency-driven saccades provide a scaffold for the development of vergence control during infancy.</p>
</abstract>
<kwd-group>
<kwd>active efficient coding</kwd>
<kwd>saccades</kwd>
<kwd>vergence</kwd>
<kwd>binocular saliency map</kwd>
<kwd>generative adaptive subspace self-organizing map</kwd>
<kwd>reinforcement learning</kwd>
</kwd-group>
<counts>
<fig-count count="12"/>
<table-count count="0"/>
<equation-count count="26"/>
<ref-count count="40"/>
<page-count count="15"/>
<word-count count="10959"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<title>Introduction</title>
<p>Biological vision systems are often active and rely on a number of eye movements to sense the environment. Remarkably, these vision systems have the ability to autonomously self-calibrate, but the underlying mechanisms are still poorly understood. Here, we focus on vergence and saccadic eye movements. Vergence eye movements are slow and disconjugate (the two eyes move in opposite directions). They serve to align the images acquired by the two eyes so that they can be binocularly fused. Saccadic eye movements are rapid and conjugate (the two eyes move in the same direction). They serve to direct gaze so that the fovea, the region with highest visual acuity, falls on objects of interest. The two types of eye movements often cooccur. For example, they are both involved when eye movements are made to direct gaze toward different objects in a 3D scene (Yang et al., <xref ref-type="bibr" rid="B34">2002</xref>). The association between vergence and saccades facilitates the imaging of objects of interest onto the fovea of both eyes (Zee et al., <xref ref-type="bibr" rid="B36">1992</xref>).</p>
<p>Saccades are of great importance for human vision. At any time, the visual system receives a large amount of information from the environment, but has limited capacity for sensing and signal processing. Humans use saccades to direct foveal vision toward places with relevant information (Yarbus, <xref ref-type="bibr" rid="B35">1967</xref>; Renninger et al., <xref ref-type="bibr" rid="B26">2007</xref>). These saccades can be driven by top-down (Gao et al., <xref ref-type="bibr" rid="B11">2009</xref>; Kanan et al., <xref ref-type="bibr" rid="B16">2009</xref>; Yang and Yang, <xref ref-type="bibr" rid="B33">2012</xref>) or bottom-up (Itti et al., <xref ref-type="bibr" rid="B15">1998</xref>; Hou and Zhang, <xref ref-type="bibr" rid="B14">2007</xref>; Zhang et al., <xref ref-type="bibr" rid="B39">2008</xref>; Bruce and Tsotsos, <xref ref-type="bibr" rid="B6">2009</xref>; Han et al., <xref ref-type="bibr" rid="B13">2011</xref>) attention mechanisms. Top-down attention is voluntary and task-driven, whereas bottom-up attention is involuntary and stimulus-driven. We focus here on the bottom-up mechanism, where saccades are assumed to be generated according to a saliency map, which assigns salience to different points on an image by combining a number of low-level features. For example, Itti et al. (<xref ref-type="bibr" rid="B15">1998</xref>) proposed to generate a saliency map by combining the outputs of feature maps that are sensitive to different features, such as color, intensity, and orientation. Since the primary visual cortex (area V1) is one of the first stages of visual information processing, many models have been inspired by the processing found there. For example, Li (<xref ref-type="bibr" rid="B18">2002</xref>) proposed to generate the saliency map by combining the responses of model V1 neurons tuned to input features such as orientation and color. The attention based on information maximization (AIM) saliency model, proposed by Bruce and Tsotsos (<xref ref-type="bibr" rid="B6">2009</xref>), combines feature maps generated by a set of learned basis functions that are similar to the receptive fields of V1 neurons.</p>
<p>Most proposed saliency models, including those described earlier, assume monocular images, ignoring the importance of depth in human vision. Depth cues play an important role in visual attention (Wolfe and Horowitz, <xref ref-type="bibr" rid="B32">2004</xref>) and have a strong relationship with objects, since depth discontinuities suggest object boundaries. Although some saliency models have incorporated depth cues, they have typically processed depth and 2D texture information separately, e.g., by combining saliency maps computed by considering each cue in isolation. For example, Wang et al. (<xref ref-type="bibr" rid="B31">2013</xref>) proposed a visual saliency model that combines a saliency map computed from disparity with a saliency map computed from monocular visual features. Liu et al. (<xref ref-type="bibr" rid="B19">2012</xref>) proposed a saliency model where disparity information is extracted by taking the difference between the left and right images. They compute the overall saliency as the weighted average of the saliencies computed from disparity, color, and intensity separately.</p>
<p>We describe here an integrated vision system that combines binocular vergence control and binocularly driven saccadic eye movements. This system extends prior work on learning binocular vergence control using the active efficient coding (AEC) framework, first proposed by Zhao et al. (<xref ref-type="bibr" rid="B40">2012</xref>). The AEC framework is an extension of Barlow&#x02019;s (Barlow, <xref ref-type="bibr" rid="B2">1961</xref>) efficient coding hypothesis, which states that the activity of the sensory-processing neurons encodes their input using as few spikes as possible. A primary prediction of the efficient encoding hypothesis is that the properties of the sensory-processing neurons adapt to the statistics of the input stimuli. The AEC framework extends the efficient coding hypothesis to include the effect of behavior. It posits that in addition, the organism&#x02019;s behavior adapts so that the input can be efficiently encoded. By combining unsupervised and reinforcement learning, AEC simultaneously learns both a distributed representation of the sensory input and a policy for mapping this representation to motor commands. Thus, it jointly learns both perception and action as the organism behaves in the environment. In previous work, AEC has been shown to model the development of many reflexive eye movements and other behaviors, such as vergence control (Zhao et al., <xref ref-type="bibr" rid="B40">2012</xref>; Lonini et al., <xref ref-type="bibr" rid="B20">2013</xref>; Klimmasch et al., <xref ref-type="bibr" rid="B17">2017</xref>), smooth pursuit (Zhang et al., <xref ref-type="bibr" rid="B38">2014</xref>; Teuli&#x000E8;re et al., <xref ref-type="bibr" rid="B27">2015</xref>), optokinetic nystagmus (Zhang et al., <xref ref-type="bibr" rid="B37">2016</xref>), the combination of vergence and smooth pursuit (Vikram et al., <xref ref-type="bibr" rid="B30">2014</xref>), and imitation learning (Triesch, <xref ref-type="bibr" rid="B29">2013</xref>).</p>
<p>We make several contributions in this work. First, in the original work by Zhao et al. (<xref ref-type="bibr" rid="B40">2012</xref>), saccades were generated completely randomly. This paper integrates the vergence control process with a more realistic model of saccade generation. Second, we extend Bruce and Tsotsos&#x02019;s (Bruce and Tsotsos, <xref ref-type="bibr" rid="B6">2009</xref>) AIM saliency model, which was formulated for monocular images, to binocular images. In particular, instead of using a set of fixed pre-trained monocular basis functions learned on a separate database, our model uses a set of binocular basis functions that are learned as the model agent interacts with the environment. These low-level binocular features integrate depth and texture information much earlier than in the prior work described earlier and are consistent with what is known about the visual cortex. Poggio and Fischer (<xref ref-type="bibr" rid="B24">1977</xref>) claimed that most cortical neurons (84%) are sensitive to the depth of a stimulus. Third, we propose a saliency model, where the saliency of a given point depends upon the current fixation point, whereas most prior saliency models have assigned saliency independently of the current fixation point. In this version of the saliency model, image points that are different from the current fixation point in terms of appearance or depth are more salient. Fourth, rather than treating saliency and vergence control as two separate problems, our model exhibits a very close coupling between the two. Not only are the two behaviors learned at the same time but they also share the same set of low-level feature detectors.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2-1">
<title>Architecture Overview</title>
<p>Our model assumes the robot is in an environment that has multiple objects located at different depths. We have made a video<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> to demonstrate how our system works in the iCub simulator (Tikhanoff et al., <xref ref-type="bibr" rid="B28">2008</xref>). One frame of the video is shown in Figure <xref ref-type="fig" rid="F1">1</xref>. The system drives the robot eyes to saccade to different fixations chosen according to a probability distribution over the points in the scene. This probability distribution is derived from the saliency map of the scene. Fixations last 400&#x02009;ms. During each fixation, the system controls the vergence eye movement. After each fixation, the robot saccades to another fixation point in the environment.</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>The virtual environment in the iCub simulator. The two red rays indicate the eye gaze vectors. The inset at the lower right hand corner outlined in red shows a red-cyan anaglyph of the stereo images.</p></caption>
<graphic xlink:href="fnbot-11-00058-g001.tif"/>
</fig>
<p>The architecture of the integrated active vision system is illustrated in Figure <xref ref-type="fig" rid="F2">2</xref>. It consists of three main parts: the perceptual representation mechanism, the saccade control mechanism, and the vergence control mechanism.</p>
<fig position="float" id="F2">
<label>Figure 2</label>
<caption><p>The architecture of the model integrating vergence and saccadic eye movements. Red regions in the saliency map correspond to high values while blue regions correspond to low values. The red arrows identify the steps in generating the saccade command. The blue arrows identify the steps in generating the vergence command.</p></caption>
<graphic xlink:href="fnbot-11-00058-g002.tif"/>
</fig>
<p>The inputs to the system come from pairs of sub-windows from the left and right camera images. The input to the saccade control mechanism comes from the largest pair of sub-windows, which cover most of the images. This pair is down-sampled to generate a coarse scale representation. The input to the vergence control mechanism comes from three pairs of sub-windows: a pair of small fine scale sub-windows, a pair of medium-sized medium scale sub-windows, and a pair of large coarse scale sub-windows.</p>
<p>The perceptual representation mechanism encodes the binocular image inputs using a distributed representation learned using the generative adaptive subspace self-organizing map (GASSOM) algorithm (Chandrapala and Shi, <xref ref-type="bibr" rid="B7">2014</xref>). The GASSOM model is a statistical generative model for time-varying sensory input that combines both sparsity and slowness. The same perceptual representation is used to generate the input to both the saccade and the vergence control mechanisms.</p>
<p>The vergence control mechanism maps the GASSOM representation of the binocular inputs at all three scales to a set of discrete vergence actions. The vergence action is chosen according to a probability distribution computed using a neural network with a softmax output. The image input in the next iteration changes because of the vergence action.</p>
<p>The saccade control mechansim maps the GASSOM representation of the entire left and right eye images at the coarse scale to a fixation point by sampling the fixation point from a 2D probability distribution generated by a binocular saliency map. The saliency control policy also includes inhibition of return (IOR) to prevent the system from returning to a previously generated fixation too quickly. Saccades move the eyes to the selected fixation point while keeping the vergence angle the same.</p>
<p>Below, we describe in more detail the experimental setup and the three parts of the integrated active vision system. To avoid clutter in the notation, we do not indicate time explicitly. However, it should be understood that most quantities evolve over time either due to the agent&#x02019;s behavior in the environment, e.g., the inputs and actions, or due to learning, e.g., the network weights. Both behavior and learning progress simultaneously.</p>
</sec>
<sec id="S2-2">
<title>Experimental Setups</title>
<p>We trained and tested our system in two different simulation environments: the Tsukuba environment and the iCub environment. Most experiments are implemented in the Tsukuba environment. The results of experiments using the iCub environment are reported in Figures <xref ref-type="fig" rid="F5">5</xref> and <xref ref-type="fig" rid="F10">10</xref>.</p>
<p>The Tsukuba environment is based on the Tsukuba dataset (Martull et al., <xref ref-type="bibr" rid="B21">2012</xref>), which contains 1,800 photorealistic stereo image pairs created by rendering a virtual 3D laboratory environment. The environment contains objects located at various depths, resulting in a large range of disparities. Each image has size 640-by-480 pixels.</p>
<p>We simulated the effect of eye movements by extracting sub-windows from a single pair of stereo images, where the locations of the extracted sub-windows changes over time. In particular, we define fixation points in the left and right images. The left and right fixation points share the same vertical position, but are offset horizontally by an amount modeling the vergence angle. If the vergence angle is equal to the disparity in the original image, then the fixation points correspond to the same point in the virtual environment.</p>
<p>The stereo input to the saccade control mechanism is obtained by extracting the largest equally sized sub-windows from the left and right images such that the left and right fixation points are aligned in the two sub-windows. Let <italic>M</italic> and <italic>N</italic> denote the vertical and horizontal sizes of the images in pixels. If the horizontal offset between the left and right fixation points is <italic>d</italic>, then the upper left and lower right locations of the sub-windows are (<italic>d</italic>, 1) and (<italic>M</italic>, <italic>N</italic>) in the left image and (1,1) and (<italic>M</italic>-<italic>d</italic>&#x02009;&#x0002B;&#x02009;1, <italic>N</italic>) in the right image. These sub-windows are down-sampled by a factor of 4, resulting in a coarse scale representation. We applied bicubic interpolation to implement the image down-sampling.</p>
<p>The stereo input to the vergence control mechanism is obtained by extracting pairs of square sub-windows centered at the fixation points and then possibly down-sampling to generate pairs of 55-by-55 pixel images corresponding to three scales: coarse, medium, and fine. The coarse scale input is obtained by down-sampling 220-by-220 pixel sub-windows by a factor of 4. The medium scale input is obtained by down-sampling 110-by-110 pixel sub-window by a factor of 2. The fine scale input is obtained by extracting 55-by-55 pixel sub-windows without down-sampling.</p>
<p>Simulations consist of 10 frame periods of fixation separated by saccades. Assuming a frame rate of 25 frames per second, each fixation lasts for 400&#x02009;ms. During each fixation, the horizontal location of the right fixation point is adjusted according to the command given by the vergence control mechanism. If learned correctly, the vergence control mechanism adjusts the horizontal shift so that both sub-images are centered on the same point in the scene. We define the retinal disparity to be the difference between the shift and the original image disparity. When the retinal disparity is 0, the images in the left and right sub-windows are aligned. The binocular image pair is changed after every 30 fixations (300 frames), modeling a change in the scene.</p>
<p>Between fixations, the system uses the saccade mechanism to choose the left image fixation point. The right image fixation point is located at the same vertical location but is offset horizontally by a shift, which models the vergence angle. The initial vergence angle of each fixation is the same as the last vergence angle from the previous fixation.</p>
<p>The iCub environment runs in the iCub simulation platform (Tikhanoff et al., <xref ref-type="bibr" rid="B28">2008</xref>). The iCub is a humanoid robot with an active binocular vision system. The horizontal and vertical field of view are 64&#x000B0; and 50&#x000B0;, respectively. To simplify the simulation of the environment, we created the iCub world using some simple objects as shown in Figure <xref ref-type="fig" rid="F1">1</xref>. The virtual environment in front of the iCub robot contains a number of frontoparallel planar surfaces: a large background plane at a depth of 2&#x02009;m, and five smaller planes of size 0.6&#x02009;m&#x02009;&#x000D7;&#x02009;0.6&#x02009;m square placed at varying depths between the iCub and the background plane and at varying frontoparallel offsets. The planes are textured with images randomly chosen from the McGill natural image database (Olmos and Kingdom, <xref ref-type="bibr" rid="B23">2004</xref>). Binocular images pairs with size 320-by-240 pixels are generated by rendering this environment based on the positions and gaze angles of the two eyes.</p>
<p>In our simulations, the iCub remains stationary, except for changes in the gaze directions of its left and right eyes, which are controlled by three degrees of freedom: the version, tilt, and vergence angles. As in the Tsukuba environment, simulations consist of 10 frame periods of fixation separated by saccades. However, the left and right fixation points are both fixed at the center of the images. Saccades between fixations are implemented by changing the version and tilt angles, which are common to both eyes. During saccades, the vergence angle remains constant. During fixation, vergence eye movements are implemented by changing the vergence angle between the two eyes while keeping the version and tilt angles fixed. The iCub environment is more realistic than the Tsukuba environment, but simulations are more time consuming due to the rendering. Every 30 fixations (300 frames), the virtual environment is changed by choosing a new set of images from the database to apply to the planar surfaces and by randomizing the depths and positions of the smaller surfaces.</p>
</sec>
<sec id="S2-3">
<title>Perceptual Representation Mechanism</title>
<p>The vergence and saccade control mechanisms are based on the same perceptual representation mechanism applied to the left and right eye inputs. The left and right eye images are divided into 2D arrays of 10-by-10 pixel patches. For saccade generation, the patches are offset by a stride of one pixel. For the vergence control, the patches are offset by a stride of five pixels. At each scale <italic>s</italic> and for each pair <italic>j</italic> of corresponding patches in the left and right eye sub-windows, we concatenate the image intensities into a 200-dimensional binocular vector,
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="[" close="]"><mml:mrow><mml:mtable equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="center"><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>L</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="center"><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>R</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>200</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>
where <italic>s</italic>&#x02009;&#x02208;&#x02009;{C, M, F} (C, M, F stand for the three scales: coarse, medium, and fine scale, respectively), and the monocular vectors, <bold><italic>x</italic></bold><sub>L,</sub><italic><sub>s,j</sub></italic> and <bold><italic>x</italic></bold><sub>R,</sub><italic><sub>s,j</sub></italic>, contain the pixel intensities from the left and right image patches, which are normalized separately to have zero mean and unit variance.</p>
<p>The representation mechanism has three sets of <italic>N</italic>&#x02009;&#x0003D;&#x02009;324 binocular feature extractors, each set corresponding to one scale. The binocular stimulus <bold><italic>x</italic></bold><italic><sub>s</sub></italic><sub>,</sub><italic><sub>j</sub></italic> is encoded by the set of feature extractors at the associated scale <italic>s</italic>. The <italic>n-</italic>th feature extractor in scale <italic>s</italic> is defined by a two dimensional subspace of the input space, which is spanned by the basis defined by the columns of a matrix <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x02009;</mml:mtext><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>200</mml:mn><mml:mo class="MathClass-bin">&#x000D7;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> where <italic>n</italic>&#x02009;&#x02208;&#x02009;{1,&#x02009;&#x02026;&#x02009;, <italic>N</italic>}. Given an input patch vector <bold><italic>x</italic></bold><italic><sub>s</sub></italic><sub>,</sub><italic><sub>j</sub></italic>, the response of the <italic>n</italic>-th feature extractor is defined to be the squared length of the projection of <bold><italic>x</italic></bold><italic><sub>s</sub></italic><sub>,</sub><italic><sub>j</sub></italic> onto the subspace defined by &#x003A6;<italic><sub>s,n</sub></italic>:
<disp-formula id="E2"><label>(2)</label><mml:math id="M3"><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msup><mml:mrow><mml:mfenced separators="" open="&#x02225;" close="&#x02225;"><mml:mrow><mml:msubsup><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>
where the superscript <italic>T</italic> denotes the transpose operation. For each feature extractor &#x003A6;<italic><sub>s,n</sub></italic>, we define the response map to be the 2D set of responses of that feature detector to the 2D array of patches.</p>
<p>The operations involved in computing the response maps are similar to those used in computing the binocular energy model, which is commonly used to model the responses of orientation, scale, and disparity tuned binocular complex cells in the primary visual cortex (Ohzawa et al., <xref ref-type="bibr" rid="B22">1990</xref>). The subspace projection operation computes two weighted sums of the binocular image intensities. Thus, each basis vector (column of &#x003A6;<italic><sub>s,n</sub></italic>) is analogous to the linear spatial receptive field of a binocular simple cell. After learning, these basis vectors exhibit Gabor-like structures and are in approximate spatial phase quadrature (Chandrapala and Shi, <xref ref-type="bibr" rid="B7">2014</xref>). As in the binocular energy model, the response <italic>r<sub>n</sub></italic>(<bold><italic>x</italic></bold><italic><sub>s,j</sub></italic>) combines the squared magnitudes of two binocular simple cells. The magnitude of the response reflects the similarity between the binocular image patch and the binocular receptive fields.</p>
<p>The subspaces are initialized randomly and develop according to the update rules for the GASSOM algorithm described in Chandrapala and Shi (<xref ref-type="bibr" rid="B7">2014</xref>). The GASSOM exploits the concept of sparsity by using only one subspace to represent the input and captures the slowness by assuming that the subspace representing <bold><italic>x</italic></bold>(<italic>t</italic>) is more likely to be the same as the one that generated <bold><italic>x</italic></bold>(<italic>t</italic>&#x02009;&#x02212;&#x02009;1). The model parameters, e.g., the matrices &#x003A6;<italic><sub>s,n</sub></italic>, are learned in an unsupervised manner, by maximizing the likelihood of the observed data. The update to each subspace is calculated by
<disp-formula id="E3"><label>(3)</label><mml:math id="M4"><mml:mn>&#x00394;</mml:mn><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mstyle displaystyle='true'><mml:munder class="msub"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mo class="MathClass-op">&#x00303;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mfenced separators="" open="&#x02225;" close="&#x02225;"><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mo class="MathClass-op">&#x00302;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mfenced separators="" open="&#x02225;" close="&#x02225;"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:math></disp-formula>
where <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> determines the amount that subspace <inline-formula><mml:math id="M6"><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is updated towards the observation, <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mo class="MathClass-op">&#x00303;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x02009;</mml:mtext><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mo class="MathClass-op">&#x00302;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the difference between the input <bold><italic>x</italic></bold><italic><sub>j</sub></italic> and its projection onto subspace &#x003A6;<italic><sub>s,n</sub></italic>, where the projection is computed by <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mtext mathvariant="italic-bold">x</mml:mtext></mml:mrow><mml:mo class="MathClass-op">&#x00302;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Thus, each subspace at time <italic>t</italic> is updated by
<disp-formula id="E4"><label>(4)</label><mml:math id="M9"><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mn>&#x003BB;</mml:mn><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:math></disp-formula>
where &#x003BB;&#x02009;&#x0003E;&#x02009;0 is the learning rate.</p>
<p>Under the AEC framework, both the GASSOM model parameters and the parameters of the vergence control policy (described below) are learned simultaneously.</p>
</sec>
<sec id="S2-4">
<title>Vergence Control Mechanism</title>
<p>The vergence control mechanism maps the visual input to a probability distribution over a discrete set of 11 possible vergence actions <italic>A</italic><sub>verg</sub>. For the Tsukuba environment, the 11 discrete vergence actions are <italic>A</italic><sub>verg</sub>&#x02009;&#x0003D;&#x02009;{&#x02212;16, &#x02212;8, &#x02212;4, &#x02212;2, &#x02212;1, 0, 1, 2, 4, 8, 16} pixels, which modify the shift between the centers of the left and right sub-windows. For the iCub environment, the vergence actions are <italic>A</italic><sub>verg</sub>&#x02009;&#x0003D;&#x02009;{&#x02212;3.2, &#x02212;1.6, &#x02212;0.8, &#x02212;0.4, &#x02212;0.2, 0, 0.2, 0.4, 0.8, 1.6, 3.2} degrees.</p>
<p>The vergence control policy is implemented by a two layer neural network. The input to the network is a 3<italic>N</italic>-dimensional vector:
<disp-formula id="E5"><label>(5)</label><mml:math id="M10"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>r</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="[" close="]"><mml:mrow><mml:mtable equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="center"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>r</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="center"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>r</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>M</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="center"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>r</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced></mml:math></disp-formula>
where for each <italic>s</italic>&#x02009;&#x02208;&#x02009;{C, M, F}, <inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x02009;</mml:mtext><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is obtained by spatially pooling the response maps of the feature detectors at scale <italic>s</italic>:
<disp-formula id="E6"><label>(6)</label><mml:math id="M12"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="[" close="]"><mml:mrow><mml:mtable equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="center"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="center"><mml:mo class="MathClass-op">&#x022EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="center"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced></mml:math></disp-formula>
where <italic>P</italic>&#x02009;&#x0003D;&#x02009;100 is the number of the patches.</p>
<p>The output layer of the network contains 11 neurons, each corresponding to a possible vergence command. The vector of activations to the output neurons, <inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, is computed as:
<disp-formula id="E7"><label>(7)</label><mml:math id="M14"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msup><mml:mrow><mml:mn mathvariant="bold">&#x003B8;</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub></mml:math></disp-formula>
where <inline-formula><mml:math id="M15"><mml:msup><mml:mrow><mml:mn mathvariant="bold">&#x003B8;</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>972</mml:mn><mml:mo class="MathClass-bin">&#x000D7;</mml:mo><mml:mn>11</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is a matrix of synaptic weights.</p>
<p>The vector of probabilities for selecting the different vergence actions, <inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>&#x003C0;</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, is calculated by applying a softmax operation to the activation vector <bold>z</bold><sub>verg</sub>:
<disp-formula id="E8"><label>(8)</label><mml:math id="M17"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>&#x003C0;</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi mathvariant="normal">softmax</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02215;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003B2;</mml:mn></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:math></disp-formula>
where the softmax function <bold>y</bold>&#x02009;&#x0003D;&#x02009;softmax(<bold>z</bold>) is defined component-wise by:
<disp-formula id="E9"><label>(9)</label><mml:math id="M18"><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">exp</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msubsup><mml:mi mathvariant="normal">exp</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:math></disp-formula>
for <italic>i</italic> &#x02208;{1,&#x02009;&#x02026;&#x02009;, 11} The temperature parameter, &#x003B2;<sub>verg</sub>, balances exploration and exploitation during reinforcement learning. In the following experiments, &#x003B2;<sub>verg</sub> is set to 1.</p>
<p>The neural network weights develop according to the natural actor-critic reinforcement learning algorithm (Bhatnagar et al., <xref ref-type="bibr" rid="B3">2009</xref>). In our system, the reinforcement learner seeks a vergence control policy that minimizes the error in the perceptual representation of the sensory input, or equivalently maximizes the fidelity of the perceptual representation. We define the instantaneous reward to be the negative of the average squared reconstruction error across the three scales, which is defined by
<disp-formula id="E10"><label>(10)</label><mml:math id="M19"><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mtext>verg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mtext>avg</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munder class="msub"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>
where <italic>E<sub>s</sub></italic> is the mean squared reconstruction error at scale <italic>s</italic> averaged across all patches. The reconstruction error of each patch is defined as the squared length of the residual between the input vector <bold><italic>x</italic></bold><italic><sub>s,j</sub></italic> and its projection onto the best-fitting subspace:
<disp-formula id="E11"><label>(11)</label><mml:math id="M20"><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mfenced separators="" open="&#x02225;" close="&#x02225;"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msubsup><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>
where <italic>m<sub>s,j</sub></italic> is the index of the best-fitting subspace for <bold><italic>x</italic></bold><italic><sub>s,j</sub></italic>,
<disp-formula id="E12"><label>(12)</label><mml:math id="M21"><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:munder accentunder="true"><mml:mrow><mml:mtext>arg max</mml:mtext></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munder><mml:mtext>&#x02009;</mml:mtext><mml:msup><mml:mrow><mml:mfenced separators="" open="&#x02225;" close="&#x02225;"><mml:mrow><mml:msubsup><mml:mrow><mml:mn>&#x003A6;</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<p>The weight matrices of the value and policy networks are updated during fixation, but not across saccades.</p>
</sec>
<sec id="S2-5">
<title>Saccade Control Mechanism</title>
<sec id="S2-5-1">
<title>Binocular Saliency Model</title>
<p>The binocular saliency map is generated using binocular attention based on information maximization (BAIM), which we propose as a binocular extension of the AIM model of Bruce and Tsotsos (<xref ref-type="bibr" rid="B6">2009</xref>). The BAIM architecture is illustrated in Figure <xref ref-type="fig" rid="F3">3</xref>. In essence, we replace the monocular basis functions learned by ICA in the AIM model with the binocular basis functions learned by GASSOM. Whereas the monocular basis functions encode only texture information, the binocular basis functions used here jointly encode depth and texture information. Since our experiment use only gray scale images, we do not jointly encode color as in the experiments by Bruce and Tsotsos (<xref ref-type="bibr" rid="B6">2009</xref>). However, the extension to include color is straightforward, involving only an expansion of the size of the input vector.</p>
<fig position="float" id="F3">
<label>Figure 3</label>
<caption><p>The architecture of the binocular attention based on information maximization (BAIM). The left and right parts of image patches and basis vectors are shown as 10-by-10 pixel images and aligned vertically. Red regions in the response maps and saliency maps correspond to high values while blue regions correspond to low values. The red squares in the response maps represent the local area where response histograms are generated.</p></caption>
<graphic xlink:href="fnbot-11-00058-g003.tif"/>
</fig>
<p>The computations to obtain the salience map are performed at the coarse scale. We assign saliency values to 10-by-10 pixel patches, which are extracted with a stride of one. Given coarse scale images with size <italic>M</italic><sub>1</sub>&#x02009;&#x000D7;&#x02009;<italic>M</italic><sub>2</sub>, we obtain a 2D array of <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mtext mathvariant="italic">sal</mml:mtext></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>9</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x000D7;</mml:mo><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>9</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:math></inline-formula> binocular patches. To obtain a saliency map at the original resolution, we use zero padding to increase the size of the coarse scale map to <italic>M</italic><sub>1</sub>&#x02009;&#x000D7;&#x02009;<italic>M</italic><sub>2</sub> pixels and then upsample by a factor of 4.</p>
<p>The saliency value for the <italic>j-</italic>th coarse scale binocular image patch, <italic>S</italic>(<bold><italic>x</italic></bold><italic><sub>C,j</sub></italic>), is a measure of how informative or unlikely the responses of the GASSOM feature detectors are in the context of the responses from the other patches. More specifically, it is the sum of saliency values computed for individual feature extractors in the GASSOM representation, <italic>S<sub>n</sub></italic>(<bold><italic>x</italic></bold><italic><sub>C,j</sub></italic>):
<disp-formula id="E13"><label>(13)</label><mml:math id="M23"><mml:mi>S</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msubsup><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<p>The saliency of each feature extractor is the Shannon self-information of the response:
<disp-formula id="E14"><label>(14)</label><mml:math id="M24"><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mtext>ln&#x02009;</mml:mtext><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="[" close="]"><mml:mrow><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mo mathvariant="bold-italic">x</mml:mo></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:math></disp-formula>
where <italic>p<sub>n</sub></italic>[&#x022C5;] is the probability distribution of the responses of the <italic>n</italic>-th feature extractor at the coarse scale, which we estimate empirically using a histogram. Each response map is normalized to the range 0&#x02013;1. The histogram of each response map is generated using <inline-formula><mml:math id="M25"><mml:mi>K</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mtext>sal</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msqrt></mml:math></inline-formula> equal width bins:
<disp-formula id="E15"><label>(15)</label><mml:math id="M26"><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="[" close="]"><mml:mrow><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>&#x003B1;</mml:mn><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mn>1</mml:mn><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>&#x003B1;</mml:mn></mml:mrow></mml:mfenced><mml:mstyle displaystyle='true'><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow><mml:mrow><mml:mfenced separators="" open="[" close=")"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:mfrac><mml:mo class="MathClass-punc">,</mml:mo><mml:mfrac><mml:mrow><mml:mi>k</mml:mi><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:math></disp-formula>
where
<disp-formula id="E16"><label>(16)</label><mml:math id="M27"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow><mml:mrow><mml:mfenced separators="" open="[" close=")"><mml:mrow><mml:mi>a</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>1</mml:mn></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mtext> if</mml:mtext><mml:mi>x</mml:mi><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:mfenced separators="" open="[" close=")"><mml:mrow><mml:mi>a</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:mfenced></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>0</mml:mn></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mtext> if</mml:mtext><mml:mi>x</mml:mi><mml:mo class="MathClass-rel">&#x02209;</mml:mo><mml:mfenced separators="" open="[" close=")"><mml:mrow><mml:mi>a</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:mfenced></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced></mml:math></disp-formula>
is an indicator function with <italic>a</italic> and <italic>b</italic> as free parameters. The parameter &#x003B1;&#x02009;&#x0003D;&#x02009;10<sup>&#x02212;6</sup> is a small number that guarantees that the response probabilities are non-zero. The coefficients
<disp-formula id="E17"><label>(17)</label><mml:math id="M28"><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mtext>&#x02009;</mml:mtext><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mtext mathvariant="italic">sal</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mtext mathvariant="italic">sal</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow><mml:mrow><mml:mfenced separators="" open="[" close=")"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:mfrac><mml:mo class="MathClass-punc">,</mml:mo><mml:mfrac><mml:mrow><mml:mi>k</mml:mi><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced></mml:mrow></mml:msub><mml:mspace width="0.3em" class="thinspace"/><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mtext>&#x02009;for&#x02009;</mml:mtext><mml:mi>k</mml:mi><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:mfenced separators="" open="{" close="}"><mml:mrow><mml:mn>0</mml:mn><mml:mo class="MathClass-punc">,</mml:mo><mml:mn>1</mml:mn><mml:mo class="MathClass-punc">,</mml:mo><mml:mn>2</mml:mn><mml:mo class="MathClass-punc">,</mml:mo><mml:mo class="MathClass-punc">.</mml:mo><mml:mo class="MathClass-punc">.</mml:mo><mml:mo class="MathClass-punc">.</mml:mo><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>K</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfenced></mml:math></disp-formula>
are empirical estimates of the probability that the response falls into the <italic>k</italic>-th bin computed over <inline-formula><mml:math id="M29"><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mtext mathvariant="italic">sal</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> patches.</p>
<p>We considered both global binocular attention based on information maximization (GBAIM) and local binocular attention based on information maximization (LBAIM) versions of the saliency map, which differed according to patches used to estimate the coefficients <italic>h<sub>n</sub></italic>(<italic>k</italic>) in Eq. <xref ref-type="disp-formula" rid="E17">17</xref>. In the GBAIM model, the coefficients were computed by summing over all coarse scale patches. In the LBAIM model, the sum was over only the coefficients from a 31&#x02009;&#x000D7;&#x02009;31 array of patches centered around the current fixation point. The LBAIM model tends to favor patches where the GASSOM responses are more unlike those in the local neighborhood of the current fixation point.</p>
<p>To speed up computations, we use only a random subset of the GASSOM feature extractors to compute the sum in Eq. <xref ref-type="disp-formula" rid="E13">13</xref>. To determine the size of the subset, we computed the correlation coefficients (CCs) to measure the similarity between the saliency maps generated by random subsets of feature extractors and by all feature extractors. The CC between two saliency maps <italic>S</italic><sub>1</sub> and <italic>S</italic><sub>2</sub> is defined as:
<disp-formula id="E18"><label>(18)</label><mml:math id="M30"><mml:mtext>CC</mml:mtext><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003BC;</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003BC;</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003BC;</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold-italic">x</mml:mtext></mml:mrow><mml:mrow><mml:mtext>C</mml:mtext><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mn>&#x003BC;</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:math></disp-formula>
where &#x003BC;<sub>1</sub> and &#x003BC;<sub>2</sub> are the mean saliencies in the maps. Figure <xref ref-type="fig" rid="F4">4</xref> plots the CCs averaged over salience maps computed from 1,800 pairs of binocular images from the Tsukuba dataset. For each saliency map, the subsets of a certain number of feature extractors are chosen randomly from all 324 feature extractors. Using only 25 feature, extractors generates a BAIM saliency map that is very similar to the one using all 324 features. Thus, in our experiments, we sum over 25 randomly selected feature extractors.</p>
<fig position="float" id="F4">
<label>Figure 4</label>
<caption><p>The average correlation coefficient (CC) values between the binocular attention based on information maximization (BAIM) saliency maps generated by all feature extractors and the saliency maps generated by different subsets of feature extractors. A logarithmic scale is used on the <italic>x</italic>-axis. Error bars represent 95% confidence intervals for the mean values of the CCs.</p></caption>
<graphic xlink:href="fnbot-11-00058-g004.tif"/>
</fig>
</sec>
<sec id="S2-5-2">
<title>Inhibition of Return</title>
<p>Given the current fixation, a saccade target for the next fixation is generated by combining the saliency map at the full image resolution with a simple IOR mechanism (Dorris et al., <xref ref-type="bibr" rid="B10">2002</xref>), which prevents the system from saccading to recently visited image locations.</p>
<p>Defining the full resolution saliency map by <italic>S</italic>(<italic>j</italic>) where <italic>j</italic> indexes the patch, we choose the next fixation point by sampling from the probability distribution
<disp-formula id="E19"><label>(19)</label><mml:math id="M31"><mml:mi>p</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mi>S</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:mtext>IOR</mml:mtext><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>S</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:mtext>IOR</mml:mtext><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:math></disp-formula>
where IOR(&#x022C5;) is a mask that suppresses recently visited image locations.</p>
<p>Posner and Cohen (<xref ref-type="bibr" rid="B25">1984</xref>) indicate that the currently attended region is inhibited for approximately 500&#x02013;1,000&#x02009;ms. Since fixations last for 10 frames, and assuming a frame rate of 25 frames per second, we prevent the system from visiting the last two fixation points, resulting in a 800&#x02009;ms long IOR window. Assuming the indices of most recent and second most recent fixation locations are <italic>j</italic><sub>1</sub> and <italic>j</italic><sub>2</sub>, we set
<disp-formula id="E20"><label>(20)</label><mml:math id="M32"><mml:mtext>IOR</mml:mtext><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>f</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:msub><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo class="MathClass-punc">,</mml:mo><mml:msubsup><mml:mrow><mml:mn>&#x003C3;</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfenced><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:mi>f</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:msub><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo class="MathClass-punc">,</mml:mo><mml:msubsup><mml:mrow><mml:mn>&#x003C3;</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfenced></mml:math></disp-formula>
where
<disp-formula id="E21"><label>(21)</label><mml:math id="M33"><mml:mi>f</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>j</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>k</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:msup><mml:mrow><mml:mn>&#x003C3;</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mn>1</mml:mn><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mi mathvariant="normal">exp</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mfenced separators="" open="&#x02225;" close="&#x02225;"><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">p</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">p</mml:mtext></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mn>&#x003C3;</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula>
where <bold>p</bold><italic><sub>j</sub></italic> is the 2D image location of patch <italic>j</italic>, &#x003C3;<sub>1</sub>&#x02009;&#x0003D;&#x02009;20 pixels and &#x003C3;<sub>2</sub>&#x02009;&#x0003D;&#x02009;10 pixels.</p>
</sec>
</sec>
</sec>
<sec id="S3">
<title>Results</title>
<sec id="S3-6">
<title>BAIM-Driven Saccades Accelerate Vergence Learning</title>
<p>We compared the rate at which vergence control policies emerged and the quality of the final polices under different saccade control policies including a random policy where <italic>p</italic>(<italic>m, n</italic>) in Eq. <xref ref-type="disp-formula" rid="E19">19</xref> was uniform over all image locations and policies where <italic>S</italic>(<italic>m, n</italic>) in Eq. <xref ref-type="disp-formula" rid="E19">19</xref> was computed according to the saliency model of Itti et al. (<xref ref-type="bibr" rid="B15">1998</xref>), the AIM model (Bruce and Tsotsos, <xref ref-type="bibr" rid="B6">2009</xref>), and the GBAIM and LBAIM models proposed here.<xref ref-type="fn" rid="fn2"><sup>2</sup></xref></p>
<p>Figure <xref ref-type="fig" rid="F5">5</xref> shows the evolution of the root mean squared error (RMSE) between the learned vergence control policies during training and the ideal policy that zeros out the input disparity in both the Tsukuba and the iCub environments. For the Tsukuba environment, we ran three training trials for each saccade control policy. We set the same learning rate for the vergence control learner to make all the saccade methods comparable. Each trial used binocular inputs generated by disjoint sets of 120 randomly chosen stereo images, but we used the same image sets to train the different saccade policies. We sampled the vergence policies at 20 equally spaced checkpoints during training. At each checkpoint, we presented the policy with inputs with initial disparities ranging from &#x02212;20 to &#x0002B;20 pixels and let the vergence policy run for 10 iterations. The RMSE in pixels was computed as the square root of the mean squared retinal disparity after 10 iterations averaged over all initial disparities and 100 inputs per disparity. The images used to characterize the policies at all testing points and for all saccade policies were identical and disjoint from those used in training. For the iCub environment, we also ran three training trials for each saccade control policy. Each trial was conducted in a different randomly generated environment with disjoint sets of image textures mapped on to the surfaces. The same environments were used to train under different saccade policies. During testing, the iCub was presented with frontoparallel surfaces at depths ranging from 0.5 to 2.0&#x02009;m and textures disjoint from those used in training and allowed to verge for 10 iterations starting from initial vergence angles ranging from 0&#x000B0; to 10&#x000B0;. The RMSE of the final vergence angle was averaged over all depths and all initial vergence angles.</p>
<fig position="float" id="F5">
<label>Figure 5</label>
<caption><p>The evolution of the root mean squared error (RMSE) of the vergence control policy for different saccade control policies over training in <bold>(A)</bold> the Tsukuba environment and <bold>(B)</bold> the iCub environment. Error bars indicate the SD computed over three training runs.</p></caption>
<graphic xlink:href="fnbot-11-00058-g005.tif"/>
</fig>
<p>As shown in Figure <xref ref-type="fig" rid="F5">5</xref>, the random saccade policy performs the worst, resulting in the largest vergence control policy RMSE at the end of training. The two BAIM-driven saccade policies result in the best final performance. Although there is little difference between the two final policies, the LBAIM model exhibits faster vergence learning, with a faster decrease in the RMSE. The two monocular saccade models result in vergence control policies with RMSE values lying between those learned under random and binocularly driven saccades, with the AIM model exhibiting slightly faster learning and lower final RMSE.</p>
<p>Figure <xref ref-type="fig" rid="F6">6</xref> shows visualizations of the final policies learned in the Tsukuba environment. It is clear that the final vergence control policies learned using the BAIM-driven saccades are closer to the ground truth. The &#x0201C;blurred&#x0201D; images for the policy learned using random and monocular saliency driven saccades indicate that the policies are less reliable in zeroing out the retinal disparity.</p>
<fig position="float" id="F6">
<label>Figure 6</label>
<caption><p>Visualizations of the final vergence policies after training. Each policy is presented as an image. The horizontal axis indicates the initial disparity and the vertical axis indicates the change in vergence after 10 iterations of the policy. The intensity of each pixel corresponds to the probability of the change in vergence given the initial disparity, i.e., the entries in each column sum to 1. For the ground truth policy, the change in vergence is always the negative of the initial disparity.</p></caption>
<graphic xlink:href="fnbot-11-00058-g006.tif"/>
</fig>
</sec>
<sec id="S3-7">
<title>BAIM-Driven Saccades Select Image Regions with Higher Entropy</title>
<p>To understand better how the different saccade policies lead to vergence control policies with different performance, we examined the entropy of the image regions around the fixation points chosen by the different saccade policies. The entropy is defined as:
<disp-formula id="E22"><label>(22)</label><mml:math id="M34"><mml:mi>E</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mstyle displaystyle='true'><mml:munder class="msub"><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="normal">log</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:math></disp-formula>
where <italic>p<sub>i</sub></italic> is the probability of the pixel intensity value <italic>i</italic>, which is estimated from the histogram of pixel intensity values. The entropy is a measure of the spread of gray values and is one measure of the information content. Image regions whose pixels all have the same intensity have zero entropy. Textured regions will have more variability in gray levels, and therefore a higher entropy. Intuitively, it will be harder to learn vergence from regions without texture (with lower entropy).</p>
<p>Figure <xref ref-type="fig" rid="F7">7</xref> shows the median entropies computed over the 55-by-55 pixel fine scale sub-windows at the fixation points selected by the different saccade control methods. The statistics were collected over 1,800 sub-window pairs. For each of the 180 stereo image pairs obtained by taking every 10th frame from the 1,800 stereo image pairs in the Tsukuba dataset, we ran the saccade/vergence policy learned after 100,000 iterations (the first checkpoint in Figure <xref ref-type="fig" rid="F5">5</xref>) for 10 fixations (100 frames), and averaged the entropies of the left and right sub-windows. Our testing results for the learned systems at other checkpoints were similar.</p>
<fig position="float" id="F7">
<label>Figure 7</label>
<caption><p><bold>(A)</bold> The entropy histogram of the fine scale sub-windows selected by local binocular attention based on information maximization (LBAIM). The red line indicates the median entropy. <bold>(B)</bold> The median entropy of the patches at the fixation points selected by different saccade control methods. The error bars represent the SEM.</p></caption>
<graphic xlink:href="fnbot-11-00058-g007.tif"/>
</fig>
</sec>
<sec id="S3-8">
<title>BAIM-Driven Saccades Lead to Improved Encoding of Small Disparities</title>
<p>By selecting different fixation points, the different saccade policies expose the perceptual representation to binocular patches with different statistics. Since the feature extractors evolve to maximize the likelihood of the observed data, these differences in the input statistics will be reflected as differences in the learned feature extractors.</p>
<p>To study these differences, we learned feature detectors in the Tsukuba environment using sub-windows centered at fixation points chosen by the different saccade control policies. Since differences in the vergence control policies will affect the input disparity statistics, we chose the vergence angles so that the retinal disparities <italic>d</italic> between the fixation points followed a discrete truncated Laplacian distribution between &#x02212;40 and &#x0002B;40 pixels:
<disp-formula id="E23"><label>(23)</label><mml:math id="M35"><mml:mi>P</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>M</mml:mi><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x02009;</mml:mtext></mml:mrow><mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mfenced separators="" open="|" close="|"><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:msup><mml:mi>/</mml:mi><mml:msub><mml:mrow></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo class="MathClass-punc">,</mml:mo><mml:mspace width="1em" class="quad"/><mml:mi>d</mml:mi><mml:mo class="MathClass-rel">&#x02208;</mml:mo><mml:mfenced separators="" open="{" close="}"><mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>40</mml:mn><mml:mo class="MathClass-punc">,</mml:mo><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>39</mml:mn><mml:mo class="MathClass-punc">,</mml:mo><mml:mo class="MathClass-punc">.</mml:mo><mml:mo class="MathClass-punc">.</mml:mo><mml:mo class="MathClass-punc">.</mml:mo><mml:mo class="MathClass-punc">,</mml:mo><mml:mn>40</mml:mn></mml:mrow></mml:mfenced></mml:math></disp-formula>
where <italic>M</italic> is a normalization factor that ensures that <italic>P</italic>(<italic>d</italic>) sums to one. The parameter <italic>D</italic> controls the spread of the input disparities. This enabled us to isolate the effects of the saccade policies on the perceptual representations.</p>
<p>Figure <xref ref-type="fig" rid="F8">8</xref> shows the average reconstruction error <italic>E</italic><sub>avg</sub> of the perceptual representations after training under different saccade control policies and assuming two different disparity statistics. The values of <italic>E</italic><sub>avg</sub> are plotted as a function of the retinal disparity at the fixation points. They are computed by averaging the values of <italic>E</italic><sub>avg</sub> computed according to Eq. <xref ref-type="disp-formula" rid="E10">10</xref> over sub-windows taken at 1,000 fixation points from 100 images in the Tsukuba dataset (10 fixations per image). Images used in testing were disjoint from those used in training.</p>
<fig position="float" id="F8">
<label>Figure 8</label>
<caption><p>The average reconstruction error for the perceptual representations learned under different saccade control policies plotted as a function of the retinal disparity between the sub-windows. The retinal disparity at the fixation points followed truncated Laplacian distributions with parameters: <bold>(A)</bold> <italic>D</italic>&#x02009;&#x0003D;&#x02009;50 and <bold>(B)</bold> <italic>D</italic>&#x02009;&#x0003D;&#x02009;5.</p></caption>
<graphic xlink:href="fnbot-11-00058-g008.tif"/>
</fig>
<p>In general, the curves have a characteristic &#x0201C;V&#x0201D; shape, being symmetric around and achieving their minima at zero disparity. This suggests that the perceptual representation is adapted to best represent binocular stimuli with zero disparity and that there are more feature extractors tuned to zero disparity. There is no preference for positive or negative disparity stimuli. These observations are consistent with the distribution of retinal disparities in the input, which is peaked at and symmetric around 0. The &#x0201C;V&#x0201D; shape indicates small reconstruction error at 0 and large average reconstruction error at large disparities. The &#x0201C;V&#x0201D; shape is more pronounced the more tightly the input disparity statistics are clustered around zero disparity, which is obtained by choosing a smaller value of <italic>D</italic> in Eq. <xref ref-type="disp-formula" rid="E23">23</xref>.</p>
<p>Figure <xref ref-type="fig" rid="F9">9</xref>A shows the reconstruction error curves of the perceptual representations learned after joint learning under different saccade control policies. The differences are much more pronounced, with curves corresponding to the BAIM-driven saccade policies exhibiting much more pronounced &#x0201C;V&#x0201D; shapes. To quantify how pronounced those &#x0201C;V&#x0201D; shapes are, we fit the reconstruction error curves with the function
<disp-formula id="E24"><label>(24)</label><mml:math id="M36"><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mn>&#x003BC;</mml:mn><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>b</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>c</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mi>a</mml:mi><mml:mo class="MathClass-bin">&#x022C5;</mml:mo><mml:mspace width="0.5em" class="nbsp" /><mml:mtext>exp</mml:mtext><mml:mspace width="0.5em" class="nbsp" /><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mfenced separators="" open="|" close="|"><mml:mrow><mml:mi>d</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>&#x003BC;</mml:mn></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mfenced></mml:math></disp-formula>
where <italic>c</italic> is a vertical offset, &#x003BC; sets the location of the minimum, and <italic>a</italic> and <italic>b</italic> control the slope and depth of the &#x0201C;V,&#x0201D; respectively. Figure <xref ref-type="fig" rid="F9">9</xref>B shows one example of a fit. We measure the sharpness of the &#x0201C;V&#x0201D; by the ratio <italic>a</italic>/<italic>b</italic>, which is the absolute value of the slope at &#x003BC;. Figure <xref ref-type="fig" rid="F9">9</xref>C shows the evolution of the slope during training under the different saccade control policies. In all cases, the slope increases over time, indicating that the basis functions evolve so that they provide a better encoding of stimuli with zero disparity. However, the rate and magnitude at which the slope increases are largest under the BAIM-driven saccade control policies.</p>
<fig position="float" id="F9">
<label>Figure 9</label>
<caption><p><bold>(A)</bold> The average reconstruction error as a function of the retinal disparity. <bold>(B)</bold> The experimentally estimated reconstruction errors of local binocular attention based on information maximization (LBAIM) and the corresponding curve fit. <bold>(C)</bold> The change of the slope <italic>a</italic>/<italic>b</italic> over the training process. Larger slopes indicate that the perceptual encoding exhibits a stronger preference for inputs with zero disparity. Error bars represent the SD computed over three training runs.</p></caption>
<graphic xlink:href="fnbot-11-00058-g009.tif"/>
</fig>
</sec>
<sec id="S3-9">
<title>LBAIM-Driven Saccades Target Locations with Different Disparities</title>
<p>Figure <xref ref-type="fig" rid="F5">5</xref> indicates that the RMSE of the final vergence control policies learned under the GBAIM and LBAIM saccade policies are similar, but that the vergence control policy emerges faster under LBAIM. The similarities between the reconstruction error curves and slope trajectories of the perceptual representations learned under GBAIM and LBAIM in Figures <xref ref-type="fig" rid="F8">8</xref> and <xref ref-type="fig" rid="F9">9</xref> suggest that the faster learning is not due to differences in the perceptual representation. Rather, we suggest the learning is faster because LBAIM presents more challenging vergence control stimuli to the system.</p>
<p>As a concrete example, Figure <xref ref-type="fig" rid="F10">10</xref> shows two examples of saliency maps computed by LBAIM in the iCub environment. The maps were computed in the same environment, which had two objects with the same textures in front of the iCub: one on the left at a closer distance and one on the right at a farther distance and partially occluded by the closer object. Figures <xref ref-type="fig" rid="F10">10</xref>A,B show example images from the left and right eye cameras. The two saliency maps were computed assuming the iCub was fixating either on the closer or the farther object. Comparing their intensities, we observe that points on the closer object become more salient when the iCub is fixating on the farther object and <italic>vice versa</italic>.</p>
<fig position="float" id="F10">
<label>Figure 10</label>
<caption><p><bold>(A,B)</bold> Left <bold>(A)</bold> and right <bold>(B)</bold> eye images obtained from the iCub simulator when the robot was viewing two planar objects: one on the left which is closer and the other on the right which is farther. <bold>(C,D)</bold> Examples of the local binocular attention based on information maximization (LBAIM) saliency map computed with the iCub fixating on the closer <bold>(C)</bold> and farther <bold>(D)</bold> objects. Bright regions correspond to high saliency. The red points indicate the fixation point. The red squares indicate the local area over which the empirical response histogram is computed.</p></caption>
<graphic xlink:href="fnbot-11-00058-g010.tif"/>
</fig>
<p>For a more comprehensive and quantitative comparison, we estimated the expected absolute disparity difference between current and next fixations for different saliency models according to
<disp-formula id="E25"><label>(25)</label><mml:math id="M37"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mn>&#x00394;</mml:mn><mml:mi>D</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfenced separators="" open="|" close="|"><mml:mrow><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mi>p</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo class="MathClass-rel">&#x0007C;</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:math></disp-formula>
where <italic>i</italic> indexes all possible next fixations, <italic>j</italic> indexes the current fixation point, <italic>D<sub>i</sub></italic> is the disparity at <italic>i</italic>, and <italic>P</italic> is the number of current fixation points considered. The term
<disp-formula id="E26"><label>(26)</label><mml:math id="M38"><mml:mi>p</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>i</mml:mi><mml:mo class="MathClass-rel">&#x0007C;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:mfenced><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mi>S</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>i</mml:mi><mml:mo class="MathClass-rel">&#x0007C;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mo class="MathClass-op">&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>S</mml:mi><mml:mfenced separators="" open="(" close=")"><mml:mrow><mml:mi>k</mml:mi><mml:mo class="MathClass-rel">&#x0007C;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:math></disp-formula>
is the probability of choosing the next fixation point <italic>i</italic> given the current fixation point <italic>j</italic>. It is similar to Eq. <xref ref-type="disp-formula" rid="E19">19</xref> except that we do not include the IOR and that we make explicit the dependency of the saliency map <italic>S</italic>(<italic>i</italic>&#x0007C;<italic>j</italic>) on the current fixation. For all of the saliency mechanisms except LBAIM, the saliency map does not change with the current fixation location.</p>
<p>The LBAIM saliency model shows a clear preference toward selecting targets with disparities that are different from that at the current fixation point. Figure <xref ref-type="fig" rid="F11">11</xref> shows the expected absolute disparity difference under different saccade policies normalized by the expected absolute disparity difference under the random saccade policy. The expected differences were estimated from data in the Tsukuba data set, for which ground truth disparity data are available. We selected 200 binocular images randomly from the 1,800 image frames in the video of the Tsukuba dataset. For each binocular image, we computed the saliency maps at 10 fixation points (i.e., <italic>P</italic>&#x02009;&#x0003D;&#x02009;2,000).</p>
<fig position="float" id="F11">
<label>Figure 11</label>
<caption><p>The estimated expected absolute disparity difference, <inline-formula><mml:math id="M39"><mml:mrow><mml:mover accent='true'><mml:mrow><mml:mo>&#x0394;</mml:mo><mml:mi>D</mml:mi></mml:mrow><mml:mo stretchy='true'>&#x00AF;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, for saccades generated by the different saliency models normalized by the expected difference estimated using random saccades. The error bar represents the SD.</p></caption>
<graphic xlink:href="fnbot-11-00058-g011.tif"/>
</fig>
</sec>
<sec id="S3-10">
<title>Change of Reconstruction Error within One Fixation</title>
<p>The average reconstruction error of the perceptual representation, <italic>E</italic><sub>avg</sub>, plays a critical role in this system. Both the learning of the perceptual representation and the learning of the vergence control policy seek to minimize <italic>E</italic><sub>avg</sub>. Large decreases in <italic>E</italic><sub>avg</sub> during fixation suggest that the vergence control reinforcement learning is being exposed to &#x0201C;challenging&#x0201D; situations where there is potential for large changes in the reward.</p>
<p>Figure <xref ref-type="fig" rid="F12">12</xref>A shows how the average decrease in <italic>E</italic><sub>avg</sub> during fixation evolves over training under the different saccade control policies. We define the normalized decrease across one fixation (10 frames) as the difference between the values of <italic>E</italic><sub>avg</sub> at the start and end of a fixation normalized by the value of <italic>E</italic><sub>avg</sub> at the start of fixation, and compute the running average across 3,000 fixations. The normalized decreases in <italic>E</italic><sub>avg</sub> under the BAIM-driven saccade policies exhibit much larger and faster increases across training than the monocular saliency and random policies.</p>
<fig position="float" id="F12">
<label>Figure 12</label>
<caption><p><bold>(A)</bold> The running average of the normalized decrease in the average reconstruction error during fixation increases during training. <bold>(B)</bold> To isolate the effect of the saccade policy from differences in the perceptual representation and the vergence policy, we plot the normalized decrease in average reconstruction error when the same perceptual representation and vergence policy were used.</p></caption>
<graphic xlink:href="fnbot-11-00058-g012.tif"/>
</fig>
<p>We isolated the effect of the choice of saccade target on the average normalized decrease in <italic>E</italic><sub>avg</sub> by using the same perceptual representation and vergence policies during fixation, but choosing the initial fixation points according to the different saccade control policies. In the results reported here, we used the perceptual representation and vergence policies learned by LBAIM, but other choices gave similar results. Figure <xref ref-type="fig" rid="F12">12</xref>B shows that the ordering of the policies is preserved, but the relative differences between the curves for different saccade policies are smaller.</p>
</sec>
</sec>
<sec id="S4" sec-type="discussion">
<title>Discussion</title>
<p>We have described an integrated active vision system that combines vergence control learning with saccade control based on two novel binocular saliency models: LBAIM and GBAIM. These models are based on binocular feature extractors that simultaneously encode both texture and depth information and that are computed in a way similar to the binocular energy model, a common computational model for disparity and orientation selective cells in the primary visual cortex. The algorithm assigns high saliency to regions with high information content considering either the LBAIM or GBAIM context.</p>
<p>Similar to the development of human infants, both the saccadic and vergence control policies in our model are immature at the start of simulation and emerge through interaction with the environment. For the saccade policy, this is because the basis functions are randomly initialized. For the vergence policy, this is because both the basis functions and the weights in the actor and critic networks are randomly initialized. The order in which these policies emerge is the same as in humans. In our simulation, the saccade policy develops before the vergence policy. In humans, infants begin to look at edges or places with sharp and high-contrast features from 1 to 2&#x02009;months of age (Bronson, <xref ref-type="bibr" rid="B4">1990</xref>, <xref ref-type="bibr" rid="B5">1991</xref>; Colombo, <xref ref-type="bibr" rid="B8">2001</xref>). Vergence develops at about 4&#x02009;months of age (Aslin, <xref ref-type="bibr" rid="B1">1977</xref>; Hainline and Riddell, <xref ref-type="bibr" rid="B12">1995</xref>). However, the absolute time scale at which these behaviors emerge differs. We measure the time it takes these policies to emerge in terms of fixations, since the number of iterations required by the model is somewhat arbitrary, as it depends upon the choice of the time step. In our simulations, the vergence policies when following the LBAIM saccade policy emerge after about 30,000 fixations (300,000 iterations) (Figure <xref ref-type="fig" rid="F5">5</xref>). Adult humans execute around 10<sup>5</sup> fixations/day,<xref ref-type="fn" rid="fn3"><sup>3</sup></xref> so this corresponds to 3&#x02009;days for an adult, but we expect this number to be a bit larger as infants spend less time awake and their saccade policies are immature. The saccade policy develops very quickly, after about 100 fixations (see <xref ref-type="sec" rid="S7">Supplementary Material</xref>), but this is largely due to the fact that it depends only on the basis functions, as much of the processing is hard coded in our model. We expect that if we incorporated learning into more aspects of the saccade control policy as is likely the case in humans, we would obtain a much slower rate of emergence.</p>
<p>We could obtain simulations where vergence policies emerge on the same timescales as in humans by lowering the learning rate. However, we did not do so to avoid excessively long computation time. We are primarily interested in the relative rates, rather than absolute rates at which the policies emerge. Since we use the same learning rate in all simulations, the faster rate at which the vergence policy emerges when following the BAIM saccade policies indicates that the system is actively choosing inputs that allow for the most effective learning. While it is clear that saccades are driven by a number of factors beyond obtaining effective inputs to train vergence control, our model does suggest a complementary, and hitherto largely unappreciated, potential role for saliency-driven saccadic eye movements.</p>
<p>Our experimental results show that vergence control policies learned with saliency-driven saccades all exhibit higher accuracy and are learned faster than when saccades are driven randomly (Figures <xref ref-type="fig" rid="F5">5</xref> and <xref ref-type="fig" rid="F6">6</xref>). The BAIM-driven saccade policies result in the highest accuracy and fastest learning. The primary function of saccades (and attention in general) is commonly thought to direct the limited neural processing of an organism to more important stimuli. Our results suggest a new complementary role of saccades in aiding in the learning of behavior.</p>
<p>Through our experiments with this model, we have identified a number of different interacting factors that account for the improved performance and faster learning.</p>
<p>First, differences in the saccade policies expose the system to input patches with different statistics. All of the attention-based saccade control models direct gaze toward image regions with higher entropy than encountered with randomly generated saccades (Figure <xref ref-type="fig" rid="F7">7</xref>). Patches selected by the LBAIM models have the highest entropy, followed by the GBAIM, Itti, and AIM models.</p>
<p>Second, since the perceptual representations adapt to the input statistics, these differences lead to differences in the perceptual representations. Higher entropy patches contain more texture, which provides more visual cues to disparity. Perceptual representations learned using higher entropy patches encode differences between zero and non-zero disparities better. Figure <xref ref-type="fig" rid="F8">8</xref> shows the dependency between the reconstruction error and the input disparity for the different perceptual representations. The differences between the reconstruction errors for zero and non-zero disparities follow the same trend as the entropy, being the smallest for the random policy, and the largest for the LBAIM and GBAIM policies. The difference also depends upon the statistics of the disparities of the input patches, increasing the more the disparities are concentrated around zero disparity (smaller values of <italic>D</italic> in Eq. <xref ref-type="disp-formula" rid="E23">23</xref>). The magnitude of the reconstruction error at zero disparity shows the opposite trend, achieving the largest value for the random policy and the smallest values for the BAIM-based policies.</p>
<p>Third, the differences between the reconstruction error curves in Figure <xref ref-type="fig" rid="F8">8</xref> are amplified during joint learning by a positive feedback loop setup by the interaction between the learning of the perceptual representation and the learning of vergence control. Both learners seek to maximize the same reward: the average negative reconstruction error of the perceptual representation. Initially, the disparity statistics will be similar, since the vergence policies are initialized with random weights. The different saccade policies will result in slightly different reconstruction error curves. The lower reconstruction error at zero disparity and the larger difference between the reconstruction error at non-zero and zero disparities for the BAIM-driven saccade control will cause the reinforcement learner to favor more strongly the emergence of vergence policies that seek to 0 out the retinal disparity, resulting in a slightly better vergence control policy. In turn, the better vergence policies cause the distribution of retinal disparities presented to the perceptual representation to be more tightly clustered around 0. The perceptual representation will respond by allocating more basis functions to represent zero disparity inputs. This further reduces the reconstruction error for zero disparity inputs and increases the difference at non-zero and zero disparities. This in turn improves vergence control and the cycle continues. Figure <xref ref-type="fig" rid="F9">9</xref>A shows the net effect of this positive feedback loop by plotting the reconstruction error curves of the perceptual representations learned after joint learning under different saccade control policies. The differences between the perceptual representations are much more pronounced, with curves corresponding to the BAIM-driven saccades policies exhibiting much more pronounced &#x0201C;V&#x0201D; shapes. Figure <xref ref-type="fig" rid="F9">9</xref>B shows the dynamic evolution of the slope of the &#x0201C;V&#x0201D; shape at its minimum point, which is a measure of the difference in reconstruction error at zero and non-zero disparities. Small initial differences expand rapidly under the positive feedback.</p>
<p>Finally, the BAIM-driven saccades direct the system to focus on more &#x0201C;challenging&#x0201D; situations, i.e., those with larger initial retinal disparity or those where the potential change in the reward are larger. In particular, we find that the LBAIM algorithm, by emphasizing saccade targets that are different from the current fixation, exposes the system to a wider diversity of input patch textures and input disparities, which drives faster learning. Intuitively, saccades between targets at different depths will present more challenges for vergence control, since they require larger change in vergence angle between the two eyes. In our system, vergence angle is preserved across saccades. Thus, saccades to locations with the same absolute disparity as the current fixation will require no change in vergence angle, presenting less of a challenge to the vergence control policy than saccades to targets with a different absolute disparity.</p>
<p>The primary difference between the LBAIM and GBAIM saliency models is that for the LBAIM model, the saliency depends upon the current fixation point, whereas for the GBAIM model, saliency is independent of the current fixation point. This dependency is introduced due to the data used to compute the coefficients of the empirical response histogram (Eq. <xref ref-type="disp-formula" rid="E17">17</xref>). For the LBAIM model, points whose feature extractor responses are different from those around the current fixation point will be more salient. Since the feature extractors encode disparity, this implies that points with disparity different from the current fixation point will be more salient under LBAIM. Thus, we observe greater differences in disparity between adjacent fixations (Figure <xref ref-type="fig" rid="F11">11</xref>).</p>
<p>We also observe larger reductions in reconstruction error during fixations (Figure <xref ref-type="fig" rid="F12">12</xref>). These larger decreases are due to the combination of a number of factors identified earlier. First, the same reduction in retinal disparity will result in a larger decrease in the reconstruction error for the BAIM policies, due to the more pronounced &#x0201C;V&#x0201D; shape of the reconstruction error curves for the BAIM policy. Second, the better quality of the vergence control policies learned under BAIM will result in larger changes in retinal disparity. Third, the choice of saccade targets will influence the change in two ways. By choosing saccade targets with larger entropy, the effect of changes in the retinal disparity on the visual input will be more pronounced, leading to larger changes in <italic>E</italic><sub>avg</sub>. In addition, choosing initial fixation points with larger retinal disparity will result in larger changes in <italic>E</italic><sub>avg</sub>. This final factor likely accounts for much of the difference between LBAIM and GBAIM.</p>
<p>Future work will focus on learning all aspects of the saccade and vergence policies simultaneously and under a common parsimonious framework provided by active efficient encoding. For example, although the BAIM saliency maps adapt to the statistics of the sensory input because of changes in the binocular feature detectors learned by the GASSOM algorithm, the way in which the feature detector outputs are integrated to construct the saliency maps is hard coded. We are currently investigating how to learn how to combine feature map outputs to generate saccade policies. In addition, our model of saccades and vergence can be made more realistic. In humans, some of the required vergence change takes place during the saccade (Coubard, <xref ref-type="bibr" rid="B9">2013</xref>), with the remaining disparity canceled by vergence changes after the saccade. Our current model is a simplification of this, since there is no change in vergence during the saccade. It will also be interesting to extend the framework here to include these initial vergence changes by incorporating an estimate of the disparity of the target.</p>
</sec>
<sec id="S5">
<title>Author Contributions</title>
<p>All authors contributed to the design of the experiments and the paper writing. QZ conducted the experiments.</p>
</sec>
<sec id="S6">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was supported in part by the Hong Kong Research Grants Council under grant 618713 and in part by the Quandt Foundation and the German Federal Ministry of Education and Research (BMBF) under Grants 01GQ1414 and 01EW1603A.</p></fn>
</fn-group>
<sec id="S7" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at <uri xlink:href="http://www.frontiersin.org/article/10.3389/fnbot.2017.00058/full&#x00023;supplementary-material">http://www.frontiersin.org/article/10.3389/fnbot.2017.00058/full&#x00023;supplementary-material</uri>.</p>
<supplementary-material xlink:href="image_1.pdf" id="SM1" mimetype="applicationn/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="video_1.mp4" id="SM2" mimetype="applicationn/mp4" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aslin</surname> <given-names>R. N.</given-names></name></person-group> (<year>1977</year>). <article-title>Development of binocular fixation in human infants</article-title>. <source>J. Exp. Child Psychol.</source> <volume>23</volume>, <fpage>133</fpage>&#x02013;<lpage>150</lpage>.<pub-id pub-id-type="doi">10.1016/0022-0965(77)90080-7</pub-id></citation></ref>
<ref id="B2"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Barlow</surname> <given-names>H. B.</given-names></name></person-group> (<year>1961</year>). <source>Possible Principles Underlying the Transformations of Sensory Messages</source>, ed. <person-group person-group-type="editor"><name><surname>Rosenblith</surname> <given-names>W. A.</given-names></name></person-group> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT press</publisher-name>), <fpage>216</fpage>&#x02013;<lpage>234</lpage>.</citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bhatnagar</surname> <given-names>S.</given-names></name> <name><surname>Sutton</surname> <given-names>R. S.</given-names></name> <name><surname>Ghavamzadeh</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Natural actor-critic algorithms</article-title>. <source>Automatica</source> <volume>45</volume>, <fpage>2471</fpage>&#x02013;<lpage>2482</lpage>.<pub-id pub-id-type="doi">10.1016/j.automatica.2009.07.008</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bronson</surname> <given-names>G. W.</given-names></name></person-group> (<year>1990</year>). <article-title>Changes in infants&#x02019; visual scanning across the 2-to 14-week age period</article-title>. <source>J. Exp. Child Psychol.</source> <volume>49</volume>, <fpage>101</fpage>&#x02013;<lpage>125</lpage>.<pub-id pub-id-type="doi">10.1016/0022-0965(90)90051-9</pub-id><pub-id pub-id-type="pmid">2303772</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bronson</surname> <given-names>G. W.</given-names></name></person-group> (<year>1991</year>). <article-title>Infant differences in rate of visual encoding</article-title>. <source>Child Dev.</source> <volume>62</volume>, <fpage>44</fpage>&#x02013;<lpage>54</lpage>.<pub-id pub-id-type="doi">10.2307/1130703</pub-id><pub-id pub-id-type="pmid">2022137</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bruce</surname> <given-names>N. D.</given-names></name> <name><surname>Tsotsos</surname> <given-names>J. K.</given-names></name></person-group> (<year>2009</year>). <article-title>Saliency, attention, and visual search: an information theoretic approach</article-title>. <source>J. Vis.</source> <volume>9</volume>, <fpage>5.1</fpage>&#x02013;<lpage>24</lpage>.<pub-id pub-id-type="doi">10.1167/9.3.5</pub-id><pub-id pub-id-type="pmid">19757944</pub-id></citation></ref>
<ref id="B7"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Chandrapala</surname> <given-names>T. N.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name></person-group> (<year>2014</year>). &#x0201C;<article-title>The generative adaptive subspace self-organizing map</article-title>,&#x0201D; in <conf-name>International Joint Conference on Neural Networks</conf-name>, <conf-loc>Beijing</conf-loc>.<pub-id pub-id-type="doi">10.1109/IJCNN.2014.6889796</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Colombo</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>The development of visual attention in infancy</article-title>. <source>Annu. Rev. Psychol.</source> <volume>52</volume>, <fpage>337</fpage>&#x02013;<lpage>367</lpage>.<pub-id pub-id-type="doi">10.1146/annurev.psych.52.1.337</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coubard</surname> <given-names>O. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Saccade and vergence eye movements: a review of motor and premotor commands</article-title>. <source>Eur. J. Neurosci.</source> <volume>38</volume>, <fpage>3384</fpage>&#x02013;<lpage>3397</lpage>.<pub-id pub-id-type="doi">10.1111/ejn.12356</pub-id><pub-id pub-id-type="pmid">24103028</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dorris</surname> <given-names>M. C.</given-names></name> <name><surname>Klein</surname> <given-names>R. M.</given-names></name> <name><surname>Everling</surname> <given-names>S.</given-names></name> <name><surname>Munoz</surname> <given-names>D. P.</given-names></name></person-group> (<year>2002</year>). <article-title>Contribution of the primate superior colliculus to inhibition of return</article-title>. <source>J. Cogn. Neurosci.</source> <volume>14</volume>, <fpage>1256</fpage>&#x02013;<lpage>1263</lpage>.<pub-id pub-id-type="doi">10.1162/089892902760807249</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>D.</given-names></name> <name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Vasconcelos</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <article-title>Discriminant saliency, the detection of suspicious coincidences, and applications to visual recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>31</volume>, <fpage>989</fpage>&#x02013;<lpage>1005</lpage>.<pub-id pub-id-type="doi">10.1109/tpami.2009.27</pub-id><pub-id pub-id-type="pmid">19372605</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hainline</surname> <given-names>L.</given-names></name> <name><surname>Riddell</surname> <given-names>P. M.</given-names></name></person-group> (<year>1995</year>). <article-title>Binocular alignment and vergence in early infancy</article-title>. <source>Vision Res.</source> <volume>35</volume>, <fpage>3229</fpage>&#x02013;<lpage>3236</lpage>.<pub-id pub-id-type="doi">10.1016/0042-6989(95)00074-o</pub-id></citation></ref>
<ref id="B13"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>B.</given-names></name> <name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Ding</surname> <given-names>Y.</given-names></name></person-group> (<year>2011</year>). &#x0201C;<article-title>Bottom-up saliency based on weighted sparse coding residual</article-title>,&#x0201D; in <conf-name>Proceedings of the 19th ACM International Conference on Multimedia</conf-name>, <conf-loc>Scottsdale</conf-loc>.<pub-id pub-id-type="doi">10.1145/2072298.2071952</pub-id></citation></ref>
<ref id="B14"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2007</year>). &#x0201C;<article-title>Saliency detection: a spectral residual approach</article-title>,&#x0201D; in <conf-name>IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Minneapolis</conf-loc>.</citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Itti</surname> <given-names>L.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name> <name><surname>Niebur</surname> <given-names>E.</given-names></name></person-group> (<year>1998</year>). <article-title>A model of saliency-based visual attention for rapid scene analysis</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>20</volume>, <fpage>1254</fpage>&#x02013;<lpage>1259</lpage>.<pub-id pub-id-type="doi">10.1109/34.730558</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kanan</surname> <given-names>C.</given-names></name> <name><surname>Tong</surname> <given-names>M. H.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Cottrell</surname> <given-names>G. W.</given-names></name></person-group> (<year>2009</year>). <article-title>SUN: top-down saliency using natural statistics</article-title>. <source>Vis. cogn.</source> <volume>17</volume>, <fpage>979</fpage>&#x02013;<lpage>1003</lpage>.<pub-id pub-id-type="doi">10.1080/13506280902771138</pub-id><pub-id pub-id-type="pmid">21052485</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klimmasch</surname> <given-names>L.</given-names></name> <name><surname>Lelais</surname> <given-names>A.</given-names></name> <name><surname>Lichtenstein</surname> <given-names>A.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name> <name><surname>Triesch</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <source>Learning of Active Binocular Vision in a Biomechanical Model of the Oculomotor System</source>. bioRxiv: 160721.</citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Z.</given-names></name></person-group> (<year>2002</year>). <article-title>A saliency map in primary visual cortex</article-title>. <source>Trends Cogn. Sci.</source> <volume>6</volume>, <fpage>9</fpage>&#x02013;<lpage>16</lpage>.<pub-id pub-id-type="doi">10.1016/S1364-6613(00)01817-9</pub-id><pub-id pub-id-type="pmid">11849610</pub-id></citation></ref>
<ref id="B19"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Zou</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>X.</given-names></name></person-group> (<year>2012</year>). &#x0201C;<article-title>Salient region detection based on binocular vision</article-title>,&#x0201D; in <conf-name>7th IEEE Conference on Industrial Electronics and Applications</conf-name>, <conf-loc>Singapore</conf-loc>.<pub-id pub-id-type="doi">10.1109/ICIEA.2012.6361031</pub-id></citation></ref>
<ref id="B20"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Lonini</surname> <given-names>L.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Chandrashekhariah</surname> <given-names>P.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name> <name><surname>Triesch</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). &#x0201C;<article-title>Autonomous learning of active multi-scale binocular vision</article-title>,&#x0201D; in <conf-name>IEEE Third Joint International Conference on Development and Learning and Epigenetic Robotics</conf-name>, <conf-loc>Osaka</conf-loc>.<pub-id pub-id-type="doi">10.1109/DevLrn.2013.6652541</pub-id></citation></ref>
<ref id="B21"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Martull</surname> <given-names>S.</given-names></name> <name><surname>Peris</surname> <given-names>M.</given-names></name> <name><surname>Fukui</surname> <given-names>K.</given-names></name></person-group> (<year>2012</year>). &#x0201C;<article-title>Realistic CG stereo image dataset with ground truth disparity maps</article-title>,&#x0201D; in <source>ICPR Workshop TrakMark</source>, <publisher-loc>Tsukuba</publisher-loc>.</citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohzawa</surname> <given-names>I.</given-names></name> <name><surname>DeAngelis</surname> <given-names>G. C.</given-names></name> <name><surname>Freeman</surname> <given-names>R. D.</given-names></name></person-group> (<year>1990</year>). <article-title>Stereoscopic depth discrimination in the visual cortex: neurons ideally suited as disparity detectors</article-title>. <source>Science</source> <volume>249</volume>, <fpage>1037</fpage>&#x02013;<lpage>1041</lpage>.<pub-id pub-id-type="doi">10.1126/science.2396096</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olmos</surname> <given-names>A.</given-names></name> <name><surname>Kingdom</surname> <given-names>F. A.</given-names></name></person-group> (<year>2004</year>). <article-title>A biologically inspired algorithm for the recovery of shading and reflectance images</article-title>. <source>Perception</source> <volume>33</volume>, <fpage>1463</fpage>&#x02013;<lpage>1473</lpage>.<pub-id pub-id-type="doi">10.1068/p5321</pub-id><pub-id pub-id-type="pmid">15729913</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poggio</surname> <given-names>G. F.</given-names></name> <name><surname>Fischer</surname> <given-names>B.</given-names></name></person-group> (<year>1977</year>). <article-title>Binocular interaction and depth sensitivity in striate and prestriate cortex of behaving rhesus monkey</article-title>. <source>J. Neurophysiol.</source> <volume>40</volume>, <fpage>1392</fpage>&#x02013;<lpage>1405</lpage>.</citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Posner</surname> <given-names>M. I.</given-names></name> <name><surname>Cohen</surname> <given-names>Y.</given-names></name></person-group> (<year>1984</year>). <article-title>Components of visual orienting</article-title>. <source>Atten. Perform. X Control Lang. Processes</source> <volume>32</volume>, <fpage>531</fpage>&#x02013;<lpage>556</lpage>.</citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Renninger</surname> <given-names>L. W.</given-names></name> <name><surname>Verghese</surname> <given-names>P.</given-names></name> <name><surname>Coughlan</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Where to look next? Eye movements reduce local uncertainty</article-title>. <source>J. Vis.</source> <volume>7</volume>, <fpage>6</fpage>.<pub-id pub-id-type="doi">10.1167/7.3.6</pub-id><pub-id pub-id-type="pmid">17461684</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teuli&#x000E8;re</surname> <given-names>C.</given-names></name> <name><surname>Forestier</surname> <given-names>S.</given-names></name> <name><surname>Lonini</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Shi</surname> <given-names>B.</given-names></name> <etal/></person-group> (<year>2015</year>). <article-title>Self-calibrating smooth pursuit through active efficient coding</article-title>. <source>Rob. Auton. Syst.</source> <volume>71</volume>, <fpage>3</fpage>&#x02013;<lpage>12</lpage>.<pub-id pub-id-type="doi">10.1016/j.robot.2014.11.006</pub-id></citation></ref>
<ref id="B28"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Tikhanoff</surname> <given-names>V.</given-names></name> <name><surname>Cangelosi</surname> <given-names>A.</given-names></name> <name><surname>Fitzpatrick</surname> <given-names>P.</given-names></name> <name><surname>Metta</surname> <given-names>G.</given-names></name> <name><surname>Natale</surname> <given-names>L.</given-names></name> <name><surname>Nori</surname> <given-names>F.</given-names></name></person-group> (<year>2008</year>). &#x0201C;<article-title>An open-source simulator for cognitive robotics research: the prototype of the iCub humanoid robot simulator</article-title>,&#x0201D; in <conf-name>Proceedings of the 8th Workshop on Performance Metrics for Intelligent Systems</conf-name>, <conf-loc>Gaithersburg</conf-loc>.<pub-id pub-id-type="doi">10.1145/1774674.1774684</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Triesch</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Imitation learning based on an intrinsic motivation mechanism for efficient coding</article-title>. <source>Front. Psychol.</source> <volume>4</volume>:<fpage>800</fpage>.<pub-id pub-id-type="doi">10.3389/fpsyg.2013.00800</pub-id><pub-id pub-id-type="pmid">24204350</pub-id></citation></ref>
<ref id="B30"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Vikram</surname> <given-names>T. N.</given-names></name> <name><surname>Teuli&#x000E8;re</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name> <name><surname>Triesch</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). &#x0201C;<article-title>Autonomous learning of smooth pursuit and vergence through active efficient coding</article-title>,&#x0201D; in <conf-name>4th International Conference on Development and Learning and on Epigenetic Robotics</conf-name>, <conf-loc>Genoa</conf-loc>.<pub-id pub-id-type="doi">10.1109/devlrn.2014.6983022</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>DaSilva</surname> <given-names>M. P.</given-names></name> <name><surname>LeCallet</surname> <given-names>P.</given-names></name> <name><surname>Ricordel</surname> <given-names>V.</given-names></name></person-group> (<year>2013</year>). <article-title>Computational model of stereoscopic 3D visual saliency</article-title>. <source>IEEE Trans. Image Process.</source> <volume>22</volume>, <fpage>2151</fpage>&#x02013;<lpage>2165</lpage>.<pub-id pub-id-type="doi">10.1109/TIP.2013.2246176</pub-id><pub-id pub-id-type="pmid">23412612</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wolfe</surname> <given-names>J. M.</given-names></name> <name><surname>Horowitz</surname> <given-names>T. S.</given-names></name></person-group> (<year>2004</year>). <article-title>What attributes guide the deployment of visual attention and how do they do it?</article-title> <source>Nat. Rev. Neurosci.</source> <volume>5</volume>, <fpage>495</fpage>&#x02013;<lpage>501</lpage>.<pub-id pub-id-type="doi">10.1038/nrn1411</pub-id></citation></ref>
<ref id="B33"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>M. H.</given-names></name></person-group> (<year>2012</year>). &#x0201C;<article-title>Top-down visual saliency via joint CRF and dictionary learning</article-title>,&#x0201D; in <conf-name>IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Providence, Rhode Island</conf-loc>.<pub-id pub-id-type="doi">10.1109/cvpr.2012.6247940</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Q.</given-names></name> <name><surname>Bucci</surname> <given-names>M. P.</given-names></name> <name><surname>Kapoula</surname> <given-names>Z.</given-names></name></person-group> (<year>2002</year>). <article-title>The latency of saccades, vergence, and combined eye movements in children and in adults</article-title>. <source>Invest. Ophthalmol. Vis. Sci.</source> <volume>43</volume>, <fpage>2939</fpage>&#x02013;<lpage>2949</lpage>. Available at: <uri xlink:href="http://iovs.arvojournals.org/article.aspx?articleid&#x0003D;2162756">http://iovs.arvojournals.org/article.aspx?articleid&#x0003D;2162756</uri></citation></ref>
<ref id="B35"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Yarbus</surname> <given-names>A. L.</given-names></name></person-group> (<year>1967</year>). &#x0201C;<article-title>Eye movements during perception of complex objects</article-title>,&#x0201D; in <source>Eye Movements and Vision</source>, ed. <person-group person-group-type="editor"><name><surname>Riggs</surname> <given-names>L. A.</given-names></name></person-group> (<publisher-loc>New York</publisher-loc>: <publisher-name>Plenum Press</publisher-name>), <fpage>171</fpage>&#x02013;<lpage>211</lpage>.<pub-id pub-id-type="doi">10.1007/978-1-4899-5379-7_8</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zee</surname> <given-names>D. S.</given-names></name> <name><surname>Fitzgibbon</surname> <given-names>E. J.</given-names></name> <name><surname>Optican</surname> <given-names>L. M.</given-names></name></person-group> (<year>1992</year>). <article-title>Saccade-vergence interactions in humans</article-title>. <source>J. Neurophysiol.</source> <volume>68</volume>, <fpage>1624</fpage>&#x02013;<lpage>1641</lpage>.<pub-id pub-id-type="pmid">1479435</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Triesch</surname> <given-names>J.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name></person-group> (<year>2016</year>). <article-title>An active-efficient-coding model of optokinetic nystagmus</article-title>. <source>J. Vis.</source> <volume>16</volume>, <fpage>10</fpage>&#x02013;<lpage>10</lpage>.<pub-id pub-id-type="doi">10.1167/16.14.10</pub-id><pub-id pub-id-type="pmid">27832268</pub-id></citation></ref>
<ref id="B38"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Triesch</surname> <given-names>J.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name></person-group> (<year>2014</year>). &#x0201C;<article-title>Intrinsically motivated learning of visual motion perception and smooth pursuit</article-title>,&#x0201D; in <conf-name>International Conference on Robotics and Automation</conf-name>, <conf-loc>Hong Kong</conf-loc>.<pub-id pub-id-type="doi">10.1109/icra.2014.6907110</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Tong</surname> <given-names>M. H.</given-names></name> <name><surname>Marks</surname> <given-names>T. K.</given-names></name> <name><surname>Shan</surname> <given-names>H.</given-names></name> <name><surname>Cottrell</surname> <given-names>G. W.</given-names></name></person-group> (<year>2008</year>). <article-title>SUN: a Bayesian framework for saliency using natural statistics</article-title>. <source>J. Vis.</source> <volume>8</volume>, <fpage>32</fpage>&#x02013;<lpage>32</lpage>.<pub-id pub-id-type="doi">10.1167/8.7.32</pub-id><pub-id pub-id-type="pmid">19146264</pub-id></citation></ref>
<ref id="B40"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Rothkopf</surname> <given-names>C. A.</given-names></name> <name><surname>Triesch</surname> <given-names>J.</given-names></name> <name><surname>Shi</surname> <given-names>B. E.</given-names></name></person-group> (<year>2012</year>). &#x0201C;<article-title>A unified model of the joint development of disparity selectivity and vergence control</article-title>,&#x0201D; in <conf-name>IEEE International Conference on Development and Learning and Epigenetic Robotics</conf-name>, <conf-loc>San Diego</conf-loc>.<pub-id pub-id-type="doi">10.1109/DevLrn.2012.6400876</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup>The video is available online at <uri xlink:href="https://youtu.be/axgbhDER1ow">https://youtu.be/axgbhDER1ow</uri>.</p></fn>
<fn id="fn2"><p><sup>2</sup>We implemented Itti&#x02019;s and AIM saliency model by using the code available at <uri xlink:href="http://www.vision.caltech.edu/harel/share/gbvs.php">http://www.vision.caltech.edu/harel/share/gbvs.php</uri> and <uri xlink:href="http://www-sop.inria.fr/members/Neil.Bruce/&#x00023;SOURCECODE">http://www-sop.inria.fr/members/Neil.Bruce/&#x00023;SOURCECODE</uri>. Our implementations of the GBAIM and LBAIM are based on the code of the AIM saliency model.</p></fn>
<fn id="fn3"><p><sup>3</sup>Here, we assume that adult humans are awake for 15&#x02009;h/day and make 2 saccades/s.</p></fn>
</fn-group>
</back>
</article>