<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2017.00007</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Modeling the Dynamics of Human Brain Activity with Recurrent Neural Networks</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>Umut</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/156164/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>van Gerven</surname> <given-names>Marcel A. J.</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/28677/overview"/>
</contrib>
</contrib-group>
<aff><institution>Donders Institute for Brain, Cognition and Behaviour, Radboud University</institution> <country>Nijmegen, Netherlands</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Alexandre Gramfort, T&#x000E9;l&#x000E9;com ParisTech, France</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Iris I. A. Groen, National Institutes of Health, USA; Karl Friston, University College London, UK; Seyed-Mahdi Khaligh-Razavi, Massachusetts Institute of Technology, USA</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Umut G&#x000FC;&#x000E7;l&#x000FC; <email>u.guclu&#x00040;donders.ru.nl</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>09</day>
<month>02</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>11</volume>
<elocation-id>7</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>10</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>01</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 G&#x000FC;&#x000E7;l&#x000FC; and van Gerven.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>G&#x000FC;&#x000E7;l&#x000FC; and van Gerven</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Encoding models are used for predicting brain activity in response to sensory stimuli with the objective of elucidating how sensory information is represented in the brain. Encoding models typically comprise a nonlinear transformation of stimuli to features (feature model) and a linear convolution of features to responses (response model). While there has been extensive work on developing better feature models, the work on developing better response models has been rather limited. Here, we investigate the extent to which recurrent neural network models can use their internal memories for nonlinear processing of arbitrary feature sequences to predict feature-evoked response sequences as measured by functional magnetic resonance imaging. We show that the proposed recurrent neural network models can significantly outperform established response models by accurately estimating long-term dependencies that drive hemodynamic responses. The results open a new window into modeling the dynamics of brain activity in response to sensory stimuli.</p></abstract>
<kwd-group>
<kwd>encoding</kwd>
<kwd>fMRI</kwd>
<kwd>RNN</kwd>
<kwd>LSTM</kwd>
<kwd>GRU</kwd>
</kwd-group>
<contract-num rid="cn001">639.072.513</contract-num>
<contract-sponsor id="cn001">Nederlandse Organisatie voor Wetenschappelijk Onderzoek<named-content content-type="fundref-id">10.13039/501100003246</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="1"/>
<equation-count count="20"/>
<ref-count count="60"/>
<page-count count="14"/>
<word-count count="9534"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Encoding models (Naselaris et al., <xref ref-type="bibr" rid="B44">2011</xref>) are used for predicting brain activity in response to naturalistic stimuli (Felsen and Dan, <xref ref-type="bibr" rid="B10">2005</xref>) with the objective of understanding how sensory information is represented in the brain. Encoding models typically comprise two main components. The first component is a feature model that nonlinearly transforms stimuli to features (i.e., the independent variables used in fMRI time series analyses). The second component is a response model that linearly transforms features to responses. While encoding models have been successfully used to characterize the relationship between stimuli in different modalities and responses in different brain regions, their performance usually falls short of the expected performance of the true encoding model given the noise in the analyzed data (noise ceiling). This means that there usually is unexplained variance in the analyzed data that can be explained solely by improving the encoding models.</p>
<p>One way to reach the noise ceiling is the development of better feature models. Recently, there has been extensive work in this direction. One example is the use of convolutional neural network representations of natural images or natural movies to explain low-, mid- and high-level representations in different brain regions along the ventral (Agrawal et al., <xref ref-type="bibr" rid="B1">2014</xref>; Cadieu et al., <xref ref-type="bibr" rid="B3">2014</xref>; Khaligh-Razavi and Kriegeskorte, <xref ref-type="bibr" rid="B32">2014</xref>; Yamins et al., <xref ref-type="bibr" rid="B59">2014</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B22">2015a</xref>; Cichy et al., <xref ref-type="bibr" rid="B5">2016</xref>) and dorsal streams (G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B23">2015b</xref>; Eickenberg et al., <xref ref-type="bibr" rid="B8">2016</xref>) of the human visual system. Another example is the use of manually constructed or statistically estimated representations of words and phrases to explain the semantic representations in different brain regions (Mitchell et al., <xref ref-type="bibr" rid="B42">2008</xref>; Huth et al., <xref ref-type="bibr" rid="B28">2012</xref>; Murphy et al., <xref ref-type="bibr" rid="B43">2012</xref>; Fyshe et al., <xref ref-type="bibr" rid="B15">2013</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B24">2015c</xref>; Nishida et al., <xref ref-type="bibr" rid="B46">2015</xref>).</p>
<p>Another way to reach the noise ceiling is the development of better response models. There is a long history of estimating hemodynamic response functions (HRFs) in fMRI time series modeling. The standard general linear (convolution) model used in procedures like statistical parametric mapping (SPM) expands the HRF in terms of orthogonal kernels or temporal basis functions that have been motivated in terms of Volterra expansions. Indeed, commonly used software packages such as the SPM software have (hidden) facilities to model second-order Volterra kernels that enable modeling of non-linear hemodynamic effects such as saturation. In reality, the transformation from stimulus features to observed responses is exceedingly complex because of various temporal dependencies that are caused by neurovascular coupling (Logothetis and Wandell, <xref ref-type="bibr" rid="B37">2004</xref>; Norris, <xref ref-type="bibr" rid="B49">2006</xref>) and other more elusive cognitive or neural factors.</p>
<p>Here, our objective is to develop a model that can be trained end to end, captures temporal dependencies and processes arbitrary input sequences for time-continuous fMRI experiments such as watching movies, listening to music or playing video games. Such time-continuous designs are characterized by the absence of discrete experimental events as those found in their block or event-related counterparts. To this end, we use recurrent neural networks (RNNs) as response models in the encoding framework. Recently, RNNs in general and two RNN variants&#x02014;long short-term memory (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B27">1997</xref>) and gated recurrent units (Cho et al., <xref ref-type="bibr" rid="B4">2014</xref>)&#x02014;in particular have been shown to be extremely successful in various tasks that involve processing of arbitrary input sequences such as handwriting recognition (Graves et al., <xref ref-type="bibr" rid="B18">2009</xref>; Graves, <xref ref-type="bibr" rid="B17">2013</xref>), language modeling (Sutskever et al., <xref ref-type="bibr" rid="B54">2011</xref>; Graves, <xref ref-type="bibr" rid="B17">2013</xref>), machine translation (Cho et al., <xref ref-type="bibr" rid="B4">2014</xref>) and speech recognition (Sak et al., <xref ref-type="bibr" rid="B52">2014</xref>). These models use their internal memories to capture the temporal dependencies that are informative about solving the task at hand. That is, these models base their predictions not only to the information available at a given time, but also to the information that was available in the past. They accomplish this by maintaining an explicit or implicit representation of the past input sequences and use it to make their predictions at each time point. If these models can be used as response models in the encoding framework, it will open a new window into modeling brain activity in response to sensory stimuli since the brain activity is modulated by long temporal dependencies.</p>
<p>While the use of RNNs in the encoding framework has been proposed a number of times (G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B22">2015a</xref>,<xref ref-type="bibr" rid="B23">b</xref>; Kriegeskorte, <xref ref-type="bibr" rid="B34">2015</xref>; Yamins and DiCarlo, <xref ref-type="bibr" rid="B57">2016a</xref>,<xref ref-type="bibr" rid="B58">b</xref>), these proposals mainly focused on using RNNs as feature models. In contrast, we have framed our approach in terms of response models used in characterizing distributed or multivariate responses to stimuli in the encoding framework. The key thing that we bring to the table is a generic and potentially useful response model that transforms features to observed (hemodynamic) responses. From the perspective of conventional analyses of functional magnetic resonance imaging (fMRI) time series, this response model corresponds to the convolution model used to map stimulus features (e.g., the presence of biological motion) to fMRI responses. In other words, the stimulus features correspond to conventional stimulus functions that enter standard convolution models of fMRI time series (e.g., the GLM used in statistical parametric mapping).</p>
<p>In brief, we know that the transformation from neuronal responses to fMRI signals is mediated by neuronal and hemodynamic factors that can always be expressed in terms of a non-linear convolution. A general form for these convolutions has been previously considered in the form of Volterra kernels or functional Taylor expansions (Friston et al., <xref ref-type="bibr" rid="B14">2000</xref>). Crucially, it is also well known that RNNs are universal non-linear approximators that can reproduce any Volterra expansion (Wray and Green, <xref ref-type="bibr" rid="B56">1994</xref>). This means that we can use RNNs as an inclusive and flexible way to parameterize the convolution of stimulus features generating hemodynamic responses. Furthermore, we can use RNNs to model not just response of a single voxel but distributed responses over multiple voxels. Having established the parametric form of this convolution, the statistical evidence or significance of each regionally specific convolution can then be assessed using standard (cross-validation) machine learning techniques by comparing the accuracy of the convolution when applied to test data after optimization with training data.</p>
<p>We test our approach by comparing how well a family of RNN models and a family of ridge regression models can predict blood-oxygen-level dependent (BOLD) hemodynamic responses to high-level and low-level features of natural movies using cross-validation. We show that the proposed recurrent neural network models can significantly outperform the standard ridge regression models and accurately estimate hemodynamic response functions by capturing temporal dependencies in the data.</p></sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and methods</title>
<sec>
<title>2.1. Data set</title>
<p>We analyzed the vim-2 data set (Nishimoto et al., <xref ref-type="bibr" rid="B48">2014</xref>), which was originally published by Nishimoto et al. (<xref ref-type="bibr" rid="B47">2011</xref>). The experimental procedures are identical to those in Nishimoto et al. (<xref ref-type="bibr" rid="B47">2011</xref>). Briefly, the data set has twelve 600 s blocks of stimulus and response sequences in a training set and nine 60 s blocks of stimulus and response sequences in a test set. The stimulus sequences are videos (512 px &#x000D7; 512 px or 20&#x000B0; &#x000D7; 20&#x000B0;, 15 FPS) that were drawn from various sources. The response sequences are BOLD responses (voxel size &#x0003D; 2 &#x000D7; 2 &#x000D7; 2.5 mm<sup>3</sup>, TR &#x0003D; 1 s) that were acquired from the occipital cortices of three subjects (S1, S2, and S3). The stimulus sequences in the test set were repeated ten times. The corresponding response sequences were averaged over the repetitions. The response sequences have already been preprocessed as described in Nishimoto et al. (<xref ref-type="bibr" rid="B47">2011</xref>). Briefly, they have been realigned to compensate for motion, detrended to compensate for drift and z-scored. Additionally, the first six seconds of the blocks were discarded. No further preprocessing was performed. Regions of interests were localized using the multifocal retinotopic mapping technique on retinotopic mapping data that were acquired in separate sessions (Hansen et al., <xref ref-type="bibr" rid="B25">2004</xref>). As a result, the voxels were grouped into 16 areas. However, not all areas were identified in all subjects (Table <xref ref-type="table" rid="T1">1</xref>). The last 45 seconds of the blocks in the training set were used as the validation set.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Number of voxels per subject and area</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>V2</bold></th>
<th valign="top" align="center"><bold>V3</bold></th>
<th valign="top" align="center"><bold>V1</bold></th>
<th valign="top" align="center"><bold>IPS</bold></th>
<th valign="top" align="center"><bold>V4</bold></th>
<th valign="top" align="center"><bold>LOC</bold></th>
<th valign="top" align="center"><bold>V7</bold></th>
<th valign="top" align="center"><bold>MT&#x0002B;</bold></th>
<th valign="top" align="center"><bold>V3A</bold></th>
<th valign="top" align="center"><bold>V3B</bold></th>
<th valign="top" align="center"><bold>VO</bold></th>
<th valign="top" align="center"><bold>EBA</bold></th>
<th valign="top" align="center"><bold>OFA</bold></th>
<th valign="top" align="center"><bold>RSC</bold></th>
<th valign="top" align="center"><bold>pSTS</bold></th>
<th valign="top" align="center"><bold>TOS</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">S1</td>
<td valign="top" align="center">1,477</td>
<td valign="top" align="center">1,141</td>
<td valign="top" align="center">994</td>
<td valign="top" align="center">2,251</td>
<td valign="top" align="center">734</td>
<td valign="top" align="center">885</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">466</td>
<td valign="top" align="center">252</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">410</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">71</td>
<td valign="top" align="center">45</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">S2</td>
<td valign="top" align="center">1,659</td>
<td valign="top" align="center">1,360</td>
<td valign="top" align="center">1,043</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1032</td>
<td valign="top" align="center">614</td>
<td valign="top" align="center">400</td>
<td valign="top" align="center">174</td>
<td valign="top" align="center">337</td>
<td valign="top" align="center">223</td>
<td valign="top" align="center">267</td>
<td valign="top" align="center">319</td>
<td valign="top" align="center">246</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">S3</td>
<td valign="top" align="center">1,377</td>
<td valign="top" align="center">1,131</td>
<td valign="top" align="center">1,366</td>
<td valign="top" align="center">893</td>
<td valign="top" align="center">750</td>
<td valign="top" align="center">408</td>
<td valign="top" align="center">583</td>
<td valign="top" align="center">263</td>
<td valign="top" align="center">282</td>
<td valign="top" align="center">225</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">131</td>
<td valign="top" align="center">91</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">16</td>
<td valign="top" align="center">41</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>2.2. Problem statement</title>
<p>Let <bold>x</bold><sup><italic>t</italic></sup> &#x02208; &#x0211D;<sup><italic>n</italic></sup> and <bold>y</bold><sup><italic>t</italic></sup> &#x02208; &#x0211D;<sup><italic>m</italic></sup> be a stimulus and a response at temporal interval [<italic>t, t</italic> &#x0002B; 1], where <italic>n</italic> is the number of stimulus dimensions and <italic>m</italic> is the number of voxel responses. We are interested in predicting the most likely response <bold>y</bold><sup><italic>t</italic></sup> given the stimulus history <bold>X</bold><sup><italic>t</italic></sup> &#x0003D; (<bold>x</bold><sup>0</sup>,&#x02026;,<bold>x</bold><sup><italic>t</italic></sup>):
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mo class="qopname">arg</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mo class="qopname">Pr</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mtext mathvariant="bold">g</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where Pr is an encoding distribution, &#x003D5; is a feature model such that &#x003D5; (&#x000B7;) &#x02208; &#x0211D;<sup><italic>p</italic></sup>, <italic>p</italic> is the number of feature dimensions, and <bold>g</bold> is a response model such that <bold>g</bold> (&#x000B7;) &#x02208; &#x0211D;<sup><italic>m</italic></sup>.</p>
<p>In order to solve this problem, we must define the feature model that transforms stimuli to features and the response model that transforms features to responses. We used two alternative feature models; a scene description model that codes for low-level visual features (Oliva and Torralba, <xref ref-type="bibr" rid="B50">2001</xref>) and a word embedding model that codes for high-level semantic content. We used two response model families that differ in architecture (recurrent neural network family and feedforward ridge regression family) (Figure <xref ref-type="fig" rid="F1">1</xref>). In contrast to standard convolution models for fMRI time series, we are dealing with potentially very large feature spaces. This means that in the absence of constraints the optimization of model parameters can be ill posed. Therefore, we use dropout and early stopping for the recurrent models, and <italic>L</italic><sup>2</sup> regularization for the feedforward models.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Overview of the response models</bold>. <bold>(A)</bold> Response models in the RNN family. All RNN models process feature sequences via two (recurrent) nonlinear layers and one (nonrecurrent) linear layer but differ in the type and number of artificial neurons. <italic>L-10/50/10</italic> models have 10, 50, or 100 long short-term memory units in both of their hidden layers, respectively. Similarly, <italic>G-10/50/10</italic> models have 10, 50, or 100 gated recurrent units in both of their hidden layers, respectively. <bold>(B)</bold> First-layer long short-term memory and gated recurrent units. Squares indicate linear combination and nonlinearity. Circles indicate elementwise operations. Gates in the units control the information flow between the time points. <bold>(C)</bold> Response models in the ridge regression family. All ridge regression models process feature sequences via one (nonrecurrent) linear layer but differ in how they account for the hemodynamic delay. <italic>R-C(TD)</italic> models convolve the feature sequence with the canonical hemodynamic response function (and its time and dispersion derivatives). <italic>R-F</italic> model lags the feature sequence for 3, 4, 5, and 6 s and concatenates the lagged sequences.</p></caption>
<graphic xlink:href="fncom-11-00007-g0001.tif"/>
</fig></sec>
<sec>
<title>2.3. Feature models</title>
<sec>
<title>2.3.1. High-level semantic model</title>
<p>As a high-level semantic model we used the word2vec (W2V) model by Mikolov et al. (<xref ref-type="bibr" rid="B38">2013a</xref>,<xref ref-type="bibr" rid="B39">b</xref>,<xref ref-type="bibr" rid="B40">c</xref>). This is a one-layer feedforward neural network that is trained for predicting either target words/phrases from source-context words (continuous bag-of-words) or source context-words from target words/phrases (skip-gram). Once trained, its hidden states are used as continuous distributed representations of words/phrases. These representations capture many semantic regularities. We used the pretrained (skip-gram) W2V model to avoid training from scratch (<ext-link ext-link-type="uri" xlink:href="https://code.google.com/archive/p/word2vec/">https://code.google.com/archive/p/word2vec/</ext-link>). It was trained on 100 billion-word Google News dataset. It contains 300-dimensional continuous distributed representations of three million words/phrases.</p>
<p>We used the W2V model for transforming a stimulus sequence to a feature sequence on a second-by-second basis as follows: First, each one second of the stimulus sequence is assigned 20 categories (words/phrases). We used the <italic>Clarifai</italic> service (<ext-link ext-link-type="uri" xlink:href="http://www.clarifai.com/">http://www.clarifai.com/</ext-link>) to automatically assign the categories rather than annotating them by hand. <italic>Clarifai</italic> provides a web-based video recognition application, which internally uses a pretrained deep neural network to automatically tag the contents of the video frames on a second-by-second basis. Then, each category is transformed into continuous distributed representations of words/phrases. Next, these representations are averaged over the categories. This resulted in a 300-dimensional feature vector per second of stimulus sequence (<italic>p</italic> &#x0003D; 300).</p></sec>
<sec>
<title>2.3.2. Low-level visual feature model</title>
<p>As a low-level visual feature model we used the GIST model (Oliva and Torralba, <xref ref-type="bibr" rid="B50">2001</xref>). The GIST model transforms scenes into spatial envelope representations. These representations capture many perceptual dimensions that represent the dominant spatial structure of a scene and have been used to study neural representations in a number of earlier work (Groen et al., <xref ref-type="bibr" rid="B20">2013</xref>; Leeds et al., <xref ref-type="bibr" rid="B36">2013</xref>; Cichy et al., <xref ref-type="bibr" rid="B5">2016</xref>). We used the implementation that is provided at: <ext-link ext-link-type="uri" xlink:href="http://people.csail.mit.edu/torralba/code/spatialenvelope/">http://people.csail.mit.edu/torralba/code/spatialenvelope/</ext-link>.</p>
<p>We used the GIST model for transforming a stimulus sequence to a feature sequence on a second-by-second basis as follows: First, each 16 non-overlapping 8 &#x000D7; 8 regions of all 15 128 &#x000D7; 128 frames in one second of the stimulus sequence are filtered with 32 Gabor filters that have eight orientations and four scales. Then, their energies are averaged over the frames. This resulted in a 512-dimensional feature vector per second of stimulus sequence (<italic>p</italic> &#x0003D; 512).</p></sec></sec>
<sec>
<title>2.4. Response models</title>
<sec>
<title>2.4.1. Ridge regression family</title>
<p>The response models in the ridge regression family predict feature-evoked responses as a linear combination of features. Each member of this family differs in how it accounts for the hemodynamic delay.</p>
<p>The <italic>R-C</italic> model (i) convolves the features with the canonical hemodynamic response function (Friston et al., <xref ref-type="bibr" rid="B12">1994</xref>) and (ii) predicts the responses as a linear combination of these features:
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>B</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the Toeplitz matrix of the canonical HRF. That is, it is a diagonal-constant matrix that contains the shifted versions of the HRF in its columns. Multiplying it with a signal corresponds to convolution of the HRF with the signal. Furthermore, <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <bold>B</bold> &#x02208; &#x0211D;<sup><italic>m</italic>&#x000D7;<italic>p</italic></sup> is the matrix of regression coefficients.</p>
<p>The <italic>R-CTD</italic> model (i) convolves the features with the canonical hemodynamic response function, its temporal derivative and its dispersion derivative (Friston et al., <xref ref-type="bibr" rid="B13">1998</xref>), (ii) concatenates these features and (iii) predicts the responses as a linear combination of these features:
<disp-formula id="E4"><label>(4)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>B</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the Toeplitz matrix of the the temporal derivative of the canonical HRF, <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the Toeplitz matrix of the the dispersion derivative of the canonical HRF and <bold>B</bold> &#x02208; &#x0211D;<sup><italic>m</italic>&#x000D7;3<italic>p</italic></sup> is the matrix of regression coefficients.</p>
<p>The <italic>R-F</italic> model is a finite impulse response (FIR) model that (i) lags the features for 3, 4, 5, and 6 s (Nishimoto et al., <xref ref-type="bibr" rid="B47">2011</xref>), (ii) concatenates these features and (iii) predicts the responses as a linear combination of these features:
<disp-formula id="E5"><label>(5)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>B</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M10"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>F</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mn>4</mml:mn><mml:mi>p</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <bold>B</bold> &#x02208; &#x0211D;<sup><italic>m</italic>&#x000D7;4<italic>p</italic></sup> is the matrix of regression coefficients.</p>
<p>We used the validation set for model selection (a regularization parameter per voxel) and the training set for model estimation (a row of <bold>B</bold> per voxel). Regularization parameters were selected as explained in G&#x000FC;&#x000E7;l&#x000FC; and van Gerven (<xref ref-type="bibr" rid="B21">2014</xref>). The rows of <bold>B</bold> were estimated by analytically minimizing the <italic>L</italic><sup>2</sup>-penalized least squares loss function. In related Bayesian models, this corresponds to applying shrinkage priors to the parameters (weights) of our model.</p></sec>
<sec>
<title>2.4.2. Recurrent neural network family</title>
<p>The response models in the RNN family are two-layer recurrent neural network models. They use their internal memories for nonlinearly processing arbitrary feature sequences and predicting feature-evoked responses as a linear combination of their second-layer hidden states:
<disp-formula id="E6"><label>(6)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>y</mml:mtext></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the hidden states in the second layer, and <bold>W</bold> are the weights. The RNN models differ in the type and number of artificial neurons.</p>
<p>The <italic>L-10, L-50</italic>, and <italic>L-100</italic> models are two-layer recurrent neural networks that have 10, 50, and 100 long short-term memory (LSTM) units (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B27">1997</xref>) in their hidden layers, respectively. Each LSTM unit has a cell state that acts as its internal memory by storing information from previous time points. The contents of the cell state are modulated by the gates of the unit and in turn modulate its outputs. As a result, the output of the unit is not only controlled by the present stimulus alone, but also by the stimulus history. The gates are implemented as multiplicative sigmoid functions of the inputs of the unit at the current time point and the outputs of the unit at the previous time point. That is, the gates produce values between zero and one, which are multiplied by (a function of) the cell state to determine the amount of information to store, forget or retrieve at each time point. The first-layer hidden states of an LSTM unit are defined as follows:
<disp-formula id="E7"><label>(7)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02299;</mml:mo><mml:mo class="qopname">tanh</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E8"><label>(8)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>o</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where &#x02299; denotes elementwise multiplication, <bold>c</bold><sup><italic>t</italic></sup> is the cell state, and <bold>o</bold><sup><italic>t</italic></sup> are the output gate activities. The cell state maintains information about the previous time points. The output gate controls what information will be retrieved from the cell state. The cell state of an LSTM unit is defined as:
<disp-formula id="E9"><label>(9)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>f</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02299;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>i</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02299;</mml:mo><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>f</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E11"><label>(11)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>i</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E12"><label>(12)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <bold>f</bold><sup><italic>t</italic></sup> are the forget gate activities, <bold>i</bold><sup><italic>t</italic></sup> are the input gate activities, and <inline-formula><mml:math id="M19"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is an auxiliary variable. Forget gates control what old information will be discarded from the cell states. Input gates control what new information will be stored in the cell states. Furthermore, <bold>U</bold>s and <bold>W</bold>s are the weights and <bold>b</bold>s are the biases that determine the behavior of the gates (i.e., the learnable parameters of the model).</p>
<p>The <italic>G-10, G-50</italic>, and <italic>G-100</italic> models are two-layer recurrent neural networks that have 10, 50, and 100 gated recurrent units (GRU) (Cho et al., <xref ref-type="bibr" rid="B4">2014</xref>) in the their hidden layers, respectively. The GRU units are simpler alternatives to the LSTM units. They combine hidden states with cell states and input gates with forget gates. The first-layer hidden states of a GRU unit is defined as follows:
<disp-formula id="E13"><label>(13)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02299;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02299;</mml:mo><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E14"><label>(14)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E15"><label>(15)</label><mml:math id="M22"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>r</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E16"><label>(16)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mo class="qopname">tanh</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>U</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>r</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02299;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>W</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>b</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <bold>z</bold><sup><italic>t</italic></sup> are update gate activities, <bold>r</bold><sup><italic>t</italic></sup> are reset gate activities and <inline-formula><mml:math id="M24"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>h</mml:mtext></mml:mstyle></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is an auxiliary variable. Like the gates in LSTM units, those in GRU units control the information flow between the time points. As before, <bold>U</bold>s and <bold>W</bold>s are the weights and <bold>b</bold>s are the biases that determine the behavior of the gates (i.e., the learnable parameters of the model).</p>
<p>The second-layer hidden states are defined similarly to the first-layer hidden states except for replacing the input features with the first-layer hidden states. For each previously identified brain area of each subject, a separate model was trained. That is, the voxels in a given brain area of a given subject shared the same recurrent layers but had different weights for linearly transforming the hidden states of the second recurrent layer to the response predictions. We used truncated backpropagation through time in conjunction with the optimization method Adam (Kingma and Ba, <xref ref-type="bibr" rid="B33">2014</xref>) to train the models on the training set by iteratively minimizing the mean squared error loss function. Dropout (Hinton et al., <xref ref-type="bibr" rid="B26">2012</xref>) was used to regularize the hidden layers. The epoch in which the validation performance was the highest was taken as the best model. The <italic>Chainer</italic> framework (<ext-link ext-link-type="uri" xlink:href="http://chainer.org/">http://chainer.org/</ext-link>) was used to implement the models.</p></sec></sec>
<sec>
<title>2.5. HRF estimation</title>
<p>Voxel-specific HRFs were estimated by stimulating the RNN model with an impulse. Let <bold>x</bold><sup>&#x02212;<italic>t</italic></sup>, &#x02026;, <bold>x</bold><sup>0</sup>, &#x02026;, <bold>x</bold><sup><italic>t</italic></sup> be an impulse such that <bold>x</bold> is a vector of zeros at times other than time 0 and a vector of ones at time 0. The period of the impulse before time 0 is used to stabilize the baseline of the impulse response. First, the response of the model to the impulse is simulated:
<disp-formula id="E17"><label>(17)</label><mml:math id="M25"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>g</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>-</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Then, the baseline of the impulse response before time 0 is subtracted from itself:
<disp-formula id="E18"><label>(18)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
Next, the impulse response is divided by its maximum:
<disp-formula id="E19"><label>(19)</label><mml:math id="M28"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mrow><mml:mo stretchy='false'>[</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mi>r</mml:mi><mml:mo>*</mml:mo></mml:msubsup><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup></mml:mrow></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:msubsup><mml:mrow><mml:mo stretchy='false'>[</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mi>r</mml:mi><mml:mo>*</mml:mo></mml:msubsup><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>/</mml:mo><mml:mi>max</mml:mi><mml:msubsup><mml:mrow><mml:mo stretchy='false'>[</mml:mo><mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>H</mml:mi></mml:mstyle><mml:mi>r</mml:mi><mml:mo>*</mml:mo></mml:msubsup><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
Finally, the period of the impulse response before time 0 is discarded, and the remaining period of the impulse response is taken as the HRF of the voxels:
<disp-formula id="E20"><label>(20)</label><mml:math id="M29"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>H</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
The time when the HRF is at its maximum was taken as the delay of the response, and the time after the delay of the response when the HRF was at its minimum was taken as the delay of undershoot.</p></sec>
<sec>
<title>2.6. Performance assessment</title>
<p>The performance of a model for a voxel was defined as the cross-validated Pearson&#x00027;s product-moment correlation coefficient between the observed and predicted responses of the voxel (<italic>r</italic>)<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>. Its performance for a group of voxels was defined as the median of its performance over the voxels in the group (<inline-formula><mml:math id="M30"><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:math></inline-formula>). The data of all subjects were concatenated prior to analyzing the performance of the models.</p>
<p>In order to make sure that the differences in the performance of a model in different areas are not caused by the differences in the signal-to-noise ratios of the areas, the performance of the model in an area was corrected for the median of the noise ceilings of the voxels in the area (<inline-formula><mml:math id="M31"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>) (Kay et al., <xref ref-type="bibr" rid="B31">2013</xref>). Briefly, we performed Monte Carlo simulations in which the correlation coefficient between a signal and a noisy signal is estimated. In each simulation, both the signal and the noise were drawn from a Gaussian distribution. The noisy signal was taken to be the summation of the signal sample and the noise sample. The parameters of the signal and the noise distributions were estimated from the 10 repeated measurements of the responses to the same stimuli. The noise distribution was assumed to be zero mean, and its variance was taken to be the variance of the standard errors of the data. The mean and the variance of the signal distribution were given as the mean of the data, and the difference between the variance of the data and the noise distribution, respectively. The medians of the correlation coefficients that were estimated in the simulations were taken to be the noise ceilings of the voxels, indicating the maximum performance that can be expected from the perfect model due to the noise in the data.</p>
<p>Permutation tests were used for comparing the performance of a model against chance level. First, data were randomly permuted over time for 200 times. Then, a separate model was trained and tested for each of the 200 permutations. Finally, the <italic>p</italic>-value was taken to be the fraction of the 200 permutations whose performance was greater than the actual performance. The performance was considered significant at &#x003B1; &#x0003D; 0.05 if the <italic>p</italic>-value was less than 0.05 (Bonferroni corrected for number of areas).</p>
<p>Bootstrapping was used for comparing the performance of two models over voxels in a ROI (i.e., all voxels or voxels in an area). For 10,000 repetitions, bootstrap samples (i.e., voxels) were drawn from the ROI with replacement, and the performance difference between the models over these voxels were estimated. The performance difference was considered significant at &#x003B1; &#x0003D; 0.05 if the 95% confidence interval of the sampled statistic did not cover zero (Bonferroni corrected for number of models).</p></sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Comparison of response models</title>
<p>We evaluated the response models by comparing the performance of the response models in the (recurrent) RNN family and (feed-forward) ridge regression family in combination with the (high-level) W2V model and the (low-level) GIST model. Using two feature models of different levels ruled out any potential biases in the performance difference of the response models that can be caused by the feature models. Recall that the models in the RNN family (<italic>G/L-10</italic>/<italic>50</italic>/<italic>100</italic> models) differed in the type and number of artificial neurons, whereas the models in the ridge regression family (<italic>R-C</italic>/<italic>R-CTD</italic>/<italic>R-F</italic> models) differed in how they account for the hemodynamic delay.</p>
<p>Once the best response models among the RNN family and the ridge regression family were identified, we first compared their performance in detail. Particular attention was paid to the voxels where the performance of the models differed by more than an arbitrary threshold of <italic>r</italic> &#x0003D; 0.1. We then compared the performance of the best response model among the RNN family over the areas along the visual pathway.</p>
<sec>
<title>3.1.1. Comparison of the response models in combination with the semantic model</title>
<p>Figure <xref ref-type="fig" rid="F2">2</xref> compares the performance of all response models in combination with the W2V model. The performance of the models in the RNN family that had 50 or 100 artificial neurons was always significantly higher than that of all models in the ridge regression family (<italic>p</italic> &#x02264; 0.05, bootstrapping). However, the performance of the models in the same family was not always significantly different from each other. The performance of the <italic>G-100</italic> model was the highest among the RNN family (<inline-formula><mml:math id="M32"><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>16</mml:mn></mml:math></inline-formula>), and that of the <italic>R-C</italic> model was the highest among the ridge regression family (<inline-formula><mml:math id="M33"><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>12</mml:mn></mml:math></inline-formula>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Comparison of the response models in combination with the W2V model</bold>. <bold>(A)</bold> Median performance of response models in RNN (<italic>G-X</italic> and <italic>L-X</italic>) and ridge regression (<italic>R-X</italic>) families over all voxels. Error bars indicate 95% confidence intervals (bootstrapping). Asterisks indicate significant performance difference. All of the individual bars depict significantly above chance-level performance (<italic>p</italic> &#x0003C; 0.05, permutation test). <bold>(B)</bold> Performance of best response models in RNN (<italic>G-100</italic> model) and ridge regression (<italic>R-C</italic> model) families over individual voxels. Points indicate voxels. Gray points indicate voxels where the performance difference is less than <italic>r</italic> &#x0003D; 0.1. Lines indicate (median) performance over all voxels.</p></caption>
<graphic xlink:href="fncom-11-00007-g0002.tif"/>
</fig>
<p>The performance of the <italic>G-100</italic> model and the <italic>R-C</italic> model differed from each other by more than the chosen threshold of <italic>r</italic> &#x0003D; 0.1 in 30% of the voxels. The performance of the <italic>G-100</italic> model was higher in 78% of these voxels (<inline-formula><mml:math id="M34"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>17</mml:mn></mml:math></inline-formula>), and that of the <italic>R-C</italic> model was higher in 22% of these voxels (<inline-formula><mml:math id="M35"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>14</mml:mn></mml:math></inline-formula>).</p>
<p>Figure <xref ref-type="fig" rid="F3">3</xref> compares the performance of the <italic>G-100</italic> model in combination with the W2V model over the areas along the visual stream. While the performance of the model was significantly higher than chance throughout the areas (<italic>p</italic> &#x02264; 0.05, permutation test), it was particularly high in downstream areas. For example, it was the highest in TOS (<inline-formula><mml:math id="M36"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>55</mml:mn></mml:math></inline-formula>), OFA (<inline-formula><mml:math id="M37"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>38</mml:mn></mml:math></inline-formula>) and EBA (<inline-formula><mml:math id="M38"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>35</mml:mn></mml:math></inline-formula>), and the lowest in pSTS (<inline-formula><mml:math id="M39"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>14</mml:mn></mml:math></inline-formula>), IPS (<inline-formula><mml:math id="M40"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>20</mml:mn></mml:math></inline-formula>) and V1 (<inline-formula><mml:math id="M41"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>24</mml:mn></mml:math></inline-formula>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Comparison of the <italic>G-100</italic> model in combination with the W2V model in different areas</bold>. <bold>(A)</bold> Median noise ceiling controlled performance over all voxels in different areas. Error bars indicate 95% confidence intervals (bootstrapping). All of the individual bars depict significantly above chance-level performance (<italic>p</italic> &#x0003C; 0.05, permutation test). <bold>(B)</bold> Projection of performance to cortical surfaces of S3.</p></caption>
<graphic xlink:href="fncom-11-00007-g0003.tif"/>
</fig></sec>
<sec>
<title>3.1.2. Comparison of the response models in combination with the low-level feature model</title>
<p>Figure <xref ref-type="fig" rid="F4">4</xref> compares the performance of the all response models in combination with the GIST model. The trends that were observed in this figure were similar to those that were observed in Figure <xref ref-type="fig" rid="F2">2</xref>. The <italic>G-100</italic> model was the best among the RNN family (<inline-formula><mml:math id="M42"><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>18</mml:mn></mml:math></inline-formula>), and the <italic>R-C</italic> model was the best among the ridge regression family (<inline-formula><mml:math id="M43"><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>14</mml:mn></mml:math></inline-formula>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Comparison of the response models in combination with the GIST model</bold>. <bold>(A)</bold> Median performance of response models in RNN (<italic>G-X</italic> and <italic>L-X</italic>) and ridge regression (<italic>R-X</italic>) families over all voxels. Error bars indicate 95% confidence intervals (bootstrapping). Asterisks indicate significant performance difference. All of the individual bars depict significantly above chance-level performance (<italic>p</italic> &#x0003C; 0.05, permutation test). <bold>(B)</bold> Performance of best response models in RNN (<italic>G-100</italic> model) and ridge regression (<italic>R-C model</italic>) families over individual voxels. Points indicate voxels. Gray points indicate voxels where the performance difference is less than <italic>r</italic> &#x0003D; 0.1. Lines indicate median performance over all voxels.</p></caption>
<graphic xlink:href="fncom-11-00007-g0004.tif"/>
</fig>
<p>The <italic>G-100</italic> model and the <italic>R-C</italic> differed from each other by more than the threshold of <italic>r</italic> &#x0003D; 0.1 in 27% of the voxels. The <italic>G-100</italic> model was better in 66% of these voxels (<inline-formula><mml:math id="M44"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>17</mml:mn></mml:math></inline-formula>). The <italic>R-C</italic> model was better in 34% of these voxels (<inline-formula><mml:math id="M45"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>14</mml:mn></mml:math></inline-formula>).</p>
<p>Figure <xref ref-type="fig" rid="F5">5</xref> compares the performance of the <italic>G-100</italic> model in combination with the GIST model over the areas along the visual pathway. While the <italic>G-100</italic> model performed significantly better than chance throughout the areas (<italic>p</italic> &#x02264; 0.05, permutation test), it performed particularly well in upstream visual areas. For example, it performed the best in V1 (<inline-formula><mml:math id="M46"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>39</mml:mn></mml:math></inline-formula>), V2 (<inline-formula><mml:math id="M47"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>35</mml:mn></mml:math></inline-formula>) and V3 (<inline-formula><mml:math id="M48"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>35</mml:mn></mml:math></inline-formula>), and the worst in TOS (<inline-formula><mml:math id="M49"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>13</mml:mn></mml:math></inline-formula>), IPS (<inline-formula><mml:math id="M50"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>16</mml:mn></mml:math></inline-formula>) and pSTS (<inline-formula><mml:math id="M51"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>16</mml:mn></mml:math></inline-formula>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>Comparison of the <italic>G-100</italic> model in combination with the GIST model in different areas</bold>. <bold>(A)</bold> Median noise ceiling controlled performance over all voxels in different areas. Error bars indicate 95% confidence intervals (bootstrapping). All of the individual bars depict significantly above chance-level performance (<italic>p</italic> &#x0003C; 0.05, permutation test). <bold>(B)</bold> Projection of performance to cortical surfaces of S3.</p></caption>
<graphic xlink:href="fncom-11-00007-g0005.tif"/>
</fig></sec></sec>
<sec>
<title>3.2. Comparison of feature models</title>
<p>Once the efficacy of the proposed RNN models was positively assessed, we performed a validation experiment in which we assessed the extent to which these models can replicate the earlier findings on the low-level and high-level subdivision of the visual cortex. This was accomplished by identifying the voxels that prefer semantic representations vs. low-level representations. Concretely, we compared the performance of the W2V model and the GIST model in combination with the <italic>G-100</italic> model (Figure <xref ref-type="fig" rid="F6">6</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Comparison of the feature models in combination with the <italic>G-100</italic> model</bold>. <bold>(A)</bold> Median performance difference over all voxels in different areas. Asterisks indicate significant performance difference. Error bars indicate 95% confidence intervals (bootstrapping). <bold>(B)</bold> Performance over individual voxels. Points indicate voxels. Gray points indicate voxels where performance difference is less than <italic>r</italic> &#x0003D; 0.1. Lines indicate median performance over all voxels. <bold>(C)</bold> Projection of performance difference to cortical surfaces of S3.</p></caption>
<graphic xlink:href="fncom-11-00007-g0006.tif"/>
</fig>
<p>The performance of the models was significantly different in all areas along the visual stream except for pSTS and V3A (<italic>p</italic> &#x02264; 0.05, bootstrapping). This difference was in favor of semantic representations in downstream areas and low-level representations in upstream areas. The largest difference in favor of semantic representations was in TOS (<inline-formula><mml:math id="M52"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>11</mml:mn></mml:math></inline-formula>), OFA (<inline-formula><mml:math id="M53"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>08</mml:mn></mml:math></inline-formula>) and MT&#x0002B; (<inline-formula><mml:math id="M54"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>04</mml:mn></mml:math></inline-formula>), and low-level representations was in V1 (<inline-formula><mml:math id="M55"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>), V2 (<inline-formula><mml:math id="M56"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>07</mml:mn></mml:math></inline-formula>) and V3 (<inline-formula><mml:math id="M57"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>05</mml:mn></mml:math></inline-formula>).</p>
<p>Thirty-nine percent of the voxels preferred either representation by more than the arbitrary threshold of <italic>r</italic> &#x0003D; 0.1. Thirty-four percent of these voxels preferred semantic representations (<inline-formula><mml:math id="M58"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>16</mml:mn></mml:math></inline-formula>), and 66% percent of these voxels preferred low-level representations (<inline-formula><mml:math id="M59"><mml:mi>&#x00394;</mml:mi><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>18</mml:mn></mml:math></inline-formula>).</p>
<p>These results are in line with a large number of earlier work that showed similar dissociations between the representations of the upstream and downstream visual areas (Mishkin et al., <xref ref-type="bibr" rid="B41">1983</xref>; Naselaris et al., <xref ref-type="bibr" rid="B45">2009</xref>; DiCarlo et al., <xref ref-type="bibr" rid="B7">2012</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B22">2015a</xref>).</p></sec>
<sec>
<title>3.3. Analysis of internal representations</title>
<p>Next, to gain insight into the temporal dependencies captured by the <italic>G-100</italic> model, we analyzed its internal representations (Figure <xref ref-type="fig" rid="F7">7</xref>).</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Internal representations of the <italic>G-100</italic> model</bold>. <bold>(A)</bold> Correlation between representational dissimilarity matrices of layer 1 and layer 2 hidden states with each other as well as with those of features and predicted responses. <bold>(B)</bold> Temporal selectivity of layer 1 and layer 2 hidden units. Points indicate lags at which cross-correlations between hidden states and features are highest.</p></caption>
<graphic xlink:href="fncom-11-00007-g0007.tif"/>
</fig>
<p>First, we investigated how the hidden states of the RNN depend on its inputs and output. We constructed representational dissimilarity matrices (RDMs) of the stimulus sequence in the test set at different stages of the processing pipeline and averaged them over subjects (Kriegeskorte et al., <xref ref-type="bibr" rid="B35">2008</xref>). Per feature model, this resulted in one RDM for the features, two RDMs for the layer 1 and layer 2 hidden states and one RDM for the predicted responses. We correlated the upper triangular parts of the RDMs with one another, which resulted in a value indicating how much the hidden states of the RNN were modulated by its inputs and how much they modulated its outputs at a given time point. We found a gradual increase in correlations of the RDMs. That is, the RDMs at each stage were more correlated with those at the next stage compared to those at the previous stages. Importantly, the hidden state RDMs were highly correlated with the predicted response RDMs (<italic>r</italic> &#x0003D; 0.61 and <italic>r</italic> &#x0003D; 0.93 for layers 1 and 2, respectively) but less so with the feature RDMs (<italic>r</italic> &#x0003D; 0.39 and <italic>r</italic> &#x0003D; 0.21 for layers 1 and 2, respectively). This means that while the hidden states of the RNN modulated its outputs at a given time point, they were not modulated by its inputs to the same extent at the same time point. This suggests that a substantial part of the output at a given time-point is not directly related to the input at the same time-point, but instead to previous time-points. That is, the RNN learned to use the input history to make its predictions as expected.</p>
<p>Then, we investigated which time points in the input history were used by the RNN to make its predictions. We cross-correlated each hidden state with each stimulus feature, and averaged the cross-correlations over the features, which resulted in a value indicating how much a hidden state is selective to different time points in the input history. The time point at which this value was at its maximum was taken as the optimal lag of that hidden unit. We found that different hidden units had different optimal lags. The majority of the hidden units had optimal lags up to -20 s, which are likely capturing the hemodynamic factors. However, there was a non-negligible number of hidden units with optimal lags beyond this period, which might be capturing other cognitive/neuronal factors or factors related to stimulus/feature statistics. It should be noted that not all hidden units, in particular those with extensive lags, can be attributed to any of these factors, and their behavior might be induced by model definition or estimation. Furthermore, the optimal lags of the hidden units in the <italic>W2V</italic> based model were on average significantly higher than those in the <italic>GIST</italic> based model (&#x003BC; &#x0003D; &#x02212;9.6 s vs. &#x003BC; &#x0003D; &#x02212;4.9 s, <italic>p</italic> &#x0003C; 0.05, two-sample <italic>t</italic>-test), which might reflect the differences in the statistics of the features that the models are based on. That is, high-level semantic features tend to be more persistent than the low-level structural features across the input sequence. For example, over a given video sequence, distribution of objects in a scene change relatively slowly compared to that of the edges in the scene.</p></sec>
<sec>
<title>3.4. Estimation of voxel-specific HRFs</title>
<p>Traditionally, models have used analytically derived (Friston et al., <xref ref-type="bibr" rid="B13">1998</xref>) or statistically estimated (Dale, <xref ref-type="bibr" rid="B6">1999</xref>; Glover, <xref ref-type="bibr" rid="B16">1999</xref>) HRFs such as the linear models considered here. Estimation of voxel-specific HRFs is an important problem since using the same HRF for all voxels ignores the variability of the hemodynamic response across the brain, which might adversely affect the model performance. Recent developments have focused on the derivation and estimation of more accurate HRFs. For example, Aquino et al. (<xref ref-type="bibr" rid="B2">2014</xref>) has shown that HRFs can be analytically derived from physiology, and Pedregosa et al. (<xref ref-type="bibr" rid="B51">2015</xref>) has shown that HRFs can be efficiently estimated from data. Note that, while the methods for statistically estimating HRFs are particularly suited for use in block designs and event related designs, they are less straightforward to use in continuous designs such as the one considered here.</p>
<p>As demonstrated in the previous subsection, one important advantage of the response models in the RNN family is that they can capture certain temporal dependencies in the data, which might correspond to the HRFs of voxels. Here, we evaluate the voxel-specific HRFs that are obtained by stimulating the <italic>G-100</italic> model with an impulse. We used both feature models in combination with the <italic>G-100</italic> model to estimate the HRFs of the voxels where the performance of any model combination was significantly higher than chance (51% of the voxels, <italic>p</italic> &#x02264; 0.05, Student&#x00027;s <italic>t</italic>-test, Bonferroni correction) (Figure <xref ref-type="fig" rid="F8">8</xref>). The W2V and <italic>G-100</italic> models were used to estimate the HRFs of the voxels where their performance was higher than that of the GIST and <italic>G-100</italic> models, and vice versa.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>Estimation of the hemodynamic response functions</bold>. The <italic>G-100</italic> model was stimulated with an impulse. The impulse response was processed by normalizing its baseline and scale. The result was taken as the HRF. <bold>(A)</bold> Median hemodynamic response functions of all voxels in different areas. Error bands indicate 68% confidence intervals (bootstrapping). Different colors indicate different areas. Dashed line indicates canonical hemodynamic response function. <bold>(B)</bold> Delays of responses and undershoots of all voxels. Black lines indicate canonical delays.</p></caption>
<graphic xlink:href="fncom-11-00007-g0008.tif"/>
</fig>
<p>It was found that the global shape of the estimated HRFs was similar to that of the canonical HRF. However, there was a considerable spread in the estimated delays of responses and the delays of undershoots (median delay of response &#x0003D; 6.57 &#x000B1; 0.02 s, median delay of undershoot &#x0003D; 16.95 &#x000B1; 0.04 s), with the delays of responses being significantly correlated with the delays of undershoots (Pearson&#x00027;s <italic>r</italic> &#x0003D; 0.45, <italic>p</italic> &#x02264; 0.05, Student&#x00027;s <italic>t</italic>-test).</p>
<p>These results demonstrate that RNNs can not only learn (stimulus) feature-response relationships but also can estimate HRFs of voxels, which in turn demonstrate that the nonlinear temporal dynamics that are learned by the RNNs capture biologically relevant temporal dependencies. Furthermore, the variability in the estimated voxel-specific HRFs revealed by the recurrent models might provide a partial explanation of the performance difference between the recurrent and ridge regression models since the ridge regression models use fixed or restricted HRFs, making it difficult for them to take such variability into account.</p></sec></sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>Understanding how the human brain responds to its environment is a key objective in neuroscience. This study has shown that recurrent neural networks are exquisitely capable of capturing how brain responses are induced by sensory stimulation, outperforming established approaches augmented with ridge regression. This increased sensitivity has important consequences for future studies in this area.</p>
<sec>
<title>4.1. Testing hypotheses about brain function</title>
<p>Like any other encoding model, RNN based encoding models can be used to test hypotheses about neural representations (Naselaris et al., <xref ref-type="bibr" rid="B44">2011</xref>). That is, they can be used to test whether a particular feature model outperforms alternative feature models when it comes to explaining observed data. As such, we have shown that a low-level visual feature model explains responses in upstream visual areas well, whereas a high-level semantic model explains responses in downstream visual areas well, conforming to the well established early and high-level subdivision of the visual cortex (Mishkin et al., <xref ref-type="bibr" rid="B41">1983</xref>; Naselaris et al., <xref ref-type="bibr" rid="B45">2009</xref>; DiCarlo et al., <xref ref-type="bibr" rid="B7">2012</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B22">2015a</xref>).</p>
<p>Furthermore, RNN-based encoding models can also be used to test hypotheses about the temporal dependencies between features and responses. For example, by constraining the temporal memory capacities of the RNN units, one can identify the optimal scale of the temporal dependencies that different brain regions are selective to.</p>
<p>Here, we used RNNs as response models in an encoding framework. That is, they were used to predict responses to features that were extracted from stimuli with separate feature models. However, use cases of RNNs are not limited to this setting. For example, RNN models can be used as feature models instead of response models in the encoding framework. Like CNNs, RNNs are being used to solve various problems in fields ranging from computer vision (Gregor et al., <xref ref-type="bibr" rid="B19">2015</xref>) to computational linguistics (Zaremba et al., <xref ref-type="bibr" rid="B60">2014</xref>). Internal representations of task-optimized CNNs were shown to correspond to neural representations in different brain regions (Kriegeskorte, <xref ref-type="bibr" rid="B34">2015</xref>; Yamins and DiCarlo, <xref ref-type="bibr" rid="B58">2016b</xref>). It would be interesting to see if the internal representations of task-based RNNs have similar correlates in the brain. For example, it was recently shown that RNNs develop representations that are reminiscent of their biological counterparts when they learn to solve a spatial navigation task (Kanitscheider and Fiete, <xref ref-type="bibr" rid="B29">2016</xref>). Such representations may turn out to be predictive of brain responses recorded during similar tasks.</p></sec>
<sec>
<title>4.2. Limitations of RNNs for investigating neural representations</title>
<p>RNNs can process arbitrary input sequences in theory. However, they have an important limitation in practice. Like any other contemporary neural network architecture, typical RNN architectures have a very large number of free parameters. Therefore, a very large amount of training data is required for accurately estimating RNN models without overfitting. While there are several methods to combat overfitting in RNNs like different variants of dropout (Hinton et al., <xref ref-type="bibr" rid="B26">2012</xref>; Zaremba et al., <xref ref-type="bibr" rid="B60">2014</xref>; Semeniuta et al., <xref ref-type="bibr" rid="B53">2016</xref>), it is still an important issue to which particular attention needs to be paid.</p>
<p>This can also be the reason why gated recurrent unit architectures were shown to outperform LSTM architectures. That is, the performance difference between the two types of architectures is likely to be caused by difficulties in model estimation in the current data regime rather than one architecture being better suited to the problem at hand than the other.</p>
<p>This also means that RNN models will face difficulties when trying to predict responses to very high-dimensional stimulus features such as the internal representations of convolutional neural networks which range from thousands to hundreds of thousands dimensions. For such features, dimensionality reduction techniques can be utilized for reducing the feature dimensionality to a range that can be handled with RNNs in scenarios with either insufficient computational resources or training data.</p>
<p>Linear response models have been used with great success in the past for gaining insights into neural representations. They have been particularly useful since linear mappings make it easy to interpret factors driving response predictions. One might argue that the nonlinearities introduced by RNNs make the interpretation harder compared to linear mappings. However, the relative difficulty of interpretation is a direct consequence of more accurate response predictions, which can be beneficial in certain scenarios. For example, it was shown that systematic nonlinearities that are not taken into account by linear mappings can lead to less accurate response predictions and tuning functions of V1 voxels (Vu et al., <xref ref-type="bibr" rid="B55">2011</xref>). Furthermore, since more accurate response predictions lead to higher statistical power, the improved model fit afforded by RNNs might make detection of more subtle effects possible. Moreover, when the goal is to compare different feature models, such as the GIST and W2V models used here, maximizing explained variance might become the main criterion of interest. That is, linear models might lead to misleading performance differences between the encoding models in the cases where their assumptions about the underlying temporal dynamics do not hold. In such cases, it would be particularly important to fit the response models as accurately as possible as to ensure that the observed performance difference between two encoding models is driven by their underlying feature representations and not suboptimal model fits. Therefore, RNNs will be particularly useful in settings where temporal dynamics are of primary interest. Finally, combining the present work with recent developments on understanding RNN representations (Karpathy et al., <xref ref-type="bibr" rid="B30">2015</xref>) is expected to improve the interpretations of factors driving response predictions.</p></sec>
<sec>
<title>4.3. Capturing temporal dependencies</title>
<p>RNNs can use their internal memories to capture the temporal dependencies in data. In the context of modeling the dynamics of brain activity in response to naturalistic stimuli, these dependencies can be caused by factors such as neurovascular coupling or stimulus-induced cognitive processes. By providing an RNN with an impulse on the input side, it was shown that, effectively, the RNN learns to represent voxel-specific hemodynamic responses. Importantly, the RNNs allowed us to estimate these HRFs from data collected under a continuous design. To the best of our knowledge this is the first time it has been shown that this is possible in practice. By analyzing the internal representations of an RNN, it was also shown that the RNN learns to represent information from stimulus features at past time points beyond the range of neurovascular coupling. Hence, the predictions of observed brain responses are likely induced by stimulus-related, cognitive or neural factors on top of the hemodynamic response.</p></sec>
<sec>
<title>4.4. Isolating neural and hemodynamic components</title>
<p>In the introduction, we motivated the use of RNNs as a generic parameterization of any non-linear convolution of stimulus features to hemodynamic responses. Crucially, this could cover both neuronal and hemodynamic convolution. In other words, our black box approach allows for a neuronal convolution of stimulus feature input to produce a neuronal response that is subsequently convolved by hemodynamic operators to produce the observed outcome. This facility may explain the increased cross-validation accuracy observed in our analyses (over and above more restricted models of hemodynamic convolution). In other words, the procedure detailed in this paper can accommodate neuronal convolutions that may be precluded in conventional models.</p>
<p>The cost of this flexibility is that we cannot separate the neuronal and hemodynamic components of the convolution. This follows from the fact that the RNN parameterization does not make an explicit distinction between neuronal and hemodynamic processes. To properly understand the relative contribution of these formally distinct processes, one would have to use a generative model approach with biologically plausible prior constraints on the neuronal and hemodynamic parts of the convolution. This is precisely the objective of dynamic causal modeling that equips a system of neuronal dynamics (and implicit recurrent connectivity) with a hemodynamic model based upon known biophysics (Friston et al., <xref ref-type="bibr" rid="B11">2003</xref>). It would therefore be interesting to examine the form of RNNs in relation to existing dynamic causal models that have a similar architecture.</p></sec>
<sec>
<title>4.5. Conclusions</title>
<p>We have shown for the first time that RNNs can be used to predict how the human brain processes sensory information. Whereas classical connectionist research has focused on the use of RNNs as models of cognitive processing (Elman, <xref ref-type="bibr" rid="B9">1993</xref>), the present work has shown that RNNs can also be used to probe the hemodynamic correlates of ongoing cognitive processes induced by dynamically changing naturalistic sensory stimuli. The ability of RNNs to learn about long-range temporal dependencies provides the flexibility to couple ongoing sensory stimuli that induce various cognitive processes with delayed measurements of brain activity that depend on such processes. This end-to-end training approach can be applied to any neuroscientific experiment in which sensory inputs are coupled to observed neural responses.</p></sec>
<sec>
<title>4.6. Data sharing</title>
<p>The data set that was used in this paper was originally published in Nishimoto et al. (<xref ref-type="bibr" rid="B47">2011</xref>) and is available at Nishimoto et al. (<xref ref-type="bibr" rid="B48">2014</xref>). The code that was used in this paper is provided at <ext-link ext-link-type="uri" xlink:href="http://www.ccnlab.net/">http://www.ccnlab.net/</ext-link>.</p></sec></sec>
<sec id="s5">
<title>Ethics statement</title>
<p>Human fMRI data set that was used in this study was taken from the public data sharing repository <ext-link ext-link-type="uri" xlink:href="http://crcns.org/">http://crcns.org/</ext-link>. The original study was approved by the local ethics committee (Committee for the Protection of Human Subjects at University of California, Berkeley).</p></sec>
<sec id="s6">
<title>Author contributions</title>
<p>UG and MvG designed research; UG performed research; UG and MvG contributed unpublished reagents/analytic tools; UG analyzed data; UG and MvG wrote the paper.</p></sec>
<sec id="s7">
<title>Funding</title>
<p>This research was supported by VIDI grant number 639.072.513 of the Netherlands Organization for Scientific Research (NWO).</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Agrawal</surname> <given-names>P.</given-names></name> <name><surname>Stansbury</surname> <given-names>D.</given-names></name> <name><surname>Malik</surname> <given-names>J.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2014</year>). <article-title>Pixels to voxels: modeling visual representation in the human brain</article-title>. arXiv:1407.5104 [q-bio.NC].</citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aquino</surname> <given-names>K. M.</given-names></name> <name><surname>Robinson</surname> <given-names>P. A.</given-names></name> <name><surname>Drysdale</surname> <given-names>P. M.</given-names></name></person-group> (<year>2014</year>). <article-title>Spatiotemporal hemodynamic response functions derived from physiology</article-title>. <source>J. Theor. Biol.</source> <volume>347</volume>, <fpage>118</fpage>&#x02013;<lpage>136</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2013.12.027</pub-id><pub-id pub-id-type="pmid">24398024</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cadieu</surname> <given-names>C. F.</given-names></name> <name><surname>Hong</surname> <given-names>H.</given-names></name> <name><surname>Yamins</surname> <given-names>D. L.</given-names></name> <name><surname>Pinto</surname> <given-names>N.</given-names></name> <name><surname>Ardila</surname> <given-names>D.</given-names></name> <name><surname>Solomon</surname> <given-names>E. A.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Deep neural networks rival the representation of primate IT cortex for core visual object recognition</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003963</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003963</pub-id><pub-id pub-id-type="pmid">25521294</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>van Merrienboer</surname> <given-names>B.</given-names></name> <name><surname>Gulcehre</surname> <given-names>C.</given-names></name> <name><surname>Bahdanau</surname> <given-names>D.</given-names></name> <name><surname>Bougares</surname> <given-names>F.</given-names></name> <name><surname>Schwenk</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Learning phrase representations using RNN encoder-decoder for statistical machine translation</article-title>. arXiv:1406.1078 [cs.CL].</citation></ref>
<ref id="B5">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Cichy</surname> <given-names>R. M.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Pantazis</surname> <given-names>D.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep neural networks predict hierarchical spatio-temporal cortical dynamics of human visual object recognition</article-title>. arXiv:1601.02970 [cs.CV].</citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dale</surname> <given-names>A. M.</given-names></name></person-group> (<year>1999</year>). <article-title>Optimal experimental design for event-related fMRI</article-title>. <source>Hum. Brain Mapp.</source> <volume>8</volume>, <fpage>109</fpage>&#x02013;<lpage>114</lpage>. <pub-id pub-id-type="pmid">10524601</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name> <name><surname>Zoccolan</surname> <given-names>D.</given-names></name> <name><surname>Rust</surname> <given-names>N. C.</given-names></name></person-group> (<year>2012</year>). <article-title>How does the brain solve visual object recognition?</article-title> <source>Neuron</source> <volume>73</volume>, <fpage>415</fpage>&#x02013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.01.010</pub-id><pub-id pub-id-type="pmid">22325196</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eickenberg</surname> <given-names>M.</given-names></name> <name><surname>Gramfort</surname> <given-names>A.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>Seeing it all: convolutional network layers map the function of the human visual system</article-title>. <source>NeuroImage</source>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2016.10.001</pub-id> [Epub ahead of print]. <pub-id pub-id-type="pmid">27777172</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name></person-group> (<year>1993</year>). <article-title>Learning and development in neural networks - the importance of prior experience</article-title>. <source>Cognition</source> <volume>48</volume>, <fpage>71</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(93)90058-4</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Felsen</surname> <given-names>G.</given-names></name> <name><surname>Dan</surname> <given-names>Y.</given-names></name></person-group> (<year>2005</year>). <article-title>A natural approach to studying vision</article-title>. <source>Nat. Neurosci.</source> <volume>8</volume>, <fpage>1643</fpage>&#x02013;<lpage>1646</lpage>. <pub-id pub-id-type="doi">10.1038/nn1608</pub-id><pub-id pub-id-type="pmid">16306891</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Harrison</surname> <given-names>L.</given-names></name> <name><surname>Penny</surname> <given-names>W.</given-names></name></person-group> (<year>2003</year>). <article-title>Dynamic causal modelling</article-title>. <source>Neuroimage</source> <volume>19</volume>, <fpage>1273</fpage>&#x02013;<lpage>1302</lpage>. <pub-id pub-id-type="doi">10.1016/S1053-8119(03)00202-7</pub-id><pub-id pub-id-type="pmid">12948688</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Holmes</surname> <given-names>A. P.</given-names></name> <name><surname>Worsley</surname> <given-names>K. J.</given-names></name> <name><surname>Poline</surname> <given-names>J.-P.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name> <name><surname>Frackowiak</surname> <given-names>R. S. J.</given-names></name></person-group> (<year>1994</year>). <article-title>Statistical parametric maps in functional imaging: a general linear approach</article-title>. <source>Hum. Brain Mapp.</source> <volume>2</volume>, <fpage>189</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.460020402</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Josephs</surname> <given-names>O.</given-names></name> <name><surname>Rees</surname> <given-names>G.</given-names></name> <name><surname>Turner</surname> <given-names>R.</given-names></name></person-group> (<year>1998</year>). <article-title>Nonlinear event-related responses in fMRI</article-title>. <source>Magn. Reson. Med.</source> <volume>39</volume>, <fpage>41</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1002/mrm.1910390109</pub-id><pub-id pub-id-type="pmid">9438436</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Mechelli</surname> <given-names>A.</given-names></name> <name><surname>Turner</surname> <given-names>R.</given-names></name> <name><surname>Price</surname> <given-names>C. J.</given-names></name></person-group> (<year>2000</year>). <article-title>Nonlinear responses in fMRI: the Balloon model, Volterra kernels, and other hemodynamics</article-title>. <source>Neuroimage</source> <volume>12</volume>, <fpage>466</fpage>&#x02013;<lpage>477</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.2000.0630</pub-id><pub-id pub-id-type="pmid">10988040</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fyshe</surname> <given-names>A.</given-names></name> <name><surname>Talukdar</surname> <given-names>P.</given-names></name> <name><surname>Murphy</surname> <given-names>B.</given-names></name> <name><surname>Mitchell</surname> <given-names>T.</given-names></name></person-group> (<year>2013</year>). <article-title>Documents and dependencies: an exploration of vector space models for semantic composition</article-title>, in <source>Documents and dependencies: an exploration of vector space models for semantic composition</source>, (<publisher-loc>Sofia</publisher-loc>).</citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glover</surname> <given-names>G. H.</given-names></name></person-group> (<year>1999</year>). <article-title>Deconvolution of impulse response in event-related BOLD fMRI</article-title>. <source>NeuroImage</source> <volume>9</volume>, <fpage>416</fpage>&#x02013;<lpage>429</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.1998.0419</pub-id><pub-id pub-id-type="pmid">10191170</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Graves</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Generating sequences with recurrent neural networks</article-title>. arXiv:1308.0850 [cs.NE].</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Liwicki</surname> <given-names>M.</given-names></name> <name><surname>Fern&#x000E1;ndez</surname> <given-names>S.</given-names></name> <name><surname>Bertolami</surname> <given-names>R.</given-names></name> <name><surname>Bunke</surname> <given-names>H.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>A novel connectionist system for unconstrained handwriting recognition</article-title>. <source>IEEE Trans. Patt. Anal. Mach. Intell.</source> <volume>31</volume>, <fpage>855</fpage>&#x02013;<lpage>868</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2008.137</pub-id><pub-id pub-id-type="pmid">19299860</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Gregor</surname> <given-names>K.</given-names></name> <name><surname>Danihelka</surname> <given-names>I.</given-names></name> <name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Jimenez Rezende</surname> <given-names>D.</given-names></name> <name><surname>Wierstra</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>DRAW: A recurrent neural network for image generation</article-title>. arXiv:1502.04623 [cs.CV].</citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Groen</surname> <given-names>I. I.</given-names></name> <name><surname>Ghebreab</surname> <given-names>S.</given-names></name> <name><surname>Prins</surname> <given-names>H.</given-names></name> <name><surname>Lamme</surname> <given-names>V. A.</given-names></name> <name><surname>Scholte</surname> <given-names>H. S.</given-names></name></person-group> (<year>2013</year>). <article-title>From image statistics to scene gist: evoked neural activity reveals transition from low-level natural image structure to scene category</article-title>. <source>J. Neurosci.</source> <volume>33</volume>, <fpage>18814</fpage>&#x02013;<lpage>18824</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3128-13.2013</pub-id><pub-id pub-id-type="pmid">24285888</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A.</given-names></name></person-group> (<year>2014</year>). <article-title>Unsupervised feature learning improves prediction of human brain activity in response to natural images</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003724</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003724</pub-id><pub-id pub-id-type="pmid">25101625</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2015a</year>). <article-title>Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>10005</fpage>&#x02013;<lpage>10014</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5023-14.2015</pub-id><pub-id pub-id-type="pmid">26157000</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2015b</year>). <article-title>Increasingly complex representations of natural movies across the dorsal stream are shared between subjects</article-title>. <source>NeuroImage</source>. <volume>145</volume>, <fpage>329</fpage>&#x02013;<lpage>336</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2015.12.036</pub-id><pub-id pub-id-type="pmid">26724778</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2015c</year>). <article-title>Semantic vector space models predict neural responses to complex visual stimuli</article-title>. arXiv:1510.04738 [q-bio.NC].</citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hansen</surname> <given-names>K. A.</given-names></name> <name><surname>David</surname> <given-names>S. V.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2004</year>). <article-title>Parametric reverse correlation reveals spatial linearity of retinotopic human V1 BOLD response</article-title>. <source>NeuroImage</source> <volume>23</volume>, <fpage>233</fpage>&#x02013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2004.05.012</pub-id><pub-id pub-id-type="pmid">15325370</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Srivastava</surname> <given-names>N.</given-names></name> <name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name></person-group> (<year>2012</year>). <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>. arXiv:1207.0580 [cs.NE].</citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput.</source> <volume>9</volume>, <fpage>1735</fpage>&#x02013;<lpage>1780</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huth</surname> <given-names>A. G.</given-names></name> <name><surname>Nishimoto</surname> <given-names>S.</given-names></name> <name><surname>Vu</surname> <given-names>A. T.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2012</year>). <article-title>A continuous semantic space describes the representation of thousands of object and action categories across the human brain</article-title>. <source>Neuron</source> <volume>76</volume>, <fpage>1210</fpage>&#x02013;<lpage>1224</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2012.10.014</pub-id><pub-id pub-id-type="pmid">23259955</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kanitscheider</surname> <given-names>I.</given-names></name> <name><surname>Fiete</surname> <given-names>I.</given-names></name></person-group> (<year>2016</year>). <article-title>Training recurrent networks to generate hypotheses about how the brain solves hard navigation problems</article-title>. arXiv:1609.09059 [q-bio.nc].</citation></ref>
<ref id="B30">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Karpathy</surname> <given-names>A.</given-names></name> <name><surname>Johnson</surname> <given-names>J.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Visualizing and understanding recurrent networks</article-title>. arXiv:1506.02078 [cs.LG].</citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kay</surname> <given-names>K. N.</given-names></name> <name><surname>Winawer</surname> <given-names>J.</given-names></name> <name><surname>Mezer</surname> <given-names>A.</given-names></name> <name><surname>Wandell</surname> <given-names>B. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Compressive spatial summation in human visual cortex</article-title>. <source>J. Neurophysiol.</source> <volume>110</volume>, <fpage>481</fpage>&#x02013;<lpage>494</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00105.2013</pub-id><pub-id pub-id-type="pmid">23615546</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khaligh-Razavi</surname> <given-names>S.-M.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2014</year>). <article-title>Deep supervised, but not unsupervised, models may explain IT cortical representation</article-title>. <source>PLoS Comput. Biol.</source> <volume>10</volume>:<fpage>e1003915</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1003915</pub-id><pub-id pub-id-type="pmid">25375136</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization</article-title>. arXiv:1412.6980 [cs.LG].</citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks: a new framework for modeling biological vision and brain information processing</article-title>. <source>Ann. Rev. Vis. Sci.</source> <volume>1</volume>, <fpage>417</fpage>&#x02013;<lpage>446</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-vision-082114-035447</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Mur</surname> <given-names>M.</given-names></name> <name><surname>Bandettini</surname> <given-names>P.</given-names></name></person-group> (<year>2008</year>). <article-title>Representational similarity analysis - connecting the branches of systems neuroscience</article-title>. <source>Front. Syst. Neurosci.</source> <volume>2</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.3389/neuro.06.004.2008</pub-id><pub-id pub-id-type="pmid">19104670</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leeds</surname> <given-names>D. D.</given-names></name> <name><surname>Seibert</surname> <given-names>D. A.</given-names></name> <name><surname>Pyles</surname> <given-names>J. A.</given-names></name> <name><surname>Tarr</surname> <given-names>M. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Comparing visual representations across human fMRI and computational vision</article-title>. <source>J. Vis.</source> <volume>13</volume>, <fpage>25</fpage>. <pub-id pub-id-type="doi">10.1167/13.13.25</pub-id><pub-id pub-id-type="pmid">24273227</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Logothetis</surname> <given-names>N. K.</given-names></name> <name><surname>Wandell</surname> <given-names>B. A.</given-names></name></person-group> (<year>2004</year>). <article-title>Interpreting the BOLD signal</article-title>. <source>Ann. Rev. Physiol.</source> <volume>66</volume>, <fpage>735</fpage>&#x02013;<lpage>769</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.physiol.66.082602.092845</pub-id><pub-id pub-id-type="pmid">14977420</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013a</year>). <article-title>Efficient estimation of word representations in vector space</article-title>. arXiv:1301.3781 [cs.CL].</citation></ref>
<ref id="B39">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013b</year>). <article-title>Distributed representations of words and phrases and their compositionality</article-title>. arXiv:1310.4546 [cs.CL].</citation></ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Yih</surname> <given-names>W.-T.</given-names></name> <name><surname>Zweig</surname> <given-names>G.</given-names></name></person-group> (<year>2013c</year>). <article-title>Linguistic regularities in continuous space word representations</article-title>, in <source>Linguistic regularities in continuous space word representations</source>, (<publisher-loc>Atlanta, GA</publisher-loc>).</citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mishkin</surname> <given-names>M.</given-names></name> <name><surname>Ungerleider</surname> <given-names>L. G.</given-names></name> <name><surname>Macko</surname> <given-names>K. A.</given-names></name></person-group> (<year>1983</year>). <article-title>Object vision and spatial vision: two cortical pathways</article-title>. <source>Trends Neurosci.</source> <volume>6</volume>, <fpage>414</fpage>&#x02013;<lpage>417</lpage>. <pub-id pub-id-type="doi">10.1016/0166-2236(83)90190-X</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mitchell</surname> <given-names>T. M.</given-names></name> <name><surname>Shinkareva</surname> <given-names>S. V.</given-names></name> <name><surname>Carlson</surname> <given-names>A.</given-names></name> <name><surname>Chang</surname> <given-names>K. M.</given-names></name> <name><surname>Malave</surname> <given-names>V. L.</given-names></name> <name><surname>Mason</surname> <given-names>R. A.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Predicting human brain activity associated with the meanings of nouns</article-title>. <source>Science</source> <volume>320</volume>, <fpage>1191</fpage>&#x02013;<lpage>1195</lpage>. <pub-id pub-id-type="doi">10.1126/science.1152876</pub-id><pub-id pub-id-type="pmid">18511683</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Murphy</surname> <given-names>B.</given-names></name> <name><surname>Talukdar</surname> <given-names>P.</given-names></name> <name><surname>Mitchell</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>Selecting corpus-semantic models for neurolinguistic decoding</article-title>, in <source>Proceedings of First Joint Conference on Lexical and Computational Semantics</source> (<publisher-loc>Montr&#x000E9;al, QC</publisher-loc>).</citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naselaris</surname> <given-names>T.</given-names></name> <name><surname>Kay</surname> <given-names>K. N.</given-names></name> <name><surname>Nishimoto</surname> <given-names>S.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2011</year>). <article-title>Encoding and decoding in fMRI</article-title>. <source>NeuroImage</source> <volume>56</volume>, <fpage>400</fpage>&#x02013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.07.073</pub-id><pub-id pub-id-type="pmid">20691790</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naselaris</surname> <given-names>T.</given-names></name> <name><surname>Prenger</surname> <given-names>R. J.</given-names></name> <name><surname>Kay</surname> <given-names>K. N.</given-names></name> <name><surname>Oliver</surname> <given-names>M.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2009</year>). <article-title>Bayesian reconstruction of natural images from human brain activity</article-title>. <source>Neuron</source> <volume>63</volume>, <fpage>902</fpage>&#x02013;<lpage>915</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2009.09.006</pub-id><pub-id pub-id-type="pmid">19778517</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nishida</surname> <given-names>S.</given-names></name> <name><surname>Huth</surname> <given-names>A.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name> <name><surname>Nishimoto</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Word statistics in large-scale texts explain the human cortical semantic representation of objects, actions, and impressions</article-title>, in <source>The 45th Annual Meeting of the Society for Neuroscience</source> (<publisher-loc>Chicago, IL</publisher-loc>).</citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nishimoto</surname> <given-names>S.</given-names></name> <name><surname>Vu</surname> <given-names>A. T.</given-names></name> <name><surname>Naselaris</surname> <given-names>T.</given-names></name> <name><surname>Benjamini</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>B.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2011</year>). <article-title>Reconstructing visual experiences from brain activity evoked by natural movies</article-title>. <source>Curr. Biol.</source> <volume>21</volume>, <fpage>1641</fpage>&#x02013;<lpage>1646</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2011.08.031</pub-id><pub-id pub-id-type="pmid">21945275</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Nishimoto</surname> <given-names>S.</given-names></name> <name><surname>Vu</surname> <given-names>A. T.</given-names></name> <name><surname>Naselaris</surname> <given-names>T.</given-names></name> <name><surname>Benjamini</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>B.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L</given-names></name></person-group> (<year>2014</year>). <source>Gallant Lab Natural Movie 4T fMRI Data</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://CRCNS.org">http://CRCNS.org</ext-link></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Norris</surname> <given-names>D. G.</given-names></name></person-group> (<year>2006</year>). <article-title>Principles of magnetic resonance assessment of brain function</article-title>. <source>J. Magn. Reson. Imaging</source> <volume>23</volume>, <fpage>794</fpage>&#x02013;<lpage>807</lpage>. <pub-id pub-id-type="doi">10.1002/jmri.20587</pub-id><pub-id pub-id-type="pmid">16649206</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2001</year>). <article-title>Modeling the shape of the scene: a holistic representation of the spatial envelope</article-title>. <source>Int. J. Comput. Vis.</source> <volume>42</volume>, <fpage>145</fpage>&#x02013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1023/A:1011139631724</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedregosa</surname> <given-names>F.</given-names></name> <name><surname>Eickenberg</surname> <given-names>M.</given-names></name> <name><surname>Ciuciu</surname> <given-names>P.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Gramfort</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Data-driven HRF estimation for encoding and decoding models</article-title>. <source>NeuroImage</source> <volume>104</volume>, <fpage>209</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2014.09.060</pub-id><pub-id pub-id-type="pmid">25304775</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Sak</surname> <given-names>H.</given-names></name> <name><surname>Senior</surname> <given-names>A.</given-names></name> <name><surname>Beaufays</surname> <given-names>F.</given-names></name></person-group> (<year>2014</year>). <article-title>Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition</article-title>. arXiv:1402.1128 [cs.NE].</citation></ref>
<ref id="B53">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Semeniuta</surname> <given-names>S.</given-names></name> <name><surname>Severyn</surname> <given-names>A.</given-names></name> <name><surname>Barth</surname> <given-names>E.</given-names></name></person-group> (<year>2016</year>). <article-title>Recurrent dropout without memory loss</article-title>. arXiv:1603.05118 [cs.CL].</citation></ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Martens</surname> <given-names>J.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <article-title>Generating text with recurrent neural networks</article-title>, in <source>Proceedings of the 28th International Conference on Machine Learning</source> (<publisher-loc>Bellevue, WA</publisher-loc>).</citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vu</surname> <given-names>V. Q.</given-names></name> <name><surname>Ravikumar</surname> <given-names>P.</given-names></name> <name><surname>Naselaris</surname> <given-names>T.</given-names></name> <name><surname>Kay</surname> <given-names>K. N.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name> <name><surname>Yu</surname> <given-names>B.</given-names></name></person-group> (<year>2011</year>). <article-title>Encoding and decoding v1 fMRI responses to natural images with sparse nonparametric models</article-title>. <source>Ann. Appl. Stat.</source> <volume>5</volume>, <fpage>1159</fpage>&#x02013;<lpage>1182</lpage>. <pub-id pub-id-type="doi">10.1214/11-AOAS476</pub-id><pub-id pub-id-type="pmid">22523529</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wray</surname> <given-names>J.</given-names></name> <name><surname>Green</surname> <given-names>G. G. R.</given-names></name></person-group> (<year>1994</year>). <article-title>Calculation of the Volterra kernels of non-linear dynamic systems using an artificial neural network</article-title>. <source>Biol. Cybern.</source> <volume>71</volume>, <fpage>187</fpage>&#x02013;<lpage>195</lpage>. <pub-id pub-id-type="doi">10.1007/BF00202758</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L. K.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2016a</year>). <article-title>Eight open questions in the computational modeling of higher sensory cortex</article-title>. <source>Curr. Opin. Neurobiol.</source> <volume>37</volume>, <fpage>114</fpage>&#x02013;<lpage>120</lpage>. <pub-id pub-id-type="doi">10.1016/j.conb.2016.02.001</pub-id><pub-id pub-id-type="pmid">26921828</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L. K.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2016b</year>). <article-title>Using goal-driven deep learning models to understand sensory cortex</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>356</fpage>&#x02013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4244</pub-id><pub-id pub-id-type="pmid">26906502</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L. K.</given-names></name> <name><surname>Hong</surname> <given-names>H.</given-names></name> <name><surname>Cadieu</surname> <given-names>C. F.</given-names></name> <name><surname>Solomon</surname> <given-names>E. A.</given-names></name> <name><surname>Seibert</surname> <given-names>D.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2014</year>). <article-title>Performance-optimized hierarchical models predict neural responses in higher visual cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>111</volume>, <fpage>8619</fpage>&#x02013;<lpage>8624</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1403112111</pub-id><pub-id pub-id-type="pmid">24812127</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Zaremba</surname> <given-names>W.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name></person-group> (<year>2014</year>). <article-title>Recurrent neural network regularization</article-title>. arXiv:1409.2329 [cs.NE].</citation></ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>The cross-validated correlation coefficient automatically penalizes for model complexity and therefore can be used as a proxy for model evidence.</p></fn>
</fn-group>
</back>
</article>