<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Hum. Neurosci.</journal-id>
<journal-title>Frontiers in Human Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Hum. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5161</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnhum.2017.00075</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Variation in Event-Related Potentials by State Transitions</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Higashi</surname> <given-names>Hiroshi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/257724/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Minami</surname> <given-names>Tetsuto</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/37501/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Nakauchi</surname> <given-names>Shigeki</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Computer Science and Engineering, Toyohashi University of Technology</institution> <country>Aichi, Japan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Electronics-Inspired Interdisciplinary Research Institute, Toyohashi University of Technology</institution> <country>Aichi, Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Juliana Yordanova, Bulgarian Academy of Sciences, Bulgaria</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Rolf Verleger, University of L&#x000FC;beck, Germany; M&#x000E1;rk Moln&#x000E1;r, Hungarian Academy of Sciences, Hungary</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Hiroshi Higashi <email>higashi&#x00040;tut.jp</email></p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>02</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>11</volume>
<elocation-id>75</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>11</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>02</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Higashi, Minami and Nakauchi.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Higashi, Minami and Nakauchi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>The probability of an event&#x00027;s occurrence affects event-related potentials (ERPs) on electroencephalograms. The relation between probability and potentials has been discussed by using a quantity called surprise that represents the self-information that humans receive from the event. Previous studies have estimated surprise based on the probability distribution in a stationary state. Our hypothesis is that state transitions also play an important role in the estimation of surprise. In this study, we compare the effects of surprise on the ERPs based on two models that generate an event sequence: a model of a stationary state and a model with state transitions. To compare these effects, we generate the event sequences with Markov chains to avoid a situation that the state transition probability converges with the stationary probability by the accumulation of the event observations. Our trial-by-trial model-based analysis showed that the stationary probability better explains the P3b component and the state transition probability better explains the P3a component. The effect on P3a suggests that the internal model, which is constantly and automatically generated by the human brain to estimate the probability distribution of the events, approximates the model with state transitions because Bayesian surprise, which represents the degree of updating of the internal model, is highly reflected in P3a. The global effect reflected in P3b, however, may not be related to the internal model because P3b depends on the stationary probability distribution. The results suggest that an internal model can represent state transitions and the global effect is generated by a different mechanism than the one for forming the internal model.</p></abstract>
<kwd-group>
<kwd>Event-Related Potentials (ERPs)</kwd>
<kwd>Electroencephalography (EEG)</kwd>
<kwd>predictive surprise</kwd>
<kwd>model-based analysis</kwd>
<kwd>single-trial analysis</kwd>
</kwd-group>
<contract-sponsor id="cn001">Japan Society for the Promotion of Science<named-content content-type="fundref-id">10.13039/501100001691</named-content></contract-sponsor>
<counts>
<fig-count count="10"/>
<table-count count="1"/>
<equation-count count="18"/>
<ref-count count="49"/>
<page-count count="11"/>
<word-count count="7255"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Humans make predictions by using prior information (Doya et al., <xref ref-type="bibr" rid="B12">2007</xref>; Friston, <xref ref-type="bibr" rid="B16">2008</xref>), and the prior information is derived from what humans have experienced. The prediction of an event&#x00027;s occurrence is equivalent to the estimation of the generative model for the event (Robert, <xref ref-type="bibr" rid="B40">2007</xref>). However, the manner in which one utilizes that experience in estimating the generative model remains unclear.</p>
<p>One approach for revealing how experience affects human prediction is the observation of event-related potentials (ERPs). ERPs are the measured brain responses for a specific internal or external event (Clark, <xref ref-type="bibr" rid="B4">2013</xref>). Some ERP components observed on electroencephalograms (EEGs) are affected by the probability of the occurrence of the event (Sutton et al., <xref ref-type="bibr" rid="B49">1965</xref>; Squires et al., <xref ref-type="bibr" rid="B47">1977</xref>; Picton, <xref ref-type="bibr" rid="B37">1992</xref>). In particular, a slow variation observed about 300 ms after the event on the EEG potentials is called P300, and its peak amplitude depends on the probability of the event&#x00027;s occurrence (Duncan-Johnson and Donchin, <xref ref-type="bibr" rid="B13">1977</xref>). Such ERP components are considered to reflect the process for predicting events and are widely used as a medium for analyzing human cognition (Horovitz et al., <xref ref-type="bibr" rid="B21">2002</xref>; Sanmiguel et al., <xref ref-type="bibr" rid="B42">2013</xref>).</p>
<p>The relation between these ERP components and probability has been discussed by using a quantity called <italic>the degree of surprise</italic> or simply <italic>surprise</italic> (Ostwald et al., <xref ref-type="bibr" rid="B35">2012</xref>). One concept of surprise is called <italic>predictive surprise</italic> that represents the subjective self-information, information content, or surprisal (Shannon, <xref ref-type="bibr" rid="B44">1948</xref>) that an observer receives from an observed event (Donchin and Coles, <xref ref-type="bibr" rid="B11">1988</xref>). Recently, Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>) and Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>) addressed the question of which factors in a preceding stimulus sequence affect predictive surprise to a present stimulus by investigating the relation between the stimulus sequence and P300 properties. To identify these factors, Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>) and Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>) used regression models in which the input was a stimulus history and the output was the P300 amplitude. Their approach with these regression models is called <italic>the model-based approach</italic>. Assuming that the observed brain activities reflect the prediction process, the model-based approach can confirm which factors human prediction depends on by finding a model that accurately estimates the amplitude of P300. Their results suggest that surprise estimated by the integration of three factors (long-term history, short-term history, and alternating expectations of the stimulus sequence) adequately predicts the P300 amplitude (Squires et al., <xref ref-type="bibr" rid="B48">1976</xref>; Kolossa et al., <xref ref-type="bibr" rid="B25">2012</xref>).</p>
<p>The other concept of surprise, proposed by Baldi and Itti (<xref ref-type="bibr" rid="B1">2010</xref>), is called <italic>Bayesian surprise</italic> that represents the degree of updating in the beliefs of an observer who experiences a new event. Recent studies (Kolossa et al., <xref ref-type="bibr" rid="B26">2015</xref>; Seer et al., <xref ref-type="bibr" rid="B43">2016</xref>) showed that predictive and Bayesian surprises affect different subcomponents of P300 called P3a and P3b (Polich, <xref ref-type="bibr" rid="B38">2007</xref>). Predictive surprise better explains P3b, which has a long latency among the subcomponents. However, Bayesian surprise better explains P3a, which has a short latency. The results suggest that the subcomponents reflect distinct neural mechanisms for prediction.</p>
<p>To reveal the relation between ERPs and surprise, theoretical frameworks, such as the context-updating model (Donchin, <xref ref-type="bibr" rid="B10">1981</xref>; Donchin and Coles, <xref ref-type="bibr" rid="B11">1988</xref>; Polich, <xref ref-type="bibr" rid="B38">2007</xref>), predictive coding (Friston, <xref ref-type="bibr" rid="B15">2002</xref>; Spratling, <xref ref-type="bibr" rid="B46">2010</xref>), and Bayesian brain hypothesis (Hampton et al., <xref ref-type="bibr" rid="B19">2006</xref>; Kopp, <xref ref-type="bibr" rid="B27">2006</xref>; Ostwald et al., <xref ref-type="bibr" rid="B35">2012</xref>; Lieder et al., <xref ref-type="bibr" rid="B29">2013</xref>), have been convincingly established. The frameworks explain human behavior or brain responses by positing the existence of an internal model that humans constantly and automatically generate about the external world (Donchin, <xref ref-type="bibr" rid="B10">1981</xref>). Different processes of the response of the internal model to an external event lead to different brain activities, such as the P3a and P3b variations (Kolossa et al., <xref ref-type="bibr" rid="B26">2015</xref>).</p>
<p>What state transition the internal model builds is discussed in this study. Previous studies, such as Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>) and Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>), assumed that the internal model is without state transitions. Accordingly, surprises were estimated based on a generative model in a stationary state (a stationary-state model). In contrast, the purpose of our study is to find evidence of a brain mechanism that codes state transitions. If the brain can generate an internal model with state transitions (the state transition model), then humans would not acquire a probability distribution of events but would acquire a model that describes how different states or situations of the world are connected to each other (Gl&#x000E4;scher et al., <xref ref-type="bibr" rid="B18">2010</xref>).</p>
<p>The possibility that the state transition models explain some effects in ERP components motivated this study. These properties of an event sequence, such as stationary-state models, alternation, and repetition (Matt et al., <xref ref-type="bibr" rid="B31">1992</xref>; Rac-Lubashevsky and Kessler, <xref ref-type="bibr" rid="B39">2016</xref>) that explain the variation in some ERP components can be generalized with a state transition model. Moreover, Gl&#x000E4;scher et al. (<xref ref-type="bibr" rid="B18">2010</xref>) suggested that probability distributions with state transitions are coded in the brain during the performance of reinforcement learning tasks (Saito et al., <xref ref-type="bibr" rid="B41">2015</xref>). Therefore, we hypothesized that, for prediction, a mechanism for coding state transitions exist; that is, surprise is modeled not only with the probability distribution of the stationary state but also with the probability distribution with state transitions.</p>
<p>In the present study, we investigated the relation between predictive surprise in a generative model that has state transitions and electrophysiological signals via a model-based analysis. To distinguish predictive surprises in the different state models, predictive surprise with state transitions is called <italic>predictive transition surprise</italic>, and predictive surprise in a stationary state is called <italic>predictive stationary surprise</italic>. Although the EEG signals were recorded with a two-choice response time task, the same one used by Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>) and Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>), the generative models for the event sequences were different. In previous studies, predictive transition surprise converged with predictive stationary surprise as the number of trials increased; therefore, that type of setting cannot isolate the effects of predictive transition surprise. To avoid this situation, we used state transition models for the generative models; we controlled generation of the event sequence with a simple Markov chain (Norris, <xref ref-type="bibr" rid="B34">1998</xref>). In the model-based analysis, we used predictive stationary or transition surprise of the Markov chain as the explanatory variable. As the response variable, we used the EEG potentials, which were observed in various electrodes and latencies. This analysis enabled us to visualize the effects of these surprises on various ERP components, such as P3a and P3b. The results show different brain activities that seem to be associated with the stationary-state model and the state transition model.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and methods</title>
<sec>
<title>2.1. Measurement</title>
<sec>
<title>2.1.1. Participants</title>
<p>Twelve individuals (10 male and 2 female) participated in the experiment. Their ages ranged from 21 to 27 years (<italic>M</italic> &#x0003D; 23.6; <italic>SD</italic> &#x0003D; 1.7). The participants had normal or corrected-to-normal visual acuity. All participants provided written informed consent. The experimental protocols were approved by the Committee for Human Research at the Toyohashi University of Technology, Aichi, Japan, and the experiment was conducted in accordance with the committee&#x00027;s approved guidelines.</p>
</sec>
<sec>
<title>2.1.2. Experimental design</title>
<p>The participants performed a two-choice response time (TCRT) task (Figure <xref ref-type="fig" rid="F1">1</xref>) without feedback about response accuracy (Mars et al., <xref ref-type="bibr" rid="B30">2008</xref>; Kolossa et al., <xref ref-type="bibr" rid="B25">2012</xref>): Two visual stimuli were presented about every 1.5 s, and the participants were required to respond to each stimulus with the previously associated button as quickly as possible. Visual stimuli were presented on an LCD display [VIEWPixx EEG (VPixx Technologies)] with Psychtoolbox-3 and MATLAB R2011b (The MathWorks, Inc.). The participants were seated in front of the display. They touched their left or right hand to the left or right button of a four-button trackball on the desk between the participant and the display.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>The experimental procedure</bold>. This represents the participant&#x00027;s task for a case in which the square corresponds to the right click and the circle corresponds to the left click.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0001.tif"/>
</fig>
<p>A single trial consisted of the following procedures. First, the participant gazed at the fixation cross (2.0&#x000B0; &#x000D7; 2.0&#x000B0;) presented at the center of the display for 400&#x02013;600 ms. Then, the fixation cross vanished, and a circle or a square (2.0&#x000B0; &#x000D7; 2.0&#x000B0;) was presented for 200 ms at the location where the fixation cross had originally been presented. The participant clicked the left or right button as quickly as possible. The participant&#x00027;s response was accepted from 200 ms to 1000 ms after the symbol was presented. If the participant did not respond during this period, then the trial was recorded as a no-response trial. The fixation cross was presented when the symbol vanished.</p>
<p>The symbol-response assignment and instructions were given to the participants before the experiment. For example, participants were instructed &#x0201C;to click the left (right) button when the circle (square) appears.&#x0201D; Before the measurement blocks, the participants underwent a training blocks for the response task consisting of 50 trials; the training blocks was repeated until the participants&#x00027; response accuracy reached 90%. During the training blocks, the stimulus sequences were generated randomly with a uniform probability distribution.</p>
<p>The symbol stimuli (circle and square) were generated with simple Markov chains. There were two conditions (<italic>C1</italic> and <italic>C2</italic>) with different Markov chains. Let Event <italic>a</italic> and Event <italic>b</italic> correspond to either the circle or the square, respectively, and let <italic>E</italic><sub><italic>n</italic></sub> be the present event and <italic>E</italic><sub><italic>n</italic>&#x02212;1</sub> be the preceding event. This assignment is denoted as the event-symbol assignment. Condition <italic>C1</italic> can be represented with the transition probabilities <italic>P</italic>(<italic>E</italic><sub><italic>n</italic></sub> | <italic>E</italic><sub><italic>n</italic>&#x02212;1</sub>) as <italic>P</italic>(<italic>a</italic> | <italic>a</italic>) &#x0003D; <italic>P</italic>(<italic>b</italic> | <italic>b</italic>) &#x0003D; 0.3 and <italic>P</italic>(<italic>b</italic> | <italic>a</italic>) &#x0003D; <italic>P</italic>(<italic>a</italic> | <italic>b</italic>) &#x0003D; 0.7. For Condition <italic>C2</italic>, <italic>P</italic>(<italic>a</italic> | <italic>a</italic>) &#x0003D; 0.3, <italic>P</italic>(<italic>b</italic> | <italic>a</italic>) &#x0003D; 0.7, and <italic>P</italic>(<italic>a</italic> | <italic>b</italic>) &#x0003D; <italic>P</italic>(<italic>b</italic> | <italic>b</italic>) &#x0003D; 0.5. The Markov chains for the two conditions are summarized in Figure <xref ref-type="fig" rid="F2">2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>The Markov chains used to generate the event sequences in the two conditions. (A)</bold> Condition <italic>C1</italic>. <bold>(B)</bold> Condition <italic>C2</italic>.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0002.tif"/>
</fig>
<p>A single block consisted of 300 trials, and the participants executed a total of four blocks. Two blocks were of Condition <italic>C1</italic>, and two were of Condition <italic>C2</italic>. The event sequence was the same for the two blocks that had the same condition. The participants took a break of at least two minutes between blocks. The participants were not told that there were two conditions for the generative model, and according to a question that we asked the participants after the experiment, none noticed that there were two conditions.</p>
<p>The symbol-response and event-symbol assignments and the block orders were decided randomly as follows. The symbol-response assignment was different for each participant. The event-symbol assignment was chosen randomly during the blocks. The order of the blocks with the two conditions was random.</p>
</sec>
<sec>
<title>2.1.3. EEG acquisition</title>
<p>The EEG signals were recorded using a BioSemi ActiveTwo system. The EEG recording was performed at a sampling rate of 512 Hz with a 64-electrode cap, referenced to the common mode sense (CMS) active electrode. The 64 active electrodes were positioned to cover the whole head according to the extended International 10/10 system. The signals in the electrodes placed on the left and right earlobes, on the right side of the right eye (on the temple), and at the left, upper, and lower sides of the left eye were also measured. For preprocessing, the signals were re-referenced with the averaged potential of both earlobes. Moreover, a Butterworth bandpass filter (passband: 0.3&#x02013;30 Hz, order: 4) was applied to the signals. Epochs were corrected using the &#x02212;100 to 0 ms period as the baseline. The epochs in which the vertical electroculograms were more than &#x000B1;80 &#x003BC;V were removed.</p>
</sec>
</sec>
<sec>
<title>2.2. Analysis of behavior and EEG</title>
<p>In our analysis, the symbols <italic>aa</italic>, <italic>ab</italic>, <italic>ba</italic>, and <italic>bb</italic> represent the data for the present stimulus after the preceding stimulus. For example, <italic>ab</italic> represents the data for Event <italic>b</italic> after Event <italic>a</italic>.</p>
<p>For the behavioral data, the clicked buttons and the response times of the participants&#x00027; responses for all trials were acquired. The trials in which the participants responded with the wrong button were removed from the analysis of the response time. We tested the averaged behavioral data for each participant with a two-way repeated-measures analysis of variance (ANOVA) (Cohen et al., <xref ref-type="bibr" rid="B5">2003</xref>; Rac-Lubashevsky and Kessler, <xref ref-type="bibr" rid="B39">2016</xref>) with the factors <italic>Present</italic> (the present stimulus with two levels: Event <italic>a</italic> and Event <italic>b</italic>) and <italic>Preceding</italic> (the preceding stimulus with two levels: Same and Different as the present stimulus). The combinations of the two factors resulted in the sequences <italic>aa</italic> for (Event <italic>a</italic>, Same), <italic>ba</italic> for (Event <italic>a</italic>, Difference), <italic>bb</italic> for (Event <italic>b</italic>, Same), and <italic>ab</italic> for (Event <italic>b</italic>, Difference).</p>
<p>For conventional analysis of ERPs, the trials in which the participants responded with the wrong button were removed from the analysis. We tested the averaged EEG potential at each channel and latency for each participant with an ANOVA in which the factors were the same as those for the behavioral data.</p>
</sec>
<sec>
<title>2.3. Model-based analysis</title>
<p>For finding the time periods and electrodes in the EEG signals that are well explained by trial-by-trial surprises, we used a model-based analysis of regression with a generalized linear model (GLM) (Bolker et al., <xref ref-type="bibr" rid="B3">2009</xref>). Trial-by-trial surprises were estimated based on the preceding series of the stimulus. The two types of surprise were compared in their effects on the EEG potentials: surprise generated by stationary-state models and by state transition models. A model-based analysis evaluates the relation between two variables with an indicator representing how much one variable accurately explains and/or predicts the other. In this analysis, the variable used to explain the other variable is called the explanatory variable, and the variable to be explained is called the response variable (Haykin, <xref ref-type="bibr" rid="B20">2005</xref>; Dobson and Barnett, <xref ref-type="bibr" rid="B9">2011</xref>).</p>
<p>Surprise concerning the present event (called predictive surprise in Kolossa et al., <xref ref-type="bibr" rid="B25">2012</xref>, <xref ref-type="bibr" rid="B26">2015</xref>) was used as the explanatory variable. Predictive stationary surprise is defined as the logarithm of the overall probability given the preceding series of events. Figure <xref ref-type="fig" rid="F3">3</xref> shows the trial-by-trial change in predictive stationary surprise for each condition. Additionally, predictive transition surprise is defined as the logarithm of the probability transitioning from the preceding event to the present event given the preceding series of events. Figure <xref ref-type="fig" rid="F4">4</xref> shows the trial-by-trial change in predictive transition surprise for each condition. The detailed definitions of surprises are described in Section A.1.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Predictive stationary surprise <italic>I</italic><sub>S</sub> at the <italic>n</italic>th trial</bold>. <bold>(A)</bold> Condition <italic>C1</italic>. <bold>(B)</bold> Condition <italic>C2</italic>.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Transition stationary surprise <italic>I</italic><sub>T</sub> at the <italic>n</italic>th trial</bold>. <bold>(A)</bold> Condition <italic>C1</italic>. <bold>(B)</bold> Condition <italic>C2</italic>.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0004.tif"/>
</fig>
<p>The EEG potentials of each trial were used as the samples of the response variable for the model-based approach. Before the potentials were extracted, the EEG signals over the two blocks of the same condition in each trial were averaged. If either sample in the two blocks was missing because of a wrong task response, no response, or artifact rejection, those trials were removed from the analysis (Kolossa et al., <xref ref-type="bibr" rid="B25">2012</xref>). The trial-by-trial ERPs were extracted by averaging the preprocessed EEG signals over a temporal window of &#x000B1;50 ms around every 20 ms from 0 to 700 ms from the onset of the stimulus.</p>
<p>The samples of Conditions <italic>C1</italic> and <italic>C2</italic> were merged into a sample set. The merging reduced specific effects of the generative models, such as alternation expectation (Squires et al., <xref ref-type="bibr" rid="B47">1977</xref>; Mars et al., <xref ref-type="bibr" rid="B30">2008</xref>; Kolossa et al., <xref ref-type="bibr" rid="B25">2012</xref>). The number of the samples for each channel and latency was 5578 (the participants&#x00027; mean &#x0003D; 464.83; and <italic>SD</italic> &#x0003D; 44.56).</p>
<p>A GLM (Bolker et al., <xref ref-type="bibr" rid="B3">2009</xref>) was adopted for the regression of the explanatory and response variables. The model is summarized in Figure <xref ref-type="fig" rid="F5">5</xref>. The <italic>N</italic> observed samples of the set of explanatory and response variables (<italic>N</italic> &#x0003D; 5578) are represented as <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which corresponds to predictive surprise and the EEG potential at a channel and time period for the <italic>n</italic>th sample. The EEG potential is modeled by the linear model of estimated surprise with consideration of individual differences formulated as &#x003BC;<sub><italic>n</italic></sub> &#x0003D; <italic>b</italic><sub>1</sub><italic>x</italic><sub><italic>n</italic></sub> &#x0002B; <italic>b</italic><sub>2</sub> &#x0002B; <italic>r</italic><sub><italic>n</italic></sub> and an additive Gaussian noise, <inline-formula><mml:math id="M20"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, with the variance formulated as <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The unknown parameters were found with the maximum log-likelihood method (Gelman et al., <xref ref-type="bibr" rid="B17">2013</xref>). Details of the model and the fitting procedure are given in Section A.2.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>The graphical model for the model-based analysis</bold>. The variables, the values of which are determined stochastically, are represented by the dashed lines. The deterministic variables are represented by the solid lines. The distributions surrounded by circles represent the non-informative prior probability distributions.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0005.tif"/>
</fig>
<p>The fitting accuracy of the regression model was evaluated using the log-Bayes factor of the estimated model. The model <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> was estimated with predictive stationary surprise, and <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> was estimated with predictive transition surprise. Moreover, a common reference model <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>NULL</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> was also estimated with a set in which all samples for the response variables were 1 (Neyman and Pearson, <xref ref-type="bibr" rid="B33">1933</xref>; Kolossa et al., <xref ref-type="bibr" rid="B26">2015</xref>). The log-likelihood of <inline-formula><mml:math id="M25"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denoted by log <italic>L</italic><sub><italic>M</italic></sub> for <italic>M</italic> is S, T, or NULL. As an indicator for the fitting accuracy, the log-Bayes factor <italic>B</italic><sub><italic>M</italic></sub> with the common reference model (Kass and Raftery, <xref ref-type="bibr" rid="B23">1995</xref>; Kolossa et al., <xref ref-type="bibr" rid="B26">2015</xref>) was adopted:</p>
<disp-formula id="E18"><label>(1)</label><mml:math id="M18"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mtext>NULL</mml:mtext></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>for <italic>M</italic> is S or T.</p>
<p>The log-Bayes factor was evaluated with a likelihood-ratio test (Neyman and Pearson, <xref ref-type="bibr" rid="B33">1933</xref>) that evaluates how more accurately the model fits than the common reference model. A parametric bootstrap method (Davison and Hinkley, <xref ref-type="bibr" rid="B8">1997</xref>) (the number of sampling &#x0003D; 1000) was used for the test.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Behavioral data</title>
<p>The response time for the button-clicking task was defined as the duration between the display of the stimulus and the clicking of the button. Mean values are displayed in Figure <xref ref-type="fig" rid="F6">6</xref>. The ANOVA showed a main effect of the factor <italic>Present</italic> in Condition <italic>C2</italic> [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 15.8042, <italic>p</italic> &#x0003D; 0.0022].</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Response time for the single stimulus or sequence</bold>. The error bars show the standard error. <bold>(A)</bold> Condition <italic>C1</italic>. <bold>(B)</bold> Condition <italic>C2</italic>.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0006.tif"/>
</fig>
<p>The response accuracy for the button-clicking task was defined as whether or not the participant clicked the assigned button correctly. Mean values are displayed in Figure <xref ref-type="fig" rid="F7">7</xref>. The ANOVA showed a main effect of the factor <italic>Preceding</italic> in Condition <italic>C1</italic> [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 16.5023, <italic>p</italic> &#x0003D; 0.0019]. In Condition <italic>C2</italic>, main effects were found for the factors <italic>Present</italic> [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 10.1730, <italic>p</italic> &#x0003D; 0.0086] and <italic>Preceding</italic> [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 6.5105, <italic>p</italic> &#x0003D; 0.0269]. An interaction of the two factors [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 6.2011, <italic>p</italic> &#x0003D; 0.03] was also found in Condition <italic>C2</italic>. Simple effects for the interaction were found for <italic>Present</italic> at the level Same [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 14.8977, <italic>p</italic> &#x0003D; 0.0027] and for <italic>Preceding</italic> at the level Event <italic>a</italic> [<italic>F</italic><sub>(1, 11)</sub> &#x0003D; 14.9957, <italic>p</italic> &#x0003D; 0.0026].</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Response accuracy for the single stimulus or sequence</bold>. The error bars show the standard error. <bold>(A)</bold> Condition <italic>C1</italic>. <bold>(B)</bold> Condition <italic>C2</italic>.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0007.tif"/>
</fig>
<p>The results of the statistical analysis suggest that the behavior (response time and accuracy) was affected by state transitions. The main effect of <italic>Preceding</italic> on response accuracy in Condition <italic>C1</italic> reflects the high transition probability for Sequences <italic>ab</italic> and <italic>ba</italic> in the generative model. The difference in the stationary probability between Events <italic>a</italic> and <italic>b</italic> in Condition <italic>C2</italic> can explain the main effects by <italic>Present</italic> for the response time and accuracy. In the response accuracy, the simple effect by <italic>Present</italic> at the level Same corresponds to the difference in the transition probabilities between Sequences <italic>aa</italic> (0.3) and <italic>bb</italic> (0.5) in Condition <italic>C2</italic>. The simple effect by <italic>Preceding</italic> at the level Event <italic>a</italic> corresponds to the difference between Sequences <italic>aa</italic> (0.3) and <italic>ba</italic> (0.5). The differences in the transition probabilities between Sequences <italic>ab</italic> (0.7) and <italic>bb</italic> (0.5), and Sequences <italic>ab</italic> (0.7) and <italic>ba</italic> (0.5) in the generative model of Condition <italic>C2</italic>, however, did not appear in behavior.</p>
</sec>
<sec>
<title>3.2. Event-related potentials</title>
<p>Figure <xref ref-type="fig" rid="F8">8</xref> depicts the grand-averaged ERP waveforms. The ANOVA showed an effect of <italic>Preceding</italic> in FCz (Conditions <italic>C1</italic> and <italic>C2</italic>) and CPz (Condition <italic>C1</italic>) at a latency around 340&#x02013;400 ms. An effect of <italic>Present</italic> was found in CPz, Condition <italic>C2</italic> at a latency around 370&#x02013;400 ms. In Figure <xref ref-type="fig" rid="F8">8A</xref>, the peak amplitude for the sequences that have a transition probability of 0.3 (<italic>aa</italic> and <italic>bb</italic>) is higher than that for those that have a transition probability of 0.7 (<italic>ab</italic> and <italic>ba</italic>). Sequence <italic>aa</italic> in Condition <italic>C2</italic>, which has a transition probability of 0.3, leads to the highest peak amplitude, as shown in Figures <xref ref-type="fig" rid="F8">8B,D</xref>.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>The grand-averaged EEG potentials observed at the channels FCz and CPz for each sequence with two successive stimuli</bold>. The vertical dashed lines located at 0 and 200 ms represent the onset and offset of the symbol stimulus. The vertical solid lines represents the latency when the maximum amplitude was observed. The time periods in which the main effects are found in the ANOVA at each channel are shown by the red (<italic>Present</italic>), green (<italic>Preceding</italic>), and blue (interaction) unfilled bars (<italic>p</italic> &#x0003C; 0.05). The time periods in which the main effects are found in the log-Bayes factors at each channel are shown by red (the stationary-state model <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>) and green (the state transition model <inline-formula><mml:math id="M27"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>) filled bars (<italic>p</italic> &#x0003C; 0.05). <bold>(A)</bold> FCz (Condition <italic>C1</italic>). <bold>(B)</bold> FCz (Condition <italic>C2</italic>). <bold>(C)</bold> CPz (Condition <italic>C1</italic>). <bold>(D)</bold> CPz (Condition <italic>C2</italic>).</p></caption>
<graphic xlink:href="fnhum-11-00075-g0008.tif"/>
</fig>
<p>The peaks of P300 are at around 400 ms, which are 100 ms later than the peak latencies reported by Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>), who employed a color discrimination task. This difference could be caused by the difference in the stimulus features that an observer should detect for the TCRT tasks (Smid et al., <xref ref-type="bibr" rid="B45">1999</xref>). Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>), who employed a shape discrimination task, reported a similar peak latency as in this analysis, around 400 ms.</p>
</sec>
<sec>
<title>3.3. Model-based analysis</title>
<p>Figure <xref ref-type="fig" rid="F9">9</xref> displays the log-Bayes factors. We found high log-Bayes factors (shown in red), which indicate high fitting accuracy, in some channels and latencies.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>Log-Bayes factors <italic>B</italic><sub><italic>M</italic></sub> for each stage of the trials (left to right: whole trials (1st&#x02013;300th) and early (1st&#x02013;100th), middle (101st&#x02013;200th), and last (201st&#x02013;300th) stages)</bold>. The factor decreases as the color turns from red to blue (see the color bar beside each figure). <bold>(A)</bold> Model based on predictive stationary surprise <inline-formula><mml:math id="M28"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. <bold>(B)</bold> Model based on predictive transition surprise <inline-formula><mml:math id="M29"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0009.tif"/>
</fig>
<p>For the stationary-state model <inline-formula><mml:math id="M30"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, the likelihood-ratio test showed that the models that reached a log-Bayes factor &#x02265; 1487 fitted significantly more accurately than the common reference model (<italic>p</italic> &#x0003C; 0.05). The latencies at the channels FCz and CPz in which the statistically significant differences were found are shown as the red bars with the ERP waveforms in Figure <xref ref-type="fig" rid="F8">8</xref>. At FCz, effects are found within 300&#x02013;360 and 500&#x02013;580 ms. At CPz, effects are found within 380&#x02013;400 and 480&#x02013;600 ms.</p>
<p>For the state transition model <inline-formula><mml:math id="M31"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, the likelihood-ratio test showed that the models that reached a log-Bayes factor &#x02265; 2241 were statistically significant (<italic>p</italic> &#x0003C; 0.05). The latencies at the channels FCz and CPz in which the statistically significant differences were found are shown as the green bars with the ERP waveforms in Figure <xref ref-type="fig" rid="F8">8</xref>. At FCz, an effect is found within 340&#x02013;480 ms. At CPz, an effect is found within 340&#x02013;440 ms.</p>
<p>In Figure <xref ref-type="fig" rid="F10">10A</xref>, which shows the change in the log-Bayes factors according to the stage of the trials, an increase at the middle stage and a decrease at the last stage in the log-Bayes factors for <inline-formula><mml:math id="M32"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> (<italic>t</italic> &#x0003D; 360 ms) were observed at FCz and Cz. Figure <xref ref-type="fig" rid="F10">10B</xref> shows that the log-Bayes factor for <inline-formula><mml:math id="M33"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> at Cz and CPz increases as the trials accumulate.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><bold>Log-Bayes factors <italic>B</italic><sub><italic>M</italic></sub> at the channels, FPz and CPz</bold>. The factors were derived using the &#x000B1;50th trials from the index of the <italic>x</italic>-axis (e.g., for the trial index 150, the 100th to 200th trials were used). <bold>(A)</bold> Predictive transition surprise, <italic>t</italic> &#x0003D; 360 ms. <bold>(B)</bold> Predictive stationary surprise, <italic>t</italic> &#x0003D; 480 ms.</p></caption>
<graphic xlink:href="fnhum-11-00075-g0010.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>This study investigated the effects of predictive stationary surprise and predictive transition surprise on EEG potentials under the assumption that the internal model is formed with state transitions and predictive surprise is based not only on a stationary-state model but also on a state transition model. For this, we applied Markov chains to generate event sequences in order to isolate the effects of stationary and transition surprises. The results show that predictive stationary surprise better explains P3b and predictive transition surprise better explains P3a. This suggests two distinct mechanisms in human prediction. The effect of predictive transition surprise on P3a suggests that a mechanism for estimating the generative model exists and that the internal model forms a state transition model. The result also indicates a mechanism for processing a stationary-state model as observed by the variability of P3b. The dependencies on time (the number of observed events) of these effects could reflect the process to form the observer&#x00027;s prediction.</p>
<p>We adopted a simple procedure in which predictive surprise was estimated as the self-information of the present event. The self-information for the event was estimated from the preceding event sequences. This procedure is equivalent to the procedure proposed by Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>). However, the optimization problem for the parameters in the DIF model is very complex, and the optimization needs to use an empirical procedure, which does not have the guarantee of a global optimum. Since we focused on the effects on the brain activity that differs between the stationary-state and state transition probabilities, we adopted a fairly simple model, one that does not have any parameters that need to be optimized. Moreover, the DIF model does not accurately produce surprise associated with the state transition model because it is based on a linear combination of the three factors.</p>
<p>The results show that the behavioral data (response accuracy and response time), ERP waveforms, and log-Bayes factors depend on predictive stationary surprise. The behavioral data results correspond to those of Miller (<xref ref-type="bibr" rid="B32">1998</xref>) and Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>). In the ERP results, as Polich (<xref ref-type="bibr" rid="B38">2007</xref>) suggested, the ERP at 390 ms in the centro-parietal region dependent on the stationary probability can be observed. From the model-based analysis, the high fitting accuracy with a centro-parietal focus within 480&#x02013;600 ms can be considered to be a result of a variation in the P3b component. This speculation is supported by Kolossa et al. (<xref ref-type="bibr" rid="B26">2015</xref>), who suggested that P3b is more strongly associated with predictive stationary surprise than P3a.</p>
<p>The effects of predictive transition surprise can be seen in the present results. The behavioral results suggest that response accuracy and response time depend on the preceding event even if the present event is the same: The difficulty of the response depends on the transition probability distribution. The feature observed in the ERP waveforms (Figure <xref ref-type="fig" rid="F8">8</xref>), that high transition surprise leads to a high peak, is similar to ERP responses to stationary surprise. We suggest here that the effect of the transition probability on behavior and ERPs has not been revealed clearly. In the model-based analysis, the high fitting accuracy with a central focus within 340&#x02013;480 ms can be considered to be caused by a variation in P3a because similar features in its area (Kopp and Lange, <xref ref-type="bibr" rid="B28">2013</xref>) and latency (Kolossa et al., <xref ref-type="bibr" rid="B26">2015</xref>) have been reported.</p>
<p>The dependence of the P3a component on predictive transition surprise suggests that the participants estimated state transition models as the generative model. This can be explained by introducing Bayesian surprise. This suggests that the variation in P3a occurs via the updating of the internal model. Because the P3a, and thus the update, can be modeled better by predictive transition surprise than by predictive stationary surprise in this experimental setting, it appears that the internal model is associated more strongly with a model with state transitions than with a stationary model. Namely, the internal model approximates a Markov chain. This speculation is supported by the decrease in the fitting accuracy for P3a in <inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">M</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> at the last stage (Figure <xref ref-type="fig" rid="F10">10A</xref>) because the internal model converges by accumulating the event observations and Bayesian surprise is slight in the last stage. This convergence corresponds to the convergence in predictive transition surprise shown in Figure <xref ref-type="fig" rid="F4">4</xref>.</p>
<p>The effect of predictive stationary surprise shows that human prediction has a mechanism different from that of the generative model. Although the internal model is built based on the state transition model, the P3b components depend on the stationary probability distribution. This result suggests that P3b is not affected by the state transitions and mainly reflects the stationary state. The effect of stationary-state models on P3b has been confirmed by El Karoui et al. (<xref ref-type="bibr" rid="B14">2015</xref>) and Bekinschtein et al. (<xref ref-type="bibr" rid="B2">2009</xref>) as <italic>the global effect</italic>. The increase of the fitting accuracy for P3b at the last stage is consistent with a feature of the global effect that is related to the accumulation of an event on longer time scales (d&#x00027;Acremont et al., <xref ref-type="bibr" rid="B7">2013</xref>; El Karoui et al., <xref ref-type="bibr" rid="B14">2015</xref>).</p>
<p>Although Kolossa et al. (<xref ref-type="bibr" rid="B26">2015</xref>) showed that P3b is distributed in the centro-parietal region, the effect of the stationary-state model is observed also in fronto-central region. This effect could be caused by a P300 latency shift that novelty detection (Courchesne et al., <xref ref-type="bibr" rid="B6">1975</xref>; Knight, <xref ref-type="bibr" rid="B24">1984</xref>) and attention (Kahneman, <xref ref-type="bibr" rid="B22">1973</xref>) affect. This hypothesis is supported by the effect observed in both time periods of the ascending and descending flanks of P300 as shown in Figures <xref ref-type="fig" rid="F8">8A,B</xref>.</p>
<p>The electrophysiological effects of the two generative models we tested in this study, the stationary-state and state transition models, support theoretical frameworks regarding ERPs, such as the context-updating model, predictive coding, and the Bayesian brain hypothesis. Moreover, the effects suggest the following hypotheses about what kind of models the brain adopts as the internal model. (1) If the external generative model has state transitions, then the internal model can represent the state transitions. (2) If the external generative model does not change over time, then updating of internal model ceases at a certain point. (3) P3b considered to be affected by prediction errors (Spratling, <xref ref-type="bibr" rid="B46">2010</xref>; Kolossa et al., <xref ref-type="bibr" rid="B25">2012</xref>) does not directly reflect the errors between a present event and a prediction generated by the internal model&#x02014;the P3b variability is caused by the prediction errors for a stationary-state model translated from the internal model or is led by a different process from the generating of the internal model.</p>
<p>As pointed out in Mars et al. (<xref ref-type="bibr" rid="B30">2008</xref>) and Kolossa et al. (<xref ref-type="bibr" rid="B25">2012</xref>), the TCRT task requires motor responses. Therefore, it is still an open problem whether the cause of the variation in the ERP components is surprise conveyed by the stimulus or surprise associated with the motor responses.</p>
<p>We conclude that our approach using Markov chains provides observation of the different effects on ERPs produced by surprises on the stationary-state and state transition models. The differences in the effects suggest that an internal model in the brain can form a probability model with state transitions. The effects of a stationary-state model suggest the existence of a different brain mechanism from that for forming the internal model. Moreover, a change in these effects by the accumulation of events was observed. This shows the part of neural responses that reflects a brain mechanism by which humans gain predictions from their experiences.</p>
</sec>
<sec id="s5">
<title>Author contributions</title>
<p>HH, TM, and SN designed the work. HH and TM collected data. HH analyzed data. HH drafted the manuscript. TM and SN revised the manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</sec>
</body>
<back>
<ack><p>This work was supported by the Japan Society for the Promotion of Science (JSPS) KAKENHI [Grant numbers 15K21079, 26240043].</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baldi</surname> <given-names>P.</given-names></name> <name><surname>Itti</surname> <given-names>L.</given-names></name></person-group> (<year>2010</year>). <article-title>Of bits and wows: a Bayesian theory of surprise with applications to attention</article-title>. <source>Neural Netw.</source> <volume>23</volume>, <fpage>649</fpage>&#x02013;<lpage>666</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2009.12.007</pub-id><pub-id pub-id-type="pmid">20080025</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bekinschtein</surname> <given-names>T. A.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name> <name><surname>Rohaut</surname> <given-names>B.</given-names></name> <name><surname>Tadel</surname> <given-names>F.</given-names></name> <name><surname>Cohen</surname> <given-names>L.</given-names></name> <name><surname>Naccache</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>Neural signature of the conscious processing of auditory regularities</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>106</volume>, <fpage>1672</fpage>&#x02013;<lpage>1677</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0809667106</pub-id><pub-id pub-id-type="pmid">19164526</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bolker</surname> <given-names>B. M.</given-names></name> <name><surname>Brooks</surname> <given-names>M. E.</given-names></name> <name><surname>Clark</surname> <given-names>C. J.</given-names></name> <name><surname>Geange</surname> <given-names>S. W.</given-names></name> <name><surname>Poulsen</surname> <given-names>J. R.</given-names></name> <name><surname>Stevens</surname> <given-names>M. H. H.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Generalized linear mixed models: a practical guide for ecology and evolution</article-title>. <source>Trends Ecol. Evol.</source> <volume>24</volume>, <fpage>127</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1016/j.tree.2008.10.008</pub-id><pub-id pub-id-type="pmid">19185386</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Whatever next? Predictive brains, situated agents, and the future of cognitive science</article-title>. <source>Behav. Brain Sci.</source> <volume>36</volume>, <fpage>181</fpage>&#x02013;<lpage>204</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X12000477</pub-id><pub-id pub-id-type="pmid">23663408</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name> <name><surname>Cohen</surname> <given-names>P.</given-names></name> <name><surname>West</surname> <given-names>S. G.</given-names></name> <name><surname>Aiken</surname> <given-names>L. S.</given-names></name></person-group> (<year>2003</year>). <source>Applied Multiple Regression/Correlation Analysis for the Behavioral Science, 3rd Edn.</source> <publisher-loc>Mahwah, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates</publisher-name>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Courchesne</surname> <given-names>E.</given-names></name> <name><surname>Hillyard</surname> <given-names>S. A.</given-names></name> <name><surname>Galambos</surname> <given-names>R.</given-names></name></person-group> (<year>1975</year>). <article-title>Stimulus novelty, task relevance and the visual evoked potential in man</article-title>. <source>Electroencephal. Clin. Neurophysiol.</source> <volume>39</volume>, <fpage>131</fpage>&#x02013;<lpage>143</lpage>. <pub-id pub-id-type="pmid">50210</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>d&#x00027;Acremont</surname> <given-names>M.</given-names></name> <name><surname>Schultz</surname> <given-names>W.</given-names></name> <name><surname>Bossaerts</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>The human brain encodes event frequencies while forming subjective beliefs</article-title>. <source>J. Neurosci.</source> <volume>33</volume>, <fpage>10887</fpage>&#x02013;<lpage>10897</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5829-12.2013</pub-id><pub-id pub-id-type="pmid">23804108</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Davison</surname> <given-names>A. C.</given-names></name> <name><surname>Hinkley</surname> <given-names>D. V.</given-names></name></person-group> (<year>1997</year>). <source>Bootstrap Methods and Their Application</source>. Cambridge Series in Statistical and Probabilistic Mathematics. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dobson</surname> <given-names>A. J.</given-names></name> <name><surname>Barnett</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <source>An Introduction to Generalized Linear Models</source>. <publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Donchin</surname> <given-names>E.</given-names></name></person-group> (<year>1981</year>). <article-title>Surprise!&#x02026; Surprise?</article-title> <source>Psychophysiology</source> <volume>18</volume>, <fpage>493</fpage>&#x02013;<lpage>513</lpage>. <pub-id pub-id-type="doi">10.1111/j.1469-8986.1981.tb01815.x</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Donchin</surname> <given-names>E.</given-names></name> <name><surname>Coles</surname> <given-names>M. G. H.</given-names></name></person-group> (<year>1988</year>). <article-title>Is the P300 component a manifestation of context updating?</article-title> <source>Behav. Brain Sci.</source> <volume>11</volume>, <fpage>357</fpage>&#x02013;<lpage>374</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X00058027</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Doya</surname> <given-names>K.</given-names></name> <name><surname>Ishii</surname> <given-names>S.</given-names></name> <name><surname>Pouget</surname> <given-names>A.</given-names></name> <name><surname>Rao</surname> <given-names>R. P. N.</given-names></name></person-group> (<year>2007</year>). <source>Bayesian Brain: Probabilistic Approaches to Neural Coding</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duncan-Johnson</surname> <given-names>C. C.</given-names></name> <name><surname>Donchin</surname> <given-names>E.</given-names></name></person-group> (<year>1977</year>). <article-title>On quantifying surprise: the variation of event-related potentials with subjective probability</article-title>. <source>Psychophysiology</source> <volume>14</volume>, <fpage>456</fpage>&#x02013;<lpage>467</lpage>. <pub-id pub-id-type="doi">10.1111/j.1469-8986.1977.tb01312.x</pub-id><pub-id pub-id-type="pmid">905483</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>El Karoui</surname> <given-names>I.</given-names></name> <name><surname>King</surname> <given-names>J.-R.</given-names></name> <name><surname>Sitt</surname> <given-names>J.</given-names></name> <name><surname>Meyniel</surname> <given-names>F.</given-names></name> <name><surname>Van Gaal</surname> <given-names>S.</given-names></name> <name><surname>Hasboun</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Event-related potential, time-frequency, and functional connectivity facets of local and global auditory novelty processing: an intracranial study in humans</article-title>. <source>Cereb. Cortex</source> <volume>25</volume>, <fpage>4203</fpage>&#x02013;<lpage>4212</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhu143</pub-id><pub-id pub-id-type="pmid">24969472</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2002</year>). <article-title>Beyond phrenology: what can neuroimaging tell us about distributed circuitry?</article-title> <source>Ann. Rev. Neurosci.</source> <volume>25</volume>, <fpage>221</fpage>&#x02013;<lpage>250</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.neuro.25.112701.142846</pub-id><pub-id pub-id-type="pmid">12052909</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K.</given-names></name></person-group> (<year>2008</year>). <article-title>Hierarchical models in the brain</article-title>. <source>PLoS Comput. Biol.</source> <volume>4</volume>:<fpage>e1000211</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000211</pub-id><pub-id pub-id-type="pmid">18989391</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gelman</surname> <given-names>A.</given-names></name> <name><surname>Carlin</surname> <given-names>J. B.</given-names></name> <name><surname>Stern</surname> <given-names>H. S.</given-names></name> <name><surname>Dunson</surname> <given-names>D. B.</given-names></name> <name><surname>Vehtari</surname> <given-names>A.</given-names></name> <name><surname>Rubin</surname> <given-names>D. B.</given-names></name></person-group> (<year>2013</year>). <source>Bayesian Data Analysis, 3rd Edn</source>. Chapman &#x00026; Hall/CRC Texts in Statistical Science. <publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gl&#x000E4;scher</surname> <given-names>J.</given-names></name> <name><surname>Daw</surname> <given-names>N.</given-names></name> <name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>O&#x00027;Doherty</surname> <given-names>J. P.</given-names></name></person-group> (<year>2010</year>). <article-title>States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning</article-title>. <source>Neuron</source> <volume>66</volume>, <fpage>585</fpage>&#x02013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2010.04.016</pub-id><pub-id pub-id-type="pmid">20510862</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hampton</surname> <given-names>A. N.</given-names></name> <name><surname>Bossaerts</surname> <given-names>P.</given-names></name> <name><surname>O&#x00027;Doherty</surname> <given-names>J. P.</given-names></name></person-group> (<year>2006</year>). <article-title>The role of the ventromedial prefrontal cortex in abstract state-based inference during decision making in humans</article-title>. <source>J. Neurosci.</source> <volume>26</volume>, <fpage>8360</fpage>&#x02013;<lpage>8367</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1010-06.2006</pub-id><pub-id pub-id-type="pmid">16899731</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haykin</surname> <given-names>S.</given-names></name></person-group> (<year>2005</year>). <source>Adaptive Filter Theory, 4th Edn</source>. <publisher-loc>Upper Saddle River, NJ</publisher-loc>: <publisher-name>Prentice-Hall, Inc.</publisher-name></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Horovitz</surname> <given-names>S. G.</given-names></name> <name><surname>Skudlarski</surname> <given-names>P.</given-names></name> <name><surname>Gore</surname> <given-names>J. C.</given-names></name></person-group> (<year>2002</year>). <article-title>Correlations and dissociations between BOLD signal and P300 amplitude in an auditory oddball task: a parametric approach to combining fMRI and ERP</article-title>. <source>Mag. Reson. Imaging</source> <volume>20</volume>, <fpage>319</fpage>&#x02013;<lpage>325</lpage>. <pub-id pub-id-type="doi">10.1016/S0730-725X(02)00496-4</pub-id><pub-id pub-id-type="pmid">12165350</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kahneman</surname> <given-names>D.</given-names></name></person-group> (<year>1973</year>). <source>Attention and Effort</source>. <publisher-loc>Englewood Cliffs, NJ</publisher-loc>: <publisher-name>Prentice-Hall, Inc</publisher-name>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kass</surname> <given-names>R. E.</given-names></name> <name><surname>Raftery</surname> <given-names>A. E.</given-names></name></person-group> (<year>1995</year>). <article-title>Bayes factors</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>90</volume>, <fpage>773</fpage>&#x02013;<lpage>795</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1995.10476572</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Knight</surname> <given-names>R. T.</given-names></name></person-group> (<year>1984</year>). <article-title>Decreased response to novel stimuli after prefrontal lesions in man</article-title>. <source>Electroencephalogr. Clin. Neurophysiol.</source> <volume>59</volume>, <fpage>9</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="pmid">6198170</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kolossa</surname> <given-names>A.</given-names></name> <name><surname>Fingscheidt</surname> <given-names>T.</given-names></name> <name><surname>Wessel</surname> <given-names>K.</given-names></name> <name><surname>Kopp</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <article-title>A model-based approach to trial-by-trial P300 amplitude fluctuations</article-title>. <source>Front. Hum. Neurosci.</source> <volume>6</volume>:<fpage>359</fpage>. <pub-id pub-id-type="doi">10.3389/fnhum.2012.00359</pub-id><pub-id pub-id-type="pmid">23404628</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kolossa</surname> <given-names>A.</given-names></name> <name><surname>Kopp</surname> <given-names>B.</given-names></name> <name><surname>Fingscheidt</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>A computational analysis of the neural bases of Bayesian inference</article-title>. <source>NeuroImage</source> <volume>106</volume>, <fpage>222</fpage>&#x02013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2014.11.007</pub-id><pub-id pub-id-type="pmid">25462794</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kopp</surname> <given-names>B.</given-names></name></person-group> (<year>2006</year>). <article-title>The P300 component of the event-related brain potential and Bayes&#x00027; theorem</article-title>. <source>Cogn. Sci.</source> <volume>2</volume>, <fpage>113</fpage>&#x02013;<lpage>125</lpage>. <pub-id pub-id-type="doi">10.13140/2.1.4049.4402</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kopp</surname> <given-names>B.</given-names></name> <name><surname>Lange</surname> <given-names>F.</given-names></name></person-group> (<year>2013</year>). <article-title>Electrophysiological indicators of surprise and entropy in dynamic task-switching environments</article-title>. <source>Front. Hum. Neurosci.</source> <volume>7</volume>:<fpage>300</fpage>. <pub-id pub-id-type="doi">10.3389/fnhum.2013.00300</pub-id><pub-id pub-id-type="pmid">23840183</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lieder</surname> <given-names>F.</given-names></name> <name><surname>Daunizeau</surname> <given-names>J.</given-names></name> <name><surname>Garrido</surname> <given-names>M. I.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Stephan</surname> <given-names>K. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Modelling trial-by-trial changes in the mismatch negativity</article-title>. <source>PLoS Comput. Biol.</source> <volume>9</volume>:<fpage>e1002911</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002911</pub-id><pub-id pub-id-type="pmid">23436989</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mars</surname> <given-names>R. B.</given-names></name> <name><surname>Debener</surname> <given-names>S.</given-names></name> <name><surname>Gladwin</surname> <given-names>T. E.</given-names></name> <name><surname>Harrison</surname> <given-names>L. M.</given-names></name> <name><surname>Haggard</surname> <given-names>P.</given-names></name> <name><surname>Rothwell</surname> <given-names>J. C.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Trial-by-trial fluctuations in the event-related electroencephalogram reflect dynamic changes in the degree of surprise</article-title>. <source>J. Neurosci.</source> <volume>28</volume>, <fpage>12539</fpage>&#x02013;<lpage>12545</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2925-08.2008</pub-id><pub-id pub-id-type="pmid">19020046</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matt</surname> <given-names>J.</given-names></name> <name><surname>Leuthold</surname> <given-names>H.</given-names></name> <name><surname>Sommer</surname> <given-names>W.</given-names></name></person-group> (<year>1992</year>). <article-title>Differential effects of voluntary expectancies on reaction times and event-related potentials: evidence for automatic and controlled expectancies</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>18</volume>, <fpage>810</fpage>&#x02013;<lpage>822</lpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.18.4.810</pub-id><pub-id pub-id-type="pmid">1385618</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>J.</given-names></name></person-group> (<year>1998</year>). <article-title>Effects of stimulus-response probability on choice reaction time: evidence from the lateralized readiness potential</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform.</source> <volume>24</volume>, <fpage>1521</fpage>&#x02013;<lpage>1534</lpage>. <pub-id pub-id-type="doi">10.1037/0096-1523.24.5.1521</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neyman</surname> <given-names>J.</given-names></name> <name><surname>Pearson</surname> <given-names>E. S.</given-names></name></person-group> (<year>1933</year>). <article-title>On the problem of the most efficient tests of statistical hypotheses</article-title>. <source>Philos. Trans. R. Soc. Lond. Ser. A Contain. Pap. Math. Phys. Char.</source> <volume>231</volume>, <fpage>289</fpage>&#x02013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.1098/rsta.1933.0009</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Norris</surname> <given-names>J. R.</given-names></name></person-group> (<year>1998</year>). <source>Markov Chains</source>. Cambridge Series in Statistical and Probabilistic Mathematics. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ostwald</surname> <given-names>D.</given-names></name> <name><surname>Spitzer</surname> <given-names>B.</given-names></name> <name><surname>Guggenmos</surname> <given-names>M.</given-names></name> <name><surname>Schmidt</surname> <given-names>T. T.</given-names></name> <name><surname>Kiebel</surname> <given-names>S. J.</given-names></name> <name><surname>Blankenburg</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>). <article-title>Evidence for neural encoding of Bayesian surprise in human somatosensation</article-title>. <source>NeuroImage</source> <volume>62</volume>, <fpage>177</fpage>&#x02013;<lpage>188</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.04.050</pub-id><pub-id pub-id-type="pmid">22579866</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Patil</surname> <given-names>A.</given-names></name> <name><surname>Huard</surname> <given-names>D.</given-names></name> <name><surname>Fonnesbeck</surname> <given-names>C. J.</given-names></name></person-group> (<year>2010</year>). <article-title>PyMC: Bayesian stochastic modelling in Python</article-title>. <source>J. Stat. Softw.</source> <volume>35</volume>, <fpage>1</fpage>&#x02013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v035.i04</pub-id><pub-id pub-id-type="pmid">21603108</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Picton</surname> <given-names>T. W.</given-names></name></person-group> (<year>1992</year>). <article-title>The P300 wave of the human event-related potential</article-title>. <source>J. Clin. Neurophysiol.</source> <volume>9</volume>, <fpage>456</fpage>&#x02013;<lpage>479</lpage>. <pub-id pub-id-type="doi">10.1097/00004691-199210000-00002</pub-id><pub-id pub-id-type="pmid">1464675</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Polich</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Updating P300: an integrative theory of P3a and P3b</article-title>. <source>Clin. Neurophysiol.</source> <volume>118</volume>, <fpage>2128</fpage>&#x02013;<lpage>2148</lpage>. <pub-id pub-id-type="doi">10.1016/j.clinph.2007.04.019</pub-id><pub-id pub-id-type="pmid">17573239</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rac-Lubashevsky</surname> <given-names>R.</given-names></name> <name><surname>Kessler</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>Dissociating working memory updating and automatic updating: the reference-back paradigm</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>42</volume>, <fpage>951</fpage>&#x02013;<lpage>969</lpage>. <pub-id pub-id-type="doi">10.1037/xlm0000219</pub-id><pub-id pub-id-type="pmid">26618910</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Robert</surname> <given-names>C.</given-names></name></person-group> (<year>2007</year>). <source>The Bayesian Choice: From Decision-Theoretic Foundations to Computational Implementation</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saito</surname> <given-names>H.</given-names></name> <name><surname>Takiyama</surname> <given-names>K.</given-names></name> <name><surname>Okada</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Estimation of state transition probabilities: a neural network model</article-title>. <source>J. Phys. Soc. Jpn.</source> <volume>84</volume>:<fpage>5</fpage>. <pub-id pub-id-type="doi">10.7566/JPSJ.84.124801</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanmiguel</surname> <given-names>I.</given-names></name> <name><surname>Saupe</surname> <given-names>K.</given-names></name> <name><surname>Schr&#x000F6;ger</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>I know what is missing here: electrophysiological prediction error signals elicited by omissions of predicted &#x0201C;what&#x0201D; but not &#x0201C;when&#x0201D;</article-title>. <source>Front. Hum. Neurosci.</source> <volume>7</volume>:<fpage>407</fpage>. <pub-id pub-id-type="doi">10.3389/fnhum.2013.00407</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seer</surname> <given-names>C.</given-names></name> <name><surname>Lange</surname> <given-names>F.</given-names></name> <name><surname>Boos</surname> <given-names>M.</given-names></name> <name><surname>Dengler</surname> <given-names>R.</given-names></name> <name><surname>Kopp</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>Prior probabilities modulate cortical surprise responses: a study of event-related potentials</article-title>. <source>Brain Cogn.</source> <volume>106</volume>, <fpage>78</fpage>&#x02013;<lpage>89</lpage>. <pub-id pub-id-type="doi">10.1016/j.bandc.2016.04.011</pub-id><pub-id pub-id-type="pmid">27266394</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shannon</surname> <given-names>C. E.</given-names></name></person-group> (<year>1948</year>). <article-title>A mathematical theory of communication</article-title>. <source>Bell Syst. Tech. J.</source> <volume>27</volume>, <fpage>379</fpage>&#x02013;<lpage>423</lpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smid</surname> <given-names>H.</given-names></name> <name><surname>Jakob</surname> <given-names>A.</given-names></name> <name><surname>Heinze</surname> <given-names>H.-J.</given-names></name></person-group> (<year>1999</year>). <article-title>An event-related brain potential study of visual selective attention to conjunctions of color and shape</article-title>. <source>Psychophysiology</source> <volume>36</volume>, <fpage>264</fpage>&#x02013;<lpage>279</lpage>. <pub-id pub-id-type="pmid">10194973</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spratling</surname> <given-names>M. W.</given-names></name></person-group> (<year>2010</year>). <article-title>Predictive coding as a model of response properties in cortical area V1</article-title>. <source>J. Neurosci.</source> <volume>30</volume>, <fpage>3531</fpage>&#x02013;<lpage>3543</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4911-09.2010</pub-id><pub-id pub-id-type="pmid">20203213</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Squires</surname> <given-names>K.</given-names></name> <name><surname>Petuchowski</surname> <given-names>S.</given-names></name> <name><surname>Wickens</surname> <given-names>C.</given-names></name> <name><surname>Donchin</surname> <given-names>E.</given-names></name></person-group> (<year>1977</year>). <article-title>The effects of stimulus sequence on event related potentials: a comparison of visual and auditory sequences</article-title>. <source>Percept. Psychophys.</source> <volume>22</volume>, <fpage>31</fpage>&#x02013;<lpage>40</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Squires</surname> <given-names>K. C.</given-names></name> <name><surname>Wickens</surname> <given-names>C.</given-names></name> <name><surname>Squires</surname> <given-names>N. K.</given-names></name> <name><surname>Donchin</surname> <given-names>E.</given-names></name></person-group> (<year>1976</year>). <article-title>The effect of stimulus sequence on the waveform of the cortical event-related potential</article-title>. <source>Science</source> <volume>193</volume>, <fpage>1142</fpage>&#x02013;<lpage>1146</lpage>. <pub-id pub-id-type="pmid">959831</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sutton</surname> <given-names>S.</given-names></name> <name><surname>Braren</surname> <given-names>M.</given-names></name> <name><surname>Zubin</surname> <given-names>J.</given-names></name> <name><surname>John</surname> <given-names>E. R.</given-names></name></person-group> (<year>1965</year>). <article-title>Evoked-potential correlates of stimulus uncertainty</article-title>. <source>Science</source> <volume>150</volume>, <fpage>1187</fpage>&#x02013;<lpage>1188</lpage>. <pub-id pub-id-type="pmid">5852977</pub-id></citation>
</ref>
</ref-list>
<app-group>
<app>
<title>A. Details of the model-based analysis</title>
<sec>
<title>A.1. Predictive surprise</title>
<p>Given an event sequence composed of <italic>n</italic> stimuli <italic>E</italic><sub>0</sub>, <italic>E</italic><sub>1</sub>, &#x02026;, <italic>E</italic><sub><italic>n</italic>&#x02212;1</sub> where <inline-formula><mml:math id="M35"><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> for <italic>n</italic>&#x02032; &#x0003D; 0, &#x02026;, <italic>n</italic> &#x02212; 1, surprise concerning the <italic>n</italic>th trial that is based on the stationary or transition probability is defined as follows.</p>
<p>The stationary probability that <italic>E</italic><sub><italic>n</italic></sub> is <italic>X</italic> is denoted by <italic>P</italic><sub><italic>n</italic></sub>(<italic>X</italic>), defined by</p>
<disp-formula id="E1"><label>(A1)</label><mml:math id="M1"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x02223;</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub></mml:mrow><mml:mo>&#x0007D;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x0007D;</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mi>n</mml:mi></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>X</italic> &#x02208; {<italic>a, b</italic>}. Then, predictive stationary surprise <italic>P</italic><sub><italic>n</italic></sub>(<italic>X</italic>) is defined by</p>
<disp-formula id="E2"><label>(A2)</label><mml:math id="M2"><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mtext>S</mml:mtext><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>log</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>The transition probability that <italic>E</italic><sub><italic>n</italic></sub> is <italic>X</italic> is denoted by <italic>P</italic><sub><italic>n</italic></sub>(<italic>X</italic> | <italic>E</italic><sub><italic>n</italic>&#x02212;1</sub>), defined as</p>
<disp-formula id="E3"><label>(A3)</label><mml:math id="M3"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub></mml:mrow><mml:mo>&#x0007D;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x0007D;</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x0007D;</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Then predictive transition surprise <italic>P</italic><sub><italic>n</italic></sub>(<italic>X</italic> | <italic>E</italic><sub><italic>n</italic>&#x02212;1</sub>) is defined as</p>
<disp-formula id="E4"><label>(A4)</label><mml:math id="M4"><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mtext>T</mml:mtext><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>log</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>&#x0007C;</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
</sec>
<sec>
<title>A.2. Regression model</title>
<p>In this section, we describe the regression model summarized in Figure <xref ref-type="fig" rid="F5">5</xref>. We assume that the samples of the response variable are generated with a Gaussian distribution by</p>
<disp-formula id="E5"><label>(A5)</label><mml:math id="M5"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>~</mml:mo><mml:mi mathvariant="-tex-caligraphic">N</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mi>n</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M36"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003BC;</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> represents a Gaussian distribution with mean &#x003BC; and variance &#x003C3;<sup>2</sup>. The parameters of the distribution are assumed to be</p>
<disp-formula id="E6"><label>(A6)</label><mml:math id="M6"><mml:mrow><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E7"><label>(A7)</label><mml:math id="M7"><mml:mrow><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mi>n</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>b</italic><sub>1</sub> and <italic>g</italic><sub>1</sub> are the slopes, <italic>b</italic><sub>2</sub> and <italic>g</italic><sub>2</sub> are the intercepts of the model, and <italic>r</italic><sub><italic>n</italic></sub> and <italic>u</italic><sub><italic>n</italic></sub> are the individual differences of the participants. Let <inline-formula><mml:math id="M37"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> be the set of the indexes of the samples obtained from the participant <italic>p</italic>, where <inline-formula><mml:math id="M38"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02229;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x02205;</mml:mi></mml:math></inline-formula> (<italic>i, j</italic> &#x0003D; 1, &#x02026;, <italic>N</italic><sub><italic>P</italic></sub>, <italic>i</italic> &#x02260; <italic>j</italic>), <inline-formula><mml:math id="M39"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0222A;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0222A;</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>&#x0222A;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>N</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, and <italic>N</italic><sub><italic>P</italic></sub> is the number of participants. Then, <italic>r</italic><sub><italic>n</italic></sub> and <italic>u</italic><sub><italic>n</italic></sub> are defined as</p>
<disp-formula id="E8"><label>(A8)</label><mml:math id="M8"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>p</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x0200A;</mml:mtext><mml:mi>n</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mi mathvariant="-tex-caligraphic">P</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E9"><label>(A9)</label><mml:math id="M9"><mml:mrow><mml:msub><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>u</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>p</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x0200A;</mml:mtext><mml:mi>n</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mi mathvariant="-tex-caligraphic">P</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>Therefore, the unknown parameters for the individual difference in the model are <inline-formula><mml:math id="M40"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x000FB;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, not <inline-formula><mml:math id="M41"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Moreover, we assume the priors for <italic>b</italic><sub>1</sub>, <italic>b</italic><sub>2</sub>, <italic>g</italic><sub>1</sub>, <italic>g</italic><sub>2</sub>, <inline-formula><mml:math id="M42"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M43"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x0015D;</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula> to be</p>
<disp-formula id="E10"><label>(A10)</label><mml:math id="M10"><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>~</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi mathvariant="-tex-caligraphic">N</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="E11"><label>(A11)</label><mml:math id="M11"><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>~</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi mathvariant="-tex-caligraphic">N</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="E12"><label>(A12)</label><mml:math id="M12"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>~</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003B1;</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="E13"><label>(A13)</label><mml:math id="M13"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>~</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003B1;</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="E14"><label>(A14)</label><mml:math id="M14"><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>p</mml:mi></mml:msub><mml:mo>~</mml:mo><mml:mi mathvariant="-tex-caligraphic">N</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>P</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E15"><label>(A15)</label><mml:math id="M15"><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>u</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>p</mml:mi></mml:msub><mml:mo>~</mml:mo><mml:mi mathvariant="-tex-caligraphic">G</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x003B1;</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003B2;</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>P</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M44"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">G</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is a gamma distribution with shape &#x003B1; and scale &#x003B2;. Furthermore, we define the hyper priors for <inline-formula><mml:math id="M45"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and &#x003B2;<sub><italic>u</italic><sub><italic>p</italic></sub></sub> as</p>
<disp-formula id="E16"><label>(A16)</label><mml:math id="M16"><mml:mrow><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>~</mml:mo><mml:mi mathvariant="-tex-caligraphic">U</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E17"><label>(A17)</label><mml:math id="M17"><mml:mrow><mml:msubsup><mml:mi>&#x003C3;</mml:mi><mml:mi>u</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>~</mml:mo><mml:mi mathvariant="-tex-caligraphic">U</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>u</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M46"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is the uniform distribution over the interval [<italic>a, b</italic>]. The undefined parameters for the model were assumed to be constants. The parameters shown in Table <xref ref-type="table" rid="TA1">A1</xref> were set to give a non-informative prior distribution for the parameters.</p>
<table-wrap position="float" id="TA1">
<label>Table A1</label>
<caption><p><bold>The parameters that we consider to be constants in the model (Figure <xref ref-type="fig" rid="F5">5</xref>) and their values</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Parameter</bold></th>
<th valign="top" align="center" style="border-right: thin solid #000000;"><bold>Value</bold></th>
<th valign="top" align="left"><bold>Parameter</bold></th>
<th valign="top" align="center"><bold>Value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">&#x003BC;<sub><italic>b</italic><sub>1</sub></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">0</td>
<td valign="top" align="left">&#x003C3;<sub><italic>b</italic><sub>1</sub></sub></td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left">&#x003BC;<sub><italic>b</italic><sub>2</sub></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">0</td>
<td valign="top" align="left">&#x003C3;<sub><italic>b</italic><sub>2</sub></sub></td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left">&#x003B1;<sub><italic>g</italic><sub>1</sub></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">1</td>
<td valign="top" align="left">&#x003B2;<sub><italic>g</italic><sub>1</sub></sub></td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left">&#x003B1;<sub><italic>g</italic><sub>2</sub></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">1</td>
<td valign="top" align="left">&#x003B2;<sub><italic>g</italic><sub>2</sub></sub></td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left">&#x003BC;<sub><italic>r</italic></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">0</td>
<td valign="top" align="left">&#x003B1;<sub><italic>u</italic></sub></td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left"><italic>a</italic><sub><italic>r</italic></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">0</td>
<td valign="top" align="left"><italic>b</italic><sub><italic>r</italic></sub></td>
<td valign="top" align="center">100</td>
</tr>
<tr>
<td valign="top" align="left"><italic>a</italic><sub><italic>u</italic></sub></td>
<td valign="top" align="center" style="border-right: thin solid #000000;">0.1</td>
<td valign="top" align="left"><italic>b</italic><sub><italic>u</italic></sub></td>
<td valign="top" align="center">100</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We found the unknown parameters for the hierarchical model by sampling with the Markov chain Monte Carlo (MCMC) method (Gelman et al., <xref ref-type="bibr" rid="B17">2013</xref>). In particular, we used the Metropolis&#x02013;Hastings algorithm implemented in PyMC 2.3.6 (Patil et al., <xref ref-type="bibr" rid="B36">2010</xref>) for the sampling. The number of sampling was 100,000 (burn-in: 10,000).</p>
</sec>
</app>
</app-group>
</back>
</article>