<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2017.01225</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Single-Boundary Accumulator Model of Response Times in an Addition Verification Task</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Faulkenberry</surname> <given-names>Thomas J.</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/92973/overview"/>
</contrib>
</contrib-group>
<aff><institution>Department of Psychological Sciences, Tarleton State University</institution> <country>Stephenville, TX, United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Bert Reynvoet, KU Leuven, Belgium</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Catherine Thevenot, Universit&#x000E9; de Gen&#x000E8;ve, Switzerland; Koen Luwel, KU Leuven, Belgium</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Thomas J. Faulkenberry <email>faulkenberry&#x00040;tarleton.edu</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Cognition, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>18</day>
<month>07</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>1225</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>05</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>07</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Faulkenberry.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Faulkenberry</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Current theories of mathematical cognition offer competing accounts of the interplay between encoding and calculation in mental arithmetic. Additive models propose that manipulations of problem format do not interact with the cognitive processes used in calculation. Alternatively, interactive models suppose that format manipulations have a direct effect on calculation processes. In the present study, we tested these competing models by fitting participants&#x00027; RT distributions in an arithmetic verification task with a single-boundary accumulator model (the shifted Wald distribution). We found that in addition to providing a more complete description of RT distributions, the accumulator model afforded a potentially more sensitive test of format effects. Specifically, we found that format affected drift rate, which implies that problem format has a direct impact on calculation processes. These data give further support for an interactive model of mental arithmetic.</p></abstract>
<kwd-group>
<kwd>mental arithmetic</kwd>
<kwd>format effects</kwd>
<kwd>accumulator model</kwd>
<kwd>shifted Wald distribution</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="0"/>
<equation-count count="5"/>
<ref-count count="48"/>
<page-count count="12"/>
<word-count count="9627"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Response times (RTs) have long held a privileged status as one of the primary behavioral measures in cognitive research (Luce, <xref ref-type="bibr" rid="B23">1986</xref>). Their role in inferring mental processes has become so ubiquitous that the justification of their use is rarely questioned. As Luce (<xref ref-type="bibr" rid="B23">1986</xref>) himself put it, &#x0201C;we surely do not understand a choice process very thoroughly until we can account for the time required for it to be carried out&#x0201D; (p. vii). This is particularly evident in the study of mathematical cognition, and in particular, the study of mental arithmetic processes. Since the seminal work of Groen and Parkman (<xref ref-type="bibr" rid="B15">1972</xref>), RTs have provided the primary behavioral signatures used to theorize about the nature of mental calculation. The purpose of the present paper is to extend this work and weigh in on a long-standing debate concerning the independence of encoding and calculation. We accomplish this by fitting distributions of RTs in a mental addition task with a mathematical model known as a shifted Wald distribution and subsequently assessing the effects of format and problem size manipulations on the parameters of these distributions.</p>
<sec>
<title>Models of mental arithmetic</title>
<p>A central question in mathematical cognition concerns the nature of the processes involved in mental arithmetic. Over the years, several competing models of mental arithmetic have been proposed. While most models share a serial architecture of encoding, calculation (which may include retrieval), and production (Ashcraft, <xref ref-type="bibr" rid="B2">1992</xref>), these competing models differ with respect to the proposed independence of these stages. The <italic>abstract code model</italic> (McCloskey, <xref ref-type="bibr" rid="B25">1992</xref>; McCloskey et al., <xref ref-type="bibr" rid="B27">1992</xref>; McCloskey and Macaruso, <xref ref-type="bibr" rid="B26">1995</xref>) proposes separate encoding, calculation, and production modules that communicate via an abstract semantic representation. Each module is specialized for a particular type of input; that is, there are separate encoding modules for verbal numerals (e.g., &#x0201C;four&#x0201D;) and Arabic numerals (e.g., &#x0201C;4&#x0201D;). For example, consider the problem 3 &#x0002B; 4. According to the abstract code model, this problem would be solved by first entering through a comprehension module, specialized for the type of stimulus (in this case, Arabic numerals). After this initial encoding, the problem would then be converted to an amodal, semantic representation (an &#x0201C;abstract code&#x0201D;). Calculation (e.g., retrieval) would operate on this abstract code. The result of the calculation (the answer 7, still in the form of an abstract representation) would then feed into a production module, specialized for the type of production required in the task (either verbal or Arabic).</p>
<p>On the other hand, the <italic>triple code model</italic> (Dehaene, <xref ref-type="bibr" rid="B10">1992</xref>; Dehaene and Cohen, <xref ref-type="bibr" rid="B11">1995</xref>) proposes three separate modules, each specialized for a specific type of representational format: an analog magnitude representation (e.g., a mental number line), an auditory/verbal module, and a visual Arabic numeral module. The triple code model differs from the abstract code model in that calculation and production occur within each module. For comparison, consider again our example 3 &#x0002B; 4. In the triple code model, this problem would be input into one of two modules, each specialized for the type of input code (either an auditory-verbal word frame or a visual-Arabic number form). Calculation would then take place within one of these modules, depending on the nature of the problem. As an example, a retrieval-based calculation to a visually-presented problem (e.g., &#x0201C;3 &#x0002B; 4&#x0201D;) could proceed by first transcoding the input to the auditory-verbal word frame, where the appropriate arithmetic fact could then be retrieved and then the verbal answer produced (e.g., retrieving the answer as &#x0201C;three plus four equals seven&#x0201D; and then verbally producing the answer &#x0201C;seven&#x0201D;).</p>
<p>While the abstract code model and the triple code model differ with respect to the issue of functional vs. representational modularity, they do share a fundamentally <italic>additive</italic> architecture. That is, any performance differences related to problem format (e.g., faster RTs for problems written in Arabic digits compared to words) simply reflect processes related to encoding. In the context of the abstract code model, such performance differences would be explained as the cost of converting a specific stimulus type (digits or words) into an amodal, abstract semantic representation that can be further fed into an appropriate calculation mechanism. In the context of the triple code model, these performance differences would reflect a cost of converting from one representational format (e.g., verbal, word-based representation) into another format (e.g., visual, digit-based representation). Critically, both models predict that format manipulations do not interact with calculation processes.</p>
<p>As an alternative to such additive models of mental arithmetic, Campbell and colleagues (e.g., Campbell and Clark, <xref ref-type="bibr" rid="B4">1988</xref>; Campbell, <xref ref-type="bibr" rid="B3">1994</xref>; Campbell and Epp, <xref ref-type="bibr" rid="B5">2004</xref>) have argued for an <italic>interactive</italic> architecture called the <italic>encoding complex</italic> model. In this model, performance differences due to manipulation of format are posited to stem from a difference in the degree of encoding-retrieval integration. For example, Arabic digits are frequently encountered in the context of calculation, and hence, strong bi-directional pathways are developed between encoding and retrieval of arithmetic facts in this format. However, number words are less frequently encountered, and hence weaker encoding-retrieval connections are formed for such inputs. While sharing some similarities with the additive models described earlier, this model differs in one critical aspect; changes in format are hypothesized to impact <italic>both</italic> encoding and calculation processes.</p>
<p>Support for such an interactive model has primarily appeared in the form of an interaction between the variables of problem format and problem size. As one of the classic &#x0201C;effects&#x0201D; in mathematical cognition, the problem size effect refers to the finding that responses for small problems (e.g., problems for which the sum of the operands is no larger than 10) are significantly faster than responses for larger problems (Ashcraft, <xref ref-type="bibr" rid="B2">1992</xref>; Zbrodoff and Logan, <xref ref-type="bibr" rid="B48">2005</xref>). One possible reason for the problem size effect is that large problems tend to be solved by procedural strategies, resulting in slower and more error prone responses (Campbell and Xue, <xref ref-type="bibr" rid="B8">2001</xref>). While the exact mechanism underlying the problem size effect is still up for debate, the more salient finding is that the problem size effect is larger for problems presented in word format compared to digit format (Campbell and Fugelsang, <xref ref-type="bibr" rid="B6">2001</xref>; Campbell and Penner-Wilger, <xref ref-type="bibr" rid="B7">2006</xref>). Campbell and colleagues have argued that this interaction between problem size and format implies that format directly impacts calculation processes, providing support for the interactive model.</p>
<p>Nonetheless, recent research has not settled the debate regarding the independence of encoding and calculation in mental arithmetic. On one hand, some researchers have argued that people form notation-independent representations of numbers. For example, Libertus et al. (<xref ref-type="bibr" rid="B21">2007</xref>) recorded ERPs (event-related potentials) during a symbolic and nonsymbolic number comparison task. Adults were presented with single numbers (shown either in Arabic digit format or in nonsymbolic dot format) and asked to decide whether each was less than or greater than 15. They found that the amplitude of the P2 component (210&#x02013;250 ms after stimulus presentation) increased as the distance between the number and the comparison standard decreased. Moreover, this pattern did not differ between number formats. This led Libertus et al. (<xref ref-type="bibr" rid="B21">2007</xref>) to conclude that number comparison proceeds via an abstract processing stage that is independent of number format. This finding mirrored previous work by Pinel et al. (<xref ref-type="bibr" rid="B33">2001</xref>), who used fMRI to identify regions in the parietal lobes whose activation was highly correlated with semantic properties of numbers (i.e., numerical distance), but invariant as to whether the number was presented in word or Arabic numeral format.</p>
<p>Similar results have been also found in behavioral experiments. For example, Ganor-Stern and Tzelgov (<xref ref-type="bibr" rid="B14">2008</xref>) used a size-congruity paradigm to investigate automaticity of numerical processing. In this paradigm, numbers are presented in differing physical sizes; this results in pairs of number symbols in which the physical comparison is congruent with numerical size (e.g., small 2 and large 8) or incongruent (e.g., large 2 and small 8). The usual finding is that incongruent pairs take longer than congruent pairs. This size-congruity effect (e.g., Henik and Tzelgov, <xref ref-type="bibr" rid="B17">1982</xref>) is often taken as evidence for automatic processing of number magnitude. In their experiment, Ganor-Stern and Tzelgov (<xref ref-type="bibr" rid="B14">2008</xref>) presented Arabic speakers with number pairs written in two different notations: Arabic and Indian digits. Ganor-Stern and Tzelgov found that even in mixed pairs (one Arabic digit and one Indian digit), there was still a substantial size-congruity effect. They interpreted this result as support for the notion that both notations are automatically converted to a common representation independent of format (see also Ganor-Stern, <xref ref-type="bibr" rid="B13">2009</xref>).</p>
<p>As mentioned earlier, evidence against this additive view of arithmetic processing has been presented by Campbell and colleagues in the form of a substantial problem size by format interaction on response times (e.g., Campbell and Fugelsang, <xref ref-type="bibr" rid="B6">2001</xref>). Some researchers (e.g., No&#x000EB;l et al., <xref ref-type="bibr" rid="B32">1997</xref>) argue that this signature on RTs does not necessarily imply that format has a direct impact on calculation processes. No&#x000EB;l et al. (<xref ref-type="bibr" rid="B32">1997</xref>) argued that such an interaction may be the result of encoding differences between digits and words that feed into the <italic>output</italic> stage, not the calculation stage. Finally, at least one recent study indicates that the interaction between problem size and format may not be as robust as first thought. For example, Megias and Macizo (<xref ref-type="bibr" rid="B28">2016</xref>) failed to find an interaction between problem size and format in a mental arithmetic task<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>. Taken together, these issues warrant further investigation of the processes in volved in mental arithmetic, as it appears that we still have more to learn about the potential interplay between encoding and calculation. In the sections below, we outline a new approach to investigating this issue, based on modeling distributions of RTs in a mental arithmetic task.</p>
</sec>
<sec>
<title>Accumulator models of RT</title>
<p>Most of the studies mentioned above have employed a similar approach to analyzing the effects of experimental manipulations on RTs. Namely, for each participant, the collection of RTs for correct trials in each experimental condition is collapsed to one number, usually the arithmetic mean. This collection of means is then analyzed via an analysis of variance to determine the effect, if any, of each manipulation on RTs. Though popular, this approach is suboptimal for two reasons. First, by collapsing RTs by condition to a single numerical summary (e.g., the mean RT), we lose much information about the <italic>distribution</italic> of RTs. Second, this procedure is usually carried out only on <italic>correct</italic> trials. As such, RTs and response accuracy are analyzed separately, even though they are not necessarily independent (e.g., the speed-accuracy tradeoff, Schouten and Bekker, <xref ref-type="bibr" rid="B41">1967</xref>; Wickelgren, <xref ref-type="bibr" rid="B47">1977</xref>). In both cases, ease of analysis comes at the price of lost information about the original patterns of RTs.</p>
<p>One solution to this problem is to employ a mathematical model such as an <italic>accumulator model</italic>, a model for decision processes that posits a continuous uptake of noisy information that continues until the accumulated evidence exceeds a decision threshold, at which point a response is initiated (Link and Heath, <xref ref-type="bibr" rid="B22">1975</xref>; Luce, <xref ref-type="bibr" rid="B23">1986</xref>; Ratcliff and McKoon, <xref ref-type="bibr" rid="B36">2008</xref>; Ratcliff et al., <xref ref-type="bibr" rid="B38">2016</xref>). One advantage of such an approach is that instead of modeling participants&#x00027; RTs in each experimental condition by a single mean RT, we can fit a model to the entire <italic>distribution</italic> of RTs for each participant in each experimental condition. This results in finding a set of parameters that not only <italic>describes</italic> the distribution mathematically, but also are indicative of the underlying cognitive processes. The advantage is that we can then directly test the effects of our experimental manipulations on the cognitive processes, not just the effects on RTs and/or errors. Hence, RTs and errors become for us a proxy to the underlying cognitive processes, not the sole object of study.</p>
<p>One popular example of a widely used accumulator model is the drift diffusion model of Ratcliff and colleagues (Ratcliff and Murdock, <xref ref-type="bibr" rid="B37">1976</xref>; Ratcliff, <xref ref-type="bibr" rid="B35">1978</xref>; Ratcliff and McKoon, <xref ref-type="bibr" rid="B36">2008</xref>; Ratcliff et al., <xref ref-type="bibr" rid="B38">2016</xref>), which describes a two-choice decision task as the result of such a noisy accumulation process. Specifically, a decision process is modeled as a continuous random walk {<italic>X</italic><sub><italic>t</italic></sub>} with absorbing boundaries 0 and &#x003B1;. This means that the initial term of the walk <italic>X</italic><sub>0</sub> begins somewhere between 0 and &#x003B1; (i.e., 0 &#x0003C; <italic>X</italic><sub>0</sub> &#x0003C; &#x003B1;), and the walk terminates whenever <italic>X</italic><sub><italic>t</italic></sub> &#x0003D; 0 (an incorrect response) or <italic>X</italic><sub><italic>t</italic></sub> &#x0003D; &#x003B1; (a correct response). Moreover, the random walk terms <italic>X</italic><sub><italic>t</italic></sub> tend to <italic>drift</italic> toward one boundary or the other. That is, <inline-formula><mml:math id="M1"><mml:mfrac><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is assumed to be normal with mean &#x003B3;; we refer to &#x003B3; as the <italic>drift rate</italic>. Finally, the decision time <italic>DT</italic> is modeled as the first time <italic>t</italic> for which <italic>X</italic><sub><italic>t</italic></sub> hits either boundary; that is, <italic>X</italic><sub><italic>t</italic></sub> &#x02264; 0 (an incorrect response) or <italic>X</italic><sub><italic>t</italic></sub> &#x02265; &#x003B1; (a correct response). The total response time <italic>RT</italic> is then expressed as <italic>RT</italic> &#x0003D; <italic>DT</italic> &#x0002B; &#x003B8;, where &#x003B8; represents the nondecision component of <italic>RT</italic> (e.g., stimulus encoding and motor execution).</p>
<p>Modeling RT distributions via this diffusion process results in a set of parameters that can be mapped onto underlying latent cognitive processes. The interpretation of these parameters as indices for cognitive processes has been the subject of much investigation over the past 40 years (see Ratcliff and McKoon, <xref ref-type="bibr" rid="B36">2008</xref>, for a review). Though the full Ratcliff diffusion model results in 7 such parameters (Wagenmakers et al., <xref ref-type="bibr" rid="B45">2007</xref>), for simplicity we restrict our discussion to the following three parameters: &#x003B1; (boundary separation), &#x003B3; (drift rate), and &#x003B8; (nondecision time). The boundary separation parameter &#x003B1; represents response caution; a high value of &#x003B1; means that more evidence needs to accrue before a decision can be made. In other words, large values of &#x003B1; reflect conservative decision criteria, whereas small values of &#x003B1; reflect more liberal decision criteria. The drift rate parameter &#x003B3; represents the <italic>quality of information</italic> provided by the stimulus. Larger drift rates reflect unambiguous stimuli, resulting in quicker decisions. Smaller drift rates reflect ambiguous stimuli, resulting in longer decision times. Finally, the nondecision time parameter &#x003B8; reflects encoding and response processes; large values of &#x003B8; reflect slower encoding and/or execution, whereas small values of &#x003B8; reflect fast encoding and/or execution.</p>
<p>While fitting RT distributions with these parameters results in a much more detailed description of the underlying RT distributions than using the mean alone, it is not always possible to fit such a model to experimental data. For instance, one problem with the Ratcliff diffusion model is that it is not well suited to tasks with very low error rates (Anders et al., <xref ref-type="bibr" rid="B1">2016</xref>). Consequently, tasks in which participants perform quite well are not fit well by the diffusion model parameters. However, an alternative accumulator model called the shifted Wald model may be well suited to such situations (Carpenter and Williams, <xref ref-type="bibr" rid="B9">1995</xref>; Schwarz, <xref ref-type="bibr" rid="B42">2001</xref>; Heathcote, <xref ref-type="bibr" rid="B16">2004</xref>). The shifted Wald model is a model of RTs based on the Wald (<xref ref-type="bibr" rid="B46">1947</xref>) distribution, which represents the density of first passage times of a continuous diffusion process that drifts toward a <italic>single</italic> absorbing boundary. Mathematically, the probability density function of the shifted Wald model is given by:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M2"><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>&#x0007C;</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x003B3;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x003B8;</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mo>&#x003B1;</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x003C0;</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B8;</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>3</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>&#x000B7;</mml:mo><mml:mi>exp</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x003B1;</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B3;</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B8;</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mo>&#x003B8;</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where &#x003B1; represents the height of the single response boundary, &#x003B3; represents the drift rate, and &#x003B8; represents a positive (rightward) shift of the entire distribution. As proxies for cognitive processes, each of these parameters has a straightforward interpretation similar to (but not quite equivalent; see Matzke and Wagenmakers, <xref ref-type="bibr" rid="B24">2009</xref>) that of the Ratcliff diffusion model (Schwarz, <xref ref-type="bibr" rid="B42">2001</xref>; Heathcote, <xref ref-type="bibr" rid="B16">2004</xref>; Anders et al., <xref ref-type="bibr" rid="B1">2016</xref>): drift rate &#x003B3; reflects task difficulty via quality of information; response threshold &#x003B1; reflects response caution (amount of information required before response initiation), and &#x003B8; reflects the nondecision time (e.g., encoding and response processes not directly related to the post-encoding decision process).</p>
<p>Aside from the measurement advantages of modeling RTs via distributions rather than via single point means, the shifted Wald model has further methodological and theoretical advantages in the context of mathematical cognition. First, as error rates in mental arithmetic tasks tend to be quite low, fitting RT distributions with the full Ratcliff diffusion model will be difficult. Instead, we can use the shifted Wald model to describe RT distributions for the correct responses. Further, compared to the Ratcliff diffusion model, the shifted Wald model can be fit with a relatively small number of experimental trials. For example, Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>) found that shifted Wald parameters can be recovered for as few as 50 observations per experimental condition. Finally, as the parameters of the shifted Wald distribution have specific cognitive interpretations, systematic variation in these parameters as a function of a stimulus manipulation (e.g., problem format) can tell us the exact locus of the effect. As described above, it is well known that problem format has an effect on RTs &#x02013; problems in word format take longer to solve than problems in digit format. However, it is unclear whether this RT effect is localized to the encoding stage, the calculation stage, or both. By modeling participants&#x00027; RT distributions via shifted Wald models, we can obtain a measure of how format affects each of the parameters &#x003B1;, &#x003B3;, and &#x003B8;. In the present study, we are mainly concerned with the question of whether problem format affects calculation. As such, we can make solid predictions about the effects of our manipulations on the drift rate &#x003B3;, which we assume reflects the calculation process. This assumption comes from the idea that &#x003B3; is a parameter related directly to a decision process, which in this experiment is a decision about the truth of a proposed addition equation. Critically, if format affects drift rate &#x003B3;, this will give evidence that format also has a direct effect on calculation over and above the previously established effects of format on encoding. Such a result would favor the interactive encoding complex model over an additive model (e.g., abstract code model or triple code model), where effects of format are isolated to only the encoding stage. Note that at present, clear predictions cannot be drawn regarding the effects of our manipulations on &#x003B1; and &#x003B8;, so for the purposes of this study, we will focus on the drift rate &#x003B3;.</p>
</sec>
<sec>
<title>The present study</title>
<p>There were two main goals in the present study. First, we sought to extend previous work in mathematical cognition by replicating the arithmetic verification task of Campbell and Fugelsang (<xref ref-type="bibr" rid="B6">2001</xref>) and applying an accumulator model (the shifted Wald model) to model the resulting RT distributions. Whereas many mental arithmetic experiments use a production task, a verification task is advantageous for this modeling approach. In addition to being the task used in Campbell and Fugelsang (<xref ref-type="bibr" rid="B6">2001</xref>), the verification task allows us to measure mental arithmetic processes in the framework of a two-choice decision task, which is the framework employed in most studies that employ accumulator models to study RTs. Second, we aimed to use the results of this modeling to test between two competing models: an additive model, where the stages of problem encoding and answer calculation are functionally independent and effects of problem format are isolated to the encoding stage only, and an interactive encoding-complex model, where effects of problem format are spread between both the encoding stage as well as the answer calculation stage.</p>
</sec>
</sec>
<sec id="s2">
<title>Method</title>
<sec>
<title>Participants</title>
<p>Twenty undergraduate students (15 female, mean age &#x0003D; 25.2 years, age range &#x0003D; 19&#x02013;60 years) participated in this experiment in exchange for partial course credit in their psychology courses. The experiment was reviewed and approved by the institutional review board at Tarleton State University.</p>
</sec>
<sec>
<title>Stimuli and apparatus</title>
<p>Our stimuli were adapted from Campbell and Fugelsang (<xref ref-type="bibr" rid="B6">2001</xref>). Each participant completed 288 experimental trials, consisting of four repetitions of a block of 72 single-digit addition verification problems. We manipulated both problem size and problem format. On even numbered trials, the problems were presented in word format using lower case English words (e.g., &#x0201C;five &#x0002B; seven &#x0003D; twelve&#x0201D;). On odd numbered trials, problems were presented in Arabic digit format (e.g., &#x0201C;5 &#x0002B; 7 &#x0003D; 12&#x0201D;). All problems (regardless of format) were composed of operands between 2 to 9, resulting in a set of 36 problems ranging between 2 &#x0002B; 2 &#x0003D; 4 and 9 &#x0002B; 9 &#x0003D; 18. Note that this assumes that commuted pairs such as 2 &#x0002B; 6 and 6 &#x0002B; 2 are counted as one problem. For each commuted pair, both operand orders were presented with equal frequency throughout the experiment. The order of the operands for each problem was alternated across blocks. Within each of the four blocks, each of these 36 problems was presented once in digit format and once in word format. Problem size was defined in terms of the product of operands: small problems had product operands less than or equal to 25, whereas large problems had operands greater than 25. Within each block, half of the problems of each problem size were presented with the smaller operand first; for the remaining half, the larger operand was presented first.</p>
<p>We further manipulated the truth value of each addition problem. Within each set of 36 problems, half were presented as true equations (e.g., &#x0201C;2 &#x0002B; 4 &#x0003D; 6&#x0201D;) and half were presented as false equations (e.g., &#x0201C;2 &#x0002B; 4 &#x0003D; 7&#x0201D;). Across all four blocks, each addition problem was tested in each format twice in a true equation and twice with a different false answer. False answers were generated pseudo-randomly to be within &#x000B1;4 of the correct answer and never corresponded to either the difference or the product of the operands. Within each set of false answers, each of the numbers 4&#x02013;18 (i.e., the range of true answers) occurred at least once but no more than four times.</p>
<p>All stimuli were presented using Superlab 5.0 (Cedrus Corporation), appearing as white characters against a black background. Responses were recorded using an RB-740 USB response box with &#x000B1;2 ms timing accuracy. The experiment was run on a 21.5-inch iMac desktop computer with 1,024 &#x000D7; 768 screen resolution. Text was displayed in 36 point Lucida Grande font. In each problem, the two operands were separated by a single space on either side of the addition sign. The answer to be verified appeared simultaneously with the problem operands to the right and after the equal sign (e.g., 2 &#x0002B; 4 &#x0003D; 6). No other characters appeared on the screen.</p>
</sec>
<sec>
<title>Procedure</title>
<p>We counterbalanced two response rules across our participants. Even-numbered participants indicated true responses by pressing the rightmost button of the response box and false responses by pressing the leftmost button. Odd-numbered participants used a reversed response mapping: they indicated true responses with the leftmost button and false responses with the rightmost button. Each participant was instructed to respond quickly but accurately.</p>
<p>Prior to the first block, we gave each participant a practice block, consisting of 12 trials in alternating word and digit format, using the operand 0 or 1 paired with 6 randomly selected digits ranging from 0 to 9. At the beginning of each trial, a fixation cross appeared at the center of the screen. When ready to begin, each participant initiated the presentation of the equation with a single button press. The fixation dot flashed for 1 s and was then replaced with one of the 72 addition verification problems. Timing began with the presentation of this equation and ended as soon as the participant pressed a button indicating whether the problem was true or false. After each response, feedback was given in the form of a green C (for correct trials) or red E (for incorrect trials), displayed in the center of the screen for 300ms. After feedback, the fixation cross reappeared, signaling to the participant that the next trial was ready to be initiated. After each block of 72 trials, participants were given an opportunity for a short rest. At the conclusion of the fourth block (288 trials completed), the experiment ended and participants were thanked for their participation.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>Participants completed 5,760 experimental trials. Of these, 394 trials contained an incorrect response (error rate &#x0003D; 6.8%); these trials were removed from further analysis. To facilitate model fitting by removing potential contaminant trials, we removed any trial for which RT was below three and above six median absolute deviations (MAD) from the overall median (median RT &#x0003D; 1,394 ms; MAD &#x0003D; 633 ms) (Leys et al., <xref ref-type="bibr" rid="B20">2013</xref>). This resulted in the removal of an additional 61 trials (1.1%). All subsequent modeling was done on the remaining 5,305 trials.</p>
<p>The general approach to modeling was as follows. First, we modeled true problems (2,656 trials) and false problems (2,649 trials) separately. Within each problem type, trials were divided into 80 design cells defined by the factorial combination of 20 participants with 2 problem size conditions (small, large) and 2 format conditions (digit, word). Afterward, two models were fitted. In the first model we employed the traditional approach where each cell is collapsed to a single mean RT. The effects of problem size and format on these mean RTs were then analyzed using a 2 &#x000D7; 2 analysis of variance, which relies upon null hypothesis significance testing. In addition, we computed Bayes factors using a Bayesian analysis of variance (Rouder et al., <xref ref-type="bibr" rid="B39">2012</xref>); this permitted a quantitative estimation of the extent to which the observed data updated our beliefs in the underlying hypotheses that were tested with the ANOVA (including null effects).</p>
<p>In the second model, we fitted a shifted Wald distribution to the RTs in each design cell using the method of Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>). This resulted in three parameters per cell, (&#x003B1;, &#x003B3;, &#x003B8;); the effects of problems size and format on these parameters were then analyzed using traditional and Bayesian ANOVA. Technical details of the fitting algorithm can be found in the <bold>Appendix</bold>. All modeling was done using R (R Core Team, <xref ref-type="bibr" rid="B34">2016</xref>) and the BayesFactor package (Morey and Rouder, <xref ref-type="bibr" rid="B30">2015</xref>). All raw data and R scripts can be downloaded from the author&#x00027;s GitHub page.</p>
<sec>
<title>Modeling mean RTs</title>
<sec>
<title>True problems</title>
<p>Mean RTs for true problems were submitted to a 2 (problem size: small, large) &#x000D7; 2 (format: digits, words) repeated-measures ANOVA (see left pane of Figure <xref ref-type="fig" rid="F1">1</xref>). As expected, there was a main effect of problem size, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 62.1, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>77</mml:mn></mml:math></inline-formula>. Small problems were verified significantly faster than large problems (1,274 ms vs. 1,704 ms, respectively). There was also a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 219.5, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M4"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>92</mml:mn></mml:math></inline-formula>. Digit problems were verified significantly faster than word problems (1,241 ms vs. 1,737 ms, respectively). The interaction between problem size and format was not significant (<italic>F</italic> &#x0003D; 0.243). A Bayesian ANOVA confirms these results: the best fitting model was the additive model containing factors of problem size and format (<inline-formula><mml:math id="M5"><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>4</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>140</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and this model was preferred over the model containing an interaction term by a factor of 14.9. Using the convention of Jeffreys (<xref ref-type="bibr" rid="B18">1961</xref>), this is considered strong evidence <italic>against</italic> an interaction between problem size and format.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Mean RTs as a function of problem size (small, large), format (digits, words), and truth value (true, false). Error bars represent within-subject 95% confidence intervals as recommended by Morey (<xref ref-type="bibr" rid="B29">2008</xref>).</p></caption>
<graphic xlink:href="fpsyg-08-01225-g0001.tif"/>
</fig>
</sec>
<sec>
<title>False problems</title>
<p>A similar picture emerged for false problems. Mean RTs for false problems were submitted to a 2 (problem size: small, large) &#x000D7; 2 (format: digits, words) repeated-measures ANOVA (see right pane of Figure <xref ref-type="fig" rid="F1">1</xref>). There was a main effect of problem size, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 46.4, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>71</mml:mn></mml:math></inline-formula>. Small problems were verified significantly faster than large problems (1,551 ms vs. 1,903 ms, respectively). There was also a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 178.3, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M7"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>90</mml:mn></mml:math></inline-formula>. Digit problems were verified significantly faster than word problems (1,496 ms vs. 1,958 ms, respectively). As with true problems, the interaction between problem size and format was not significant (<italic>F</italic> &#x0003D; 0.461). A Bayesian ANOVA confirmed that the best fitting model was again the additive model containing factors of problem size and format (<inline-formula><mml:math id="M8"><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>.</mml:mo><mml:mn>9</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>97</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and this model was preferred over the model containing an interaction term by a factor of 13.1.</p>
<p>The picture that emerges from modeling only mean RTs is clear; whereas the expected effects of problem size and format are quite robust, there is strong evidence against an interaction between problem size and format.</p>
</sec>
</sec>
<sec>
<title>Modeling RT distributions</title>
<sec>
<title>True problems</title>
<p>The distributions of RTs for true problems in each design cell [(2 (problem size: small, large) &#x000D7; 2 format: digits, words) &#x000D7; 20 (participants)] were fitted with shifted Wald distributions using the method of Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>) (see <bold>Appendix</bold>). Specifically, this method estimates values for three parameters (&#x003B3;, drift rate; &#x003B1;, response threshold; and &#x003B8;, nondecision time) for each of the 80 design cells. We will first describe the overall model fit, then separately analyze the effects of problem size and format on each of these fitted parameters.</p>
<p>To assess model fit, three diagnostic plots were constructed (see Figure <xref ref-type="fig" rid="F2">2</xref>). The leftmost plot displays a QQ plot comparing observed RT deciles against model-predicted RT deciles. There is no obvious curvature in the plot, which is indicative of a strong model fit. The center plot displays for each RT decile the distribution of standard residuals (difference between observed data deciles and model-predicted deciles, divided by standard deviation of the distribution). The plot indicates that the residual magnitudes tend to increase with RT magnitude. Such behavior is a property of positive-skewed distributions (and in particular, simulated shifted Wald distributions; Anders et al., <xref ref-type="bibr" rid="B1">2016</xref>), and is again indicative of a strong model fit. Finally, the rightmost plot displays a goodness of fit measure &#x00394; for each of the 80 design cells, along with the average cell goodness of fit (<inline-formula><mml:math id="M9"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mo>&#x00394;</mml:mo></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula>), the 5% and 95% quantile range for the &#x00394;-values, the mean standard deviation of the observed data cells &#x003C3;<sub><italic>X</italic></sub>, and the Pearson correlation between &#x00394; and &#x003C3;<sub><italic>X</italic></sub>. The plot and reported values are in line with the recommendations of Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>). Overall, the three diagnostic plots indicate that the data is fit quite well by the shifted Wald model.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Three diagnostic plots for assessing model fit for true problems. The leftmost plot shows a QQ plot comparing observed RT deciles against model-predicted RT deciles. The center plot shows distributions of standardized residuals for each RT decile. The rightmost plot shows overall goodness of fit for each of the 80 fitted cells.</p></caption>
<graphic xlink:href="fpsyg-08-01225-g0002.tif"/>
</fig>
<p>Given that the shifted Wald model is a good fit of the RT distributions, we can proceed with testing the effects of our experimental manipulations (problem size and format) on the three shifted Wald parameters. To this end, we separately submitted each parameter to a 2 (problem size: small, large) &#x000D7; 2 (format: digits, words) repeated-measures ANOVA. The results can be seen in Figure <xref ref-type="fig" rid="F3">3</xref>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Mean shifted Wald parameters for true problems, plotted as a function of problem size (small, large), format (digits, words). <bold>(A)</bold> shows mean drift rate &#x003B3;, <bold>(B)</bold> shows mean response threshold &#x003B1;, and <bold>(C)</bold> shows mean nondecision time &#x003B8;. Error bars represent within-subject 95% confidence intervals as recommended by Morey (<xref ref-type="bibr" rid="B29">2008</xref>).</p></caption>
<graphic xlink:href="fpsyg-08-01225-g0003.tif"/>
</fig>
<p>For drift rate &#x003B3;, there was a main effect of problem size, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 71.6, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M10"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>79</mml:mn></mml:math></inline-formula>. As can be seen in Figure <xref ref-type="fig" rid="F3">3A</xref>, small problems had a significantly larger drift rate (0.07) compared to large problems (0.05). There was also a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 11.1, <italic>p</italic> &#x0003D; 0.003, <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>37</mml:mn></mml:math></inline-formula>. Digit problems exhibited a significantly larger drift rate (0.064) than word problems (0.053). Finally, there was a significant interaction between problem size and format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 7.3, <italic>p</italic> &#x0003D; 0.014, <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula>. As is evident from Figure <xref ref-type="fig" rid="F3">3A</xref>, the effect of format on drift rate was restricted to small problems. A Bayesian ANOVA gives moderate support for this pattern of results, as the best fitting model included the interaction between problem size and format (<inline-formula><mml:math id="M13"><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>6</mml:mn><mml:mo>.</mml:mo><mml:mn>27</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and this model was preferred over the additive-only model by a factor of 3.15.</p>
<p>For response threshold &#x003B1;, there was only a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 62.7, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>77</mml:mn></mml:math></inline-formula>. As can be seen in Figure <xref ref-type="fig" rid="F3">3B</xref>, word problems had a significantly larger mean response threshold (50.7) than digit problems (32.4). No other terms in the ANOVA model were significant (all <italic>F</italic>-values less than 0.86). A Bayesian ANOVA yielded a best fitting model that included only a term for format (<inline-formula><mml:math id="M15"><mml:mi>B</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>7</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and this model was preferred over a model that contained an additional term for problem size by a factor of 4.25.</p>
<p>A similar picture emerges with nondecision time &#x003B8;; again, there was only a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 11.3, <italic>p</italic> &#x0003D; 0.003, <inline-formula><mml:math id="M16"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>37</mml:mn></mml:math></inline-formula>. As can be seen in Figure <xref ref-type="fig" rid="F3">3C</xref>, word problems had a significantly longer mean nondecision time (693 ms) than digit problems (604 ms). No other terms in the ANOVA model were significant (all <italic>F</italic>-values less than 0.47). A Bayesian ANOVA yielded a best fitting model that included only a term for format (<italic>BF</italic><sub>10</sub> &#x0003D; 40.4), and this model was preferred over a model that contained an additional term for problem size by a factor of 2.65.</p>
</sec>
<sec>
<title>False problems</title>
<p>As with true problems, the distributions of RTs for false problems in each design cell [2 (problem size: small, large) &#x000D7; 2 (format: digits, words) &#x000D7; 20 (participants)] were fitted with shifted Wald distributions. As can be seen in Figure <xref ref-type="fig" rid="F4">4</xref>, the three diagnostic plots are again in line with the recommendations of Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>), thus indicating that these data are fit quite well by the shifted Wald model.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Three diagnostic plots for assessing model fit for false problems. The leftmost plot shows a QQ plot comparing observed RT deciles against model-predicted RT deciles. The center plot shows distributions of standardized residuals for each RT decile. The rightmost plot shows overall goodness of fit for each of the 80 fitted cells.</p></caption>
<graphic xlink:href="fpsyg-08-01225-g0004.tif"/>
</fig>
<p>Given the acceptable fit of the shifted Wald model, we submitted each parameter to a 2 (problem size: small, large) &#x000D7; 2 (format: digits, words) repeated-measures ANOVA. The results can be seen in Figure <xref ref-type="fig" rid="F5">5</xref>. For drift rate &#x003B3;, there was a main effect of problem size, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 17.0, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M17"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>47</mml:mn></mml:math></inline-formula>. As can be seen in Figure <xref ref-type="fig" rid="F5">5A</xref>, small problems had a significantly larger drift rate (0.06) compared to large problems (0.045). The main effect of format was not statistically significant, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 3.78, <italic>p</italic> &#x0003D; 0.07, but there was a significant interaction between problem size and format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 5.86, <italic>p</italic> &#x0003D; 0.026, <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>23</mml:mn></mml:math></inline-formula>. Figure <xref ref-type="fig" rid="F5">5A</xref> reveals a similar picture to the situation we saw with true problems; the (albeit marginal) effect of format on drift rate was again restricted to small problems. A Bayesian ANOVA indicated that the best fitting model included the interaction between problem size and format (<italic>BF</italic><sub>10</sub> &#x0003D; 338), and this model was preferred over the model with two main effects (problem size and format) by a factor of 2.90.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Mean shifted Wald parameters for false problems, plotted as a function of problem size (small, large), format (digits, words). <bold>(A)</bold> shows mean drift rate &#x003B3;, <bold>(B)</bold> shows mean response threshold &#x003B1;, and <bold>(C)</bold> shows mean nondecision time &#x003B8;. Error bars represent within-subject 95% confidence intervals as recommended by Morey (<xref ref-type="bibr" rid="B29">2008</xref>).</p></caption>
<graphic xlink:href="fpsyg-08-01225-g0005.tif"/>
</fig>
<p>For response threshold &#x003B1;, we found a significant main effect of problem size, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 10.3, <italic>p</italic> &#x0003D; 0.005, <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>35</mml:mn></mml:math></inline-formula>. As can be seen in Figure <xref ref-type="fig" rid="F5">5B</xref>, large problems exhibited a larger response threshold (47.4) than small problems (39.8). We also saw a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 24.8, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>57</mml:mn></mml:math></inline-formula>; word problems exhibited a larger response threshold (50.9) than digit problems (36.3). The interaction between format and problem size was not significant, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 0.008, <italic>p</italic> &#x0003D; 0.93. A Bayesian ANOVA indicated that the best fitting model was the model with two main effects (problem size and format) (<italic>BF</italic><sub>10</sub> &#x0003D; 701, 677), and this model was preferred over a model that contained an additional interaction term by a factor of 3.22.</p>
<p>The results for nondecision time &#x003B8; mirrored those for true problems. As before, there was only a main effect of format, <italic>F</italic><sub>(1, 19)</sub> &#x0003D; 20.8, <italic>p</italic> &#x0003C; 0.001, <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mo>&#x003B7;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>52</mml:mn></mml:math></inline-formula>. As can be seen in Figure <xref ref-type="fig" rid="F5">5C</xref>, word problems exhibited a significantly longer nondecision time (847 ms) than digit problems (718 ms). No other terms in the ANOVA model were significant (all <italic>F</italic>-values less than 1.6). A Bayesian ANOVA yielded a best fitting model that included only a term for format (<italic>BF</italic><sub>10</sub> &#x0003D; 3, 039), and this model was preferred over a model that contained an additional term for problem size by a factor of 3.60.</p>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>The purpose of the present study was to use a single boundary accumulator model (the shifted Wald distribution) to investigate the independence of encoding and calculation processes in mental arithmetic. Previous research has presented evidence in favor of two competing models: an additive model (Dehaene, <xref ref-type="bibr" rid="B10">1992</xref>; McCloskey, <xref ref-type="bibr" rid="B25">1992</xref>), where encoding processes are isolated from calculation processes, and an interactive model (Campbell and Clark, <xref ref-type="bibr" rid="B4">1988</xref>; Campbell and Epp, <xref ref-type="bibr" rid="B5">2004</xref>), where encoding processes interact with calculation processes. Past studies have attempted to decide between these models by looking for an interaction between the effects of format and problem size on mean RTs. In this paper, we extended this approach and fit a shifted Wald model (Anders et al., <xref ref-type="bibr" rid="B1">2016</xref>) to the distribution of RTs in each experimental condition. This resulted in a collection of three parameters (drift rate, response threshold, and nondecision time) on which we could then test the effects of format and problem size. We found that drift rate was affected by both problem size and format, but response threshold and nondecision time were generally affected only by format. As we will explain below, such results (in particular, the effect of format on drift rate), favor an interactive model of mental arithmetic (e.g., Campbell, <xref ref-type="bibr" rid="B3">1994</xref>).</p>
<p>One of the primary advantages of modeling <italic>distributions</italic> of RTs (compared to analyzing mean RTs alone) is that this modeling approach provides a substantial increase in measurement resolution. To see this, consider that in our experiment, we found no interaction between format and problem size when restricting our attention to mean RTs. If we limited our analysis to this null effect on mean RTs, it would seem that we have found support for an additive model, where encoding and calculation are independent from each other. Moreover, this result would not likely be due to a Type II error (i.e., a false negative), where our null effect would be simply the result of failing to find a significant interaction due to inherently low power. On the contrary, we computed a Bayesian analysis of variance which indicated a Bayes factor of 13.1 in favor of the additive model. This means that after observing the data, we should update the ratio of our belief in the additive model (over the interactive model) by a factor of 13.1, which is considered fairly strong evidence (Jeffreys, <xref ref-type="bibr" rid="B18">1961</xref>). This is a surprising result, as our results certainly do not match the previous findings of Campbell and colleagues (e.g., Campbell and Fugelsang, <xref ref-type="bibr" rid="B6">2001</xref>), who consistently find strong format by problem size interactions.</p>
<p>However, we saw a different picture emerge when we analyzed the <italic>distributions</italic> of RTs in each experimental condition. First, we found that problem size and format interacted to impact drift rate. Specifically, drift rates decreased for as problem size increased. Also, when restricting to small problems, drift rates were larger for digit problems compared to word problems, but this advantage disappeared for large problems. If we assume that our estimated drift rates reflect the cognitive processes related to calculation, the fact that format affects drift rate implies that our format manipulation has a direct impact on calculation. This supports an interactive architecture of mental arithmetic (Campbell, <xref ref-type="bibr" rid="B3">1994</xref>).</p>
<p>In addition to the various effects on drift rate, we saw format effects on response threshold &#x003B1; and nondecision time &#x003B8;. Compared to digit problems, problems presented in word format required more accumulated information before response initiation (i.e., larger response threshold &#x003B1;). Similarly, word problems also exhibited larger nondecision times &#x003B8; than digit problems. One possible explanation is that encoding costs might be incurred when using less familiar word-based stimuli compared to the more familiar digit stimuli. At present, this is speculative, as the format effects on &#x003B8; reflect a general cost of format on processes external to signal accumulation (not just encoding). This leaves open the possibility that format may affect only encoding processes, only response processes, or both. The second option is unlikely, as all existing models of mental arithmetic predict that format affects encoding processes. However, it is not currently clear whether format additionally affects the later processes involved in responding. Several recent studies have indicated that manipulations of number encoding can feed forward into the response phase, at least in simple number decision tasks (e.g., Santens and Verguts, <xref ref-type="bibr" rid="B40">2011</xref>; Faulkenberry et al., <xref ref-type="bibr" rid="B12">2016</xref>; Sobel et al., <xref ref-type="bibr" rid="B43">2016</xref>, <xref ref-type="bibr" rid="B44">2017</xref>). As such, the third option remains viable; hopefully future studies can further address the effects of format on response processes in mental arithmetic.</p>
<p>We also found an interesting, yet unexpected interaction between the effects of format and problem size on drift rate. Specifically, small problems exhibited a large format effect, where drift rates were significantly larger for digits than for words. However, this effect disappeared for large problems. While we can only speculate at this time, this interaction could be due to a difference in solution strategies employed for small and large problems. Indeed, small problems tend to use long term memory retrieval (Ashcraft, <xref ref-type="bibr" rid="B2">1992</xref>; Campbell and Xue, <xref ref-type="bibr" rid="B8">2001</xref>), whereas larger problems tend to use nonretrieval strategies (LeFevre et al., <xref ref-type="bibr" rid="B19">1996</xref>). Our results could reflect the idea that for retrieval-based calculation processes, Arabic digits result in better quality of stimulus information than word format problems (perhaps because the Arabic digit format better matches the way in which such small arithmetic facts were originally learned). However, such format effects could be erased when nonretrieval strategies are used, most likely because the underlying cognitive processes are more complex in this case and are not completely reflected by drift rate. At this point, this is an excellent open question for future research.</p>
<p>We think that these results are an important first step for a new approach to studying problems in mathematical cognition. By fitting the distribution of RTs instead of collapsing all data to a single measure, we were able to capture behavioral phenomena that we would have simply missed by focusing on the traditional mean RT measures. In other words, using accumulator models of RTs results in better measurement fidelity than that obtained by mean RTs, a perspective long advocated in the field of mathematical psychology (e.g., Luce, <xref ref-type="bibr" rid="B23">1986</xref>). A second advantage of our modeling approach is that we were able to get more direct measures of the cognitive processes involved in mathematical decision making. We hope that other researchers will extend this approach to a more general framework for building and testing theories of the cognitive processes involved in mathematical thinking.</p>
<p>We opted to use the shifted Wald distribution (a single boundary accumulator model) in this study, but there is no reason that future studies could not use other accumulator models, such as the diffusion model of Ratcliff and colleagues (Ratcliff and McKoon, <xref ref-type="bibr" rid="B36">2008</xref>; Ratcliff et al., <xref ref-type="bibr" rid="B38">2016</xref>). One reason we opted for the shifted Wald distribution is because the diffusion model is difficult to fit when participants make few errors (Anders et al., <xref ref-type="bibr" rid="B1">2016</xref>). On the other hand, the shifted Wald distribution is perfectly suited to such tasks. The only concession that we had to make is that we had to remove errors from analysis, which prevents us from being able to assess speed-accuracy tradeoffs. Clearly, experiments designed to test predictions about speed-accuracy tradeoffs should consider the more general diffusion framework, which allows modeling both correct and incorrect responses. Another advantage to the shifted Wald distribution is that it requires a relatively small number of trials in each experimental condition. We were able to fit the shifted Wald distribution to our participants&#x00027; data with 72 trials per condition. Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>) showed that shifted Wald parameters can be recovered quite well for as few as 50 observations per condition.</p>
<p>We should note that using the shifted Wald distribution to theorize about cognitive processes should be done with care. While it is tempting to make direct associations between the drift rate, response threshold, and nondecision time of the shifted Wald model and the similarly-defined drift rate, response threshold, and nondecision time of the diffusion model, such a mapping is not directly obvious. For example, Matzke and Wagenmakers (<xref ref-type="bibr" rid="B24">2009</xref>) simulated data using a two-boundary diffusion process, then subsequently fit the data with a single-boundary shifted Wald model. They found that the recovered shifted Wald parameters did not correspond uniquely with the diffusion parameters used to simulate the data. Thus, it is not entirely clear that shifted Wald parameters should be interpreted the same way that diffusion parameters are interpreted. However, Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>) notes that in situations like the ones modeled by Matzke and Wagenmakers (<xref ref-type="bibr" rid="B24">2009</xref>), the shifted Wald would exhibit a poor model fit anyway. Thus, since our data exhibited a reasonably good model fit, we are cautiously confident in our cognitive interpretations of our obtained shifted Wald parameters. Of course, more research is needed in order to better understand when (and how) the shifted Wald distribution can be used as a cognitive process model.</p>
<p>Finally, we note that the choice of task is important to studies in mathematical cognition. A verification task was ideal for the present study, as it framed our mental arithmetic task as a two-choice decision task, which is standard in studies involving accumulator models of RTs. However, it is important to note a verification task might not necessarily be the best reflection of the processes involved in arithmetic. One reason is that decisions might not always be derived from the same calculation processes involved in production. For example, on some problems, participants could rely on a shortcut strategy for detecting false problems, such as knowing that the outcome can only be odd if only one of the addends is odd. As such, the processes involved in this decision would be quite different from the processes involved if the problem was solved by first calculating the answer, then comparing the calculated answer to the one presented. For future studies, it will be important to consider this type of modeling for production tasks as well. We think a single-boundary accumulator model is ideal for this.</p>
<p>In summary, we used a single boundary accumulator model (the shifted Wald distribution) to fit RT distributions in an arithmetic verification task. While we found no interaction between problem format and problem size on mean RTs, we did find that format directly affected drift rate. Thus, we conclude that format affects are not isolated to the encoding stage (as predicted by additive models, e.g., Dehaene, <xref ref-type="bibr" rid="B10">1992</xref>; McCloskey, <xref ref-type="bibr" rid="B25">1992</xref>). Instead, our data supports an interactive model of arithmetic processing (Campbell, <xref ref-type="bibr" rid="B3">1994</xref>), where the effects of problem format extend beyond the encoding stage to have direct impacts on the processes involved in calculation.</p>
</sec>
<sec id="s5">
<title>Ethics statement</title>
<p>This study was carried out in accordance with the recommendations of the APA ethics code, with written informed consent from all subjects. All subjects gave written informed consent in accordance with the Declaration of Helsinki. The protocol was approved by the Institutional Review Board at Tarleton State University.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>TF designed and conducted the experiment, analyzed and modeled the data, and wrote the manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The reviewer KL and handling Editor declared their shared affiliation, and the handling Editor states that the process nevertheless met the standards of a fair and objective review.</p>
</sec>
</sec>
</body>
<back>
<ack><p>The author would like to thank Adam Frampton for his assistance with data collection.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anders</surname> <given-names>R.</given-names></name> <name><surname>Alario</surname> <given-names>F.-X.</given-names></name> <name><surname>Maanen</surname> <given-names>L. V.</given-names></name></person-group> (<year>2016</year>). <article-title>The shifted Wald distribution for 574 response time data analysis</article-title>. <source>Psycholog. Methods</source>, <volume>21</volume>, <fpage>309</fpage>&#x02013;<lpage>327</lpage>. <pub-id pub-id-type="doi">10.1037/575met0000066</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashcraft</surname> <given-names>M. H.</given-names></name></person-group> (<year>1992</year>). <article-title>Cognitive arithmetic: a review of data and theory</article-title>. <source>Cognition</source> <volume>44</volume>, <fpage>75</fpage>&#x02013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(92)90051-i</pub-id><pub-id pub-id-type="pmid">1511587</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name></person-group> (<year>1994</year>). <article-title>Architectures for numerical cognition</article-title>. <source>Cognition</source> <volume>53</volume>, <fpage>1</fpage>&#x02013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(94)90075-2</pub-id><pub-id pub-id-type="pmid">7988104</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name> <name><surname>Clark</surname> <given-names>J. M.</given-names></name></person-group> (<year>1988</year>). <article-title>An encoding-complex view of cognitive number processing: Comment on McCloskey, Sokol, and Goodman (1986)</article-title>. <source>J. Exp. Psychol. Gen</source>. <volume>117</volume>, <fpage>204</fpage>&#x02013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1037/0096-3445.117.2.204</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name> <name><surname>Epp</surname> <given-names>L. J.</given-names></name></person-group> (<year>2004</year>). <article-title>An encoding-complex approach to numerical cognition in Chinese-English bilinguals</article-title>. <source>Can. J. Exp. Psychol.</source> <volume>58</volume>, <fpage>229</fpage>&#x02013;<lpage>244</lpage>. <pub-id pub-id-type="doi">10.1037/h0087447</pub-id><pub-id pub-id-type="pmid">15648727</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name> <name><surname>Fugelsang</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>Strategy choice for arithmetic verification: effects of numerical surface form</article-title>. <source>Cognition</source> <volume>80</volume>, <fpage>B21</fpage>&#x02013;<lpage>B30</lpage>. <pub-id pub-id-type="doi">10.1016/S0010-0277(01)00115-9</pub-id><pub-id pub-id-type="pmid">11274988</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name> <name><surname>Penner-Wilger</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>Calculation latency: the &#x003BC; of memory and the &#x003C4; of transformation</article-title>. <source>Mem. Cogn</source> <volume>34</volume>, <fpage>217</fpage>&#x02013;<lpage>226</lpage>. <pub-id pub-id-type="doi">10.3758/BF03193400</pub-id><pub-id pub-id-type="pmid">16686120</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name> <name><surname>Xue</surname> <given-names>Q.</given-names></name></person-group> (<year>2001</year>). <article-title>Cognitive arithmetic across cultures</article-title>. <source>J. Exp. Psychol. Gen.</source> <volume>130</volume>, <fpage>299</fpage>&#x02013;<lpage>315</lpage>. <pub-id pub-id-type="doi">10.1037/0096-3445.130.2.299</pub-id><pub-id pub-id-type="pmid">11409105</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carpenter</surname> <given-names>R. H. S.</given-names></name> <name><surname>Williams</surname> <given-names>M. L. L.</given-names></name></person-group> (<year>1995</year>). <article-title>Neural computation of log likelihood in control of saccadic eye movements</article-title>. <source>Nature</source> <volume>377</volume>, <fpage>59</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1038/377059a0</pub-id><pub-id pub-id-type="pmid">7659161</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dehaene</surname> <given-names>S.</given-names></name></person-group> (<year>1992</year>). <article-title>Varieties of numerical abilities</article-title>. <source>Cognition</source> <volume>44</volume>, <fpage>1</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(92)90049-n</pub-id><pub-id pub-id-type="pmid">1511583</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dehaene</surname> <given-names>S.</given-names></name> <name><surname>Cohen</surname> <given-names>L.</given-names></name></person-group> (<year>1995</year>). <article-title>Towards an anatomical and functional model of number processing</article-title>. <source>Math. Cogn.</source> <volume>1</volume>, <fpage>83</fpage>&#x02013;<lpage>120</lpage>.</citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Faulkenberry</surname> <given-names>T. J.</given-names></name> <name><surname>Cruise</surname> <given-names>A.</given-names></name> <name><surname>Lavro</surname> <given-names>D.</given-names></name> <name><surname>Shaki</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Response trajectories capture the continuous dynamics of the size congruity effect</article-title>. <source>Acta Psychol.</source> <volume>163</volume>, <fpage>114</fpage>&#x02013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1016/j.actpsy.2015.11.010</pub-id><pub-id pub-id-type="pmid">26647112</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ganor-Stern</surname> <given-names>D.</given-names></name></person-group> (<year>2009</year>). <article-title>Automatic numerical processing is based on an abstract representation</article-title>. <source>Behav. Brain Sci.</source> <volume>32</volume>, <fpage>337</fpage>&#x02013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X09990781</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ganor-Stern</surname> <given-names>D.</given-names></name> <name><surname>Tzelgov</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Across-notation automatic numerical processing</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>34</volume>, <fpage>430</fpage>-<lpage>437</lpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.34.2.430</pub-id><pub-id pub-id-type="pmid">18315418</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Groen</surname> <given-names>G. J.</given-names></name> <name><surname>Parkman</surname> <given-names>J. M.</given-names></name></person-group> (<year>1972</year>). <article-title>A chronometric analysis of simple addition</article-title>. <source>Psychol. Rev.</source> <volume>79</volume>, <fpage>329</fpage>&#x02013;<lpage>343</lpage>. <pub-id pub-id-type="doi">10.1037/h0032950</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heathcote</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Fitting Wald and ex-Wald distributions to response time data: an example using functions for the S-PLUS package</article-title>. <source>Behav. Res. Methods Instrum. Comput.</source> <volume>36</volume>, <fpage>678</fpage>&#x02013;<lpage>694</lpage>. <pub-id pub-id-type="doi">10.3758/BF03206550</pub-id><pub-id pub-id-type="pmid">15641415</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henik</surname> <given-names>A.</given-names></name> <name><surname>Tzelgov</surname> <given-names>J.</given-names></name></person-group> (<year>1982</year>). <article-title>Is three greater than five: the relation between physical and semantic size in comparison tasks</article-title>. <source>Mem. Cogn.</source> <volume>10</volume>, <fpage>389</fpage>&#x02013;<lpage>395</lpage>. <pub-id pub-id-type="doi">10.3758/BF03202431</pub-id><pub-id pub-id-type="pmid">7132716</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jeffreys</surname> <given-names>H.</given-names></name></person-group> (<year>1961</year>). <source>The Theory of Probability, 3rd Edn.</source> <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeFevre</surname> <given-names>J.-A.</given-names></name> <name><surname>Sadesky</surname> <given-names>G. S.</given-names></name> <name><surname>Bisanz</surname> <given-names>J.</given-names></name></person-group> (<year>1996</year>). <article-title>Selection of procedures in mental addition: Reassessing the problem size effect in adults</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn.</source> <volume>22</volume>, <fpage>216</fpage>&#x02013;<lpage>230</lpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.22.1.216</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leys</surname> <given-names>C.</given-names></name> <name><surname>Ley</surname> <given-names>C.</given-names></name> <name><surname>Klein</surname> <given-names>O.</given-names></name> <name><surname>Bernard</surname> <given-names>P.</given-names></name> <name><surname>Licata</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>Detecting outliers: do not use standard deviation around the mean, use absolute deviation around the median</article-title>. <source>J. Exp. Soc. Psychol.</source> <volume>49</volume>, <fpage>764</fpage>&#x02013;<lpage>766</lpage>. <pub-id pub-id-type="doi">10.1016/j.jesp.2013.03.013</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Libertus</surname> <given-names>M. E.</given-names></name> <name><surname>Woldorff</surname> <given-names>M. G.</given-names></name> <name><surname>Brannon</surname> <given-names>E. M.</given-names></name></person-group> (<year>2007</year>). <article-title>Electrophysiological evidence for notation independence in numerical processing</article-title>. <source>Behav. Brain Funct.</source> <volume>3</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.1186/1744-9081-3-1</pub-id><pub-id pub-id-type="pmid">17214890</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Link</surname> <given-names>S. W.</given-names></name> <name><surname>Heath</surname> <given-names>R. A.</given-names></name></person-group> (<year>1975</year>). <article-title>A sequential theory of psychological discrimination</article-title>. <source>Psychometrika</source> <volume>40</volume>, <fpage>77</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1007/BF02291481</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Luce</surname> <given-names>R. D.</given-names></name></person-group> (<year>1986</year>). <source>Response Times: Their Role in Inferring Elementary Mental Organization.</source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matzke</surname> <given-names>D.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name></person-group> (<year>2009</year>). <article-title>Psychological interpretation of the ex-Gaussian and shifted Wald parameters: a diffusion model analysis</article-title>. <source>Psychon. Bull. Rev.</source> <volume>16</volume>, <fpage>798</fpage>&#x02013;<lpage>817</lpage>. <pub-id pub-id-type="doi">10.3758/PBR.16.5.798</pub-id><pub-id pub-id-type="pmid">19815782</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCloskey</surname> <given-names>M.</given-names></name></person-group> (<year>1992</year>). <article-title>Cognitive mechanisms in numerical processing: evidence from acquired dyscalculia</article-title>. <source>Cognition</source> <volume>44</volume>, <fpage>107</fpage>&#x02013;<lpage>157</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0277(92)90052-j</pub-id><pub-id pub-id-type="pmid">1511584</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCloskey</surname> <given-names>M.</given-names></name> <name><surname>Macaruso</surname> <given-names>P.</given-names></name></person-group> (<year>1995</year>). <article-title>Representing and using numerical information</article-title>. <source>Am. Psychol.</source> <volume>50</volume>, <fpage>351</fpage>&#x02013;<lpage>363</lpage>. <pub-id pub-id-type="doi">10.1037/0003-066X.50.5.351</pub-id><pub-id pub-id-type="pmid">7762888</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>McCloskey</surname> <given-names>M.</given-names></name> <name><surname>Macaruso</surname> <given-names>P.</given-names></name> <name><surname>Whetstone</surname> <given-names>T.</given-names></name></person-group> (<year>1992</year>). <article-title>The functional architecture of numerical processing mechanisms: defending the modular model</article-title>, in <source>The Nature and Origins of Mathematical Skills</source>, ed <person-group person-group-type="editor"><name><surname>Camp-bell</surname> <given-names>J. I. D.</given-names></name></person-group> (<publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Elsevier</publisher-name>), <fpage>493</fpage>&#x02013;<lpage>537</lpage>.</citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Megias</surname> <given-names>P.</given-names></name> <name><surname>Macizo</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Activation and selection of arithmetic facts: the role of numerical format</article-title>. <source>Mem. Cogn.</source> <volume>44</volume>, <fpage>350</fpage>&#x02013;<lpage>364</lpage>. <pub-id pub-id-type="doi">10.3758/s13421-015-0559-6</pub-id><pub-id pub-id-type="pmid">26438234</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morey</surname> <given-names>R. D.</given-names></name></person-group> (<year>2008</year>). <article-title>Confidence intervals from normalized data: a correction to cousineau (2005)</article-title>. <source>Tutor. Quant. Methods Psychol.</source> <volume>4</volume>, <fpage>61</fpage>&#x02013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.20982/tqmp.04.2.p061</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Morey</surname> <given-names>R. D.</given-names></name> <name><surname>Rouder</surname> <given-names>J. N.</given-names></name></person-group> (<year>2015</year>). <source>BayesFactor: Computation of Bayes Factors for Common Designs. R package version 0.9.12-2</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=BayesFactor">https://CRAN.R-project.org/package=BayesFactor</ext-link></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nagatsuka</surname> <given-names>H.</given-names></name> <name><surname>Balakrishnan</surname> <given-names>N.</given-names></name></person-group> (<year>2013</year>). <article-title>A method for estimating parameters and quantiles of the three-parameter inverse Gaussian distribution based on statistics invariant to unknown location</article-title>. <source>J. Stat. Comput. Simul.</source> <volume>84</volume>, <fpage>2361</fpage>&#x02013;<lpage>2377</lpage>. <pub-id pub-id-type="doi">10.1080/00949655.2013.795564</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>No&#x000EB;l</surname> <given-names>M.-P.</given-names></name> <name><surname>Fias</surname> <given-names>W.</given-names></name> <name><surname>Brysbaert</surname> <given-names>M.</given-names></name></person-group> (<year>1997</year>). <article-title>About the influence of the presentation format on arithmetical-fact retrieval processes</article-title>. <source>Cognition</source> <volume>63</volume>, <fpage>335</fpage>&#x02013;<lpage>374</lpage>. <pub-id pub-id-type="doi">10.1016/S0010-0277(97)00009-7</pub-id><pub-id pub-id-type="pmid">9265874</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pinel</surname> <given-names>P.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name> <name><surname>Rivi&#x000E8;re</surname> <given-names>D.</given-names></name> <name><surname>LeBihan</surname> <given-names>D.</given-names></name></person-group> (<year>2001</year>). <article-title>Modulation of parietal activation by semantic distance in a number comparison task</article-title>. <source>Neuroimage</source> <volume>14</volume>, <fpage>1013</fpage>&#x02013;<lpage>1026</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.2001.0913</pub-id><pub-id pub-id-type="pmid">11697933</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="web"><person-group person-group-type="author"><collab>R Core Team</collab></person-group> (<year>2016</year>). <source>R: A Language and environment for Statistical Computing</source>. <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Foundation for Statistical Computing</publisher-name>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.R-project.org/">https://www.R-project.org/</ext-link></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ratcliff</surname> <given-names>R.</given-names></name></person-group> (<year>1978</year>). <article-title>A theory of memory retrieval</article-title>. <source>Psychol. Rev.</source> <volume>85</volume>, <fpage>59</fpage>&#x02013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.85.2.59</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ratcliff</surname> <given-names>R.</given-names></name> <name><surname>McKoon</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>The diffusion decision model: theory and data for two-choice decision tasks</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>873</fpage>&#x02013;<lpage>922</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.12-06-420</pub-id><pub-id pub-id-type="pmid">18085991</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ratcliff</surname> <given-names>R.</given-names></name> <name><surname>Murdock</surname> <given-names>B. B.</given-names></name></person-group> (<year>1976</year>). <article-title>Retrieval processes in recognition memory</article-title>. <source>Psychol. Rev.</source> <volume>83</volume>, <fpage>190</fpage>&#x02013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.83.3.190</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ratcliff</surname> <given-names>R.</given-names></name> <name><surname>Smith</surname> <given-names>P. L.</given-names></name> <name><surname>Brown</surname> <given-names>S. D.</given-names></name> <name><surname>McKoon</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Diffusion decision model: current issues and history</article-title>. <source>Trends Cogn. Sci.</source> <volume>20</volume>, <fpage>260</fpage>&#x02013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2016.01.007</pub-id><pub-id pub-id-type="pmid">26952739</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rouder</surname> <given-names>J. N.</given-names></name> <name><surname>Morey</surname> <given-names>R. D.</given-names></name> <name><surname>Speckman</surname> <given-names>P. L.</given-names></name> <name><surname>Province</surname> <given-names>J. M.</given-names></name></person-group> (<year>2012</year>). <article-title>Default Bayes factors for ANOVA designs</article-title>. <source>J. Math. Psychol.</source> <volume>56</volume>, <fpage>356</fpage>&#x02013;<lpage>374</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmp.2012.08.001</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santens</surname> <given-names>S.</given-names></name> <name><surname>Verguts</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>The size congruity effect: is bigger always more?</article-title> <source>Cognition</source> <volume>118</volume>, <fpage>94</fpage>&#x02013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2010.10.014</pub-id><pub-id pub-id-type="pmid">21074146</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schouten</surname> <given-names>J.</given-names></name> <name><surname>Bekker</surname> <given-names>J.</given-names></name></person-group> (<year>1967</year>). <article-title>Reaction time and accuracy</article-title>. <source>Acta Psychol.</source> <volume>27</volume>, <fpage>143</fpage>&#x02013;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1016/0001-6918(67)90054-6</pub-id><pub-id pub-id-type="pmid">6062205</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwarz</surname> <given-names>W.</given-names></name></person-group> (<year>2001</year>). <article-title>The ex-Wald distribution as a descriptive model of response times</article-title>. <source>Behav. Res. Methods Instrum. Comput.</source> <volume>33</volume>, <fpage>457</fpage>&#x02013;<lpage>469</lpage>. <pub-id pub-id-type="doi">10.3758/BF03195403</pub-id><pub-id pub-id-type="pmid">11816448</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sobel</surname> <given-names>K. V.</given-names></name> <name><surname>Puri</surname> <given-names>A. M.</given-names></name> <name><surname>Faulkenberry</surname> <given-names>T. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Bottom-up and top-down attentional contributions to the size congruity effect</article-title>. <source>Attent. Percept. Psychophys.</source> <volume>78</volume>, <fpage>1324</fpage>&#x02013;<lpage>1336</lpage>. <pub-id pub-id-type="doi">10.3758/s13414-016-1098-3</pub-id><pub-id pub-id-type="pmid">27052836</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sobel</surname> <given-names>K. V.</given-names></name> <name><surname>Puri</surname> <given-names>A. M.</given-names></name> <name><surname>Faulkenberry</surname> <given-names>T. J.</given-names></name> <name><surname>Dague</surname> <given-names>T. D.</given-names></name></person-group> (<year>2017</year>). <article-title>Visual search for conjunctions of physical and numerical size shows that they are processed independently</article-title>. <source>J. Exp. Psychol. Hum. Percept. Perform.</source> <volume>43</volume>, <fpage>444</fpage>&#x02013;<lpage>453</lpage>. <pub-id pub-id-type="doi">10.1037/xhp0000323</pub-id><pub-id pub-id-type="pmid">27893271</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name> <name><surname>Maas</surname> <given-names>H. L. J. V. D.</given-names></name> <name><surname>Grasman</surname> <given-names>R. P. P. P.</given-names></name></person-group> (<year>2007</year>). <article-title>An EZ-diffusion model for response time and accuracy</article-title>. <source>Psychon. Bull. Rev.</source> <volume>14</volume>, <fpage>3</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.3758/BF03194023</pub-id><pub-id pub-id-type="pmid">17546727</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wald</surname> <given-names>A.</given-names></name></person-group> (<year>1947</year>). <source>Sequential Analysis.</source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Wiley</publisher-name>.</citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wickelgren</surname> <given-names>W. A.</given-names></name></person-group> (<year>1977</year>). <article-title>Speed-accuracy tradeoff and information processing dynamics</article-title>. <source>Acta Psychol.</source> <volume>41</volume>, <fpage>67</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1016/0001-6918(77)90012-9</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zbrodoff</surname> <given-names>N. J.</given-names></name> <name><surname>Logan</surname> <given-names>G. D.</given-names></name></person-group> (<year>2005</year>). <article-title>What everyone finds: the problem size effect</article-title>, in <source>Handbook of Mathematical Cognition</source>, ed <person-group person-group-type="editor"><name><surname>Campbell</surname> <given-names>J. I. D.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Psychology Press</publisher-name>), <fpage>331</fpage>&#x02013;<lpage>346</lpage>.</citation></ref>
</ref-list>
<app-group>
<app id="A1">
<title>Appendix</title>
<sec>
<title>Fitting the shifted wald model</title>
<p>The approach used for fitting the shifted Wald model in this paper is similar to that first reported in Anders et al. (<xref ref-type="bibr" rid="B1">2016</xref>). The idea is to map the raw data onto three parameter estimates <inline-formula><mml:math id="M22"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M23"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M24"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> that produce predicted RTs that are as close as possible to the actual RTs. This is done in the following manner.</p>
<p>Let <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> be a set of <italic>M</italic> independent observations; in our case, this would be a set of RTs for <italic>M</italic> trials. Let <inline-formula><mml:math id="M26"><mml:mo>&#x003B2;</mml:mo><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Given a value for &#x003B2;, one can iteratively calculate maximum likelihood estimators (MLEs) for <inline-formula><mml:math id="M27"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M28"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> (Nagatsuka and Balakrishnan, <xref ref-type="bibr" rid="B31">2013</xref>). To begin, we first compute a &#x0201C;seed&#x0201D;:</p>
<disp-formula id="E2"><label>(A1)</label><mml:math id="M29"><mml:msubsup><mml:mover accent='true'><mml:mi>a</mml:mi><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover><mml:mn>0</mml:mn><mml:mo>*</mml:mo></mml:msubsup><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover accent='true'><mml:mi>X</mml:mi><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>M</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mover accent='true'><mml:mi>X</mml:mi><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac></mml:mrow></mml:msqrt></mml:math></disp-formula>
<p>Then, using this seed, we can obtain the MLEs for <inline-formula><mml:math id="M30"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M31"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> using the equations</p>
<disp-formula id="E3"><label>(A2)</label><mml:math id="M32"><mml:mover accent='true'><mml:mo>&#x003B8;</mml:mo><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mtext>min</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msubsup><mml:mover accent='true'><mml:mi>a</mml:mi><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover><mml:mn>0</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:msup><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x0221E;</mml:mo></mml:msubsup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy='false'>[</mml:mo><mml:mi>z</mml:mi><mml:mo>;</mml:mo><mml:mover accent='true'><mml:mo>&#x003B2;</mml:mo><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mstyle><mml:mi>M</mml:mi></mml:msup><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E4"><label>(A3)</label><mml:math id="M33"><mml:mover accent='true'><mml:mo>&#x003B1;</mml:mo><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>M</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:munderover><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mover accent='true'><mml:mo>&#x003B8;</mml:mo><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mstyle><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mover accent='true'><mml:mi>X</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:mover accent='true'><mml:mo>&#x003B8;</mml:mo><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>
<p>where <italic>F</italic>(&#x000B7;) is the cumulative distribution function of the shifted Wald distribution (Equation 1). Specifically, the initial seed <inline-formula><mml:math id="M34"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> obtained in Equation (A1) is input into Equation (A2) in place of <inline-formula><mml:math id="M35"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, which then produces an estimate <inline-formula><mml:math id="M36"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>. We then compute <inline-formula><mml:math id="M37"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> via Equation (A3) by setting <inline-formula><mml:math id="M38"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>. Afterward, this value of <inline-formula><mml:math id="M39"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is used to compute <inline-formula><mml:math id="M40"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> and then <inline-formula><mml:math id="M41"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> via Equations (A2) and (A3), respectively. Finally, <inline-formula><mml:math id="M42"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> can be computed easily as <inline-formula><mml:math id="M43"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mo>&#x003B2;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>.</p>
<p>From these MLEs (<inline-formula><mml:math id="M44"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B1;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M45"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B3;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M46"><mml:mover accent="true"><mml:mrow><mml:mo>&#x003B8;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>), one uses the Wald pdf (Equation 1) to calculate a set of <italic>M</italic> predicted quantiles <inline-formula><mml:math id="M47"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> which can then be compared to the observed quantiles (from our data) <inline-formula><mml:math id="M48"><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> via an <italic>l</italic><sub>1</sub>-norm:</p>
<disp-formula id="E5"><label>(A4)</label><mml:math id="M49"><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mi>P</mml:mi></mml:msup><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>Q</mml:mi><mml:mi>O</mml:mi></mml:msup><mml:msub><mml:mo>&#x02016;</mml:mo><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:munderover><mml:mo stretchy='false'>&#x0007C;</mml:mo></mml:mstyle><mml:msubsup><mml:mi>Q</mml:mi><mml:mi>k</mml:mi><mml:mi>P</mml:mi></mml:msubsup><mml:mo>&#x02212;</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mi>k</mml:mi><mml:mi>O</mml:mi></mml:msubsup><mml:mo stretchy='false'>&#x0007C;</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>Thus, for every chosen value of &#x003B2;, we obtain a deviance measure <inline-formula><mml:math id="M50"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></inline-formula> This procedure is repeated over a plausible range of values for &#x003B2; (e.g., 0.01 &#x0003C; &#x003B2; &#x0003C; 0.99), and the parameter set that minimizes <inline-formula><mml:math id="M51"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is chosen as the fit.</p>
</sec>
</app>
</app-group>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>Though they did not report any inferential statistics for this specific interaction, they did report that problem size did not interact with any other variable (including format), with all <italic>p</italic>-values greater than 0.12 (p. 356).</p></fn>
</fn-group>
</back>
</article>