<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2017.01847</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Hypothesis and Theory</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>From Discovery to Justification: Outline of an Ideal Research Program in Empirical Psychology</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Witte</surname> <given-names>Erich H.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/483872/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zenker</surname> <given-names>Frank</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/379981/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Social and Economic Psychology, University of Hamburg</institution>, <addr-line>Hamburg</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Philosophy and Cognitive Science, Lund University</institution>, <addr-line>Lund</addr-line>, <country>Sweden</country></aff>
<aff id="aff3"><sup>3</sup><institution>Institute of Philosophy, Slovak Academy of Sciences (SAS)</institution>, <addr-line>Bratislava</addr-line>, <country>Slovakia</country></aff>
<aff id="aff4"><sup>4</sup><institution>Philosophy, Konstanz University</institution>, <addr-line>Konstanz</addr-line>, <country>Germany</country></aff>
<aff id="aff5"><sup>5</sup><institution>Institute of Logic and Cognition, Sun Yat-sen University</institution>, <addr-line>Guangzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Holmes Finch, Ball State University, United States</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Evgueni Borokhovski, Concordia University, Canada; Paul T. Barrett, Advanced Projects R&#x0026;D Ltd., New Zealand</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>Frank Zenker, <email>frank.zenker@fil.lu.se</email></italic></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Quantitative Psychology and Measurement, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>10</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>1847</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>09</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>10</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2017 Witte and Zenker.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Witte and Zenker</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The gold standard for an empirical science is the replicability of its research results. But the estimated average replicability rate of key-effects that top-tier psychology journals report falls between 36 and 39% (objective vs. subjective rate; <xref ref-type="bibr" rid="B50">Open Science Collaboration, 2015</xref>). So the standard mode of applying null-hypothesis significance testing (NHST) fails to adequately separate stable from random effects. Therefore, NHST does not fully convince as a statistical inference strategy. We argue that the replicability crisis is &#x201C;home-made&#x201D; because more sophisticated strategies can deliver results the successful replication of which is sufficiently probable. Thus, we can overcome the replicability crisis by integrating empirical results into genuine research programs. Instead of continuing to narrowly evaluate only the stability of data against random fluctuations (<italic>discovery context</italic>), such programs evaluate rival hypotheses against stable data (<italic>justification context</italic>).</p>
</abstract>
<kwd-group>
<kwd>confirmation</kwd>
<kwd>knowledge accumulation</kwd>
<kwd>meta-analysis</kwd>
<kwd>psi-hypothesis</kwd>
<kwd>replicability crisis</kwd>
<kwd>research programs</kwd>
<kwd>significance-test</kwd>
<kwd>test-power</kwd>
</kwd-group>
<contract-num rid="cn001">90 531</contract-num>
<contract-num rid="cn002">1225/02/03</contract-num>
<contract-sponsor id="cn001">Volkswagen Foundation<named-content content-type="fundref-id">10.13039/501100001663</named-content></contract-sponsor>
<contract-sponsor id="cn002">Seventh Framework Programme<named-content content-type="fundref-id">10.13039/501100004963</named-content></contract-sponsor>
<contract-sponsor id="cn003">Ragnar S&#x00F6;derbergs stiftelse<named-content content-type="fundref-id">10.13039/100007459</named-content></contract-sponsor>
<counts>
<fig-count count="2"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="85"/>
<page-count count="12"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>Empirical psychology and the social sciences at large remain in crisis today, because (too) many key-results cannot be replicated (<xref ref-type="bibr" rid="B3">Baker, 2015</xref>; <xref ref-type="bibr" rid="B50">Open Science Collaboration, 2015</xref>; <xref ref-type="bibr" rid="B16">Etz and Vandekerckhove, 2016</xref>). Having diagnosed a disciplinary crisis as early as <xref ref-type="bibr" rid="B73">Willy (1889)</xref>, psychologists did so most recently in a special issue of <italic>Perspectives on Psychological Science</italic> (<xref ref-type="bibr" rid="B76">Witte, 1996a</xref>; <xref ref-type="bibr" rid="B51">Pashler and Wagenmakers, 2012</xref>; <xref ref-type="bibr" rid="B63">Spellman, 2012</xref>; <xref ref-type="bibr" rid="B66">Sturm and M&#x00FC;lberger, 2012</xref>). Particularly the crisis of significance-testing is about as old as the test itself (<xref ref-type="bibr" rid="B74">Witte, 1980</xref>; <xref ref-type="bibr" rid="B13">Cowles, 1989</xref>; <xref ref-type="bibr" rid="B28">Harlow et al., 1997</xref>). Setting the current crisis apart is the insight that null-hypothesis significance testing (NHST) has broadly failed to deliver the stable effects that should characterize empirical knowledge. Many researchers are therefore (rightly) concerned that <italic>all</italic> published effects are under doubt. The perhaps most pressing question today is how our field might regain trust.<sup><xref ref-type="fn" rid="fn01">1</xref></sup></p>
<p>In our view, the ongoing replicability crisis reflects a goal-conflict between publishing statistically significant results as an individual researcher and increasing the trustworthiness of scientific knowledge as a community (<xref ref-type="bibr" rid="B4">Bakker et al., 2012</xref>; <xref ref-type="bibr" rid="B49">Nosek et al., 2012</xref>; <xref ref-type="bibr" rid="B32">Ioannidis, 2014</xref>). Acknowledging that we can separate the corresponding research activities only analytically, we map both goals onto the terms &#x2018;discovery&#x2019; and &#x2018;justification&#x2019; (aka &#x2018;DJ-distinction&#x2019;). Since the <italic>status quo</italic> favors &#x201C;making discoveries,&#x201D; we submit, the balance between these goals must be redressed. As regards variously proposed &#x201C;minimally invasive&#x201D; remedies, however, we find that such &#x201C;soft&#x201D; measures are insufficient to regain trust.<sup><xref ref-type="fn" rid="fn02">2</xref></sup></p>
<p>We rest our case on the observation that psychologists typically deploy statistical inference methods in underpowered studies (<xref ref-type="bibr" rid="B41">Maxwell, 2004</xref>). This praxis generates theoretically disconnected &#x201C;one-off&#x201D; discoveries whose replication is improbable. But such results <italic>should</italic> not be trusted, because they are insufficiently stable to justify, or corroborate, a theoretical hypothesis (<xref ref-type="bibr" rid="B53">Rosnow and Rosenthal, 1989</xref>). By contrast, corroboration <italic>is</italic> possible within a research program. To overcome the crisis, therefore, our community should come to coordinate itself on <italic>joint</italic> long-term research endeavors.</p>
</sec>
<sec><title>From Discovery to Justification</title>
<sec><title>Overview</title>
<p>Constructing a psychological theory begins with discovering non-random relations between antecedent variables and their (causal) consequences, aka <italic>stable</italic> effects. Relying on a probabilistic version of the lean DJ-distinction, this section contrasts the discovery and the justification context, explains the replicability of empirical results, and shows why underpowered discoveries cannot be trusted. We then define two key concepts: <italic>induction quality of data</italic> and <italic>corroboration quality of hypotheses</italic>, and formulate a brief upshot.</p>
</sec>
<sec><title>Stable Effects</title>
<p>As late 19th-century psychologists transformed their field into an empirical science, the guiding idea was to base empirical hypotheses on stable non-random effects. Indeed, only stable effects guide researchers toward the <italic>explananda</italic>. Otherwise, explanation would be pointless&#x2014;for what should be explained? This pedestrian insight makes the discovery of a stable effect as necessary as the managing of random influences is regularly unavoidable (e.g., in measurement, sampling, or situation-construction). We therefore discover an effect but in a <italic>probabilistic</italic> sense.</p>
<p>As we thus evaluate our chances of having made a discovery, we must gauge the effect&#x2019;s deviation from random against a statistical significance threshold. Of course, if we cannot discover an effect with certainty, then it follows by parity of reasoning that we cannot falsify it with certainty either. We must therefore ever invest <italic>some</italic> trust that the effect in fact surpasses potential random influences.</p>
<p>Also known as &#x2018;stable observations,&#x2019; such effects register as highly probable deviations from a content-free random or null-hypothesis (H<sub>0</sub>). But even a stable effect (in this probabilistic sense) may be subsumed under distinct explanatory hypotheses. We must therefore corroborate such diverging explanations via a theory that predicts the effect&#x2019;s probabilistic signature from initial and boundary conditions.</p>
<p>If we base hypothesis construction and validation on uncertain observation, then hypothesis corroboration likewise entails the probabilistic comparison of a hypotheses pair. <italic>Pace</italic> efforts by the likes of Rudolf Carnap and Karl Popper, this is an immediate consequence of recognizing that uncontrollable random influences are relevant to theory acceptance. As we now show, it is this interplay between an effect&#x2019;s probabilistic discovery and its probabilistic corroboration that connects the discovery and the justification contexts (aka DJ-distinction).</p>
</sec>
<sec><title>The Lean DJ-Distinction</title>
<p>Discovery and justification are fairly self-evident concepts. The dominant mode of deploying statistical inference methods nevertheless reflects them inadequately. Indeed, a review of the textbook literature would show that the NHST approach largely ignores the DJ-distinction. To set this right, we rely on a probabilistic model to make a non-controversial version of the DJ-distinction precise.<sup><xref ref-type="fn" rid="fn03">3</xref></sup> <xref ref-type="bibr" rid="B30">Hoyningen-Huene (2006)</xref> calls it the &#x2018;lean DJ-distinction&#x2019; to denote &#x201C;an abstract distinction between the factual [&#x2026;] and the normative or evaluative [&#x2026;]&#x201D; (ibid., p. 128).</p>
<p>In adapting this distinction, &#x2018;discovery&#x2019; refers to research activities in the <italic>data space</italic> that employ probabilities, while &#x2018;justification&#x2019; denotes activities in the <italic>hypotheses space</italic> that employ likelihoods. Unlike a measure of the probability, <italic>P</italic>, of data given a hypothesis [where 0 &#x2264;<italic>P</italic>(D,H) &#x2264; 1], likelihoods, <italic>L</italic>, aka &#x2018;inverse probabilities,&#x2019; are a sort of probability measure for hypotheses given data [where 0 &#x003C; <italic>L</italic>(H&#x007C;D) &#x003C; &#x221E;]. &#x2018;Data space&#x2019; and &#x2018;hypotheses space&#x2019; are analytical constructs, respectively denoting the collection of stable data and their subsumption under hypotheses.</p>
<p>Of course, research activities alternate between both spaces, witness metaphorical ideas, and such practices as data-&#x201C;torturing&#x201D; or inventing <italic>ad hoc</italic> hypotheses, etc. Similar heuristics are fine, but they do not amount to a hypothesis test. After all, discovery context activities focus on phenomena we are yet to discover. So a probabilistic model defines a discovery as a (theory-laden) observation of a non-random effect. This yields as relevant elements the H<sub>0</sub>-hypothesis and the actual data distribution. If data have a low probability given the H<sub>0</sub> (aka the effect&#x2019;s <italic>p</italic>-value; see <xref ref-type="bibr" rid="B20">Fisher, 1956</xref>), this is called a &#x2018;discovered effect.&#x2019;</p>
<p>But when we next evaluate the <italic>trustworthiness</italic> of data, the crucial question is this: if we treat the data distribution that had been obtained in a given test-condition as a theoretical parameter, and moreover hold constant the <italic>p</italic>-value and the number of observations, what is the probability of obtaining the <italic>same</italic> distribution given the H<sub>0</sub> in a subsequent instance of this test-condition?<sup><xref ref-type="fn" rid="fn04">4</xref></sup> So to assess the trustworthiness of data <italic>is</italic> to measure the probability of replicating a non-random effect. This leads us to consider test-power.</p>
</sec>
<sec><title>Test-Power and Statistical Significance vs. Theoretical Importance</title>
<p><xref ref-type="bibr" rid="B11">Cohen (1962)</xref> had suggested that empirical studies in the social sciences tend to be underpowered. This entails a <italic>low</italic> probability of successfully replicating a non-random effect. His guiding assumption was that the true amount of influence from independent onto dependent variables be of medium size (<italic>d</italic> = 0.50).</p>
<p>Provided a typical sample size around n1 = n2 = 30, if we assume <italic>d</italic> = 0.50 and a two-sided &#x03B1;-error = 0.05, then average test-power comes to 1&#x2013;&#x03B2;-error = 0.46. Some 25 years later, <xref ref-type="bibr" rid="B61">Sedlmeier and Gigerenzer (1989)</xref> arrived at a similar value. (The &#x03B1;-error denotes the chance of obtaining a false positive test-result, the &#x03B2;-error the chance of obtaining a false negative result; both errors are normally non-zero, and should be small for data to be trustworthy.) Test-power = 0.46 implies that samples are typically too small to expect a <italic>stable</italic> effect. Therefore, empirical results obtained with similar test-power may at best issue invitations to more closely study such &#x201C;discoveries.&#x201D;</p>
<p>In psychology as elsewhere in the social sciences, however, researchers tend to over-report such underpowered results as hypothesis <italic>confirmations</italic>. Already some 50 years ago, this praxis was identified as a cause of publication bias (<xref ref-type="bibr" rid="B65">Sterling, 1959</xref>). To more fully appreciate why this overstates the capabilities of the method applied, consider that NHST does normally not specify the H<sub>1</sub>. One thus fails to assign precise semantic content to it. By contrast, the (random) H<sub>0</sub> does tend to be well-specified.<sup><xref ref-type="fn" rid="fn05">5</xref></sup></p>
<p>If data now display a sufficiently large deviation from random, this is (erroneously) interpreted as a discovery of theoretical importance. But on the reasonable assumption that theoretically important discoveries are stable rather than random, a <italic>necessary</italic> condition in order to meaningfully speak of a theoretically important discovery is to have properly employed a reproducibility measure such as test-power. This in turn presupposes a specified effect size, which the standard mode of deploying NHST, however, cannot offer. So NHST cannot warrant an immediate transition from &#x2018;statistically significant effect&#x2019; to &#x2018;theoretically important discovery.&#x2019;</p>
<p>Many researchers appear to be satisfied knowing that test-power is maximal, but fail to check its exact value. So they do not <italic>properly</italic> deploy a reproducibility measure. The one &#x201C;good&#x201D; reason for this is their failure to specify the H<sub>1</sub>, because only its specification renders test-power quantifiable. This failure contributes to the confidence loss in our research results. Understandably so, too, for some 50 years after Cohen had hinted at <italic>d</italic> = 0.46, test-power typically registers even lower, at <italic>d</italic> = 0.35. So <italic>most</italic> studies are underpowered (<xref ref-type="bibr" rid="B4">Bakker et al., 2012</xref>; <xref ref-type="bibr" rid="B3">Baker, 2015</xref>).</p>
<p>Against the background of our probabilistic model, we proceed to explain why so many empirical results should not be trusted.</p>
</sec>
<sec><title>Trustworthy Discoveries</title>
<p>In his influential textbook Cohen had recommended that &#x201C;[w]hen the investigator has no other basis for setting the desired power value, the value [1&#x2013;&#x03B2;-error = 0.80] is used&#x201D; (<xref ref-type="bibr" rid="B12">Cohen, 1977</xref>, p. 56). This set a widely accepted standard. Together with &#x03B1; = 0.05, it implies a weighing of epistemological values that makes the preliminarily discovery of an effect <italic>four</italic> times more important than its stable replication. (This alone goes some way toward explaining 1&#x2013;&#x03B2;-error = 0.35.) Similarly-powered &#x201C;discoveries&#x201D; thus are typically instable.</p>
<p>It follows that few effects which arise in the discovery context are <italic>known</italic> to deserve a theoretical explanation. Hence, the unsophisticated application of NHST as a discovery method regularly fails to inform the evaluation of hypotheses in the justification context. Since justification contexts activities aim at developing the comparatively best-corroborated hypotheses into theories, it can hardly surprise that psychology offers so few genuine theories (we return to this in Section &#x201C;Precise Theoretical Constructs?&#x201D;).</p>
<p>To set this right, discovery context activities must establish non-random effects that also feature a high replication probability, i.e., <italic>stable</italic> effects. Yet, the current research- and publication-praxis does not fully reflect that insight. For instance, though &#x201C;mere&#x201D; replications do serve to evaluate whether an effect is stable, until recently one could not publish such work in a top-tier journal (see <xref ref-type="bibr" rid="B29">Holcombe, 2016</xref> on <italic>Perspectives on Psychological Science</italic>&#x2019;s new replication section).</p>
<p>In summary, we can quantify the replication probability of data only if we specify <italic>two</italic> point hypotheses (H<sub>0</sub>, H<sub>1</sub>). To improve the trustworthiness of empirical knowledge, we must therefore increase the precision of theoretical assumptions (<xref ref-type="bibr" rid="B35">Klein, 2014</xref>). For only this yields knowledge of test-power, and only then has the comparative corroboration of the H<sub>1</sub> by a data-set D (as compared to a rival H<sub>0</sub>) <italic>not</italic> been indirectly deduced from our estimation of D given the H<sub>0</sub>.</p>
<p>This puts us in a position to offer two central definitions.</p>
</sec>
<sec><title>Induction Quality of Data, Corroboration Quality of Hypotheses</title>
<p>Our probabilified version of the lean DJ-distinction suggests that two analytically distinct activities govern the research process. Roughly, one first creates an empirical set-up serving as a test-condition to obtain data of sufficient induction quality (discovery context). Next, one tests point-hypotheses against such data (justification context).</p>
<p>In more detail, discovery context activities evaluate data by means of descriptive and inferential statistics given fixed hypotheses. Since this gauges <italic>induction quality of data</italic>, we perform an evaluation in the data space. (In fact, we proceed hypothetically, effectively assessing if data would be sufficiently trustworthy if they were obtained.) Exactly this is expressed by &#x2018;gauging the probability that data are replicable given two point-hypotheses.&#x2019;</p>
<p>Proceeding to the justification context, we now evaluate point-hypotheses in order to gauge <italic>corroboration quality of hypotheses</italic> given data. So we perform an evaluation in the hypotheses space. Crucially, only if we in fact obtain data of sufficient induction quality can we properly quantify the inductive support that actual data lend to a hypothesis. So &#x2018;gauging corroboration quality&#x2019; refers to evaluating the degree to which probably replicable data support one hypothesis more than another.</p>
<p>Before, we apply these distinctions in the next section, we can define as follows:</p>
<list list-type="simple" prefix-word="simple">
<list-item><p><italic>Def. induction quality</italic>: A measure of the sensitivity of an empirical set-up (given two specified point-hypotheses and a fixed sample size) that is stated as &#x03B1;- and &#x03B2;-error. Though a set-up&#x2019;s acceptability rests on convention, equating both errors (&#x03B1; = &#x03B2;) avoids a bias <italic>pro</italic> detection (&#x03B1;-error) and <italic>con</italic> replicability (&#x03B2;-error). Currently, &#x03B1; = &#x03B2; = 0.05 or &#x03B1; = &#x03B2; = 0.01 are common standards. Based on Neyman&#x2013;Pearson theory, this measure is restricted to the discovery context; it qualifies the test-condition itself. Since we can gauge induction quality <italic>without</italic> actual data, this has nothing to do with a hypothesis test.</p></list-item>
<list-item><p><italic>Def. corroboration quality</italic>: A comparative measure of the inductive support that data lend to hypotheses, stated as the likelihood-ratio (aka &#x2018;Bayes-factor&#x2019;) of two point-hypotheses given data of sufficient induction quality. The support threshold is the ratio (1&#x2013;&#x03B2;-error)/&#x03B1;-error, and so depends on induction quality of data. For instance, setting &#x03B1; = &#x03B2; = 0.05 yields a threshold of 19 (or log 19 = 1.28), and &#x03B1; = &#x03B2; = 0.01 yields 99 (log 99 = 2.00), etc. Based on Wald&#x2019;s non-sequential testing theory the measure tests hypotheses against <italic>actual</italic> data in the justification context (<xref ref-type="bibr" rid="B1">Azzalini, 1996</xref>; <xref ref-type="bibr" rid="B56">Royall, 1997</xref>).</p></list-item></list>
<p>Of course, the final letter in NHST continues to abbreviate the term &#x2018;test.&#x2019; After all, NHST does test the <italic>probability</italic> of data given a hypothesis, <italic>P</italic>(D,H). But our definitions imply that data of low replication probability are insufficient to test a hypothesis in the sense of gauging its <italic>likelihood, L</italic>(H&#x007C;D). This is because a low-powered &#x201C;discovery&#x201D; of effect E in test-condition C (as indicated by a large &#x03B2;-error) entails the <italic>improbability</italic> of redetecting E in subsequent instances of C. So even in view of a confirmatory likelihood-ratio, data of insufficient induction quality may well-initially support a hypothesis. But similar support need not arise in new data of sufficient induction quality. So a given hypothesis may subsequently fail to be corroborated.</p>
</sec>
<sec><title>Upshot</title>
<p>On this background, the ongoing debates between statistical &#x201C;schools&#x201D; seem to be academic ones. After all, most extant estimation procedures for statistical significance operate squarely in the data space (<xref ref-type="bibr" rid="B56">Royall, 1997</xref>; <xref ref-type="bibr" rid="B24">Gelman, 2011</xref>; <xref ref-type="bibr" rid="B72">Wetzels et al., 2011</xref>). But if future theoretical developments must come to rely on coordinated activities that integrate the data with the hypotheses space, then the current crisis of empirical psychology would (at least partially) have arisen as a consequence of a methodologically <italic>unsound</italic> transition from discovery to justification. A sound version thereof, as we saw, leads from stable effects to trustworthy discoveries and on to acceptable forms of hypothesis corroboration.</p>
<p>As we also saw, integrating both spaces presupposes that we specify the expected empirical observation as a point-value. This states a theoretically sound minimum effect size (whether derived from a theory, or not); in uncertain cases it states a two-point interval placed around that value. By contrast, all alternative strategies simply let data have the &#x201C;last word&#x201D; on how we should construct a data-saving hypothesis. But this runs directly into the unmet challenge of validating induction.</p>
<p>Indeed, the risk of being &#x201C;perfectly wrong&#x201D; should be accepted even for precise theoretical assumptions. After all, being &#x201C;broadly right&#x201D; under merely vague assumptions is to accept virtually all non-random data-saving hypotheses. But that obviously fails to inform theoretical knowledge.</p>
</sec>
</sec>
<sec><title>Case Study: Psi-Research</title>
<sec><title>Overview</title>
<p>To clarify the relation between hypothesis corroboration and data replication, we exemplify our distinctions with a fairly controversial effect, treat its size, point to future research needs, and summarize the main insight.</p>
</sec>
<sec><title>Replicating Bem&#x2019;s Psi-hypothesis</title>
<p><xref ref-type="bibr" rid="B6">Bem&#x2019;s (2011)</xref> infamous results on precognition allegedly support the hypothesis that future expectations influence present behavior (aka &#x2018;psi-hypothesis&#x2019;). Seeking to replicate Bem&#x2019;s data, <xref ref-type="bibr" rid="B70">Wagenmakers et al. (2012)</xref> claimed to pursue a confirmatory research agenda. They could stop inquiry after 200 sessions with 100 subjects (see their Figure 2; ibid., p. 636). For by then their data had lent 6.2 times more support to the H<sub>0</sub> (read: <italic>no</italic> influence from future expectations) than to the H<sub>1</sub> (read: influence). This is considered <italic>substantial</italic> evidence for the H<sub>0</sub> (<xref ref-type="bibr" rid="B33">Jeffrey, 1961</xref>; <xref ref-type="bibr" rid="B69">Wagenmakers et al., 2011</xref>). Since Wagenmakers and colleagues base their inquiry on <xref ref-type="bibr" rid="B71">Wald&#x2019;s (1947)</xref> sequential analysis, however, we see reasons to treat their result with caution.</p>
<p>Rather than in order to evaluate theoretical hypotheses, <xref ref-type="bibr" rid="B71">Wald (1947)</xref> had developed sequential testing during WWII as a quality-control method in ammunition production. Its immediate purpose was to estimate how many shells in a lot deviate from a margin of error, M. While any deviation exceeding M provides a sufficient reason to discard the whole lot, measuring large deviations from M is of course less cumbersome than measuring small ones. So it saves effort to infer <italic>probable but unobserved</italic> small deviations from large observed deviations.</p>
<p>Wald&#x2019;s sequential testing strategy provides a rather brilliant solution to the classical problem of inducing properties of the whole from its parts. But it cannot generalize to <italic>additional</italic> lots. Nor was it intended to induce over abstract categories, but rather over material objects. Applied to the case of Wagenmakers and colleagues testing the specified hypotheses <italic>p</italic> = 0.50 vs. <italic>p</italic> = 0.531 (based on <xref ref-type="bibr" rid="B6">Bem</xref>&#x2019;s <xref ref-type="bibr" rid="B6">2011</xref> first experiment), this means that neither the &#x03B1;- nor the &#x03B2;-error were known. After all, the number of observations keeps varying with the observed result. Without at least stipulating both errors, however, we cannot quantify the replicability of their result. So we should rather not trust it.</p>
<p>To explain, our previous section had shown that hypothesis testing requires trustworthy data. To more fully appreciate that trustworthiness is largely owed to knowledge of errors (<xref ref-type="bibr" rid="B42">Mayo, 1996</xref>, <xref ref-type="bibr" rid="B43">2011</xref>), recall that induction quality and corroboration quality are related: if induction quality is unknown, then corroboration quality remains diffuse, and so can at best facilitate a vague form of justification. After all, even if a new and larger sample includes &#x201C;old&#x201D; data, a subsequent sample may nevertheless lead to a contrary decision as to whether a hypothesis is confirmed, or not.</p>
<p>As <xref ref-type="bibr" rid="B54">Rouder&#x2019;s (2014)</xref> discussion of the stopping rule shows (nicely), it is for this reason that Bayesians recommend that we <italic>keep</italic> adjusting our confidence level as data come in. Indeed, Bayesian inference puts the &#x201C;focus on the [<italic>current</italic>] degree of belief for considered models [here: H<sub>1</sub> and H<sub>0</sub>], which need not and should not be calibrated relative to some hypothetical truth&#x201D; (ibid., p. 308). Hence, the authors could reject Bem&#x2019;s hypothesis (H<sub>1</sub>), and instead accept the H<sub>0</sub> at a confidence level of (1&#x2013;&#x03B1;-error)/&#x03B2;-error = 1.64, where 1&#x2013;&#x03B1; = 0.95 and &#x03B2; = 0.58.</p>
<p>But the matter is more intricate yet. After all, corroboration quality would change if we altered the presumed distribution of <italic>possible</italic> data (aka &#x2018;the priors&#x2019;), for instance from a Cauchy- to a normal-distribution (<xref ref-type="bibr" rid="B7">Bem et al., 2011</xref>). Consequently, the <italic>same</italic> actual data would now rather support the alternative hypothesis. In general, which hypothesis it is that data confirm can be manipulated&#x2014;intentionally or not&#x2014;by suitably selecting the distribution of possible data. So &#x201C;that different priors result in different Bayes factors should [indeed] not come as a surprise&#x201D; (<xref ref-type="bibr" rid="B40">Ly et al., 2016</xref>, p. 12).</p>
<p>The selected type of prior distribution, however, is logically independent of the hypothesis we wish to test, ever entails weighing one hypothesis against another, and ultimately reflects a subjective decision. This provides reasons against giving Bayes-factors <italic>alone</italic> the final say in hypothesis corroboration. By contrast, to point-specify the hypothesis one does test eliminates this caveat by avoiding the priors, and no other method does. In the absence of a specified H<sub>1</sub>, then, whether a test-condition suffices for a clear justification of the H<sub>0</sub> depends on it featuring acceptably low &#x03B1;- and &#x03B2;-errors.</p>
<p>It follows that, had Wagenmakers and colleagues stopped their test after a mere 38 sessions, the same sequential testing-strategy should have led them to accept the H<sub>1</sub> on the basis of <italic>nearly substantial</italic> evidence (a likelihood-ratio of 3). This would have &#x201C;confirmed&#x201D; Bem&#x2019;s result by replicating his data. (See the curve in Figure 2 of <xref ref-type="bibr" rid="B70">Wagenmakers et al., 2012</xref>, p. 636; <xref ref-type="bibr" rid="B62">Simmons et al., 2011</xref> also illustrate this issue.) Given the small effect size <italic>g</italic> = 3.1% as a theoretical specification of prior results (50% against 53.1%, according to <xref ref-type="bibr" rid="B6">Bem, 2011</xref>, p. 409, experiment 1), one should therefore construct a sufficiently strong test-condition to obtain data of sufficient induction quality.<sup><xref ref-type="fn" rid="fn06">6</xref></sup></p>
<p>Neyman&#x2013;Pearson theory defines the necessary sample size (or number of observations, subjects, sessions, etc.) to firmly decide between two hypotheses with a difference of <italic>g</italic> = 3.1%. Since this is a test <italic>against</italic> the H<sub>1</sub>, we should treat both errors equally, for instance by setting &#x03B1; = &#x03B2; = 0.05. It follows that, for the H<sub>0</sub> to be accepted, it must be 19 times (0.95/0.05) more probable than the H<sub>1</sub>. The necessary sample size (for a proportion difference measured against a theoretical constant of 0.50) then comes to <italic>n</italic> = 2829 (see <xref ref-type="bibr" rid="B12">Cohen, 1977</xref>, p. 169). Comparing this tall figure to the <italic>n</italic> = 200 that <xref ref-type="bibr" rid="B70">Wagenmakers et al. (2012)</xref> report should make clear why one cannot trust their result.</p>
<p>A well-suited approximation of the necessary sample size, n, given specified errors and a postulated effect size of mean differences, d, is (2(<italic>z</italic><sub>(1-&#x03B1;)</sub>+ <italic>z</italic><sub>(1-&#x03B2;)</sub>)<sup>2</sup>)/d<sup>2</sup>= n, where <italic>z</italic><sub>(1-&#x03B1;)</sub> and <italic>z</italic><sub>(1-&#x03B2;)</sub> increase provided &#x03B1; and &#x03B2; decrease, with <italic>z</italic> taking values greater than 1, and <italic>d</italic> mostly remaining below 1. Given acceptable errors, a very small d (or g) thus generates a large n. So to achieve reasonable certainty under specified errors, we may incur an extremely large number of data points.</p>
<p>The main reason for the large n is the small difference between the two rivaling hypotheses and our rigor in controlling both errors. Therefore, it does not suffice to publish the testing-strategy before and a stopping-rule after data inspection (<xref ref-type="bibr" rid="B48">Nosek and Bar-Anan, 2012</xref>; <xref ref-type="bibr" rid="B49">Nosek et al., 2012</xref>, <xref ref-type="bibr" rid="B47">2015</xref>). Similarly, though sequential testing is a less problematic way of inflating the &#x03B1;-error than &#x201C;double-dipping&#x201D; (<xref ref-type="bibr" rid="B36">Kriegeskorte et al., 2009</xref>), it necessarily inflates the &#x03B2;-error, given that we hold effects constant. So it increases the chance of not detecting a true difference, making it improbable to replicate data in independent studies.</p>
</sec>
<sec><title>Gauging the Psi-effect</title>
<p>In the case of replicating the psi-hypothesis, the &#x03B2;-error was at least &#x03B2; = 0.58, given <italic>n</italic> = 200, &#x03B1; = 0.05 (one-sided), and <italic>g</italic> = 0.05. Though this is a slightly higher value than was in fact observed, it still qualifies as a small effect (see <xref ref-type="bibr" rid="B12">Cohen, 1977</xref>, p. 155). In fact, G<sup>&#x2217;</sup>Power software (<xref ref-type="bibr" rid="B18">Faul et al., 2007</xref>) estimates an even larger error of &#x03B2; = 0.64. At any rate, the large &#x03B2;-error renders Wagenmakers and colleagues&#x2019; test-condition unacceptable.</p>
<p>Based on extant psi-studies, Bem had formulated a H<sub>1</sub> of <italic>d</italic> = 0.25. (Bem prefers <italic>t</italic>-tests and more &#x201C;classical&#x201D; effect size measures over a binomial test; though results are stated in percentages, the output of both kinds of tests is equivalent.) Having specified d, we can thus calculate the error probabilities: for &#x03B1; = 0.05 (one-sided), <italic>d</italic> = 0.25 and <italic>n</italic> = 100, we find &#x03B2;-error = 0.18. Now setting &#x03B1; = &#x03B2; = 0.05, the necessary sample is <italic>n</italic> = 175, smaller than immediately above because d now is comparably large. (<italic>n</italic> = 175 is about half the sample the above formula approximates, since a difference between a constant and an empirical mean is evaluated.)<sup><xref ref-type="fn" rid="fn07">7</xref></sup> The Bayes-factor thus registers some 22 times in favor of a H<sub>1</sub> postulating <italic>d</italic> = 0.25. On a Bayesian view, this is a clear corroboration of the H<sub>1</sub> over the H<sub>0</sub>.</p>
<p>Though the critical value (1&#x2013;&#x03B2;-error)/&#x03B1;-error = 0.80/0.05 = 16 has now been surpassed, upon inspecting induction quality it transpires that the result is insufficiently stable. To explain, some of Bem&#x2019;s trials sought to induce arousal by displaying erotic and non-erotic pictures in random order, measured the degree of arousal before displaying a picture-type, and interpreted heightened arousal <italic>before</italic> showing an erotic picture as evidence of precognition. (This suffices to interpret the set-up as including a control group of sorts.) The observed effect was <italic>d</italic> = 0.19 with <italic>n</italic> = 100 in both samples (erotic vs. non-erotic). Comparing this with the hypotheses <italic>d</italic> = 0.00 and <italic>d</italic> = 0.25, however, a Bayes-factor of 2.18 is now too low. So Bem&#x2019;s first experiment indeed &#x201C;discovers&#x201D; a deviation from random (given &#x03B1; = 0.05, one-sided, and 1&#x2013;&#x03B2; = 0.82). But the effect isn&#x2019;t stable (i.e., its reproducibility is insufficiently probable), particularly given that both hypotheses had been specified by recourse to meta-analytical results.</p>
<p>This goes to show that a Bayes-factor may well be extremely large although the effect is not trustworthy. Indeed, if we interpret data from the display of non-erotic pictures as a control group&#x2014;as we should, because data from the display of erotic pictures did not significantly deviate from random&#x2014;then a likelihood-ratio of 2.18 (i.e., a logarithm of 0.34) is hardly any evidence for a psi-effect, even if we ignore its low replication probability. Rather, a firm decision under sufficient induction quality requires another sample of <italic>n</italic> = 75 (see below).</p>
</sec>
<sec><title>Meta-analyses of Additional Replication Attempts</title>
<p><xref ref-type="bibr" rid="B6">Bem&#x2019;s (2011)</xref> thought-provoking research has meanwhile initiated something <italic>like</italic> a research program. But we cannot elucidate its contradictory results by relying on <italic>either</italic> frequentist or Bayesian approaches to statistical inference. It should therefore be of interest to clarify the psi-debate by integrating both approaches.</p>
<p><xref ref-type="bibr" rid="B23">Galak et al. (2012)</xref> and <xref ref-type="bibr" rid="B5">Bem et al. (2016)</xref> have conducted two independent meta-analyses of psi-studies. Their <italic>combination</italic> in fact yields the necessary sample size. The first study concludes negatively:</p>
<list list-type="simple" prefix-word="simple">
<list-item><p>&#x201C;Across seven experiments (<italic>N</italic> = 3,298), we replicate the procedure of experiments 8 and 9 from <xref ref-type="bibr" rid="B6">Bem (2011)</xref>, which had originally demonstrated retroactive facilitation of recall. <italic>We failed to replicate that finding</italic>. We further conduct a meta-analysis of all replication attempts of these experiments and find that the average effect size (<italic>d</italic> = 0.04) is not different from 0&#x201D; (<xref ref-type="bibr" rid="B23">Galak et al., 2012</xref>, p. 933; <italic>italics added</italic>).</p></list-item></list>
<p>But the second meta-analysis arrives at a quite different result:</p>
<list list-type="simple" prefix-word="simple">
<list-item><p>&#x201C;We here report a meta-analysis of 90 experiments from 33 laboratories in 14 countries which yielded an overall effect greater than six sigma, <italic>z</italic> = 6.40, <italic>p</italic> = 1.2 &#x00D7; 10<sup>-10</sup> with an effect size (Hedges&#x2019; g) of 0.09. A Bayesian analysis yielded a Bayes Factor [BF] of 5.1 &#x00D7; 10<sup>9</sup>, greatly exceeding the criterion value of 100 for &#x2018;decisive evidence&#x2019; in support of the experimental hypothesis. When Bem&#x2019;s experiments are excluded the combined effect size for replications by independent investigators is 0.06, <italic>z</italic> = 4.16, <italic>p</italic> = 1.1 &#x00D7; 10<sup>-5</sup>, and the BF value is 3.853, again exceeding the criterion of &#x2018;decisive evidence.&#x2019; [&#x2026;] <italic>P</italic>-curve analysis, a recently introduced statistical technique, estimates the true effect size of the experiments to be 0.20 for the complete database and 0.24 for the independent replications&#x201D; (<xref ref-type="bibr" rid="B5">Bem et al., 2016</xref>, p. 1).</p></list-item></list>
<p>To prepare for a critical discussion, consider that both meta-analyses sought to discover a non-random effect, but neither <italic>tested</italic> the psi-hypothesis in the sense of gauging <italic>L</italic>(H&#x007C;D); effect sizes are heterogeneous, suggesting that uncontrolled influences are at play; Bem&#x2019;s own studies report larger effects than their independent replications, suggesting a self-fulfilling prophecy; Bayes-<italic>t</italic>-tests, as we saw, depend on the prior distribution and different priors can lead to contradictory results; most studies included in these meta-analyses are individually underpowered.</p>
<p>Of course, to simply aggregate various underpowered studies will not yield a trustworthy inductive basis. After all, almost all mean differences become statistically significant if we arbitrarily divide a sufficiently large sample (<italic>n</italic> &#x2265; 60.000) into two subgroups (<xref ref-type="bibr" rid="B2">Bakan, 1966</xref>). So provided that only the H<sub>0</sub> is specified, we can almost always obtain a non-random result by increasing the sample. [This contrasts with the methodology of physics (<xref ref-type="bibr" rid="B44">Meehl, 1967</xref>), for instance, where a theoretical parameter is fixed and increasing the sample size eventually <italic>disproves</italic> a theory.]</p>
<p>In view of the more than 90 psi-studies that both meta-analyses reviewed, researchers did particularly consider the point-hypotheses <italic>d</italic> = 0.00 (random, H<sub>0</sub>) and <italic>d</italic> = 0.25 (specified, H<sub>1</sub>). (Bem assumed the latter <italic>d</italic>-value to plan studies with power = 0.80, after &#x201C;predicting&#x201D; <italic>d</italic> = 0.24 by analyzing independent replications; see <xref ref-type="bibr" rid="B5">Bem et al., 2016</xref>.) Since this &#x201C;research program&#x201D; requires amendment before it can successfully address the challenges the replicability crisis has made apparent, we now exemplify the inference strategy a <italic>genuine</italic> psi-research program would pursue.</p>
<p>Among the 90 psi-studies, we consider most trustworthy those that arose independently of Bem&#x2019;s research group, that are classified as exact replications, and that are peer-reviewed (These admittedly rigorous criteria leave but nine of the studies listed in Table A1 in <xref ref-type="bibr" rid="B5">Bem et al., 2016</xref>, p. 7, Dataset S1).</p>
<p>Our first critical question must pertain to induction quality: given &#x03B1; = &#x03B2; = 0.05 and <italic>d</italic> = 0.25 (see above), the necessary sample is <italic>n</italic> = 175. The total sample from the nine studies comes to <italic>n</italic> = 520. Since this far exceeds our requirements, we can in fact assess corroboration quality more severely. Given &#x03B1; = &#x03B2; and <italic>n</italic> = 520, the critical value to corroborate the H<sub>1</sub> now is (1&#x2013;&#x03B2;-error)/&#x03B1;-error = 0.998/0.002 = 499 (or log 2.70), and correspondingly for the H<sub>0</sub>, i.e., (1 - &#x03B1;-error)/&#x03B2;-error. Adding the log-likelihood-ratios for the H<sub>0</sub> and subtracting those for the H<sub>1</sub> then yields 5.13. This value is much higher than 2.70, making it 135,000 times more probable that data inductively support the H<sub>0</sub>, rather than a psi-effect of size <italic>d</italic> = 0.25.</p>
<p>When we gauge the average effect size of our nine studies (weighed by the sample size), moreover, maximum likelihood estimation yields <italic>d</italic> = 0.07 as a psi-effect that is 3.64 times more probable than a random effect. Rather than a final verdict on the true hypothesis, of course, this provides but a relative corroboration of one hypothesis against another. A maximum likelihood estimate that registers so close to the hypothetical parameter, however, fails to provide <italic>any</italic> hint for further research.</p>
<p>The H<sub>0</sub> is now much better corroborated than the psi-hypothesis. Of course, this does not falsify a psi-hypothesis postulating a <italic>yet smaller</italic> effect. It is nevertheless reasonable to reject the psi hypothesis. After all, to theoretically explain an effect only becomes more difficult the smaller it is. But again, one cannot be certain.<sup><xref ref-type="fn" rid="fn08">8</xref></sup></p>
</sec>
<sec><title>Summary</title>
<p>Our discussion shows that a <italic>firm and transparent</italic> decision between two specified hypotheses requires knowing the effect size and selecting tolerable &#x03B1;- and &#x03B2;-errors. We also saw how this determines the necessary sample size. Specifically, the smaller an effect is the larger is the necessary sample. Theories that predict small effects should therefore be confronted with much larger samples than is typical. Conversely, the small effects and samples that top-tier journals often report let few published effects count as probably replicable. Indeed, the average test-power reported in <xref ref-type="bibr" rid="B4">Bakker et al. (2012)</xref> &#x201C;predicts&#x201D; the lower bound (36%) of the average replication rate reported by <xref ref-type="bibr" rid="B50">Open Science Collaboration (2015)</xref>.</p>
<p>Since many studies avoid specifying the alternative hypothesis, researchers may at best <italic>hope</italic> to find a significant effect among their results. This partially explains why they sometimes &#x201C;torture&#x201D; data until significance is achieved. As we saw, such results frequently surface as allegedly important effects, that is, as <italic>genuine</italic> discoveries (see <xref ref-type="bibr" rid="B17">Fanelli and Gl&#x00E4;nzel, 2013</xref>; <xref ref-type="bibr" rid="B82">Witte and Strohmeier, 2013</xref>). Predictably, top-tier journals regularly publish studies that fail to report probably replicable discoveries, namely when their test-power is too low (&#x003C;0.80 or &#x003C;0.95) to safely reject the H<sub>0</sub> (<xref ref-type="bibr" rid="B4">Bakker et al., 2012</xref>; <xref ref-type="bibr" rid="B21">Francis, 2012</xref>; <xref ref-type="bibr" rid="B50">Open Science Collaboration, 2015</xref>; <xref ref-type="bibr" rid="B16">Etz and Vandekerckhove, 2016</xref>).</p>
<p>With these insufficiently designed studies as evidence for a goal-conflict between psychologists and their field (<xref ref-type="bibr" rid="B78">Witte, 2005</xref>; <xref ref-type="bibr" rid="B15">Eriksson and Simpson, 2013</xref>), we go on to show how research programs improve the <italic>status quo</italic>.</p>
</sec>
</sec>
<sec><title>Research Programs</title>
<sec><title>Four Developmental Steps</title>
<p>This section outlines the four steps a progressive research program takes in order to improve empirical knowledge (<xref ref-type="bibr" rid="B27">Hacking, 1978</xref>; <xref ref-type="bibr" rid="B38">Lakatos, 1978</xref>; <xref ref-type="bibr" rid="B39">Larvor, 1998</xref>; <xref ref-type="bibr" rid="B46">Motterlini, 2002</xref>). Adopting such programs is a natural consequence of recognizing the precisification of hypotheses as a necessary condition to obtain trustworthy results, and of accepting that prior (fallible) knowledge informs future research.</p>
</sec>
<sec><title>Step One: Ideas without Controlled Observation</title>
<p>The first step involves an idea or intuition, perhaps acquired in what C.S. Peirce called <italic>retroduction</italic>. Interesting in itself, <italic>having</italic> that intuition is less relevant for developing a scientific construct. Though an intuition is by definition not based on conscious observation, to be further explored it must nonetheless sufficiently impress us. For instance, we might seek hints in subjective experience, theoretical observation, or collegial discussion. Once we are convinced that the idea <italic>is</italic> relevant, we can engage with it systematically, perhaps moving from material at hand to thought-experiments, computer simulations, or an &#x201C;idea-paper&#x201D; (<italic>sans</italic> significance tests, etc.). Importantly, such activities are possible without collecting data.</p>
</sec>
<sec><title>Step Two: Devising an Empirical Set-up</title>
<p>The second step establishes our idea with a method, so that it may potentially count as a <italic>genuine</italic> discovery. This requires controlling an empirical set-up in order to &#x201C;observe&#x201D; the phenomenon. But we saw that observations may mislead because of random effects and sampling- or measurement-error. So for a discovery to be established, the phenomenon must significantly deviate from random. After all, random variation lacks <italic>specific</italic> semantic content and so does not explain anything but itself.</p>
<p>We also saw that proofs under probabilistic variation are based on inferential statistics that consider the observation&#x2019;s <italic>p</italic>-value given the H<sub>0</sub>. The classical claim ascribed to Fisher is that a small <italic>p-</italic>value reflects a rare enough event in a random model. (At this step, we cannot properly call the search for small <italic>p</italic>-values &#x2018;p-hacking,&#x2019; an often observed strategy after having obtained data.) As a consequence, a &#x201C;large p&#x201D; (<italic>p</italic> > 0.05) signals our failure to measure a significant deviation from random. This <italic>prima facie</italic> suggests complicity with a random model. But a large p is expected, of course, if the set-up is insufficiently sensitive to detect an effect that nevertheless <italic>is</italic> present. So it would be na&#x00EF;ve to treat the absence of a statistically significant deviation from random as conclusive evidence for the effect&#x2019;s absence.</p>
<p>This sensitivity caveat calls on us to adjust our logic of decision-making. After all, statistically <italic>insignificant</italic> deviations from random may still signal epistemic value. Historically, that insight led Neyman-Pearson-test-theory to provide us also with an estimate of the &#x03B2;-error. This should be sufficiently small for the chance of not detecting a true effect to be low. In other words, the set-up&#x2019;s sensitivity should be sufficiently high.</p>
<p>A set-up will generally be <italic>optimal</italic> if evidence features the smallest possible <italic>p</italic>-value given the random model, on one hand, and the largest probability of registering each true deviation from random using a minimum number of observations, on the other. Moreover, an optimal test should be unbiased, so that we can almost certainly detect a true effect as we increase the number of observations. In short, we should select the <italic>most powerful</italic> test-condition. Test-power considerations thus inform how we gauge the relative inductive support that data provide for a content-free H<sub>0</sub>, compared to the support that the same data provided for a (one- or two-sided) H<sub>1</sub> of substantial content.</p>
<p>At this point, however, rather than manage a point-hypothesis, we still deal with a vague H<sub>1</sub>. Moreover, we know induction quality of data but <italic>partially</italic> as long as we merely hold the &#x03B1;-error constant, but not the &#x03B2;-error. As a consequence, data may now <italic>seem</italic> to corroborate or falsify a hypothesis, but they may not be replicable. Lastly, corroboration and falsification equally depend on the distribution of possible data. So corroboration quality is at best <italic>diffuse</italic>.</p>
<p>Even if data should now &#x201C;eliminate&#x201D; the H<sub>0</sub>, as it were, a genuine <italic>bundle</italic> of substantial alternative H<sub>1</sub>-hypotheses remains to be eliminated. Since each such H<sub>1</sub>-hypothesis postulates a distinct effect size, the whole bundle provides mutually exclusive explanatory candidates. Each such H<sub>1</sub> therefore associates to a distinct likelihood-ratio as its corroboration measure. So the unsolved problem of validating induction merely let data <italic>hint</italic> at the correct theoretical parameter, but data alone cannot determine it. Instead, it is <italic>our</italic> having further specified this parameter that eventually lets data decide firmly between any two such point-hypotheses.</p>
</sec>
<sec><title>Step Three: Replication and Meta-analysis</title>
<p>As we achieve replication-success several times, we can give a better size-estimate of a significant effect. But this estimate will be unbiased only if we also base it on (often unpublished) statistically non-significant results (<xref ref-type="bibr" rid="B65">Sterling, 1959</xref>; <xref ref-type="bibr" rid="B57">Scargle, 2000</xref>; <xref ref-type="bibr" rid="B60">Schonemann and Scargle, 2008</xref>; <xref ref-type="bibr" rid="B19">Ferguson and Heene, 2012</xref>). Moreover, an effect size that has remained heterogeneous across several studies should eventually be differentiated from its test-condition. In fact, this is the <italic>genuine</italic> purpose of a meta-analysis. But extant analyses are heavily biased toward published results. Despite various bias detection-tools, such analyses therefore often present a skewed picture of all available data, and hence exaggerate the true effect size (<xref ref-type="bibr" rid="B21">Francis, 2012</xref>).</p>
<p>When we reproduce an effect, we should therefore correct a plain induction over prior findings by more theoretical strategies (see, e.g., <xref ref-type="bibr" rid="B75">Witte, 1994</xref>, <xref ref-type="bibr" rid="B77">1996b</xref>, <xref ref-type="bibr" rid="B78">2005</xref>). This includes data-inspection and -reconstruction by means of a theory with a mathematical core that predicts quantitative results while retaining an adequate connection to data. Data reconstruction thus becomes a stepping stone to formulating point-hypotheses.</p>
<p>The forgoing exhausts the relevant discovery context activities. Subsequent work should be guided by specifying effect sizes as point-hypotheses, and by testing them against <italic>new</italic> data that arise as retrodictions or predictions.</p>
</sec>
<sec><title>Step Four: Precisification of Effect Sizes and Theoretical Construction</title>
<p>As we saw, we can quantitatively assess the quality of an empirical set-up only <italic>after</italic> a point-hypothesis is available. Subsequently, a set-up&#x2019;s induction quality can serve as a criterion to probabilistically corroborate, or falsify, a theoretical construct. So the fourth step directly concerns formulating precise point-hypotheses.</p>
<p>A clear indicator that we have reached the justification context is to inspect likelihoods of hypotheses given data, <italic>L</italic>(H&#x007C;D), rather than probabilities of data given hypotheses, <italic>P</italic>(D,H). Whether precision-gains then arise from a parameter-estimation or by combining significant results (<xref ref-type="bibr" rid="B83">Witte and Zenker, 2016a</xref>,<xref ref-type="bibr" rid="B84">b</xref>, <xref ref-type="bibr" rid="B85">2017</xref>), here we either induce over empirical results, or we perform a quantitative reanalysis. This marks the onset of an explanatory construction (<xref ref-type="bibr" rid="B77">Witte, 1996b</xref>; <xref ref-type="bibr" rid="B80">Witte and Heitkamp, 2006</xref>). As a rule, if we have corroborated an effect size, we should next provide a theoretical (semantic) explanation.</p>
<p>Generating such explanations takes time, of course, and often incurs unforeseen problems. Indeed, it need not succeed. With <xref ref-type="bibr" rid="B38">Lakatos (1978)</xref>, we consider a research program <italic>progressive</italic> as long as a theoretical construction or its core-preserving modification generate predictions that are at least <italic>partially</italic> corroborated by new data of sufficient induction quality. Any such construction may therefore lead to further discoveries, e.g., in the form of observations that deviate from random. So it remains overly simple to treat theory-development as a linear process (see <bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>). But that scientific progress grounds in fallible knowledge of an original phenomenon deserves acceptance.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Only stable data (&#x201C;phenomena&#x201D;) can be meaningfully subsumed under empirical hypotheses (bottom arrow); such hypotheses can be systematized into theories (left arrow) that retrodict old and predict new data; this may lead to constructing empirical set-ups (top arrow) for new phenomena (right arrow), which potentially improve extant theories.</p></caption>
<graphic xlink:href="fpsyg-08-01847-g001.tif"/>
</fig>
<p>A basic principle is to continue the precisification of hypotheses, since this improves both induction quality (from <italic>unknown</italic> to <italic>known to be probably reproducible</italic>) as well as corroboration quality (from <italic>diffuse</italic> to <italic>precise</italic>). A progressive research program thus shifts the focus, away from vague justification against the H<sub>0</sub>, toward precise justification of a point-specified H<sub>1</sub>. That said, HARKing and double dipping (<xref ref-type="bibr" rid="B34">Kerr, 1998</xref>; <xref ref-type="bibr" rid="B36">Kriegeskorte et al., 2009</xref>; <xref ref-type="bibr" rid="B62">Simmons et al., 2011</xref>) remain non-cogent justifications because each so generated result must lead to a confirmation.</p>
</sec>
</sec>
<sec><title>Move Over, Please!</title>
<p>Statistical inference is a tool to corroborate theoretical assumptions rather than a machinery to generate theoretically grounded empirical knowledge&#x2014;which an empirical science must justify with trustworthy data. Therefore, a loose discovery under vague hypotheses (based on sample-estimations) can never lead to a firmly corroborated theoretical assumption. In fact, to even only disconfirm a precise assumption is already far more informative than to never achieve adequate precision.</p>
<p>Considering the scarcity of similarly well-hardened knowledge in current psychological research, the field&#x2019;s progress demands justification context-activities (<xref ref-type="bibr" rid="B31">Ioannidis, 2012</xref>). Indeed, as long as the field as a whole remains in the discovery context, it is doubtful that lenient reviewers, full data disclosure, or removing the publication-bottleneck, etc. even address the true challenges. Instead, we would presumably see meta-analyses lead to meta-meta-analyses, which however cannot serve a progressive research program (<xref ref-type="bibr" rid="B59">Schmidt et al., 2009</xref>; <xref ref-type="bibr" rid="B9">Cafri et al., 2010</xref>; <xref ref-type="bibr" rid="B64">Stegenga, 2011</xref>; <xref ref-type="bibr" rid="B10">Chan and Arvey, 2012</xref>; <xref ref-type="bibr" rid="B19">Ferguson and Heene, 2012</xref>; <xref ref-type="bibr" rid="B45">Mitchell, 2012</xref>). Rather, if a meta-analysis discovers an effect, one should seek to confirm a precisified version thereof by successfully predicting it in <italic>new</italic> data.<sup><xref ref-type="fn" rid="fn09">9</xref></sup></p>
<p>The likelihood ratio, as we saw, is the measure of corroboration quality. For a tolerable error-range, reasonable certainty that a H<sub>1</sub> is justified thus requires that its likelihood-ratio exceeds a conventional threshold. We can compare the error terms from Neyman-Pearson-theory to <xref ref-type="bibr" rid="B33">Jeffrey&#x2019;s (1961)</xref> qualitative classification: &#x03B1; = &#x03B2; = 0.05 translates into a likelihood-ratio of 19 (&#x201C;strong evidence&#x201D;) and &#x03B1; = &#x03B2; = 0.01 into a ratio of 99 (&#x201C;very strong evidence,&#x201D; &#x201C;nearly extreme evidence&#x201D;). Further, provided that &#x03B1; = &#x03B2; &#x2264; 0.01, the degree of corroboration for the supported hypothesis will in the long run approach the test-power value (<xref ref-type="bibr" rid="B71">Wald, 1947</xref>).</p>
<p>As a summary, we now list six increasingly important results of a research program and their measures. The last result corroborates a point-hypothesis.</p>
<list list-type="simple" prefix-word="simple">
<list-item><label>(1)</label><p><italic>preliminary discovery</italic>: &#x03B1;-error or merely a <italic>p</italic>-value (to establish that data are non-random)</p></list-item>
<list-item><label>(2)</label><p><italic>substantial discovery</italic>: &#x03B1; and 1&#x2013;&#x03B2;-error (to gauge replicability based on a specific effect size and a particular sample size)</p></list-item>
<list-item><label>(3)</label><p><italic>preliminary falsification</italic> of the H<sub>0</sub>: <italic>L</italic>(d<sub>H1</sub> > 0)/<italic>L</italic>(d<sub>H0</sub> = 0) (to establish a H<sub>1</sub> that deviates in one direction from the H<sub>0</sub> as more likely than the random parameter <italic>d</italic> = 0)</p></list-item>
<list-item><label>(4)</label><p><italic>substantial falsification</italic> of the H<sub>0</sub>: <italic>L</italic>(d<sub>H1</sub> > &#x0394;)/<italic>L</italic>(d<sub>H0</sub> = 0), where &#x0394; is the theoretical minimum effect size value (to establish a H<sub>1</sub> as non-random and also exceeding &#x0394;)</p></list-item>
<list-item><label>(5)</label><p><italic>preliminary verification</italic> of the H<sub>1</sub>: <italic>L</italic>(d<sub>H1</sub> = &#x0394;)/<italic>L</italic>(d<sub>H0</sub> = 0) (to corroborate the theoretical parameter &#x0394; against the random parameter)</p></list-item>
<list-item><label>(6)</label><p><italic>substantial verification</italic> of the H<sub>1</sub>: <italic>L</italic>(d<sub>emp, H1</sub>)/L(d<sub>H2</sub> = &#x0394;) &#x003C; 4, given approximately normally distributed data, where d<sub>emp</sub> is the empirical effect size<sup><xref ref-type="fn" rid="fn010">10</xref></sup> (to indirectly corroborate &#x0394; against the maximum-likelihood estimate of empirical data, d<sub>emp</sub>).</p></list-item></list>
<p>The justification context starts with line (3). So the unsophisticated application of NHST sees large parts of empirical psychological research &#x201C;stuck&#x201D; with making preliminary or substantial discoveries. Notice, too, that the laudable proposal by <xref ref-type="bibr" rid="B8">Benjamin et al. (2017)</xref> to drastically lower the &#x03B1;-error merely addresses what shall count as a preliminary discovery. But it leaves unaddressed how we establish a substantial discovery as well as the justification context as a whole.</p>
</sec>
<sec><title>Precise Theoretical Constructs?</title>
<p>Allow us to briefly speculate why the goal conflict between generating statistically significant results and generating trustworthy fallible knowledge has played out in favor of the former goal. We saw that trustworthy data is primarily relevant toward developing and testing more refined theoretical constructs. But theory-construction is rarely taught at universities (<xref ref-type="bibr" rid="B25">Gigerenzer, 2010</xref>). So the replicability crisis also showcases our inability to in fact erect the constructs that statistically significant effects should lead to (<xref ref-type="bibr" rid="B14">Ellemers, 2013</xref>; <xref ref-type="bibr" rid="B35">Klein, 2014</xref>).</p>
<p>To give but two examples, the 51 theories collected in a recent social psychology handbook (<xref ref-type="bibr" rid="B67">van Lange et al., 2012</xref>) cannot achieve equally precise predictions as those that a random model offers. But not only should we base a fair decision between two hypotheses on data that are known to be probably replicable; we should also require <italic>parity of precision</italic> (<xref ref-type="bibr" rid="B81">Witte and Kaufman, 1997</xref>). By contrast, a two volume edition on small group behavior (<xref ref-type="bibr" rid="B79">Witte and Davis, 1996</xref>) contains eleven theories that are sufficiently elaborated to predict precise effects, while another ten theories make vague predictions. So precise constructions <italic>are</italic> possible, yet they are rare. Therefore, calling the replicability crisis &#x201C;home-made&#x201D; does not directly implicate career aspiration or resource shortage. Rather, ignorance of how one constructs a theory seems to be an important intermediate factor.</p>
<p>The crisis will probably seem insurmountable to those who disbelieve that a developing research program is even possible. Alas, might one not first ask to show us a <italic>failed</italic> research program? Indeed, would not a research program culture need to have been established, and to have broadly failed, too? If so, then an urgent challenge is to coordinate our research away from the individualistically organized but statistically underpowered short-term efforts that have produced the crises, toward jointly managed and well-powered long-term research programs.</p>
<p>Of course, it takes more than writing papers. Indeed, funding-, career- and incentive-structures may also need to change, including how empirical psychologists understand their field. It may even require that larger groups break with what most do quasi-habitually. This, perhaps, would establish what our paper only described.</p>
</sec>
<sec><title>Conclusion</title>
<p>The replicability crisis in psychology is in large part a consequence of applying an unsophisticated version of NHST. This praxis merely generates theoretically disconnected &#x201C;one-off&#x201D; discoveries&#x2014;that is, effects which deviate statistically significantly from a random model. Parallel to it runs the interpretative or rhetorical praxis of publishing such effects as scientifically important results, rather than as the parameter estimations they are. The former praxis fails to maximize the utility of a sophisticated version of NHST, which nevertheless remains a useful and elegant approach to gauging <italic>P</italic>(D,H). The latter praxis regularly over-reports <italic>P</italic>(D,H) as <italic>L</italic>(H&#x007C;D), which even the most sophisticated application of NHST, however, cannot warrant. Both praxes may plausibly have arisen from not fully understanding the limits of NHST.</p>
<p>The field has thus &#x201C;discovered&#x201D; many small effects. But whenever empirical studies remain underpowered, because they comprise too few data-points, research efforts remain in the discovery context. Here, vague effect sizes and vague alternative hypotheses are normal. But such results fail to convince as trustworthy effects that deserve theoretical explanation. By contrast, a clear indicator that research has shifted to the justification context is a likelihood-ratio-based decision regarding a precisified effect size that is expectable in <italic>new</italic> data of known induction quality (entailing a Bayes-factor for fixed hypotheses).</p>
<p>This requires a diachronic notion of the research process&#x2014;a developing <italic>research program</italic>&#x2014;that adapts statistical inference methods to prior knowledge. Schematically (<bold>Figure <xref ref-type="fig" rid="F2">2</xref></bold>), we start with <italic>p</italic>-values (Fisher), move on to an optimal test against a random-model (Neyman&#x2013;Pearson with &#x03B1;- and 1-&#x03B2;-error), accompanied by parameter estimation via meta-analysis, to achieve&#x2014;entering the justification context&#x2014;a corroboration of a theoretically specified effect size based on stable data of known induction quality (under tolerable errors), against a random model or against another point-specified hypothesis.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Salient results of a research program.</p></caption>
<graphic xlink:href="fpsyg-08-01847-g002.tif"/>
</fig>
<p>In particular, we should deduce from a theory a specified effect size that goes beyond the simple assumption of a minimal effect. A quantitative specification without theoretical explanation may well be a first step (assuming known induction quality and precise corroboration quality). But a successful research program must derive a prediction from a theoretical model <italic>and</italic> provide an explanation of the effect&#x2019;s magnitude. For instance, psi-research would only gain from such an explanation.</p>
<p>Since many theories only offer vague predictions, moreover, the lack of confidence among psychologists might at least partially result from failed theory-development. But the replicability crisis itself is narrowly owed to underpowered studies. Making headway takes researchers who join forces and resources, who coordinate research efforts under a long-term perspective, and who adapt statistical inference methods to prior knowledge. To this end, we have also seen a strategy for combining data from various studies that avoids the pitfalls of extant meta-analyses. It is in the <italic>integrated</italic> long run, then, that empirical psychology may improve.</p>
</sec>
<sec><title>Author Contributions</title>
<p>Both authors have jointly drafted this manuscript; the conceptual part of this work originates with EW; FZ supplied additional explanatory material and edited the manuscript.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>For valuable comments that served to improve earlier versions of this manuscript, we thank Peter Killeen and Moritz Heene as well as EB and PB for their reviews. Thanks also to HF for overseeing the peer review process. All values were calculated using G<sup>&#x2217;</sup>Power, V3.1.9.2 (<xref ref-type="bibr" rid="B18">Faul et al., 2007</xref>). FZ acknowledges funding from the Ragnar S&#x00F6;derberg Foundation, an &#x201C;Understanding China&#x201D;-Fellowship from the Confucius Institute (HANBAN), as well as funding through the European Union&#x2019;s FP 7 framework program (No. 1225/02/03) and the Volkswagen Foundation (No. 90/531). Finally, thanks to Frontiers for a partial fee waiver, and to Lund University&#x2019;s Library and the Department of Philosophy for providing open access funding.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Azzalini</surname> <given-names>A.</given-names></name></person-group> (<year>1996</year>). <source><italic>Statistical Inference. Based on the Likelihood.</italic></source> <publisher-loc>London</publisher-loc>: <publisher-name>Chapman &#x0026; Hall</publisher-name>.</citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bakan</surname> <given-names>D.</given-names></name></person-group> (<year>1966</year>). <article-title>The test of significance in psychological research.</article-title> <source><italic>Psychol. Bull.</italic></source> <volume>66</volume> <fpage>423</fpage>&#x2013;<lpage>437</lpage>. <pub-id pub-id-type="doi">10.1037/h0020412</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baker</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>First results from psychology&#x2019;s largest reproducibility test.</article-title> <source><italic>Nat. News.</italic></source> <pub-id pub-id-type="doi">10.1038/nature.2015.17433</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bakker</surname> <given-names>M.</given-names></name> <name><surname>van Dijk</surname> <given-names>A.</given-names></name> <name><surname>Wicherts</surname> <given-names>J. M.</given-names></name></person-group> (<year>2012</year>). <article-title>The rules of the game called psychological science.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>543</fpage>&#x2013;<lpage>554</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612459060</pub-id> <pub-id pub-id-type="pmid">26168111</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bem</surname> <given-names>D.</given-names></name> <name><surname>Tressoldi</surname> <given-names>P.</given-names></name> <name><surname>Rabeyron</surname> <given-names>T. H.</given-names></name> <name><surname>Duggan</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Feeling the future: a meta-analysis of 90 experiments on the anticipation of random future events [version 2; referees: 2 approved].</article-title> <source><italic>F1000 Res.</italic></source> <volume>4</volume> <issue>1188</issue>. <pub-id pub-id-type="doi">10.12688/f1000research.7177.2</pub-id> <pub-id pub-id-type="pmid">26834996</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bem</surname> <given-names>D. J.</given-names></name></person-group> (<year>2011</year>). <article-title>Feeling the future: experimental evidence for anomalous retroactive influences on cognition and affect.</article-title> <source><italic>J. Pers. Soc. Psychol.</italic></source> <volume>100</volume> <fpage>407</fpage>&#x2013;<lpage>425</lpage>. <pub-id pub-id-type="doi">10.1037/a0021524</pub-id> <pub-id pub-id-type="pmid">21280961</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bem</surname> <given-names>D. J.</given-names></name> <name><surname>Utts</surname> <given-names>J.</given-names></name> <name><surname>Johnson</surname> <given-names>W. O.</given-names></name></person-group> (<year>2011</year>). <article-title>Reply. Must psychologists change the way they analyze their data?</article-title> <source><italic>J. Pers. Soc. Psychol.</italic></source> <volume>101</volume> <fpage>716</fpage>&#x2013;<lpage>719</lpage>. <pub-id pub-id-type="doi">10.1037/a0024777</pub-id> <pub-id pub-id-type="pmid">21928916</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benjamin</surname> <given-names>D. J.</given-names></name> <name><surname>Berger</surname> <given-names>J.</given-names></name> <name><surname>Johannesson</surname> <given-names>M.</given-names></name> <name><surname>Nosek</surname> <given-names>B. A.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E. -J.</given-names></name> <name><surname>Berk</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2017</year>). <source><italic>Redefine Statistical Significance.</italic></source> <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://psyarxiv.com/mky9j">psyarxiv.com/mky9j</ext-link></comment></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cafri</surname> <given-names>G.</given-names></name> <name><surname>Kromrey</surname> <given-names>J. D.</given-names></name> <name><surname>Brannick</surname> <given-names>M. T.</given-names></name></person-group> (<year>2010</year>). <article-title>A meta-meta-analysis: empirical review of statistical power, type I error rates, effect sizes, and model selection of meta-analyses published in psychology.</article-title> <source><italic>Multivariate Behav. Res.</italic></source> <volume>45</volume> <fpage>239</fpage>&#x2013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1080/00273171003680187</pub-id> <pub-id pub-id-type="pmid">26760285</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>M.-L. E.</given-names></name> <name><surname>Arvey</surname> <given-names>R. D.</given-names></name></person-group> (<year>2012</year>). <article-title>Meta-analysis and the development of knowledge.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>79</fpage>&#x2013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1177/1745691611429355</pub-id> <pub-id pub-id-type="pmid">26168427</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1962</year>). <article-title>The statistical power analysis for the behavioral sciences: a review.</article-title> <source><italic>J. Abnorm. Soc. Psychol.</italic></source> <volume>65</volume> <fpage>145</fpage>&#x2013;<lpage>153</lpage>.</citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1977</year>). <source><italic>Statistical Power Analysis for the Behavioral Sciences</italic></source> <comment>(Rev. ed.)</comment>. <publisher-loc>London</publisher-loc>: <publisher-name>Academic Press</publisher-name>.</citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cowles</surname> <given-names>M.</given-names></name></person-group> (<year>1989</year>). <source><italic>Statistics in Psychology. An Historical Perspective.</italic></source> <publisher-loc>Hillsdale</publisher-loc>: <publisher-name>Erlbaum</publisher-name>.</citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ellemers</surname> <given-names>N.</given-names></name></person-group> (<year>2013</year>). <article-title>Connecting the dots: mobilizing theory to reveal the big picture in social psychology (and why we should do this).</article-title> <source><italic>Eur. J. Soc. Psychol.</italic></source> <volume>43</volume> <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1002/ejsp.1932</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eriksson</surname> <given-names>K.</given-names></name> <name><surname>Simpson</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>Editorial decisions may perpetuate belief in invalid research findings.</article-title> <source><italic>PLOS ONE</italic></source> <volume>8</volume>:<issue>e73364</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0073364</pub-id> <pub-id pub-id-type="pmid">24023863</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Etz</surname> <given-names>A.</given-names></name> <name><surname>Vandekerckhove</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>A Bayesian perspective on the reproducibility project: psychology.</article-title> <source><italic>PLOS ONE</italic></source> <volume>11</volume>:<issue>0149794</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0149794</pub-id> <pub-id pub-id-type="pmid">26919473</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fanelli</surname> <given-names>D.</given-names></name> <name><surname>Gl&#x00E4;nzel</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>Bibliometric evidence for a hierarchy of the sciences.</article-title> <source><italic>PLOS ONE</italic></source> <volume>8</volume>:<issue>e66938</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0066938</pub-id> <pub-id pub-id-type="pmid">23840557</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Faul</surname> <given-names>F.</given-names></name> <name><surname>Erdfelder</surname> <given-names>E.</given-names></name> <name><surname>Lang</surname> <given-names>A.-G.</given-names></name> <name><surname>Buchner</surname> <given-names>A.</given-names></name></person-group> (<year>2007</year>). <article-title>G<sup>&#x2217;</sup>Power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences.</article-title> <source><italic>Behav. Res. Methods</italic></source> <volume>39</volume> <fpage>175</fpage>&#x2013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.3758/BF03193146</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferguson</surname> <given-names>C. J.</given-names></name> <name><surname>Heene</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>A vast graveyard of undead theories: publication bias and psychological science&#x2019;s aversion to the null.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>555</fpage>&#x2013;<lpage>561</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612459059</pub-id> <pub-id pub-id-type="pmid">26168112</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>R. A.</given-names></name></person-group> (<year>1956</year>). <source><italic>Statistical Methods and Scientific Inference.</italic></source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Hafner</publisher-name>.</citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Francis</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>The psychology of replication and replication in psychology.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>585</fpage>&#x2013;<lpage>594</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612459520</pub-id> <pub-id pub-id-type="pmid">26168115</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fuchs</surname> <given-names>H. M.</given-names></name> <name><surname>Jenny</surname> <given-names>M.</given-names></name> <name><surname>Fiedler</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Psychologists are open to change, yet wary of rules.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>639</fpage>&#x2013;<lpage>642</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612459521</pub-id> <pub-id pub-id-type="pmid">26168123</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galak</surname> <given-names>J.</given-names></name> <name><surname>LeBoeuf</surname> <given-names>R. A.</given-names></name> <name><surname>Nelson</surname> <given-names>L. D.</given-names></name> <name><surname>Simmons</surname> <given-names>J. P.</given-names></name></person-group> (<year>2012</year>). <article-title>Correcting the past: failures to replicate Psi.</article-title> <source><italic>J. Pers. Soc. Psychol.</italic></source> <volume>103</volume> <fpage>933</fpage>&#x2013;<lpage>948</lpage>. <pub-id pub-id-type="doi">10.1037/a0029709</pub-id> <pub-id pub-id-type="pmid">22924750</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gelman</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Induction and deduction in Bayesian data analysis.</article-title> <source><italic>Ration. Mark. Morals</italic></source> <volume>2</volume> <fpage>67</fpage>&#x2013;<lpage>78</lpage>.</citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gigerenzer</surname> <given-names>G.</given-names></name></person-group> (<year>2010</year>). <article-title>Personal reflections on theory and psychology.</article-title> <source><italic>Theory Psychol.</italic></source> <volume>20</volume> <fpage>733</fpage>&#x2013;<lpage>743</lpage>. <pub-id pub-id-type="doi">10.1177/0959354310378184</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gilbert</surname> <given-names>D. T.</given-names></name> <name><surname>King</surname> <given-names>G.</given-names></name> <name><surname>Pettigrew</surname> <given-names>S.</given-names></name> <name><surname>Wilson</surname> <given-names>T. D.</given-names></name></person-group> (<year>2016</year>). <article-title>Comment on &#x201C;Estimating the reproducibility of psychological science.</article-title> <source><italic>Science</italic></source> <volume>351</volume> <issue>1037</issue>. <pub-id pub-id-type="doi">10.1126/science.aad7243</pub-id> <pub-id pub-id-type="pmid">26941311</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hacking</surname> <given-names>I.</given-names></name></person-group> (<year>1978</year>). <article-title>Imre Lakatos&#x2019;s philosophy of science.</article-title> <source><italic>Br. J. Philos. Sci.</italic></source> <volume>30</volume> <fpage>381</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1093/bjps/30.4.381</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harlow</surname> <given-names>L. L.</given-names></name> <name><surname>Mulaik</surname> <given-names>S. A.</given-names></name> <name><surname>Steiger</surname> <given-names>J. H.</given-names></name></person-group> <comment>(eds)</comment> (<year>1997</year>). <source><italic>What If There Were No Significance Tests?</italic></source> <publisher-loc>Mahwah</publisher-loc>: <publisher-name>Erlbaum</publisher-name>.</citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holcombe</surname> <given-names>A. O.</given-names></name></person-group> (<year>2016</year>). <article-title>Introduction to a registered replication report on ego depletion.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>11</volume> <issue>545</issue>. <pub-id pub-id-type="doi">10.1177/1745691616652871</pub-id> <pub-id pub-id-type="pmid">27474141</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoyningen-Huene</surname> <given-names>P.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x201C;Context of discovery vs. context of justification and Thomas Kuhn,&#x201D; in</article-title> <source><italic>Revisiting Discovery and Justification</italic></source> <role>eds</role> <person-group person-group-type="editor"><name><surname>Schickore</surname> <given-names>J.</given-names></name> <name><surname>Steinle</surname> <given-names>F.</given-names></name></person-group> (<publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>119</fpage>&#x2013;<lpage>131</lpage>.</citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ioannidis</surname> <given-names>J. P. A.</given-names></name></person-group> (<year>2012</year>). <article-title>Why science is not necessarily self-correcting.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>645</fpage>&#x2013;<lpage>654</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612464056</pub-id> <pub-id pub-id-type="pmid">26168125</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ioannidis</surname> <given-names>J. P. A.</given-names></name></person-group> (<year>2014</year>). <article-title>How to make more published research true.</article-title> <source><italic>PLOS Med.</italic></source> <volume>11</volume>:<issue>e1001747</issue>. <pub-id pub-id-type="doi">10.1371/journal.pmed.1001747</pub-id> <pub-id pub-id-type="pmid">25334033</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jeffrey</surname> <given-names>H.</given-names></name></person-group> (<year>1961</year>). <source><italic>The Theory of Probability.</italic></source> <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kerr</surname> <given-names>N. L.</given-names></name></person-group> (<year>1998</year>). <article-title>HARKing: hypothesizing after the results are known.</article-title> <source><italic>Pers. Soc. Psychol. Rev.</italic></source> <volume>2</volume> <fpage>196</fpage>&#x2013;<lpage>217</lpage>. <pub-id pub-id-type="doi">10.1207/s15327957pspr0203_4</pub-id> <pub-id pub-id-type="pmid">15647155</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klein</surname> <given-names>S. B.</given-names></name></person-group> (<year>2014</year>). <article-title>What can recent replication failures tell us about the theoretical commitments of psychology?</article-title> <source><italic>Theory Psychol.</italic></source> <volume>24</volume> <fpage>326</fpage>&#x2013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1177/0959354314529616</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Simmons</surname> <given-names>W. K.</given-names></name> <name><surname>Bellgowan</surname> <given-names>P. S. F.</given-names></name> <name><surname>Baker</surname> <given-names>C. I.</given-names></name></person-group> (<year>2009</year>). <article-title>Circular analysis in systems neuroscience: the dangers of double dipping.</article-title> <source><italic>Nat. Neurosci.</italic></source> <volume>12</volume> <fpage>535</fpage>&#x2013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1038/nn.2303</pub-id> <pub-id pub-id-type="pmid">19396166</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>T. S.</given-names></name></person-group> (<year>1970</year>). <source><italic>The Structure of Scientific Revolutions</italic></source> <edition>2nd Edn.</edition> <publisher-loc>Chicago</publisher-loc>: <publisher-name>University of Chicago Press</publisher-name>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lakatos</surname> <given-names>I.</given-names></name></person-group> (<year>1978</year>). <source><italic>The Methodology of Scientific Research Programs.</italic></source> <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Larvor</surname> <given-names>B.</given-names></name></person-group> (<year>1998</year>). <source><italic>Lakatos: An Introduction.</italic></source> <publisher-loc>London</publisher-loc>: <publisher-name>Routledge</publisher-name>.</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ly</surname> <given-names>A.</given-names></name> <name><surname>Verhagen</surname> <given-names>J.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name></person-group> (<year>2016</year>). <article-title>Harold Jeffrey&#x2019;s default Bayes factor hypothesis tests: explanation, extension, and application in psychology.</article-title> <source><italic>J. Math. Psychol.</italic></source> <volume>72</volume> <fpage>19</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmp.2015.06.004</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maxwell</surname> <given-names>S. E.</given-names></name></person-group> (<year>2004</year>). <article-title>The persistence of underpowered studies in psychological research: causes, consequences, and remedies.</article-title> <source><italic>Psychol. Methods</italic></source> <volume>9</volume> <fpage>147</fpage>&#x2013;<lpage>163</lpage>. <pub-id pub-id-type="doi">10.1037/1082-989X.9.2.147</pub-id> <pub-id pub-id-type="pmid">15137886</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mayo</surname> <given-names>D. G.</given-names></name></person-group> (<year>1996</year>). <source><italic>Error and the Growth of Experimental Knowledge.</italic></source> <publisher-loc>Chicago</publisher-loc>: <publisher-name>University of Chicago Press</publisher-name>. <pub-id pub-id-type="doi">10.7208/chicago/9780226511993.001.0001</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mayo</surname> <given-names>D. G.</given-names></name></person-group> (<year>2011</year>). <article-title>Statistical science and philosophy of science: where do/should they meet in 2011 (and beyond)?</article-title> <source><italic>Ration. Mark. Morals</italic></source> <volume>2</volume> <fpage>79</fpage>&#x2013;<lpage>102</lpage>.</citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meehl</surname> <given-names>P. E.</given-names></name></person-group> (<year>1967</year>). <article-title>Theory testing in psychology and physics: a methodological paradox.</article-title> <source><italic>Philos. Sci.</italic></source> <volume>34</volume> <fpage>103</fpage>&#x2013;<lpage>115</lpage>. <pub-id pub-id-type="doi">10.1086/288135</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mitchell</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>Revisiting truth or triviality: the external validity of research in the psychological laboratory.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>109</fpage>&#x2013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1177/1745691611432343</pub-id> <pub-id pub-id-type="pmid">26168439</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Motterlini</surname> <given-names>M.</given-names></name></person-group> (<year>2002</year>). <article-title>Reconstructing Lakatos: a reassessment of Lakatos&#x2019; epistemological project in the light of the Lakatos Archive.</article-title> <source><italic>Stud. Hist. Philos. Sci.</italic></source> <volume>33</volume> <fpage>487</fpage>&#x2013;<lpage>509</lpage>. <pub-id pub-id-type="doi">10.1016/S0039-3681(02)00024-9</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nosek</surname> <given-names>B. A.</given-names></name> <name><surname>Alter</surname> <given-names>G.</given-names></name> <name><surname>Banks</surname> <given-names>G. C.</given-names></name> <name><surname>Borsboom</surname> <given-names>D.</given-names></name> <name><surname>Bowman</surname> <given-names>S. D.</given-names></name> <name><surname>Breckler</surname> <given-names>S. J.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Promoting an open research culture.</article-title> <source><italic>Science</italic></source> <volume>348</volume> <fpage>1422</fpage>&#x2013;<lpage>1425</lpage>. <pub-id pub-id-type="doi">10.1126/science.aab2374</pub-id> <pub-id pub-id-type="pmid">26113702</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nosek</surname> <given-names>B. A.</given-names></name> <name><surname>Bar-Anan</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>Scientific utopia: I. Opening scientific communication.</article-title> <source><italic>Psychol. Inq.</italic></source> <volume>23</volume> <fpage>217</fpage>&#x2013;<lpage>243</lpage>. <pub-id pub-id-type="doi">10.1080/1047840X.2012.692215</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nosek</surname> <given-names>B. A.</given-names></name> <name><surname>Spies</surname> <given-names>J. R.</given-names></name> <name><surname>Motyl</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Scientific utopia: II. Restructuring incentives and practices to promote truth over publishability.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>615</fpage>&#x2013;<lpage>631</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612459058</pub-id> <pub-id pub-id-type="pmid">26168121</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><collab>Open Science Collaboration.</collab> (<year>2015</year>). <article-title>Estimating the reproducibility of psychological science.</article-title> <source><italic>Science</italic></source> <volume>349</volume>:<issue>acc4716</issue>. <pub-id pub-id-type="doi">10.1126/science.aac4716</pub-id> <pub-id pub-id-type="pmid">26315443</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pashler</surname> <given-names>H.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name></person-group> (<year>2012</year>). <article-title>Editors&#x2019; introduction to the special section on replicability in psychological science: a crisis of confidence?</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>528</fpage>&#x2013;<lpage>530</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612465253</pub-id> <pub-id pub-id-type="pmid">26168108</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reichenbach</surname> <given-names>H.</given-names></name></person-group> (<year>1938</year>). <source><italic>Experience and Prediction.</italic></source> <publisher-loc>Chicago</publisher-loc>: <publisher-name>The University of Chicago Press</publisher-name>.</citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosnow</surname> <given-names>R.</given-names></name> <name><surname>Rosenthal</surname> <given-names>R.</given-names></name></person-group> (<year>1989</year>). <article-title>Statistical procedures and the justification of knowledge in psychological science.</article-title> <source><italic>Am. Psychol.</italic></source> <volume>44</volume> <fpage>1276</fpage>&#x2013;<lpage>1284</lpage>. <pub-id pub-id-type="doi">10.1037/0003-066X.44.10.1276</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rouder</surname> <given-names>J. N.</given-names></name></person-group> (<year>2014</year>). <article-title>Optional stopping: no problem for Bayesians.</article-title> <source><italic>Psychon. Bull. Rev.</italic></source> <volume>21</volume> <fpage>301</fpage>&#x2013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-014-0595-4</pub-id> <pub-id pub-id-type="pmid">24659049</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rouder</surname> <given-names>J. N.</given-names></name> <name><surname>Speckman</surname> <given-names>P. L.</given-names></name> <name><surname>Sun</surname> <given-names>D.</given-names></name> <name><surname>Morey</surname> <given-names>R. D.</given-names></name> <name><surname>Iverson</surname> <given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>Baysian t-tests for accepting and rejecting the null hypothesis.</article-title> <source><italic>Psychon. Bull. Rev.</italic></source> <volume>16</volume> <fpage>225</fpage>&#x2013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.3758/PBR.16.2.225</pub-id> <pub-id pub-id-type="pmid">19293088</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Royall</surname> <given-names>R.</given-names></name></person-group> (<year>1997</year>). <source><italic>Statistical Evidence. A Likelihood Paradigm.</italic></source> <publisher-loc>London</publisher-loc>: <publisher-name>Chapman &#x0026; Hall</publisher-name>.</citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scargle</surname> <given-names>J. D.</given-names></name></person-group> (<year>2000</year>). <article-title>Publication bias: the &#x201C;file-drawer&#x201D; problem in scientific inference.</article-title> <source><italic>J. Sci. Explor.</italic></source> <volume>14</volume> <fpage>91</fpage>&#x2013;<lpage>106</lpage>.</citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schickore</surname> <given-names>J.</given-names></name> <name><surname>Steinle</surname> <given-names>F.</given-names></name></person-group> (eds) (<year>2006</year>). <source><italic>Revisiting Discovery and Justification: Historical and Philosophical Perspectives on the Contest Distinction.</italic></source> <publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidt</surname> <given-names>F. L.</given-names></name> <name><surname>Oh</surname> <given-names>I.</given-names></name> <name><surname>Hayes</surname> <given-names>T. L.</given-names></name></person-group> (<year>2009</year>). <article-title>Fixed versus random-effect models in meta-analysis: model properties and an empirical comparison of differences in results.</article-title> <source><italic>Br. J. Math. Stat. Psychol.</italic></source> <volume>62</volume> <fpage>97</fpage>&#x2013;<lpage>128</lpage>. <pub-id pub-id-type="doi">10.1348/000711007X255327</pub-id> <pub-id pub-id-type="pmid">18001516</pub-id></citation></ref>
<ref id="B60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schonemann</surname> <given-names>P. H.</given-names></name> <name><surname>Scargle</surname> <given-names>J. D.</given-names></name></person-group> (<year>2008</year>). <article-title>A generalized publication bias model.</article-title> <source><italic>Chin. J. Psychol.</italic></source> <volume>50</volume> <fpage>21</fpage>&#x2013;<lpage>29</lpage>.</citation></ref>
<ref id="B61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sedlmeier</surname> <given-names>P.</given-names></name> <name><surname>Gigerenzer</surname> <given-names>G.</given-names></name></person-group> (<year>1989</year>). <article-title>Do studies of statistical power have an effect on the power of studies?</article-title> <source><italic>Psychol. Bull.</italic></source> <volume>105</volume> <fpage>309</fpage>&#x2013;<lpage>316</lpage>. <pub-id pub-id-type="doi">10.1037/0033-2909.105.2.309</pub-id></citation></ref>
<ref id="B62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simmons</surname> <given-names>J. P.</given-names></name> <name><surname>Nelson</surname> <given-names>L. D.</given-names></name> <name><surname>Simonsohn</surname> <given-names>U.</given-names></name></person-group> (<year>2011</year>). <article-title>False-positive psychology.</article-title> <source><italic>Psychol. Sci.</italic></source> <volume>22</volume> <fpage>1359</fpage>&#x2013;<lpage>1366</lpage>. <pub-id pub-id-type="doi">10.1177/0956797611417632</pub-id> <pub-id pub-id-type="pmid">22006061</pub-id></citation></ref>
<ref id="B63"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spellman</surname> <given-names>B. A.</given-names></name></person-group> (<year>2012</year>). <article-title>Introduction to the special section on research practices.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>655</fpage>&#x2013;<lpage>656</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612465075</pub-id> <pub-id pub-id-type="pmid">26168126</pub-id></citation></ref>
<ref id="B64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stegenga</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>Is meta-analysis the platinum standard of evidence?</article-title> <source><italic>Stud. Hist. Philos. Biol. Sci.</italic></source> <volume>42</volume> <fpage>497</fpage>&#x2013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1016/j.shpsc.2011.07.003</pub-id> <pub-id pub-id-type="pmid">22035723</pub-id></citation></ref>
<ref id="B65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sterling</surname> <given-names>T. D.</given-names></name></person-group> (<year>1959</year>). <article-title>Publication decisions and their possible effects on inferences drawn from tests of significance&#x2014;or vice versa.</article-title> <source><italic>J. Am. Stat. Assoc.</italic></source> <volume>54</volume> <fpage>30</fpage>&#x2013;<lpage>34</lpage>.</citation></ref>
<ref id="B66"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sturm</surname> <given-names>T. H.</given-names></name> <name><surname>M&#x00FC;lberger</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Crisis discussions in psychology: new historical and philosophical perspectives.</article-title> <source><italic>Stud. Hist. Philos. Biol. Sci.</italic></source> <volume>43</volume> <fpage>425</fpage>&#x2013;<lpage>433</lpage>. <pub-id pub-id-type="doi">10.1016/j.shpsc.2011.11.001</pub-id> <pub-id pub-id-type="pmid">22520191</pub-id></citation></ref>
<ref id="B67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Lange</surname> <given-names>P. A. M.</given-names></name> <name><surname>Kruglanski</surname> <given-names>A. W.</given-names></name> <name><surname>Higgins</surname> <given-names>E. T.</given-names></name></person-group> (<year>2012</year>). <source><italic>Handbook of Theories of Social Psychology</italic></source> <volume>Vol. 1+2</volume>. <publisher-loc>London</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation></ref>
<ref id="B68"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verhagen</surname> <given-names>A. J.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name></person-group> (<year>2014</year>). <article-title>Bayesian tests to quantify the result of a replication attempt.</article-title> <source><italic>J. Exp. Psychol. Gen.</italic></source> <volume>143</volume> <fpage>1457</fpage>&#x2013;<lpage>1475</lpage>. <pub-id pub-id-type="doi">10.1037/a0036731</pub-id> <pub-id pub-id-type="pmid">24867486</pub-id></citation></ref>
<ref id="B69"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name> <name><surname>Wetzels</surname> <given-names>R.</given-names></name> <name><surname>Borsboom</surname> <given-names>D.</given-names></name> <name><surname>van der Maas</surname> <given-names>H. L. J.</given-names></name></person-group> (<year>2011</year>). <article-title>Why psychologists must change the way they analyze their data: the psi case: comment on Bem (2011)</article-title>. <source><italic>J. Pers. Soc. Psychol.</italic></source> <volume>100</volume> <fpage>426</fpage>&#x2013;<lpage>432</lpage>. <pub-id pub-id-type="doi">10.1037/a0022790</pub-id> <pub-id pub-id-type="pmid">21280965</pub-id></citation></ref>
<ref id="B70"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name> <name><surname>Wetzels</surname> <given-names>R.</given-names></name> <name><surname>Borsboom</surname> <given-names>D.</given-names></name> <name><surname>van der Maas</surname> <given-names>H. L. J.</given-names></name> <name><surname>Kievit</surname> <given-names>R. A.</given-names></name></person-group> (<year>2012</year>). <article-title>An agenda for purely confirmatory research.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>7</volume> <fpage>632</fpage>&#x2013;<lpage>638</lpage>. <pub-id pub-id-type="doi">10.1177/1745691612463078</pub-id> <pub-id pub-id-type="pmid">26168122</pub-id></citation></ref>
<ref id="B71"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wald</surname> <given-names>A.</given-names></name></person-group> (<year>1947</year>). <source><italic>Sequential Analysis.</italic></source> <publisher-loc>New York</publisher-loc>: <publisher-name>Wiley</publisher-name>.</citation></ref>
<ref id="B72"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wetzels</surname> <given-names>R.</given-names></name> <name><surname>Matzke</surname> <given-names>D.</given-names></name> <name><surname>Lee</surname> <given-names>M. D.</given-names></name> <name><surname>Rouder</surname> <given-names>J. N.</given-names></name> <name><surname>Iverson</surname> <given-names>G. J.</given-names></name> <name><surname>Wagenmakers</surname> <given-names>E.-J.</given-names></name></person-group> (<year>2011</year>). <article-title>Statistical evidence in experimental psychology: an empirical comparison using 855 t-tests.</article-title> <source><italic>Perspect. Psychol. Sci.</italic></source> <volume>6</volume> <fpage>291</fpage>&#x2013;<lpage>298</lpage>. <pub-id pub-id-type="doi">10.1177/1745691611406923</pub-id> <pub-id pub-id-type="pmid">26168519</pub-id></citation></ref>
<ref id="B73"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Willy</surname> <given-names>R.</given-names></name></person-group> (<year>1889</year>). <source><italic>Die Krisis in der Psychologie.</italic></source> <publisher-loc>Leipzig</publisher-loc>: <publisher-name>Reisland</publisher-name>.</citation></ref>
<ref id="B74"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name></person-group> (<year>1980</year>). <source><italic>Signifikanztest und statistische Inferenz. Analysen, Probleme, Alternativen [Significance test and statistical inference: Analyses, problems, alternatives].</italic></source> <publisher-loc>Stuttgart</publisher-loc>: <publisher-name>Enke</publisher-name>.</citation></ref>
<ref id="B75"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name></person-group> (<year>1994</year>). <article-title>&#x201C;Minority influences and innovations: the search for an integrated explanation of psychological and sociological models,&#x201D; in</article-title> <source><italic>Minority Influence</italic></source> <role>eds</role> <person-group person-group-type="editor"><name><surname>Moscovici-Faina</surname> <given-names>S.</given-names></name> <name><surname>Mucchi</surname> <given-names>A.</given-names></name> <name><surname>Maass</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Chicago</publisher-loc>: <publisher-name>Nelson-Hall</publisher-name>) <fpage>67</fpage>&#x2013;<lpage>93</lpage>.</citation></ref>
<ref id="B76"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name></person-group> (<year>1996a</year>). <article-title>&#x201C;Small-group research and the crisis of social psychology: an introduction,&#x201D; in</article-title> <source><italic>Understanding Group Behavior</italic></source> <volume>Vol. 2</volume> <role>eds</role> <person-group person-group-type="editor"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Davis</surname> <given-names>J. H.</given-names></name></person-group> (<publisher-loc>Mahwah</publisher-loc>: <publisher-name>Erlbaum</publisher-name>) <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</citation></ref>
<ref id="B77"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name></person-group> (<year>1996b</year>). <article-title>&#x201C;The extended group situation theory (EGST): explaining the amount of change,&#x201D; in</article-title> <source><italic>Understanding Group Behavior</italic></source> <volume>Vol. 1</volume> <role>eds</role> <person-group person-group-type="editor"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Davis</surname> <given-names>J. H.</given-names></name></person-group> (<publisher-loc>Mahwah</publisher-loc>: <publisher-name>Erlbaum</publisher-name>) <fpage>253</fpage>&#x2013;<lpage>291</lpage>.</citation></ref>
<ref id="B78"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x201C;Theorienentwicklung und -konstruktion in der Sozialpsychologie [Theory development and theory construction in social psychology],&#x201D; in</article-title> <source><italic>Entwicklungsperspektiven der Sozialpsychologie [Developmental Perspectives of Social Psychology]</italic></source> <role>ed.</role> <person-group person-group-type="editor"><name><surname>Witte</surname> <given-names>E. H.</given-names></name></person-group> (<publisher-loc>Lengerich</publisher-loc>: <publisher-name>Pabst</publisher-name>) <fpage>172</fpage>&#x2013;<lpage>188</lpage>.</citation></ref>
<ref id="B79"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Davis</surname> <given-names>J. H.</given-names></name></person-group> <comment>(eds)</comment> (<year>1996</year>). <source><italic>Understanding Group Behavior</italic></source> <volume>Vol. 1 and 2</volume>. <publisher-loc>Mahwah</publisher-loc>: <publisher-name>Erlbaum</publisher-name>.</citation></ref>
<ref id="B80"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Heitkamp</surname> <given-names>I.</given-names></name></person-group> (<year>2006</year>). <article-title>Quantitative rekonstruktionen (retrognosen) als instrument der theorienbildung und theorienpr&#x00FC;fung in der sozialpsychologie [Quantitative reconstructions (retrognoses) as an instrument for theory construction and theory assessment in social psychology].</article-title> <source><italic>Z. Sozialpsychol.</italic></source> <volume>37</volume> <fpage>205</fpage>&#x2013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1024/0044-3514.37.3.205</pub-id></citation></ref>
<ref id="B81"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Kaufman</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <source><italic>The Stepwise Hybrid Statistical Inference Strategy: FOSTIS. HAFOS, 18.</italic></source> <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://psydok.sulb.uni-saarland.de/frontdoor.php?source_opus=2286&#x0026;la=de">http://psydok.sulb.uni-saarland.de/frontdoor.php?source_opus=2286&#x0026;la=de</ext-link> [accessed February 22 2017]</comment>.</citation></ref>
<ref id="B82"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Strohmeier</surname> <given-names>C. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Forschung in der psychologie. Ihre disziplin&#x00E4;re matrix im vergleich zu physik, biologie und sozialwissenschaft [Research in psychology. Its disciplinary matrix as compared to physics, biology, and the social sciences].</article-title> <source><italic>Psychol. Rundsch.</italic></source> <volume>64</volume> <fpage>16</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1026/0033-3042/a0000145</pub-id></citation></ref>
<ref id="B83"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Zenker</surname> <given-names>F.</given-names></name></person-group> (<year>2016a</year>). <article-title>Beyond schools&#x2014;reply to Marsman, Ly &#x0026; Wagenmakers.</article-title> <source><italic>Basic Appl. Soc. Psychol.</italic></source> <volume>38</volume> <fpage>313</fpage>&#x2013;<lpage>317</lpage>. <pub-id pub-id-type="doi">10.1080/01973533.2016.1227710</pub-id></citation></ref>
<ref id="B84"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Zenker</surname> <given-names>F.</given-names></name></person-group> (<year>2016b</year>). <article-title>Reconstructing recent work on macro-social stress as a research program.</article-title> <source><italic>Basic Appl. Soc. Psychol.</italic></source> <volume>38</volume> <fpage>301</fpage>&#x2013;<lpage>307</lpage>. <pub-id pub-id-type="doi">10.1080/01973533.2016.1207077</pub-id></citation></ref>
<ref id="B85"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witte</surname> <given-names>E. H.</given-names></name> <name><surname>Zenker</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>Extending a multilab preregistered replication of the ego-depletion effect to a research program.</article-title> <source><italic>Basic Appl. Soc. Psychol.</italic></source> <volume>39</volume> <fpage>74</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1080/01973533.2016.1269286</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn01"><label>1</label><p>In fact, OSC reported an average replicability-rate of some 36%. Worse still, a Bayesian approach finds clear and consistent results in merely 11% of 72 reanalyzed datasets (<xref ref-type="bibr" rid="B16">Etz and Vandekerckhove, 2016</xref>). Though <xref ref-type="bibr" rid="B26">Gilbert et al. (2016</xref>, p. 1037-a) submit that &#x201C;[i]f OSC (Open Science Collaboration) had limited their analyses to endorsed studies, they would have found 59.7% [95% confidence interval: 47.5, 70.9%] were replicated successfully," it is clear that even 59.7% is insufficient to regain trust.</p></fn>
<fn id="fn02"><label>2</label><p>As a recent survey indicates, psychologists are open to milder changes (<xref ref-type="bibr" rid="B22">Fuchs et al., 2012</xref>). The accepted rules of best research practice, for instance, shall not be turned into binding publication conditions. Moreover, 84% among respondents find that reviewers should be more tolerant of imperfections in what their peers submit for publication (ibid., p. 640). Additional proposals include intensifying communication, pre-registering hypotheses, and exchanging data- and design-characteristics (e.g., <xref ref-type="bibr" rid="B48">Nosek and Bar-Anan, 2012</xref>; <xref ref-type="bibr" rid="B47">Nosek et al., 2015</xref>).</p></fn>
<fn id="fn03"><label>3</label><p>Introduced by <xref ref-type="bibr" rid="B52">Reichenbach (1938)</xref>, the DJ-distinction has been contended from diverse perspectives, notably since the reception of <xref ref-type="bibr" rid="B37">Kuhn (1970)</xref>. Extant discussion of the distinction&#x2019;s multiple versions questions the separability of both contexts on temporal, methodological, goal- or question-related criteria. See <xref ref-type="bibr" rid="B58">Schickore and Steinle (2006)</xref> and references provided there.</p></fn>
<fn id="fn04"><label>4</label><p>The inferential strategy generating this question entails an inductive transition. After all, by formulating a specific non-random hypothesis the probability model, as it were, transforms &#x201C;observational&#x201D; data into a quasi-theoretical hypothesis. So we run into the unsolved issue of justifying induction as a <italic>valid</italic> inference (see below).</p></fn>
<fn id="fn05"><label>5</label><p>This praxis is not restricted to NHST but also implicates proponents of Bayes-factor testing. <xref ref-type="bibr" rid="B68">Verhagen and Wagenmakers (2014</xref>, p. 1461), for instance, state that &#x201C;[t]he problem with this analysis [likelihood-testing] is that the exact alternative effect size &#x03B4;<sub>a</sub> is never known beforehand. In Bayesian statistics, this uncertainty about &#x03B4; is addressed by assigning it a prior distribution.&#x201D; Indeed, &#x201C;[t]he major drawback of this procedure is that it is based on a point estimate, thereby ignoring the precision with which the effect size is estimated&#x201D; (ibid., p. 1463). Similarly, &#x201C;[w]e assumed that the alternative was at a single point. This assumption, however, is too restrictive to be practical&#x201D; (<xref ref-type="bibr" rid="B55">Rouder et al., 2009</xref>, p. 229).</p></fn>
<fn id="fn06"><label>6</label><p>We use &#x2018;g&#x2019; to express the effect size as the difference between a stipulated proportion of 0.50 and an observed proportion, using the sign test; &#x2018;d&#x2019; denotes a difference between numerical means (<xref ref-type="bibr" rid="B12">Cohen, 1977</xref>). Notice that when errors and likelihoods are calculated via a normal curve approximation of the effect size, the model remains constant under different proof distributions.</p></fn>
<fn id="fn07"><label>7</label><p>Here restricting the focus to a two sample (Neyman&#x2013;Pearson) <italic>t</italic>-test with &#x03B1;- and &#x03B2;-errors, a draft on the one sample <italic>t</italic>-test and the negative consequences of various &#x201C;saving-strategies&#x201D; to reduce the sample size can be obtained from the authors.</p></fn>
<fn id="fn08"><label>8</label><p>It is natural to object that very few people may in fact command psi-abilities. This might seem to explain why Bem can measure only a small effect. However, this explanation-sketch presupposes that we could (somehow) aggregate effect sizes from individually underpowered studies into a &#x201C;pooled&#x201D; effect size. Instead, what we can safely aggregate are log-likelihoods (see <xref ref-type="bibr" rid="B83">Witte and Zenker, 2016a</xref>,<xref ref-type="bibr" rid="B84">b</xref>, <xref ref-type="bibr" rid="B85">2017</xref>). The explanation-sketch might nevertheless lead to a new hypothesis, namely: can we reliably separate subjects into those who command and those who lack psi-abilities? If so, we should next test among the former group if the psi-effect is stable. At any rate, research addressing such hypotheses qualifies as a discovery context activity.</p></fn>
<fn id="fn09"><label>9</label><p>Such research clearly differs from parameter-estimation because it can bury &#x201C;undead&#x201D; theories. But this is impossible given how standard methods are mostly used today (<xref ref-type="bibr" rid="B19">Ferguson and Heene, 2012</xref>; <xref ref-type="bibr" rid="B15">Eriksson and Simpson, 2013</xref>).</p></fn>
<fn id="fn010"><label>10</label><p>If the likelihood value of the theoretical parameter does <italic>not</italic> fall outside of the 95%-interval placed around the maximum-likelihood-estimate, then we view the theoretical parameter as substantially verified, and thus corroborated. After all, both the empirical result and the theoretical assumption now lie within an acceptable interval. The corroboration threshold is given by the ratio of the two likelihood-values, i.e., the maximum ordinate of the normal curve (0.3989) and the ordinate at the 95%-interval (0.10). To good approximation, this yields 4.</p></fn>
</fn-group>
</back>
</article>