<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="article-commentary">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2017.01715</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>General Commentary</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Commentary: Psychological Science&#x00027;s Aversion to the Null</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Perezgonzalez</surname> <given-names>Jose D.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/197438/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Fr&#x000ED;as-Navarro</surname> <given-names>Dolores</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/355485/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Pascual-Llobell</surname> <given-names>Juan</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/445141/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Business School, Massey University</institution>, <addr-line>Palmerston North</addr-line>, <country>New Zealand</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Methodology of the Behavioral Sciences, Universitat de Val&#x000E8;ncia</institution>, <addr-line>Valencia</addr-line>, <country>Spain</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Hannes Schr&#x000F6;ter, German Institute for Adult Education (LG), Germany</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Daniel Bratzke, Universit&#x000E4;t T&#x000FC;bingen, Germany</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Jose D. Perezgonzalez <email>j.d.perezgonzalez&#x00040;massey.ac.nz</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Educational Psychology, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>09</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>1715</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>05</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>09</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Perezgonzalez, Fr&#x000ED;as-Navarro and Pascual-Llobell.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Perezgonzalez, Fr&#x000ED;as-Navarro and Pascual-Llobell</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<kwd-group>
<kwd>data testing</kwd>
<kwd>hypothesis testing</kwd>
<kwd>null hypothesis significance testing</kwd>
<kwd>effect size</kwd>
<kwd>falsificationism</kwd>
<kwd>statistics</kwd>
</kwd-group>
<counts>
<fig-count count="0"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="16"/>
<page-count count="2"/>
<word-count count="1653"/>
</counts>
</article-meta>
</front>
<body>
<p>A commentary on <bold>Psychological Science&#x00027;s Aversion to the Null</bold> <italic>by Heene, M., and Ferguson, C. J. (2017). Psychological Science under Scrutiny: Recent Challenges and Proposed Solutions, eds S. O. Lilienfeld and I. D. Waldman (Chichester: John Wiley &#x00026; Sons), 34&#x02013;52.</italic></p>
<p>Heene and Ferguson (<xref ref-type="bibr" rid="B5">2017</xref>) contributed important epistemological, ethical and didactical ideas to the debate on null hypothesis significance testing, chief among them ideas about falsificationism, statistical power, dubious statistical practices, and publication bias. Important as those contributions are, the authors do not fully resolve four confusions which we would like to clarify.</p>
<p>One confusion is equating the null hypothesis (H<sub>0</sub>) with randomness when &#x0201C;chance&#x0201D; actually resides in the sample. We can, indeed, read three different instances of randomness in the text: associated with the sample on pages 36 (trial performance) and 37; associated with the alternative hypothesis (H<sub>A</sub>) on page 41 (&#x0201C;less likely to observe mean differences&#x02026;far off the true&#x02026;mean difference of 0.7&#x0201D;); and associated with H<sub>0</sub> throughout the text, starting on page 36. In reality, H<sub>0</sub> simply claims a population non-effect (H<sub>0</sub>: &#x00394; &#x0003D; 0) while H<sub>A</sub> claims a constant effect (e.g., H<sub>A</sub>: &#x00394; &#x0003D; 0.7), their corresponding distributions assuming random sampling variation in both cases. It is in the (random) sample where &#x0201C;chance&#x0201D; resides, as by chance we may pick a sample which shows a given effect (e.g., &#x003B4; &#x0003D; 0.3) when the true effect in the population is either &#x0201C;0&#x0201D; (H<sub>0</sub>) or &#x0201C;0.7&#x0201D; (H<sub>A</sub>). Frequentist tests only assess the probability of getting the observed sample effect under H<sub>0</sub> while Bayesian statistics also assesses the probability of such effect under H<sub>A</sub> (e.g., Rouder et al., <xref ref-type="bibr" rid="B16">2009</xref>). Therefore, the <italic>p</italic>-value does not inform about a hypothesis of chance but about the probability of the data under H<sub>0</sub> (Fisher, <xref ref-type="bibr" rid="B3">1954</xref>).</p>
<p>A second issue confuses power with missing true effects, something explicitly expressed on page 42 but also suggested when discussing sample sizes throughout the text (p. 36 onwards). The underlying argument is that larger sample sizes allow for achieving statistical significance so that a true effect may not be missed&#x02014;something which is, at the same time, portrayed as unethical, e.g., p. 36, and ludicrous, e.g., p. 44. In reality, &#x0201C;we cannot manipulate population effect sizes&#x0201D; (p. 41), as they are deemed constant in the population (e.g., H<sub>A</sub>: &#x00394; &#x0003D; 0.7), and a significant result at 50% power will not be missed at 80% power. As Heene and Ferguson&#x00027;s Figures 3.1A,C show, power simply moves the goalposts on the real line, reducing the Type II error (&#x003B2;), while the larger sample size also reduces the standard error. By moving the goalposts, smaller (by chance) sample effects get associated with H<sub>A</sub>, which is a correct association as long as there is a true population effect. Thus, power is there not to prevent missing effects due to small sample sizes but to be able to justify whether we could plausibly accept H<sub>0</sub> when results are not significant (Neyman, <xref ref-type="bibr" rid="B11">1955</xref>; Cohen, <xref ref-type="bibr" rid="B1">1988</xref>).</p>
<p>A third issue is about falsificationism (pp. 35&#x02013;37), which the authors argue cannot happen in psychology because we never accept H<sub>0</sub>, only reject it or fail to reject it. In reality, frequentist tests are logically based on <italic>modus tollens</italic>, the valid argument form for the falsification of statements (Perezgonzalez, <xref ref-type="bibr" rid="B14">2017a</xref>). H<sub>0</sub> is simply the contrapositive of our research hypothesis, and denying H<sub>0</sub> allows us to affirm the latter. Therefore, frequentist tests are eminently falsificationist, attempting to disprove H<sub>0</sub> via <italic>reductio</italic> arguments (<italic>p</italic>, &#x003B1;; Mayo, <xref ref-type="bibr" rid="B8">2017</xref>). Indeed, H<sub>0</sub> does not even need to be &#x0201C;zero&#x0201D; in the population: We could perfectly substitute the actual value of our H<sub>A</sub>, so that we may prove the theory false with a significant result (the &#x0201C;strong&#x0201D; test purported by Meehl, <xref ref-type="bibr" rid="B10">1997</xref>).</p>
<p>A fourth issue is whether we always need to be in the position of accepting H<sub>0</sub> (something argued on pages 36&#x02013;37). This is not necessarily so. Just testing H<sub>0</sub> as for rejecting it is suitable when we are only interested in learning about our research hypothesis (e.g., does the treatment have an effect?&#x02014;Perezgonzalez, <xref ref-type="bibr" rid="B13">2016</xref>). In such context, H<sub>0</sub> provides a precise statistical hypothesis for carrying out the test and, because the actual parameter (&#x00394;) is unknown, it only provides informative value via its rejection (Fisher, <xref ref-type="bibr" rid="B3">1954</xref>), H<sub>0</sub> acting merely as a &#x0201C;straw man&#x0201D; (Cortina and Dunlap, <xref ref-type="bibr" rid="B2">1997</xref>). This testing procedure was not only developed in the context of small samples (Fisher, <xref ref-type="bibr" rid="B3">1954</xref>) but the lack of a specific H<sub>A</sub> precludes the control of Type II errors and of power. (A way forward would be to assess the effects warranted under H<sub>0</sub>&#x02014;Mayo and Spanos, <xref ref-type="bibr" rid="B9">2006</xref>&#x02014;or to control sample size via a sensitiveness analysis&#x02014;Perezgonzalez, <xref ref-type="bibr" rid="B15">2017b</xref>).</p>
<p>If we wish to be able to accept H<sub>0</sub>, then we are stating that we are also interested in the potential demise of our intervention (i.e., if the treatment has no effect, we want to make sure it is akin to placebo; Perezgonzalez, <xref ref-type="bibr" rid="B13">2016</xref>). This testing seems similar to Fisher&#x00027;s, but it requires active control of the severity with which the alternative hypothesis is to be tested (ideally, &#x02265;80% power; Neyman, <xref ref-type="bibr" rid="B11">1955</xref>; Cohen, <xref ref-type="bibr" rid="B1">1988</xref>). Such control necessarily means more information&#x02014;a precise alternative hypothesis (e.g., H<sub>A</sub>: &#x003BC;<sub>1</sub> &#x02013; &#x003BC;<sub>2</sub> &#x0003D; 0.7, vs. H<sub>0</sub>: &#x003BC;<sub>1</sub> &#x02013; &#x003BC;<sub>2</sub> &#x0003D; 0) and a specified Type II error for H<sub>A</sub> (e.g., &#x003B2; &#x0003D; 0.20)&#x02014;so that the power of the test can be managed (given &#x003B1;, &#x003B2;, and <italic>N</italic>). This approach not only allows for accepting H<sub>0</sub> but also illustrates that power is only relevant for such purpose, not for rejecting H<sub>0</sub>. Such approach, and similar ones, have also been available since Fisher&#x00027;s tests of significance (e.g., Neyman and Pearson, <xref ref-type="bibr" rid="B12">1928</xref>; Jeffreys, <xref ref-type="bibr" rid="B6">1939</xref>).</p>
<p>As final note, frequentist approaches only deal with the probability of data under H<sub>0</sub> [p(D|H<sub>0</sub>)]. If we want to say anything about the (posterior) probability of the hypotheses, then a Bayesian approach is needed in order to confirm which hypothesis is most likely given both the likelihood of the data and the prior probabilities of the hypotheses themselves (Jeffreys, <xref ref-type="bibr" rid="B7">1961</xref>; Gelman et al., <xref ref-type="bibr" rid="B4">2013</xref>).</p>
<sec id="s1">
<title>Author contributions</title>
<p>JDP initiated and drafted the general commentary. DF and JP contributed theoretical background and feedback. All authors approved the final version of the manuscript for submission.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1988</year>). <source>Statistical Power Analysis for the Behavioral Sciences, 2nd Edn</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Psychology Press</publisher-name>.</citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cortina</surname> <given-names>J. M.</given-names></name> <name><surname>Dunlap</surname> <given-names>W. P.</given-names></name></person-group> (<year>1997</year>). <article-title>On the logic and purpose of significance testing</article-title>. <source>Psychol. Methods</source> <volume>2</volume>, <fpage>161</fpage>&#x02013;<lpage>172</lpage>. <pub-id pub-id-type="doi">10.1037/1082-989X.2.2.161</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>R. A.</given-names></name></person-group> (<year>1954</year>). <source>Statistical Methods for Research Workers, 12th Edn.</source> <publisher-loc>Edinburgh</publisher-loc>: <publisher-name>Oliver and Boyd</publisher-name>.</citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gelman</surname> <given-names>A.</given-names></name> <name><surname>Carlin</surname> <given-names>J. B.</given-names></name> <name><surname>Stern</surname> <given-names>H. S.</given-names></name> <name><surname>Dunson</surname> <given-names>D. B.</given-names></name> <name><surname>Vehtari</surname> <given-names>A.</given-names></name> <name><surname>Rubin</surname> <given-names>D. B.</given-names></name></person-group> (<year>2013</year>). <source>Bayesian Data Analysis, 3rd Edn</source>. <publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Heene</surname> <given-names>M.</given-names></name> <name><surname>Ferguson</surname> <given-names>C. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Psychological science&#x00027;s aversion to the null, and why many of the things you think are true, aren&#x00027;t</article-title>, in <source>Psychological Science under Scrutiny: Recent Challenges and Proposed Solutions</source>, eds <person-group person-group-type="editor"><name><surname>Lilienfeld</surname> <given-names>S. O.</given-names></name> <name><surname>Waldman</surname> <given-names>I. D.</given-names></name></person-group> (<publisher-loc>Chichester</publisher-loc>: <publisher-name>John Wiley &#x00026; Sons</publisher-name>), <fpage>34</fpage>&#x02013;<lpage>52</lpage>.</citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jeffreys</surname> <given-names>H.</given-names></name></person-group> (<year>1939</year>). <source>Theory of Probability.</source> <publisher-loc>Oxford</publisher-loc>: <publisher-name>Clarendon Press</publisher-name>.</citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jeffreys</surname> <given-names>H.</given-names></name></person-group> (<year>1961</year>). <source>Theory of Probability, 3rd Edn.</source> <publisher-loc>Oxford</publisher-loc>: <publisher-name>Clarendon Press</publisher-name>.</citation></ref>
<ref id="B8">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Mayo</surname> <given-names>D. G.</given-names></name></person-group> (<year>2017</year>). <source>If you&#x00027;re Seeing Limb-Sawing in p-Value Logic, You&#x00027;re Sawing Off the Limbs of Reductio Arguments [Web log post]</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://errorstatistics.com/2017/04/15/if-youre-seeing-limb-sawing-in-p-value-logic-youre-sawing-off-the-limbs-of-reductio-arguments/">https://errorstatistics.com/2017/04/15/if-youre-seeing-limb-sawing-in-p-value-logic-youre-sawing-off-the-limbs-of-reductio-arguments/</ext-link>.</citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mayo</surname> <given-names>D. G.</given-names></name> <name><surname>Spanos</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Severe testing as a basic concept in a Neyman-Pearson philosophy of induction</article-title>. <source>Br. J. Philos. Sci.</source> <volume>57</volume>, <fpage>323</fpage>&#x02013;<lpage>357</lpage>. <pub-id pub-id-type="doi">10.1093/bjps/axl003</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Meehl</surname> <given-names>P.E.</given-names></name></person-group> (<year>1997</year>). <article-title>The problem is epistemology, not statistics: replace significance tests by confidence intervals and quantify accuracy of risky numerical predictions</article-title>, in <source>What If There Were No Significance Tests?</source> eds <person-group person-group-type="editor"><name><surname>Harlow</surname> <given-names>L. L.</given-names></name> <name><surname>Mulaik</surname> <given-names>S. A.</given-names></name> <name><surname>Steiger</surname> <given-names>J. H.</given-names></name></person-group> (<publisher-loc>Mahwah</publisher-loc>: <publisher-name>Erlbaum</publisher-name>), <fpage>393</fpage>&#x02013;<lpage>425</lpage>.</citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neyman</surname> <given-names>J.</given-names></name></person-group> (<year>1955</year>). <article-title>The problem of inductive inference</article-title>. <source>Commun. Pure Appl. Math.</source> <volume>8</volume>, <fpage>13</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1002/cpa.3160080103</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neyman</surname> <given-names>J.</given-names></name> <name><surname>Pearson</surname> <given-names>E. S.</given-names></name></person-group> (<year>1928</year>). <article-title>On the use and interpretation of certain test criteria for purposes of statistical inference: part I</article-title>. <source>Biometrika</source> <volume>20A</volume>, <fpage>175</fpage>&#x02013;<lpage>240</lpage>. <pub-id pub-id-type="doi">10.2307/2331945</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perezgonzalez</surname> <given-names>J. D.</given-names></name></person-group> (<year>2016</year>). <article-title>Commentary: how Bayes factors change scientific practice</article-title>. <source>Front. Psychol.</source> <volume>7</volume>:<fpage>1504</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2016.01504</pub-id><pub-id pub-id-type="pmid">27775731</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perezgonzalez</surname> <given-names>J. D.</given-names></name></person-group> (<year>2017a</year>). <article-title>Commentary: the need for Bayesian hypothesis testing in psychological science</article-title>. <source>Front. Psychol.</source> <volume>8</volume>:<fpage>1434</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2017.01434</pub-id><pub-id pub-id-type="pmid">28878724</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Perezgonzalez</surname> <given-names>J. D.</given-names></name></person-group> (<year>2017b</year>). <source>Statistical Sensitiveness for the Behavioral Sciences</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://osf.io/preprints/psyarxiv/qd3gu">https://osf.io/preprints/psyarxiv/qd3gu</ext-link>.</citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rouder</surname> <given-names>J. N.</given-names></name> <name><surname>Speckman</surname> <given-names>P. L.</given-names></name> <name><surname>Sun</surname> <given-names>D.</given-names></name> <name><surname>Morey</surname> <given-names>R. D.</given-names></name> <name><surname>Iverson</surname> <given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>Bayesian t-tests for accepting and rejecting the null hypothesis</article-title>. <source>Psychon. Bull. Rev.</source> <volume>16</volume>, <fpage>225</fpage>&#x02013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.3758/PBR.16.2.225</pub-id><pub-id pub-id-type="pmid">19293088</pub-id></citation></ref>
</ref-list>
</back>
</article>
