<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="review-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2017.00543</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Classical Statistics and Statistical Learning in Imaging Neuroscience</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Bzdok</surname> <given-names>Danilo</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/49944/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Psychiatry, Psychotherapy and Psychosomatics, Medical Faculty, RWTH Aachen University</institution>, <addr-line>Aachen</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Translational Brain Medicine, J&#x000FC;lich-Aachen Research Alliance (JARA)</institution>, <addr-line>Aachen</addr-line>, <country>Germany</country></aff>
<aff id="aff3"><sup>3</sup><institution>Parietal Team, Institut National de Recherche en Informatique et en Automatique (INRIA)</institution>, <addr-line>Gif-sur-Yvette</addr-line>, <country>France</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Yaroslav O. Halchenko, Dartmouth College, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Matthew Brett, University of Cambridge, United Kingdom; Jean-Baptiste Poline, University of California, Berkeley, United States</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Danilo Bzdok <email>danilo.bzdok&#x00040;rwth-aachen.de</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Brain Imaging Methods, a section of the journal Frontiers in Neuroscience</p></fn></author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>10</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>11</volume>
<elocation-id>543</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>04</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>09</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Bzdok.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Bzdok</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Brain-imaging research has predominantly generated insight by means of classical statistics, including regression-type analyses and null-hypothesis testing using <italic>t</italic>-test and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity especially for applications in rich and complex data, including cross-validated out-of-sample prediction using pattern classification and sparsity-inducing regression. This concept paper discusses the implications of inferential justifications and algorithmic methodologies in common data analysis scenarios in neuroimaging. It is retraced how classical statistics and statistical learning originated from different historical contexts, build on different theoretical foundations, make different assumptions, and evaluate different outcome metrics to permit differently nuanced conclusions. The present considerations should help reduce current confusion between model-driven classical hypothesis testing and data-driven learning algorithms for investigating the brain with imaging techniques.</p></abstract>
<kwd-group>
<kwd>neuroimaging</kwd>
<kwd>data science</kwd>
<kwd>epistemology</kwd>
<kwd>statistical inference</kwd>
<kwd>machine learning</kwd>
<kwd><italic>p</italic>-value</kwd>
<kwd>Rosetta Stone</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="222"/>
<page-count count="23"/>
<word-count count="19941"/>
</counts>
</article-meta>
</front>
<body>
<p><disp-quote>
<p><italic>&#x0201C;The trick to being a scientist is to be open to using a wide variety of tools.&#x0201D;</italic></p>
<attrib><italic>Breiman (<xref ref-type="bibr" rid="B20">2001</xref>)</italic></attrib>
</disp-quote></p>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Among the greatest challenges humans face are cultural misunderstandings between individuals, groups, and institutions (Hall, <xref ref-type="bibr" rid="B106">1976</xref>). The topic of the present paper is the culture clash between knowledge generation based on null-hypothesis testing and out-of-sample pattern generalization (Friedman, <xref ref-type="bibr" rid="B78">1998</xref>; Breiman, <xref ref-type="bibr" rid="B20">2001</xref>; Shmueli, <xref ref-type="bibr" rid="B185">2010</xref>; Donoho, <xref ref-type="bibr" rid="B57">2015</xref>). These statistical paradigms are now increasingly combined in brain-imaging studies (Kriegeskorte et al., <xref ref-type="bibr" rid="B136">2009</xref>; Varoquaux and Thirion, <xref ref-type="bibr" rid="B204">2014</xref>). Ensuing inter-cultural misunderstandings are unfortunate because the invention and application of new research methods has always been a driving force in the neurosciences (Greenwald, <xref ref-type="bibr" rid="B101">2012</xref>; Yuste, <xref ref-type="bibr" rid="B220">2015</xref>). Here the goal is to disentangle the contexts underlying <italic>classical statistical inference</italic> and <italic>out-of-sample generalization</italic> by providing a direct comparison of their historical trajectories, modeling philosophies, conceptual frameworks, and performance metrics.</p>
<p>During recent years, neuroscience has transitioned from qualitative reports of few patients with neurological brain lesions to quantitative lesion-symptom mapping on the voxel level in hundreds of patients (Gl&#x000E4;scher et al., <xref ref-type="bibr" rid="B96">2012</xref>). We have gone from manually staining and microscopically inspecting single brain slices to 3D models of neuroanatomy at micrometer scale (Amunts et al., <xref ref-type="bibr" rid="B3">2013</xref>). We have also gone from experimental studies conducted by a single laboratory to automatized knowledge aggregation across thousands of previously isolated neuroimaging findings (Yarkoni et al., <xref ref-type="bibr" rid="B216">2011</xref>; Fox et al., <xref ref-type="bibr" rid="B75">2014</xref>). Rather than laboriously collecting in-house data published in a single paper, investigators are now routinely reanalyzing multi-modal data repositories (Derrfuss and Mar, <xref ref-type="bibr" rid="B54">2009</xref>; Markram, <xref ref-type="bibr" rid="B146">2012</xref>; Van Essen et al., <xref ref-type="bibr" rid="B199">2012</xref>; Kandel et al., <xref ref-type="bibr" rid="B129">2013</xref>; Poldrack and Gorgolewski, <xref ref-type="bibr" rid="B172">2014</xref>). The detail of neuroimaging datasets is hence growing in terms of information resolution, sample size, and complexity of meta-information (Van Horn and Toga, <xref ref-type="bibr" rid="B200">2014</xref>; Eickhoff et al., <xref ref-type="bibr" rid="B65">2016</xref>; Bzdok and Yeo, <xref ref-type="bibr" rid="B30">2017</xref>). As a consequence of the data demand of many pattern-recognition algorithms, the scope of neuroimaging analyses has expanded beyond the predominance of regression-type analyses combined with null-hypothesis testing (Figure <xref ref-type="fig" rid="F1">1</xref>). Applications of statistical learning methods (i) are more data-driven due to particularly flexible models, (ii) have scaling properties compatible with high-dimensional data with myriads of input variables, and (iii) follow a heuristic agenda by prioritizing useful approximations to patterns in data (Jordan and Mitchell, <xref ref-type="bibr" rid="B126">2015</xref>; LeCun et al., <xref ref-type="bibr" rid="B139">2015</xref>; Blei and Smyth, <xref ref-type="bibr" rid="B18">2017</xref>). <italic>Statistical learning</italic> (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>) henceforth comprises the umbrella of &#x0201C;machine learning,&#x0201D; &#x0201C;data mining,&#x0201D; &#x0201C;pattern recognition,&#x0201D; &#x0201C;knowledge discovery,&#x0201D; &#x0201C;high-dimensional statistics,&#x0201D; and bears close relation to &#x0201C;data science.&#x0201D;</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Application areas of two statistical paradigms. Lists examples of research domains which apply relatively more classical statistics (blue) or learning algorithms (red). The co-occurrence of increased computational resources, growing data repositories, and improving pattern-learning techniques have initiated a shift toward less hypothesis-driven and more algorithmic methodologies. As a broad intuition, researchers in the empirical sciences on the left tend to use statistics to evaluate a pre-assumed model on the data. Researchers in the application domains on the right tend to derive a model directly from the data: A new function with potentially many parameters is created that can predict the output from the input alone without explicit programming model. One of the key differences becomes apparent when thinking of the neurobiological phenomenon under study as a black box (Breiman, <xref ref-type="bibr" rid="B20">2001</xref>). ClSt typically aims at modeling the black box by making a set of formal assumptions about its content, such as the nature of the signal distribution. Gaussian distributional assumptions have been very useful in many instances to enhance mathematical convenience and, hence, computational tractability. Instead, StLe takes a brute-force approach to model the output of the black box (e.g., tell healthy and schizophrenic people apart) from its input (e.g., volumetric brain measurements) while making a possible minimum of assumptions (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>). In ClSt the stochastic processes that generated the data is therefore treated as partly known, whereas in StLe the phenomenon is treated as complex, largely unknown, and partly unknowable.</p></caption>
<graphic xlink:href="fnins-11-00543-g0001.tif"/>
</fig>
<p>From a technical perspective, one should make a note of caution that holds across application domains such as neuroscience: While the research question often precedes the choice of statistical model, perhaps no single criterion exists that alone allows for a clear-cut distinction between classical statistics and statistical learning in all cases. For decades, the two statistical cultures have evolved in partly independent sociological niches (Breiman, <xref ref-type="bibr" rid="B20">2001</xref>). There is currently a scarcity of scientific papers and books that would provide an explicit account on how concepts and tools from classical statistics and statistical learning are exactly related to each other. Efron and Hastie are perhaps among the first to discuss the issue in their book &#x0201C;Computer-Age Statistical Inference&#x0201D; (2016). The authors cautiously conclude that statistical learning inventions, such as support vector machines, random-forest algorithms, and &#x0201C;deep&#x0201D; neural networks, can not be easily situated in the classical theory of twentieth century statistics. They go on to say that &#x0201C;pessimistically or optimistically, one can consider this as a bipolar disorder of the field or as a healthy duality that is bound to improve both branches&#x0201D; (Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>, p. 447). In the current absence of a commonly agreed-upon theoretical account from the technical literature, the present concept paper examines applications of classical statistics vs. statistical learning in the concrete context of neuroimaging analysis questions.</p>
<p>More generally, ensuring that a statistical effect discovered in one set of data extrapolates to new observations in the brain can take different forms (Efron, <xref ref-type="bibr" rid="B59">2012</xref>). As one possible definition, &#x0201C;the goal of statistical inference is to say what we have learned about the population <italic>X</italic> from the observed data x&#x0201D; (Efron and Tibshirani, <xref ref-type="bibr" rid="B62">1994</xref>). In a similar spirit, a committee report to the National Academies of the USA stated (Committee on the Analysis of Massive Data et al., <xref ref-type="bibr" rid="B44">2013</xref>, p. 8): &#x0201C;Inference is the problem of turning data into knowledge, where knowledge often is expressed in terms of variables [&#x02026;] that are not present in the data <italic>per se</italic>, but are present in models that one uses to interpret the data.&#x0201D; According to these definitions, <italic>statistical inference can be understood as encompassing not only the classical null-hypothesis testing framework but also Bayesian model inversion to compute posterior distributions as well as more recently emerged pattern-learning algorithms relying on out-of-sample generalization</italic> (cf. Gigerenzer and Murray, <xref ref-type="bibr" rid="B94">1987</xref>; Cohen, <xref ref-type="bibr" rid="B41">1990</xref>; Efron, <xref ref-type="bibr" rid="B59">2012</xref>; Ghahramani, <xref ref-type="bibr" rid="B91">2015</xref>). The important consequence for the present considerations is that classical statistics and statistical learning can give rise to different categories of inferential thinking (Chamberlin, <xref ref-type="bibr" rid="B32">1890</xref>; Platt, <xref ref-type="bibr" rid="B168">1964</xref>; Efron and Tibshirani, <xref ref-type="bibr" rid="B62">1994</xref>)&#x02014;an investigator may ask an identical neuroscientific question in different mathematical contexts.</p>
<p>For a long time, knowledge generation in psychology, neuroscience, and medicine has been dominated by classical statistics with <italic>estimation</italic> of linear-regression-like models and subsequent <italic>statistical significance testing</italic> whether an effect exists in the sample. In contrast, computation-intensive pattern learning methods have always had a strong focus on <italic>prediction</italic> in frequently extensive data with more modest concern for interpretability and the &#x0201C;right&#x0201D; underlying question (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>; Ghahramani, <xref ref-type="bibr" rid="B91">2015</xref>). In many statistical learning applications, it is standard practice to quantify the ability of a predictive pattern to extrapolate to other samples, possibly in individual subjects. In a two-step procedure, a learning algorithm is fitted on a typically bigger amount of available data (<italic>training data</italic>) and the ensuing fitted model is empirically evaluated on a commonly smaller amount of independent data (<italic>test data</italic>). This stands in contrast to classical statistical inference where the investigator seeks to reject the null hypothesis by considering the entirety of a data sample (Wasserstein and Lazar, <xref ref-type="bibr" rid="B209">2016</xref>), typically all available subjects. In this case, the desired relevance of a statistical relationship in the underlying population is ensured by formal mathematical proofs and is not commonly ascertained by explicit evaluations on new data (Breiman, <xref ref-type="bibr" rid="B20">2001</xref>; Wasserstein and Lazar, <xref ref-type="bibr" rid="B209">2016</xref>). As such, generating insight according to classical statistics and statistical learning serves rather distinct modeling purposes. Classical statistics and statistical learning do therefore not judge data on the same aspects of evidence (Breiman, <xref ref-type="bibr" rid="B20">2001</xref>; Shmueli, <xref ref-type="bibr" rid="B185">2010</xref>; Arbabshirani et al., <xref ref-type="bibr" rid="B6">2017</xref>; Bzdok and Yeo, <xref ref-type="bibr" rid="B30">2017</xref>). The two statistical cultures perform different types of principled assessment for successful extrapolation of a statistical relationship beyond the particular observations at hand.</p>
<p>Taking an epistemological perspective helps appreciating that scientific research is rarely an entirely objective process but deeply depends on the beliefs and expectations of the investigator. A new &#x0201C;scientific fact&#x0201D; about the brain is probably not established in vacuo (Fleck et al., <xref ref-type="bibr" rid="B74">1935</xref>; terms in quotes taken from source). Rather, a research &#x0201C;object&#x0201D; is recognized and accepted by the &#x0201C;subject&#x0201D; according to socially conditioned &#x0201C;thought styles&#x0201D; that are cultivated among members of &#x0201C;thought collectives.&#x0201D; A witnessed and measured neurobiological phenomenon tends to only become &#x0201C;true&#x0201D; if not at odds with the constructed &#x0201C;thought history&#x0201D; and &#x0201C;closed opinion system&#x0201D; shared by that subject. The present paper will revisit and reintegrate two such thought milieus in the context of imaging neuroscience: classical statistics (ClSt) and statistical learning (StLe).</p>
</sec>
<sec id="s2">
<title>Different histories: the origins of classical hypothesis testing and pattern-learning algorithms</title>
<p>One of many possible ways to group statistical methods is by framing them along the lines of ClSt and StLe. The incongruent historical developments of the two statistical communities are even evident from their basic terminology. Inputs to statistical models are usually called <italic>independent variables, explanatory variables</italic>, or <italic>predictors</italic> in the ClSt community, but are typically called <italic>features</italic> collected in a <italic>feature space</italic> in the StLe community. The model outputs are typically called <italic>dependent variables, explained variable</italic>, or <italic>responses</italic> in ClSt, while these are often called <italic>target variables</italic> in StLe. It follows a summary of characteristic events in the development of what can today be considered as ClSt and StLe (Figure <xref ref-type="fig" rid="F2">2</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Developments in the history of classical statistics and statistical learning. Examples of important inventions in statistical methodology. Roughly, a number of statistical methods taught in today&#x00027;s textbooks in psychology and medicine have emerged in the first half of the twentieth century (blue). Instead, many algorithmic techniques and procedures have emerged in the second half of the twentieth century (red). &#x0201C;The postwar era witnessed a massive expansion of statistical methodology, responding to the data-driven demands of modern scientific technology.&#x0201D; (Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>).</p></caption>
<graphic xlink:href="fnins-11-00543-g0002.tif"/>
</fig>
<p>Around 1900 the notions of <italic>standard deviation, goodness of fit</italic>, and the <italic>p</italic> &#x0003C; 0.05 threshold emerged (Cowles and Davis, <xref ref-type="bibr" rid="B45">1982</xref>). This was also the period when William S. Gosset published the <italic>t</italic>-test under the incognito name &#x0201C;Student&#x0201D; to quantify production quality in Guinness breweries. Motivated by concrete problems such as the interaction between potato varieties and fertilizers, Ronald A. Fisher invented the <italic>analysis of variance</italic> (ANOVA), <italic>null-hypothesis testing</italic>, promoted <italic>p-values</italic>, and devised principles of proper experimental conduct (Fisher and Mackenzie, <xref ref-type="bibr" rid="B72">1923</xref>; Fisher, <xref ref-type="bibr" rid="B70">1925</xref>, <xref ref-type="bibr" rid="B71">1935</xref>). Another framework by Jerzy Neyman and Egon S. Pearson proposed the <italic>alternative hypothesis</italic>, which allowed for the statistical notions of <italic>power, false positives</italic> and <italic>false negatives</italic>, but left out the concept of <italic>p</italic>-values (Neyman and Pearson, <xref ref-type="bibr" rid="B154">1933</xref>). This was a time before electrical calculators emerged after World War II (Efron and Tibshirani, <xref ref-type="bibr" rid="B61">1991</xref>; Gigerenzer, <xref ref-type="bibr" rid="B92">1993</xref>). Student&#x00027;s <italic>t</italic>-test and Fisher&#x00027;s inference framework were institutionalized by American psychology textbooks widely read in the 40s and 50s, while Neyman and Pearson&#x00027;s framework only became increasingly known in the 50s and 60s. Today&#x00027;s applied statistics textbooks have inherited a mixture of the Fisher and Neyman-Pearson approaches to statistical inference.</p>
<p>It is a topic of current debate<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref><sup>,</sup><xref ref-type="fn" rid="fn0002"><sup>2</sup></xref><sup>,</sup><xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> whether ClSt is a discipline that is separate from StLe (e.g., Chambers, <xref ref-type="bibr" rid="B33">1993</xref>; Breiman, <xref ref-type="bibr" rid="B20">2001</xref>; Friedman, <xref ref-type="bibr" rid="B79">2001</xref>; Bishop and Lasserre, <xref ref-type="bibr" rid="B17">2007</xref>; Shalev-Shwartz and Ben-David, <xref ref-type="bibr" rid="B184">2014</xref>; Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>) or if &#x0201C;statistics&#x0201D; denotes a broader methodological class that includes both ClSt and StLe tools as its members (e.g., Tukey, <xref ref-type="bibr" rid="B196">1962</xref>; Cleveland, <xref ref-type="bibr" rid="B38">2001</xref>; Jordan and Mitchell, <xref ref-type="bibr" rid="B126">2015</xref>; Blei and Smyth, <xref ref-type="bibr" rid="B18">2017</xref>). StLe methods may be more often adopted by computer scientists, physicists, engineers, and others who typically have less formal statistical background and may be more frequently working in industry rather than academia. In fact, John W. Tukey foresaw many of the developments that led up to what one might today call statistical learning (Tukey, <xref ref-type="bibr" rid="B196">1962</xref>). He early proposed a &#x0201C;peaceful collision of computing and statistics&#x0201D;. A modern reformulation of the same idea states (Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>): &#x0201C;If the inference/algorithm race is a tortoise-and-hare affair, then modern electronic computation has bred a bionic hare.&#x0201D; Indeed, kernel methods, decision trees, nearest-neighbor algorithms, graphical models, and various other statistical tools actually emerged in the ClSt community, but largely continued to develop in the StLe community (Friedman, <xref ref-type="bibr" rid="B79">2001</xref>).</p>
<p>As often cited beginnings of statistical learning approaches, the <italic>perceptron</italic> was an early brain-inspired computing algorithm (Rosenblatt, <xref ref-type="bibr" rid="B176">1958</xref>), and Arthur Samuel created a checker board program that succeeded in beating its own creator (Samuel, <xref ref-type="bibr" rid="B179">1959</xref>). Such studies toward <italic>artificial intelligence</italic> (AI) led to enthusiastic optimism and subsequent periodes of disappointment during the so-called &#x0201C;AI winters&#x0201D; in the late 70s and around the 90s (Russell and Norvig, <xref ref-type="bibr" rid="B178">2002</xref>; Kurzweil, <xref ref-type="bibr" rid="B137">2005</xref>; Cox and Dean, <xref ref-type="bibr" rid="B46">2014</xref>), while the increasingly available computers in the 80s encouraged a new wave of statistical algorithms (Efron and Tibshirani, <xref ref-type="bibr" rid="B61">1991</xref>). Later, the use of StLe methods increased steadily in many quantitative scientific domains as they underwent an increase in data richness from classical &#x0201C;long data&#x0201D; (samples <italic>n</italic> &#x0003E; variables <italic>p</italic>) to increasingly encountered &#x0201C;wide data&#x0201D; (<italic>n</italic> &#x0003C; &#x0003C; <italic>p</italic>) (Tibshirani, <xref ref-type="bibr" rid="B195">1996</xref>; Hastie et al., <xref ref-type="bibr" rid="B113">2015</xref>). The emerging field of StLe has received conceptual consolidation by the seminal book &#x0201C;The Elements of Statistical Learning&#x0201D; (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>). The coincidence of changing data properties, increasing computational power, and cheaper memory resources encouraged a still ongoing resurge in StLe research and applications approximately since 2000 (Manyika et al., <xref ref-type="bibr" rid="B145">2011</xref>; UK House of Common S.a.T, <xref ref-type="bibr" rid="B197">2016</xref>). For instance, over the last 15 years, <italic>sparsity</italic> assumptions gained increasing relevance for statistical and computational tractability as well as for domain interpretability when using <italic>supervised</italic> and <italic>unsupervised</italic> learning algorithms (i.e., with and without target variables) in the high-dimensional &#x0201C;<italic>n</italic> &#x0003C; &#x0003C; <italic>p</italic>&#x0201D; setting (B&#x000FC;hlmann and Van De Geer, <xref ref-type="bibr" rid="B25">2011</xref>; Hastie et al., <xref ref-type="bibr" rid="B113">2015</xref>). More recently, improvements in training very &#x0201C;deep&#x0201D; (i.e., many non-linear hidden layers) neural-networks architectures (Hinton and Salakhutdinov, <xref ref-type="bibr" rid="B120">2006</xref>) have much improved automatized feature selection (Bengio et al., <xref ref-type="bibr" rid="B13">2013</xref>) and have exceeded human-level performance in several application domains (LeCun et al., <xref ref-type="bibr" rid="B139">2015</xref>).</p>
<p>In sum, &#x0201C;the biggest difference between pre- and post-war statistical practice is the degree of automation&#x0201D; (Efron and Tibshirani, <xref ref-type="bibr" rid="B62">1994</xref>) up to a point where &#x0201C;almost all topics in twenty-first-century statistics are now computer-dependent&#x0201D; (Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>). ClSt has seen many important inventions in the first half of the twentieth century, which have often developed at statistical departments of academic institutions and remain in nearly unchanged form in current textbooks of psychology and other empirical sciences. The emergence of StLe as a coherent field has mostly taken place in the second half of the twentieth century as a number of disjoint developments in industry and often non-statistical departments in academia (e.g., AT&#x00026;T Bell Laboratories), which lead for instance to artificial neural networks, support vector machines, and boosting algorithms (Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>). Today, systematic education in StLe is still rare at the large majority of universities, in contrast to the many consistently offered ClSt courses (Cleveland, <xref ref-type="bibr" rid="B38">2001</xref>; Vanderplas, <xref ref-type="bibr" rid="B198">2013</xref>; Burnham and Anderson, <xref ref-type="bibr" rid="B26">2014</xref>; Donoho, <xref ref-type="bibr" rid="B57">2015</xref>).</p>
<p>In neuroscience, the advent of brain-imaging techniques, including positron emission tomography (PET) and functional magnetic resonance imaging (fMRI), allowed for the <italic>in-vivo</italic> characterization of the neural correlates underlying sensory, cognitive, or affective tasks. Brain scanning enabled <italic>quantitative</italic> brain measurements with <italic>many variables per observation</italic> (analogous to the advent of high-dimensional microarrays in genetics; Efron, <xref ref-type="bibr" rid="B59">2012</xref>). Since the inception of PET and fMRI, deriving topographical localization of neural activity changes was dominated by analysis approaches from ClSt, especially the general linear model (Scheff&#x000E9;, <xref ref-type="bibr" rid="B181">1959</xref>; Poline and Brett, <xref ref-type="bibr" rid="B173">2012</xref>; GLM). The classical approach to neuroimaging analysis is probably best exemplified by the statistical parametric mapping (SPM) software package that implements the GLM to provide a mass-univariate characterization of regionally specific effects.</p>
<p>As distributed information over voxels is less well captured by many ClSt approaches, including common GLM applications, StLe models were proposed early on for neuroimaging investigations. For instance, principal component analysis was used to distinguish globally distributed neural activity changes (Moeller et al., <xref ref-type="bibr" rid="B150">1987</xref>) as well as to study Alzheimer&#x00027;s disease (Grady et al., <xref ref-type="bibr" rid="B100">1990</xref>). Canonical correlation analysis was used to quantify complex relationships between task-free neural activity and schizophrenia symptoms (Friston et al., <xref ref-type="bibr" rid="B86">1992</xref>). However, these first approaches to &#x0201C;multivariate&#x0201D; brain-behavior associations did not ignite a major research trend (cf. Worsley et al., <xref ref-type="bibr" rid="B212">1997</xref>; Friston et al., <xref ref-type="bibr" rid="B84">2008</xref>). As a seminal contribution, Haxby and colleagues devised an innovative across-voxel correlation analysis to provide evidence against the widely assumed face-specificity of neural responses in the ventral temporal cortex (2001). This ClSt realization of one-nearest neighbor classification based on correlation distance foreshadowed several important developments, including (i) joint analysis of sets of brain locations to capture &#x0201C;distributed and overlapping representations&#x0201D;, (ii) repeated analysis in different splits of the data sample to compare against chance performance, and (iii) analysis across multiple stimulus categories to assess the specificity of neural responses. The finding of distributed face representation was confirmed in independent, similar data (Cox and Savoy, <xref ref-type="bibr" rid="B47">2003</xref>) and based on neural network algorithms (Hanson et al., <xref ref-type="bibr" rid="B111">2004</xref>).</p>
<p>The application of StLe methods in neuroimaging increased further after rebranding as &#x0201C;mind-reading,&#x0201D; &#x0201C;brain decoding,&#x0201D; and &#x0201C;MVPA&#x0201D; (Haynes and Rees, <xref ref-type="bibr" rid="B117">2005</xref>; Kamitani and Tong, <xref ref-type="bibr" rid="B128">2005</xref>). Note that &#x0201C;MVPA&#x0201D; initally referred to &#x0201C;multi<italic>voxel</italic> pattern analysis&#x0201D; (Kamitani and Tong, <xref ref-type="bibr" rid="B128">2005</xref>; Norman et al., <xref ref-type="bibr" rid="B160">2006</xref>) and later changed to &#x0201C;multi<italic>variate</italic> pattern analysis&#x0201D; (Haynes and Rees, <xref ref-type="bibr" rid="B117">2005</xref>; Hanke et al., <xref ref-type="bibr" rid="B109">2009</xref>; Haxby, <xref ref-type="bibr" rid="B114">2012</xref>). Up to that point, the term <italic>prediction</italic> had less often been used by imaging neuroscientists in the sense of out-of-sample generalization of a learning algorithm and more often in the incompatible sense of (in-sample) linear correlation such as using Pearson&#x00027;s or Spearman&#x00027;s method (Shmueli, <xref ref-type="bibr" rid="B185">2010</xref>; Gabrieli et al., <xref ref-type="bibr" rid="B88">2015</xref>). While there was scarce discussion of the position of &#x0201C;decoding&#x0201D; models in formal statistical terms, growing interest was manifested in first review publications and tutorial papers on applying StLe methods to neuroimaging data (Haynes and Rees, <xref ref-type="bibr" rid="B118">2006</xref>; Mur et al., <xref ref-type="bibr" rid="B151">2009</xref>; Pereira et al., <xref ref-type="bibr" rid="B166">2009</xref>). The interpretational gains of this new access to the neural representation of behavior and its disturbances in disease was flanked by the availability of necessary computing power and memory resources. Although challenging to realize, &#x0201C;deep&#x0201D; neural network algorithms have recently been introduced to neuroimaging research (Plis et al., <xref ref-type="bibr" rid="B169">2014</xref>; de Brebisson and Montana, <xref ref-type="bibr" rid="B52">2015</xref>; G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B102">2015</xref>). These computation-intensive models might help in approximating and deciphering the nature of neural processing in brain circuits (Cox and Dean, <xref ref-type="bibr" rid="B46">2014</xref>; Yamins and DiCarlo, <xref ref-type="bibr" rid="B214">2016</xref>). As the dimensionality and complexity of neuroimaging datasets are constantly increasing, neuroscientific investigations will be always more likely to benefit from StLe methods given their natural scaling to large-scale data analysis (Efron, <xref ref-type="bibr" rid="B59">2012</xref>; Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>; Blei and Smyth, <xref ref-type="bibr" rid="B18">2017</xref>).</p>
<p>From a conceptual viewpoint (Figure <xref ref-type="fig" rid="F3">3</xref>), a large majority of statistical methods can be situated somewhere on a continuum between the two poles of ClSt and StLe (Committee on the Analysis of Massive Data et al., <xref ref-type="bibr" rid="B44">2013</xref>; Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>; p. 61). ClSt was mostly fashioned for problems with small samples that can be grasped by plausible models with a small number of parameters chosen by the investigator in an analytical fashion. StLe was mostly fashioned for problems with many variables in potentially large samples with little knowledge of the data-generating process that gets emulated by a mathematical function derived from data in a heuristic fashion. Tools from ClSt therefore typically assume that the data behave according to certain known mechanisms, whereas StLe exploits algorithmic techniques to avoid many a-priori specifications of data-generating mechanisms. Neither ClSt or StLe nor any of the other categories of statistical models can be considered generally superior. This relativism is captured by the so-called <italic>no free lunch theorem</italic><xref ref-type="fn" rid="fn0004"><sup>4</sup></xref> (Wolpert, <xref ref-type="bibr" rid="B210">1996</xref>): no single statistical strategy can consistently do better in all circumstances (cf. Gigerenzer, <xref ref-type="bibr" rid="B93">2004</xref>). As a very general rule of thumb, ClSt preassumes and formally tests <italic>a model for the data</italic>, whereas StLe extracts and empirically evaluates <italic>a model from the data</italic>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Key differences in the modeling philosophy of classical statistics and statistical learning. Ten modeling intuitions that tend to be relatively more characteristic for classical statistical methods (blue) or pattern-learning methods (red). In comparison to ClSt, StLe &#x0201C;is essentially a form of applied statistics with increased emphasis on the use of computers to statistically estimate complicated functions and a decreased emphasis on proving confidence intervals around these functions&#x0201D; (Goodfellow et al., <xref ref-type="bibr" rid="B98">2016</xref>). Broadly, ClSt tends to be more analytical by imposing mathematical rigor on the phenomenon, whereas StLe tends to be more heuristic by finding useful approximations. In practice, ClSt is probably more often applied to experimental data, where a set of target variables are systematically controlled by the investigator and the brain system under studied has been subject to experimental perturbation. Instead, StLe is probably more often applied to observational data without such structured influence and where the studied system has been left unperturbed. ClSt fully specifies the statistical model at the beginning of the investigation, whereas in StLe there is a bigger emphasis on models that can flexibly adapt to the data (e.g., learning algorithms creating decision trees).</p></caption>
<graphic xlink:href="fnins-11-00543-g0003.tif"/>
</fig>
</sec>
<sec id="s3">
<title>Case study one: cognitive contrast analysis and decoding mental states</title>
<p>Vignette: A neuroimaging investigator wants to reveal the neural correlates underlying face processing in humans. 40 healthy, right-handed adults are recruited and undergo a block design experiment run in a 3T MRI scanner with whole-brain coverage. In a passive viewing paradigm, 60 colored stimuli of unfamiliar faces are presented, which have forward head and gaze position. The control condition presents colored pictures of 60 different houses to the participants. In the experimental paradigm, a picture of a face or a house is presented for 2 s in each trial and the interval between trials within each block is randomly jittered varying from 2 to 7 s. The picture stimuli are presented in pseudo-randomized fashion and are counterbalanced in each passively watching participant. Despite the blocked presentation of stimuli, each experiment trial is modeled separately. The fMRI data are analyzed using a GLM as implemented in the SPM software package. Two task regressors are included in the model for the face and house conditions based on the stimulus onsets and viewing durations and using a canonical hemodynamic response function. In the GLM design matrix, the face column and house column are hence set to 1 for brain scans from the corresponding task condition and set to 0 otherwise. Separately in each brain voxel, the GLM parameters are estimated, which fits beta<sub>face</sub> and beta<sub>house</sub> regression coefficients to explain the contribution of each experimental task to the neural activity increases and decreases observed in that voxel. A <italic>t</italic>-test can then formally assess whether the fMRI signal in the current voxel is significantly more involved in viewing faces as opposed to the house control condition.</p>
<p>Question: What is the statistical difference between <italic>subtracting</italic> the neural activity from the face vs. house conditions and <italic>decoding</italic> the neural activity during face vs. house processing?</p>
<p>Computing cognitive contrasts is a ClSt approach that was and still is routinely performed in the <italic>mass-univariate regime:</italic> it fits a separate GLM model for each individual voxel in the brain scans and then tests for significant differences between the obtained condition coefficients (Friston et al., <xref ref-type="bibr" rid="B85">1994</xref>). Instead, decoding cognitive processes from neural activity is a StLe approach that is typically performed in a <italic>multivariate regime</italic>: a learning algorithm is trained on a large number of voxel observations in brain scans and then the model&#x00027;s prediction accuracy is evaluated on sets of new brain scans. These ClSt and StLe approaches to identifying the neural correlates underlying cognitive processes of interest are closely related to the notions of <italic>encoding models</italic> and <italic>decoding models</italic>, respectively (Kriegeskorte, <xref ref-type="bibr" rid="B133">2011</xref>; Naselaris et al., <xref ref-type="bibr" rid="B153">2011</xref>; Pedregosa et al., <xref ref-type="bibr" rid="B164">2015</xref>; but see G&#x000FC;&#x000E7;l&#x000FC; and van Gerven, <xref ref-type="bibr" rid="B102">2015</xref>).</p>
<p>Encoding models regress the brain data against a design matrix with indicators of the face vs. house condition and formally test whether the difference is statistically significant. Decoding models typically aim to predict these indicators by training and empirically evaluating classification algorithms on different splits from the whole dataset. In ClSt parlance, the model <italic>explains</italic> the neural activity, the <italic>dependent or explained variable</italic>, measured in each separate brain voxel, by the <italic>beta coefficients</italic> according to the experimental condition indicators in the <italic>design matrix</italic> columns, the <italic>independent or explanatory variables</italic>. That is, the GLM can be used to explain neural activity changes by a linear combination of experimental variables (Naselaris et al., <xref ref-type="bibr" rid="B153">2011</xref>). Answering the same neuroscientific question with decoding models in StLe jargon, the <italic>model weights</italic> of a <italic>classifier</italic> are fitted on the <italic>training set</italic> of the <italic>input data</italic> to <italic>predict</italic> the <italic>class labels</italic>, the <italic>target variables</italic>, and are subsequently evaluated on the <italic>test set</italic> by <italic>cross-validation</italic> to obtain their <italic>out-of-sample generalization performance</italic>. Here, classification algorithms are used to predict entries of the design matrix by identifying a linear or more complicated combination between the many simultaneously considered brain voxels (Pereira et al., <xref ref-type="bibr" rid="B166">2009</xref>). More broadly, ClSt applications in functional neuroimaging tend to estimate the location of cognitive processes from neural activity, whereas many StLe applications estimate properties of neural activity underlying different cognitive tasks.</p>
<p>A key difference between many ClSt-mediated encoding models and StLe-mediated decoding models thus pertains to the direction of statistical estimation between brain space and behavior space (Friston et al., <xref ref-type="bibr" rid="B84">2008</xref>; Varoquaux and Thirion, <xref ref-type="bibr" rid="B204">2014</xref>). It was noted (Friston et al., <xref ref-type="bibr" rid="B84">2008</xref>) that the direction of brain-behavior association is related to the question whether the stimulus indicators in the model act as causes by representing deterministic experimental variables of an encoding model or consequences by representing probabilistic outputs of a decoding model. Such considerations also reveal the intimate relationship of ClSt models to the notion of <italic>forward inference</italic>, while StLe methods are probably more often used for formal <italic>reverse inference</italic> in functional neuroimaging (Poldrack, <xref ref-type="bibr" rid="B170">2006</xref>; Eickhoff et al., <xref ref-type="bibr" rid="B63">2011</xref>; Yarkoni et al., <xref ref-type="bibr" rid="B216">2011</xref>; Varoquaux and Thirion, <xref ref-type="bibr" rid="B204">2014</xref>). On the one hand, <italic>forward inference</italic> relates to encoding models by testing the probability of observing activity in a brain location given knowledge of a psychological process. On the other hand, <italic>reverse inference</italic> relates to brain decoding to the extent that classification algorithms can learn to distinguish experimental fMRI data to belong to two psychological conditions and subsequently be used to estimate the presence of specific cognitive processes based on new neural activity observations (cf. Poldrack, <xref ref-type="bibr" rid="B170">2006</xref>). Finally, establishing a brain-behavior association has been argued to be more important than the actual direction of the mapping function (Friston, <xref ref-type="bibr" rid="B82">2009</xref>). This author stated that &#x0201C;showing that one can decode activity in the visual cortex to classify [&#x02026;] a subject&#x00027;s percept is exactly the same as demonstrating significant visual cortex responses to perceptual changes&#x0201D; and, conversely, &#x0201C;all demonstrations of functionally specialized responses represent an implicit mindreading.&#x0201D;</p>
<p>Conceptually, GLM-based encoding models follow a <italic>localization agenda</italic> by testing hypotheses on <italic>regional effects of functional specialization</italic> in the brain (where?). A <italic>t</italic>-test is used to compare pairs of neural activity estimates to statistically distinguish the target face and the non-target house condition (Friston et al., <xref ref-type="bibr" rid="B87">1996</xref>). Essentially, this test for significant differences between the fitted beta coefficients corresponds to two stimulus indicators based on well-founded arguments from cognitive theory. This statistical approach assumes that <italic>cognitive subtraction</italic> is possible, that is, the regional brain responses of interest can be isolated by contrasting two sets of brain scans that are believed to differ in the cognitive facet of interest (Friston et al., <xref ref-type="bibr" rid="B87">1996</xref>; Stark and Squire, <xref ref-type="bibr" rid="B190">2001</xref>). For one voxel location at a time, an attempt is made to reject the null hypothesis of no difference between the averaged <italic>neural activity level</italic> of a target brain state and the averaged neural activity of a control brain state. It is important to appreciate that the localization agenda thus emphasizes the <italic>relative difference</italic> in fMRI signal during tasks and may neglect the individual neural activity information of each particular task (Logothetis et al., <xref ref-type="bibr" rid="B144">2001</xref>). Note that the univariate GLM analysis can be extended to more than one output (dependent or explained) variable within the ClSt regime by performing a multivariate analysis of covariance (MANCOVA). This allows for tests of more complex hypotheses but incurs multivariate normality assumptions (Kriegeskorte, <xref ref-type="bibr" rid="B133">2011</xref>).</p>
<p>More generally, it is seldom mentioned that the standard GLM would not have been solvable for unique solutions in the high-dimensional &#x0201C;<italic>n</italic> &#x0003C; &#x0003C; <italic>p</italic>&#x0201D; regime, instead of fitting one model for each voxel in the brain scans. This is because the number of brain voxels <italic>p</italic> exceed by far the number of data samples n (i.e., leading to an under-determined system of equations), which incapacitates many statistical estimators from ClSt (cf. Giraud, <xref ref-type="bibr" rid="B95">2014</xref>; Hastie et al., <xref ref-type="bibr" rid="B113">2015</xref>). Regularization by sparsity-inducing norms, such as in modern <italic>penalized</italic> regression analysis using the LASSO and ElasticNet, emerged only later (Tibshirani, <xref ref-type="bibr" rid="B195">1996</xref>; Zou and Hastie, <xref ref-type="bibr" rid="B221">2005</xref>) as a principled StLe strategy to de-escalate the need for dimensionality reduction or preliminary filtering of important voxels and to enable the tractability of the high-dimensional analysis setting.</p>
<p>Because hypothesis testing for significant differences between beta coefficients of fitted GLMs relies on comparing the means of neural activity measurements, the results from statistical tests are not corrupted by the conventionally applied spatial smoothing with a Gaussian filter. On the contrary, this image preprocessing step even helps the correction for multiple comparisons based on random fields theory (cf. below), alleviates inter-individual neuroanatomical variability, and can thus increases sensitivity. Spatial smoothing however discards fine-grained neural activity patterns spatially distributed across voxels that potentially carry information associated with mental operations (cf. Kamitani and Sawahata, <xref ref-type="bibr" rid="B127">2010</xref>; Haynes, <xref ref-type="bibr" rid="B116">2015</xref>). Indeed, some authors believe that sensory, cognitive, and motor processes manifest themselves as &#x0201C;neuronal population codes&#x0201D; (Averbeck et al., <xref ref-type="bibr" rid="B7">2006</xref>). Relevance of such population codes in human neuroimaging was for instance suggested by revealing subject-specific neural responses in the fusiform gyrus to facial stimuli (Saygin et al., <xref ref-type="bibr" rid="B180">2012</xref>). In applications of StLe models, the spatial smoothing step is therefore often skipped because the &#x0201C;decoding&#x0201D; algorithms precisely exploit the locally varying structure of the salt-and-pepper patterns in fMRI signals.</p>
<p>In so doing, decoding models use learning algorithms in an <italic>information agenda</italic> by showing <italic>generalization of robust patterns</italic> to new brain activity acquisitions (Kriegeskorte et al., <xref ref-type="bibr" rid="B134">2006</xref>; Mur et al., <xref ref-type="bibr" rid="B151">2009</xref>; de-Wit et al., <xref ref-type="bibr" rid="B55">2016</xref>). Information that is weak in one voxel but spatially distributed across voxels can be effectively harvested in a structure-preserving fashion (Haynes and Rees, <xref ref-type="bibr" rid="B118">2006</xref>; Haynes, <xref ref-type="bibr" rid="B116">2015</xref>). This modeling agenda is focused on the whole <italic>neural activity pattern</italic>, in contrast to the localization agenda dedicated to separate increases or decreases in <italic>neural activity level</italic>. For instance, the default mode network typically exhibits activity <italic>decreases</italic> at the onset of many psychological tasks with visual or other sensory stimuli, whereas the induced activity <italic>patterns</italic> in that less activated network may nevertheless functionally subserve task execution (Bzdok et al., <xref ref-type="bibr" rid="B29">2016</xref>; Christoff et al., <xref ref-type="bibr" rid="B36">2016</xref>). Some brain-behavior associations might only emerge when simultaneously capturing neural activity in a group of voxels but disappear in single-voxel approaches, such as mass-univariate GLM analyses (cf. Davatzikos, <xref ref-type="bibr" rid="B50">2004</xref>). Note that, analogous to multivariate variants of the GLM, decoding could also be replaced by classical statistical approaches (cf. Haxby et al., <xref ref-type="bibr" rid="B115">2001</xref>; Brodersen et al., <xref ref-type="bibr" rid="B23">2011a</xref>). For many linear classification algorithm trained to predict face vs. house stimuli based on many brain voxels, model fitting typically searches iteratively through the <italic>hypothesis space</italic> (&#x0003D; <italic>function space</italic>) of the chosen learning model. In our case, the final hypothesis selected by the linear classifier commonly corresponds to one specific combination of model weights (i.e., a weighted contribution of individual brain measurements) that equates with one mapping function from the neural activity features to the face vs. house target variable.</p>
<p>Among other views, it has previously been proposed (Brodersen, <xref ref-type="bibr" rid="B21">2009</xref>) that four types of neuroscientific questions become readily quantifiable through StLe applications to neuroimaging: (i) <italic>Where</italic> is an information category neurally processed? This can extend the interpretational spectrum from increase and decrease of neural activity to the existence of complex combinations of activity variations distributed across voxels. For instance, across-voxel linear correlation could decode object categories from the ventral temporal cortex even after excluding the fusiform gyrus, which is known to be responsive to object stimuli (Haxby et al., <xref ref-type="bibr" rid="B115">2001</xref>). (ii) <italic>Whether</italic> a given information category is reflected by neural activity? This can extend the interpretational spectrum to topographically similar but neurally distinct processes that potentially underlie different cognitive facets. For instance, linear classifiers could successfully distinguish whether a subject is attending to the first or second of two simultaneously presented stimuli (Kamitani and Tong, <xref ref-type="bibr" rid="B128">2005</xref>). (iii) <italic>When</italic> is an information category generated (i.e., onset), processed (i.e., duration), and bound (i.e., alteration)? When applying classifiers to neural time series, the interpretational spectrum can be extended to the beginning, evolution, and end of distinct cognitive facets. For instance, different classifiers have been demonstrated to map the decodability time structure of mental operation sequences (King and Dehaene, <xref ref-type="bibr" rid="B131">2014</xref>). (iv) More controversially, <italic>how</italic> is an information category neurally processed? The interpretational spectrum can be extended to computational properties of the neural processes, including processing in brain regions vs. brain networks or isolated vs. partially shared processing facets. For instance, a classifier trained for evolutionarily conserved eye gaze movement was able to decode evolutionarily more recent mathematical calculation processes as a possible case of &#x0201C;neural recycling&#x0201D; in the human brain (Knops et al., <xref ref-type="bibr" rid="B132">2009</xref>; Anderson, <xref ref-type="bibr" rid="B5">2010</xref>). As an important caveat in interpreting StLe models, the particular technical properties of a chosen learning algorithm (e.g., linear vs. non-linear support vector machines) can probably seldom serve as a convincing argument for reverse-engineering mechanisms of neural information processing as measured by fMRI scanning (cf. Misaki et al., <xref ref-type="bibr" rid="B149">2010</xref>).</p>
<p>In sum, the statistical properties of ClSt and StLe methods have characteristic consequences in neuroimaging analysis and interpretation. They can hence offer different access routes and complementary answers to identical neuroscientific questions.</p>
</sec>
<sec id="s4">
<title>Case study two: small volume correction and searchlight analysis</title>
<p>Vignette: The neuroimaging experiment from case study 1 successfully identified the fusiform gyrus of the ventral visual stream to be more responsive to face stimuli than house stimuli. However, the investigator&#x00027;s initial hypothesis of also observing face-responsive neural activity in the ventromedial prefrontal cortex could not be confirmed in the <italic>whole-brain</italic> analyses. The investigator therefore wants to follow up with a <italic>topographically focused</italic> approach that examines differences in neural activity between the face and house conditions exclusively in the ventromedial prefrontal cortex.</p>
<p>Question: What are the statistical implications of delineating task-relevant neural responses in a spatially constrained search space rather than analyzing brain measurements of the entire brain?</p>
<p>A popular ClSt approach to corroborate less pronounced neural activity findings is <italic>small volume correction</italic>. This region of interest (ROI) analysis involves application of the mass-univariate GLM approach only to the ventromedial prefrontal cortex as a preselected biological compartment, rather than considering the gray-matter voxels of the entire brain in a na&#x000EF;ve, topographically unconstrained fashion. Small volume correction allows for significant findings in the ROI that remain sub-threshold after accounting for the tens of thousands of multiple comparisons in the whole-brain GLM analysis. Small volume correction is therefore a simple means to alleviate the multiple-comparisons problem that motivated more than two decades of still ongoing methodological developments in the neuroimaging domain (Worsley et al., <xref ref-type="bibr" rid="B211">1992</xref>; Smith et al., <xref ref-type="bibr" rid="B188">2001</xref>; Friston, <xref ref-type="bibr" rid="B81">2006</xref>; Nichols, <xref ref-type="bibr" rid="B155">2012</xref>). Whole-brain GLM results were initially reported as uncorrected findings without accounting for multiple comparisons, then with Bonferroni&#x00027;s family wise error (FWE) correction, later by random field theory correction using neural activity height (or clusters), followed by false discovery rate (FDR) (Genovese et al., <xref ref-type="bibr" rid="B89">2002</xref>) and slowly increasing adoption of cluster-thresholding for voxel-level inference via permutation testing (Smith and Nichols, <xref ref-type="bibr" rid="B189">2009</xref>). Rather than the isolated voxel, it has early been discussed that a possibly better unit of interest should be spatially neighboring voxel groups (see here for an overview: Chumbley and Friston, <xref ref-type="bibr" rid="B37">2009</xref>). The setting of high regional correlation of neural activity was successfully addressed by random field theory that provide inferences not about individual voxels but topological features in the underlying (spatially continuous) effects. This topological inference is used to identify clusters of relevant neural activity changes from their peak, size, or mass (Worsley et al., <xref ref-type="bibr" rid="B211">1992</xref>). Importantly, the spatial dependencies of voxel observations were not incorporated into the GLM estimation step, but instead taken into account during the subsequent model inference step to alleviate the multiple-comparisons problem.</p>
<p>A related cousin of small volume correction in the StLe world would be to apply classification algorithms to a subset of voxels to be considered as input to the model (i.e., <italic>feature selection</italic>). In particular, <italic>searchlight analysis</italic> is an increasingly popular learning technique that can identify <italic>locally constrained multivariate patterns</italic> in neural activity (Friman et al., <xref ref-type="bibr" rid="B80">2001</xref>; Kriegeskorte et al., <xref ref-type="bibr" rid="B134">2006</xref>). For each voxel in the ventromedial prefrontal cortex, the brain measurements of the immediate neighborhood are first collected (e.g., radius of 10 mm voxels). In each such searchlight, a classification algorithm, for instance linear support vector machines, is then trained on one part of the brain scans (<italic>training set</italic>) and subsequently applied to determine the prediction accuracy in the remaining, unseen brain scans (<italic>test set</italic>). In this StLe approach, the excess of brain voxels is handled by performing pattern recognition analysis in only dozens of locally adjacent voxel neighborhoods at a time. Finally, the mean classification accuracy of face vs. house stimuli across all permutations over the brain data is mapped to the center of each considered sphere. The searchlight is then moved through the ROI until each seed voxel had once been the center voxel of the searchlight. This yields a voxel-wise classification map of accuracy estimates for the entire ventromedial prefrontal cortex. Consistent with the information agenda (cf. above), searchlight analysis quantifies the extent to which (local) neural activity <italic>patterns</italic> can <italic>predict</italic> the difference between the house and face conditions. It contrasts small volume correction that determines whether one experimental condition exhibited a significant neural activity <italic>increase</italic> or <italic>decrease</italic> relative to a particular other experimental condition, consistent with the localization agenda. Further, searchlight analysis alleviates the burden of abundant input variables by fitting learning algorithms restricted to the voxels in small sphere neighborhoods. However, the searchlight procedure thus yields many prediction performances for many brain locations, which motivates correction for multiple comparisons across the considered neighborhoods.</p>
<p>When considering high-dimensional brain scans through the ClSt lens, the statistical challenge resides in solving the <italic>multiple-comparisons problem</italic> (Nichols and Hayasaka, <xref ref-type="bibr" rid="B156">2003</xref>; Nichols, <xref ref-type="bibr" rid="B155">2012</xref>). From the StLe stance, however, it is the <italic>curse of dimensionality</italic> and <italic>overfitting</italic> that statistical analyses need to tackle (Friston et al., <xref ref-type="bibr" rid="B84">2008</xref>; Domingos, <xref ref-type="bibr" rid="B56">2012</xref>). Many neuroimaging analyses based on ClSt methods can be viewed as testing a particular hypothesis (i.e., the null hypothesis) repeatedly in a large number of separate voxels. In contrast, testing whether learning algorithm extrapolate to new brain data can be viewed as searching through thousands of different hypotheses in a single process (i.e., walking through the hypothesis space; cf. above) (Shalev-Shwartz and Ben-David, <xref ref-type="bibr" rid="B184">2014</xref>).</p>
<p>As common brain scans offer measurements of &#x0003E;100,000 brain locations, a mass-univariate GLM analysis typically entails the same statistical test to be applied &#x0003E;100,000 times. The more often the investigator tests a hypothesis of relevance for a brain location, the more locations will be falsely detected as relevant (false positive, Type I error), especially in the noisy neuroimaging data. All dimensions in the brain data (i.e., voxel variables) are implicitly treated as equally important and no neighborhoods of most expected variation are statistically exploited (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>). Hence, the absence of restrictions on observable structure in the set of data variables during the statistical modeling of neuroimaging data takes a heavy toll at the final inference step. This is where <italic>random field theory</italic> comes to the rescue. As noted above, this form of topological inference dispenses with the problem of inferring which voxels are significant and tries to identify significant topological features in the underlying distributed responses. By definition, topological features like maxima are sparse events and can be thought of as a form of dimensionality reduction&#x02014;not in data space but in the statistical characterization of where neural responses occur.</p>
<p>This is contrasted by the high-dimensional StLe regime, where the initial model family chosen by the investigator determines the complexity restrictions to all data dimensions (i.e., all voxels, not single voxels) that are imposed explicitly or implicitly by the model structure. Model choice predisposes existing but unknown low-dimensional neighborhoods in the full voxel space to achieve the prediction task. Here, the toll is taken at the beginning of the investigation because there are so many different alternative model choices that would impose a different set of complexity constraints to the high-dimensional measurements in the brain. For instance, signals from &#x0201C;brain regions&#x0201D; are likely to be well approximated by models that impose discrete, locally constant compartments on the data (e.g., <italic>k</italic>-means or spatially constrained Ward clustering). Instead, tuning model choice to signals from macroscopical &#x0201C;brain networks&#x0201D; should impose overlapping, locally continuous data compartments (e.g., independent component analysis or sparse principal component analysis) (Yeo et al., <xref ref-type="bibr" rid="B218">2014</xref>; Bzdok and Yeo, <xref ref-type="bibr" rid="B30">2017</xref>; Bzdok et al., <xref ref-type="bibr" rid="B28">2017</xref>).</p>
<p>Exploiting such <italic>effective dimensions</italic> in the neuroimaging data (i.e., coherent brain-behavior associations involving many distributed brain voxels) is a rare opportunity to simultaneously reduce the <italic>model bias</italic> and <italic>model variance</italic>, despite their typical inverse relationship (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>). Model bias relates to prediction failures incurred because the learning algorithm can systematically not represent certain parts of the underlying relationship between brain scans and experimental conditions (formally, the deviation between the target function and the average function space of the model). Model variance relates to prediction failures incurred by noise in the estimation of the optimal brain-behavior association (formally, the difference between the best-choice input-output relation and the average function space of the model). A model that is too simple to capture a brain-behavior association probably underfits due to high bias. Yet, an overly complex model probably overfits due to high variance. Generally, high-variance approaches are better at <italic>approximating</italic> the &#x0201C;true&#x0201D; brain-behavior relation (i.e., in-sample model estimation), while high-bias approaches have a higher chance of <italic>generalizing</italic> the identified pattern to new observations (i.e., out-of-sample model evaluation). The bias-variance tradeoff can be useful in explaining why applications of statistical models intimately depend on (i) the amount of available data, (ii) the typically not known amount of noise in the data, and (iii) the unknown complexity of the target function in nature (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>).</p>
<p>Learning algorithms that overcome the curse of dimensionality&#x02014;extracting coherent patterns from all considered brain voxels at once&#x02014;typically incorporate an implicit bias for anisotropic neighborhoods in the data (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>; Bach, <xref ref-type="bibr" rid="B8">2014</xref>; Bzdok et al., <xref ref-type="bibr" rid="B27">2015</xref>). Put differently, prediction models successful in the high-dimensional setting have an in-built specialization to representing types of functions that are compatible with the structure to be uncovered in the brain data. Knowledge embodied in a learning algorithm suited to a particular application domain can better calibrate the sweet spot between underfitting and overfitting. When applying a model without any complexity restrictions to high-dimensional data generalization becomes difficult to impossible because all directions in the data (i.e., individual brain voxels) are treated equally with isotropic structure. At the root of the problem, all data samples look virtually identical to the learning algorithm in high-dimensional data scenarios (Bellman, <xref ref-type="bibr" rid="B11">1961</xref>). The learning algorithm will struggle to see through the idiosyncracies in the data, will tend to overfit, and thus be unlikely to generalize to new observations. Such considerations provide insight into why the multiple-comparisons problem is more often an issue in encoding studies, while overfitting is more closely related to decoding studies (Friston et al., <xref ref-type="bibr" rid="B84">2008</xref>). The juxtaposition of ClSt and StLe views offers insights into why restricting neural data analysis to an ROI with fewer voxels, rather than the whole brain, simultaneously alleviates both the multiple-comparisons problem (ClSt) and the curse of dimensionality (StLe).</p>
<p>As an practical summary, drawing classical inference in neuroimaging data has largely been performed by considering each voxel independently and by massive simultaneous testing of a same null hypothesis in all observed voxels. This has incurred a multiple-comparisons problem difficult enough that common approaches may still be prone to incorrect results (Efron, <xref ref-type="bibr" rid="B59">2012</xref>). In contrast, aiming for generalization of a pattern in high-dimensional neuroimaging data to new observations in the brain incurs the equally challenging curse of dimensionality. Successfully accounting for the high number of input dimensions will probably depend on learning models that impose neurobiologically justified bias and keeping the variance under control by dimensionality reduction and regularization techniques.</p>
<p>More broadly, asking at what point new neurobiological knowledge is arising during ClSt and StLe investigations relies on largely distinct theoretical frameworks that revolve around <italic>null-hypothesis testing</italic> and <italic>statistical learning theory</italic> (Figure <xref ref-type="fig" rid="F4">4</xref>). Both ClSt and StLe methods share the common goal of demonstrating relevance of a given effect in the data beyond the sample brain scans at hand. However, the attempt to show successful extrapolation of a statistical relationship at the general population is embedded in different mathematical contexts. Knowledge generation in ClSt and StLe is hence rooted in different notions of statistical inference.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Key concepts in classical statistics and statistical learning. Schematic with statistical notions that are relatively more associated with classical statistical methods (left column) or pattern-learning methods (right column). As there is a smooth transition between the classical statistical toolkit and learning algorithms, some notions may be closely associated with both statistical cultures (middle column).</p></caption>
<graphic xlink:href="fnins-11-00543-g0004.tif"/>
</fig>
<p>ClSt laid down its most important inferential framework in the Popperian spirit of critical empiricism (Popper, <xref ref-type="bibr" rid="B174">1935/2005</xref>): scientific progress is to be made by continuous replacement of current hypotheses by ever more pertinent hypotheses using <italic>falsification</italic>. The rationale behind hypothesis falsification is that one counterexample can reject a theory by <italic>deductive reasoning</italic>, while any quantity of evidence can not confirm a given theory by inductive reasoning (Goodman, <xref ref-type="bibr" rid="B99">1999</xref>). The investigator verbalizes two mutually exclusive hypotheses by domain-informed judgment. The <italic>alternative hypothesis</italic> should be conceived as the outcome intended by the investigator and to contradict the state of the art of the research topic. The <italic>null hypothesis</italic> represents the devil&#x00027;s advocate argument that the investigator wants to reject (i.e., falsify) and it should automatically deduce from the newly articulated alternative hypothesis. A conventional 5%-threshold (i.e., equating with roughly two standard deviations) guards against rejection due to the idiosyncrasies of the sample that are not representative of the general population. If the data have a probability of &#x02264;5% given the null hypothesis [P(result|H<sub>0</sub>)], it is evaluated to be significant. Such a <italic>test for statistical significance</italic> indicates a difference between two means with a 5% chance of being a false positive finding. If the null hypothesis can not be rejected (which depends on power), then the test yields no conclusive result, rather than a null result (Schmidt, <xref ref-type="bibr" rid="B182">1996</xref>). In this way, classical hypothesis testing continuously replaces currently embraced hypotheses explaining a phenomenon in nature by better hypotheses with more empirical support in a Darwinian selection process. Finally, Fisher, Neyman, and Pearson intended hypothesis testing as a marker for further investigation, rather than an off-the-shelf decision-making instrument (Cohen, <xref ref-type="bibr" rid="B43">1994</xref>; Nuzzo, <xref ref-type="bibr" rid="B161">2014</xref>).</p>
<p>In StLe instead, answers to how neurobiological conclusions can be drawn from a dataset at hand are provided by the <italic>Vapnik-Chervonenkis dimensions</italic> (VC dimensions) from <italic>statistical learning theory</italic> (Vapnik, <xref ref-type="bibr" rid="B201">1989</xref>, <xref ref-type="bibr" rid="B202">1996</xref>). The VC dimensions of a pattern-learning algorithm quantify the probability at which the distinction between the neural correlates underlying the face vs. house conditions can be captured and used for correct predictions in new, possibly later acquired brain scans from the same cognitive experiment (i.e., <italic>out-of-sample generalization</italic>). Such statistical approaches implement the <italic>inductive</italic> strategy to learn general principles (i.e., the neural signature associated with given cognitive processes) from a series of exemplary brain measurements, which contrasts the <italic>deductive</italic> strategy of rejecting a certain null hypothesis based on counterexamples (cf. Tenenbaum et al., <xref ref-type="bibr" rid="B193">2011</xref>; Bengio, <xref ref-type="bibr" rid="B12">2014</xref>; Lake et al., <xref ref-type="bibr" rid="B138">2015</xref>). The VC dimensions measure how complicated the examined relationship between brain scans and experimental conditions could become&#x02014;in other words, the richness of the representation which can be instantiated by the used model, the complexity capacity of its <italic>hypothesis space</italic>, the &#x0201C;wiggliness&#x0201D; of the decision boundary used to distinguish examples from several classes, or, more intuitively, the &#x0201C;currency&#x0201D; of learnability. VC dimensions are derived from the maximal number of different brain scans that can be correctly detected to belong to either the house condition or the face condition by a given model. The VC dimensions thus provide a theoretical guideline for the largest set of brain scan examples fed into a learning algorithm such that this model is able to guarantee zero classification errors.</p>
<p>As one of the most important results from statistical learning theory, in any intelligent learning system, the opportunity to derive abstract patterns in the world by reducing the discrepancy between prediction error from training data (in-sample estimate) and prediction error from independent test data (out-of-sample estimate) decreases with the higher model capacity and increases with the number of available training observations (Vapnik and Kotz, <xref ref-type="bibr" rid="B203">1982</xref>; Vapnik, <xref ref-type="bibr" rid="B202">1996</xref>). In brain imaging, a learning algorithm is hence theoretically backed up to successfully predict outcomes in future brain scans with high probability if the choosen model ignores structure that is overly complicated, such as higher-order non-linearities between many brain voxels, and if the model is provided with a sufficient number of training brain scans. Hence, VC dimensions provide explanations why increasing the number of considered brain voxels as input features (i.e., entailing increased number of model parameters) or using a more sophisticated prediction model, requires more training data for successful generalization. Notably, the VC dimensions (analogous to null-hypothesis testing) are unrelated to the <italic>target function</italic>, as the &#x0201C;true&#x0201D; mechanisms underlying the studied phenomenon in nature. Nevertheless, the VC dimensions provide justification that a certain learning model can be used to approximate that target function by fitting a model to a collection of input-output pairs. In short, VC dimensions is among the best frameworks to derive theoretical errors bounds for predictive models (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>).</p>
<p>Further, some common invalidations of the ClSt and StLe statistical concern in neuroimaging studies performing classical inference is <italic>double dipping</italic> or <italic>circular analysis</italic> (Kriegeskorte et al., <xref ref-type="bibr" rid="B136">2009</xref>). This occurs when, for instance, first correlating a behavioral measure with brain activity and then using the identified subset of brain voxels for a second correlation analysis with that same behavioral measurement (Lieberman et al., <xref ref-type="bibr" rid="B141">2009</xref>; Vul et al., <xref ref-type="bibr" rid="B206">2009</xref>). In this scenario, voxels are submitted to two statistical tests with the same goal in a nested, non-independent fashion<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> (Freedman, <xref ref-type="bibr" rid="B77">1983</xref>). This corrupts the <italic>validity of the null hypothesis</italic> on which the reported test results conditionally depend. Importantly, this case of repeating a same statistical estimation with iteratively pruned data selections (on the training data split) is a valid routine in the StLe framework, such as in recursive feature extraction (Guyon et al., <xref ref-type="bibr" rid="B104">2002</xref>; Hanson and Halchenko, <xref ref-type="bibr" rid="B110">2008</xref>). However, double-dipping or circular analysis in ClSt applications to neuroimaging data have an analog in StLe analyses aiming at out-of-sample generalization: <italic>data-snooping</italic> or <italic>peeking</italic> (Pereira et al., <xref ref-type="bibr" rid="B166">2009</xref>; Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>; Fithian et al., <xref ref-type="bibr" rid="B73">2014</xref>). This can occur, for instance, when performing simple (e.g., mean-centering) or more involved (e.g., <italic>k</italic>-means clustering) target-variable-dependent or -independent preprocessing on the entire dataset if it should be applied separately to the training sets and test sets. Data-snooping can lead to overly optimistic cross-validation estimates and a trained learning algorithm that fails on fresh data drawn from the same distribution (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>). Rather than a corrupted null hypothesis, it is the <italic>error bounds of the VC dimensions that are loosened</italic> and, ultimately, invalidated because information from the concealed test set influences model selection on the training set.</p>
<p>In sum, statistical inference in ClSt is drawn by using the <italic>entire data</italic> at hand to <italic>formally test</italic> for <italic>theoretically guaranteed</italic> extrapolation of an effect to the general population. In stark contrast, inferential conclusions in StLe are typically drawn by fitting a model on a <italic>larger part of the data</italic> at hand (i.e., in-sample model selection) and <italic>empirically testing</italic> for successful extrapolation to an independent, smaller part of the data (i.e., out-of-sample model evaluation). As such, ClSt has a focus on <italic>in-sample estimates</italic> and <italic>explained-variance</italic> metrics that measure some form of goodness of fit, while StLe has a focus on <italic>out-of-sample estimates</italic> and <italic>prediction accuracy</italic>.</p>
</sec>
<sec id="s5">
<title>Case study three: significant group differences and predicting the group of participants</title>
<p>Vignette: After isolating the neural correlates underlying face processing, the neuroimaging investigator wants to examine their relevance in psychiatric disease. In addition to the 40 healthy participants, 40 patients diagnosed with schizophrenia are recruited and administered the same experimental paradigm and set of face and house pictures. In this clinical fMRI study on group differences, the investigator wants to explore possible imaging-derived markers that index deficits in social-affective processing in patients carrying a diagnosis of schizophrenia.</p>
<p>Question: Can metrics of statistical relevance from ClSt and StLe be combined to corroborate a given candidate biomarker?</p>
<p>Many investigators in imaging neuroscience share a background in psychology, biology, or medicine, which includes training in traditional &#x0201C;textbook&#x0201D; statistics. Many neuroscientists have thus adopted a natural habit of assessing the quality of statistical relationships by means of <italic>p</italic>-values, effect sizes, confidence intervals, and statistical power. These are ubiquitously taught and used at many universities, although they are not the only coherent set of statistical diagnostics (Figure <xref ref-type="fig" rid="F5">5</xref>). These outcome metrics from ClSt may for instance be less familiar to some scientists with a background in computer science, physics, engineering, or philosophy. As an equally legitimate and internally coherent, yet less widely known diagnostic toolkit from the StLe community, prediction accuracy, precision, recall, confusion matrices, F1 score, and learning curves can also be used to measure the relevance of statistical relationships (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>; Yarkoni and Westfall, <xref ref-type="bibr" rid="B217">2017</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Key differences between measuring outcomes in classical statistics and statistical learning. Ten intuitions on quantifying statistical modeling outcomes that tend to be relatively more true for classical statistical methods (blue) or pattern-learning methods (red). ClSt typically yields point estimates and interval estimates (e.g., <italic>p</italic>-values, variances, confidence intervals), whereas StLe frequently outputs a function or a program that can yield point and interval estimates on new observations (e.g., the <italic>k</italic>-means centroids or a trained classifier&#x00027;s decision function can be applied to new data). In many cases, classical inference is a judgment about an entire data sample, whereas a trained predictive model can obtain quantitative answers from a single data point.</p></caption>
<graphic xlink:href="fnins-11-00543-g0005.tif"/>
</fig>
<p>On a general basis, applications of ClSt and StLe methods may not judge findings on identical grounds (Breiman, <xref ref-type="bibr" rid="B20">2001</xref>; Shmueli, <xref ref-type="bibr" rid="B185">2010</xref>; Lo et al., <xref ref-type="bibr" rid="B142">2015</xref>). There is an often-overlooked misconception that models with high explanatory performance do necessarily exhibit high predictive performance (Wu et al., <xref ref-type="bibr" rid="B213">2009</xref>; Lo et al., <xref ref-type="bibr" rid="B142">2015</xref>; Yarkoni and Westfall, <xref ref-type="bibr" rid="B217">2017</xref>). For instance, brain voxels in ventral visual stream found to well <italic>explain</italic> the difference between face processing in healthy and schizophrenic participants based on an ANOVA may not in all cases be the best brain features to train a support vector machine to <italic>predict</italic> this group effect in new participants. An important outcome measure in ClSt is the quantified <italic>significance</italic> associated with a statistical relationship between few variables given a pre-specified model. ClSt tends to <italic>test for a particular structure</italic> in the brain data based on <italic>analytical guarantees</italic>, in form of as mathematical convergence theorems about approximating the population properties with increasing sample size. The outcome measure for StLe is the quantified <italic>generalization of patterns</italic> between many variables or, more generally, the robustness of special structure in the data (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>). In the neuroimaging literature, reports of statistical outcomes have previously been noted to confuse diagnostic measures from classical statistics and statistical learning (Friston, <xref ref-type="bibr" rid="B83">2012</xref>).</p>
<p>For neuroscientists adopting a ClSt culture computing <italic>p</italic>-values takes a central position. The <italic>p-value</italic> denotes the probability of observing a result at least as extreme as a test statistic, assuming the null hypothesis is true. Results are considered significant when it is equal or below a pre-specified value, like <italic>p</italic> &#x0003D; 0.05 (Anderson et al., <xref ref-type="bibr" rid="B4">2000</xref>). Under the condition of sufficiently high power (cf. below), it quantifies the strength of evidence against the null hypothesis as a continuous function (Rosnow and Rosenthal, <xref ref-type="bibr" rid="B177">1989</xref>). Counterintuitively, it is not an immediate judgment on the alternative hypothesis H<sub>1</sub> preferred by the investigator (Cohen, <xref ref-type="bibr" rid="B43">1994</xref>; Anderson et al., <xref ref-type="bibr" rid="B4">2000</xref>). <italic>P</italic>-values do also not qualify the possibility of replication. It is another important caveat that a finding in the brain becomes more statistically significant (i.e., lower <italic>p</italic>-value) with increasing sample size (Berkson, <xref ref-type="bibr" rid="B15">1938</xref>; Miller et al., <xref ref-type="bibr" rid="B148">2016</xref>).</p>
<p>The essentially binary <italic>p</italic>-value (i.e., significant vs. not significant) is therefore often complemented by continuous <italic>effect size</italic> measures for the importance of rejecting H<sub>0</sub>. The effect size allows the identification of marginal effects that pass the statistical significance threshold but are not practically relevant in the real world. The <italic>p</italic>-value is a deductive <italic>inferential</italic> measure, whereas the effect size is a <italic>descriptive</italic> measure that follows neither inductive nor deductive reasoning. The (normalized) effect size can be viewed as the strength of a statistical relationship&#x02014;how much H<sub>0</sub> deviates from H<sub>1</sub>, or the likely presence of an effect in the general population (Chow, <xref ref-type="bibr" rid="B35">1998</xref>; Ferguson, <xref ref-type="bibr" rid="B68">2009</xref>; Kelley and Preacher, <xref ref-type="bibr" rid="B130">2012</xref>). This diagnostic measure is often unit-free, sample-size independent, and typically standardized. As a property of the actual statistical test, the effect size can be essential to report for biological understanding, but has different names and takes various forms, such as <italic>rho</italic> in Pearson correlation, <italic>eta</italic><sup><italic>2</italic></sup> in explained variances, and <italic>Cohen&#x00027;s d</italic> in differences between group averages.</p>
<p>Additionally, the certainty of a <italic>point estimate</italic> (i.e., the outcome is a value) can be expressed by an <italic>interval estimate</italic> (i.e., the outcome is a value range) using <italic>confidence intervals</italic> (Casella and Berger, <xref ref-type="bibr" rid="B31">2002</xref>). These variability diagnostics indicate a range of values between which the true value will fall a given proportion of the time (Estes, <xref ref-type="bibr" rid="B66">1997</xref>; Nickerson, <xref ref-type="bibr" rid="B158">2000</xref>; Cumming, <xref ref-type="bibr" rid="B49">2009</xref>). Typically, a 95% confidence interval is spanned around the population mean in 19 out of 20 cases across all observed samples. The tighter the confidence interval, the smaller the variance of the point estimate of the population parameter in each drawn sample. The estimation of confidence intervals is influenced by sample size and population variability. Confidence intervals may be asymmetrical (ignored by Gaussianity assumptions; Efron, <xref ref-type="bibr" rid="B59">2012</xref>), can be reported for different statistics and with different percentage borders. Notably, they can be used as a viable surrogate for formal tests of statistical significance in many scenarios (Cumming, <xref ref-type="bibr" rid="B49">2009</xref>).</p>
<p>Some confidence intervals can be computed in various data scenarios and statistical regimes, whereas the <italic>power</italic> may be especially meaningful within the culture of classical hypothesis testing (Cohen, <xref ref-type="bibr" rid="B40">1977</xref>, <xref ref-type="bibr" rid="B42">1992</xref>; Oakes, <xref ref-type="bibr" rid="B162">1986</xref>). To estimate power the investigator needs to specify the true effect size and variance under H<sub>1</sub>. The ClSt-minded investigator can then estimate the probability for rejecting null hypotheses that should be rejected, at the given threshold alpha and given that H<sub>1</sub> is true. A high power thus ensures that statistically significant and non-significant tests indeed reflect a property of the population (Chow, <xref ref-type="bibr" rid="B35">1998</xref>). Intuitively, a small confidence interval around a relevant effect suggests high statistical power. False negatives (i.e., Type II errors, beta error) become less likely with higher power (&#x0003D; 1&#x02014;beta error) (cf. Ioannidis, <xref ref-type="bibr" rid="B121">2005</xref>). Concretely, an underpowered investigation means that the investigator is less likely to be able to distinguish between H<sub>0</sub> and H<sub>1</sub> at the specified significance threshold alpha. Power calculations depend on several factors, including significance threshold alpha, the effect size in the population, variation in the population, sample size <italic>n</italic>, and experimental design (Cohen, <xref ref-type="bibr" rid="B42">1992</xref>).</p>
<p>While neuroimaging studies based on classical statistical inference ubiquitously report <italic>p</italic>-values and confidence intervals, there have however been few reports of effect size in the neuroimaging literature (Kriegeskorte et al., <xref ref-type="bibr" rid="B135">2010</xref>). Effect sizes are however necessary to compute power estimates. This explains the even rarer occurrence of power calculations in the neuroimaging literature (Yarkoni and Braver, <xref ref-type="bibr" rid="B215">2010</xref>; but see Poldrack et al., <xref ref-type="bibr" rid="B171">2017</xref>). Given the importance of <italic>p</italic>-values <italic>and</italic> effect sizes, the goal of computing both these useful statistics, such as for group differences in the neural processing of face stimuli, can be achieved based on two independent samples of these experimental data (especially if some selection process has been used). One sample would be used to perform statistical inference on the neural activity change yielding a <italic>p</italic>-value and one sample to obtain unbiased effect sizes. Further, it has been previously emphasized (Friston, <xref ref-type="bibr" rid="B83">2012</xref>) that <italic>p</italic>-values and effect sizes reflect in-sample estimates in a retrospective inference regime (ClSt). These metrics find an analog in out-of-sample estimates issued from cross-validation in a prospective prediction regime (StLe). In-sample effect sizes are typically an <italic>optimistic</italic> estimate of the &#x0201C;true&#x0201D; effect size (inflated by high significance thresholds), whereas out-of-sample effect sizes are <italic>unbiased</italic> estimates of the &#x0201C;true&#x0201D; effect size.</p>
<p>In the high-dimensional scenario, the StLe-minded investigator analyzing &#x0201C;wide&#x0201D; neuroimaging data in our case, computing, and judging statistical significance by <italic>p</italic>-values can become challenging (B&#x000FC;hlmann and Van De Geer, <xref ref-type="bibr" rid="B25">2011</xref>; Efron, <xref ref-type="bibr" rid="B59">2012</xref>; James et al., <xref ref-type="bibr" rid="B124">2013</xref>). Instead, <italic>classification accuracy</italic> on fresh data is a frequently reported performance metric in neuroimaging studies using learning algorithms. The <italic>classification accuracy</italic> is a simple summary statistic that captures the fraction of correct prediction instances among all performed applications of a fitted model. Basing interpretation on accuracy alone can be an insufficient diagnostic because it is frequently influenced by the number of samples, the local characteristics of hemodynamic responses, efficiency of experimental design, data folding into train and test sets, and differences in the feature number <italic>p</italic> (Haynes, <xref ref-type="bibr" rid="B116">2015</xref>). A potentially under-exploited data-driven tool in this context is <italic>bootstrapping</italic>. The archetypical example of computer-intensive statistical method enables population-level inference of unknown distributions largely independent of model complexity by repeated random draws from the neuroimaging data sample at hand (Efron, <xref ref-type="bibr" rid="B58">1979</xref>; Efron and Tibshirani, <xref ref-type="bibr" rid="B62">1994</xref>). This opportunity to equip various point estimates by an interval estimate of certainty (e.g., the possibly asymmetrical interval for the &#x0201C;true&#x0201D; accuracy of a classifier) is unfortunately seldom embraced in neuroimaging today (but see Bellec et al., <xref ref-type="bibr" rid="B10">2010</xref>; Pernet et al., <xref ref-type="bibr" rid="B167">2011</xref>; Vogelstein et al., <xref ref-type="bibr" rid="B205">2014</xref>). Besides providing confidence intervals, bootstrapping can also perform non-parametric null hypothesis testing. This may be one of few examples of a direct connection between ClSt and StLe methodology. Alternatively, <italic>binomial tests</italic> have been used to obtain a <italic>p</italic>-value estimate of statistical significance from accuracies and other performance scores (Pereira et al., <xref ref-type="bibr" rid="B166">2009</xref>; Brodersen et al., <xref ref-type="bibr" rid="B22">2013</xref>; Hanke et al., <xref ref-type="bibr" rid="B108">2015</xref>) in the binary classification setting. It has frequently been employed to reject the null hypothesis that two categories occur equally often. There are however increasing concerns about the validity of this approach if statistical independence between the performance estimates (e.g., prediction accuracies from each cross-validation fold) is in question (Pereira and Botvinick, <xref ref-type="bibr" rid="B165">2011</xref>; Noirhomme et al., <xref ref-type="bibr" rid="B159">2014</xref>; Jamalabadi et al., <xref ref-type="bibr" rid="B123">2016</xref>). Yet another option to derive <italic>p</italic>-values from classification performances of two groups is <italic>label permutation</italic> based on non-parametric resampling procedures (Nichols and Holmes, <xref ref-type="bibr" rid="B157">2002</xref>; Golland and Fischl, <xref ref-type="bibr" rid="B97">2003</xref>). This algorithmic significance-testing tool can serve to reject the null hypothesis that the neuroimaging data do not contain relevant information about the group labels in many complex data analysis settings.</p>
<p>The neuroscientist who adopted a StLe culture is in the habit of corroborating prediction accuracies using <italic>cross-validation:</italic> the de facto standard to obtain an unbiased estimate of a model&#x00027;s capacity to generalize beyond the brain scans at hand (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>; Bishop, <xref ref-type="bibr" rid="B16">2006</xref>). <italic>Model assessment</italic> is commonly done by training on a bigger subset of the available data (i.e., <italic>training set</italic> for <italic>in-sample performance</italic>) and subsequent application of the trained model to the typically smaller remaining part of data (i.e., <italic>test set</italic> for <italic>out-of-sample performance</italic>), both assumed to be drawn from the same distribution. Cross-validation typically divides the sample into data splits such that the class label (i.e., healthy vs. schizophrenic) of each data point is to be predicted once. The pairs of model-predicted label and the corresponding true label for each data point (i.e., brain scan) in the dataset can then be submitted to the quality measures (Powers, <xref ref-type="bibr" rid="B175">2011</xref>), including <italic>prediction accuracy</italic> (inversely related to <italic>prediction error</italic>), <italic>precision, recall</italic>, and <italic>F1 score</italic>. Accuracy and the other performance metrics are often computed separately on the training set and the test set. Additionally, the measures from training and testing can be expressed by their inverse (e.g., <italic>training error</italic> as <italic>in-sample error</italic> and <italic>test error</italic> as <italic>out-of-sample error</italic>) because the positive and negative cases are interchangeable.</p>
<p>The classification accuracy can be further decomposed into group-wise metrics based on the so-called <italic>confusion matrix</italic>, the juxtaposition of the true and predicted group memberships. The <italic>precision</italic> measures (Table <xref ref-type="table" rid="T1">1</xref>) how many of the labels predicted from brain scans are correct, that is, how many participants predicted to belong to a certain class really belong to that class. Put differently, among the participants predicted to suffer from schizophrenia, how many have really been diagnosed with that disease? On the other hand, the <italic>recall</italic> measures how many labels are correctly predicted, that is, how many members of a class were predicted to really belong to that class. Hence, among the participants known to be affected by schizophrenia, how many were actually detected as such? Precision can be viewed as a measure of &#x0201C;exactness&#x0201D; and recall as a measure of &#x0201C;completeness&#x0201D; (Powers, <xref ref-type="bibr" rid="B175">2011</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Metrics used to create ROC curves.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Notion</bold></th>
<th valign="top" align="left"><bold>Formula</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Specificity</td>
<td valign="top" align="left">true negative/(true negative &#x0002B; false positive)</td>
</tr>
<tr>
<td valign="top" align="left">Sensitivity/Recall</td>
<td valign="top" align="left">true positive/(true positive &#x0002B; false negative)</td>
</tr>
<tr>
<td valign="top" align="left">Precision</td>
<td valign="top" align="left">true positive/(true positive &#x0002B; false positive)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Neither accuracy, precision, or recall allow injecting subjective importance into the evaluation process of the learning algorithm. This disadvantage is addressed by the <italic>F</italic><sub><italic>beta</italic></sub> <italic>score</italic>: a weighted combination of the precision and recall prediction scores. Concretely, the F<sub>1</sub> score would equally weigh precision and recall of class predictions, while the F<sub>0.5</sub> score puts more emphasis on precision and the F<sub>2</sub> score more on recall. Moreover, applications of recall, precision, and F<sub>beta</sub> scores have been noted to ignore the true negative cases as well as to be highly susceptible to estimator bias (Powers, <xref ref-type="bibr" rid="B175">2011</xref>). Needless to say, no single outcome metric can be equally optimal in all contexts.</p>
<p>Extending from the setting of healthy-diseased classification to the <italic>multi-class setting</italic> (e.g., comparing healthy, schizophrenic, bipolar, and autistic participants) injects ambiguity into the interpretation of accuracy scores. Rather than reporting mere better-than-chance findings in StLe analyses, it becomes more important to evaluate the F<sub>1</sub>, precision and recall scores for each class to be predicted in the brain scans (e.g., Brodersen et al., <xref ref-type="bibr" rid="B24">2011b</xref>; Schwartz et al., <xref ref-type="bibr" rid="B183">2013</xref>). It is important to appreciate that the sensitivity/specificity metrics, perhaps more frequently reported in ClSt communities, and the precision/recall metrics, probably more frequently reported in StLe communities, tell slightly different stories about identical neuroscientific findings. In fact, sensitivity equates with recall. Specificity does however not equate with precision. Further, a ClSt view on the StLe metrics would be that maximum precision corresponds to absent Type I errors (i.e., no false positives), whereas maximum recall corresponds to absent false negatives (i.e., no Type II errors). Again, Type I and II errors are related to the entirety of data points in a ClSt regime and prediction is only evaluated on a test data split of the sample in an StLe regime. Moreover, many empirical sciences usually aggregate results in <italic>ROC</italic> (receiver operating characteristic) curves plotting sensitivity against specificity scores, whereas other scientific domains tend to report analogous yet different <italic>recall-precision curves</italic> instead (Altman and Bland, <xref ref-type="bibr" rid="B2">1994</xref>; Davis and Goadrich, <xref ref-type="bibr" rid="B51">2006</xref>; Dem&#x00161;ar, <xref ref-type="bibr" rid="B53">2006</xref>).</p>
<p>Finally, StLe-minded investigators use <italic>learning curves</italic> (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>; Murphy, <xref ref-type="bibr" rid="B152">2012</xref>) as an important diagnostic tool for empirical estimates of the <italic>sample complexity</italic>, that is, the achieved model fit and prediction accuracy as a function of the available sample size n. For increasingly bigger subsets of the training set, a classification algorithm is trained on that current share of the training set and then evaluated for accuracy on the always-same test set. Across subset instances, simple models display relatively high in-sample error because they can not approximate the target function very well (underfitting) but exhibit good generalization to unseen data with relatively low out-of-sample error. Yet, complex models display relatively low in-sample error because they adapt too well to the data (overfitting) with difficulty to extrapolate to newly sampled data with high out-of-sample error. Put differently, a big gap between high in-sample and low out-of-sample performance is typically observed for high-variance models, such as artificial neural network algorithms or random forests. These performance metrics from different data splits often converge for high-bias models, such as linear support vector machines and logistic regression.</p>
<p>In sum, the ClSt and StLe communities rely on diagnostic metrics that are largely incongruent and may therefore not lend themselves for direct comparison in all practical analysis settings.</p>
</sec>
<sec id="s6">
<title>Case study four: out-of-sample generalization and subsequent classical inference</title>
<p>Vignette: The investigator is interested in potential differences in brain volume that are associated with an individual&#x00027;s age (<italic>continuous target variable</italic>). A LASSO (<italic>often considered as StLe arsenal</italic>) is computed on the voxel-based morphometry data from the brain&#x00027;s gray matter of the 1,200-subject HCP release (Human Connectome Project; Van Essen et al., <xref ref-type="bibr" rid="B199">2012</xref>). This L1-penalized residual-sum-of-squares regression performs automatic variable selection (i.e., <italic>effectively eliminates coefficients by setting them to zero</italic>) on all gray-matter voxels&#x00027; volume information in a <italic>high-dimensional regime</italic> (i.e., no mass-univariate analysis). Assessing <italic>generalization</italic> performance of different sparse models using 5-fold cross-validation yields the non-zero coefficients for few brain voxels whose volumetric information is most <italic>predictive</italic> of an individual&#x00027;s age.</p>
<p>Question: How can the investigator perform <italic>classical inference</italic> to know which of the gray-matter voxels selected to be predictive for biological age are <italic>statistically significant</italic>?</p>
<p>This is an important concern because most statistical methods currently applied to large datasets perform some explicit or implicit form of variable selection (Jenatton et al., <xref ref-type="bibr" rid="B125">2011</xref>; Committee on the Analysis of Massive Data et al., <xref ref-type="bibr" rid="B44">2013</xref>; Hastie et al., <xref ref-type="bibr" rid="B113">2015</xref>). There are even many different forms of preliminary selection of variables before performing significance tests on them. First, LASSO is a widely used estimator in engineering, compressive sensing, various &#x0201C;omics&#x0201D; branches and other sciences, where it is often applied without an additional significance test. Beyond neuroscience, generalization-approved statistical learning models are routinely solving a diverse set of real-world challenges. This includes algorithmic trading in financial markets, fraud detection in credit card transactions, real-time speech translation, SPAM filtering for e-mails, face recognition in digital cameras, and piloting self-driving cars (Jordan and Mitchell, <xref ref-type="bibr" rid="B126">2015</xref>; LeCun et al., <xref ref-type="bibr" rid="B139">2015</xref>). In all these examples, statistical learning algorithms successfully generalize to unseen, later acquired data and thus tackle the problem heuristically without classical significance test on specific variables or for overall model performance.</p>
<p>Second, the LASSO has been introduced as an elegant solution to the combinatorial problem of what subset of gray-matter voxels is sufficient for predicting an individual&#x00027;s age by <italic>automatic variable selection</italic> (Tibshirani, <xref ref-type="bibr" rid="B195">1996</xref>). Computing voxel-wise <italic>p</italic>-values would recast this high-dimensional pattern-learning setting (i.e., considering all brain voxels at once) into a mass-univariate hypothesis-testing problem (i.e., considering one voxel after the other) where relevance would be computed independently for each voxel and correction for multiple comparisons would become necessary. Yet, recasting into the mass-univariate setting would ignore the sophisticated selection process that led to the predictive model with a reduced number of variables (Wu et al., <xref ref-type="bibr" rid="B213">2009</xref>). Put differently, variable selection via the LASSO is itself a stochastic process that is however not accounted for by the theoretical guarantees of classical inference for statistical significance (Berk et al., <xref ref-type="bibr" rid="B14">2013</xref>). Put in yet another way, data-driven model selection is corrupting the null hypothesis of classical statistical inference because the sampling distribution of the parameter estimates is altered. The important consequence is that na&#x000EF;ve classical inference expects a non-adaptive model chosen before data acquisition and can therefore not be readily used along LASSO in particular or arbitrary selection procedures in general<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref>.</p>
<p>Third, the portrayed conflict between more exploratory model selection by cross-validation (StLe) and more confirmatory classical inference (ClSt) is currently at the frontier of statistical development (Loftus, <xref ref-type="bibr" rid="B143">2015</xref>; Taylor and Tibshirani, <xref ref-type="bibr" rid="B192">2015</xref>). New methods for so-called <italic>post-selection inference</italic> (or <italic>selective inference</italic>) allow computing <italic>p</italic>-values for a set of features that have previously been chosen to be meaningful predictors by some criterion, one example being sparsity-incuding prediction algorithms such as LASSO. According to the theory of ClSt, the statistical model is to be chosen before visiting the data. Classical statistical tests and confidence intervals therefore become invalidated and the <italic>p</italic>-values become optimistically biased (Berk et al., <xref ref-type="bibr" rid="B14">2013</xref>). Consequently, the association between a predictor and the target variable must be even stronger to certify the same level of significance. Selective inference for modern adaptive regression thus replaces loose <italic>na&#x000EF;ve p-values</italic> by more rigorous <italic>selection-adjusted p-values</italic>. As an ordinary null hypothesis can hardly be adopted in this adaptive testing setting, conceptual extension is also prompted on the level of ClSt theory itself (Hastie et al., <xref ref-type="bibr" rid="B113">2015</xref>). For instance, closed-form solutions to adjusted classical inference after variable selection already exist for principal component analysis (Choi et al., <xref ref-type="bibr" rid="B34">2014</xref>) and forward stepwise regression (Taylor et al., <xref ref-type="bibr" rid="B191">2014</xref>). Moreover, a simple alternative to formally account for preceding model selection is <italic>data splitting</italic> (Cox, <xref ref-type="bibr" rid="B48">1975</xref>; Wasserman and Roeder, <xref ref-type="bibr" rid="B208">2009</xref>; Fithian et al., <xref ref-type="bibr" rid="B73">2014</xref>), which is frequent practice in genetics (e.g., Sladek et al., <xref ref-type="bibr" rid="B186">2007</xref>). In this procedure, the variable selection procedure is computed on one data split and <italic>p</italic>-values are computed on the remaining second data split. However, such data splitting is not always possible and will incur power losses.</p>
<p>In sum, in many analysis settings, the same data should typically not be used to first apply supervised learning algorithms for automatic selection of the most predictive variables and to then test for statistical significance of the variables already found to be most predictive based on these data points. The recent developments for post-selection inference can be viewed as an attempt to reconcile certain aspects of how the StLe and ClSt paradigms draw conclusions from data.</p>
</sec>
<sec id="s7">
<title>Case study five: classical inference and subsequent out-of-sample generalization</title>
<p>Vignette: The investigator is interested in potential brain structure differences that are associated with an individual&#x00027;s gender (<italic>categorical target variable</italic>) in the voxel-based morphometry data of the 1,200-subject HCP release (Human Connectome Project; Van Essen et al., <xref ref-type="bibr" rid="B199">2012</xref>). First, the &#x0003E;100,000 voxels per brain scan are reduced to the most important 10,000 voxels to lower the computational cost and facilitate estimation of a prediction model. To this end, ANOVA (<italic>univariate test for statistical significance belonging to ClSt</italic>) is initially used to obtain a ranking of the most relevant 10,000 features from the gray matter. This selects the 10,000 out of the original &#x0003E;100,000 voxel variables with highest variance explaining volume differences between males and females (i.e., <italic>the gender information associated with each brain scan is used in the univariate test</italic>). Second, support vector machine classification (&#x0201C;<italic>multivariate&#x0201D; pattern-learning algorithm belonging to StLe</italic>) is performed by cross-validation on a feature space with the 10,000 preselected gray-matter measurements to predict the gender from each subject&#x00027;s brain scan.</p>
<p>Question: Is an analysis pipeline with <italic>univariate classical inference</italic> and subsequent <italic>high-dimensional prediction</italic> valid if both steps rely on gender as the target variables?</p>
<p>The implications of feature engineering procedures applied before training a learning algorithm is a frequent concern and can require subtle answers (Guyon and Elisseeff, <xref ref-type="bibr" rid="B103">2003</xref>; Kriegeskorte et al., <xref ref-type="bibr" rid="B136">2009</xref>; Lemm et al., <xref ref-type="bibr" rid="B140">2011</xref>; Hanke et al., <xref ref-type="bibr" rid="B108">2015</xref>). In most applications of predictive models the large majority of brain voxels will not be very informative (Brodersen et al., <xref ref-type="bibr" rid="B23">2011a</xref>). The described scenario of <italic>dimensionality reduction</italic> by feature selection to focus prediction is clearly allowed under the condition that the ANOVA is not computed on the entire data sample. Rather, the initial identification of voxels explaining most variance between the male and female individuals should be computed only on the training set in each cross-validation fold. In the training set and test set of each fold the same identified candidate voxels are then regrouped into a feature space that is fed into the support vector machine algorithm. This ensures an identical feature space for model training and model testing but its construction only depends on structural brain scans from the training set. Generally, voxel preprocessing performed before model training is authorized if the feature space construction is not influenced by properties of the concealed test set. In the present scenario, the Vapnik-Chervonenkis bounds of the cross-validation estimator are therefore not loosened or invalidated if class labels have been exploited for feature selection or depending on whether the feature selection procedure is univariate or multivariate (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>; Shalev-Shwartz and Ben-David, <xref ref-type="bibr" rid="B184">2014</xref>). Put differently, the cross-validation procedure simply evaluates the entire prediction process including the automatized and potentially nested dimensionality reduction approaches. In sum, in an StLe regime, using class information during feature preprocessing for a cross-validated supervised estimator is not an instance of <italic>data-snooping</italic> (or <italic>peeking</italic>) if done exclusively on the training set (Abu-Mostafa et al., <xref ref-type="bibr" rid="B1">2012</xref>).</p>
<p>At the core of this explanation is the goal of cross-validation to yield <italic>out-of-sample estimates</italic>. In stark contrast, remember that null-hypothesis testing yields <italic>in-sample estimates</italic> as it needs all available data points to take its decision. Using the class labels for a variable selection step just before null-hypothesis testing on a same data sample would invalidate the null hypothesis (Kriegeskorte et al., <xref ref-type="bibr" rid="B136">2009</xref>, <xref ref-type="bibr" rid="B135">2010</xref>). Consequently, in a ClSt regime, using class information to select variables before null-hypothesis testing will incur an instance of <italic>double-dipping</italic> (or <italic>circular analysis</italic>). This also occurs when, for instance, first correlating a behavioral measure with brain activity and then using the identified subset of brain voxels for a second correlation analysis with that same behavioral measurement (Lieberman et al., <xref ref-type="bibr" rid="B141">2009</xref>; Vul et al., <xref ref-type="bibr" rid="B206">2009</xref>). In this scenario, voxels are submitted to two statistical tests with the same goal in a nested, non-independent fashion (Freedman, <xref ref-type="bibr" rid="B77">1983</xref>). This corrupts the <italic>validity of the null hypothesis</italic> on which the reported test results conditionally depend.</p>
<p>Regarding interpretation of the results, the classifier will miss some brain voxels that only carry relevant information when considered in voxel ensembles. This is because the ANOVA filter has kept voxels that are independently relevant (Brodersen et al., <xref ref-type="bibr" rid="B23">2011a</xref>). Univariate feature selection in high-dimensional brain scans may therefore systematically encourage model selection (i.e., each weight combination equates with a model hypothesis from the classifier&#x00027;s function space) that is not tuned to neurobiological meaningfulness. Concretely, in the discussed scenario the classifier learns <italic>complex patterns between voxels that were previously chosen to be individually important</italic>. This may considerably weaken the interpretability and conclusions on &#x0201C;whole-brain multivariate patterns&#x0201D;. Remember also that variables that have a <italic>statistically significant association</italic> with a target variable do not necessarily have good <italic>generalization performance</italic>, and vice versa (Shmueli, <xref ref-type="bibr" rid="B185">2010</xref>; Lo et al., <xref ref-type="bibr" rid="B142">2015</xref>; Bzdok and Yeo, <xref ref-type="bibr" rid="B30">2017</xref>). On the upside, it is frequently observed that the combination of whole-brain univariate feature selection and linear classification is among the best approaches if the primary goal is maximizing <italic>prediction performance</italic> as opposed to maximizing <italic>interpretability</italic>.</p>
<p>Finally, it is interesting to consider that ANOVA-mediated feature selection to a subset of <italic>p</italic> &#x0003C; 500 voxel variables would reduce the &#x0201C;wide&#x0201D; neuroimaging data (&#x0201C;<italic>n</italic> &#x0003C; &#x0003C; <italic>p</italic>&#x0201D; setting) down to &#x0201C;long&#x0201D; neuroimaging data with fewer features than observations (&#x0201C;<italic>n</italic> &#x0003E; <italic>p</italic>&#x0201D; setting) given the <italic>n</italic> &#x0003D; 500 subjects (Wainwright, <xref ref-type="bibr" rid="B207">2014</xref>). This allows recasting the StLe regime into a ClSt regime in order to fit a GLM and perform classical statistical tests instead of training a predictive classification algorithm (Brodersen et al., <xref ref-type="bibr" rid="B23">2011a</xref>).</p>
<p>In sum, in many analysis settings, prediction algorithms can be trained after choosing the input variables most significantly associated with an explanatory target variable if the initial classical inference (<italic>p</italic>-values) is performed only in the training set and the ensuing evaluation of algorithm generalization (prediction performance) is performed on the independent test set.</p>
</sec>
<sec id="s8">
<title>Case study six: structure discovery by clustering algorithms</title>
<p>Vignette: Each functionally specialized region in the human brain probably has a unique set of long-range connections (Passingham et al., <xref ref-type="bibr" rid="B163">2002</xref>). This notion has prompted connectivity-based parcellation methods in neuroimaging that segregate an ROI (can be locally circumscribed or brain global; Eickhoff et al., <xref ref-type="bibr" rid="B64">2015</xref>) into distinct cortical modules (Behrens et al., <xref ref-type="bibr" rid="B9">2003</xref>). The whole-brain connectivity for each ROI voxel is computed and the voxel-wise connectional fingerprints are submitted to a clustering algorithm (i.e., <italic>individual brain voxels in the ROI are the elements to group; the connectivity strength values are the features of each element for similarity assessment</italic>). The investigator wants to apply connectivity-based parcellation to the fusiform gyrus to segregate this ROI into cortical modules that exhibit similar connectivity patterns with the rest of the brain and are, thus potentially, functionally distinct. That is, voxels within the same cluster in the ROI will have more similar whole-brain connectivity properties than voxels from different clusters in the fusiform gyrus.</p>
<p>Question: Is it possible to decide whether the obtained brain <italic>clusters</italic> are <italic>statistically significant</italic>?</p>
<p>In essence, the aim of connectivity-guided brain parcellation is to find useful, simplified structure by imposing circumscribed compartments on brain topography (Yeo et al., <xref ref-type="bibr" rid="B219">2011</xref>; Smith et al., <xref ref-type="bibr" rid="B187">2013</xref>; Frackowiak and Markram, <xref ref-type="bibr" rid="B76">2015</xref>). This is typically achieved by using k-means, hierarchical, Ward, or spectral clustering algorithms (Thirion et al., <xref ref-type="bibr" rid="B194">2014</xref>; Eickhoff et al., <xref ref-type="bibr" rid="B64">2015</xref>). Putting on the ClSt hat, an ROI clustering result would be deemed statistically significant if the obtained data are incompatible with the null hypothesis that the investigator seeks to reject (Everitt, <xref ref-type="bibr" rid="B67">1979</xref>; Halkidi et al., <xref ref-type="bibr" rid="B105">2001</xref>). Choosing a test statistic for clustering solutions to obtain <italic>p</italic>-values is difficult (Vogelstein et al., <xref ref-type="bibr" rid="B205">2014</xref>) because of the need to find a meaningful null hypothesis to test against (Jain et al., <xref ref-type="bibr" rid="B122">1999</xref>). Put differently, for classical inference based on statistical hypothesis testing one may need to pick an arbitrary null hypothesis to falsify. It follows that neither the ClSt notions of effect size and power do seem to apply in the case of brain parcellation (also a frequent question by paper reviewers). Instead of classical inference to formally <italic>test</italic> for a particular structure in the clustering results, the investigator actually needs to resort to exploratory approaches that discover and assess structure in the neuroimaging data (Tukey, <xref ref-type="bibr" rid="B196">1962</xref>; Efron and Tibshirani, <xref ref-type="bibr" rid="B61">1991</xref>; Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>). Although statistical methods span a continuum between the two poles of ClSt and StLe, finding a clustering model with the highest fit in the sense of explaining the regional connectivity differences at hand is perhaps more naturally situated in the StLe community.</p>
<p>Putting on the StLe hat, the investigator realizes that the problem of brain parcellation constitutes an <italic>unsupervised</italic> learning setting without any target variable y to predict (e.g., cognitive tasks, the age or gender of the participants). The learning problem does therefore not consist in estimating a supervised predictive model y &#x0003D; f(X), but to estimate an unsupervised descriptive model for the connectivity data X themselves. Solving such unsupervised estimation problems is generally recognized to be ill-posed because it is generally unclear what the best way is to quantify how well relevant structure has been captured and what notion of &#x0201C;relevance&#x0201D; is most pertinent (Hastie et al., <xref ref-type="bibr" rid="B112">2001</xref>; Ghahramani, <xref ref-type="bibr" rid="B90">2004</xref>; Bishop, <xref ref-type="bibr" rid="B16">2006</xref>; Shalev-Shwartz and Ben-David, <xref ref-type="bibr" rid="B184">2014</xref>). In clustering analysis, there are many possible transformations, projections, and compressions of X but there is usually no unique criterion of optimality that clearly suggests itself. On the one hand, the &#x0201C;true&#x0201D; <italic>shape of clusters</italic> is unknown for most real-world clustering problems, including brain parcellation studies. On the other hand, finding an &#x0201C;optimal&#x0201D; <italic>number of clusters</italic> represents an unresolved issue (<italic>cluster validity problem</italic>) in statistics in general and in brain neuroimaging in particular (Jain et al., <xref ref-type="bibr" rid="B122">1999</xref>; Handl et al., <xref ref-type="bibr" rid="B107">2005</xref>). In other words, &#x0201C;the clustering problem is inherently ill posed, in the sense that there is no single criterion that measures how well a clustering of data corresponds to the real world&#x0201D; (Goodfellow et al., <xref ref-type="bibr" rid="B98">2016</xref>). Evaluating the adequacy of clustering results is therefore conventionally addressed by applying different <italic>cluster validity criteria</italic> (Thirion et al., <xref ref-type="bibr" rid="B194">2014</xref>; Eickhoff et al., <xref ref-type="bibr" rid="B64">2015</xref>). These heuristic metrics are useful and necessary because clustering algorithms will always find some subregions in the investigator&#x00027;s ROI, that is, find relevant structure with respect to the particular optimization objective of the clustering algorithm whether such structure truly exists in nature or not. The various clustering validity criteria, possibly based on information theory, topology, or consistency (Eickhoff et al., <xref ref-type="bibr" rid="B64">2015</xref>), typically encourage cluster solutions with low within-cluster and high between-cluster differences according to a certain notion of optimality. Given that the notions of optimality are not coherent with each other (Shalev-Shwartz and Ben-David, <xref ref-type="bibr" rid="B184">2014</xref>; Thirion et al., <xref ref-type="bibr" rid="B194">2014</xref>), investigators should evaluate cluster findings and choose the cluster number by relying on a set of complementary cluster validity criteria, such as reproducibility and goodness of fit or bias and variance.</p>
<p>Evidently, the discovered set of connectivity-derived clusters only represent hints to candidate brain modules. Their &#x0201C;existence&#x0201D; in neurobiology requires further scrutiny (Thirion et al., <xref ref-type="bibr" rid="B194">2014</xref>; Eickhoff et al., <xref ref-type="bibr" rid="B64">2015</xref>). Nevertheless, such clustering solutions provide important means to narrow down high-dimensional neuroimaging data. Preliminary clustering results broaden the space of research hypotheses that the investigator can articulate. For instance, unexpected discovery of a candidate brain region (cf. Mars et al., <xref ref-type="bibr" rid="B147">2012</xref>; zu Eulenburg et al., <xref ref-type="bibr" rid="B222">2012</xref>) can provide an argument for future experimental investigations. Brain parcellation can thus be viewed as an exploratory unsupervised method outlining relevant structure in neuroimaging data that can subsequently be tested as research hypotheses in targeted future neuroimaging studies on classical inference or out-of-sample generalization.</p>
<p>In sum, in most analysis settings, quantifying the importance of clustering solutions is inherently ill-posed because, without an explanatory target variable, many different low-dimensional reexpressions of high-dimensional input data can be useful. Choosing the right variant among the possible dimensionality reductions by clustering algorithms alone can typically not be done based on extrapolation metrics from ClSt (<italic>p</italic>-values, effect size, power) or StLe (out-of-sample prediction performance, learning curves).</p>
</sec>
<sec sec-type="conclusions" id="s9">
<title>Conclusion</title>
<p>A novel scientific fact about the brain is only valid in the context of the complexity restrictions that have been imposed on the studied phenomenon during the investigation (Box, <xref ref-type="bibr" rid="B19">1976</xref>). Tools of the imaging neuroscientist&#x00027;s statistical arsenal can be placed on a continuum between <italic>classical inference</italic> by hypothesis falsification and increasingly used <italic>out-of-sample generalization</italic> by extrapolating complex patterns to independent data (Efron and Hastie, <xref ref-type="bibr" rid="B60">2016</xref>). While null-hypothesis testing has been dominating academic milieus in the empirical sciences and statistics departments for several decades, statistical learning methods are perhaps still more prevalent in data-intensive industries (Breiman, <xref ref-type="bibr" rid="B20">2001</xref>; Vanderplas, <xref ref-type="bibr" rid="B198">2013</xref>; Henke et al., <xref ref-type="bibr" rid="B119">2016</xref>). This sociological segregation may contribute to the existing confusion about the mutual relationship between the ClSt and StLe camps in application domains such as imaging neuroscience. Despite the incongruent historical trajectories and theoretical foundations, both statistical cultures aim at inferential conclusions by extracting new knowledge from data using mathematical models (Friston et al., <xref ref-type="bibr" rid="B84">2008</xref>; Committee on the Analysis of Massive Data et al., <xref ref-type="bibr" rid="B44">2013</xref>). However, an observed effect in the brain with a statistically significant <italic>p</italic>-value does not in all cases generalize to future brain recordings (Shmueli, <xref ref-type="bibr" rid="B185">2010</xref>; Arbabshirani et al., <xref ref-type="bibr" rid="B6">2017</xref>; Yarkoni and Westfall, <xref ref-type="bibr" rid="B217">2017</xref>). Conversely, a neurobiological effect that can be successfully captured by a learning algorithm as evidenced by out-of-sample generalization does not invariably entail a significant <italic>p</italic>-value when submitted to null-hypothesis testing. The distributional properties of brain data important for high statistical significance and for high prediction accuracy are not identical (Efron, <xref ref-type="bibr" rid="B59">2012</xref>; Lo et al., <xref ref-type="bibr" rid="B142">2015</xref>; Arbabshirani et al., <xref ref-type="bibr" rid="B6">2017</xref>). The goal and permissible conclusions of a neuroscientific investigation are therefore conditioned by the adopted statistical framework (cf. Feyerabend, <xref ref-type="bibr" rid="B69">1975</xref>). Awareness of the <italic>prediction-inference distinction</italic> will be criticial to keep pace with the increasing information detail of neuroimaging data repositories (Eickhoff et al., <xref ref-type="bibr" rid="B65">2016</xref>; Bzdok and Yeo, <xref ref-type="bibr" rid="B30">2017</xref>). Ultimately, statistical inference is not a uniquely defined concept.</p>
</sec>
<sec id="s10">
<title>Author contributions</title>
<p>The author confirms being the sole contributor of this work and approved it for publication.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abu-Mostafa</surname> <given-names>Y. S.</given-names></name> <name><surname>Magdon-Ismail</surname> <given-names>M.</given-names></name> <name><surname>Lin</surname> <given-names>H. T.</given-names></name></person-group> (<year>2012</year>). <source>Learning from Data</source>. <publisher-name>AMLBook</publisher-name>.</citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altman</surname> <given-names>D. G.</given-names></name> <name><surname>Bland</surname> <given-names>J. M.</given-names></name></person-group> (<year>1994</year>). <article-title>Statistics notes: diagnostic tests 2: predictive values</article-title>. <source>BMJ</source> <volume>309</volume>:<fpage>102</fpage>. <pub-id pub-id-type="doi">10.1136/bmj.309.6947.102</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amunts</surname> <given-names>K.</given-names></name> <name><surname>Lepage</surname> <given-names>C.</given-names></name> <name><surname>Borgeat</surname> <given-names>L.</given-names></name> <name><surname>Mohlberg</surname> <given-names>H.</given-names></name> <name><surname>Dickscheid</surname> <given-names>T.</given-names></name> <name><surname>Rousseau</surname> <given-names>M. E.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>BigBrain: an ultrahigh-resolution 3D human brain model</article-title>. <source>Science</source> <volume>340</volume>, <fpage>1472</fpage>&#x02013;<lpage>1475</lpage>. <pub-id pub-id-type="doi">10.1126/science.1235381</pub-id><pub-id pub-id-type="pmid">23788795</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>D. R.</given-names></name> <name><surname>Burnham</surname> <given-names>K. P.</given-names></name> <name><surname>Thompson</surname> <given-names>W. L.</given-names></name></person-group> (<year>2000</year>). <article-title>Null hypothesis testing: problems, prevalence, and an alternative</article-title>. <source>J. Wildl. Manage.</source> <fpage>912</fpage>&#x02013;<lpage>923</lpage>. <pub-id pub-id-type="doi">10.2307/3803199</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>M. L.</given-names></name></person-group> (<year>2010</year>). <article-title>Neural reuse: a fundamental organizational principle of the brain</article-title>. <source>Behav. Brain Sci.</source> <volume>33</volume>, <fpage>245</fpage>&#x02013;<lpage>266</lpage>; discussion 266&#x02013;313. <pub-id pub-id-type="doi">10.1017/S0140525X10000853</pub-id><pub-id pub-id-type="pmid">20964882</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arbabshirani</surname> <given-names>M. R.</given-names></name> <name><surname>Plis</surname> <given-names>S.</given-names></name> <name><surname>Sui</surname> <given-names>J.</given-names></name> <name><surname>Calhoun</surname> <given-names>V. D.</given-names></name></person-group> (<year>2017</year>). <article-title>Single subject prediction of brain disorders in neuroimaging: promises and pitfalls</article-title>. <source>Neuroimage</source> <volume>145</volume>, <fpage>137</fpage>&#x02013;<lpage>165</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2016.02.079</pub-id><pub-id pub-id-type="pmid">27012503</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Averbeck</surname> <given-names>B. B.</given-names></name> <name><surname>Latham</surname> <given-names>P. E.</given-names></name> <name><surname>Pouget</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Neural correlations, population coding and computation</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>7</volume>, <fpage>358</fpage>&#x02013;<lpage>366</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1888</pub-id><pub-id pub-id-type="pmid">16760916</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bach</surname> <given-names>F.</given-names></name></person-group> (<year>2014</year>). <article-title>Breaking the curse of dimensionality with convex neural networks</article-title>. arXiv:1412.8690.</citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Behrens</surname> <given-names>T. E.</given-names></name> <name><surname>Johansen-Berg</surname> <given-names>H.</given-names></name> <name><surname>Woolrich</surname> <given-names>M. W.</given-names></name> <name><surname>Smith</surname> <given-names>S. M.</given-names></name> <name><surname>Wheeler-Kingshott</surname> <given-names>C. A.</given-names></name> <name><surname>Boulby</surname> <given-names>P. A.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Non-invasive mapping of connections between human thalamus and cortex using diffusion imaging</article-title>. <source>Nat. Neurosci.</source> <volume>6</volume>, <fpage>750</fpage>&#x02013;<lpage>757</lpage>. <pub-id pub-id-type="doi">10.1038/nn1075</pub-id><pub-id pub-id-type="pmid">12808459</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bellec</surname> <given-names>P.</given-names></name> <name><surname>Rosa-Neto</surname> <given-names>P.</given-names></name> <name><surname>Lyttelton</surname> <given-names>O. C.</given-names></name> <name><surname>Benali</surname> <given-names>H.</given-names></name> <name><surname>Evans</surname> <given-names>A. C.</given-names></name></person-group> (<year>2010</year>). <article-title>Multi-level bootstrap analysis of stable clusters in resting-state fMRI</article-title>. <source>Neuroimage</source> <volume>51</volume>, <fpage>1126</fpage>&#x02013;<lpage>1139</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.02.082</pub-id><pub-id pub-id-type="pmid">20226257</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bellman</surname> <given-names>R. E.</given-names></name></person-group> (<year>1961</year>). <source>Adaptive Control Processes: A Guided Tour</source>. <publisher-loc>Princeton, NJ</publisher-loc>: <publisher-name>Princeton University Press</publisher-name>.</citation></ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Evolving culture versus local minima,</article-title> in <source>Growing Adaptive Machines</source>, <volume>Vol. 557</volume>, eds <person-group person-group-type="editor"><name><surname>Kowaliw</surname> <given-names>T.</given-names></name> <name><surname>Bredeche</surname> <given-names>N.</given-names></name> <name><surname>Doursat</surname> <given-names>R.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>109</fpage>&#x02013;<lpage>138</lpage>.</citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Courville</surname> <given-names>A.</given-names></name> <name><surname>Vincent</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>Representation learning: a review and new perspectives</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>35</volume>, <fpage>1798</fpage>&#x02013;<lpage>1828</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2013.50</pub-id><pub-id pub-id-type="pmid">23787338</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berk</surname> <given-names>R.</given-names></name> <name><surname>Brown</surname> <given-names>L.</given-names></name> <name><surname>Buja</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>Valid post-selection inference</article-title>. <source>Ann. Stat.</source> <volume>41</volume>, <fpage>802</fpage>&#x02013;<lpage>837</lpage>. <pub-id pub-id-type="doi">10.1214/12-AOS1077</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berkson</surname> <given-names>J.</given-names></name></person-group> (<year>1938</year>). <article-title>Some difficulties of interpretation encountered in the application of the chi-square test</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>33</volume>, <fpage>526</fpage>&#x02013;<lpage>536</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1938.10502329</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bishop</surname> <given-names>C. M.</given-names></name></person-group> (<year>2006</year>). <source>Pattern Recognition and Machine Learning</source>. <publisher-loc>Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bishop</surname> <given-names>C. M.</given-names></name> <name><surname>Lasserre</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Generative or discriminative? getting the best of both worlds</article-title>. <source>Bayesian Stat.</source> <volume>8</volume>, <fpage>3</fpage>&#x02013;<lpage>24</lpage>.</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blei</surname> <given-names>D. M.</given-names></name> <name><surname>Smyth</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Science and data science</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>114</volume>, <fpage>8689</fpage>&#x02013;<lpage>8692</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1702076114</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Box</surname> <given-names>G. E. P.</given-names></name></person-group> (<year>1976</year>). <article-title>Science and statistics</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>71</volume>, <fpage>791</fpage>&#x02013;<lpage>799</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1976.10480949</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L.</given-names></name></person-group> (<year>2001</year>). <article-title>Statistical modeling: the two cultures</article-title>. <source>Stat. Sci.</source> <volume>16</volume>, <fpage>199</fpage>&#x02013;<lpage>231</lpage>. <pub-id pub-id-type="doi">10.1214/ss/1009213726</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brodersen</surname> <given-names>K. H.</given-names></name></person-group> (<year>2009</year>). <source>Decoding Mental Activity from Neuroimaging Data &#x02014; the Science Behind Mind-Reading</source>, <volume>Vol. 4</volume>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>The New Collection</publisher-name>.</citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brodersen</surname> <given-names>K. H.</given-names></name> <name><surname>Daunizeau</surname> <given-names>J.</given-names></name> <name><surname>Mathys</surname> <given-names>C.</given-names></name> <name><surname>Chumbley</surname> <given-names>J. R.</given-names></name> <name><surname>Buhmann</surname> <given-names>J. M.</given-names></name> <name><surname>Stephan</surname> <given-names>K. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Variational Bayesian mixed-effects inference for classification studies</article-title>. <source>Neuroimage</source> <volume>76</volume>, <fpage>345</fpage>&#x02013;<lpage>361</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.03.008</pub-id><pub-id pub-id-type="pmid">23507390</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brodersen</surname> <given-names>K. H.</given-names></name> <name><surname>Haiss</surname> <given-names>F.</given-names></name> <name><surname>Ong</surname> <given-names>C. S.</given-names></name> <name><surname>Jung</surname> <given-names>F.</given-names></name> <name><surname>Tittgemeyer</surname> <given-names>M.</given-names></name> <name><surname>Buhmann</surname> <given-names>J. M.</given-names></name> <etal/></person-group>. (<year>2011a</year>). <article-title>Model-based feature construction for multivariate decoding</article-title>. <source>Neuroimage</source> <volume>56</volume>, <fpage>601</fpage>&#x02013;<lpage>615</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.04.036</pub-id><pub-id pub-id-type="pmid">20406688</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brodersen</surname> <given-names>K. H.</given-names></name> <name><surname>Schofield</surname> <given-names>T. M.</given-names></name> <name><surname>Leff</surname> <given-names>A. P.</given-names></name> <name><surname>Ong</surname> <given-names>C. S.</given-names></name> <name><surname>Lomakina</surname> <given-names>E. I.</given-names></name> <name><surname>Buhmann</surname> <given-names>J. M.</given-names></name> <etal/></person-group>. (<year>2011b</year>). <article-title>Generative embedding for model-based classification of fMRI data</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>:<fpage>e1002079</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002079</pub-id><pub-id pub-id-type="pmid">21731479</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>B&#x000FC;hlmann</surname> <given-names>P.</given-names></name> <name><surname>Van De Geer</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <source>Statistics for High-Dimensional Data: Methods, Theory and Applications.</source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer Science &#x00026; Business Media</publisher-name>.</citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burnham</surname> <given-names>K. P.</given-names></name> <name><surname>Anderson</surname> <given-names>D. R.</given-names></name></person-group> (<year>2014</year>). <article-title>P values are only an index to evidence: 20th-vs. 21st-century statistical science</article-title>. <source>Ecology</source> <volume>95</volume>, <fpage>627</fpage>&#x02013;<lpage>630</lpage>. <pub-id pub-id-type="doi">10.1890/13-1066.1</pub-id><pub-id pub-id-type="pmid">24804444</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bzdok</surname> <given-names>D.</given-names></name> <name><surname>Eickenberg</surname> <given-names>M.</given-names></name> <name><surname>Grisel</surname> <given-names>O.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Semi-supervised factored logistic regression for high-dimensional neuroimaging Data,</article-title> in <source>NIPS&#x00027;15 Proceedings of the 28th International Conference on Neural Information Processing Systems</source>, (<publisher-loc>Cambridge, MA</publisher-loc>), <fpage>3348</fpage>&#x02013;<lpage>3356</lpage>.</citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bzdok</surname> <given-names>D.</given-names></name> <name><surname>Eickenberg</surname> <given-names>M.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>Hierarchical region-network sparsity for high-dimensional inference in brain imaging,</article-title> in <source>International Conference on Information Processing in Medical Imaging (IPMI)</source> (<publisher-loc>Boone, NC</publisher-loc>).</citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bzdok</surname> <given-names>D.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Grisel</surname> <given-names>O.</given-names></name> <name><surname>Eickenberg</surname> <given-names>M.</given-names></name> <name><surname>Poupon</surname> <given-names>C.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>Formal models of the network co-occurrence underlying mental operations</article-title>. <source>PLoS Comput. Biol.</source> <volume>12</volume>:<fpage>e1004994</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1004994</pub-id><pub-id pub-id-type="pmid">27310288</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bzdok</surname> <given-names>D.</given-names></name> <name><surname>Yeo</surname> <given-names>B. T. T.</given-names></name></person-group> (<year>2017</year>). <article-title>Inference in the age of big data: future perspectives on neuroscience</article-title>. <source>Neuroimage</source> <volume>155</volume>, <fpage>549</fpage>&#x02013;<lpage>564</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2017.04.061</pub-id><pub-id pub-id-type="pmid">28456584</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Casella</surname> <given-names>G.</given-names></name> <name><surname>Berger</surname> <given-names>R. L.</given-names></name></person-group> (<year>2002</year>). <source>Statistical Inference</source>. <publisher-loc>Pacific Grove, CA</publisher-loc>: <publisher-name>Duxbury</publisher-name>.</citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chamberlin</surname> <given-names>T. C.</given-names></name></person-group> (<year>1890</year>). <article-title>The method of multiple working hypotheses</article-title>. <source>Science</source> <volume>15</volume>, <fpage>92</fpage>&#x02013;<lpage>96</lpage>.</citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chambers</surname> <given-names>J. M.</given-names></name></person-group> (<year>1993</year>). <article-title>Greater or lesser statistics: a choice for future research</article-title>. <source>Stat. Comput.</source> <volume>3</volume>, <fpage>182</fpage>&#x02013;<lpage>184</lpage>. <pub-id pub-id-type="doi">10.1007/BF00141776</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Choi</surname> <given-names>Y.</given-names></name> <name><surname>Taylor</surname> <given-names>J.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Selecting the number of principal components: estimation of the true rank of a noisy matrix</article-title>. arXiv:1410.8260.</citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chow</surname> <given-names>S. L.</given-names></name></person-group> (<year>1998</year>). <article-title>Precis of statistical significance: rationale, validity, and utility</article-title>. <source>Behav. Brain Sci.</source> <volume>21</volume>, <fpage>169</fpage>&#x02013;<lpage>194</lpage>; discussion 194&#x02013;239. <pub-id pub-id-type="doi">10.1017/S0140525X98001162</pub-id><pub-id pub-id-type="pmid">10097013</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Christoff</surname> <given-names>K.</given-names></name> <name><surname>Irving</surname> <given-names>Z. C.</given-names></name> <name><surname>Fox</surname> <given-names>K. C. R.</given-names></name> <name><surname>Spreng</surname> <given-names>R. N.</given-names></name> <name><surname>Andrews-Hanna</surname> <given-names>J. R.</given-names></name></person-group> (<year>2016</year>). <article-title>Mind-wandering as spontaneous thought: a dynamic framework</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>17</volume>, <fpage>718</fpage>&#x02013;<lpage>731</lpage>. <pub-id pub-id-type="doi">10.1038/nrn.2016.113</pub-id><pub-id pub-id-type="pmid">27654862</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chumbley</surname> <given-names>J. R.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2009</year>). <article-title>False discovery rate revisited: FDR and topological inference using Gaussian random fields</article-title>. <source>Neuroimage</source> <volume>44</volume>, <fpage>62</fpage>&#x02013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2008.05.021</pub-id><pub-id pub-id-type="pmid">18603449</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cleveland</surname> <given-names>W. S.</given-names></name></person-group> (<year>2001</year>). <article-title>Data science: an action plan for expanding the technical areas of the field of statistics</article-title>. <source>Int. Stat. Rev.</source> <volume>69</volume>, <fpage>21</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1111/j.1751-5823.2001.tb00477.x</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Coase</surname> <given-names>R. H.</given-names></name></person-group> (<year>1982</year>). <source>How Should Economists Choose? The G. Warren Nutter Lectures in Political Economy.</source> <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>American Enterprise Institute for Public Policy Research</publisher-name>.</citation></ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1977</year>). <source>Statistical Power Analysis for the Behavioral Sciences</source>. <publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates, Inc</publisher-name>.</citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1990</year>). <article-title>Things I have learned (so far)</article-title>. <source>Am. Psychol.</source> <volume>45</volume>:<fpage>1304</fpage>. <pub-id pub-id-type="doi">10.1037/0003-066X.45.12.1304</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1992</year>). <article-title>A power primer</article-title>. <source>Psychol. Bull.</source> <volume>112</volume>:<fpage>155</fpage>. <pub-id pub-id-type="doi">10.1037/0033-2909.112.1.155</pub-id><pub-id pub-id-type="pmid">19565683</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1994</year>). <article-title>The Earth Is Round (p &#x0003C; 0.05)</article-title>. <source>Am. Psychol.</source> <volume>49</volume>, <fpage>997</fpage>&#x02013;<lpage>1003</lpage>. <pub-id pub-id-type="doi">10.1037/0003-066X.49.12.997</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><collab>Committee on the Analysis of Massive Data, Committee on Applied and Theoretical Statistics, Board on Mathematical Sciences and their Applications, Division on Engineering and Physical Sciences, and National Research Council.</collab></person-group> (<year>2013</year>). <source>Frontiers in Massive Data Analysis</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cowles</surname> <given-names>M.</given-names></name> <name><surname>Davis</surname> <given-names>C.</given-names></name></person-group> (<year>1982</year>). <article-title>On the origins of the.05 level of statistical significance</article-title>. <source>Am. Psychol.</source> <volume>37</volume>, <fpage>553</fpage>&#x02013;<lpage>558</lpage>.</citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cox</surname> <given-names>D. D.</given-names></name> <name><surname>Dean</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural networks and neuroscience-inspired computer vision</article-title>. <source>Curr. Biol.</source> <volume>24</volume>, <fpage>R921</fpage>&#x02013;<lpage>R929</lpage>. <pub-id pub-id-type="doi">10.1016/j.cub.2014.08.026</pub-id><pub-id pub-id-type="pmid">25247371</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cox</surname> <given-names>D. D.</given-names></name> <name><surname>Savoy</surname> <given-names>R. L.</given-names></name></person-group> (<year>2003</year>). <article-title>Functional magnetic resonance imaging (fMRI) &#x0201C;brain reading&#x0201D;: detecting and classifying distributed patterns of fMRI activity in human visual cortex</article-title>. <source>Neuroimage</source> <volume>19</volume>, <fpage>261</fpage>&#x02013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1016/S1053-8119(03)00049-1</pub-id><pub-id pub-id-type="pmid">12814577</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cox</surname> <given-names>D. R.</given-names></name></person-group> (<year>1975</year>). <article-title>A note on data-splitting for the evaluation of significance levels</article-title>. <source>Biometrika</source> <volume>62</volume>, <fpage>441</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/62.2.441</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cumming</surname> <given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>Inference by eye: reading the overlap of independent confidence intervals</article-title>. <source>Stat. Med.</source> <volume>28</volume>, <fpage>205</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1002/sim.3471</pub-id><pub-id pub-id-type="pmid">18991332</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Davatzikos</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <article-title>Why voxel-based morphometric analysis should be used with great caution when characterizing group differences</article-title>. <source>Neuroimage</source> <volume>23</volume>, <fpage>17</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2004.05.010</pub-id><pub-id pub-id-type="pmid">15325347</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Davis</surname> <given-names>J.</given-names></name> <name><surname>Goadrich</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <source>The relationship between Precision-Recall and ROC curves</source>,&#x0201D; in <italic>Proceedings of the 23rd International Conference on Machine Learning</italic> (<publisher-loc>Pittsburgh, PA</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>233</fpage>&#x02013;<lpage>240</lpage>.</citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Brebisson</surname> <given-names>A.</given-names></name> <name><surname>Montana</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks for anatomical brain segmentation</article-title>. arXiv:1502.02445.</citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dem&#x00161;ar</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Statistical comparisons of classifiers over multiple data sets</article-title>. <source>J. Mach. Learn. Res.</source> <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>30</lpage>.</citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Derrfuss</surname> <given-names>J.</given-names></name> <name><surname>Mar</surname> <given-names>R. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Lost in localization: the need for a universal coordinate database</article-title>. <source>Neuroimage</source> <volume>48</volume>, <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2009.01.053</pub-id><pub-id pub-id-type="pmid">19457374</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de-Wit</surname> <given-names>L.</given-names></name> <name><surname>Alexander</surname> <given-names>D.</given-names></name> <name><surname>Ekroll</surname> <given-names>V.</given-names></name> <name><surname>Wagemans</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Is neuroimaging measuring information in the brain?</article-title> <source>Psychon. Bull. Rev.</source> <volume>23</volume>, <fpage>1415</fpage>&#x02013;<lpage>1428</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-016-1002-0</pub-id><pub-id pub-id-type="pmid">26833316</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Domingos</surname> <given-names>P.</given-names></name></person-group> (<year>2012</year>). <article-title>A few useful things to know about machine learning</article-title>. <source>Commun. ACM</source> <volume>55</volume>, <fpage>78</fpage>&#x02013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1145/2347736.2347755</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Donoho</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>50 years of data science,</article-title> in <source>Based on a Presentation at the Tukey Centennial Workshop</source> (<publisher-loc>Princeton</publisher-loc>: <publisher-name>NJ</publisher-name>).</citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Efron</surname> <given-names>B.</given-names></name></person-group> (<year>1979</year>). <article-title>Bootstrap methods: another look at the jackknife</article-title>. <source>Ann. Stat.</source> <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1214/aos/1176344552</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Efron</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <source>Large-Scale Inference: Empirical Bayes Methods for Estimation, Testing, and Prediction.</source> <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Efron</surname> <given-names>B.</given-names></name> <name><surname>Hastie</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <source>Computer-Age Statistical Inference</source>. <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Efron</surname> <given-names>B.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R. J.</given-names></name></person-group> (<year>1991</year>). <article-title>Statistical data analysis in the computer age</article-title>. <source>Science</source> <volume>253</volume>, <fpage>390</fpage>&#x02013;<lpage>395</lpage>. <pub-id pub-id-type="doi">10.1126/science.253.5018.390</pub-id><pub-id pub-id-type="pmid">17746394</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Efron</surname> <given-names>B.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R. J.</given-names></name></person-group> (<year>1994</year>). <source>An Introduction to the Bootstrap</source>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>CRC press</publisher-name>.</citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eickhoff</surname> <given-names>S. B.</given-names></name> <name><surname>Bzdok</surname> <given-names>D.</given-names></name> <name><surname>Laird</surname> <given-names>A. R.</given-names></name> <name><surname>Roski</surname> <given-names>C.</given-names></name> <name><surname>Caspers</surname> <given-names>S.</given-names></name> <name><surname>Zilles</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Co-activation patterns distinguish cortical modules, their connectivity and functional differentiation</article-title>. <source>Neuroimage</source> <volume>57</volume>, <fpage>938</fpage>&#x02013;<lpage>949</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2011.05.021</pub-id><pub-id pub-id-type="pmid">21609770</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eickhoff</surname> <given-names>S. B.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Bzdok</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>Connectivity-based parcellation: critique and implications</article-title>. <source>Hum. Brain Mapp.</source> <volume>36</volume>, <fpage>4771</fpage>&#x02013;<lpage>4792</lpage> <pub-id pub-id-type="doi">10.1002/hbm.22933</pub-id><pub-id pub-id-type="pmid">26409749</pub-id></citation></ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eickhoff</surname> <given-names>S.</given-names></name> <name><surname>Turner</surname> <given-names>J. A.</given-names></name> <name><surname>Nichols</surname> <given-names>T. E.</given-names></name> <name><surname>Van Horn</surname> <given-names>J. D.</given-names></name></person-group> (<year>2016</year>). <article-title>Sharing the wealth: neuroimaging data repositories</article-title>. <source>Neuroimage</source> <volume>124</volume>, <fpage>1065</fpage>&#x02013;<lpage>1068</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2015.10.079</pub-id><pub-id pub-id-type="pmid">26574120</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Estes</surname> <given-names>W. K.</given-names></name></person-group> (<year>1997</year>). <article-title>On the communication of information by displays of standard errors and confidence intervals</article-title>. <source>Psychon. Bull. Rev.</source> <volume>4</volume>, <fpage>330</fpage>&#x02013;<lpage>341</lpage>. <pub-id pub-id-type="doi">10.3758/BF03210790</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Everitt</surname> <given-names>B. S.</given-names></name></person-group> (<year>1979</year>). <article-title>Unresolved problems in cluster analysis</article-title>. <source>Biometrics</source> <volume>35</volume>, <fpage>169</fpage>&#x02013;<lpage>181</lpage>. <pub-id pub-id-type="doi">10.2307/2529943</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferguson</surname> <given-names>C. J.</given-names></name></person-group> (<year>2009</year>). <article-title>An effect size primer: a guide for clinicians and researchers</article-title>. <source>Prof. Psychol.</source> <volume>40</volume>:<fpage>532</fpage>. <pub-id pub-id-type="doi">10.1037/a0015808</pub-id></citation></ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Feyerabend</surname> <given-names>P.</given-names></name></person-group> (<year>1975</year>). <source>Against Method: Outline of an Anarchist Theory of Knowledge</source>. <publisher-loc>London</publisher-loc>: <publisher-name>New Left Books</publisher-name>.</citation></ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>R. A.</given-names></name></person-group> (<year>1925</year>). <source>Statistical Methods of Research Workers</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Oliver and Boyd</publisher-name>.</citation></ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>R. A.</given-names></name></person-group> (<year>1935</year>). <source>The Design of Experiments.</source> <publisher-loc>Edinburgh</publisher-loc>: <publisher-name>Oliver and Boyd</publisher-name>. <pub-id pub-id-type="pmid">13823235</pub-id></citation></ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>R. A.</given-names></name> <name><surname>Mackenzie</surname> <given-names>W. A.</given-names></name></person-group> (<year>1923</year>). <article-title>Studies in crop variation. II. The manurial response of different potato varieties</article-title>. <source>J. Agric. Sci.</source> <volume>13</volume>, <fpage>311</fpage>&#x02013;<lpage>320</lpage>. <pub-id pub-id-type="doi">10.1017/S0021859600003592</pub-id></citation></ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fithian</surname> <given-names>W.</given-names></name> <name><surname>Sun</surname> <given-names>D.</given-names></name> <name><surname>Taylor</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Optimal inference after model selection</article-title>. arXiv:1410.2597.</citation></ref>
<ref id="B74">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fleck</surname> <given-names>L.</given-names></name> <name><surname>Sch&#x000E4;fer</surname> <given-names>L.</given-names></name> <name><surname>Schnelle</surname> <given-names>T.</given-names></name></person-group> (<year>1935</year>). <source>Entstehung und Entwicklung einer Wissenschaftlichen Tatsache</source>. <publisher-loc>Basel</publisher-loc>: <publisher-name>Schwabe</publisher-name>.</citation></ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fox</surname> <given-names>P. T.</given-names></name> <name><surname>Lancaster</surname> <given-names>J. L.</given-names></name> <name><surname>Laird</surname> <given-names>A. R.</given-names></name> <name><surname>Eickhoff</surname> <given-names>S. B.</given-names></name></person-group> (<year>2014</year>). <article-title>Meta-analysis in human neuroimaging: computational modeling of large-scale databases</article-title>. <source>Annu. Rev. Neurosci.</source> <volume>37</volume>, <fpage>409</fpage>&#x02013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-neuro-062012-170320</pub-id><pub-id pub-id-type="pmid">25032500</pub-id></citation></ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frackowiak</surname> <given-names>R.</given-names></name> <name><surname>Markram</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>The future of human cerebral cartography: a novel approach</article-title>. <source>Philos. Trans. R. Soc. Lond. B Biol. Sci.</source> <volume>370</volume>:<fpage>20140171</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2014.0171</pub-id><pub-id pub-id-type="pmid">25823868</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Freedman</surname> <given-names>D. A.</given-names></name></person-group> (<year>1983</year>). <article-title>A note on screening regression equations</article-title>. <source>Am. Stat.</source> <volume>37</volume>, <fpage>152</fpage>&#x02013;<lpage>155</lpage>.</citation></ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friedman</surname> <given-names>J. H.</given-names></name></person-group> (<year>1998</year>). <article-title>Data mining and statistics: what&#x00027;s the connection?</article-title> <source>Comput. Sci. Stat.</source> <volume>29</volume>, <fpage>3</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friedman</surname> <given-names>J. H.</given-names></name></person-group> (<year>2001</year>). <article-title>The role of statistics in the data revolution?</article-title> <source>Int. Stat. Rev.</source> <volume>69</volume>, <fpage>5</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friman</surname> <given-names>O.</given-names></name> <name><surname>Cedefamn</surname> <given-names>J.</given-names></name> <name><surname>Lundberg</surname> <given-names>P.</given-names></name> <name><surname>Borga</surname> <given-names>M.</given-names></name> <name><surname>Knutsson</surname> <given-names>H.</given-names></name></person-group> (<year>2001</year>). <article-title>Detection of neural activity in functional MRI using canonical correlation analysis</article-title>. <source>Magn. Reson. Med.</source> <volume>45</volume>, <fpage>323</fpage>&#x02013;<lpage>330</lpage>. <pub-id pub-id-type="doi">10.1002/1522-2594(200102)45:2&#x0003C;323::AID-MRM1041&#x0003E;3.0.CO;2-&#x00023;</pub-id><pub-id pub-id-type="pmid">11180440</pub-id></citation></ref>
<ref id="B81">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2006</year>). <source>Statistical Parametric Mapping: The Analysis of Functional Brain Images</source>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Academic Press</publisher-name>.</citation></ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Modalities, modes, and models in functional neuroimaging</article-title>. <source>Science</source> <volume>326</volume>, <fpage>399</fpage>&#x02013;<lpage>403</lpage>. <pub-id pub-id-type="doi">10.1126/science.1174521</pub-id><pub-id pub-id-type="pmid">19833961</pub-id></citation></ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name></person-group> (<year>2012</year>). <article-title>Ten ironic rules for non-statistical reviewers</article-title>. <source>Neuroimage</source> <volume>61</volume>, <fpage>1300</fpage>&#x02013;<lpage>1310</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.04.018</pub-id><pub-id pub-id-type="pmid">22521475</pub-id></citation></ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Chu</surname> <given-names>C.</given-names></name> <name><surname>Mourao-Miranda</surname> <given-names>J.</given-names></name> <name><surname>Hulme</surname> <given-names>O.</given-names></name> <name><surname>Rees</surname> <given-names>G.</given-names></name> <name><surname>Penny</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Bayesian decoding of brain images</article-title>. <source>Neuroimage</source> <volume>39</volume>, <fpage>181</fpage>&#x02013;<lpage>205</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2007.08.013</pub-id><pub-id pub-id-type="pmid">17919928</pub-id></citation></ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Holmes</surname> <given-names>A. P.</given-names></name> <name><surname>Worsley</surname> <given-names>K. J.</given-names></name> <name><surname>Poline</surname> <given-names>J. P.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name> <name><surname>Frackowiak</surname> <given-names>R. S.</given-names></name></person-group> (<year>1994</year>). <article-title>Statistical parametric maps in functional imaging: a general linear approach</article-title>. <source>Hum. Brain Mapp.</source> <volume>2</volume>, <fpage>189</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.460020402</pub-id></citation></ref>
<ref id="B86">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Liddle</surname> <given-names>P. F.</given-names></name> <name><surname>Frith</surname> <given-names>C. D.</given-names></name> <name><surname>Hirsch</surname> <given-names>S. R.</given-names></name> <name><surname>Frackowiak</surname> <given-names>R. S. J.</given-names></name></person-group> (<year>1992</year>). <article-title>The left medial temporal region and schizophrenia</article-title>. <source>Brain</source> <volume>115</volume>, <fpage>367</fpage>&#x02013;<lpage>382</lpage>. <pub-id pub-id-type="doi">10.1093/brain/115.2.367</pub-id><pub-id pub-id-type="pmid">1606474</pub-id></citation></ref>
<ref id="B87">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Price</surname> <given-names>C. J.</given-names></name> <name><surname>Fletcher</surname> <given-names>P.</given-names></name> <name><surname>Moore</surname> <given-names>C.</given-names></name> <name><surname>Frackowiak</surname> <given-names>R. S. J.</given-names></name> <name><surname>Dolan</surname> <given-names>R. J.</given-names></name></person-group> (<year>1996</year>). <article-title>The trouble with cognitive subtraction</article-title>. <source>Neuroimage</source> <volume>4</volume>, <fpage>97</fpage>&#x02013;<lpage>104</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.1996.0033</pub-id><pub-id pub-id-type="pmid">9345501</pub-id></citation></ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gabrieli</surname> <given-names>J. D.</given-names></name> <name><surname>Ghosh</surname> <given-names>S. S.</given-names></name> <name><surname>Whitfield-Gabrieli</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Prediction as a humanitarian and pragmatic contribution from human cognitive neuroscience</article-title>. <source>Neuron</source> <volume>85</volume>, <fpage>11</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2014.10.047</pub-id><pub-id pub-id-type="pmid">25569345</pub-id></citation></ref>
<ref id="B89">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Genovese</surname> <given-names>C. R.</given-names></name> <name><surname>Lazar</surname> <given-names>N. A.</given-names></name> <name><surname>Nichols</surname> <given-names>T.</given-names></name></person-group> (<year>2002</year>). <article-title>Thresholding of statistical maps in functional neuroimaging using the false discovery rate</article-title>. <source>Neuroimage</source> <volume>15</volume>, <fpage>870</fpage>&#x02013;<lpage>878</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.2001.1037</pub-id><pub-id pub-id-type="pmid">11906227</pub-id></citation></ref>
<ref id="B90">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ghahramani</surname> <given-names>Z.</given-names></name></person-group> (<year>2004</year>). <article-title>Unsupervised learning,</article-title> in <source>Advanced Lectures on Machine Learning</source>, eds <person-group person-group-type="editor"><name><surname>Bousquet</surname> <given-names>O.</given-names></name> <name><surname>von Luxburg</surname> <given-names>U.</given-names></name> <name><surname>R&#x000E4;tsch</surname> <given-names>G.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>72</fpage>&#x02013;<lpage>112</lpage>.</citation></ref>
<ref id="B91">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ghahramani</surname> <given-names>Z.</given-names></name></person-group> (<year>2015</year>). <article-title>Probabilistic machine learning and artificial intelligence</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>452</fpage>&#x02013;<lpage>459</lpage>. <pub-id pub-id-type="doi">10.1038/nature14541</pub-id><pub-id pub-id-type="pmid">26017444</pub-id></citation></ref>
<ref id="B92">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gigerenzer</surname> <given-names>G.</given-names></name></person-group> (<year>1993</year>). <article-title>The superego, the ego, and the id in statistical reasoning,</article-title> in <source>A Handbook for Data Analysis in the Behavioral Sciences: Methodological issues</source>, eds <person-group person-group-type="editor"><name><surname>Keren</surname> <given-names>G.</given-names></name> <name><surname>Lewis</surname> <given-names>C.</given-names></name></person-group> (<publisher-loc>Hillslade, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates</publisher-name>), <fpage>311</fpage>&#x02013;<lpage>339</lpage>.</citation></ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gigerenzer</surname> <given-names>G.</given-names></name></person-group> (<year>2004</year>). <article-title>Mindless statistics</article-title>. <source>J. Soc. Econ.</source> <volume>33</volume>, <fpage>587</fpage>&#x02013;<lpage>606</lpage>. <pub-id pub-id-type="doi">10.1016/j.socec.2004.09.033</pub-id></citation></ref>
<ref id="B94">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gigerenzer</surname> <given-names>G.</given-names></name> <name><surname>Murray</surname> <given-names>D. J.</given-names></name></person-group> (<year>1987</year>). <source>Cognition as Intuitive Statistics</source>. <publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Erlbaum</publisher-name>.</citation></ref>
<ref id="B95">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Giraud</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <source>Introduction to High-Dimensional Statistics</source>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</citation></ref>
<ref id="B96">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gl&#x000E4;scher</surname> <given-names>J.</given-names></name> <name><surname>Adolphs</surname> <given-names>R.</given-names></name> <name><surname>Damasio</surname> <given-names>H.</given-names></name> <name><surname>Bechara</surname> <given-names>A.</given-names></name> <name><surname>Rudrauf</surname> <given-names>D.</given-names></name> <name><surname>Calamia</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Lesion mapping of cognitive control and value-based decision making in the prefrontal cortex</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>109</volume>, <fpage>14681</fpage>&#x02013;<lpage>14686</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1206608109</pub-id><pub-id pub-id-type="pmid">22908286</pub-id></citation></ref>
<ref id="B97">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Golland</surname> <given-names>P.</given-names></name> <name><surname>Fischl</surname> <given-names>B.</given-names></name></person-group> (<year>2003</year>). <article-title>Permutation tests for classification: towards statistical significance in image-based studies,</article-title> in <source>Information Processing in Medical Imaging</source>, eds <person-group person-group-type="editor"><name><surname>Taylor</surname> <given-names>C.</given-names></name> <name><surname>Noble</surname> <given-names>J. A.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>330</fpage>&#x02013;<lpage>341</lpage>. <pub-id pub-id-type="pmid">15344469</pub-id></citation></ref>
<ref id="B98">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Goodfellow</surname> <given-names>I. J.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Courville</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <source>Deep Learning</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B99">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goodman</surname> <given-names>S. N.</given-names></name></person-group> (<year>1999</year>). <article-title>Toward evidence-based medical statistics. 1: the P value fallacy</article-title>. <source>Ann. Int. Med.</source> <volume>130</volume>, <fpage>995</fpage>&#x02013;<lpage>1004</lpage>. <pub-id pub-id-type="doi">10.7326/0003-4819-130-12-199906150-00008</pub-id><pub-id pub-id-type="pmid">10383371</pub-id></citation></ref>
<ref id="B100">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grady</surname> <given-names>C. L.</given-names></name> <name><surname>Haxby</surname> <given-names>J. V.</given-names></name> <name><surname>Schapiro</surname> <given-names>M. B.</given-names></name> <name><surname>Gonzalez-Aviles</surname> <given-names>A.</given-names></name> <name><surname>Kumar</surname> <given-names>A.</given-names></name> <name><surname>Ball</surname> <given-names>M. J.</given-names></name> <etal/></person-group>. (<year>1990</year>). <article-title>Subgroups in dementia of the Alzheimer type identified using positron emission tomography</article-title>. <source>J. Neuropsychiatry Clin. Neurosci.</source> <volume>2</volume>, <fpage>373</fpage>&#x02013;<lpage>384</lpage>. <pub-id pub-id-type="doi">10.1176/jnp.2.4.373</pub-id><pub-id pub-id-type="pmid">2136389</pub-id></citation></ref>
<ref id="B101">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Greenwald</surname> <given-names>A. G.</given-names></name></person-group> (<year>2012</year>). <article-title>There is nothing so theoretical as a good method</article-title>. <source>Perspect. Psychol. Sci.</source> <volume>7</volume>, <fpage>99</fpage>&#x02013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1177/1745691611434210</pub-id><pub-id pub-id-type="pmid">26168438</pub-id></citation></ref>
<ref id="B102">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x000FC;&#x000E7;l&#x000FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream</article-title>. <source>J. Neurosci.</source> <volume>35</volume>, <fpage>10005</fpage>&#x02013;<lpage>10014</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5023-14.2015</pub-id><pub-id pub-id-type="pmid">26157000</pub-id></citation></ref>
<ref id="B103">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guyon</surname> <given-names>I.</given-names></name> <name><surname>Elisseeff</surname> <given-names>A.</given-names></name></person-group> (<year>2003</year>). <article-title>An introduction to variable and feature selection</article-title>. <source>J. Mach. Learn. Res.</source> <volume>3</volume>, <fpage>1157</fpage>&#x02013;<lpage>1182</lpage>.</citation></ref>
<ref id="B104">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guyon</surname> <given-names>I.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name> <name><surname>Barnhill</surname> <given-names>S.</given-names></name> <name><surname>Vapnik</surname> <given-names>V.</given-names></name></person-group> (<year>2002</year>). <article-title>Gene selection for cancer classification using support vector machines</article-title>. <source>Mach. Learn.</source> <volume>46</volume>, <fpage>389</fpage>&#x02013;<lpage>422</lpage>. <pub-id pub-id-type="doi">10.1023/A:1012487302797</pub-id></citation></ref>
<ref id="B105">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Halkidi</surname> <given-names>M.</given-names></name> <name><surname>Batistakis</surname> <given-names>Y.</given-names></name> <name><surname>Vazirgiannis</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>On clustering validation techniques</article-title>. <source>J. Intell. Inf. Syst.</source> <volume>17</volume>, <fpage>107</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1023/A:1012801612483</pub-id></citation></ref>
<ref id="B106">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hall</surname> <given-names>E. T.</given-names></name></person-group> (<year>1976</year>). <source>Beyond Culture</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Anchor Books</publisher-name>.</citation></ref>
<ref id="B107">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Handl</surname> <given-names>J.</given-names></name> <name><surname>Knowles</surname> <given-names>J.</given-names></name> <name><surname>Kell</surname> <given-names>D. B.</given-names></name></person-group> (<year>2005</year>). <article-title>Computational cluster validation in post-genomic data analysis</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>3201</fpage>&#x02013;<lpage>3212</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti517</pub-id><pub-id pub-id-type="pmid">15914541</pub-id></citation></ref>
<ref id="B108">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Hanke</surname> <given-names>M.</given-names></name> <name><surname>Halchenko</surname> <given-names>Y. O.</given-names></name> <name><surname>Oosterhof</surname> <given-names>N. N.</given-names></name></person-group> (<year>2015</year>). <source>PyMVPA Manuel</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.pymvpa.org/">http://www.pymvpa.org/</ext-link></citation></ref>
<ref id="B109">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanke</surname> <given-names>M.</given-names></name> <name><surname>Halchenko</surname> <given-names>Y. O.</given-names></name> <name><surname>Sederberg</surname> <given-names>P. B.</given-names></name> <name><surname>Hanson</surname> <given-names>S. J.</given-names></name> <name><surname>Haxby</surname> <given-names>J. V.</given-names></name> <name><surname>Pollmann</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>PyMVPA: a python toolbox for multivariate pattern analysis of fMRI data</article-title>. <source>Neuroinformatics</source> <volume>7</volume>, <fpage>37</fpage>&#x02013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1007/s12021-008-9041-y</pub-id><pub-id pub-id-type="pmid">19184561</pub-id></citation></ref>
<ref id="B110">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanson</surname> <given-names>S. J.</given-names></name> <name><surname>Halchenko</surname> <given-names>Y. O.</given-names></name></person-group> (<year>2008</year>). <article-title>Brain reading using full brain support vectormachines for object recognition: there is no &#x0201C;Face&#x0201D; Identification Area</article-title>. <source>Neural Comput.</source> <volume>20</volume>, <fpage>486</fpage>&#x02013;<lpage>503</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2007.09-06-340</pub-id><pub-id pub-id-type="pmid">18047411</pub-id></citation></ref>
<ref id="B111">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hanson</surname> <given-names>S. J.</given-names></name> <name><surname>Matsuka</surname> <given-names>T.</given-names></name> <name><surname>Haxby</surname> <given-names>J. V.</given-names></name></person-group> (<year>2004</year>). <article-title>Combinatorial codes in ventral temporal lobe for object recognition: Haxby (2001) revisited: is there a &#x0201C;face&#x0201D; area?</article-title> <source>Neuroimage</source> <volume>23</volume>, <fpage>156</fpage>&#x02013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2004.05.020</pub-id><pub-id pub-id-type="pmid">15325362</pub-id></citation></ref>
<ref id="B112">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name> <name><surname>Friedman</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <source>The Elements of Statistical Learning. Springer Series in Statistics</source>. <publisher-loc>Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B113">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name> <name><surname>Wainwright</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <source>Statistical Learning with Sparsity. The Lasso and Generalizations.</source> <publisher-loc>London, UK</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</citation></ref>
<ref id="B114">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haxby</surname> <given-names>J. V.</given-names></name></person-group> (<year>2012</year>). <article-title>Multivariate pattern analysis of fMRI: the early beginnings</article-title>. <source>Neuroimage</source> <volume>62</volume>, <fpage>852</fpage>&#x02013;<lpage>855</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.03.016</pub-id><pub-id pub-id-type="pmid">22425670</pub-id></citation></ref>
<ref id="B115">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haxby</surname> <given-names>J. V.</given-names></name> <name><surname>Gobbini</surname> <given-names>M. I.</given-names></name> <name><surname>Furey</surname> <given-names>M. L.</given-names></name> <name><surname>Ishai</surname> <given-names>A.</given-names></name> <name><surname>Schouten</surname> <given-names>J. L.</given-names></name> <name><surname>Pietrini</surname> <given-names>P.</given-names></name></person-group> (<year>2001</year>). <article-title>Distributed and overlapping representations of faces and objects in ventral temporal cortex</article-title>. <source>Science</source> <volume>293</volume>, <fpage>2425</fpage>&#x02013;<lpage>2430</lpage>. <pub-id pub-id-type="doi">10.1126/science.1063736</pub-id><pub-id pub-id-type="pmid">11577229</pub-id></citation></ref>
<ref id="B116">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haynes</surname> <given-names>J.-D.</given-names></name></person-group> (<year>2015</year>). <article-title>A primer on pattern-based approaches to fMRI: principles, pitfalls, and perspectives</article-title>. <source>Neuron</source> <volume>87</volume>, <fpage>257</fpage>&#x02013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.05.025</pub-id><pub-id pub-id-type="pmid">26182413</pub-id></citation></ref>
<ref id="B117">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haynes</surname> <given-names>J. D.</given-names></name> <name><surname>Rees</surname> <given-names>G.</given-names></name></person-group> (<year>2005</year>). <article-title>Predicting the orientation of invisible stimuli from acitvity in human primary visual cortex</article-title>. <source>Nat. Neurosci.</source> <volume>8</volume>, <fpage>686</fpage>&#x02013;<lpage>691</lpage>. <pub-id pub-id-type="doi">10.1038/nn1445</pub-id><pub-id pub-id-type="pmid">15852013</pub-id></citation></ref>
<ref id="B118">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haynes</surname> <given-names>J. D.</given-names></name> <name><surname>Rees</surname> <given-names>G.</given-names></name></person-group> (<year>2006</year>). <article-title>Decoding mental states from brain activity in humans</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>7</volume>, <fpage>523</fpage>&#x02013;<lpage>534</lpage>. <pub-id pub-id-type="doi">10.1038/nrn1931</pub-id><pub-id pub-id-type="pmid">16791142</pub-id></citation></ref>
<ref id="B119">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henke</surname> <given-names>N.</given-names></name> <name><surname>Bughin</surname> <given-names>J.</given-names></name> <name><surname>Chui</surname> <given-names>M.</given-names></name> <name><surname>Manyika</surname> <given-names>J.</given-names></name> <name><surname>Saleh</surname> <given-names>T.</given-names></name> <name><surname>Wiseman</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2016</year>). <source>The Age of Analytics: Competing in a data-driven world</source>. Technical Report, McKinsey Global Institute.</citation></ref>
<ref id="B120">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name></person-group> (<year>2006</year>). <article-title>Reducing the dimensionality of data with neural networks</article-title>. <source>Science</source> <volume>313</volume>, <fpage>504</fpage>&#x02013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127647</pub-id><pub-id pub-id-type="pmid">16873662</pub-id></citation></ref>
<ref id="B121">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ioannidis</surname> <given-names>J. P.</given-names></name></person-group> (<year>2005</year>). <article-title>Why most published research findings are false</article-title>. <source>PLoS Med.</source> <volume>2</volume>:<fpage>e124</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pmed.0020124</pub-id></citation></ref>
<ref id="B122">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jain</surname> <given-names>A. K.</given-names></name> <name><surname>Murty</surname> <given-names>M. N.</given-names></name> <name><surname>Flynn</surname> <given-names>P. J.</given-names></name></person-group> (<year>1999</year>). <article-title>Data clustering: a review</article-title>. <source>ACN Comput. Surv.</source> <volume>31</volume>, <fpage>264</fpage>&#x02013;<lpage>323</lpage>. <pub-id pub-id-type="doi">10.1145/331499.331504</pub-id></citation></ref>
<ref id="B123">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jamalabadi</surname> <given-names>H.</given-names></name> <name><surname>Alizadeh</surname> <given-names>S.</given-names></name> <name><surname>Sch&#x000F6;nauer</surname> <given-names>M.</given-names></name> <name><surname>Leibold</surname> <given-names>C.</given-names></name> <name><surname>Gais</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Classification based hypothesis testing in neuroscience: below-chance level classification rates and overlooked statistical properties of linear parametric classifiers</article-title>. <source>Hum. Brain Mapp.</source> <volume>37</volume>, <fpage>1842</fpage>&#x02013;<lpage>1855</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.23140</pub-id><pub-id pub-id-type="pmid">27015748</pub-id></citation></ref>
<ref id="B124">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>James</surname> <given-names>G.</given-names></name> <name><surname>Witten</surname> <given-names>D.</given-names></name> <name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <source>An Introduction to Statistical Learning</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B125">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jenatton</surname> <given-names>R.</given-names></name> <name><surname>Audibert</surname> <given-names>J.-Y.</given-names></name> <name><surname>Bach</surname> <given-names>F.</given-names></name></person-group> (<year>2011</year>). <article-title>Structured variable selection with sparsity-inducing norms</article-title>. <source>J. Mach. Learn. Res.</source> <volume>12</volume>, <fpage>2777</fpage>&#x02013;<lpage>2824</lpage>.</citation></ref>
<ref id="B126">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jordan</surname> <given-names>M. I.</given-names></name> <name><surname>Mitchell</surname> <given-names>T. M.</given-names></name></person-group> (<year>2015</year>). <article-title>Machine learning: trends, perspectives, and prospects</article-title>. <source>Science</source> <volume>349</volume>, <fpage>255</fpage>&#x02013;<lpage>260</lpage>. <pub-id pub-id-type="doi">10.1126/science.aaa8415</pub-id><pub-id pub-id-type="pmid">26185243</pub-id></citation></ref>
<ref id="B127">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kamitani</surname> <given-names>Y.</given-names></name> <name><surname>Sawahata</surname> <given-names>Y.</given-names></name></person-group> (<year>2010</year>). <article-title>Spatial smoothing hurts localization but not information: pitfalls for brain mappers</article-title>. <source>Neuroimage</source> <volume>49</volume>, <fpage>1949</fpage>&#x02013;<lpage>1952</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2009.06.040</pub-id></citation></ref>
<ref id="B128">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kamitani</surname> <given-names>Y.</given-names></name> <name><surname>Tong</surname> <given-names>F.</given-names></name></person-group> (<year>2005</year>). <article-title>Decoding the visual and subjective contents of the human brain</article-title>. <source>Nat. Neurosci.</source> <volume>8</volume>, <fpage>679</fpage>&#x02013;<lpage>685</lpage>. <pub-id pub-id-type="doi">10.1038/nn1444</pub-id><pub-id pub-id-type="pmid">15852014</pub-id></citation></ref>
<ref id="B129">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kandel</surname> <given-names>E. R.</given-names></name> <name><surname>Markram</surname> <given-names>H.</given-names></name> <name><surname>Matthews</surname> <given-names>P. M.</given-names></name> <name><surname>Yuste</surname> <given-names>R.</given-names></name> <name><surname>Koch</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>Neuroscience thinks big (and collaboratively)</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>14</volume>, <fpage>659</fpage>&#x02013;<lpage>664</lpage>. <pub-id pub-id-type="doi">10.1038/nrn3578</pub-id><pub-id pub-id-type="pmid">23958663</pub-id></citation></ref>
<ref id="B130">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kelley</surname> <given-names>K.</given-names></name> <name><surname>Preacher</surname> <given-names>K. J.</given-names></name></person-group> (<year>2012</year>). <article-title>On effect size</article-title>. <source>Psychol. Methods</source> <volume>17</volume>, <fpage>137</fpage>. <pub-id pub-id-type="doi">10.1037/a0028086</pub-id></citation></ref>
<ref id="B131">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>King</surname> <given-names>J. R.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Characterizing the dynamics of mental representations: the temporal generalization method</article-title>. <source>Trends Cogn. Sci.</source> <volume>18</volume>, <fpage>203</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2014.01.002</pub-id><pub-id pub-id-type="pmid">24593982</pub-id></citation></ref>
<ref id="B132">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Knops</surname> <given-names>A.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Hubbard</surname> <given-names>E. M.</given-names></name> <name><surname>Michel</surname> <given-names>V.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Recruitment of an area involved in eye movements during mental arithmetic</article-title>. <source>Science</source> <volume>324</volume>, <fpage>1583</fpage>&#x02013;<lpage>1585</lpage>. <pub-id pub-id-type="doi">10.1126/science.1171599</pub-id><pub-id pub-id-type="pmid">19423779</pub-id></citation></ref>
<ref id="B133">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2011</year>). <article-title>Pattern-information analysis: from stimulus decoding to computational-model testing</article-title>. <source>Neuroimage</source> <volume>56</volume>, <fpage>411</fpage>&#x02013;<lpage>421</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2011.01.061</pub-id><pub-id pub-id-type="pmid">21281719</pub-id></citation></ref>
<ref id="B134">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Goebel</surname> <given-names>R.</given-names></name> <name><surname>Bandettini</surname> <given-names>P.</given-names></name></person-group> (<year>2006</year>). <article-title>Information-based functional brain mapping</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>103</volume>, <fpage>3863</fpage>&#x02013;<lpage>3868</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0600244103</pub-id><pub-id pub-id-type="pmid">16537458</pub-id></citation></ref>
<ref id="B135">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Lindquist</surname> <given-names>M. A.</given-names></name> <name><surname>Nichols</surname> <given-names>T. E.</given-names></name> <name><surname>Poldrack</surname> <given-names>R. A.</given-names></name> <name><surname>Vul</surname> <given-names>E.</given-names></name></person-group> (<year>2010</year>). <article-title>Everything you never wanted to know about circular analysis, but were afraid to ask</article-title>. <source>J. Cereb. Blood Flow Metab.</source> <volume>30</volume>, <fpage>1551</fpage>&#x02013;<lpage>1557</lpage>. <pub-id pub-id-type="doi">10.1038/jcbfm.2010.86</pub-id><pub-id pub-id-type="pmid">20571517</pub-id></citation></ref>
<ref id="B136">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name> <name><surname>Simmons</surname> <given-names>W. K.</given-names></name> <name><surname>Bellgowan</surname> <given-names>P. S.</given-names></name> <name><surname>Baker</surname> <given-names>C. I.</given-names></name></person-group> (<year>2009</year>). <article-title>Circular analysis in systems neuroscience: the dangers of double dipping</article-title>. <source>Nat. Neurosci.</source> <volume>12</volume>, <fpage>535</fpage>&#x02013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1038/nn.2303</pub-id><pub-id pub-id-type="pmid">19396166</pub-id></citation></ref>
<ref id="B137">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kurzweil</surname> <given-names>R.</given-names></name></person-group> (<year>2005</year>). <source>The Singularity is Near: When Humans Transcend Biology</source>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>Penguin</publisher-name>.</citation></ref>
<ref id="B138">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lake</surname> <given-names>B. M.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name></person-group> (<year>2015</year>). <article-title>Human-level concept learning through probabilistic program induction</article-title>. <source>Science</source> <volume>350</volume>, <fpage>1332</fpage>&#x02013;<lpage>1338</lpage>. <pub-id pub-id-type="doi">10.1126/science.aab3050</pub-id><pub-id pub-id-type="pmid">26659050</pub-id></citation></ref>
<ref id="B139">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B140">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lemm</surname> <given-names>S.</given-names></name> <name><surname>Blankertz</surname> <given-names>B.</given-names></name> <name><surname>Dickhaus</surname> <given-names>T.</given-names></name> <name><surname>Muller</surname> <given-names>K. R.</given-names></name></person-group> (<year>2011</year>). <article-title>Introduction to machine learning for brain imaging</article-title>. <source>Neuroimage</source> <volume>56</volume>, <fpage>387</fpage>&#x02013;<lpage>399</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.11.004</pub-id><pub-id pub-id-type="pmid">21172442</pub-id></citation></ref>
<ref id="B141">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lieberman</surname> <given-names>M. D.</given-names></name> <name><surname>Berkman</surname> <given-names>E. T.</given-names></name> <name><surname>Wager</surname> <given-names>T. D.</given-names></name></person-group> (<year>2009</year>). <article-title>Correlations in social neuroscience aren&#x00027;t Voodoo: Commentary on Vul et al</article-title>. <source>Perspect. Psychol. Sci.</source> <volume>4</volume>, <fpage>299</fpage>&#x02013;<lpage>307</lpage>. <pub-id pub-id-type="doi">10.1111/j.1745-6924.2009.01128.x</pub-id><pub-id pub-id-type="pmid">26158967</pub-id></citation></ref>
<ref id="B142">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lo</surname> <given-names>A.</given-names></name> <name><surname>Chernoff</surname> <given-names>H.</given-names></name> <name><surname>Zheng</surname> <given-names>T.</given-names></name> <name><surname>Lo</surname> <given-names>S. H.</given-names></name></person-group> (<year>2015</year>). <article-title>Why significant variables aren&#x00027;t automatically good predictors</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>13892</fpage>&#x02013;<lpage>13897</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1518285112</pub-id><pub-id pub-id-type="pmid">26504198</pub-id></citation></ref>
<ref id="B143">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Loftus</surname> <given-names>J. R.</given-names></name></person-group> (<year>2015</year>). <article-title>Selective inference after cross-validation</article-title>. arXiv:1511.08866.</citation></ref>
<ref id="B144">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Logothetis</surname> <given-names>N. K.</given-names></name> <name><surname>Pauls</surname> <given-names>J.</given-names></name> <name><surname>Augath</surname> <given-names>M.</given-names></name> <name><surname>Trinath</surname> <given-names>T.</given-names></name> <name><surname>Oeltermann</surname> <given-names>A.</given-names></name></person-group> (<year>2001</year>). <article-title>Neurophysiological investigation of the basis of the fMRI signal</article-title>. <source>Nature</source> <volume>412</volume>, <fpage>150</fpage>&#x02013;<lpage>157</lpage>. <pub-id pub-id-type="doi">10.1038/35084005</pub-id><pub-id pub-id-type="pmid">11449264</pub-id></citation></ref>
<ref id="B145">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Manyika</surname> <given-names>J.</given-names></name> <name><surname>Chui</surname> <given-names>M.</given-names></name> <name><surname>Brown</surname> <given-names>B.</given-names></name> <name><surname>Bughin</surname> <given-names>J.</given-names></name> <name><surname>Dobbs</surname> <given-names>R.</given-names></name> <name><surname>Roxburgh</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2011</year>). <source>Big Data: The Next Frontier for Innovation, Competition, and Productivity</source>. Technical Report, McKinsey Global Institute.</citation></ref>
<ref id="B146">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markram</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>The human brain project</article-title>. <source>Sci. Am.</source> <volume>306</volume>, <fpage>50</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1038/scientificamerican0612-50</pub-id><pub-id pub-id-type="pmid">22649994</pub-id></citation></ref>
<ref id="B147">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mars</surname> <given-names>R. B.</given-names></name> <name><surname>Sallet</surname> <given-names>J.</given-names></name> <name><surname>Schuffelgen</surname> <given-names>U.</given-names></name> <name><surname>Jbabdi</surname> <given-names>S.</given-names></name> <name><surname>Toni</surname> <given-names>I.</given-names></name> <name><surname>Rushworth</surname> <given-names>M. F.</given-names></name></person-group> (<year>2012</year>). <article-title>Connectivity-based subdivisions of the human right &#x0201C;Temporoparietal Junction Area&#x0201D;: evidence for different areas participating in different cortical networks</article-title>. <source>Cereb. Cortex</source> <volume>22</volume>, <fpage>1894</fpage>&#x02013;<lpage>1903</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhr268</pub-id><pub-id pub-id-type="pmid">21955921</pub-id></citation></ref>
<ref id="B148">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>K. L.</given-names></name> <name><surname>Alfaro-Almagro</surname> <given-names>F.</given-names></name> <name><surname>Bangerter</surname> <given-names>N. K.</given-names></name> <name><surname>Thomas</surname> <given-names>D. L.</given-names></name> <name><surname>Yacoub</surname> <given-names>E.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Multimodal population brain imaging in the UK Biobank prospective epidemiological study</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>1523</fpage>&#x02013;<lpage>1536</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4393</pub-id><pub-id pub-id-type="pmid">27643430</pub-id></citation></ref>
<ref id="B149">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Misaki</surname> <given-names>M.</given-names></name> <name><surname>Kim</surname> <given-names>Y.</given-names></name> <name><surname>Bandettini</surname> <given-names>P. A.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2010</year>). <article-title>Comparison of multivariate classifiers and response normalizations for pattern-information fMRI</article-title>. <source>Neuroimage</source> <volume>53</volume>, <fpage>103</fpage>&#x02013;<lpage>118</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.05.051</pub-id><pub-id pub-id-type="pmid">20580933</pub-id></citation></ref>
<ref id="B150">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moeller</surname> <given-names>J. R.</given-names></name> <name><surname>Strother</surname> <given-names>S. C.</given-names></name> <name><surname>Sidtis</surname> <given-names>J. J.</given-names></name> <name><surname>Rottenberg</surname> <given-names>D. A.</given-names></name></person-group> (<year>1987</year>). <article-title>Scaled subprofile model: a statistical approach to the analysis of functional patterns in positron emission tomographic data</article-title>. <source>J. Cereb. Blood Flow Metab.</source> <volume>7</volume>, <fpage>649</fpage>&#x02013;<lpage>658</lpage>. <pub-id pub-id-type="doi">10.1038/jcbfm.1987.118</pub-id><pub-id pub-id-type="pmid">3498733</pub-id></citation></ref>
<ref id="B151">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mur</surname> <given-names>M.</given-names></name> <name><surname>Bandettini</surname> <given-names>P. A.</given-names></name> <name><surname>Kriegeskorte</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <article-title>Revealing representational content with pattern-information fMRI&#x02013;an introductory guide</article-title>. <source>Soc. Cogn. Affect. Neurosci.</source> <volume>4</volume>, <fpage>101</fpage>&#x02013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1093/scan/nsn044</pub-id><pub-id pub-id-type="pmid">19151374</pub-id></citation></ref>
<ref id="B152">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Murphy</surname> <given-names>K. P.</given-names></name></person-group> (<year>2012</year>). <source>Machine Learning: A Probabilistic Perspective</source>. <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation></ref>
<ref id="B153">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naselaris</surname> <given-names>T.</given-names></name> <name><surname>Kay</surname> <given-names>K. N.</given-names></name> <name><surname>Nishimoto</surname> <given-names>S.</given-names></name> <name><surname>Gallant</surname> <given-names>J. L.</given-names></name></person-group> (<year>2011</year>). <article-title>Encoding and decoding in fMRI</article-title>. <source>Neuroimage</source> <volume>56</volume>, <fpage>400</fpage>&#x02013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.07.073</pub-id><pub-id pub-id-type="pmid">20691790</pub-id></citation></ref>
<ref id="B154">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Neyman</surname> <given-names>J.</given-names></name> <name><surname>Pearson</surname> <given-names>E. S.</given-names></name></person-group> (<year>1933</year>). <article-title>On the problem of the most efficient tests for statistical hypotheses</article-title>. <source>Philos. Trans. R. Soc. A</source> <volume>231</volume>, <fpage>289</fpage>&#x02013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.1098/rsta.1933.0009</pub-id></citation></ref>
<ref id="B155">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nichols</surname> <given-names>T. E.</given-names></name></person-group> (<year>2012</year>). <article-title>Multiple testing corrections, nonparametric methods, and random field theory</article-title>. <source>Neuroimage</source> <volume>62</volume>, <fpage>811</fpage>&#x02013;<lpage>815</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.04.014</pub-id><pub-id pub-id-type="pmid">22521256</pub-id></citation></ref>
<ref id="B156">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nichols</surname> <given-names>T. E.</given-names></name> <name><surname>Hayasaka</surname> <given-names>S.</given-names></name></person-group> (<year>2003</year>). <article-title>Controlling the familywise error rate in functional neuroimaging: a comparative review</article-title>. <source>Stat. Methods Med. Res.</source> <volume>12</volume>, <fpage>419</fpage>&#x02013;<lpage>446</lpage>. <pub-id pub-id-type="doi">10.1191/0962280203sm341ra</pub-id><pub-id pub-id-type="pmid">14599004</pub-id></citation></ref>
<ref id="B157">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nichols</surname> <given-names>T. E.</given-names></name> <name><surname>Holmes</surname> <given-names>A. P.</given-names></name></person-group> (<year>2002</year>). <article-title>Nonparametric permutation tests for functional neuroimaging: a primer with examples</article-title>. <source>Hum. Brain Mapp.</source> <volume>15</volume>, <fpage>1</fpage>&#x02013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.1058</pub-id><pub-id pub-id-type="pmid">11747097</pub-id></citation></ref>
<ref id="B158">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nickerson</surname> <given-names>R. S.</given-names></name></person-group> (<year>2000</year>). <article-title>Null hypothesis significance testing: a review of an old and continuing controversy</article-title>. <source>Psychol. Methods</source> <volume>5</volume>, <fpage>241</fpage>&#x02013;<lpage>301</lpage>. <pub-id pub-id-type="doi">10.1037/1082-989X.5.2.241</pub-id><pub-id pub-id-type="pmid">10937333</pub-id></citation></ref>
<ref id="B159">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noirhomme</surname> <given-names>Q.</given-names></name> <name><surname>Lesenfants</surname> <given-names>D.</given-names></name> <name><surname>Gomez</surname> <given-names>F.</given-names></name> <name><surname>Soddu</surname> <given-names>A.</given-names></name> <name><surname>Schrouff</surname> <given-names>J.</given-names></name> <name><surname>Garraux</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Biased binomial assessment of cross-validated estimation of classification accuracies illustrated in diagnosis predictions</article-title>. <source>Neuroimage</source> <volume>4</volume>, <fpage>687</fpage>&#x02013;<lpage>694</lpage>. <pub-id pub-id-type="doi">10.1016/j.nicl.2014.04.004</pub-id><pub-id pub-id-type="pmid">24936420</pub-id></citation></ref>
<ref id="B160">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Norman</surname> <given-names>K. A.</given-names></name> <name><surname>Polyn</surname> <given-names>S. M.</given-names></name> <name><surname>Detre</surname> <given-names>G. J.</given-names></name> <name><surname>Haxby</surname> <given-names>J. V.</given-names></name></person-group> (<year>2006</year>). <article-title>Beyond mind-reading: multi-voxel pattern analysis of fMRI data</article-title>. <source>Trends Cogn. Sci.</source> <volume>10</volume>, <fpage>424</fpage>&#x02013;<lpage>430</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2006.07.005</pub-id><pub-id pub-id-type="pmid">16899397</pub-id></citation></ref>
<ref id="B161">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nuzzo</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Scientific method: statistical errors</article-title>. <source>Nature</source> <volume>506</volume>, <fpage>150</fpage>&#x02013;<lpage>152</lpage>. <pub-id pub-id-type="doi">10.1038/506150a</pub-id><pub-id pub-id-type="pmid">24522584</pub-id></citation></ref>
<ref id="B162">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oakes</surname> <given-names>M.</given-names></name></person-group> (<year>1986</year>). <source>Statistical Inference: A Commentary for the Social and Behavioral Sciences</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Wiley</publisher-name>.</citation></ref>
<ref id="B163">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Passingham</surname> <given-names>R. E.</given-names></name> <name><surname>Stephan</surname> <given-names>K. E.</given-names></name> <name><surname>Kotter</surname> <given-names>R.</given-names></name></person-group> (<year>2002</year>). <article-title>The anatomical basis of functional localization in the cortex</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>3</volume>, <fpage>606</fpage>&#x02013;<lpage>616</lpage>. <pub-id pub-id-type="doi">10.1038/nrn893</pub-id><pub-id pub-id-type="pmid">12154362</pub-id></citation></ref>
<ref id="B164">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedregosa</surname> <given-names>F.</given-names></name> <name><surname>Eickenberg</surname> <given-names>M.</given-names></name> <name><surname>Ciuciu</surname> <given-names>P.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Gramfort</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Data-driven HRF estimation for encoding and decoding models</article-title>. <source>Neuroimage</source> <volume>104</volume>, <fpage>209</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2014.09.060</pub-id><pub-id pub-id-type="pmid">25304775</pub-id></citation></ref>
<ref id="B165">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pereira</surname> <given-names>F.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Information mapping with pattern classifiers: a comparative study</article-title>. <source>Neuroimage</source> <volume>56</volume>, <fpage>476</fpage>&#x02013;<lpage>496</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.05.026</pub-id><pub-id pub-id-type="pmid">20488249</pub-id></citation></ref>
<ref id="B166">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pereira</surname> <given-names>F.</given-names></name> <name><surname>Mitchell</surname> <given-names>T.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Machine learning classifiers and fMRI: a tutorial overview</article-title>. <source>Neuroimage</source> <volume>45</volume>, <fpage>199</fpage>&#x02013;<lpage>209</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2008.11.007</pub-id><pub-id pub-id-type="pmid">19070668</pub-id></citation></ref>
<ref id="B167">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pernet</surname> <given-names>C. R.</given-names></name> <name><surname>Chauveau</surname> <given-names>N.</given-names></name> <name><surname>Gaspar</surname> <given-names>C.</given-names></name> <name><surname>Rousselet</surname> <given-names>G. A.</given-names></name></person-group> (<year>2011</year>). <article-title>LIMO EEG: a toolbox for hierarchical LInear MOdeling of ElectroEncephaloGraphic data</article-title>. <source>Comput. Intell. Neurosci.</source> <volume>2011</volume>:<fpage>831409</fpage>. <pub-id pub-id-type="doi">10.1155/2011/831409</pub-id><pub-id pub-id-type="pmid">21403915</pub-id></citation></ref>
<ref id="B168">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Platt</surname> <given-names>J. R.</given-names></name></person-group> (<year>1964</year>). <article-title>Strong inference: certain systematic methods of scientific thinking may produce much more rapid progress than others</article-title>. <source>Science</source> <volume>146</volume>, <fpage>347</fpage>&#x02013;<lpage>353</lpage>. <pub-id pub-id-type="doi">10.1126/science.146.3642.347</pub-id><pub-id pub-id-type="pmid">17739513</pub-id></citation></ref>
<ref id="B169">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Plis</surname> <given-names>S. M.</given-names></name> <name><surname>Hjelm</surname> <given-names>D. R.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name> <name><surname>Allen</surname> <given-names>E. A.</given-names></name> <name><surname>Bockholt</surname> <given-names>H. J.</given-names></name> <name><surname>Long</surname> <given-names>J. D.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Deep learning for neuroimaging: a validation study</article-title>. <source>Front. Neurosci.</source> <volume>8</volume>:<fpage>229</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2014.00229</pub-id><pub-id pub-id-type="pmid">25191215</pub-id></citation></ref>
<ref id="B170">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poldrack</surname> <given-names>R. A.</given-names></name></person-group> (<year>2006</year>). <article-title>Can cognitive processes be inferred from neuroimaging data?</article-title> <source>Trends Cogn. Sci.</source> <volume>10</volume>, <fpage>59</fpage>&#x02013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2005.12.004</pub-id><pub-id pub-id-type="pmid">16406760</pub-id></citation></ref>
<ref id="B171">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poldrack</surname> <given-names>R. A.</given-names></name> <name><surname>Baker</surname> <given-names>C. I.</given-names></name> <name><surname>Durnez</surname> <given-names>J.</given-names></name> <name><surname>Gorgolewski</surname> <given-names>K. J.</given-names></name> <name><surname>Matthews</surname> <given-names>P. M.</given-names></name> <name><surname>Munaf&#x000F2;</surname> <given-names>M. R.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Scanning the horizon: towards transparent and reproducible neuroimaging research</article-title>. <source>Nat. Rev. Neurosci.</source> <volume>18</volume>, <fpage>115</fpage>&#x02013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1038/nrn.2016.167</pub-id><pub-id pub-id-type="pmid">28053326</pub-id></citation></ref>
<ref id="B172">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poldrack</surname> <given-names>R. A.</given-names></name> <name><surname>Gorgolewski</surname> <given-names>K. J.</given-names></name></person-group> (<year>2014</year>). <article-title>Making big data open: data sharing in neuroimaging</article-title>. <source>Nat. Neurosci.</source> <volume>17</volume>, <fpage>1510</fpage>&#x02013;<lpage>1517</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3818</pub-id><pub-id pub-id-type="pmid">25349916</pub-id></citation></ref>
<ref id="B173">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poline</surname> <given-names>J.-B.</given-names></name> <name><surname>Brett</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>The general linear model and fMRI: does love last forever?</article-title> <source>Neuroimage</source> <volume>62</volume>, <fpage>871</fpage>&#x02013;<lpage>880</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.01.133</pub-id><pub-id pub-id-type="pmid">22343127</pub-id></citation></ref>
<ref id="B174">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Popper</surname> <given-names>K.</given-names></name></person-group> (<year>1935/2005</year>). <source>Logik der Forschung, 11th Edn</source>. <publisher-loc>T&#x000FC;bingen</publisher-loc>: <publisher-name>Mohr Siebeck</publisher-name>.</citation></ref>
<ref id="B175">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Powers</surname> <given-names>D. M.</given-names></name></person-group> (<year>2011</year>). <article-title>Evaluation: from precision, recall and f-measure to ROC, informedness, markedness and correlation</article-title>. <source>J. Mach. Learn. Technol.</source> <volume>2</volume>, <fpage>37</fpage>&#x02013;<lpage>63</lpage></citation></ref>
<ref id="B176">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosenblatt</surname> <given-names>F.</given-names></name></person-group> (<year>1958</year>). <article-title>The perceptron: a probabilistic model for information storage and organization in the brain</article-title>. <source>Psychol. Rev.</source> <volume>65</volume>, <fpage>386</fpage>. <pub-id pub-id-type="doi">10.1037/h0042519</pub-id><pub-id pub-id-type="pmid">13602029</pub-id></citation></ref>
<ref id="B177">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosnow</surname> <given-names>R. L.</given-names></name> <name><surname>Rosenthal</surname> <given-names>R.</given-names></name></person-group> (<year>1989</year>). <article-title>Statistical procedures and the justification of knowledge in psychological science</article-title>. <source>Am. Psychol.</source> <volume>44</volume>:<fpage>1276</fpage>. <pub-id pub-id-type="doi">10.1037/0003-066X.44.10.1276</pub-id></citation></ref>
<ref id="B178">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Russell</surname> <given-names>S. J.</given-names></name> <name><surname>Norvig</surname> <given-names>P.</given-names></name></person-group> (<year>2002</year>). <source>Artificial Intelligence: A Modern Approach (International Edition)</source>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>Pearson</publisher-name>.</citation></ref>
<ref id="B179">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samuel</surname> <given-names>A. L.</given-names></name></person-group> (<year>1959</year>). <article-title>Some studies in machine learning using the game of checkers</article-title>. <source>IBM J. Res. Dev.</source> <volume>3</volume>, <fpage>210</fpage>&#x02013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1147/rd.33.0210</pub-id></citation></ref>
<ref id="B180">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saygin</surname> <given-names>Z. M.</given-names></name> <name><surname>Osher</surname> <given-names>D. E.</given-names></name> <name><surname>Koldewyn</surname> <given-names>K.</given-names></name> <name><surname>Reynolds</surname> <given-names>G.</given-names></name> <name><surname>Gabrieli</surname> <given-names>J. D.</given-names></name> <name><surname>Saxe</surname> <given-names>R. R.</given-names></name></person-group> (<year>2012</year>). <article-title>Anatomical connectivity patterns predict face selectivity in the fusiform gyrus</article-title>. <source>Nat. Neurosci.</source> <volume>15</volume>, <fpage>321</fpage>&#x02013;<lpage>327</lpage>. <pub-id pub-id-type="doi">10.1038/nn.3001</pub-id><pub-id pub-id-type="pmid">22197830</pub-id></citation></ref>
<ref id="B181">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Scheff&#x000E9;</surname> <given-names>H.</given-names></name></person-group> (<year>1959</year>). <source>The Analysis of Variance</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Wiley</publisher-name>. <pub-id pub-id-type="pmid">20260627</pub-id></citation></ref>
<ref id="B182">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidt</surname> <given-names>F. L.</given-names></name></person-group> (<year>1996</year>). <article-title>Statistical significance testing and cumulative knowledge in psychology: implications for training of researchers</article-title>. <source>Psychol. Methods</source> <volume>1</volume>:<fpage>115</fpage>. <pub-id pub-id-type="doi">10.1037/1082-989X.1.2.115</pub-id></citation></ref>
<ref id="B183">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwartz</surname> <given-names>Y.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>Mapping paradigm ontologies to and from the brain,</article-title> in <source>Advances in Neural Information Processing Systems.</source> <fpage>1673</fpage>&#x02013;<lpage>1681</lpage>.</citation></ref>
<ref id="B184">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shalev-Shwartz</surname> <given-names>S.</given-names></name> <name><surname>Ben-David</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <source>Understanding Machine Learning: From Theory to Algorithms</source>. <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B185">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shmueli</surname> <given-names>G.</given-names></name></person-group> (<year>2010</year>). <article-title>To explain or to predict?</article-title> <source>Stat. Sci.</source> <volume>25</volume>, <fpage>289</fpage>&#x02013;<lpage>310</lpage>. <pub-id pub-id-type="doi">10.1214/10-STS330</pub-id></citation></ref>
<ref id="B186">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sladek</surname> <given-names>R.</given-names></name> <name><surname>Rocheleau</surname> <given-names>G.</given-names></name> <name><surname>Rung</surname> <given-names>J.</given-names></name> <name><surname>Dina</surname> <given-names>C.</given-names></name> <name><surname>Shen</surname> <given-names>L.</given-names></name> <name><surname>Serre</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>A genome-wide association study identifies novel risk loci for type 2 diabetes</article-title>. <source>Nature</source> <volume>445</volume>, <fpage>881</fpage>&#x02013;<lpage>885</lpage>. <pub-id pub-id-type="doi">10.1038/nature05616</pub-id><pub-id pub-id-type="pmid">17293876</pub-id></citation></ref>
<ref id="B187">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>S. M.</given-names></name> <name><surname>Beckmann</surname> <given-names>C. F.</given-names></name> <name><surname>Andersson</surname> <given-names>J.</given-names></name> <name><surname>Auerbach</surname> <given-names>E. J.</given-names></name> <name><surname>Bijsterbosch</surname> <given-names>J.</given-names></name> <name><surname>Douaud</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Resting-state fMRI in the human connectome project</article-title>. <source>Neuroimage</source> <volume>80</volume>, <fpage>144</fpage>&#x02013;<lpage>168</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.05.039</pub-id><pub-id pub-id-type="pmid">23702415</pub-id></citation></ref>
<ref id="B188">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>S. M.</given-names></name> <name><surname>Matthews</surname> <given-names>P. M.</given-names></name> <name><surname>Jezzard</surname> <given-names>P.</given-names></name></person-group> (<year>2001</year>). <source>Functional MRI: An Introduction to Methods</source>. <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="B189">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>S. M.</given-names></name> <name><surname>Nichols</surname> <given-names>T. E.</given-names></name></person-group> (<year>2009</year>). <article-title>Threshold-free cluster enhancement: addressing problems of smoothing, threshold dependence and localisation in cluster inference</article-title>. <source>Neuroimage</source> <volume>44</volume>, <fpage>83</fpage>&#x02013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2008.03.061</pub-id><pub-id pub-id-type="pmid">18501637</pub-id></citation></ref>
<ref id="B190">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stark</surname> <given-names>C. E.</given-names></name> <name><surname>Squire</surname> <given-names>L. R.</given-names></name></person-group> (<year>2001</year>). <article-title>When zero is not zero: the problem of ambiguous baseline conditions in fMRI</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>98</volume>, <fpage>12760</fpage>&#x02013;<lpage>12766</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.221462998</pub-id></citation></ref>
<ref id="B191">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taylor</surname> <given-names>J.</given-names></name> <name><surname>Lockhart</surname> <given-names>R.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R. J.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Exact post-selection inference for forward stepwise and least angle regression</article-title>. arXiv:1401.3889.</citation></ref>
<ref id="B192">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taylor</surname> <given-names>J.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Statistical learning and selective inference</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>7629</fpage>&#x02013;<lpage>7634</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1507583112</pub-id><pub-id pub-id-type="pmid">26100887</pub-id></citation></ref>
<ref id="B193">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tenenbaum</surname> <given-names>J. B.</given-names></name> <name><surname>Kemp</surname> <given-names>C.</given-names></name> <name><surname>Griffiths</surname> <given-names>T. L.</given-names></name> <name><surname>Goodman</surname> <given-names>N. D.</given-names></name></person-group> (<year>2011</year>). <article-title>How to grow a mind: statistics, structure, and abstraction</article-title>. <source>Science</source> <volume>331</volume>, <fpage>1279</fpage>&#x02013;<lpage>1285</lpage>. <pub-id pub-id-type="doi">10.1126/science.1192788</pub-id><pub-id pub-id-type="pmid">21393536</pub-id></citation></ref>
<ref id="B194">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Dohmatob</surname> <given-names>E.</given-names></name> <name><surname>Poline</surname> <given-names>J. B.</given-names></name></person-group> (<year>2014</year>). <article-title>Which fMRI clustering gives good brain parcellations?</article-title> <source>Front. Neurosci.</source> <volume>8</volume>:<fpage>167</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2014.00167</pub-id><pub-id pub-id-type="pmid">25071425</pub-id></citation></ref>
<ref id="B195">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tibshirani</surname> <given-names>R.</given-names></name></person-group> (<year>1996</year>). <article-title>Regression shrinkage and selection via the lasso</article-title>. <source>J. R. Stat. Soc. B</source> <volume>73</volume>, <fpage>267</fpage>&#x02013;<lpage>288</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-9868.2011.00771.x</pub-id></citation></ref>
<ref id="B196">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tukey</surname> <given-names>J. W.</given-names></name></person-group> (<year>1962</year>). <article-title>The future of data analysis</article-title>. <source>Ann. Stat.</source> <volume>33</volume>, <fpage>1</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1214/aoms/1177704711</pub-id></citation></ref>
<ref id="B197">
<citation citation-type="other"><person-group person-group-type="author"><collab>UK House of Common S.a.T</collab></person-group> (<year>2016</year>). <source>The Big Data Dilemma</source>. Committee on Applied and Theoretical Statistics.</citation></ref>
<ref id="B198">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Vanderplas</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <source>The Big Data Brain Drain: Why Science is in Trouble.</source> Pythonic Perambulations.</citation></ref>
<ref id="B199">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Essen</surname> <given-names>D. C.</given-names></name> <name><surname>Ugurbil</surname> <given-names>K.</given-names></name> <name><surname>Auerbach</surname> <given-names>E.</given-names></name> <name><surname>Barch</surname> <given-names>D.</given-names></name> <name><surname>Behrens</surname> <given-names>T. E.</given-names></name> <name><surname>Bucholz</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>The human connectome project: a data acquisition perspective</article-title>. <source>Neuroimage</source> <volume>62</volume>, <fpage>2222</fpage>&#x02013;<lpage>2231</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.02.018</pub-id><pub-id pub-id-type="pmid">22366334</pub-id></citation></ref>
<ref id="B200">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Horn</surname> <given-names>J. D.</given-names></name> <name><surname>Toga</surname> <given-names>A. W.</given-names></name></person-group> (<year>2014</year>). <article-title>Human neuroimaging as a &#x0201C;Big Data&#x0201D; science</article-title>. <source>Brain Imaging Behav.</source> <volume>8</volume>, <fpage>323</fpage>&#x02013;<lpage>331</lpage>. <pub-id pub-id-type="doi">10.1007/s11682-013-9255-y</pub-id><pub-id pub-id-type="pmid">24113873</pub-id></citation></ref>
<ref id="B201">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vapnik</surname> <given-names>V. N.</given-names></name></person-group> (<year>1989</year>). <source>Statistical Learning Theory</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Wiley-Interscience</publisher-name>. <pub-id pub-id-type="pmid">18252602</pub-id></citation></ref>
<ref id="B202">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vapnik</surname> <given-names>V. N.</given-names></name></person-group> (<year>1996</year>). <source>The Nature of Statistical learnIng Theory</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B203">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vapnik</surname> <given-names>V. N.</given-names></name> <name><surname>Kotz</surname> <given-names>S.</given-names></name></person-group> (<year>1982</year>). <source>Estimation of Dependences Based on Empirical Data</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer-Verlag New York</publisher-name>.</citation></ref>
<ref id="B204">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>How machine learning is shaping cognitive neuroimaging</article-title>. <source>Gigascience</source> <volume>3</volume>:<fpage>28</fpage>. <pub-id pub-id-type="doi">10.1186/2047-217X-3-28</pub-id><pub-id pub-id-type="pmid">25405022</pub-id></citation></ref>
<ref id="B205">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vogelstein</surname> <given-names>J. T.</given-names></name> <name><surname>Park</surname> <given-names>Y.</given-names></name> <name><surname>Ohyama</surname> <given-names>T.</given-names></name> <name><surname>Kerr</surname> <given-names>R. A.</given-names></name> <name><surname>Truman</surname> <given-names>J. W.</given-names></name> <name><surname>Priebe</surname> <given-names>C. E.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Discovery of brainwide neural-behavioral maps via multiscale unsupervised structure learning</article-title>. <source>Science</source> <volume>344</volume>, <fpage>386</fpage>&#x02013;<lpage>392</lpage>. <pub-id pub-id-type="doi">10.1126/science.1250298</pub-id><pub-id pub-id-type="pmid">24674869</pub-id></citation></ref>
<ref id="B206">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vul</surname> <given-names>E.</given-names></name> <name><surname>Harris</surname> <given-names>C.</given-names></name> <name><surname>Winkielman</surname> <given-names>P.</given-names></name> <name><surname>Pashler</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Puzzlingly high correlations in fMRI studies of emotion, personality, and social cognition</article-title>. <source>Perspect. Psychol. Sci.</source> <volume>4</volume>, <fpage>274</fpage>&#x02013;<lpage>290</lpage>. <pub-id pub-id-type="doi">10.1111/j.1745-6924.2009.01125.x</pub-id></citation></ref>
<ref id="B207">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wainwright</surname> <given-names>M. J.</given-names></name></person-group> (<year>2014</year>). <article-title>Structured regularizers for high-dimensional problems: statistical and computational issues</article-title>. <source>Annu. Rev. Stat. Appl.</source> <volume>1</volume>, <fpage>233</fpage>&#x02013;<lpage>253</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-statistics-022513-115643</pub-id></citation></ref>
<ref id="B208">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wasserman</surname> <given-names>L.</given-names></name> <name><surname>Roeder</surname> <given-names>K.</given-names></name></person-group> (<year>2009</year>). <article-title>High dimensional variable selection</article-title>. <source>Ann. Stat.</source> <volume>37</volume>:<fpage>2178</fpage>. <pub-id pub-id-type="doi">10.1214/08-AOS646</pub-id><pub-id pub-id-type="pmid">19784398</pub-id></citation></ref>
<ref id="B209">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wasserstein</surname> <given-names>R. L.</given-names></name> <name><surname>Lazar</surname> <given-names>N. A.</given-names></name></person-group> (<year>2016</year>). <article-title>The ASA&#x00027;s statement on p-values: context, process, and purpose</article-title>. <source>Am. Stat.</source> <volume>70</volume>, <fpage>129</fpage>&#x02013;<lpage>133</lpage>. <pub-id pub-id-type="doi">10.1080/00031305.2016.1154108</pub-id></citation></ref>
<ref id="B210">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wolpert</surname> <given-names>D.</given-names></name></person-group> (<year>1996</year>). <article-title>The lack of a priori distinctions between learning algorithms</article-title>. <source>Neural Comput.</source> <volume>8</volume>, <fpage>1341</fpage>&#x02013;<lpage>1390</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1996.8.7.1341</pub-id></citation></ref>
<ref id="B211">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Worsley</surname> <given-names>K. J.</given-names></name> <name><surname>Evans</surname> <given-names>A. C.</given-names></name> <name><surname>Marrett</surname> <given-names>S.</given-names></name> <name><surname>Neelin</surname> <given-names>P.</given-names></name></person-group> (<year>1992</year>). <article-title>A three-dimensional statistical analysis for CBF activation studies in human brain</article-title>. <source>J. Cereb. Blood Flow Metab.</source> <volume>12</volume>, <fpage>900</fpage>&#x02013;<lpage>918</lpage>. <pub-id pub-id-type="doi">10.1038/jcbfm.1992.127</pub-id><pub-id pub-id-type="pmid">1400644</pub-id></citation></ref>
<ref id="B212">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Worsley</surname> <given-names>K. J.</given-names></name> <name><surname>Poline</surname> <given-names>J.-B.</given-names></name> <name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Evans</surname> <given-names>A. C.</given-names></name></person-group> (<year>1997</year>). <article-title>Characterizing the response of PET and fMRI data using multivariate linear models</article-title>. <source>Neuroimage</source> <volume>6</volume>, <fpage>305</fpage>&#x02013;<lpage>319</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.1997.0294</pub-id><pub-id pub-id-type="pmid">9417973</pub-id></citation></ref>
<ref id="B213">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>T. T.</given-names></name> <name><surname>Chen</surname> <given-names>Y. F.</given-names></name> <name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Sobel</surname> <given-names>E.</given-names></name> <name><surname>Lange</surname> <given-names>K.</given-names></name></person-group> (<year>2009</year>). <article-title>Genome-wide association analysis by lasso penalized logistic regression</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>714</fpage>&#x02013;<lpage>721</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp041</pub-id><pub-id pub-id-type="pmid">19176549</pub-id></citation></ref>
<ref id="B214">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamins</surname> <given-names>D. L.</given-names></name> <name><surname>DiCarlo</surname> <given-names>J. J.</given-names></name></person-group> (<year>2016</year>). <article-title>Using goal-driven deep learning models to understand sensory cortex</article-title>. <source>Nat. Neurosci.</source> <volume>19</volume>, <fpage>356</fpage>&#x02013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1038/nn.4244</pub-id><pub-id pub-id-type="pmid">26906502</pub-id></citation></ref>
<ref id="B215">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yarkoni</surname> <given-names>T.</given-names></name> <name><surname>Braver</surname> <given-names>T. S.</given-names></name></person-group> (<year>2010</year>). <article-title>Cognitive neuroscience approaches to individual differences in working memory and executive control: conceptual and methodological issues,</article-title> in <source>Handbook of Individual Differences in Cognition</source>, eds <person-group person-group-type="editor"><name><surname>Gruszka</surname> <given-names>A.</given-names></name> <name><surname>Matthews</surname> <given-names>G.</given-names></name> <name><surname>Szymura</surname> <given-names>B.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>87</fpage>&#x02013;<lpage>107</lpage>.</citation></ref>
<ref id="B216">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yarkoni</surname> <given-names>T.</given-names></name> <name><surname>Poldrack</surname> <given-names>R. A.</given-names></name> <name><surname>Nichols</surname> <given-names>T. E.</given-names></name> <name><surname>Van Essen</surname> <given-names>D. C.</given-names></name> <name><surname>Wager</surname> <given-names>T. D.</given-names></name></person-group> (<year>2011</year>). <article-title>Large-scale automated synthesis of human functional neuroimaging data</article-title>. <source>Nat. Methods</source> <volume>8</volume>, <fpage>665</fpage>&#x02013;<lpage>670</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.1635</pub-id><pub-id pub-id-type="pmid">21706013</pub-id></citation></ref>
<ref id="B217">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yarkoni</surname> <given-names>T.</given-names></name> <name><surname>Westfall</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Choosing prediction over explanation in psychology: lessons from machine learning</article-title>. <source>Perspect. Psychol. Sci.</source> [Epub ahead of print]. <pub-id pub-id-type="doi">10.1177/1745691617693393</pub-id><pub-id pub-id-type="pmid">28841086</pub-id></citation></ref>
<ref id="B218">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeo</surname> <given-names>B. T.</given-names></name> <name><surname>Krienen</surname> <given-names>F. M.</given-names></name> <name><surname>Chee</surname> <given-names>M. W.</given-names></name> <name><surname>Buckner</surname> <given-names>R. L.</given-names></name></person-group> (<year>2014</year>). <article-title>Estimates of segregation and overlap of functional connectivity networks in the human cerebral cortex</article-title>. <source>Neuroimage</source> <volume>88</volume>, <fpage>212</fpage>&#x02013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.10.046</pub-id><pub-id pub-id-type="pmid">24185018</pub-id></citation></ref>
<ref id="B219">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeo</surname> <given-names>B. T.</given-names></name> <name><surname>Krienen</surname> <given-names>F. M.</given-names></name> <name><surname>Sepulcre</surname> <given-names>J.</given-names></name> <name><surname>Sabuncu</surname> <given-names>M. R.</given-names></name> <name><surname>Lashkari</surname> <given-names>D.</given-names></name> <name><surname>Hollinshead</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>The organization of the human cerebral cortex estimated by intrinsic functional connectivity</article-title>. <source>J. Neurophysiol.</source> <volume>106</volume>, <fpage>1125</fpage>&#x02013;<lpage>1165</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00338.2011</pub-id><pub-id pub-id-type="pmid">21653723</pub-id></citation></ref>
<ref id="B220">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuste</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>From the neuron doctrine to neural networks</article-title>. <source>Nature Reviews Neuroscience</source> <volume>16</volume>, <fpage>487</fpage>&#x02013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.1038/nrn3962</pub-id><pub-id pub-id-type="pmid">26152865</pub-id></citation></ref>
<ref id="B221">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zou</surname> <given-names>H.</given-names></name> <name><surname>Hastie</surname> <given-names>T.</given-names></name></person-group> (<year>2005</year>). <article-title>Regularization and variable selection via the elastic net</article-title>. <source>J. R. Stat. Soc.</source> <volume>67</volume>, <fpage>301</fpage>&#x02013;<lpage>320</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-9868.2005.00503.x</pub-id></citation></ref>
<ref id="B222">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>zu Eulenburg</surname> <given-names>P.</given-names></name> <name><surname>Caspers</surname> <given-names>S.</given-names></name> <name><surname>Roski</surname> <given-names>C.</given-names></name> <name><surname>Eickhoff</surname> <given-names>S. B.</given-names></name></person-group> (<year>2012</year>). <article-title>Meta-analytical definition and functional connectivity of the human vestibular cortex</article-title>. <source>Neuroimage</source> <volume>60</volume>, <fpage>162</fpage>&#x02013;<lpage>169</lpage> <pub-id pub-id-type="doi">10.1016/j.neuroimage.2011.12.032</pub-id><pub-id pub-id-type="pmid">22209784</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>&#x0201C;Data Science and Statistics: different worlds?&#x0201D; (Panel at Royal Statistical Society UK, March 2015) (<ext-link ext-link-type="uri" xlink:href="https://www.youtube.com/watch?v=C1zMUjHOLr4">https://www.youtube.com/watch?v=C1zMUjHOLr4</ext-link>&#x0201D;)</p></fn>
<fn id="fn0002"><p><sup>2</sup>&#x0201C;50 years of Data Science&#x0201D; (David Donoho, Tukey Centennial workshop, USA, September 2015)</p></fn>
<fn id="fn0003"><p><sup>3</sup>&#x0201C;Are ML and Statistics Complementary?&#x0201D; (Max Welling, 6th IMS-ISBA meeting, December 2015)</p></fn>
<fn id="fn0004"><p><sup>4</sup>In the supervised setting, there is no a priori distinction between learning algorithms evaluated by out-of-sample prediction error. In the optimization setting of finite spaces, all algorithms searching an extremum perform identically when averaged across possible cost functions. (<ext-link ext-link-type="uri" xlink:href="http://www.no-free-lunch.org/">http://www.no-free-lunch.org/</ext-link>)</p></fn>
<fn id="fn0005"><p><sup>5</sup>&#x0201C;If you torture the data enough, nature will always confess.&#x0201D; (Coase R. H., <xref ref-type="bibr" rid="B39">1982</xref> How should economists choose?)</p></fn>
<fn id="fn0006"><p><sup>6</sup>&#x0201C;Once applied only to the selected few, the interpretation of the usual measures of uncertainty do not remain intact directly, unless properly adjusted.&#x0201D; (Yoav Benjamini)</p></fn>
</fn-group>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This work was supported by the Deutsche Forschungsgemeinschaft (DFG, BZ2/2-1, BZ2/3-1, and BZ2/4-1; International Research Training Group IRTG2150), Amazon AWS Research Grant (2016 and 2017), the German National Merit Foundation, and the START-Program of the Faculty of Medicine, RWTH Aachen.</p>
</fn>
</fn-group>
</back>
</article>