<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2023.1217661</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The Methodological Quality Scale (MQS) for intervention programs: validity evidence</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Chac&#x00F3;n-Moscoso</surname>
<given-names>Salvador</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/196512/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Sanduvete-Chaves</surname>
<given-names>Susana</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/196514/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Lozano-Lozano</surname>
<given-names>Jos&#x00E9; Antonio</given-names>
</name>
<xref rid="aff3" ref-type="aff"><sup>3</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/210450/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Holgado-Tello</surname>
<given-names>Francisco Pablo</given-names>
</name>
<xref rid="aff4" ref-type="aff"><sup>4</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/225047/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Departamento de Psicolog&#x00ED;a Experimental, Facultad de Psicolog&#x00ED;a, Universidad de Sevilla</institution>, <addr-line>Sevilla</addr-line>, <country>Spain</country></aff>
<aff id="aff2"><sup>2</sup><institution>Departamento de Psicolog&#x00ED;a, Universidad Aut&#x00F3;noma de Chile</institution>, <addr-line>Santiago</addr-line>, <country>Chile</country></aff>
<aff id="aff3"><sup>3</sup><institution>Instituto de Ciencias Biom&#x00E9;dicas, Universidad Aut&#x00F3;noma de Chile</institution>, <addr-line>Santiago</addr-line>, <country>Chile</country></aff>
<aff id="aff4"><sup>4</sup><institution>Departamento de Metodolog&#x00ED;a de las Ciencias del Comportamiento, Facultad de Psicolog&#x00ED;a, Universidad Nacional de Educaci&#x00F3;n a Distancia</institution>, <addr-line>Madrid</addr-line>, <country>Spain</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Gudberg K. Jonsson, University of Iceland, Iceland</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Miguel Pic, South Ural State University, Russia; Elena Escolano-P&#x00E9;rez, University of Zaragoza, Spain</p></fn>
<corresp id="c001">&#x002A;Correspondence: Salvador Chac&#x00F3;n-Moscoso, <email>schacon@us.es</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>07</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1217661</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Chac&#x00F3;n-Moscoso, Sanduvete-Chaves, Lozano-Lozano and Holgado-Tello.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Chac&#x00F3;n-Moscoso, Sanduvete-Chaves, Lozano-Lozano and Holgado-Tello</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>A wide variety of instruments are used when assessing the methodological quality (MQ) of intervention programs. Nevertheless, studies on their metric quality are often not available. In order to address this shortcoming, the methodological quality scale (MQS) is presented as a simple and useful tool with adequate reliability, validity evidence, and metric properties.</p>
</sec>
<sec>
<title>Methods</title>
<p>Two coders independently applied the MQS to a set of primary studies. The number of MQ facets was determined in parallel analyses before performing factor analyses. For each facet of validity obtained, mean and standard deviation are presented jointly with reliability and average discrimination. Additionally, the validity facet scores are interpreted based on Shadish, Cook, and Campbell&#x2019;s validity model.</p>
</sec>
<sec>
<title>Results and discussion</title>
<p>An empirical validation of the three facets of the MQ (external, internal, and construct validity) and the interpretation of the scores were obtained based on a theoretical framework. Unlike other existing scales, MQS is easy to apply and presents adequate metric properties. In addition, MQ profiles can be obtained in different areas of intervention using different methodologies and proves useful for both researchers doing meta-analysis and for evaluators and professionals designing a new intervention.</p>
</sec>
</abstract>
<kwd-group>
<kwd>methodological quality</kwd>
<kwd>scale</kwd>
<kwd>meta-analysis</kwd>
<kwd>program evaluation</kwd>
<kwd>reliability</kwd>
<kwd>validity</kwd>
</kwd-group>
<contract-num rid="cn1">1190945</contract-num>
<contract-num rid="cn2">PID2020-115486GB-I00</contract-num>
<contract-num rid="cn3">MCIN/AEI/10.13039/501100011033</contract-num>
<contract-num rid="cn4">PID2020-114538RB-I00</contract-num>
<contract-sponsor id="cn1">Chilean national projects FONDECYT Regular 2019, Agencia Nacional de Investigaci&#x00F3;n y Desarrollo (ANID)&#x2013;, Government of Chile</contract-sponsor>
<contract-sponsor id="cn2">Ministerio de Ciencia e Innovaci&#x00F3;n<named-content content-type="fundref-id">10.13039/501100004837</named-content></contract-sponsor>
<contract-sponsor id="cn3">Government of Spain</contract-sponsor>
<contract-sponsor id="cn4">Ministerio de Ciencia e Innovaci&#x00F3;n, Government of Spain</contract-sponsor>
<counts>
<fig-count count="2"/>
<table-count count="8"/>
<equation-count count="0"/>
<ref-count count="35"/>
<page-count count="10"/>
<word-count count="7287"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Quantitative Psychology and Measurement</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="sec4" sec-type="intro">
<label>1.</label>
<title>Introduction</title>
<p>The concept methodological quality (MQ) can be defined as the degree to which a study can avoid systematic errors (bias), and the degree to which we are sure that such study can be believed (<xref ref-type="bibr" rid="ref25">Reitsma et al., 2009</xref>). Measuring MQ is important to foster accumulative knowledge given the relationship between MQ and effect size, where effect size is higher when MQ is low; i.e., low MQ studies tend to overestimate the effectiveness of interventions (<xref ref-type="bibr" rid="ref12">Hempel et al., 2011</xref>). Thus, when multiple interventions lack MQ, it becomes difficult to reach trustworthy conclusions (<xref ref-type="bibr" rid="ref5">Chac&#x00F3;n-Moscoso et al., 2014</xref>).</p>
<p>In meta-analytical research, the results of different primary studies on a specific issue or research question are quantitatively integrated (<xref ref-type="bibr" rid="ref9">Cooper et al., 2009</xref>). Generally, the MQ of primary studies is measured with the intent of evaluating the credibility of the results obtained in the meta-analysis (<xref ref-type="bibr" rid="ref21">Luhnen et al., 2019</xref>). On some occasions, low MQ is an exclusion criterion.</p>
<p>Measuring MQ is not just useful for integrating finished interventions in meta-analysis. In the context of program evaluation, it is also fundamental to increase the MQ of the design, the implementation, and the evaluation of ongoing and future intervention programs. Finally, when several different interventions are feasible, it allows the most adequate to be chosen based on the target population, the aims, and the context (<xref ref-type="bibr" rid="ref4">Cano-Garc&#x00ED;a et al., 2017</xref>).</p>
<p>Thus, a wide variety of professionals need to measure MQ. In most cases, these professionals are not experts in methodological issues, and when they seek out an instrument to gauge MQ, they encounter a wide variety of them (<xref ref-type="bibr" rid="ref14">Higgins et al., 2013</xref>). There are two reasons why experts in methodology are unable to offer a simple way to assess quality.</p>
<p>First, a plethora of strategies for assessing the MQ of primary sources can be found in the literature (see, for example, <ext-link xlink:href="http://www.equator-network.org" ext-link-type="uri">www.equator-network.org</ext-link>). At present, we can affirm that around 100 quality scales have been identified (<xref ref-type="bibr" rid="ref8">Conn and Rantz, 2003</xref>) as well as more than 550 different strategies to measure MQ (<xref ref-type="bibr" rid="ref7">Chac&#x00F3;n-Moscoso et al., 2016</xref>).</p>
<p>In some settings, such as medicine, a certain consensus has been reached on the use of individual quality components or items not combined into scales (<xref ref-type="bibr" rid="ref13">Herbison et al., 2006</xref>), and researchers in the social sciences also appear to be increasingly forgoing scales (<xref ref-type="bibr" rid="ref19">Littell et al., 2008</xref>). However, no empirical tests have been conducted in other areas such as psychology, which is problematic given how this decision affects the replicability of study results (<xref ref-type="bibr" rid="ref2">Anvari and Lakens, 2018</xref>).</p>
<p>Second, discrepant results were found when applying several measurement strategies in the same sample of primary studies (<xref ref-type="bibr" rid="ref20">Losilla et al., 2018</xref>). Depending on the instrument chosen, the assessment of MQ may vary. As a result, the choice of scale can lead us to treat a study differently, rely on its results to varying degrees, or even include/exclude it from a meta-analysis.</p>
<p>The discrepancies between the quality scales may be attributed to different causes: the scales measure different aspects of quality, have been constructed from different research contexts (<xref ref-type="bibr" rid="ref1">Albanese et al., 2020</xref>), or present metric deficiencies (<xref ref-type="bibr" rid="ref31">Sterne et al., 2016</xref>). This is because the makers of these scales did not follow the standards for developing measuring instruments, and their metric properties are generally unexplored.</p>
<p>This paper considers MQ based on its existing use and applications in the literature. Our approach, which draws on consequential validity (<xref ref-type="bibr" rid="ref3">Brussow, 2018</xref>) and is centered on the descriptive theory of valuation, aims to describe values without evaluating whether any one is better than others (<xref ref-type="bibr" rid="ref30">Shadish et al., 1991</xref>). For this purpose, we developed the 12-item MQ Scale (MQS) (<xref ref-type="bibr" rid="ref7">Chac&#x00F3;n-Moscoso et al., 2016</xref>). We did an exhaustive review of the literature (<xref ref-type="bibr" rid="ref23">Nunnally and Bernstein, 1994</xref>) and compiled 550 different strategies to assess MQ (the list of bibliographic references is available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S1</xref>). Subsequently, we selected the most frequent indicators of MQ, obtaining 23 items. A content validity study was then carried out. Thirty experts in meta-analysis and/or methodology participated voluntarily. All were methods group members of the Campbell Collaboration and/or the European Association of Methodology. Participants (12 women and 18 men, 20 from Europe and 10 from the United States) were contacted by e-mail or face-to-face in the biannual congresses of the associations. Their mean age was 42, with an average of 14&#x2009;years of experience on these issues. Participants evaluated the representativeness, utility, and feasibility of each item with respect to a hypothetical global construct of MQ. Finally, the 12 items that passed the cutoff point were selected and refined after an intercoder reliability study (the final version of the instrument, used as a coding manual for this work, is available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S2</xref>).</p>
<p>An advantage to this approach is that we specify the origin and reasons for the selection of these final items and, additionally, these were not limited to any particular intervention, methodology, or context. Furthermore, to bypass the common handicap of presenting a proposal without a thorough study of its metric properties, the aim of this paper was to analyze the metric properties of the scores obtained with the MQS in terms of reliability and validity evidence, as well as its dimensional structure.</p>
<p>As the concept of MQ is multidimensional, we hypothesized that we were not going to find a general factor that explained the set of 12 items. Instead, we approached the empirical study of the possible different facets (profiles) of the validity evidence of these 12 items based on a conceptual validity framework (<xref ref-type="bibr" rid="ref29">Shadish et al., 2002</xref>), and the structural dimensions that form the acronym UTOSTi; Units, Treatment, Output, Setting and Time (<xref ref-type="bibr" rid="ref5">Chac&#x00F3;n-Moscoso et al., 2014</xref>, <xref ref-type="bibr" rid="ref6">2021</xref>). Finally, we present an application in organizational training programs based on the interpretation of the scores obtained in such validity facets.</p>
</sec>
<sec id="sec5" sec-type="methods">
<label>2.</label>
<title>Methods</title>
<sec id="sec6">
<label>2.1.</label>
<title>Participants</title>
<p>Studies on training programs for workers in organizations were selected as a topic that has attracted substantial research interest (<xref ref-type="bibr" rid="ref26">Sanduvete-Chaves et al., 2009</xref>). A total of 299 full texts were selected. References from these studies are available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S3</xref>. Each study had to meet the following inclusion criteria: their research topic was training programs for workers at organizations; non-duplicated; written in English or Spanish; full text available; primary study; empirical study in which a training program was applied; and the training program was the aim of the study.</p>
</sec>
<sec id="sec7">
<label>2.2.</label>
<title>Instruments</title>
<p>MQS, available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S2</xref>, was applied. This scale presented 12 items, each with three alternatives (0, 0.5, and 1) representing, respectively, the null, medium, or total level of achievement of the criterion presented in the item.</p>
</sec>
<sec id="sec8">
<label>2.3.</label>
<title>Procedure</title>
<p>The search cut-off date for the primary studies was June 2020. The following databases were included because of the relevant issues they cover: Web of Science, SCOPUS, Springer, EBSCO Online, Medline, CINAHL, Econlit, MathSci Net, Current Contents, ERIC, and PsycINFO. The combined keywords were &#x201C;evaluation&#x201D; AND &#x201C;work&#x201D; AND &#x201C;training programs,&#x201D; with a search in title, abstract, keywords, and complete article. Additionally, authors who published most frequently on training programs for workers were contacted by e-mail to ask if they could share any other work, published or unpublished, on this topic.</p>
<p>During the initial screening, the inclusion criteria were applied to title, keywords, and abstract. The included studies were evaluated at a second stage, applying the inclusion criteria to the full texts. Two coders (SSC and FPHT) applied the criteria independently. In case of disagreements, a third coder (SCM) mediated to reach a consensus.</p>
<p>For the data extraction, the same two coders participated in 2 months of training sessions until the appropriate inter-coder reliability was met, with <italic>&#x03BA;</italic> agreement greater than 0.7 in a pilot study. Subsequently, they coded the entire sample of selected studies. Finally, the discrepancies between coders were resolved by involving a third researcher (SCM) to reach a consensus.</p>
</sec>
<sec id="sec9">
<label>2.4.</label>
<title>Data analyses</title>
<p>Using SPSS v.26, we conducted an intercoder reliability analysis after the selection phase and data extraction phase. Kappa (<italic>&#x03BA;</italic>) coefficient, a statistic specifically created to value inter-rater agreement which corrects the probability of concordance due to hazard (<xref ref-type="bibr" rid="ref22">McHugh, 2012</xref>), was computed for each item and 95% confidence intervals. <italic>&#x03BA;</italic> between 0.61 and 0.80 was considered substantial; and above 0.80, very good (<xref ref-type="bibr" rid="ref18">Landis and Koch, 1977</xref>). We then performed descriptive analyses of item scores, calculating the mean, the standard deviation, skewness, and kurtosis coefficients.</p>
<p>To obtain the validity facets that were implicit in the tool, the FACTOR software (<xref ref-type="bibr" rid="ref10">Ferrando and Lorenzo-Seva, 2017</xref>) was used. First, a parallel analysis was done using optimal implementation to determine the number of dimensions (<xref ref-type="bibr" rid="ref33">Timmerman and Lorenzo-Seva, 2011</xref>; <xref ref-type="bibr" rid="ref35">Yang and Xia, 2015</xref>); second, Exploratory Factor Analyses (EFA) were performed to extract the main dimensions (<xref ref-type="bibr" rid="ref11">Ferrando and Lorenzo-Seva, 2018</xref>). The polychoric correlation matrix (<xref ref-type="bibr" rid="ref15">Holgado-Tello et al., 2010</xref>) was used because of the ordinal metric of the variables and the non-normal data distribution. Unweighted least squares were applied as the estimation method and varimax rotation (<xref ref-type="bibr" rid="ref27">Sanduvete-Chaves et al., 2013</xref>, <xref ref-type="bibr" rid="ref28">2018</xref>).</p>
<p>Using the JASP version 0.16 software (<xref ref-type="bibr" rid="ref17">JASP Team, 2021</xref>), the reliability of the test scores was examined for each dimension obtained by calculating the McDonald&#x2019;s omega (<italic>&#x03C9;</italic>) coefficient. For item discrimination, we computed corrected item-total correlation coefficients.</p>
<p>In addition, the theoretical interpretation of each extracted dimension was analyzed according to its items. Thus, the correlation between items, the factor solution, the metric features of the dimensions, and the theoretical congruence were considered to obtain the different dimensions.</p>
<p>Once the validity facets were obtained, we presented their descriptive statistics (mean, standard deviation, reliability, and average discrimination). Finally, a theoretical interpretation was performed for the primary studies analyzed.</p>
</sec>
</sec>
<sec id="sec10" sec-type="results">
<label>3.</label>
<title>Results</title>
<sec id="sec11">
<label>3.1.</label>
<title>Selection of the studies</title>
<p><xref rid="fig1" ref-type="fig">Figure 1</xref> summarizes the selection process. A total of 2,886 studies were found in database searches and 39 were sent by the authors contacted by e-mail. Of the 2,878 nonduplicated papers found, 887 met the inclusion criteria, 299 of which were selected at random.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>PRISMA Flow chart of the study selection process (<xref ref-type="bibr" rid="ref24">Page et al., 2021</xref>).</p>
</caption>
<graphic xlink:href="fpsyg-14-1217661-g001.tif"/>
</fig>
</sec>
<sec id="sec12">
<label>3.2.</label>
<title>Intercoder reliability</title>
<p>In the study search, intercoder reliability was <italic>&#x03BA;</italic>&#x2009;=&#x2009;0.705. <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, 95% CI [0.674, 0.736]. <xref rid="tab1" ref-type="table">Table 1</xref> presents <italic>&#x03BA;</italic> values with their significance and confidence intervals that refer to the information extraction phase. <italic>&#x03BA;</italic> in items varied between 0.651 and 0.949, with an average of <italic>&#x03BA;</italic>&#x2009;=&#x2009;0.910. <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001, 95% CI (0.898, 0.922). All items obtained adequate results.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Intercoder reliability and descriptive statistics of the items.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top" colspan="3">Intercoder reliability</th>
<th align="center" valign="top" colspan="6">Descriptive statistics</th>
</tr>
<tr>
<th align="left" valign="top">Item</th>
<th align="center" valign="top">Kappa</th>
<th align="center" valign="top"><italic>LL</italic></th>
<th align="center" valign="top"><italic>UL</italic></th>
<th align="center" valign="top"><italic>M</italic></th>
<th align="center" valign="top"><italic>Mdn</italic></th>
<th align="center" valign="top"><italic>SD</italic></th>
<th align="center" valign="top">S</th>
<th align="center" valign="top">K</th>
<th align="center" valign="top">SW</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">1</td>
<td align="char" valign="top" char=".">0.651</td>
<td align="char" valign="top" char=".">0.543</td>
<td align="char" valign="top" char=".">0.759</td>
<td align="char" valign="top" char=".">0.89</td>
<td align="center" valign="top">1</td>
<td align="char" valign="top" char=".">0.21</td>
<td align="char" valign="top" char=".">&#x2212;1.53</td>
<td align="char" valign="top" char=".">0.77</td>
<td align="char" valign="top" char=".">0.51</td>
</tr>
<tr>
<td align="left" valign="top">2</td>
<td align="char" valign="top" char=".">0.912</td>
<td align="char" valign="top" char=".">0.869</td>
<td align="char" valign="top" char=".">0.955</td>
<td align="char" valign="top" char=".">0.27</td>
<td align="center" valign="top">0</td>
<td align="char" valign="top" char=".">0.35</td>
<td align="char" valign="top" char=".">0.90</td>
<td align="char" valign="top" char=".">&#x2212;0.45</td>
<td align="char" valign="top" char=".">0.72</td>
</tr>
<tr>
<td align="left" valign="top">3</td>
<td align="char" valign="top" char=".">0.783</td>
<td align="char" valign="top" char=".">0.679</td>
<td align="char" valign="top" char=".">0.887</td>
<td align="char" valign="top" char=".">0.89</td>
<td align="center" valign="top">1</td>
<td align="char" valign="top" char=".">0.30</td>
<td align="char" valign="top" char=".">&#x2212;2.46</td>
<td align="char" valign="top" char=".">4.32</td>
<td align="char" valign="top" char=".">0.4</td>
</tr>
<tr>
<td align="left" valign="top">4</td>
<td align="char" valign="top" char=".">0.784</td>
<td align="char" valign="top" char=".">0.633</td>
<td align="char" valign="top" char=".">0.935</td>
<td align="char" valign="top" char=".">0.78</td>
<td align="center" valign="top">1</td>
<td align="char" valign="top" char=".">0.41</td>
<td align="char" valign="top" char=".">&#x2212;1.35</td>
<td align="char" valign="top" char=".">&#x2212;0.08</td>
<td align="char" valign="top" char=".">0.55</td>
</tr>
<tr>
<td align="left" valign="top">5</td>
<td align="char" valign="top" char=".">0.798</td>
<td align="char" valign="top" char=".">0.412</td>
<td align="char" valign="top" char=".">1.184</td>
<td align="char" valign="top" char=".">0.99</td>
<td align="center" valign="top">1</td>
<td align="char" valign="top" char=".">0.04</td>
<td align="char" valign="top" char=".">&#x2212;12.14</td>
<td align="char" valign="top" char=".">145.97</td>
<td align="char" valign="top" char=".">0.05</td>
</tr>
<tr>
<td align="left" valign="top">6</td>
<td align="char" valign="top" char=".">0.938</td>
<td align="char" valign="top" char=".">0.899</td>
<td align="char" valign="top" char=".">0.977</td>
<td align="char" valign="top" char=".">0.24</td>
<td align="center" valign="top">0</td>
<td align="char" valign="top" char=".">0.38</td>
<td align="char" valign="top" char=".">1.26</td>
<td align="char" valign="top" char=".">&#x2212;0.13</td>
<td align="char" valign="top" char=".">0.63</td>
</tr>
<tr>
<td align="left" valign="top">7</td>
<td align="char" valign="top" char=".">0.949</td>
<td align="char" valign="top" char=".">0.918</td>
<td align="char" valign="top" char=".">0.980</td>
<td align="char" valign="top" char=".">0.52</td>
<td align="center" valign="top">0.5</td>
<td align="char" valign="top" char=".">0.39</td>
<td align="char" valign="top" char=".">&#x2212;0.05</td>
<td align="char" valign="top" char=".">&#x2212;1.32</td>
<td align="char" valign="top" char=".">0.81</td>
</tr>
<tr>
<td align="left" valign="top">8</td>
<td align="char" valign="top" char=".">0.86</td>
<td align="char" valign="top" char=".">0.768</td>
<td align="char" valign="top" char=".">0.952</td>
<td align="char" valign="top" char=".">0.91</td>
<td align="center" valign="top">1</td>
<td align="char" valign="top" char=".">0.22</td>
<td align="char" valign="top" char=".">&#x2212;2.309</td>
<td align="char" valign="top" char=".">4.78</td>
<td align="char" valign="top" char=".">0.46</td>
</tr>
<tr>
<td align="left" valign="top">9</td>
<td align="char" valign="top" char=".">0.86</td>
<td align="char" valign="top" char=".">0.811</td>
<td align="char" valign="top" char=".">0.909</td>
<td align="char" valign="top" char=".">0.51</td>
<td align="center" valign="top">0.5</td>
<td align="char" valign="top" char=".">0.44</td>
<td align="char" valign="top" char=".">&#x2212;0.04</td>
<td align="char" valign="top" char=".">&#x2212;1.73</td>
<td align="char" valign="top" char=".">0.75</td>
</tr>
<tr>
<td align="left" valign="top">10</td>
<td align="char" valign="top" char=".">0.775</td>
<td align="char" valign="top" char=".">0.673</td>
<td align="char" valign="top" char=".">0.877</td>
<td align="char" valign="top" char=".">0.57</td>
<td align="center" valign="top">0.5</td>
<td align="char" valign="top" char=".">0.18</td>
<td align="char" valign="top" char=".">1.85</td>
<td align="char" valign="top" char=".">2.24</td>
<td align="char" valign="top" char=".">0.44</td>
</tr>
<tr>
<td align="left" valign="top">11</td>
<td align="char" valign="top" char=".">0.809</td>
<td align="char" valign="top" char=".">0.735</td>
<td align="char" valign="top" char=".">0.883</td>
<td align="char" valign="top" char=".">0.87</td>
<td align="center" valign="top">1</td>
<td align="char" valign="top" char=".">0.23</td>
<td align="char" valign="top" char=".">&#x2212;1.52</td>
<td align="char" valign="top" char=".">1.25</td>
<td align="char" valign="top" char=".">0.55</td>
</tr>
<tr>
<td align="left" valign="top">12</td>
<td align="char" valign="top" char=".">0.884</td>
<td align="char" valign="top" char=".">0.829</td>
<td align="char" valign="top" char=".">0.939</td>
<td align="char" valign="top" char=".">0.36</td>
<td align="center" valign="top">0</td>
<td align="char" valign="top" char=".">0.48</td>
<td align="char" valign="top" char=".">0.60</td>
<td align="char" valign="top" char=".">&#x2212;1.65</td>
<td align="char" valign="top" char=".">0.61</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>S, skewness; K, kurtosis; SW, shapiro-wilk normality test. All Kappa and SW obtained <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="sec13">
<label>3.3.</label>
<title>Descriptive analysis</title>
<p>The database used in this article is available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S4</xref>. <xref rid="tab1" ref-type="table">Table 1</xref> presents descriptive statistics for the 12 items. The distributions obtained for each of the items highlighted their skewness. The median was 1 for most items. The means ranged between 0.24 and 0.99, the standard deviations were between 0.04 and 0.48, and there was no normal distribution of the items.</p>
<p>Items 5 and 8 presented means over 0.9, which implies that they lack the capacity to discriminate. Items 1, 5, 8, 10, and 11 obtained low variability, with SD below 0.25 (for example, in item 5, 99% of the studies fell into the category 1). Skewness was negative and less than &#x2212;2.3 for items 3, 5 and 8. Finally, kurtosis exceeded 4.3 for items 3, 5, and 8.</p>
<p>To analyze the relationship between items, <xref rid="tab2" ref-type="table">Table 2</xref> presents the bivariate polychoric correlation matrix. Based on the associations between items, the highest positive bivariate correlations were between items 6 and 7 (<italic>r</italic>&#x2009;=&#x2009;0.77) and between items 9 and 11 (<italic>r</italic>&#x2009;=&#x2009;0.73). Additionally, items 5 and 8 were related (<italic>r</italic>&#x2009;=&#x2009;0.47) and behaved differently than the remaining items, since their correlations with the others were negative and/or low. This may be related to the small discrimination capacity and variability that items 5 and 8 presented in <xref rid="tab1" ref-type="table">Table 1</xref>.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Polychoric correlation matrix.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Item</th>
<th align="center" valign="top">1</th>
<th align="center" valign="top">2</th>
<th align="center" valign="top">3</th>
<th align="center" valign="top">4</th>
<th align="center" valign="top">5</th>
<th align="center" valign="top">6</th>
<th align="center" valign="top">7</th>
<th align="center" valign="top">8</th>
<th align="center" valign="top">9</th>
<th align="center" valign="top">10</th>
<th align="center" valign="top">11</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">1</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">2</td>
<td align="center" valign="top">0.30</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">3</td>
<td align="center" valign="top">0.41</td>
<td align="center" valign="top">0.32</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">4</td>
<td align="center" valign="top">&#x2212;0.20</td>
<td align="center" valign="top">&#x2212;0.73</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">5</td>
<td align="center" valign="top">&#x2212;0.07</td>
<td align="center" valign="top">&#x2212;0.33</td>
<td align="center" valign="top">&#x2212;0.95</td>
<td align="center" valign="top">0.38</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">6</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.58</td>
<td align="center" valign="top">0.22</td>
<td align="center" valign="top">&#x2212;0.14</td>
<td align="center" valign="top">&#x2212;0.07</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">7</td>
<td align="center" valign="top">0.08</td>
<td align="center" valign="top">0.64</td>
<td align="center" valign="top">0.33</td>
<td align="center" valign="top">&#x2212;0.29</td>
<td align="center" valign="top">&#x2212;0.62</td>
<td align="center" valign="top">0.77</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">8</td>
<td align="center" valign="top">&#x2212;0.04</td>
<td align="center" valign="top">&#x2212;0.33</td>
<td align="center" valign="top">&#x2212;0.11</td>
<td align="center" valign="top">0.09</td>
<td align="center" valign="top">0.47</td>
<td align="center" valign="top">&#x2212;0.35</td>
<td align="center" valign="top">&#x2212;0.67</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">9</td>
<td align="center" valign="top">0.29</td>
<td align="center" valign="top">0.42</td>
<td align="center" valign="top">0.25</td>
<td align="center" valign="top">&#x2212;0.47</td>
<td align="center" valign="top">&#x2212;0.32</td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">0.42</td>
<td align="center" valign="top">&#x2212;0.21</td>
<td align="center" valign="top">1</td>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">10</td>
<td align="center" valign="top">0.48</td>
<td align="center" valign="top">0.41</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">&#x2212;0.29</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.09</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">&#x2212;0.01</td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">1</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">11</td>
<td align="center" valign="top">0.33</td>
<td align="center" valign="top">0.51</td>
<td align="center" valign="top">0.29</td>
<td align="center" valign="top">&#x2212;0.66</td>
<td align="center" valign="top">&#x2212;0.55</td>
<td align="center" valign="top">0.23</td>
<td align="center" valign="top">0.36</td>
<td align="center" valign="top">&#x2212;0.07</td>
<td align="center" valign="top">0.73</td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">1</td>
</tr>
<tr>
<td align="left" valign="top">12</td>
<td align="center" valign="top">0.35</td>
<td align="center" valign="top">&#x2212;0.07</td>
<td align="center" valign="top">0.52</td>
<td align="center" valign="top">0.08</td>
<td align="center" valign="top">&#x2212;0.22</td>
<td align="center" valign="top">&#x2212;0.16</td>
<td align="center" valign="top">&#x2212;0.15</td>
<td align="center" valign="top">0.19</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.06</td>
<td align="center" valign="top">0.19</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="sec14">
<label>3.4.</label>
<title>Study of dimensionality</title>
<p>A parallel analysis was conducted to obtain empirical evidence about the number of factors that the scale presented (see <xref rid="tab3" ref-type="table">Table 3</xref>). The results suggested no unidimensionality.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Parallel analysis.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Variable</th>
<th align="center" valign="top">% of <italic>S</italic><sup>2</sup> in real data</th>
<th align="center" valign="top">Mean random % of <italic>S</italic><sup>2</sup></th>
<th align="center" valign="top">95 <italic>P</italic> random % of <italic>S</italic><sup>2</sup></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">1</td>
<td align="char" valign="top" char=".">19.91&#x002A;</td>
<td align="char" valign="top" char=".">17.66</td>
<td align="char" valign="top" char=".">20.38</td>
</tr>
<tr>
<td align="left" valign="top">2</td>
<td align="char" valign="top" char=".">15.98&#x002A;</td>
<td align="char" valign="top" char=".">15.35</td>
<td align="char" valign="top" char=".">17.31</td>
</tr>
<tr>
<td align="left" valign="top">3</td>
<td align="char" valign="top" char=".">13.54&#x002A;</td>
<td align="char" valign="top" char=".">13.48</td>
<td align="char" valign="top" char=".">14.85</td>
</tr>
<tr>
<td align="left" valign="top">4</td>
<td align="char" valign="top" char=".">12.94&#x002A;</td>
<td align="char" valign="top" char=".">11.74</td>
<td align="char" valign="top" char=".">12.82</td>
</tr>
<tr>
<td align="left" valign="top">5</td>
<td align="char" valign="top" char=".">10.25&#x002A;</td>
<td align="char" valign="top" char=".">10.18</td>
<td align="char" valign="top" char=".">11.37</td>
</tr>
<tr>
<td align="left" valign="top">6</td>
<td align="char" valign="top" char=".">7.69</td>
<td align="char" valign="top" char=".">8.72</td>
<td align="char" valign="top" char=".">9.82</td>
</tr>
<tr>
<td align="left" valign="top">7</td>
<td align="char" valign="top" char=".">6.55</td>
<td align="char" valign="top" char=".">7.34</td>
<td align="char" valign="top" char=".">8.39</td>
</tr>
<tr>
<td align="left" valign="top">8</td>
<td align="char" valign="top" char=".">5.45</td>
<td align="char" valign="top" char=".">6.03</td>
<td align="char" valign="top" char=".">7.29</td>
</tr>
<tr>
<td align="left" valign="top">9</td>
<td align="char" valign="top" char=".">4.16</td>
<td align="char" valign="top" char=".">4.62</td>
<td align="char" valign="top" char=".">5.76</td>
</tr>
<tr>
<td align="left" valign="top">10</td>
<td align="char" valign="top" char=".">2.27</td>
<td align="char" valign="top" char=".">3.15</td>
<td align="char" valign="top" char=".">4.45</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>P</italic>, percentile.</p>
</table-wrap-foot>
</table-wrap>
<p>Next, we set out to identify the relevant factors from the twelve items, based on the chosen validity framework. Based on the previous parallel analysis, the first EFA conducted to extract dimensions was set to five factors. The rotated loadings (see <xref rid="tab4" ref-type="table">Table 4</xref>), interpreted from a theoretical point of view, led us to form a dimension composed of items 1 (inclusion and exclusion criteria for the units), 3 (attrition), 4 (attrition between groups), and 12 (statistical methods for imputing missing data). This dimension (factor 1 -F1-) could be interpreted as a measure of external validity, as items are focused on the representativeness of participants from a delimited population, selection criteria, the possible problem of a loss of participants during the study, and the method used to compute any missing data.</p>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption>
<p>Rotated matrix (exploratory factor analysis) set to five factors.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Item</th>
<th align="center" valign="top">Factor 1</th>
<th align="center" valign="top">Factor 2</th>
<th align="center" valign="top">Factor 3</th>
<th align="center" valign="top">Factor 4</th>
<th align="center" valign="top">Factor 5</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">1</td>
<td align="center" valign="top">0.48</td>
<td/>
<td align="center" valign="top">&#x2212;0.44</td>
<td/>
<td align="center" valign="top">0.34</td>
</tr>
<tr>
<td align="left" valign="top">2</td>
<td/>
<td align="center" valign="top">0.22</td>
<td/>
<td/>
<td align="center" valign="top">0.57</td>
</tr>
<tr>
<td align="left" valign="top">3</td>
<td align="center" valign="top">0.58</td>
<td align="center" valign="top">0.31</td>
<td align="center" valign="top">&#x2212;0.21</td>
<td/>
<td align="center" valign="top">&#x2212;0.24</td>
</tr>
<tr>
<td align="left" valign="top">4</td>
<td align="center" valign="top">0.98</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">5</td>
<td/>
<td/>
<td align="center" valign="top">0.97</td>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">6</td>
<td/>
<td align="center" valign="top">0.53</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">7</td>
<td/>
<td align="center" valign="top">0.97</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">8</td>
<td/>
<td/>
<td align="center" valign="top">0.33</td>
<td/>
<td/>
</tr>
<tr>
<td align="left" valign="top">9</td>
<td/>
<td align="center" valign="top">0.31</td>
<td align="center" valign="top">0.23</td>
<td align="center" valign="top">0.45</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">10</td>
<td/>
<td/>
<td/>
<td/>
<td align="center" valign="top">0.99</td>
</tr>
<tr>
<td align="left" valign="top">11</td>
<td/>
<td/>
<td/>
<td align="center" valign="top">0.97</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">12</td>
<td align="center" valign="top">0.63</td>
<td/>
<td/>
<td/>
<td align="center" valign="top">0.63</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Loadings lower than absolute 0.20 were omitted.</p>
</table-wrap-foot>
</table-wrap>
<p>To guarantee that the four items mentioned (items 1, 3, 4, and 12) could be interpreted as a single dimension, a parallel analysis was conducted, introducing these four items exclusively. According to the results, a single dimension was recommended. The reliability of this dimension was <italic>&#x03C9;</italic>&#x2009;=&#x2009;0.60, and the discrimination of the items was 0.21 for item 1, 0.45 for items 3 and 4, and 0.25 for item 12.</p>
<p>Once F1 was defined, the next step was to obtain empirical evidence that could statistically support other possible factors. Thus, a new EFA was conducted after omitting the items that formed F1 (items 1, 3, 4, and 12). The aim was to extract the next most relevant factor, avoiding redundant variability that could hamper its interpretation. Following this procedure, and after interpreting the results shown on <xref rid="tab4" ref-type="table">Table 4</xref>, a second factor (F2) that could be interpreted as internal validity was obtained (<xref ref-type="bibr" rid="ref29">Shadish et al., 2002</xref>). This second factor was formed by items 2 (methodology or design), 6 (follow-up period), 7 (measurement occasions for each dependent variable), and 10 (control techniques). These items focus on the level of manipulation, the number of groups, the measurements of relevant dependent variables to be measured, and the techniques applied to control for potential sources of error. As shown on <xref rid="tab2" ref-type="table">Table 2</xref>, the bivariate correlations between items 2, 6 and 7 are high, between 0.58 and 0.77.</p>
<p>A parallel analysis conducted only with these items (items 2, 6, 7, and 10) yielded a single dimension. The reliability coefficient of F2 was <italic>&#x03C9;</italic>&#x2009;=&#x2009;0.70. The inclusion of item 10 negatively affected the reliability of F2 (<italic>&#x03C9;</italic>&#x2009;=&#x2009;0.77 without item 10); however, from a content validity perspective, it was considered that the information contained in item 10 was relevant to define this dimension, as it referred to control techniques directly related to internal validity. The discriminations of the items were 0.56 (item 2), 0.55 (item 6), 0.65 (item 7), and 0.17 (item 10).</p>
<p>Once F1 and F2 were defined, a third EFA was performed with the remaining items 5, 8, 9, and 11. The results, strongly supported by the theory and according to results obtained in <xref rid="tab4" ref-type="table">Table 4</xref>, showed a dimension defined by items 9 and 11. Additionally, they presented a high correlation in <xref rid="tab2" ref-type="table">Table 2</xref> (<italic>r</italic>&#x2009;=&#x2009;0.73). These items measured the standardization of the dependent variables (item 9), and the construct definition (item 11). Therefore, this dimension (factor 3 -F3-) was interpreted as construct validity (<xref ref-type="bibr" rid="ref29">Shadish et al., 2002</xref>), because it is focused on explaining the concept, model, or schematic idea measured as a dependent variable, the way the theoretical dimensions are empirically defined, and the standardization of the tool used to measure the dependent variable.</p>
<p>Based on the parallel analysis conducted after including items 9 and 11, a single dimension was recommended. The reliability of F3 was <italic>&#x03C9;</italic>&#x2009;=&#x2009;0.65 and the discrimination of the two items was 0.48.</p>
<p>Based on the results obtained in <xref rid="tab4" ref-type="table">Table 4</xref>, items 5 (exclusions after assignment) and 8 (measures in pretest appear in post-test) were difficult to integrate, though they appeared to be linked. If we hypothesize that, together, these items form a dimension, its metric indices would be very low, with a reliability coefficient of <italic>&#x03C9;</italic>&#x2009;=&#x2009;0.13 and discrimination indexes of 0.07. These items were predicted to present problems, since they had no variability, did not discriminate between studies, and presented excessive skewness, as shown on <xref rid="tab1" ref-type="table">Table 1</xref>. We decided to exclude these two items from the defined F3 because they did not fit items 9 and 11, and presented low theoretical congruence in this dimension.</p>
</sec>
<sec id="sec15">
<label>3.5.</label>
<title>Interpretation of the study scores in each validity facet and acquisition of possible profiles</title>
<p><xref rid="tab5" ref-type="table">Table 5</xref> shows the descriptive statistics for each of the theoretical validity facets obtained. To interpret the items in the study, each received a score of 0 (low), 0.5 (medium) or 1 (high).</p>
<table-wrap position="float" id="tab5">
<label>Table 5</label>
<caption>
<p>Descriptive statistics for each facet.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top">F1</th>
<th align="center" valign="top">F2</th>
<th align="center" valign="top">F3</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Possible range</td>
<td align="char" valign="top" char=".">0&#x2013;4</td>
<td align="char" valign="top" char=".">0&#x2013;4</td>
<td align="char" valign="top" char=".">0&#x2013;2</td>
</tr>
<tr>
<td align="left" valign="top">Mean</td>
<td align="char" valign="top" char=".">3.98</td>
<td align="char" valign="top" char=".">1.59</td>
<td align="char" valign="top" char=".">1.38</td>
</tr>
<tr>
<td align="left" valign="top">Standard deviation</td>
<td align="char" valign="top" char=".">0.89</td>
<td align="char" valign="top" char=".">0.96</td>
<td align="char" valign="top" char=".">0.59</td>
</tr>
<tr>
<td align="left" valign="top">McDonald&#x2019;s &#x03C9;</td>
<td align="char" valign="top" char=".">0.60</td>
<td align="char" valign="top" char=".">0.70</td>
<td align="char" valign="top" char=".">0.65</td>
</tr>
<tr>
<td align="left" valign="top">Discrimination</td>
<td align="char" valign="top" char=".">0.34</td>
<td align="char" valign="top" char=".">0.48</td>
<td align="char" valign="top" char=".">0.48</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>F1, external validity facet; F2, internal validity facet; F3, construct validity facet.</p>
</table-wrap-foot>
</table-wrap>
<p>Formed by items 1 (inclusion and exclusion criteria for units), 3 (attrition), 4 (attrition between groups), and 12 (statistical methods for imputing missing data), F1 assesses external validity. F1 answers the question: how accurately are the population and the selection criteria for units defined? Studies with high scores in F1 should be characterized by a well-defined reference population, explicit selection criteria for the units that form the sample, and the monitoring of possible unit losses over the course of the study that could compromise the representativeness of the results.</p>
<p>Formed by items 2 (methodology or design), 6 (follow-up period), 7 (measurement occasions for each dependent variable), and 10 (control techniques), F2 assesses internal validity. It answers the questions: What are the relevant variables of the study? How and when are they manipulated and measured? What is done to control for possible sources of error? Studies with a great capacity to manipulate variables and control for threats to validity (<xref ref-type="bibr" rid="ref16">Holgado-Tello et al., 2016</xref>) would receive high scores in F2. Other factors that affect scores in F2 include clearly established criteria for assigning the units to study conditions and the quantity of measures before, during, and after the interventions.</p>
<p>Finally, F3 assesses construct validity. Formed by items 9 (standardization of the dependent variables) and 11 (construct definition of outcomes), it answers the question: How are dimensions empirically operationalized from their conceptual referents? Studies with high scores in F3 clearly present the referent or conceptual model, an empirical operationalization of its components, and standardized measurements.</p>
<p><xref rid="tab6" ref-type="table">Table 6</xref> provides an example of how scores can be interpreted in each study based on the scores for each item, as well as an overall assessment of the study based on its mean per facet. The scores of the facets were obtained calculating the average of the items that comprised them. This average ranged from 0 to 1 (low if &#x003C;0.5; medium if ranging from 0.5&#x2013;75, both values included; and high for values &#x003E;0.75). For example, study 3 had a score of 2 in F1 (average&#x2009;=&#x2009;0.5; medium quality); in F2, 1.5 (average&#x2009;=&#x2009;0.37; low quality); and in F3, 1.5 (average&#x2009;=&#x2009;0.75, medium quality).</p>
<table-wrap position="float" id="tab6">
<label>Table 6</label>
<caption>
<p>Scores of four studies in each item, and average values in each facet.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="left" valign="top" colspan="10">Validity facets</th>
<th align="center" valign="top" colspan="3">Global</th>
</tr>
<tr>
<th align="left" valign="middle" rowspan="2">Studies</th>
<th align="center" valign="middle" colspan="4">F1</th>
<th align="center" valign="middle" colspan="4">F2</th>
<th align="center" valign="middle" colspan="2">F3</th>
<th align="center" valign="middle" colspan="3">Facets</th>
</tr>
<tr>
<th align="center" valign="middle">I<sub>1</sub></th>
<th align="center" valign="middle">I <sub>3</sub></th>
<th align="center" valign="middle">I <sub>4</sub></th>
<th align="center" valign="middle">I <sub>12</sub></th>
<th align="center" valign="middle">I <sub>2</sub></th>
<th align="center" valign="middle">I <sub>6</sub></th>
<th align="center" valign="middle">I <sub>7</sub></th>
<th align="center" valign="middle">I <sub>10</sub></th>
<th align="center" valign="middle">I <sub>9</sub></th>
<th align="center" valign="middle">I <sub>11</sub></th>
<th align="center" valign="middle">F1</th>
<th align="center" valign="middle">F2</th>
<th align="center" valign="middle">F3</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">3</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0.37</td>
<td align="center" valign="middle">0.75</td>
</tr>
<tr>
<td align="left" valign="middle">6</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0.75</td>
<td align="center" valign="middle">0.88</td>
<td align="center" valign="middle">1</td>
</tr>
<tr>
<td align="left" valign="middle">138</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">&#x2013;</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="middle">0.33</td>
<td align="center" valign="middle">0.12</td>
<td align="center" valign="middle">0.25</td>
</tr>
<tr>
<td align="left" valign="middle">299</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">1</td>
<td align="center" valign="middle">0.75</td>
<td align="center" valign="middle">1</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>F1, external validity facet; F2, internal validity facet; F3, construct validity facet; I, item. Red, low level; yellow, medium level; green, high level of quality.</p>
</table-wrap-foot>
</table-wrap>
<p>The evaluations of the 299 coded studies, by items and facets, are available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S4</xref>. <xref rid="fig2" ref-type="fig">Figure 2</xref> represents part of this database, showing one line for each study, scores (0, 0.5 or 1) for each item, and the average score in each facet.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Example of primary study coding. 9, not applicable. IT, item; F1, external validity facet; F2, internal validity facet; F3, construct validity facet. Red, low level; yellow, medium level; green, high level of quality. The complete items are available in <xref rid="tab8" ref-type="table">Table 8</xref>.</p>
</caption>
<graphic xlink:href="fpsyg-14-1217661-g002.tif"/>
</fig>
<p><xref rid="tab7" ref-type="table">Table 7</xref> presents the frequencies and percentages of studies of the sample that had a low, medium, and high level of quality in each facet. Based on the results obtained, most of the studies that comprised the sample presented medium levels of quality in external validity, low levels in internal validity, and high levels in construct validity.</p>
<table-wrap position="float" id="tab7">
<label>Table 7</label>
<caption>
<p>Distribution of studies by quality level in each facet (frequencies and %).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Level of quality</th>
<th align="center" valign="top">F1 External v.</th>
<th align="center" valign="top">F2 Internal v.</th>
<th align="center" valign="top">F3 Construct v.</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Low</td>
<td align="char" valign="top" char="(">35 (11.7)</td>
<td align="char" valign="top" char="(">185 (61.9)</td>
<td align="char" valign="top" char="(">56 (18.7)</td>
</tr>
<tr>
<td align="left" valign="top">Medium</td>
<td align="char" valign="top" char="(">127 (42.5)</td>
<td align="char" valign="top" char="(">63 (21.1)</td>
<td align="char" valign="top" char="(">71 (23.8)</td>
</tr>
<tr>
<td align="left" valign="top">High</td>
<td align="char" valign="top" char="(">137 (45.8)</td>
<td align="char" valign="top" char="(">51 (17.0)</td>
<td align="char" valign="top" char="(">172 (57.5)</td>
</tr>
<tr>
<td align="left" valign="top"><italic>Total</italic></td>
<td align="char" valign="top" char="("><italic>299 (100)</italic></td>
<td align="char" valign="top" char="("><italic>299 (100)</italic></td>
<td align="char" valign="top" char="("><italic>299 (100)</italic></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>F, facet; v., validity. Percentages are presented in brackets. For each facet (F1, F2, and F3), the most frequent level (low, medium, or high) is marked.</p>
</table-wrap-foot>
</table-wrap>
<p><xref rid="tab8" ref-type="table">Table 8</xref> presents the resulting MQS, ready to be used to measure the MQ in primary studies (MQS is also available in a printable version in <xref ref-type="supplementary-material" rid="SM1">Supplementary material S5</xref>).</p>
<table-wrap position="float" id="tab8">
<label>Table 8</label>
<caption>
<p>Methodological quality scale (final version).</p>
</caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td align="left" valign="middle" colspan="3"><bold>Facet 1. External validity</bold></td>
</tr>
<tr>
<td align="left" valign="middle">Item 1</td>
<td align="left" valign="top" colspan="2"><bold>Inclusion and exclusion criteria for the units provided</bold>: explicit reasons provided as to why certain units (usually people) were able to participate in the study and others were not:<break/><list list-type="simple">
<list-item>
<p><bold>0. No:</bold> no explicit selection criteria for units AND with exceptions in their application; information unavailable.</p>
</list-item>
<list-item>
<p><bold>0.5. Intermediate:</bold> explicit selection criteria for units OR applied to all potential participants.</p>
</list-item>
<list-item>
<p><bold>1. Yes (replicable):</bold> explicit selection criteria for units AND applied to all potential participants.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 2</td>
<td align="left" valign="top" colspan="2"><bold>Attrition:</bold> loss of units. In randomized experiments, this refers to loss that occurred after the random assignment, i.e., the number of participants from the initial sample that did not conclude the study (e.g., N pre minus N post).<break/><list list-type="simple">
<list-item>
<p><bold>0. Unspecified:</bold> information is not available and cannot be calculated AND reasons for loss of units are not specified.</p>
</list-item>
<list-item>
<p><bold>0.5. Intermediate:</bold> number of units lost is specified or can be calculated OR reasons for loss of units are specified.</p>
</list-item>
<list-item>
<p><bold>1. Specified:</bold> no units are lost, or number of units lost is specified or can be calculated AND reasons for loss of units are specified.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 3</td>
<td align="left" valign="top" colspan="2"><bold>Attrition between groups:</bold> this item evaluated the differences in attrition between two groups.<break/><list list-type="simple">
<list-item>
<p><bold>0. Unspecified:</bold> information is not available and cannot be calculated AND reasons for attrition between groups are not specified.</p>
</list-item>
<list-item>
<p><bold>0.5. Intermediate:</bold> number of lost units is specified or can be calculated OR reasons for attrition between groups are specified.</p>
</list-item>
<list-item>
<p><bold>1. Specified:</bold> no units were lost, or number of lost units is specified or can be calculated AND reason/s for the attrition between groups is/are specified.</p>
</list-item>
<list-item>
<p><bold>9. Not applicable:</bold> no cross-group comparison.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 4</td>
<td align="left" valign="top" colspan="2"><bold>Statistical methods for imputing missing data</bold>: to estimate what the study would have yielded had there been no attrition:<break/><list list-type="simple">
<list-item>
<p><bold>0. High risk:</bold> it is not clear if there was attrition, or there was attrition and calculations to estimate effects were carried out without imputing missing data.</p>
</list-item>
<list-item>
<p><bold>0.5. Medium risk:</bold> values for the missing data points were imputed so they could be included in the analyses. The method used was specified, i.e., sample mean substitution, last value forward method for longitudinal data sets, hot deck imputation, single imputation (e.g., imputation, regression imputation), or multiple imputation (e.g., likelihood ratio test after multiple imputation). The reasons for choosing the specific method were not specified.</p>
</list-item>
<list-item>
<p><bold>1. Low risk:</bold> there was no attrition or values for the missing data points were imputed so they could be included in the analyses; and the specific method used AND the reasons for choosing the specific method were specified.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Total facet 1</td>
<td align="left" valign="top" colspan="2"><bold>External validity score:</bold><break/>Add the scores obtained in items 1&#x2013;4 and divide by the number of items. If item 3 is not applicable, do not add a score for that item and divide the sum of items 1, 2 and 4 by 3.</td>
</tr>
<tr>
<td align="left" valign="middle" colspan="3"><bold>Facet 2. Internal validity</bold></td>
</tr>
<tr>
<td align="left" valign="middle">Item 5</td>
<td align="left" valign="top" colspan="2"><bold>Methodology or design:</bold> something an experimenter could manipulate or control in an experiment to help address a threat to validity:<break/><list list-type="simple">
<list-item>
<p><bold>0. Pre-experimental/others</bold> (questionnaires/observational/naturalistic): a study with only one group and a maximum of two measurement occasions for the same dependent variable (e.g., pre-post design); or when there are two groups and only one measure (e.g., control-experimental design).</p>
</list-item>
<list-item>
<p><bold>0.5. Quasi-experimental</bold> (two groups without randomized assignment) non-equivalent control groups with pre-test and post-test; or one group with three or more measures of the same dependent variable (even without pretest): an experiment (exploration of the effects of manipulating a variable) in which units are not randomly assigned to conditions.</p>
</list-item>
<list-item>
<p><bold>1. Experimental; randomized:</bold> an experiment (exploration of the effects of manipulating a variable) in which units are randomly assigned to conditions.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 6</td>
<td align="left" valign="top" colspan="2"><bold>Follow-up period</bold>: the amount of time between the first post-intervention measurements and any additional measurements. When the study presented more than one follow-up period, the longest was considered.<break/><list list-type="simple">
<list-item>
<p><bold>0.</bold> No follow-up or less than 2 months.</p>
</list-item>
<list-item>
<p><bold>0.5.</bold> Between two and 6 months (both included).</p>
</list-item>
<list-item>
<p><bold>1.</bold> More than 6 months.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 7</td>
<td align="left" valign="top" colspan="2"><bold>Measurement occasions for each dependent variable</bold>: this item specified when the measurements were taken.<break/><list list-type="simple">
<list-item>
<p><bold>0. Post-intervention only:</bold> all measurements were taken after the intervention.</p>
</list-item>
<list-item>
<p><bold>0.5. Pre- and post-intervention:</bold> some measurements were taken before and immediately after the intervention.</p>
</list-item>
<list-item>
<p><bold>1. Pre-, post-intervention and follow-up period:</bold> some measurements were taken before, immediately after the intervention, and again at a later date.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 8</td>
<td align="left" valign="top" colspan="2"><bold>Control techniques</bold>:<break/><list list-type="simple">
<list-item>
<p><bold>0. None:</bold> no control technique is specified or described.</p>
</list-item>
<list-item>
<p><bold>0.5 Masking OR other/s:</bold> masking, also known as double-blinding, refers to a procedure that prevented participants and/or experimenters from knowing the hypotheses; OR any other control technique was used (e.g., matching, stratifying, counterbalancing, constant, participant as own experimental control -longitudinal-).</p>
</list-item>
<list-item>
<p><bold>1. Masking AND other:</bold> masking AND at least one other control technique.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Total facet 2</td>
<td align="left" valign="top" colspan="2"><bold>Internal validity score:</bold><break/>Add the scores obtained in items 5&#x2013;8 and divide by the number of items (4).</td>
</tr>
<tr>
<td align="left" valign="middle" colspan="3"><bold>Facet 3. Construct validity</bold></td>
</tr>
<tr>
<td align="left" valign="middle">Item 9</td>
<td align="left" valign="top" colspan="2"><bold>Standardization of the dependent variables:</bold> level of normalization of the tool to measure the variable that varied in response to the independent variable (also called effect or outcome).<break/><list list-type="simple">
<list-item>
<p><bold>0. Low standardization (self-reports and <italic>post hoc</italic> records)</bold>: all measurements were taken using <italic>ad hoc</italic> tools, developed in a specific situation, and without any study of their psychometric properties.</p>
</list-item>
<list-item>
<p><bold>0.5. Medium standardization</bold>: at least one measurement was taken using structured tools with ONE study of their psychometric properties (reliability or one form of validity evidence).</p>
</list-item>
<list-item>
<p><bold>1. High standardization</bold>: at least one measurement was taken using structured tools. At least TWO studies of their psychometric properties (reliability, validity, construction of scaling) were carried out.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Item 10</td>
<td align="left" valign="top" colspan="2"><bold>Construct definition of outcome</bold>: explanation of the concept, model, or schematic idea measured as a dependent variable:<break/><list list-type="simple">
<list-item>
<p><bold>0. No definition:</bold> no concept treated as a dependent variable was measured in a conceptual or empirical way.</p>
</list-item>
<list-item>
<p><bold>0.5. Vague definition:</bold> at least one concept treated as a dependent variable was defined in a conceptual and/or empirical way.</p>
</list-item>
<list-item>
<p><bold>1. Replicable by reader in own setting:</bold> all concepts treated as dependent variables were defined in a conceptual and empirical way.</p>
</list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Total facet 3</td>
<td align="left" valign="top" colspan="2"><bold>Construct validity score:</bold><break/>Add the scores obtained in items 9 and 10 and divide by the number of items (2).</td>
</tr>
<tr>
<td align="left" valign="middle" colspan="3"><bold>INTERPRETATION for each type of validity (facet):</bold></td>
</tr>
<tr>
<td align="left" valign="middle">&#x003C;0.5 Low</td>
<td align="center" valign="middle">[0.5&#x2013;0.75] Medium</td>
<td align="center" valign="middle">&#x003E;0.75 High</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="sec16" sec-type="discussions">
<label>4.</label>
<title>Discussion</title>
<p>This work offers a practical approach to solve an existing problem, i.e., how to measure the varying quality levels of primary studies. It does so by analyzing the metric properties of a scale based on the standards for the constructions of measuring instruments that guarantee validity and reliability; not only analyzing content validity and intercoder reliability, but also including validity evidence based on the internal structure of the scale and metric properties of the tool (reliability based on internal coherence and discrimination). The proposed tool is available for researchers who are planning to carry out a meta-analysis. Additionally, it presents the basic elements to assess MQ of intervention programs, so professionals who are not experts in methodology can use the tool to design a new intervention or to evaluate an ongoing or completed intervention. Thus, this tool represents a first step toward guaranteeing that meta-analyses and interventions respond to replicability criteria.</p>
<p>In terms of other advantages, it is important to highlight that the inclusion criteria for the initial items of the MQS were specified. It not only considers the risk of bias associated with internal validity, but more broadly, that associated with external and construct validity. This yielded profiles with three facets, thus facilitating interpretation. It is a tool that can be applied to any type of intervention study (i.e., not only experimental methodology or randomized control trials). It is applicable in different areas of interest (not only in a specific setting). Moreover, it is easy to apply, as it is formed by ten items with three-point Likert scales.</p>
<p>A potential limitation of this study is that the MQS has only been applied to a set of studies in a specific field of intervention. However, regardless on the field of intervention, MQ varies between studies. For example, randomized studies with a high manipulation of variables can be found in in the field of health and in the social sciences too. MQ in itself is not related to the field of intervention. For this reason, we consider that MQ indicators can be studied in any context.</p>
<p>Additionally, the fact that the construct validity is comprised of only two items could be considered another limitation. However, one of the main ideas was to reduce the number of items in the scale as much as possible without lower its metric properties; this facet presents adequate validity and reliability indexes. Additionally, other tools (e.g., <xref ref-type="bibr" rid="ref34">Valentine and Cooper, 2008</xref>), contain only one item on construct validity (i.e., face validity).</p>
<p>In relation to items 5 and 8, it was not possible to justify a single factor with adequate metric indexes, such as statistical conclusion validity, based on the results obtained. Nonetheless, the content of item 5 (exclusions after assignment) can be considered that is included in items referred to attrition (items 3 and 4 -external validity-). Moreover, the content of item 8 (measures in pretest appear in posttest) can be considered that is included in items referred to methodology and follow-up (items 2 and 6, respectively -internal validity-).</p>
<p>For further research, the same sample of primary studies will be coded using other tools available in the literature to compare the results. The Risk of Bias version 2 (RoB 2) (<xref ref-type="bibr" rid="ref32">Sterne et al., 2019</xref>) will be applied for experimental designs; and for quasi-experimental designs, the Risk Of Bias In Non-randomised Studies (ROBINS-I) (<xref ref-type="bibr" rid="ref31">Sterne et al., 2016</xref>). Additionally, a cross-disciplinary guide will be drafted to inform practitioners of the design, implementation, and evaluation of intervention programs.</p>
</sec>
<sec id="sec17" sec-type="data-availability">
<title>Data availability statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/<xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
</sec>
<sec id="sec18">
<title>Author contributions</title>
<p>SC<bold>-</bold>M conceived of and designed the study, analyzed and interpreted the data, and wrote the first draft and revised it. SS<bold>-</bold>C collaborated in the acquisition of data and critically reviewed the drafting for important intellectual content. JL<bold>-</bold>L contributed to the data acquisition, analysis, and interpretation and reviewed the paper. FH<bold>-</bold>T analyzed and interpreted the data and collaborated on the writing of the article. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec19" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by the Chilean national projects FONDECYT Regular 2019, Agencia Nacional de Investigaci&#x00F3;n y Desarrollo (ANID)&#x2013;, Government of Chile (1190945); the grant PID2020-115486GB-I00 funded by the Ministerio de Ciencia e Innovaci&#x00F3;n, MCIN/AEI/10.13039/501100011033, Government of Spain; and the grant PID2020-114538RB-I00, funded by the Ministerio de Ciencia e Innovaci&#x00F3;n, Government of Spain.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to thank Rafael R. Verdugo-Mora, Elena Garrido-Jim&#x00E9;nez, Pablo Garc&#x00ED;a-Campos, Miguel A. L&#x00F3;pez-Espinoza, Erica Villoria-Fern&#x00E1;ndez, Romina Salcedo, Isabella Fioravante, Lorena E. Salazar-Aravena, Mar&#x00ED;a E. Avenda&#x00F1;o-Abello, Laura M. Peinado-Peinado, Enrique Bahamondes, Mireya Mendoza-Dom&#x00ED;nguez, Sof&#x00ED;a Tomba, and Miguel &#x00C1;ngel Fern&#x00E1;ndez-Centeno for their collaboration applying the instrument to primary studies and providing advice on improving it. Finally, the authors would like to dedicate this work to William R. Shadish, for his valuable contributions to the original idea of this work.</p>
</ack>
<sec id="sec21" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1217661/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1217661/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.DOCX" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_2.DOCX" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_3.DOCX" id="SM3" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_4.XLSX" id="SM4" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_5.DOCX" id="SM5" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Albanese</surname> <given-names>E.</given-names></name> <name><surname>B&#x00FC;tikofer</surname> <given-names>L.</given-names></name> <name><surname>Armijo-Olivo</surname> <given-names>S.</given-names></name> <name><surname>Ha</surname> <given-names>C.</given-names></name> <name><surname>Egger</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Construct validity of the physiotherapy evidence database (PEDro) quality scale for randomized trials: item response theory and factor analyses</article-title>. <source>Res. Synth. Methods</source> <volume>11</volume>, <fpage>227</fpage>&#x2013;<lpage>236</lpage>. doi: <pub-id pub-id-type="doi">10.1002/jrsm.1385</pub-id>, PMID: <pub-id pub-id-type="pmid">31733091</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anvari</surname> <given-names>F.</given-names></name> <name><surname>Lakens</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>The replicability crisis and public trust in psychological science</article-title>. <source>Compr. Results Soc. Psychol.</source> <volume>3</volume>, <fpage>266</fpage>&#x2013;<lpage>286</lpage>. doi: <pub-id pub-id-type="doi">10.1080/23743603.2019.1684822</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Brussow</surname> <given-names>J. A.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Consequential validity evidence</article-title>&#x201D; in <source>The SAGE encyclopedia of educational research, measurement, and evaluation</source>. ed. <person-group person-group-type="editor"><name><surname>Frey</surname> <given-names>B. B.</given-names></name></person-group> (<publisher-loc>Kansas</publisher-loc>: <publisher-name>SAGE</publisher-name>)</citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cano-Garc&#x00ED;a</surname> <given-names>F. J.</given-names></name> <name><surname>Gonz&#x00E1;lez-Ortega</surname> <given-names>M. C.</given-names></name> <name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Moreno-Borrego</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Evaluation of a psychological intervention for patients with chronic pain in primary care</article-title>. <source>Front. Psychol.</source> <volume>8</volume>:<fpage>435</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2017.00435</pub-id>, PMID: <pub-id pub-id-type="pmid">28386242</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Anguera</surname> <given-names>M. T.</given-names></name> <name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>S&#x00E1;nchez-Mart&#x00ED;n</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Methodological convergence of program evaluation designs</article-title>. <source>Psicothema</source> <volume>26</volume>, <fpage>91</fpage>&#x2013;<lpage>96</lpage>. doi: <pub-id pub-id-type="doi">10.7334/psicothema2013.144</pub-id>, PMID: <pub-id pub-id-type="pmid">24444735</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>Lozano-Lozano</surname> <given-names>J. A.</given-names></name> <name><surname>Portell</surname> <given-names>M.</given-names></name> <name><surname>Anguera</surname> <given-names>M. T.</given-names></name></person-group> (<year>2021</year>). <article-title>From randomized control trial to mixed methods: a practical framework for program evaluation based on methodological quality</article-title>. <source>An. Psicol.</source> <volume>37</volume>, <fpage>599</fpage>&#x2013;<lpage>608</lpage>. doi: <pub-id pub-id-type="doi">10.6018/analesps.470021</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>S&#x00E1;nchez-Mart&#x00ED;n</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>The development of a checklist to enhance methodological quality in intervention programs</article-title>. <source>Front. Psychol.</source> <volume>7</volume>:<fpage>1811</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2016.01811</pub-id>, PMID: <pub-id pub-id-type="pmid">27917143</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conn</surname> <given-names>V. S.</given-names></name> <name><surname>Rantz</surname> <given-names>M. J.</given-names></name></person-group> (<year>2003</year>). <article-title>Research methods: managing primary study quality in meta-analyses</article-title>. <source>Res. Nurs. Health</source> <volume>26</volume>, <fpage>322</fpage>&#x2013;<lpage>333</lpage>. doi: <pub-id pub-id-type="doi">10.1002/nur.10092</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Cooper</surname> <given-names>H.</given-names></name> <name><surname>Hedges</surname> <given-names>L. V.</given-names></name> <name><surname>Valentine</surname> <given-names>J. C.</given-names></name></person-group> (<year>2009</year>). <source>The handbook of research synthesis and meta-analysis</source> (<edition>2nd ed.</edition>). <publisher-loc>New York</publisher-loc>: <publisher-name>Russell Sage</publisher-name>.</citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferrando</surname> <given-names>P. J.</given-names></name> <name><surname>Lorenzo-Seva</surname> <given-names>U.</given-names></name></person-group> (<year>2017</year>). <article-title>Program FACTOR at 10: origins, development and future directions</article-title>. <source>Psicothema</source> <volume>29</volume>, <fpage>236</fpage>&#x2013;<lpage>240</lpage>. doi: <pub-id pub-id-type="doi">10.7334/psicothema2016.304</pub-id>, PMID: <pub-id pub-id-type="pmid">28438248</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferrando</surname> <given-names>P. J.</given-names></name> <name><surname>Lorenzo-Seva</surname> <given-names>U.</given-names></name></person-group> (<year>2018</year>). <article-title>Assessing the quality and appropriateness of factor solutions and factor score estimates in exploratory item factor analysis</article-title>. <source>Educ. Psychol. Meas.</source> <volume>78</volume>, <fpage>762</fpage>&#x2013;<lpage>780</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0013164417719308</pub-id>, PMID: <pub-id pub-id-type="pmid">32655169</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Hempel</surname> <given-names>S.</given-names></name> <name><surname>Suttorp</surname> <given-names>M. J.</given-names></name> <name><surname>Miles</surname> <given-names>J. N. V.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Maglione</surname> <given-names>M.</given-names></name> <name><surname>Morton</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2011</year>). <source>Empirical evidence of associations between trial quality and effect size</source>. <publisher-loc>Rockville</publisher-loc>: <publisher-name>Agency for Healthcare Research and Quality</publisher-name>.</citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herbison</surname> <given-names>P.</given-names></name> <name><surname>Hay-Smith</surname> <given-names>J.</given-names></name> <name><surname>Gillespie</surname> <given-names>W. J.</given-names></name></person-group> (<year>2006</year>). <article-title>Adjustment of meta-analyses on the basis of quality scores should be abandoned</article-title>. <source>J. Clin. Epidemiol.</source> <volume>59</volume>, <fpage>1249.e1</fpage>&#x2013;<lpage>1249.e11</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jclinepi.2006.03.008</pub-id>, PMID: <pub-id pub-id-type="pmid">17098567</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Higgins</surname> <given-names>J. P.</given-names></name> <name><surname>Ramsay</surname> <given-names>C.</given-names></name> <name><surname>Reeves</surname> <given-names>B. C.</given-names></name> <name><surname>Deeks</surname> <given-names>J. J.</given-names></name> <name><surname>Shea</surname> <given-names>B.</given-names></name> <name><surname>Valentine</surname> <given-names>J. C.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Issues relating to study design and risk of bias when including non-randomized studies in systematic reviews on the effects of interventions</article-title>. <source>Res. Synth. Methods</source> <volume>4</volume>, <fpage>12</fpage>&#x2013;<lpage>25</lpage>. doi: <pub-id pub-id-type="doi">10.1002/jrsm.1056</pub-id>, PMID: <pub-id pub-id-type="pmid">26053536</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holgado-Tello</surname> <given-names>F. P.</given-names></name> <name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Barbero-Garc&#x00ED;a</surname> <given-names>I.</given-names></name> <name><surname>Vila-Abad</surname> <given-names>E.</given-names></name></person-group> (<year>2010</year>). <article-title>Polychoric versus Pearson correlations in exploratory and confirmatory factor analysis of ordinal variables</article-title>. <source>Qual. Quant.</source> <volume>44</volume>, <fpage>153</fpage>&#x2013;<lpage>166</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11135-008-9190-y</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holgado-Tello</surname> <given-names>F. P.</given-names></name> <name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>P&#x00E9;rez-Gil</surname> <given-names>J. A.</given-names></name></person-group> (<year>2016</year>). <article-title>A simulation study of threats to validity in quasi-experimental designs: interrelationship between design, measurement, and analysis</article-title>. <source>Front. Psychol.</source> <volume>7</volume>:<fpage>897</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2016.00897</pub-id>, PMID: <pub-id pub-id-type="pmid">27378991</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll1">JASP Team</collab></person-group>. (<year>2021</year>). JASP (0.16).</citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Landis</surname> <given-names>J. R.</given-names></name> <name><surname>Koch</surname> <given-names>G. G.</given-names></name></person-group> (<year>1977</year>). <article-title>The measurement of observer agreement for categorical data</article-title>. <source>Biom.</source> <volume>33</volume>, <fpage>159</fpage>&#x2013;<lpage>174</lpage>. doi: <pub-id pub-id-type="doi">10.2307/2529310</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Littell</surname> <given-names>J. H.</given-names></name> <name><surname>Corcoran</surname> <given-names>J.</given-names></name> <name><surname>Pillai</surname> <given-names>V.</given-names></name></person-group> (<year>2008</year>). <source>Systematic reviews and meta-analysis</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Oxford University Press</publisher-name></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Losilla</surname> <given-names>J.-M.</given-names></name> <name><surname>Oliveras</surname> <given-names>I.</given-names></name> <name><surname>Marin-Garcia</surname> <given-names>J. A.</given-names></name> <name><surname>Vives</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Three risk of bias tools lead to opposite conclusions in observational research synthesis</article-title>. <source>J. Clin. Epidemiol.</source> <volume>101</volume>, <fpage>61</fpage>&#x2013;<lpage>72</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jclinepi.2018.05.021</pub-id>, PMID: <pub-id pub-id-type="pmid">29864541</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luhnen</surname> <given-names>M.</given-names></name> <name><surname>Prediger</surname> <given-names>B.</given-names></name> <name><surname>Neugebauer</surname> <given-names>E. A. M.</given-names></name> <name><surname>Mathes</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Systematic reviews of health economic evaluations: a structured analysis of characteristics and methods applied</article-title>. <source>Res. Synth. Methods</source> <volume>10</volume>, <fpage>195</fpage>&#x2013;<lpage>206</lpage>. doi: <pub-id pub-id-type="doi">10.1002/jrsm.1342</pub-id>, PMID: <pub-id pub-id-type="pmid">30761762</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McHugh</surname> <given-names>M. L.</given-names></name></person-group> (<year>2012</year>). <article-title>Interrater reliability: the kappa statistic</article-title>. <source>Biochem. Med.</source> <volume>22</volume>, <fpage>276</fpage>&#x2013;<lpage>282</lpage>. doi: <pub-id pub-id-type="doi">10.11613/BM.2012.031</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Nunnally</surname> <given-names>J.</given-names></name> <name><surname>Bernstein</surname> <given-names>I.</given-names></name></person-group> (<year>1994</year>). <source>Psychometric theory</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>McGraw-Hill</publisher-name>.</citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Page</surname> <given-names>M. J.</given-names></name> <name><surname>McKenzie</surname> <given-names>J. E.</given-names></name> <name><surname>Bossuyt</surname> <given-names>P. M.</given-names></name> <name><surname>Boutron</surname> <given-names>I.</given-names></name> <name><surname>Hoffmann</surname> <given-names>T. C.</given-names></name> <name><surname>Mulrow</surname> <given-names>C. D.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>The PRISMA 2020 statement: an updated guideline for reporting systematic reviews</article-title>. <source>BMJ</source> <volume>372</volume>:<fpage>n71</fpage>. doi: <pub-id pub-id-type="doi">10.1136/bmj.n71</pub-id>, PMID: <pub-id pub-id-type="pmid">33782057</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Reitsma</surname> <given-names>H.</given-names></name> <name><surname>Rutjes</surname> <given-names>A.</given-names></name> <name><surname>Whiting</surname> <given-names>P.</given-names></name> <name><surname>Vlassov</surname> <given-names>V.</given-names></name> <name><surname>Deeks</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). &#x201C;<article-title>9 assessing methodological quality</article-title>&#x201D; in <source>Cochrane handbook for systematic reviews of diagnostic test accuracy version 1.0.0</source>. eds. <person-group person-group-type="editor"><name><surname>Deeks</surname> <given-names>J. J.</given-names></name> <name><surname>Bossuyt</surname> <given-names>P. M.</given-names></name> <name><surname>Gatsonis</surname> <given-names>C.</given-names></name></person-group> (<publisher-name>London: The Cochrane Collaboration</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>27</lpage>.</citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>Barbero</surname> <given-names>M. I.</given-names></name> <name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>P&#x00E9;rez-Gil</surname> <given-names>J. A.</given-names></name> <name><surname>Holgado</surname> <given-names>F. P.</given-names></name> <name><surname>S&#x00E1;nchez-Mart&#x00ED;n</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Scaling methods applied to set priorities in training programs in organizations</article-title>. <source>Psicothema</source> <volume>21</volume>, <fpage>509</fpage>&#x2013;<lpage>514</lpage>. PMID: <pub-id pub-id-type="pmid">19861090</pub-id> <comment>Available at: </comment><ext-link xlink:href="http://www.psicothema.com/pdf/3662.pdf" ext-link-type="uri">http://www.psicothema.com/pdf/3662.pdf</ext-link></citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>Holgado</surname> <given-names>F. P.</given-names></name> <name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Barbero</surname> <given-names>M. I.</given-names></name></person-group> (<year>2013</year>). <article-title>Measurement invariance study in the training satisfaction questionnaire (TSQ)</article-title>. <source>Span. J. Psychol.</source> <volume>16</volume>, <fpage>E28</fpage>&#x2013;<lpage>E12</lpage>. doi: <pub-id pub-id-type="doi">10.1017/sjp.2013.49</pub-id>, PMID: <pub-id pub-id-type="pmid">23866222</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanduvete-Chaves</surname> <given-names>S.</given-names></name> <name><surname>Lozano-Lozano</surname> <given-names>J. A.</given-names></name> <name><surname>Chac&#x00F3;n-Moscoso</surname> <given-names>S.</given-names></name> <name><surname>Holgado-Tello</surname> <given-names>F. P.</given-names></name></person-group> (<year>2018</year>). <article-title>Development of a work climate scale in emergency health services</article-title>. <source>Front. Psychol.</source> <volume>9</volume>:<fpage>10</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2018.00010</pub-id>, PMID: <pub-id pub-id-type="pmid">29403417</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Shadish</surname> <given-names>W. R.</given-names></name> <name><surname>Cook</surname> <given-names>T. D.</given-names></name> <name><surname>Campbell</surname> <given-names>D. T.</given-names></name></person-group> (<year>2002</year>). <source>Experimental and quasi-experimental designs for generalized causal inference</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Houghton Mifflin</publisher-name>.</citation></ref>
<ref id="ref30"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Shadish</surname> <given-names>W. R.</given-names></name> <name><surname>Cook</surname> <given-names>T. D.</given-names></name> <name><surname>Leviton</surname> <given-names>L. C.</given-names></name></person-group> (<year>1991</year>). <source>Foundations of program evaluation</source>. <publisher-loc>Newbury Park</publisher-loc>: <publisher-name>Sage</publisher-name>.</citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sterne</surname> <given-names>J. A. C.</given-names></name> <name><surname>Hern&#x00E1;n</surname> <given-names>M. A.</given-names></name> <name><surname>Reeves</surname> <given-names>B. C.</given-names></name> <name><surname>Savovi&#x0107;</surname> <given-names>J.</given-names></name> <name><surname>Berkman</surname> <given-names>N. D.</given-names></name> <name><surname>Viswanathan</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions</article-title>. <source>BMJ</source> <volume>355</volume>:<fpage>i4919</fpage>. doi: <pub-id pub-id-type="doi">10.1136/bmj.i4919</pub-id>, PMID: <pub-id pub-id-type="pmid">27733354</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sterne</surname> <given-names>J. A. C.</given-names></name> <name><surname>Savovi&#x0107;</surname> <given-names>J.</given-names></name> <name><surname>Page</surname> <given-names>M. J.</given-names></name> <name><surname>Elbers</surname> <given-names>R. G.</given-names></name> <name><surname>Blencowe</surname> <given-names>N. S.</given-names></name> <name><surname>Boutron</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>RoB 2: a revised tool for assessing risk of bias in randomised trials</article-title>. <source>BMJ</source> <volume>366</volume>:<fpage>l4898</fpage>. doi: <pub-id pub-id-type="doi">10.1136/bmj.l4898</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Timmerman</surname> <given-names>M. E.</given-names></name> <name><surname>Lorenzo-Seva</surname> <given-names>U.</given-names></name></person-group> (<year>2011</year>). <article-title>Dimensionality assessment of ordered polytomous items with parallel analysis</article-title>. <source>Psychol. Methods</source> <volume>16</volume>, <fpage>209</fpage>&#x2013;<lpage>220</lpage>. doi: <pub-id pub-id-type="doi">10.1037/a0023353</pub-id>, PMID: <pub-id pub-id-type="pmid">21500916</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Valentine</surname> <given-names>J. C.</given-names></name> <name><surname>Cooper</surname> <given-names>H.</given-names></name></person-group> (<year>2008</year>). <article-title>A systematic and transparent approach for assessing the methodological quality of intervention effectiveness research: the study design and implementation assessment device (study DIAD)</article-title>. <source>Psychol. Methods</source> <volume>13</volume>, <fpage>130</fpage>&#x2013;<lpage>149</lpage>. doi: <pub-id pub-id-type="doi">10.1037/1082-989X.13.2.130</pub-id>, PMID: <pub-id pub-id-type="pmid">18557682</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Xia</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>On the number of factors to retain in exploratory factor analysis for ordered categorical data</article-title>. <source>Behav. Res. Methods</source> <volume>47</volume>, <fpage>756</fpage>&#x2013;<lpage>772</lpage>. doi: <pub-id pub-id-type="doi">10.3758/s13428-014-0499-2</pub-id>, PMID: <pub-id pub-id-type="pmid">24947054</pub-id></citation></ref>
</ref-list>
</back>
</article>