<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Educ.</journal-id>
<journal-title>Frontiers in Education</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Educ.</abbrev-journal-title>
<issn pub-type="epub">2504-284X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/feduc.2020.554806</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Education</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Practical Rubrics for Informal Science Education Studies: (1) a STEM Research Design Rubric for Assessing Study Design and a (2) STEM Impact Rubric for Measuring Evidence of Impact</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Habig</surname> <given-names>Bobby</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/960678/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>American Museum of Natural History</institution>, <addr-line>New York, NY</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Biology, Queens College, City University of New York</institution>, <addr-line>Flushing, NY</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Ida Ah Chee Mok, The University of Hong Kong, Hong Kong</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Kathryn Holmes, Western Sydney University, Australia; Veronica Catete, North Carolina State University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Bobby Habig <email>bhabig&#x00040;amnh.org</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to STEM Education, a section of the journal Frontiers in Education</p></fn></author-notes>
<pub-date pub-type="epub">
<day>09</day>
<month>12</month>
<year>2020</year>
</pub-date>
<pub-date pub-type="collection">
<year>2020</year>
</pub-date>
<volume>5</volume>
<elocation-id>554806</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>04</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>11</month>
<year>2020</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2020 Habig.</copyright-statement>
<copyright-year>2020</copyright-year>
<copyright-holder>Habig</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Informal learning institutions, such as museums, science centers, and community-based organizations, play a critical role in providing opportunities for students to engage in science, technology, engineering, and mathematics (STEM) activities during out-of-school time hours. In recent years, thousands of studies, evaluations, and conference proceedings have been published measuring the impact that these programs have had on their participants. However, because studies of informal science education (ISE) programs vary considerably in how they are designed and in the quality of their designs, it is often quite difficult to assess their impact on participants. Knowing whether the outcomes reported by these studies are supported with sufficient evidence is important not only for maximizing participant impact, but also because there are considerable economic and human resources invested to support informal learning initiatives. To address this problem, I used the theories of impact analysis and triangulation as a framework for developing user-friendly rubrics for assessing quality of research designs and evidence of impact. I used two main sources, research-based recommendations from STEM governing bodies and feedback from a focus group, to identify criteria indicative of high-quality STEM research and study design. Accordingly, I developed three STEM Research Design Rubrics, one for quantitative studies, one for qualitative studies, and another for mixed methods studies, that can be used by ISE researchers, practitioners, and evaluators to assess research design quality. Likewise, I developed three STEM Impact Rubrics, one for quantitative studies, one for qualitative studies, and another for mixed methods studies, that can be used by ISE researchers, practitioners, and evaluators to assess evidence of outcomes. The rubrics developed in this study are practical tools that can be used by ISE researchers, practitioners, and evaluators to improve the field of informal science learning by increasing the quality of study design and for discerning whether studies or program evaluations are providing sufficient evidence of impact.</p></abstract>
<kwd-group>
<kwd>informal science education</kwd>
<kwd>museum education</kwd>
<kwd>research design</kwd>
<kwd>STEM</kwd>
<kwd>rubric design</kwd>
<kwd>evidence-based outcomes</kwd>
<kwd>out-of-school time</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Science Foundation<named-content content-type="fundref-id">10.13039/100000001</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="7"/>
<equation-count count="0"/>
<ref-count count="59"/>
<page-count count="21"/>
<word-count count="14035"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Informal science education (ISE) programs can be important vehicles for facilitating interest in science, technology, engineering, and mathematics (STEM) (National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>; Young et al., <xref ref-type="bibr" rid="B59">2017</xref>; Habig et al., <xref ref-type="bibr" rid="B24">2018</xref>). Indeed, in the last few decades, multiple studies and evaluations have reported evidence that involvement in informal, out-of-school time (OST) STEM programs is linked to participants&#x00027; awareness, interest, and engagement in STEM majors and careers (e.g., Fadigan and Hammrich, <xref ref-type="bibr" rid="B16">2004</xref>; Schumacher et al., <xref ref-type="bibr" rid="B48">2009</xref>; Winkleby et al., <xref ref-type="bibr" rid="B57">2009</xref>; McCreedy and Dierking, <xref ref-type="bibr" rid="B36">2013</xref>). Because many of these studies and evaluations vary considerably in how they are designed and in the quality of their designs, it is often difficult to gauge whether the outcomes reported are supported with sufficient evidence (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>). This is particularly a dilemma for studies of ISE programs because participation is voluntary and thus more fluid than formal programs, and there is also considerable variation in the number of contact hours between programs and among participants (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>). Therefore, a continuous challenge faced by the ISE community is how to gauge whether and to what extent the outcomes reported by studies and evaluations are supported by evidence. To address this problem, the goal of this study was to develop user-friendly rubrics that can be used to assess research designs and STEM outcomes of ISE studies. These rubrics, in turn, can be used to discern whether individual research studies or program evaluations are providing sufficient evidence supporting claims such as increased awareness, interest, and engagement in STEM majors and careers.</p>
<p>Over the past decade, there have been multiple initiatives carried out by several national agencies, including the National Research Council, the United States Department of Education, and the National Science Foundation, with the goal of identifying characteristics of high-impact STEM programs (e.g., U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; What Works Clearinghouse, <xref ref-type="bibr" rid="B54">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B40">2011</xref>, <xref ref-type="bibr" rid="B41">2013</xref>). Many of these agencies have established criteria for assessing a range of STEM outcomes including the mastery of twenty first Century Skills (National Research Council, <xref ref-type="bibr" rid="B40">2011</xref>), the implementation of Next Generation Science Standards (National Research Council, <xref ref-type="bibr" rid="B42">2014</xref>), and the impact of teacher professional development programs on student achievement (Yoon et al., <xref ref-type="bibr" rid="B58">2007</xref>). Additionally, the Committee on Highly Successful Schools or Programs for K-12 STEM Education established criteria for identifying the effectiveness of STEM-focused schools and associated student outcomes (National Research Council, <xref ref-type="bibr" rid="B40">2011</xref>). In recent years, various agencies have turned their attention to ISE programs. For example, the U.S. Department of Education (<xref ref-type="bibr" rid="B53">2007</xref>), in a report by the Academic Competitive Council, advanced as a national goal to improve awareness, interest, and engagement of STEM careers in the context of informal education. Additionally, the Institute for Learning Innovation (<xref ref-type="bibr" rid="B28">2007</xref>) was charged with assessing the quality and strength of evidence of STEM outcomes of ISE programs and of its participants. Further work has also been completed by several national agencies outlining criteria that can be used to identify effective OST STEM projects (Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B39">2010</xref>, <xref ref-type="bibr" rid="B42">2014</xref>, <xref ref-type="bibr" rid="B43">2015</xref>; Krishnamurthi et al., <xref ref-type="bibr" rid="B32">2014</xref>). Overall, there have been many strides made by science education stakeholders in terms of identifying characteristics of high-quality studies and for recommending criteria for assessing program outcomes. However, for studies of informal science projects, there remains a need for developing accessible methodologies for assessing the effectiveness of STEM interventions and the strength of the evidence of ISE research outcomes.</p>
<p>Due to the unique characteristics of informal learning environments, it is quite challenging to develop evidence-based criteria to assess whether a research study or evaluation has achieved specific outcomes (National Research Council, <xref ref-type="bibr" rid="B39">2010</xref>). Informal science education, defined as <italic>voluntary</italic> participation in science during out-of-school time hours, typically occurs after school, on weekends, and during the summer in a variety of settings including but not limited to museums, zoos, universities, and, non-profit organizations (Blanchard et al., <xref ref-type="bibr" rid="B3">2020</xref>). By design, ISE programs are voluntary, inquiry-based, and emphasize choice learning (National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>). Thus, for many programs, the random assignment of participants to treatment and control groups, often considered the gold standard in research design (What Works Clearinghouse, <xref ref-type="bibr" rid="B54">2008</xref>), is often logistically infeasible, potentially upsetting to learners, and may jeopardize the validity of certain studies (National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). An additional challenge is that the durations of many OST experiences are short-term making it difficult to measure program impact especially if the evidence of effects occur downstream of the experience (National Research Council, <xref ref-type="bibr" rid="B43">2015</xref>). Furthermore, because many ISE OST programs are designed with the specific intent of differentiating from formal school programs, program leaders often avoid the administration of written assessments (National Research Council, <xref ref-type="bibr" rid="B43">2015</xref>). Lastly, the Academic Competitive Council (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>) recognizes that informal learning experiences are highly individualized, complex, and multifaceted, and suggest that due to the modest scale of many of these programs, they may not warrant a costly assessment approach. Therefore, the application of criteria typically used to assess the effectiveness of outcomes of formal science programs and participants (e.g., What Works Clearinghouse, <xref ref-type="bibr" rid="B54">2008</xref>) might not be feasible for assessing the effectiveness of research studies or program evaluations designed to measure outcomes of participants of ISE programs. Promisingly, in recent years, many ISE stakeholders have developed rigorous research designs that are alternatives to random control trials including the employment of mixed methods and triangulation designs (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; Flick, <xref ref-type="bibr" rid="B17">2018a</xref>,<xref ref-type="bibr" rid="B18">b</xref>). Nonetheless, the unique nature of ISE programs must be accounted for when developing methodologies for assessing the effectiveness of STEM interventions in an informal setting.</p>
<p>Despite the inherent challenges of assessing the effectiveness of research studies and program evaluations of ISE participants, knowing whether the outcomes reported by these studies are supported with sufficient evidence is important for several reasons. First, there is an economic justification for gauging what works and what doesn&#x00027;t work because many ISE institutions, non-profit organizations, and governmental agencies invest considerable monetary and human resources to support informal STEM education initiatives (Wilkerson and Haden, <xref ref-type="bibr" rid="B55">2014</xref>). Therefore, to convince funding agencies, policy makers, and the public at-large that investing in ISE OST programs is important, the ISE community needs to show that these programs are helping young people to persist in STEM (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Wilkerson and Haden, <xref ref-type="bibr" rid="B55">2014</xref>). Second, the use of user-friendly tools can serve as a guide to inspire researchers to design rigorous studies that maximize evidence-based outcomes (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; Panadero and Jonsson, <xref ref-type="bibr" rid="B44">2013</xref>). Consequently, program leaders can more confidently use the information from these research studies to revise, redesign, and continuously improve their programs. Lastly, as more programs provide evidence of high impact, the ISE community can extract program design principles from highly effective programs and where appropriate, ISE OST program leaders can adopt and adapt these principles across institutions (Klein et al., <xref ref-type="bibr" rid="B31">2017</xref>).</p>
<p>One area of interest by STEM stakeholders related to assessment is how and to what extent ISE programs augment participants&#x00027; STEM major and STEM career outcomes (e.g., Cuddeback et al., <xref ref-type="bibr" rid="B13">2019</xref>; Chan et al., <xref ref-type="bibr" rid="B6">2020</xref>). Indeed, over the past decade, a major goal set forth by the National Research Council and the United States Department of Education is to inspire and motivate students to consider a STEM pathway (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B40">2011</xref>, <xref ref-type="bibr" rid="B41">2013</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). In the context of informal education and outreach, the Academic Competitiveness Council (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>) identified increased awareness, interest, and engagement in STEM majors and careers as priorities; each outcome is defined below:</p>
<list list-type="bullet">
<list-item><p><italic>STEM major awareness</italic>: increased knowledge and awareness of the various STEM disciplines available as fields of study at institutions of higher education</p></list-item>
<list-item><p><italic>STEM major interest:</italic> increased curiosity, motivation, and attention toward a STEM discipline as a focus of study at an institution of higher education</p></list-item>
<list-item><p><italic>STEM major engagement</italic>: a formal commitment to a STEM discipline as a focus of study in an institution of higher education</p></list-item>
<list-item><p><italic>STEM career awareness</italic>: increased knowledge and understanding of various STEM professions</p></list-item>
<list-item><p><italic>STEM career interest</italic>: increased curiosity, motivation, and attention toward STEM professions</p></list-item>
<list-item><p><italic>STEM career engagement</italic>: employment in a STEM profession.</p></list-item>
</list>
<p>In response to this challenge, many ISE programs have provided outreach and programming specifically designed to augment students&#x00027; awareness, interest, and engagement in a STEM pathway. Moreover, hundreds of studies and evaluations have been carried out to assess participants&#x00027; outcomes. Unfortunately, a tool to assess research design and to test whether these studies are supported with sufficient evidence is lacking, hence the focus of this study.</p>
<p>Our understanding of ISE OST programs is derived from two forms of published knowledge&#x02014;studies that have been published in peer-reviewed journals and studies that are the result of internal and external program evaluation (National Research Council, <xref ref-type="bibr" rid="B43">2015</xref>). Peer review is considered the cornerstone of academic research because research methods and findings are subject to critical examination by experts within a discipline. Peer-reviewed studies are essential for answering scholarly questions and for communicating meaningful research (Gannon, <xref ref-type="bibr" rid="B21">2001</xref>). As an alternative to peer review, internal and external program evaluations are also valuable for documenting program outcomes and for informing stakeholders on how to improve program design. Of these two forms of evaluation, external evaluations are often preferred by stakeholders because during the evaluation process, a program of interest is subject to independent analysis by an objective third party; however, external evaluations are typically more expensive than internal evaluations (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B43">2015</xref>). Of the many types of program evaluations, the two most common conducted in ISE research are formative and summative. Formative evaluations typically occur during the developmental stage of a program and are particularly informative for improving program design. Summative evaluations are conducted after the completion of a program and are useful for assessing whether the program outcomes align with the project goals and objectives (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>). The choice to conduct peer-reviewed research or an internal or external evaluation depends on many factors including available financial resources, the nature and duration of the program, and the goals of the stakeholders. Regardless of which form of published knowledge is selected, it is critical that research studies and program evaluations are rigorously designed to ensure the validity of research outcomes.</p>
<p>The aim of this study was to develop user-friendly rubrics to assess research design and to gauge whether a research study or program evaluation provided sufficient evidence to support specific claims (e.g., increased awareness, interest, and engagement in STEM majors and careers). To accomplish this goal, first I reviewed what experts consider to be evidence of high-quality research design and evidence of impact. Based on these research-based recommendations, I created a STEM Research Design Rubric and a STEM Impact Rubric tailored specifically for quantitative, qualitative, and mixed methods studies and evaluations of ISE OST programs. Second, I tested these rubrics for user-friendliness, reliability, and validity. Based on feedback from STEM researchers and practitioners, I made revisions to the rubrics when appropriate. Lastly, I assessed specific ISE OST studies and evaluations using these rubrics and provided case studies that illustrate how these tools can be used to evaluate research design and evidence of impact. Through this process, I developed practical tools that can be used by ISE researchers, STEM practitioners, and other stakeholder to assess the effectiveness of STEM interventions and evidence of research outcomes for both research studies and evaluations.</p>
</sec>
<sec id="s2">
<title>Theoretical Framework</title>
<p>The <italic>theory of impact analysis</italic> is described as a &#x0201C;rigorous and parsimonious&#x0201D; framework for mapping the assessment of impacts (Mohr, <xref ref-type="bibr" rid="B37">1995</xref> p. 55). The theory, which stems from the seminal work of Campbell and Stanley (<xref ref-type="bibr" rid="B4">1963</xref>) and later refined by Mohr (<xref ref-type="bibr" rid="B37">1995</xref>), considers three basic experimental designs: (1) experimental; (2) quasi-experimental; and (3) retrospective. In a true <italic>experimental design</italic>, the researcher sets up one or more subjects to receive the treatment (participation in the program) and another group in which one or more subjects do not receive the treatment (the control). As a benchmark, the control and treatment groups are assigned randomly, and an adequate number of subjects are assigned to participate in the study. As a best practice, the theory of impact analysis recommends when possible, the use of larger treatment and control groups in order to help increase the sensitivity of a study. For example, the Institute for Learning Innovation (<xref ref-type="bibr" rid="B28">2007</xref>), while assessing the impact of different program evaluations of various ISE programs, defined an adequate sample size as a minimum of 50 subjects. Two examples of a true experimental design are the pre-test, post-test design with random assignment and the post-test only design with random assignment (<xref ref-type="table" rid="T1">Table 1</xref>). The second design considered in the theory of impact analysis is the <italic>quasi-experimental design</italic>. In the quasi-experimental design, researchers set up intervention and comparison groups, but the groups are <italic>not</italic> assigned randomly. If the study is designed so that comparison groups are closely matched in key characteristics, evidence suggests that a quasi-experimental study can yield strong evidence of the intervention&#x00027;s impact. Examples of quasi-experimental studies include post-test only designs, comparative change, and comparative time series (<xref ref-type="table" rid="T1">Table 1</xref>). Lastly, in a <italic>retrospective design</italic>, also known as <italic>ex post facto</italic> (&#x0201C;after the fact&#x0201D;), the subjects that received treatment were not assigned by the experimenter. For example, if some youth register for an ISE program and others do not, the selection of participants was not assigned by the researcher; rather, the participants underwent self-selection.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Study designs for assessing the effectiveness of a STEM intervention.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Quantitative study design</bold></th>
<th valign="top" align="left"><bold>Examples</bold></th>
<th valign="top" align="left"><bold>Diagram</bold></th>
<th valign="top" align="left"><bold>Advantages/Disadvantages</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td/>
<td valign="top" align="left">X = STEM intervention<break/> C = comparison group<break/> T = treatment group<break/> R = random assignment<break/> O = outcome measure/evidence<break/> NA = not applicable</td>
<td/>
</tr> <tr>
<td valign="top" align="left">Experimental design (randomized controlled trials)</td>
<td valign="top" align="left">1. Pre-test, post-test design with random assignment <break/>2. Post-test only with random assignment <break/>3. Solomon four group design (Solomon, <xref ref-type="bibr" rid="B51">1949</xref>)</td>
<td valign="top" align="left">1. T(R): OXO <break/>2. C(R): OO <break/>3. T(R): XO<break/> C(R): O<break/> T(R): OXO<break/> C(R): OO<break/> T(R): XO<break/> C(R): O</td>
<td valign="top" align="left">1. Reduces threats to internal validity/doesn&#x00027;t control for effect of pre-test <break/>2. Controls for pre-test effects/doesn&#x00027;t measure change over time <break/>3. Strongest quantitative design for reducing threats to validity</td>
</tr>
<tr>
<td valign="top" align="left">Quasi-experimental designs (well-matched comparison groups)</td>
<td valign="top" align="left">1. Pre-test, post-test design with comparison group <break/>2. Post-test only with comparison group <break/>3. Time series with comparison group</td>
<td valign="top" align="left">1.T: OXO <break/> C: OO 2.T: XO<break/> C: O3. Example:<break/> T: OXXOXXO<break/> C: OOO</td>
<td valign="top" align="left">1. Reduces threats to internal validity/doesn&#x00027;t control for effect of pre-test <break/>2. Assesses change over time/non-random assignment increases threats to validity <break/>3. Assesses longer term change/non-random assignment increases threats to validity</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Other quantitative designs</td>
<td valign="top" align="left">1. Pre-test, post-test design without comparison group <break/>2. Post-test only without comparison group <break/>3. Time series without comparison group</td>
<td valign="top" align="left">1. T: OXO <break/>2. T: XO <break/>3. Example:<break/> T: OXXOXXO</td>
<td valign="top" align="left">1. Assesses change over time/lack of control increases threats to validity <break/>2. Provides a snapshot/lack of control increases threats to validity <break/>3. Assesses longer term change/lack of control increases threats to validity</td>
</tr> <tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>Qualitative study design</bold></td>
<td valign="top" align="left"><bold>Examples</bold></td>
<td valign="top" align="left"><bold>Description</bold></td>
<td valign="top" align="left"><bold>Advantages/Disadvantages</bold></td>
</tr> <tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">Qualitative design</td>
<td valign="top" align="left">1. Narrative study <break/>2. Phenomenological study <break/>3. Grounded theory study <break/>4. Ethnographic study <break/>5. Case study</td>
<td valign="top" align="left">1. Researcher extracts themes from narratives of one or more individuals <break/>2. Researchers study several individuals with shared experiences to analyze a phenomenon of interest <break/>3. Researcher extracts data from interviews of &#x0007E;20&#x02013;60 individuals and uses systematic coding to develop a unified theoretical explanation <break/>4. Researcher extracts themes by describing and interpreting patterns of a shared culture of group <break/>5. Researchers conduct an in-depth analysis of one or multiple cases</td>
<td valign="top" align="left">1. Allows for in-depth exploration/resource intensive, potential observer bias <break/>2. Provides a deep understanding of phenomenon experienced by multiple individuals/resources intensive, potential observer bias <break/>3. Systematic approach to data analysis/exhaustive process, potential observer bias <break/>4. Development of a complex, exhaustive description of a culture of group/resource intensive, potential observer bias <break/>5. Greater depth of analysis/limited generalizability</td>
</tr>
<tr>
<td valign="top" align="left">Mixed methods design</td>
<td valign="top" align="left">Uses one or more quantitative and qualitative design</td>
<td valign="top" align="left">Researchers collect, analyze, and integrate quantitative and qualitative data</td>
<td valign="top" align="left">Allows for triangulation of data, which counteracts disadvantages of individual designs/may be difficult to interpret if there are conflicting outcomes</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In alignment with the theory of impact analysis, the U.S. Department of Education (<xref ref-type="bibr" rid="B53">2007</xref>) proposed a &#x0201C;Hierarchy of Study Designs for Evaluating the Effectiveness of a STEM Educational Intervention&#x0201D; consisting of three hierarchical levels&#x02014;(1) experimental designs (randomized controlled trials), (2) quasi-experimental designs (well-matched comparison groups), and (3) other designs (e.g., pre/post studies, comparison groups without careful matching). In this hierarchy, a well-designed randomized controlled trial is the preferred method; quasi-experimental designs are preferred when experimental designs are not feasible, and other designs are considered when the first two designs are not feasible. Thus, according to the theory of impact analysis, the best designed studies are those that exhibit internal validity (i.e., studies that can make a causal link between treatment and outcomes) and those that exhibit external validity (studies that are generalizable to other programs and populations). Threats to internal validity include non-equivalent comparison groups, non-random assignment of subjects, and confounded experimental treatments (Fuchs and Fuchs, <xref ref-type="bibr" rid="B20">1986</xref>). Threats to external validity include the timing of the study, the setting in which it occurred, the study subjects, and treatment conditions (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>). The theory of impact analysis thus describes a framework for reducing threats to internal and external validity and for developing experimental designs that measure the extent to which a study provides evidence of impact.</p>
<p>One limitation of the theory of impact analysis is that it might underestimate the power of qualitative studies. Qualitative studies are important because they emphasize depth of understanding and provide rich information on how participants have interpreted their STEM experience (Diamond et al., <xref ref-type="bibr" rid="B15">2016</xref>). Because qualitative research relies on open-ended questions and in-depth responses, researchers often use this approach to identify patterns and to develop emerging themes (Jackson et al., <xref ref-type="bibr" rid="B29">2007</xref>). According to Creswell and Poth (<xref ref-type="bibr" rid="B12">2018</xref>), high quality qualitative studies share common characteristics including the employment of rigorous methodological, data collection, and data analysis protocols, and the incorporation of one of the five following approaches to qualitative inquiry: (1) a narrative study, (2) a phenomenological study, (3) a grounded theory study, (4) an ethnographic study, and (5) a case study (<xref ref-type="table" rid="T1">Table 1</xref>). When qualitative studies are rigorously designed, they can be effective means for probing how ISE OST experiences impact participants&#x00027; STEM pathways (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B43">2015</xref>). One way to include qualitative research in an impact analysis is to allow for the combination of quantitative and qualitative data, which is the basis of the theory of triangulation (Greene and McClintock, <xref ref-type="bibr" rid="B23">1985</xref>; Creswell and Poth, <xref ref-type="bibr" rid="B12">2018</xref>).</p>
<p>The <italic>theory of triangulation</italic> is based on the idea that the collection of multiple sources of quantitative and qualitative data helps to increase study validity allowing researchers to gain a more complete picture of participants&#x00027; outcomes (Denzin, <xref ref-type="bibr" rid="B14">1970</xref>; Ammenwerth et al., <xref ref-type="bibr" rid="B1">2003</xref>; Flick, <xref ref-type="bibr" rid="B17">2018a</xref>,<xref ref-type="bibr" rid="B18">b</xref>). The concept of triangulation can be traced to Campbell and Fiske (<xref ref-type="bibr" rid="B5">1959</xref>), who developed the idea of &#x0201C;multiple operationalism,&#x0201D; which argues that the use of multiple methods ensures that the variance reflects the study&#x00027;s outcomes and not its methodology. Triangulation has also been described in the literature as convergent methodology, convergent validation, and mixed methods (Hussein, <xref ref-type="bibr" rid="B27">2009</xref>). The triangulation metaphor stems from military and navigational strategy where multiple reference points are used to identify an object&#x00027;s exact position (Smith, <xref ref-type="bibr" rid="B50">1975</xref>). By applying basic principles of geometry, multiple perspectives allow for greater accuracy. Similarly, researchers can improve the accuracy of their interpretations by collecting multiple sources of data assessing the same phenomenon (Jick, <xref ref-type="bibr" rid="B30">1979</xref>). Together, the <italic>theory of impact analysis</italic> and the <italic>theory of triangulation</italic> are helpful lenses for informing the development of a STEM Research Design Rubric and a STEM Impact Rubric.</p>
</sec>
<sec sec-type="methods" id="s3">
<title>Methods</title>
<p>The primary objectives of this study were to develop a user-friendly STEM Research Design Rubric and a STEM Impact Rubric for ISE researchers, evaluators, and practitioners. To do so, first I identified national agencies and STEM governing bodies comprised of experts specifically charged with identifying criteria indicative of high-quality STEM research and study design. This was accomplished by searching documents in the Center for the Advancement of Informal Science Education repository (<ext-link ext-link-type="uri" xlink:href="https://informalscience.org">informalscience.org</ext-link>) and by contacting three distinguished ISE researchers (two university-based ISE researchers and one research director at a large ISE institution) for recommended sources of information. Through this process, seven agencies were identified as potential sources of information: (1) National Research Council; (2) United States Department of Education; (3) National Science Foundation; (4) Institute of Learning Innovation; (5) Afterschool Alliance; (6) The PEAR Institute: Partnerships in Education and Resilience; and (7) What Works Clearinghouse. Second, I used Google Scholar and <ext-link ext-link-type="uri" xlink:href="https://informalscience.org">informalscience.org</ext-link> to identify and extract publications from these agencies with a specific focus on documents that provide research-based recommendations on how to assess research design and evidence of impact from studies and evaluations of STEM participants. Accordingly, the criteria that I identified to develop the STEM Research Design Rubric and STEM Impact Rubric were based on a synthesis of recommendations from published reports from multiple governing bodies. Importantly, these criteria stemmed from the counsel of leading ISE researchers, statisticians, and policymakers based on our current knowledge of best practices for research design and assessment. Thus, to ensure content validity, I adopted research-based recommendations from ISE stakeholders because these experts specifically considered the unique characteristics of ISE programs when making recommendations (e.g., Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>), and I used the theories of impact analysis and triangulation to inform this process. Lastly, based on these recommendations, I developed rubrics that researchers, STEM practitioners, and other stakeholders can use to assess research design and whether a research study or program evaluation has provided sufficient evidence to support a claim made about a particular STEM outcome.</p>
<p>During the design of these rubrics, I constructed a rating scale for the STEM Research Design Rubric and another for the STEM Impact Rubric. For the STEM Research Design Rubric, the rating scale was divided into four different levels: (1) a rating of one was indicative of a study or evaluation in which there was a <italic>weak research design</italic>; (2) a rating of two was indicative of a study or evaluation in which there was an <italic>adequate research design</italic>; (3) a rating of three was indicative of a study or evaluation in which there was a <italic>strong research design</italic>; and (4) a rating of four was indicative of a study or evaluation in which there was an <italic>exemplary research design</italic>. For the STEM Impact Rubric, the rating scale was also divided into four levels: (1) a rating of one was indicative of a study or evaluation in which there was <italic>little or no evidence</italic> of impact; (2) a rating of two was indicative of a study or evaluation in which there was <italic>moderate evidence</italic> of impact; (3) a rating of three was indicative of a study or evaluation in which there was <italic>strong evidence</italic> of impact; and (4) a rating of four was indicative of a study or evaluation in which there was <italic>exemplary evidence</italic> of impact. A four-point rubric was applied because this rating scale is considered the gold standard in rubric design (Phillip, <xref ref-type="bibr" rid="B45">2002</xref>) and is commonly applied by STEM stakeholders in a myriad of contexts (e.g., What Works Clearinghouse, <xref ref-type="bibr" rid="B54">2008</xref>; Singer et al., <xref ref-type="bibr" rid="B49">2012</xref>; The PEAR Institute: Partnerships in Education and Resilience, <xref ref-type="bibr" rid="B52">2017</xref>).</p>
<p>After designing prototypes of the STEM Impact Rubric, I recruited eight ISE stakeholders to review the STEM Impact Rubric and to participate in a focus group. All eight participants are ISE educators; five are active STEM researchers with a PhD; three hold a Masters in a STEM-related field. Of the eight focus group participants, one is a director of research at one of the largest informal science education institutions in the world; two are postdoctoral fellows both with a background in informal science education and mixed methods research. The remaining five participants are program managers at various ISE institutions; three hold Master&#x00027;s in museum science education, one of these three has a background in statistics and two in anthropology. The other two participants are directors of a high school science research mentoring program; both have PhDs in Biology. All eight focus group members have published ISE research, and in addition to their status as STEM stakeholders, are also intended users of the rubrics developed in this study.</p>
<p>The focus group members were provided with rubrics and asked to conduct an initial review and to test out the rubrics on a study or evaluation on the STEM outcomes of ISE participants. Each volunteer was given 1 week to review the document, and then invited to a focus group meeting to provide their feedback. During the focus group, I conducted a semi-structured group interview that included the following questions for discussion: (1) What are your impressions of the rubrics? (2) Do you think that this tool is user-friendly for ISE stakeholders (researchers, evaluators, practitioners)? Why or why not? (3) What changes or tweaks would you make so that these rubrics are more user-friendly? The focus group discussion was also guided by the theories of impact analysis and triangulation as the focus group participants used the hierarchy of study designs and mixed methods triangulation as a lens for assessing and practicing the rubrics (Denzin, <xref ref-type="bibr" rid="B14">1970</xref>; Mohr, <xref ref-type="bibr" rid="B37">1995</xref>; Ammenwerth et al., <xref ref-type="bibr" rid="B1">2003</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Flick, <xref ref-type="bibr" rid="B17">2018a</xref>,<xref ref-type="bibr" rid="B18">b</xref>). While the focus group members did not have overlap with the experts who helped inform the initial rubric design, their feedback was important for improving the categories and metrics used for each rubric. Based on the recommendations of the focus group, I revised the rubrics and next tested the tools for reliability and validity.</p>
<p>To test for reliability, I worked with a graduate research assistant and together we used the STEM Research Design Rubric to rate 25 ISE studies (i.e., a combination of peer-reviewed studies, program evaluations, and conference proceedings) independently. Additionally, we used the STEM Impact Rubric to rate 47 outcomes independently. To assess consistency across raters (reliability), I calculated inter-rater agreement using both percent agreement (Lombard et al., <xref ref-type="bibr" rid="B34">2002</xref>) and Cohen&#x00027;s kappa (Cohen, <xref ref-type="bibr" rid="B8">1968</xref>). Percent agreement for the STEM Research Design Rubric was 92.0% and for the STEM Impact Rubric 80.9%. Cohen&#x00027;s kappa (k) for the STEM Research Design Rubric was 0.89, and for the STEM Impact Rubric 0.74. These measures were indicative of substantial agreement (Cohen, <xref ref-type="bibr" rid="B8">1968</xref>). Lastly, to increase content validity of our rubric, we incorporated triangulation methods (Creswell and Miller, <xref ref-type="bibr" rid="B11">2000</xref>; Creswell and Poth, <xref ref-type="bibr" rid="B12">2018</xref>). Because the most valid assessments stem from &#x0201C;the collective judgment of recognized experts in that field&#x0201D; (Baer and McKool, <xref ref-type="bibr" rid="B2">2014</xref>, p. 82), during triangulation, I extracted and synthesized data from multiple sources including recommendations from national agencies and STEM governing bodies comprised of experts specifically charged with identifying criteria indicative of high-quality STEM research and from focus group participants consisting of experienced ISE STEM researchers and practitioners.</p>
</sec>
<sec sec-type="results" id="s4">
<title>Results</title>
<sec>
<title>Focus Group Results</title>
<p>Four key recommendations were made by the focus group. First, the focus group recommended that I design six distinct rubrics separated into two categories: (1) three <italic>research design</italic> rubrics, one for quantitative studies, one for qualitative studies, and one for mixed methods studies, and (2) three evidence of <italic>impact</italic> rubrics, one for quantitative studies, one for qualitative studies, and one for mixed methods studies. Initially, I provided the focus group with three rubrics that combined <italic>research design</italic> and evidence of <italic>impact</italic> together. However, the focus group unanimously agreed that <italic>research design</italic> and evidence of <italic>impact</italic> are distinct criteria warranting separate analysis. Thus, based on this recommendation, in the final iteration, I developed six distinct rubrics to assess quantitative, qualitative, and mixed methods studies: three STEM Research Design Rubrics and three STEM Impact Rubrics. Second, to make the rubrics more user friendly, the focus group recommended that instead of presenting the four levels of evidence vertically, I should present these criteria horizontally. This recommendation was adopted in the final draft. Third, because some researchers might have difficulties using each rubric as a stand-alone tool, the focus group suggested that I develop a worksheet with a key to help guide researchers, practitioners, and evaluators through the assessment process when using each rubric. This recommendation was also adopted (see <xref ref-type="fig" rid="F1">Figures 1</xref>&#x02013;<xref ref-type="fig" rid="F6">6</xref>). Lastly, there was some debate between focus group members on whether a research design <italic>not</italic> grounded in theory should automatically be rated as an <italic>adequate research design</italic> (a score of 2 out of 4). The focus group did not reach consensus with respect to this question and instead recommended that if I retain this rating in the final draft, then I should also emphasize in the Discussion that these rubrics should be viewed as heuristics for researchers, practitioners, and evaluators and that in some cases, depending on the goals of a particular study, it might be appropriate to alter criteria.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>STEM Research Design Rubric worksheet for quantitative studies.</p></caption>
<graphic xlink:href="feduc-05-554806-g0001.tif"/>
</fig>
</sec>
<sec>
<title>A STEM Research Design Rubric for Quantitative Studies</title>
<p>A STEM Research Design Rubric for quantitative studies (<xref ref-type="table" rid="T2">Table 2</xref>) was designed in accordance with research-based recommendations from ISE experts (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). In alignment with the theory of impact analysis (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>), I found that a quantitative study or evaluation with evidence of an <italic>exemplary research design</italic> is one that includes a control (comparison) and treatment group (program participants) selected randomly (experimental) (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B41">2013</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). Alternatively, because of the unique nature of ISE programs, I found that it might be more appropriate to reference a comparison group that is not a strict control (National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>). Therefore, I also found that a quantitative study or evaluation that provides evidence of an <italic>exemplary research design</italic> may alternatively include a well-matched comparison group (control) and treatment group (program participants) selected non-randomly (quasi-experimental) (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B41">2013</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). ISE experts also recommend that a quantitative study or evaluation with evidence of <italic>exemplary research design</italic> should be grounded in a theoretical framework (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>) and consist of a sample size of fifty or more subjects per group (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; Diamond et al., <xref ref-type="bibr" rid="B15">2016</xref>). Based on additional recommendations from ISE experts, I also developed criteria indicative of <italic>strong, adequate</italic>, and <italic>weak</italic> research design for quantitative studies and evaluations (<xref ref-type="table" rid="T2">Table 2</xref>). Lastly, as recommended by the focus group, I developed a STEM Research Design Worksheet for quantitative studies as a key to guide researchers, practitioners, or evaluators through the process of assessing research design quality (<xref ref-type="fig" rid="F1">Figure 1</xref>).</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>STEM research design rubric for quantitative studies.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center" colspan="4"><bold>STEM Research Design Rubric (Quantitative studies)</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>(4) Exemplary research design</bold></td>
<td valign="top" align="left"><bold>(3) Strong research design</bold></td>
<td valign="top" align="left"><bold>(2) Adequate research design</bold></td>
<td valign="top" align="left"><bold>(1) Weak research design</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group (e.g., pre-test/post-test design with comparison; post-test only design with comparison) <break/>&#x02022; study grounded in a theoretical framework <break/>&#x02022; minimum of 50 subjects per group</td>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group OR <italic>quantitative design without comparison group</italic> (e.g., pre-test/post-test, post-test only, and/or time series design(s) without comparison) <break/>&#x02022; study grounded in a theoretical framework <break/>&#x02022; minimum of 40 subjects</td>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group OR <italic>quantitative design without comparison group</italic> (e.g., pre-test/post-test, post-test only, and/or time series design(s) without comparison) <break/>&#x02022; study may or may not be grounded in a theoretical framework <break/>&#x02022; minimum of 25 subjects</td>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group OR <italic>quantitative design without comparison group</italic> (e.g., pre-test/post-test, post-test only, and/or time series design(s) without comparison) <break/>&#x02022; study may or may not be grounded in a theoretical framework <break/>&#x02022; &#x0003C;25 subjects or not reported</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4"><bold>Helpful shortcuts:</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="4">&#x02713; If there is no comparison group, the starting value is 3 <break/>&#x02713; If there is no theoretical framework, the starting value is 2 <break/>&#x02713; The sample size represents the number of subjects analyzed in the study, not the number of program participants</td>
</tr>
<tr>
<td/>
<td/>
<td/>
<td valign="top" align="left"><bold>Rubric Score:</bold> <italic><bold>__________</bold></italic></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Note: In order to receive a rubric score (4, 3, 2, or 1), a study must meet the criteria of all three items in a given column</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>A STEM Impact Rubric for Quantitative Studies</title>
<p>A STEM Impact Rubric for quantitative studies (<xref ref-type="table" rid="T3">Table 3</xref>) was designed in accordance with research-based recommendations from ISE experts (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). Based on these recommendations, a study outcome indicative of <italic>exemplary evidence</italic> of impact needs to report a statistically significant difference between comparison and treatment groups (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B41">2013</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). Therefore, a study that measured STEM career interest (or some other outcome of interest) and found that program participants exhibited significantly higher interest than a well-matched comparison group would be an example of an outcome indicative of <italic>exemplary evidence</italic> of impact. Based on additional recommendations from ISE experts, I also developed criteria indicative of <italic>strong, moderate</italic>, and <italic>little or no evidence</italic> of impact for quantitative studies and evaluations (<xref ref-type="table" rid="T3">Table 3</xref>). Lastly, as recommended by the focus group, I developed a STEM Impact Worksheet for quantitative studies as a key to guide researchers, practitioners, or evaluators through the process of assessing the impact of specific outcomes (<xref ref-type="fig" rid="F2">Figure 2</xref>).</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>STEM impact rubric for quantitative studies.</p></caption>
<table frame="hsides" rules="groups">
<tbody><tr>
<td valign="top" align="left" colspan="4"><bold>Title of study or evaluation:</bold><break/> <bold>STEM outcome under review:</bold><break/> List all method(s) used to assess the STEM outcome under review (e.g., survey question 1, survey question 2, etc.):<break/> 1. ____________________ 2. ____________________ 3. ____________________ 4. ____________________ 5. ____________________<break/> 6. ____________________ 7. ____________________ 8. ____________________ 9. ____________________ 10. ____________________<break/><break/> <bold>Directions</bold>:<break/> 1. Identify the STEM outcome under review (e.g., increased STEM career interest)<break/> 2. List all criteria used to assess the STEM outcome under review. For example, if there were four survey questions that evaluated STEM career interest, list all four questions (e.g., survey question 1, survey question 2, etc.)<break/> 3. Use the rubric to evaluate each criterion used to assess the STEM outcome under review starting from left (exemplary evidence) to right (little or no evidence)<break/> 4. Calculate the average rubric score by dividing the sum of all rubric scores by the number of criteria used to assess STEM outcomes</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="center" colspan="4" style="border-bottom: thin solid #000000;"><bold>STEM Impact Rubric (Quantitative studies)</bold></td>
</tr> <tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>(4) Exemplary evidence</bold></td>
<td valign="top" align="left"><bold>(3) Strong evidence</bold></td>
<td valign="top" align="left"><bold>(2) Moderate evidence</bold></td>
<td valign="top" align="left"><bold>(1) Little or No evidence</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02022; statistically significant difference between comparison (control) and treatment (program participants) groups <break/>&#x02022; If there is no comparison group, starting value is 3</td>
<td valign="top" align="left">&#x02022; statistically significant difference between pre- and post-survey OR minimum of 75% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program survey</td>
<td valign="top" align="left">&#x02022; 40&#x02013;75% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program surveys OR STEM outcome is maintained in comparisons of pre- and post-surveys (i.e., no significant difference between pre- and post-assessments)</td>
<td valign="top" align="left">&#x02022; less than 40% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program surveys OR STEM outcome significantly decreases in comparisons of pre- and post-surveys OR outcomes of participants are the same or significantly lower than outcomes of comparison group</td>
</tr>
<tr>
<td valign="top" align="center" colspan="4"><bold>Average Rubric Score (sum of all rubric scores</bold> <bold>&#x000F7;</bold> <bold>number of criteria used to assess STEM outcomes):</bold> <italic><bold>__________</bold></italic></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>STEM Impact Rubric worksheet for quantitative studies.</p></caption>
<graphic xlink:href="feduc-05-554806-g0002.tif"/>
</fig>
</sec>
<sec>
<title>A STEM Research Design Rubric for Qualitative Studies</title>
<p>A STEM Research Design Rubric for qualitative studies (<xref ref-type="table" rid="T4">Table 4</xref>) was designed in accordance with research-based recommendations from ISE experts (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). In alignment with the theory of impact analysis (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>), I found that a qualitative study or evaluation (e.g., narrative study; phenomenological study; grounded theory study; ethnographic study) that provides evidence of <italic>exemplary research design</italic> is one that includes a control (well-matched comparison) and treatment group (program participants) selected randomly (experimental) or non-randomly (quasi-experimental) (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; National Research Council, <xref ref-type="bibr" rid="B41">2013</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). ISE experts also recommend that a qualitative study or evaluation with evidence of <italic>exemplary research design</italic> should be grounded in a theoretical framework (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>) and consist of a sample size of twenty or more subjects per group (Creswell and Poth, <xref ref-type="bibr" rid="B12">2018</xref>). Based on additional recommendations from ISE experts, I also developed criteria indicative of <italic>strong, adequate</italic>, and <italic>weak</italic> research design for qualitative studies and evaluations (<xref ref-type="table" rid="T4">Table 4</xref>). Lastly, as recommended by the focus group, I developed a STEM Research Design Worksheet for qualitative studies as a key to guide researchers, practitioners, or evaluators through the process of assessing research design quality (<xref ref-type="fig" rid="F3">Figure 3</xref>).</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>STEM research design rubric for qualitative studies.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center" colspan="4"><bold>STEM Research Design Rubric (Qualitative studies)</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>(4) Exemplary research design</bold></td>
<td valign="top" align="left"><bold>(3) Strong research design</bold></td>
<td valign="top" align="left"><bold>(2) Adequate research design</bold></td>
<td valign="top" align="left"><bold>(1) Weak research design</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study <break/>&#x02022; study grounded in a theoretical framework <break/>&#x02022; minimum of 20 subjects per group</td>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study OR <italic>qualitative design without comparison group</italic> <break/>&#x02022; study grounded in a theoretical framework <break/>&#x02022; minimum of 15 subjects</td>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study OR <italic>qualitative design without comparison group</italic> <break/>&#x02022; study may or may not be grounded in a theoretical framework <break/>&#x02022; minimum of 10 subjects</td>
<td valign="top" align="left">&#x02022; <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study OR <italic>qualitative design without comparison group</italic> <break/>&#x02022; study may or may not be grounded in a theoretical framework <break/>&#x02022; &#x0003C;10 subjects or not reported</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4"><bold>Helpful shortcuts:</bold></td>
</tr>
<tr>
<td valign="top" align="left" colspan="4">&#x02713; If there is no comparison group, the starting value is 3 <break/>&#x02713; If there is no theoretical framework, the starting value is 2 <break/>&#x02713; The sample size represents the number of subjects analyzed in the study, not the number of program participants</td>
</tr>
<tr>
<td/>
<td/>
<td/>
<td valign="top" align="left"><bold>Rubric Score:</bold> <italic><bold>__________</bold></italic></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Note: In order to receive a rubric score (4, 3, 2, or 1), a study must meet the criteria of all three items in a given column</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>STEM Research Design Rubric worksheet for qualitative studies.</p></caption>
<graphic xlink:href="feduc-05-554806-g0003.tif"/>
</fig>
</sec>
<sec>
<title>A STEM Impact Rubric for Qualitative Studies</title>
<p>A STEM Impact Rubric for qualitative studies (<xref ref-type="table" rid="T5">Table 5</xref>) was designed in accordance with research-based recommendations from ISE experts (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). Based on these recommendations, a qualitative study indicative of <italic>exemplary evidence</italic> of impact may demonstrate this in one of two ways: (1) a study or evaluation in which researchers identify the outcome of interest in the treatment but not the control as an emerging theme or (2) a study or evaluation in which a minimum of 75% of data (e.g., interview data, response to open-ended question) is indicative of the outcome of interest in the treatment but not the control; this difference needs to be statistically significant. Based on additional recommendations from ISE experts, I also developed criteria indicative of <italic>strong, moderate</italic>, and <italic>little or no evidence</italic> of impact for qualitative studies and evaluations (<xref ref-type="table" rid="T5">Table 5</xref>). Lastly, as recommended by the focus group, I developed a STEM Impact Worksheet for qualitative studies as a key to guide researchers, practitioners, or evaluators through the process of assessing the impact of specific outcomes (<xref ref-type="fig" rid="F4">Figure 4</xref>).</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>STEM impact rubric for qualitative studies.</p></caption>
<table frame="hsides" rules="groups">
<tbody><tr>
<td valign="top" align="left" colspan="4"><bold>Title of study or evaluation</bold>:<break/> <bold>STEM outcome under review</bold>:<break/> List all method(s) used to assess the STEM outcome under review (e.g., survey question 1, survey question 2, etc.):<break/> 1. ____________________ 2. ____________________ 3. ____________________ 4. ____________________ 5. ____________________<break/> 6. ____________________ 7. ____________________ 8. ____________________ 9. ____________________ 10. ____________________<break/><break/> <bold>Directions</bold>:<break/> 1. Identify the STEM outcome under review (e.g., increased STEM career interest)<break/> 2. List all criteria used to assess the STEM outcome under review. For example, if there were four survey questions that evaluated STEM career interest, list all four questions (e.g., open-ended question 1, open-ended question 2, etc.)<break/> 3. Use the rubric to evaluate each criterion used to assess the STEM outcome under review starting from left (exemplary evidence) to right (little or no evidence)<break/> 4. Calculate the average rubric score by dividing the sum of all rubric scores by the number of criteria used to assess STEM outcomes</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="center" colspan="4" style="border-bottom: thin solid #000000;"><bold>STEM Impact Rubric (Qualitative studies)</bold></td>
</tr> <tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>(4) Exemplary evidence</bold></td>
<td valign="top" align="left"><bold>(3) Strong evidence</bold></td>
<td valign="top" align="left"><bold>(2) Moderate evidence</bold></td>
<td valign="top" align="left"><bold>(1) Little or No evidence</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02022; researchers identify outcome of interest as an emerging theme during qualitative analyses for treatment but not comparison OR minimum of 75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted an outcome of interest in the treatment but not control, and there was a statistically significant difference between these two groups <break/>&#x02022; If there is no comparison group, starting value is 3</td>
<td valign="top" align="left">&#x02022; researchers identify outcome of interest as an emerging theme during qualitative analyses OR minimum of 75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted STEM outcome of interest</td>
<td valign="top" align="left">&#x02022; 40&#x02013;75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted STEM outcome of interest</td>
<td valign="top" align="left">&#x02022; less than 40% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted STEM outcome of interest OR anecdotal evidence of STEM outcomes (e.g., handful of participants quotes, but no systematic analysis) OR no difference in emerging themes between treatment and comparison groups</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4"><bold>Average Rubric Score (sum of all rubric scores</bold> <bold>&#x000F7;</bold> <bold>number of criteria used to assess STEM outcomes):</bold> <italic><bold>__________</bold></italic></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>STEM Impact Rubric worksheet for qualitative studies.</p></caption>
<graphic xlink:href="feduc-05-554806-g0004.tif"/>
</fig>
</sec>
<sec>
<title>A STEM Research Design Rubric for Mixed Methods Studies</title>
<p>A STEM Research Design Rubric for mixed methods studies (<xref ref-type="table" rid="T6">Table 6</xref>) was designed in accordance with research-based recommendations from ISE experts (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). In alignment with the theories of impact analysis (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>) and triangulation (Campbell and Fiske, <xref ref-type="bibr" rid="B5">1959</xref>), I found that a mixed methods study (i.e., a study or evaluation that incorporates both quantitative and qualitative analyses) provides exemplary evidence of quality research design if the study or evaluation meets the benchmarks for exemplary evidence described for quantitative (<xref ref-type="table" rid="T2">Table 2</xref>) and qualitative (<xref ref-type="table" rid="T4">Table 4</xref>) analyses. In terms of evidence of <italic>exemplary research design</italic>, the study or evaluation must also be grounded in a theoretical framework (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>) and the sample size needs to be comprised of fifty or more subjects per group for the quantitative analysis and twenty or more subjects per group for the qualitative analysis (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; Diamond et al., <xref ref-type="bibr" rid="B15">2016</xref>; Creswell and Poth, <xref ref-type="bibr" rid="B12">2018</xref>). Based on additional recommendations from ISE experts, I also developed criteria indicative of <italic>strong, adequate</italic>, and <italic>weak</italic> research design for mixed methods studies and evaluations (<xref ref-type="table" rid="T6">Table 6</xref>). Lastly, as recommended by the focus group, I developed a STEM Research Design Worksheet for mixed methods studies as a key to guide researchers, practitioners, or evaluators through the process of assessing research design quality (<xref ref-type="fig" rid="F5">Figure 5</xref>).</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>STEM research design rubric for mixed methods studies.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center" colspan="4"><bold>STEM Research Design Rubric (Mixed methods studies)</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>(4) Exemplary research design</bold></td>
<td valign="top" align="left"><bold>(3) Strong research design</bold></td>
<td valign="top" align="left"><bold>(2) Adequate research design</bold></td>
<td valign="top" align="left"><bold>(1) Weak research design</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02022; quantitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group (e.g., pre-test/post-test design with comparison; post-test only design with comparison) <break/>&#x02022; qualitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study <break/>&#x02022; study grounded in a theoretical framework <break/>&#x02022; quantitative analyses: minimum of 50 subjects per group <break/>&#x02022; qualitative analyses: minimum of 20 subjects per group <break/>&#x02022; Note: If there is no comparison group, starting value is 3</td>
<td valign="top" align="left">&#x02022; quantitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group OR <italic>quantitative design without comparison group</italic> (e.g., pre-test/post-test, post-test only, and/or time series design(s) without comparison) <break/>&#x02022; qualitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study OR <italic>qualitative design without comparison group</italic> <break/>&#x02022; study grounded in a theoretical framework <break/>&#x02022; quantitative analyses: minimum of 40 subjects per group <break/>&#x02022; qualitative analyses: minimum of 15 subjects per group</td>
<td valign="top" align="left">&#x02022; quantitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group OR <italic>quantitative design without comparison group</italic> (e.g., pre-test/post-test, post-test only, and/or time series design(s) without comparison) <break/>&#x02022; qualitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study OR <italic>qualitative design without comparison group</italic> <break/>&#x02022; study may or may not be grounded in a theoretical framework <break/>&#x02022; quantitative analyses: minimum of 25 subjects per group <break/>&#x02022; qualitative analyses: minimum of 10 subjects per group</td>
<td valign="top" align="left">&#x02022; quantitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group OR <italic>quantitative design without comparison group</italic> (e.g., pre-test/post-test, post-test only, and/or time series design(s) without comparison) <break/>&#x02022; qualitative analysis: <italic>random assignment</italic> (experimental) OR <italic>non-random assignment</italic> (quasi-experimental) using one of the following designs: narrative study, phenomenological study, grounded theory study, ethnographic study OR <italic>qualitative design without comparison group</italic> <break/>&#x02022; study may or may not be grounded in a theoretical framework <break/>&#x02022; quantitative analyses: &#x0003C;25 subjects per group <break/>&#x02022; qualitative analyses: &#x0003C;10 subjects per group <break/>&#x02022; sample size not reported</td>
</tr>
<tr>
<td/>
<td/>
<td/>
<td valign="top" align="left"><bold>Rubric Score:</bold> <italic><bold>__________</bold></italic></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Note: In order to receive a rubric score (4, 3, 2, or 1), a study must meet the criteria of all items in a given column</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>STEM Research Design Rubric worksheet for mixed methods studies.</p></caption>
<graphic xlink:href="feduc-05-554806-g0005.tif"/>
</fig>
</sec>
<sec>
<title>A STEM Impact Rubric for Mixed Methods Studies</title>
<p>Lastly, a STEM Impact Rubric for mixed methods studies (<xref ref-type="table" rid="T7">Table 7</xref>) was designed in accordance with research-based recommendations from ISE experts (Institute for Learning Innovation, <xref ref-type="bibr" rid="B28">2007</xref>; U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>; Friedman, <xref ref-type="bibr" rid="B19">2008</xref>; National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B43">2015</xref>). In terms of <italic>exemplary evidence</italic> of impact, the mixed methods study or evaluation must demonstrate the following: (1) a quantitative analysis in which there is a significant difference between the comparison (well-matched control) and treatment (program participants) groups and (2) a qualitative analysis in which researchers identify the outcome of interest as an emerging theme in the treatment but not in the comparison (well-matched control) group or a qualitative analysis in which a minimum of 75% of data (e.g., interview data, response to open-ended question) is indicative of the outcome of interest in the treatment but not the comparison (well-matched control) group; this difference needs to be statistically significant. Based on additional recommendations from ISE experts, I also developed criteria indicative of <italic>strong, moderate</italic>, and <italic>little or no evidence</italic> of impact for mixed methods studies and evaluations (<xref ref-type="table" rid="T7">Table 7</xref>). Lastly, as recommended by the focus group, I developed a STEM Impact Worksheet for mixed methods studies as a key to guide researchers, practitioners, or evaluators through the process of assessing the impact of specific outcomes (<xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>STEM impact rubric for mixed methods studies.</p></caption>
<table frame="hsides" rules="groups">
<tbody><tr>
<td valign="top" align="left" colspan="4"><bold>Title of study or evaluation</bold>:<break/> <bold>STEM outcome under review</bold>:<break/> List all method(s) used to assess the STEM outcome under review (e.g., survey question 1, survey question 2, etc.):<break/> 1. ____________________ 2. ____________________ 3. ____________________ 4. ____________________ 5. ____________________<break/> 6. ____________________ 7. ____________________ 8. ____________________ 9. ____________________ 10. ____________________<break/><break/> <bold>Directions</bold>:<break/> 1. Identify the STEM outcome under review (e.g., increased STEM career interest).<break/> 2. List all criteria used to assess the STEM outcome under review. For example, if there were four survey questions that evaluated STEM career interest, list all four questions (e.g., survey question 1, survey question 2, open-ended question 1, etc.)<break/> 3. Use the rubric to evaluate each criterion used to assess the STEM outcome under review starting from left (exemplary evidence) to right (little or no evidence)<break/> 4. Calculate the average rubric score by dividing the sum of all rubric scores by the number of criteria used to assess STEM outcomes</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="center" colspan="4" style="border-bottom: thin solid #000000;"><bold>STEM Impact Rubric (Mixed methods studies)</bold></td>
</tr> <tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>(4) Exemplary evidence</bold></td>
<td valign="top" align="left"><bold>(3) Strong evidence</bold></td>
<td valign="top" align="left"><bold>(2) Moderate evidence</bold></td>
<td valign="top" align="left"><bold>(1) Little or No evidence</bold></td>
</tr> <tr>
<td valign="top" align="left">&#x02022; quantitative analysis: statistically significant difference between comparison (control) and treatment groups (program participants) <break/>&#x02022; qualitative analysis: researchers identify outcome of interest as an emerging theme during qualitative analyses for treatment but not comparison OR minimum of 75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted outcome of interest and there is a statistically significant difference between the comparison and treatment groups <break/>&#x02022; Note: If there is no comparison group, starting value is 3</td>
<td valign="top" align="left">&#x02022; quantitative analysis: statistically significant difference between pre- and post-survey OR minimum of 75% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program survey <break/>&#x02022; qualitative analysis without comparison group: researchers identify outcome of interest as an emerging theme during qualitative analyses OR minimum of 75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted outcome of interest</td>
<td valign="top" align="left">&#x02022; quantitative analysis: 40-75% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program surveys OR STEM outcome is maintained in comparisons of pre- and post-surveys <break/>&#x02022; qualitative analysis without comparison group: 40-75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted STEM outcome of interest</td>
<td valign="top" align="left">&#x02022; quantitative analysis: less than 40% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program surveys OR STEM outcome significantly decreases in comparisons of pre- and post-surveys OR outcomes of participants are the same or significantly lower than outcomes of comparison group <break/>&#x02022; qualitative analysis without comparison group: 40-75% of data (e.g., interview data; responses to open-ended questions) indicate that the program impacted STEM outcome of interest OR qualitative analysis with comparison group: no difference between groups in outcome of interest</td>
</tr>
<tr>
<td valign="top" align="center" colspan="4"><bold>Average Rubric Score (sum of all rubric scores</bold> <bold>&#x000F7;</bold> <bold>number of criteria used to assess STEM outcomes):</bold> <italic><bold>__________</bold></italic></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>STEM Impact Rubric worksheet for mixed methods studies.</p></caption>
<graphic xlink:href="feduc-05-554806-g0006.tif"/>
</fig>
</sec>
<sec>
<title>Case Studies</title>
<p>In the next section, I present case studies demonstrating how the STEM Research Design and STEM Impact Rubrics can be used to assess research design quality and measure evidence of impact for specific outcomes. In accordance with recommendations from the Academic Competitiveness Council (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>), I identified three select studies that measure one of the following: STEM major or STEM career awareness, interest, or engagement. The first case study provides an example of how to use the rubrics to assess a quantitative study; the second case study provides an example of how to assess a qualitative study; the third case study provides an example of how to assess a mixed methods study. While these case studies focus specifically on awareness, interest, and engagement, a STEM practitioner may use these rubrics to assess any outcome of interest and to compare studies that assess comparable outcomes.</p>
</sec>
<sec>
<title>Case Study 1: Stanford Medical Youth Science Program</title>
<p>The Stanford Medical Youth Science Program is a biomedical pipeline program for high school students. The goal of this program is to diversify participation in the health professions (Winkleby, <xref ref-type="bibr" rid="B56">2007</xref>). This 5 week residential summer program includes classroom-based workshops, anatomy and pathology practicums, hospital field placements, research projects, and college readiness advisement. In 2009, Winkleby et al. (<xref ref-type="bibr" rid="B57">2009</xref>) published a quantitative study of the STEM outcomes of program participants. Two specific outcomes measured in this study were whether alumni of this program (1) majored in a STEM discipline or (2) engaged in a STEM career. I used the STEM Research Design Rubric (<xref ref-type="table" rid="T2">Table 2</xref>) to assess quality of research design and the STEM Impact Rubric (<xref ref-type="table" rid="T3">Table 3</xref>) to measure evidence of outcome (engagement in a STEM major; engagement in a STEM career). This process is also depicted graphically in <xref ref-type="fig" rid="F7">Figure 7A</xref>.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>(A)</bold> Graphical depiction of case study 1. <bold>(B)</bold> Graphical depiction of case study 2. <bold>(C)</bold> Graphical depiction of case study 3.</p></caption>
<graphic xlink:href="feduc-05-554806-g0007.tif"/>
</fig>
<p>In terms of research design, first I examined evidence of <italic>exemplary research design</italic> (<xref ref-type="table" rid="T2">Table 2</xref>, column 1). Since the study did not include either a random or non-random comparison group, I moved on to the second column: <italic>strong research design</italic>. First, I checked the first bullet point in column two. Since the study was a <italic>quantitative design without a comparison group</italic> (post-test only), it met the first criterion for <italic>strong research design</italic>. Next, I checked the second bullet point in column two. Since the study was grounded in two theoretical frameworks: (1) Cognitive Apprenticeship (Collins et al., <xref ref-type="bibr" rid="B9">1991</xref>) and (2) Situated Learning (Lave and Wenger, <xref ref-type="bibr" rid="B33">1991</xref>), it met the second criterion for <italic>strong research design</italic>. Lastly, I checked the third bullet point in column two. Since the sample size well-exceeded the minimum of 40, it met the third criterion for <italic>strong research design</italic>. Thus, based on the STEM Research Design Rubric, this study was classified as a <italic>strong research design</italic>.</p>
<p>In terms of evidence of outcome (engagement in a STEM major or STEM career), first I assessed criteria for <italic>exemplary evidence</italic> (<xref ref-type="table" rid="T3">Table 3</xref>, column 1). Since there was no statistical comparison between program participants and a comparison group, I moved on to the second column (<italic>strong evidence</italic>). The criteria for <italic>strong evidence</italic> of impact included two possible outcomes: (1) statistically significant difference between pre- and post-survey or (2) minimum of 75% of participants indicate higher than median (e.g., very likely; likely) outcome on post-program survey. Neither of these outcomes was found for STEM major engagement or STEM career engagement. Next, I moved on to the third column (<italic>moderate evidence</italic> of impact). A survey of alumni of the Stanford Medical Youth Science Program indicated that 57.1% were engaged in a STEM major, which met the criteria for <italic>moderate evidence</italic> (a favorable response by 40&#x02013;75% of participants). However, only 33.1% of alumni were engaged in a STEM career; this met the criteria for <italic>little or no evidence</italic> (<xref ref-type="table" rid="T3">Table 3</xref>, column 4). Overall, this study provided evidence of a <italic>strong research design</italic> and <italic>moderate evidence</italic> of impact, in terms of engagement in a STEM major, and <italic>little or no evidence</italic> of impact in terms of STEM career engagement.</p>
</sec>
<sec>
<title>Case Study 2: The Source (Game Changer Chicago Design Lab, University of Chicago)</title>
<p>The Source is a 5 week summer program that uses alternative reality games for teaching engineering concepts to high school students (Gilliam et al., <xref ref-type="bibr" rid="B22">2017</xref>). The weekly program includes 1 day of online activities off-campus and 3 to 4 days of on-campus activities focused on workshops in different STEM subject areas. Gilliam et al. (<xref ref-type="bibr" rid="B22">2017</xref>) published a qualitative study describing the STEM outcomes of high school participants of this program. To assess the quality of the research design of this study and to measure evidence of impact (in this case STEM career awareness), I used both the STEM Research Design Rubric for qualitative studies (<xref ref-type="table" rid="T4">Table 4</xref>) and the STEM Impact Rubric for qualitative studies (<xref ref-type="table" rid="T5">Table 5</xref>) as described below. This process is also depicted graphically in <xref ref-type="fig" rid="F7">Figure 7B</xref>.</p>
<p>In terms of research design, first I examined evidence of <italic>exemplary research design</italic> (<xref ref-type="table" rid="T4">Table 4</xref>, column 1). Since the study did not include either a random or nonrandom comparison group, I moved on to the second column: <italic>strong research design</italic>. First, I checked the first bullet point in column two. Since the study was a <italic>qualitative design without comparison group</italic>, it met the first criterion for <italic>strong research design</italic>. Next, I checked the second bullet point in column two. Since the study was grounded in multiple theoretical frameworks, including Situated Learning Theory (Lave and Wenger, <xref ref-type="bibr" rid="B33">1991</xref>), it met the second criterion for <italic>strong research design</italic>. Lastly, I checked the third bullet point in column two. Since 43 students were interviewed, the sample size well-exceeded the minimum of 15 subjects; thus, it met the third criterion for <italic>strong research design</italic>. Therefore, based on the STEM Research Design Rubric for qualitative studies, this study was classified as a <italic>strong research design</italic>.</p>
<p>In terms of evidence of outcome (in this case, STEM career awareness), first I used the STEM Impact Rubric for qualitative studies to assess criteria for <italic>exemplary evidence</italic> (<xref ref-type="table" rid="T5">Table 5</xref>, column 1). Since there was no comparison group in this analysis, I moved on to the second column and tested for <italic>strong evidence</italic> of impact. One criterion indicative of <italic>strong evidence</italic> of impact is if the <italic>researchers identify outcome of interest as an emerging theme during qualitative analyses</italic>. A theme that emerged from this study was &#x0201C;Mentoring and Exposure to STEM Professionals,&#x0201D; which provided evidence that participants became more aware of STEM career opportunities as a result of their experiences in the program. Thus, this study provides <italic>strong evidence</italic> of impact. Overall, this study provided evidence of a <italic>strong research design</italic> and <italic>strong evidence</italic> of outcome, in this case, STEM career awareness.</p>
</sec>
<sec>
<title>Case Study 3: The Lang Science Program (American Museum of Natural History)</title>
<p>The Lang Science Program is a comprehensive 7 year program for middle school and high school students facilitated by the American Museum of Natural History. The program takes place on alternating Saturdays during the academic year and for 3 weeks during the summer for a minimum of 165 contact hours per year. The program centers on the research disciplines of the Museum&#x02014;the biological sciences, Earth and planetary sciences, and anthropological sciences. As students transition from middle school to high school, they increasingly engage in authentic research projects alongside scientists and educators as well as career and college readiness workshops (Habig et al., <xref ref-type="bibr" rid="B24">2018</xref>). A mixed methods study of the Lang Program centered on STEM major engagement and STEM career awareness (Habig et al., <xref ref-type="bibr" rid="B24">2018</xref>). To assess the quality of the research design of this study and to measure evidence of impact, I used both the STEM Research Design Rubric for mixed methods studies (<xref ref-type="table" rid="T6">Table 6</xref>) and the STEM Impact Rubric for mixed methods studies (<xref ref-type="table" rid="T7">Table 7</xref>) as described below. This process is depicted graphically in <xref ref-type="fig" rid="F7">Figure 7C</xref>.</p>
<p>In terms of research design, first I examined evidence of <italic>exemplary research design</italic> by assessing the first bullet point in column 1 (<xref ref-type="table" rid="T6">Table 6</xref>). Since the study design met the criteria of the first bullet point (a quantitative analysis that included a comparison group), I moved on to the second bullet point in column 1. Since there was no comparison group for the qualitative analysis, this study did not meet the criteria for an <italic>exemplary research design</italic>. Therefore, I moved on to the second column: <italic>strong research design</italic>. Since the study met the criteria for the first two bullet points in column 2 (a quantitative analysis with <italic>non-random assignment</italic> (quasi-experimental) of treatment and comparison group; a <italic>qualitative design without comparison group</italic>), I next checked the third bullet point in column 2. Since the study was grounded in two theoretical frameworks: (1) communities of practice (Lave and Wenger, <xref ref-type="bibr" rid="B33">1991</xref>) and (2) possible selves (Markus and Nurius, <xref ref-type="bibr" rid="B35">1986</xref>); it met the third criterion for <italic>strong research design</italic>. Next, I checked bullet points five and six (sample sizes). Since the sample sizes exceeded the minimum of 40 subjects for quantitative analyses and 15 subjects for qualitative analyses, this study met the fifth and sixth criteria for <italic>strong research design</italic>. Therefore, based on the STEM Research Design Rubric for mixed methods studies, this study was classified as a <italic>strong research design</italic>.</p>
<p>In terms of evidence of outcome, first I assessed the first outcome of interest: STEM major engagement. Four quantitative criteria were used to measure STEM major engagement: (1) percentage of STEM majors; (2) percentage of STEM majors compared to college students nationally; (3) percentage of STEM majors compared to college students in New York City; (4) percentage of STEM majors compared to students who attended specialized science, mathematics, and technology high schools. Since 80.3% of alumni engaged in a STEM major, this outcome was indicative of <italic>strong evidence</italic> of impact (<xref ref-type="table" rid="T7">Table 7</xref>, column 2). Since the percentage of STEM majors of the Lang program was significantly higher than college students nationally and locally, these two outcomes were indicative of <italic>exemplary evidence</italic> of impact (<xref ref-type="table" rid="T7">Table 7</xref>, column 1). However, since there was no significant difference in the percentage of STEM majors when comparing Lang alumni to alumni of specialized science, mathematics, and technology high schools, this last outcome was indicative of <italic>little or no evidence</italic> of impact. After averaging the four outcomes, the mean STEM rubric score was 3 suggestive of <italic>strong evidence</italic> of impact overall. For the second outcome, STEM career awareness, a qualitative analysis was used to measure evidence of impact. Based on the STEM Impact Rubric (<xref ref-type="table" rid="T7">Table 7</xref>, column 2), one criterion indicative of <italic>strong evidence</italic> of impact is if the <italic>researchers identify outcome of interest as an emerging theme during qualitative analyses</italic>. A theme that emerged from this study was &#x0201C;Discovering Possible Selves,&#x0201D; which provided evidence that participants were exposed to and became more aware of STEM career opportunities as a result of their experiences in the program. Thus, this outcome was indicative of <italic>strong evidence</italic> of impact. Overall, this study provided evidence of a <italic>strong research design</italic> and <italic>strong evidence</italic> of outcomes in terms of STEM major engagement and STEM career awareness, although the former outcome varied considerably based on the analysis used to measure impact.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>Discussion</title>
<p>The purpose of this study was to develop user-friendly rubrics that can be used by ISE STEM researchers, practitioners, and evaluators to improve quality of research designs and to measure evidence of outcomes. By consulting with informal learning experts and by leveraging several sources of data from STEM governing bodies, non-profit organizations, and ISE institutions, I identified research-based recommendations on how to assess research designs and on how to measure evidence of impact. With the feedback of STEM practitioners and recommendations from STEM stakeholders, the end products were two types of rubrics for quantitative, qualitative, and mixed methods ISE studies: (1) a STEM Research Design Rubric and (2) a STEM Impact Rubric. These tools were found to be especially applicable for assessing informal learning studies designed to inspire and motivate students to consider a STEM pathway. Here, I discuss what was learned from this process, the general applicability of these rubrics, and their respective limitations.</p>
<p>When designing the STEM Research Design and STEM Impact Rubrics, I used the theories of impact analysis (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>) and triangulation (Denzin, <xref ref-type="bibr" rid="B14">1970</xref>; Ammenwerth et al., <xref ref-type="bibr" rid="B1">2003</xref>; Flick, <xref ref-type="bibr" rid="B17">2018a</xref>,<xref ref-type="bibr" rid="B18">b</xref>) to inform my research. By espousing the principles of impact analysis in alignment with triangulation theory, this study allowed for ISE educators to carefully consider different hierarchical levels for designing effective studies, ranging from higher-level designs (i.e., experimental and quasi-experimental designs) to lower-level designs (pre/post studies; comparison groups without careful matching, etc.), while simultaneously considering the unique characteristics of ISE programs (National Research Council, <xref ref-type="bibr" rid="B38">2009</xref>, <xref ref-type="bibr" rid="B39">2010</xref>). Moreover, because many ISE studies use qualitative designs to assess participants&#x00027; outcomes, the application of impact analysis theory, coupled with research-based recommendations, were particularly instructional during the process of developing rubrics for qualitative study design and impact.</p>
<p>Based on these results, ISE practitioners might consider the inclusion of comparison groups in qualitative study designs (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>). For example, the use of open-ended survey questions or semi-structured interviews that qualitatively compare program participants to well-matched comparisons are practices that will help to increase the internal validity of ISE studies (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>). Lastly, the theories of impact analysis and triangulation were particularly informative when designing rubrics for mixed methods studies as these studies provide multiple sources of data, both quantitative and qualitative, and help to increase validity and paint a more complete picture of participants&#x00027; STEM outcomes (Denzin, <xref ref-type="bibr" rid="B14">1970</xref>; Ammenwerth et al., <xref ref-type="bibr" rid="B1">2003</xref>).</p>
<p>The STEM Research Design and STEM Impact Rubrics developed in this study are practical tools that can be used by ISE researchers, practitioners, and evaluators to improve the field of informal science learning. First, depending on the goals of a study, the STEM Research Design Rubric is a useful tool for designing a study and for deciding which of the three hierarchical levels of study design is most appropriate for a given study (U.S. Department of Education, <xref ref-type="bibr" rid="B53">2007</xref>). In some cases, especially studies receiving external funding and that focus on long-term outcomes, it might be most appropriate to design an experimental or quasi-experimental study. For example, in the third case study presented in this paper (Habig et al., <xref ref-type="bibr" rid="B24">2018</xref>), a quasi-experimental design was applied to compare STEM major engagement outcomes between museum program participants and comparison groups. If the research team that conducted this study had access to the STEM Research Design Rubric prior to conducting this study, the authors might have opted to design their study differently. Specifically, the researchers might have matched program participants to a comparison group with shared demographic characteristics and used propensity score analysis to compare STEM major outcomes between groups (Rosenbaum and Rubin, <xref ref-type="bibr" rid="B47">1983</xref>; Hahs-Vaughn and Onwuegbuzie, <xref ref-type="bibr" rid="B25">2006</xref>). Alternatively, prior to the study, the research team might opt to randomly select students to participate in the program via lottery and then compare the outcomes between the treatment and comparison groups (e.g., Hubelbank et al., <xref ref-type="bibr" rid="B26">2007</xref>). In other cases, especially when funding is limited and the study is focusing on a short-term outcome, it might be appropriate to design a study using one of the lower hierarchical levels. For example, a study of a 1 day engineering outreach event for 4th&#x02212;7th grade Girl Scouts used a pre/post study design to assess whether participation in this program increased participants&#x00027; awareness of engineering careers (Christman et al., <xref ref-type="bibr" rid="B7">2008</xref>). Thus, for programs with short-term outcomes and/or little or no funding, it might be most appropriate to design quasi-experimental studies or more simple designs such as a pre/post study design. Practically speaking, studies that assess short-term outcomes, such as STEM major and STEM career awareness, might operate at lower hierarchical levels (e.g., pre/post study design); studies that assess intermediate STEM outcomes, such as STEM major and STEM career interest, might operate at middle hierarchical levels (e.g., quasi-experimental), and programs that assess long-terms STEM outcomes, such as STEM major and STEM career engagement, might operate at higher hierarchical levels (e.g., experimental design) (Cooper et al., <xref ref-type="bibr" rid="B10">2000</xref>; Wilkerson and Haden, <xref ref-type="bibr" rid="B55">2014</xref>).</p>
<p>Second, while the STEM Research Design Rubric is useful for informing study design, the STEM Impact Rubric is an important companion tool for helping ISE researchers, practitioners, and evaluators assess evidence of impact. Critically, the more rigorous the research design (<italic>exemplary research design</italic>), the more likely you can trust the study&#x00027;s validity (Mohr, <xref ref-type="bibr" rid="B37">1995</xref>). If a study provides evidence of <italic>exemplary research design</italic>, then the results of the STEM Impact Rubric can be used to more confidently claim evidence of outcome. Thus, the STEM Research Design Rubric is a practical tool for developing rigorous study designs and in turn, the STEM Impact Rubric is a practical tool for providing evidence of impact. Importantly, the use of these rubrics across different ISE studies will ensure consistency in research design and measurement of impact, and when applicable, the results of these rubrics can be used to refine programs especially when there is <italic>little or no evidence</italic> of impact. Notably, when a STEM Impact Rubric is informed by a study with evidence of <italic>strong</italic> or <italic>exemplary research design</italic>, these results can be used for reports to government officials and funding agencies and can be potential sources for increased funding.</p>
<p>Finally, while the development of STEM Research Design and STEM Impact Rubrics are critical for improving the field of informal science learning and for promoting consistency between studies, it should be noted that there are several limitations of these tools. First, if a study is not designed rigorously, then we cannot reliably infer evidence of outcome. In data science, the term &#x0201C;garbage-in, garbage-out&#x0201D; is used to describe a situation in which the quality of the output is linked to the quality of the input (Rose and Fischer, <xref ref-type="bibr" rid="B46">2011</xref>). Analogously, evidence of STEM outcomes (based on the STEM Impact Rubric) is linked to the quality of study design (based on the STEM Research Design Rubric). Thus, without a well-designed study, it is virtually impossible to confidently infer evidence of impact. In support, the National Research Council (<xref ref-type="bibr" rid="B41">2013</xref>) recommends that results from non-rigorous study designs (e.g., pre/post study design; comparison group without careful comparison) are appropriate for refining hypotheses that can be used to inform more rigorous study designs conducted in the future; however, these results should not be interpreted as conclusive evidence of impact. A second limitation is that STEM stakeholders might not always agree on the criteria used to assess research design or evidence of impact. For example, in the present study, there were some disagreements among the focus group members about whether or not a research design needs to be grounded in a theoretical framework as recommended by the Institute for Learning Innovation (<xref ref-type="bibr" rid="B28">2007</xref>). Thus, it is important to emphasize that these rubrics should be viewed as heuristics&#x02014;tools for aiding ISE stakeholders in evaluating research design and evidence of impact&#x02014;and that in some cases, depending on the goals of a particular study, it might be appropriate to alter certain criteria. Lastly, a third limitation is the challenges of applying these rubrics universally. Some ISE studies are very complex consisting of varied analyses where it might be quite difficult to assess study design and impact. For example, the third case study (Habig et al., <xref ref-type="bibr" rid="B24">2018</xref>) used four different methods to assess evidence of impact with respect to STEM major engagement. In two cases, the rubric indicated that there was <italic>exemplary evidence</italic> of impact; in one case, the rubric indicated that there was <italic>strong evidence</italic> of impact, and in the final case; the rubric indicated that there was <italic>little or no evidence</italic> of impact. While averaging the rubric score provided a rough estimate of overall impact, the choice of an inappropriate analysis might skew the results. One of the comparison groups in this quasi-experimental study was students who attended specialized science, mathematics, and technology high schools. While there were no significant differences in STEM major engagement between participants of the informal museum program and these high school students, this analysis might not accurately measure the outcome of interest&#x02014;whether participation in the ISE museum program had an impact on participants&#x00027; decisions to major in STEM. Thus, it is critical that a study design is aligned to the outcome of interest and that the appropriate comparison group is considered carefully.</p>
</sec>
<sec sec-type="conclusions" id="s6">
<title>Conclusions</title>
<p>The STEM Research Design and a STEM Impact Rubrics developed in this study are potentially useful tools for ISE researchers, practitioners, and evaluators for improving study design and for assessing the effectiveness of STEM interventions. There are several possible applications for these rubrics especially in terms of areas of future research. First, in future studies, ISE researchers can measure whether large scale use of the STEM Research Design Rubric results in overall improvements in research design across the informal learning community. Second, researchers can measure whether the STEM Impact Rubric helps to inform program design and in turn, helps ISE institutions to improve specific outcomes such as persistence in STEM. More specifically, STEM researchers can assess studies yielding <italic>exemplary evidence</italic> of impact, extract program design principles from these studies, and share best practices across institutions. Critically, and as recommended by the National Research Council (<xref ref-type="bibr" rid="B41">2013</xref>), researchers can also track outcomes of interest by race, ethnicity, language status, and socioeconomic status to ensure that programs are effective across different populations. Lastly, these rubrics can be used in meta-analyses to quantitatively compare research design quality and evidence of impact across studies and outcomes of interest. In summary, the large-scale application of the STEM Research Design Rubric and the STEM Impact Rubric has the potential to transform research design quality and more confidently measure evidence of outcomes.</p>
</sec>
<sec sec-type="data-availability-statement" id="s7">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s8">
<title>Ethics Statement</title>
<p>Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. Written informed consent was not required to participate in this study in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>The author conceived, designed, and wrote the manuscript for this study.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack><p>The author would like to thank the following individuals for providing feedback during the design of the STEM Research Design and STEM Impact Rubrics: Rachel Chaffee, Jennifer Cosme, Leah Golubchick, Preeti Gupta, Brian Levine, Maleha Mahmud, Nickcoles Martinez, Maria Strangas, Mark Weckel, and Alex Watford. The author would also like to thank Veronica Catete and Kathryn Holmes for providing commentary on previous drafts of this manuscript.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ammenwerth</surname> <given-names>E.</given-names></name> <name><surname>Iller</surname> <given-names>C.</given-names></name> <name><surname>Mansmann</surname> <given-names>U.</given-names></name></person-group> (<year>2003</year>). <article-title>Can evaluation studies benefit from triangulation? A case study</article-title>. <source>Int. J. Med. Inform.</source> <volume>70</volume>, <fpage>237</fpage>&#x02013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.1016/S1386-5056(03)00059-5</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baer</surname> <given-names>J.</given-names></name> <name><surname>McKool</surname> <given-names>S. S.</given-names></name></person-group> (<year>2014</year>). <article-title>The gold standard for assessing creativity</article-title>. <source>Int. J. Qual. Assur. Eng. Tech. Educ.</source> <volume>3</volume>, <fpage>81</fpage>&#x02013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.4018/ijqaete.2014010104</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Blanchard</surname> <given-names>M. R.</given-names></name> <name><surname>Gutierrez</surname> <given-names>K. S.</given-names></name> <name><surname>Habig</surname> <given-names>B.</given-names></name> <name><surname>Gupta</surname> <given-names>P.</given-names></name> <name><surname>Adams</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Informal STEM Education</article-title>, in <source>Handbook of Research on STEM Education</source>, eds <person-group person-group-type="editor"><name><surname>Johnson</surname> <given-names>C.</given-names></name> <name><surname>Mohr-Schroeder</surname> <given-names>M.</given-names></name> <name><surname>Moore</surname> <given-names>T.</given-names></name> <name><surname>English</surname> <given-names>L.</given-names></name></person-group> (<publisher-loc>Routledge</publisher-loc>: <publisher-name>Taylor &#x00026; Francis</publisher-name>) <fpage>138</fpage>&#x02013;<lpage>151</lpage>. <pub-id pub-id-type="pmid">30984064</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>D.</given-names></name> <name><surname>Stanley</surname> <given-names>J.</given-names></name></person-group> (<year>1963</year>). <source>Experimental and Quasi-Experimental Designs for Research</source>. <publisher-loc>Chicago, IL</publisher-loc>: <publisher-name>Rand McNally</publisher-name>.</citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campbell</surname> <given-names>D. T.</given-names></name> <name><surname>Fiske</surname> <given-names>D. W.</given-names></name></person-group> (<year>1959</year>). <article-title>Convergent and discriminant validation by the multitrait-multimethod matrix</article-title>. <source>Psychol. Bull.</source> <volume>56</volume>, <fpage>81</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1037/h0046016</pub-id><pub-id pub-id-type="pmid">13634291</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>H. Y.</given-names></name> <name><surname>Choi</surname> <given-names>H.</given-names></name> <name><surname>Hailu</surname> <given-names>M. F.</given-names></name> <name><surname>Whitford</surname> <given-names>M.</given-names></name> <name><surname>Duplechain DeRouen</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Participation in structured STEM-focused out-of-school time programs in secondary school: linkage to postsecondary STEM aspiration and major</article-title>. <source>J. Res. Sci. Teach</source>. <volume>57</volume>, <fpage>1250</fpage>&#x02013;<lpage>1280</lpage>. <pub-id pub-id-type="doi">10.1002/tea.21629</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Christman</surname> <given-names>K. A.</given-names></name> <name><surname>Hankemeier</surname> <given-names>S.</given-names></name> <name><surname>Hunter</surname> <given-names>J.</given-names></name> <name><surname>Jennings</surname> <given-names>J.</given-names></name> <name><surname>Moser</surname> <given-names>D.</given-names></name> <name><surname>Stiles</surname> <given-names>S.</given-names></name></person-group> (<year>2008</year>). <article-title>Overnights encourage girls&#x00027; interest in science-related careers</article-title>. <source>J. Youth Dev.</source> <volume>3</volume>, <fpage>89</fpage>&#x02013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.5195/JYD.2008.322</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1968</year>). <article-title>Weighted kappa: nominal scale agreement provision for scaled disagreement or partial credit</article-title>. <source>Psychol. Bull.</source> <volume>70</volume>, <fpage>213</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1037/h0026256</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Collins</surname> <given-names>A.</given-names></name> <name><surname>Brown</surname> <given-names>J. S.</given-names></name> <name><surname>Holum</surname> <given-names>A.</given-names></name></person-group> (<year>1991</year>). <article-title>Cognitive apprenticeship: making thinking visible</article-title>. <source>Am. Educ.</source> <volume>75</volume>, <fpage>6</fpage>&#x02013;<lpage>11</lpage>.</citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cooper</surname> <given-names>H.</given-names></name> <name><surname>Charlton</surname> <given-names>K.</given-names></name> <name><surname>Valentine</surname> <given-names>J. C.</given-names></name> <name><surname>Muhlenbruck</surname> <given-names>L.</given-names></name></person-group> (<year>2000</year>). <article-title>Making the most of summer school: a meta-analytic and narrative review</article-title>. <source>Monog. Soc. Res. Child. Dev.</source> <volume>65</volume>, <fpage>i</fpage>&#x02013;<lpage>v</lpage>. <pub-id pub-id-type="doi">10.1111/1540-5834.00058</pub-id><pub-id pub-id-type="pmid">12467098</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Creswell</surname> <given-names>J. W.</given-names></name> <name><surname>Miller</surname> <given-names>D. L.</given-names></name></person-group> (<year>2000</year>). <article-title>Determining validity in qualitative inquiry</article-title>. <source>Theor. Pract</source>. <volume>39</volume>, <fpage>124</fpage>&#x02013;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.1207/s15430421tip3903_2</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Creswell</surname> <given-names>J. W.</given-names></name> <name><surname>Poth</surname> <given-names>C. N.</given-names></name></person-group> (<year>2018</year>). <source>Qualitative Inquiry and Research Design: Choosing Among Five Approaches</source>, <edition>4th Edn</edition>. <publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cuddeback</surname> <given-names>L.</given-names></name> <name><surname>Idema</surname> <given-names>J.</given-names></name> <name><surname>Daniel</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>Lions, tigers, and teens: promoting interest in science as a career path through teen volunteering</article-title>. <source>IZE J.</source> <volume>2019</volume>:<fpage>10</fpage>.</citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Denzin</surname> <given-names>N. K.</given-names></name></person-group> (<year>1970</year>). <source>The Research Act in Sociology</source>, <publisher-loc>Chicago, IL</publisher-loc>: <publisher-name>Aldine</publisher-name>. <pub-id pub-id-type="pmid">30990805</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Diamond</surname> <given-names>J.</given-names></name> <name><surname>Horn</surname> <given-names>M.</given-names></name> <name><surname>Uttal</surname> <given-names>D. H.</given-names></name></person-group> (<year>2016</year>). <source>Practical Evaluation Guide: Tools for Museums and Other Informal Educational Settings</source>. <publisher-loc>Lanham, MD</publisher-loc>: <publisher-name>Rowman and Littlefield</publisher-name>.</citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fadigan</surname> <given-names>K. A.</given-names></name> <name><surname>Hammrich</surname> <given-names>P. L.</given-names></name></person-group> (<year>2004</year>). <article-title>A longitudinal study of the educational and career trajectories of female participants of an urban informal science education program</article-title>. <source>J. Res. Sci. Teach.</source> <volume>41</volume>, <fpage>835</fpage>&#x02013;<lpage>860</lpage>. <pub-id pub-id-type="doi">10.1002/tea.20026</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Flick</surname> <given-names>U.</given-names></name></person-group> (<year>2018a</year>). <source>Doing Triangulation and Mixed Methods</source>, Vol. <volume>8</volume>. <publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Flick</surname> <given-names>U.</given-names></name></person-group> (<year>2018b</year>). <article-title>Triangulation in data collection</article-title>, in <source>The Sage Handbook of Qualitative Data Collection</source>, ed <person-group person-group-type="editor"><name><surname>Flick</surname> <given-names>U.</given-names></name></person-group> (<publisher-loc>London</publisher-loc>: <publisher-name>SAGE Publications Ltd</publisher-name>), <fpage>527</fpage>&#x02013;<lpage>544</lpage>.</citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Friedman</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <source>Framework for Evaluating Impacts of Informal Science Education Projects</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>National Science Foundation</publisher-name>.</citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fuchs</surname> <given-names>L. S.</given-names></name> <name><surname>Fuchs</surname> <given-names>D.</given-names></name></person-group> (<year>1986</year>). <article-title>Effects of systematic formative evaluation: a meta-analysis</article-title>. <source>Except. Child.</source> <volume>53</volume>, <fpage>199</fpage>&#x02013;<lpage>208</lpage>. <pub-id pub-id-type="doi">10.1177/001440298605300301</pub-id><pub-id pub-id-type="pmid">3792417</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gannon</surname> <given-names>F.</given-names></name></person-group> (<year>2001</year>). <article-title>The essential role of peer review</article-title>. <source>EMBO Rep.</source> <volume>2</volume>, <fpage>743</fpage>&#x02013;<lpage>743</lpage>. <pub-id pub-id-type="doi">10.1093/embo-reports/kve188</pub-id><pub-id pub-id-type="pmid">11559578</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gilliam</surname> <given-names>M.</given-names></name> <name><surname>Jagoda</surname> <given-names>P.</given-names></name> <name><surname>Fabiyi</surname> <given-names>C.</given-names></name> <name><surname>Lyman</surname> <given-names>P.</given-names></name> <name><surname>Wilson</surname> <given-names>C.</given-names></name> <name><surname>Hill</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Alternate reality games as an informal learning tool for generating STEM engagement among underrepresented youth: a qualitative evaluation of the source</article-title>. <source>J. Sci. Educ. Technol.</source> <volume>26</volume>, <fpage>295</fpage>&#x02013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1007/s10956-016-9679-4</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Greene</surname> <given-names>J.</given-names></name> <name><surname>McClintock</surname> <given-names>C.</given-names></name></person-group> (<year>1985</year>). <article-title>Triangulation in evaluation: design and analysis issues</article-title>. <source>Eval. Rev.</source> <volume>9</volume>, <fpage>523</fpage>&#x02013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1177/0193841X8500900501</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Habig</surname> <given-names>B.</given-names></name> <name><surname>Gupta</surname> <given-names>P.</given-names></name> <name><surname>Levine</surname> <given-names>B.</given-names></name> <name><surname>Adams</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>An informal science education program&#x00027;s impact on STEM major and STEM career outcomes</article-title>. <source>Res. Sci. Ed.</source> <volume>50</volume>, <fpage>1051</fpage>&#x02013;<lpage>1074</lpage>. <pub-id pub-id-type="doi">10.1007/s11165-018-9722-y</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hahs-Vaughn</surname> <given-names>D. L.</given-names></name> <name><surname>Onwuegbuzie</surname> <given-names>A. J.</given-names></name></person-group> (<year>2006</year>). <article-title>Estimating and using propensity score analysis with complex samples</article-title>. <source>J. Exp. Educ.</source> <volume>75</volume>, <fpage>31</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.3200/JEXE.75.1.31-65</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Hubelbank</surname> <given-names>J.</given-names></name> <name><surname>Demetry</surname> <given-names>C.</given-names></name> <name><surname>Nicholson</surname> <given-names>M. E.</given-names></name> <name><surname>Blaisdell</surname> <given-names>S.</given-names></name> <name><surname>Quinn</surname> <given-names>P.</given-names></name> <name><surname>Rosenthal</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>Long term effects of a middle school engineering outreach program for girls: a controlled study</article-title>, in <source>Paper Presented at 2007 Annual Conference &#x00026; Exposition, Honolulu, Hawaii</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://peer.asee.org/2098">https://peer.asee.org/2098</ext-link> (accessed July 01, 2020).</citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hussein</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>The use of triangulation in social sciences research: can qualitative and quantitative methods be combined?</article-title> <source>J. Comp. Soc. Work</source> <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.31265/jcsw.v4i1.48</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="web"><person-group person-group-type="author"><collab>Institute for Learning Innovation</collab></person-group> (<year>2007</year>). <article-title>Evaluation of learning in informal learning environments</article-title>, in <source>Paper Prepared for the Committee on Science Education for Learning Science in Informal Environments</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.informalscience.org/evaluation-learning-informal-learning-environments">https://www.informalscience.org/evaluation-learning-informal-learning-environments</ext-link> (accessed July 01, 2020).</citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jackson</surname> <given-names>R. L.</given-names></name> <name><surname>Drummond</surname> <given-names>D. K.</given-names></name> <name><surname>Camara</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <article-title>What is qualitative research?</article-title> <source>Qual. Res. Rep. Commun.</source> <volume>8</volume>, <fpage>21</fpage>&#x02013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1080/17459430701617879</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jick</surname> <given-names>T. D.</given-names></name></person-group> (<year>1979</year>). <article-title>Mixing qualitative and quantitative methods: triangulation in action</article-title>. <source>Admin. Sci. Quart.</source> <volume>24</volume>, <fpage>602</fpage>&#x02013;<lpage>611</lpage>. <pub-id pub-id-type="doi">10.2307/2392366</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Klein</surname> <given-names>C.</given-names></name> <name><surname>Tisdal</surname> <given-names>C.</given-names></name> <name><surname>Hancock</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <source>Roads Taken&#x02014;Long-Term Impacts of Youth Programs</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>Association of Science Technology Centers</publisher-name>.</citation></ref>
<ref id="B32">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Krishnamurthi</surname> <given-names>A.</given-names></name> <name><surname>Ballard</surname> <given-names>M.</given-names></name> <name><surname>Noam</surname> <given-names>G. G.</given-names></name></person-group> (<year>2014</year>). <source>Examining the Impact of Afterschool STEM Programs</source>. <publisher-name>Afterschool Alliance</publisher-name>. Avaulable online at: <ext-link ext-link-type="uri" xlink:href="https://www.informalscience.org/examining-impact-afterschool-stem-programs">https://www.informalscience.org/examining-impact-afterschool-stem-programs</ext-link> (accessed July 01, 2020).</citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lave</surname> <given-names>J.</given-names></name> <name><surname>Wenger</surname> <given-names>E.</given-names></name></person-group> (<year>1991</year>). <source>Situated Learning: Legitimate Peripheral Participation</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lombard</surname> <given-names>M.</given-names></name> <name><surname>Snyder-Duch</surname> <given-names>J.</given-names></name> <name><surname>Bracken</surname> <given-names>C. C.</given-names></name></person-group> (<year>2002</year>). <article-title>Content analysis in mass communication: assessment and reporting of intercoder reliability</article-title>. <source>Hum. Commun. Res.</source> <volume>28</volume>, <fpage>587</fpage>&#x02013;<lpage>604</lpage>. <pub-id pub-id-type="doi">10.1111/j.1468-2958.2002.tb00826.x</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Markus</surname> <given-names>H.</given-names></name> <name><surname>Nurius</surname> <given-names>P.</given-names></name></person-group> (<year>1986</year>). <article-title>Possible selves</article-title>. <source>Am. Psychol.</source> <volume>41</volume>, <fpage>954</fpage>&#x02013;<lpage>969</lpage>. <pub-id pub-id-type="doi">10.1037/0003-066X.41.9.954</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>McCreedy</surname> <given-names>D.</given-names></name> <name><surname>Dierking</surname> <given-names>L. D.</given-names></name></person-group> (<year>2013</year>). <article-title>Cascading influences: long-term impacts of informal STEM experiences for girls</article-title>, in <source>Presented at 27th Annual Visitor Studies Association Conference</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.informalscience.org/cascading-influences-long-term-impacts-stem-informal-experiences-girls">https://www.informalscience.org/cascading-influences-long-term-impacts-stem-informal-experiences-girls</ext-link> (accessed July 01, 2020).</citation></ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mohr</surname> <given-names>L. B.</given-names></name></person-group> (<year>1995</year>). <source>Impact Analysis for Program Evaluation.</source> <publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation></ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><collab>National Research Council</collab></person-group> (<year>2009</year>). <source>Learning Science in Informal Environments: People, Places, and Pursuits</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation></ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><collab>National Research Council</collab></person-group> (<year>2010</year>). <source>Surrounded by Science: Learning Science in Informal Environments</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation></ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><collab>National Research Council</collab></person-group> (<year>2011</year>). <source>Successful K-12 STEM Education. Identifying Effective Approaches in Science, Technology, Engineering, and Mathematics</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation></ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><collab>National Research Council</collab></person-group> (<year>2013</year>). <source>Monitoring Progress Toward Successful K-12 STEM Education. A Nation Advancing</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><collab>National Research Council</collab></person-group> (<year>2014</year>). <source>Developing Assessments for the Next Generation Science Standards</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation></ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><collab>National Research Council</collab></person-group> (<year>2015</year>). <source>Identifying and Supporting Productive STEM Programs in Out-of-School Settings</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Panadero</surname> <given-names>E.</given-names></name> <name><surname>Jonsson</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>The use of scoring rubrics for formative assessment purposes revisited: a review</article-title>. <source>Educ. Res. Rev.</source> <volume>9</volume>, <fpage>129</fpage>&#x02013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1016/j.edurev.2013.01.002</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Phillip</surname> <given-names>C.</given-names></name></person-group> (<year>2002</year>). <article-title>Clear expectations</article-title>. <source>Knowl. Quest.</source> <volume>31</volume>:<fpage>26</fpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rose</surname> <given-names>L. T.</given-names></name> <name><surname>Fischer</surname> <given-names>K. W.</given-names></name></person-group> (<year>2011</year>). <article-title>Garbage in, garbage out: having useful data is everything</article-title>. <source>Measurement</source> <volume>9</volume>, <fpage>222</fpage>&#x02013;<lpage>226</lpage>. <pub-id pub-id-type="doi">10.1080/15366367.2011.632338</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosenbaum</surname> <given-names>P. R.</given-names></name> <name><surname>Rubin</surname> <given-names>D. B.</given-names></name></person-group> (<year>1983</year>). <article-title>The central role of the propensity score in observational studies for causal effects</article-title>. <source>Biometrika</source> <volume>70</volume>, <fpage>41</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/70.1.41</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schumacher</surname> <given-names>M.</given-names></name> <name><surname>Stansbury</surname> <given-names>K.</given-names></name> <name><surname>Johnson</surname> <given-names>M.</given-names></name> <name><surname>Floyd</surname> <given-names>S.</given-names></name> <name><surname>Reid</surname> <given-names>C.</given-names></name> <name><surname>Noland</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>The young women in science program: a five-year follow-up of an intervention to change science attitudes, academic behavior, and career aspirations</article-title>. <source>J. Women Minor. Sci. Eng.</source> <volume>15</volume>, <fpage>303</fpage>&#x02013;<lpage>321</lpage>. <pub-id pub-id-type="doi">10.1615/JWomenMinorScienEng.v15.i4.20</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Singer</surname> <given-names>S. R.</given-names></name> <name><surname>Nielsen</surname> <given-names>N. R.</given-names></name> <name><surname>Schweingruber</surname> <given-names>H. A.</given-names></name></person-group> (<year>2012</year>). <source>Discipline-Based Education Research: Understanding and Improving Learning in Undergraduate Science and Engineering.</source> <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>The National Academies Press</publisher-name>.</citation></ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>H. W.</given-names></name></person-group> (<year>1975</year>). <source>Strategies of Social Research: the Methodological Imagination</source>. <publisher-loc>Englewood Cliffs, NJ</publisher-loc>: <publisher-name>Prentice-Hall</publisher-name>.</citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Solomon</surname> <given-names>R. L.</given-names></name></person-group> (<year>1949</year>). <article-title>An extension of control group design</article-title>. <source>Psychol. Bull.</source> <volume>46</volume>, <fpage>137</fpage>&#x02013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1037/h0062958</pub-id><pub-id pub-id-type="pmid">18116724</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><collab>The PEAR Institute: Partnerships in Education and Resilience</collab></person-group> (<year>2017</year>). <source>A Guide to PEAR&#x00027;s STEM tools: Dimensions of Success and Common Instrument Suite</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Harvard University</publisher-name>.</citation></ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><collab>U.S. Department of Education</collab></person-group> (<year>2007</year>). <source>Report of the Academic Competitiveness Council</source>. <publisher-loc>Washington, DC</publisher-loc>.</citation></ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><collab>What Works Clearinghouse</collab></person-group> (<year>2008</year>). <source>What Works Clearinghouse Evidence Standards for Reviewing Studies, Version 1.0</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>US Department of Education</publisher-name>.</citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wilkerson</surname> <given-names>S. B.</given-names></name> <name><surname>Haden</surname> <given-names>C. M.</given-names></name></person-group> (<year>2014</year>). <article-title>Effective practices for evaluating STEM out-of-school time programs</article-title>. <source>Aftersch. Matt.</source> <volume>19</volume>, <fpage>10</fpage>&#x02013;<lpage>19</lpage>.</citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Winkleby</surname> <given-names>M. A.</given-names></name></person-group> (<year>2007</year>). <article-title>The stanford medical youth science program: 18 years of a biomedical program for low-income high school students</article-title>. <source>Acad. Med.</source> <volume>82</volume>, <fpage>139</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1097/ACM.0b013e31802d8de6</pub-id><pub-id pub-id-type="pmid">17264691</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Winkleby</surname> <given-names>M. A.</given-names></name> <name><surname>Ned</surname> <given-names>J.</given-names></name> <name><surname>Ahn</surname> <given-names>D.</given-names></name> <name><surname>Koehler</surname> <given-names>A.</given-names></name> <name><surname>Kennedy</surname> <given-names>J. D.</given-names></name></person-group> (<year>2009</year>). <article-title>Increasing diversity in science and health professions: a 21-year longitudinal study documenting college and career success</article-title>. <source>J. Sci. Educ. Technol.</source> <volume>18</volume>, <fpage>535</fpage>&#x02013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1007/s10956-009-9168-0</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yoon</surname> <given-names>K. S.</given-names></name> <name><surname>Duncan</surname> <given-names>T.</given-names></name> <name><surname>Lee</surname> <given-names>S. W.-Y.</given-names></name> <name><surname>Scarloss</surname> <given-names>B.</given-names></name> <name><surname>Shapley</surname> <given-names>K.</given-names></name></person-group> (<year>2007</year>). <source>Reviewing the Evidence on How Teacher Professional Development Affects Student Achievement. Issues and Answers Report, No. 033</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>U.S. Department of Education</publisher-name>.</citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Young</surname> <given-names>J. R.</given-names></name> <name><surname>Ortiz</surname> <given-names>N.</given-names></name> <name><surname>Young</surname> <given-names>J. L.</given-names></name></person-group> (<year>2017</year>). <article-title>STEMulating interest: a meta-analysis of the effects of out-of-school time on student STEM interest</article-title>. <source>Int. J. Educ. Math. Sci. Technol.</source> <volume>5</volume>, <fpage>62</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.18404/ijemst.61149</pub-id></citation>
</ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This work was supported by National Science Foundation grant no. 1710792, Postdoctoral Fellowship in Biology, Broadening Participation of Groups Underrepresented in Biology.</p>
</fn>
</fn-group>
</back>
</article>