<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Educ.</journal-id>
<journal-title>Frontiers in Education</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Educ.</abbrev-journal-title>
<issn pub-type="epub">2504-284X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/feduc.2023.1127644</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Education</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Test engagement and rapid guessing: Evidence from a large-scale state assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Svetina Valdivia</surname>
<given-names>Dubravka</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/314202/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Rutkowski</surname>
<given-names>Leslie</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Rutkowski</surname>
<given-names>David</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Canbolat</surname>
<given-names>Yusuf</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Underhill</surname>
<given-names>Stephanie</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Counseling and Educational Psychology, Indiana University</institution>, <addr-line>Bloomington, IN</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Educational Leadership and Policy Studies, Indiana University</institution>, <addr-line>Bloomington, IN</addr-line>, <country>United States</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Raman Grover, Consultant, Vancouver, Canada</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Chia-Lin Tsai, University of Northern Colorado, United States; Kaiwen Man, University of Alabama, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Dubravka Svetina Valdivia, <email>dsvetina@indiana.edu</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Assessment, Testing and Applied Measurement, a section of the journal Frontiers in Education</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>05</day>
<month>05</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>8</volume>
<elocation-id>1127644</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>03</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Svetina Valdivia, Rutkowski, Rutkowski, Canbolat and Underhill.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Svetina Valdivia, Rutkowski, Rutkowski, Canbolat and Underhill</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>A recent increase in studies related to testing behavior reignited the decades long conversation regarding score validity from assessments that have minimal stakes for students but which may have high stakes for schools or educational systems as a whole. Using data from a large-scale state assessment (with over 80 thousand students per grade), we examined rapid-guessing behavior <italic>via</italic> normative threshold (NT) approaches. We found that the response time effort (RTE) was 0.991 and 0.980 in grade 3 and grade 8, respectively, based on the maximum threshold of 10% (NT10). Similar rates were found based on methods that used 20 and 30%. Percentages of RTEs below 0.90, which indicated meaningful disengagement, were smaller in grade 3 than grade 8 in all normative threshold approaches. Overall, our results suggested that students had high levels of engagement on the assessment, although descriptive differences were found across various demographic subgroups.</p>
</abstract>
<kwd-group>
<kwd>test behavior</kwd>
<kwd>rapid guessing</kwd>
<kwd>standardized assessment</kwd>
<kwd>validity</kwd>
<kwd>effortful responses</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="5"/>
<equation-count count="0"/>
<ref-count count="28"/>
<page-count count="10"/>
<word-count count="7647"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Assessment, Testing and Applied Measurement</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>In a number of testing contexts, consequences for the examinee and other stakeholders are not in alignment. For example, Indiana&#x2019;s (US state) summative assessment for grades 3 through 8, the Indiana Learning Evaluation Assessment Readiness Network (ILEARN), is required of all students at that grade level. Student performance, however, has no impact on the student. Rather, schools and teachers are evaluated based on students&#x2019; aggregated results. The disconnect between assessment stakes for different stakeholders brings into question the validity of the test score use and interpretation. Specifically, as <xref ref-type="bibr" rid="ref16">Soland et al. (2021)</xref> reminded us, a major assumption fundamental to valid use of achievement tests is that &#x201C;examinees are providing maximal effort on the test&#x201D; (p. 1). The authors further note that this assumption is often violated when little is at stake for students (e.g., <xref ref-type="bibr" rid="ref24">Wise and Kong, 2005</xref>; <xref ref-type="bibr" rid="ref01">Rios et al., 2017</xref>; <xref ref-type="bibr" rid="ref4">Jensen et al., 2018</xref>; <xref ref-type="bibr" rid="ref13">Soland, 2018a</xref>,<xref ref-type="bibr" rid="ref14">b</xref>; <xref ref-type="bibr" rid="ref03">Wise and Kuhfeld, 2020</xref>). For these reasons, research on test taking behavior includes, among other things, investigations of students&#x2019; effort and motivation. For example, it has long been noted that on international assessments, US students typically underperform in the areas of mathematics or science, when compared to economically similar educational systems. Some have attributed the difference in performance to school system and cultural differences (e.g., <xref ref-type="bibr" rid="ref17">Stevenson and Stigler, 1994</xref>; <xref ref-type="bibr" rid="ref27">Woessmann, 2016</xref>) or levels of motivation (e.g., <xref ref-type="bibr" rid="ref1">Gneezy et al., 2019</xref>). Namely, in a recent study, <xref ref-type="bibr" rid="ref1">Gneezy et al. (2019)</xref> investigated the difference in effort put forward by students from the US and China to better understand if and what role effort plays on the test itself. In an experiment with high school students from China and the US, the authors found that levels of intrinsic motivation were different between the two groups. The authors recognized that their experiment did not represent the population (as only a handful of high schools were involved), but they have raised an important question of whether the ranking of countries on assessment programs such as the OECD&#x2019;s Programme for International Student Assessment (PISA) reflects not only differences in achievement levels, but also motivation to perform well on the test.</p>
<p>Further, in a comprehensive study, <xref ref-type="bibr" rid="ref8">Rios (2021)</xref> reviewed research on test taking behaviour and effort that spanned four decades, which showed that students&#x2019; low test-taking effort was indeed a serious threat. As Rios and others pointed out, the problem with violation of this assumption is multifold. Researchers have shown that low effort can lead to downward bias in the observed test scores (<xref ref-type="bibr" rid="ref01">Rios et al., 2017</xref>; <xref ref-type="bibr" rid="ref13">Soland, 2018a</xref>,<xref ref-type="bibr" rid="ref14">b</xref>), can affect subgroups differently, furthering the achievement gap biases in estimates, and can affect students who are disengaged from school (e.g., <xref ref-type="bibr" rid="ref13">Soland, 2018a</xref>; <xref ref-type="bibr" rid="ref15">Soland et al., 2019</xref>; <xref ref-type="bibr" rid="ref03">Wise and Kuhfeld, 2020</xref>; <xref ref-type="bibr" rid="ref22">Wise et al., 2021</xref>). One manifestation of low engagement involves rapid response behavior (<xref ref-type="bibr" rid="ref2">Guo and Ercikan, 2020</xref>), where students either quickly guess at answers or tick responses randomly or systematically without expending effort to achieve a correct response. As Guo and Ercikan suggested, rapid response behavior can compromise both the reliability and validity of test score use and interpretation and have negative impacts on estimated performance.</p>
<sec id="sec2">
<title>Aims of the current study</title>
<p>In light of the fact that students might experience low motivation or engagement, that low motivation can take the form of rapid guessing, and that rapid guessing can lead to biased estimates of proficiency, we query the following research questions. In particular, the current paper aims to understand test taking behavior on a large-scale mandatory state assessment in the US by attending to the following questions:</p>
<list list-type="order">
<list-item>
<p>How much rapid guessing is present on the ILEARN assessment?</p>
</list-item>
<list-item>
<p>Does rapid guessing occur in the same amount/rates across policy relevant and typically reported various demographic subgroups?</p>
</list-item>
<list-item>
<p>What is the relationship between effort, accuracy, and proficiency (achievement levels)?</p>
</list-item>
<list-item>
<p>Do normative threshold approaches studied here yield consistent rates of rapid guessing on the ILEARN assessment?</p>
</list-item>
</list>
<p>To answer our research questions, we use census-level assessment data in mathematics and apply a <italic>normative threshold</italic> method to detect rapid guessing. This paper is organized as follows. First, in the background section, we situate our study by discussing the literature around the notions of effort and motivation to understand test taking behavior. We also briefly introduce the ILEARN assessment used in the current study. Next, we describe the data and our analysis plan to attend to the study aims in the methods section, followed by the results. We conclude with a discussion related to the findings, implications, and future directions.</p>
</sec>
</sec>
<sec id="sec3">
<title>Background</title>
<sec id="sec4">
<title>Notions of (low) effort, response time, and rapid guessing to understand test taking behavior</title>
<p>Researchers investigating test taking behavior utilize various methods to study how examinees engage with the assessment. To orient our discussion, we provide definitions of terms such as effortful response, motivation, and rapid guessing by leading scholars in the field on the topic. We note that these definitions are related to some degree to the methods utilized to observe rapid guessing/effortfulness on the assessment. The method we chose, the <italic>normative threshold</italic> method, is no exception. Further, as suggested next, terminology across the studies is somewhat fluid, suggesting that to some extent, definitions used to describe test taking behavior overlap in meaning.</p>
<p><xref ref-type="bibr" rid="ref13">Soland (2018a</xref>,<xref ref-type="bibr" rid="ref14">b)</xref> describes a low effort response to an item as a situation when a student responds to the item faster than a defined minimum response time. <xref ref-type="bibr" rid="ref24">Wise and Kong (2005)</xref> considered rapid guessing behavior as a quick response that does not fully consider the item, which is similar to <xref ref-type="bibr" rid="ref4">Jensen et al.&#x2019;s (2018)</xref> definition of a rapid guess as &#x201C;any item response for which a student responded so rapidly he or she could not have reasonably provided an accurate response, given how long other students of similar proficiency levels took to respond to the same item&#x201D; (p. 268). Rapid guessing has been used in the literature as an indicator of low effort. Hence, effortful responding would suggest that a student attempts to provide a correct response to a test item, while non-effortful responding would assume that a student makes no attempt to respond correctly (e.g., intentional disregard for item content; <xref ref-type="bibr" rid="ref10">Rios and Guo, 2020</xref>). Thus, while low effort has been implicitly understood as a student not trying their best, scholars have connected it to the concepts of rapid guessing and (low test) motivation (e.g., <xref ref-type="bibr" rid="ref15">Soland et al., 2019</xref>).</p>
<p>Technological advances and commensurate shifts from paper-and-pencil to computer-based assessment platforms, including in ILEARN, offer the opportunity to gather test taking process data, including timing, number of actions, and other information. In other words, computerized assessment delivery allows researchers to learn about some test taking behavior, such as rapid guessing, that would typically not be possible on a paper-and-pencil test (e.g., time spent on any item). One methodological approach developed/refined by Wise and colleagues (<xref ref-type="bibr" rid="ref24">Wise and Kong, 2005</xref>; <xref ref-type="bibr" rid="ref25">Wise and Ma, 2012</xref>) with the purpose of measuring student rapid guessing behavior is known as the response time effort (RTE). Through RTE, student test taking effort is captured by examining the duration of a student&#x2019;s individual item response relative to some predetermined threshold, and a student is classified as either using solution behavior (meaning, effortful response) or rapid guessing (non-effortful response). Thus, an item response is flagged as rapid guessing when a student takes less time than the item-specific threshold to respond to an item.</p>
</sec>
<sec id="sec5">
<title>Assessments administration across states, with a focus on <italic>ILEARN</italic></title>
<p>As <xref ref-type="bibr" rid="ref12">Rutkowski et al. (2023)</xref> suggested, in most K-12 settings, the ability to complete a task in a specified amount of time is usually not the construct of interest. While the focus of the current paper is to understand test taking behavior, and not directly on timing, we concur with <xref ref-type="bibr" rid="ref5">Jurich&#x2019;s, 2020</xref> implication that time limits on standardized assessments are more often imposed for practical reasons (cost, logistics and efficiency of test administration) rather than to make an assessment speeded. Nonetheless, states and assessment systems have taken different positions on timing. For example, New York, Indiana, and the Smarter Balanced Assessments, made their standardized assessments untimed in 2016, 2019, and 2020, respectively. Texas and the Partnership for Assessment of Readiness for College and Careers (PARCC), on the other hand, impose time limits on their state assessments. We describe the state standardized assessment for Indiana next, as our study examines student test taking behavior based on the data from this standardized assessment.</p>
</sec>
<sec id="sec6">
<title>ILEARN</title>
<p>ILEARN is a criterion-referenced, summative assessment designed to measure the Indiana academic standards (Indiana Department of Education, 2020). Specifically, ILEARN measures student achievement according to Indiana Academic Standards for Mathematics and English/Language Arts (ELA) for grades three through eight, Science for grades four and six, and Social Studies for grade five. Additionally, students are required to participate in the ILEARN Biology End-of-Course Assessment (ECA) upon completion of the high school biology course to fulfill a federal participation requirement. There is also an optional US Government ECA for students who completed a high school US Government course. None of the ILEARN assessments can be retaken. The test is administered <italic>via</italic> a computer platform (i.e., desktops, laptops, and tablets) and as an item-level computerized adaptive test (CAT) for all but social studies and government, which are fixed format. The assessments are untimed during a four-week test window, and students are allowed to take breaks.</p>
<p>ILEARN contains different types of items (e.g., multiple-choice, matching, short answer, extended response items), which were drawn from licensed item banks including Smarter Balanced, Independent College and Career Ready, and previously used items from older cycles of Indiana assessments. New items were custom developed to align with Indiana educational standards.</p>
<p>Scores on ILEARN reflect statistical estimates of students&#x2019; proficiency/performance (scores are reported at the scale level as well as domain level to indicate students&#x2019; strengths and weaknesses at different content areas). Item response theory (IRT) models are used to calibrate items and derive student scores (<xref ref-type="bibr" rid="ref3">Hall, n.d.</xref>), and scores can be used for multiple purposes. Specifically, ILEARN scores can be used to form instructional strategies to enrich or remediate instruction (<xref ref-type="bibr" rid="ref3">Hall, n.d.</xref>, p. 129), to determine if a student is on track and if they have the skills essential for college-and-career readiness by the time they graduate high school,<xref rid="fn0004" ref-type="fn"><sup>1</sup></xref> or to a smaller degree (and at high-level conclusions), to track progress from year to year (i.e., monitoring student growth).<xref rid="fn0005" ref-type="fn"><sup>2</sup></xref></p>
</sec>
</sec>
<sec id="sec7" sec-type="methods">
<title>Methods</title>
<sec id="sec8">
<title>Data</title>
<p>Data used in the current study came from the mathematics domain of the 2018&#x2013;2019 ILEARN assessment in grades 3 and 8.<xref rid="fn0006" ref-type="fn"><sup>3</sup></xref> From the entire student record dataset, grade 3: <italic>N</italic>&#x2009;=&#x2009;83,095 and grade 8: <italic>N</italic>&#x2009;=&#x2009;83,044. Student responses included in analysis were those with a valid response time for a given item. Specifically, records with a response time =0 were not included, resulting in the exclusion of 228 students in grade 3 and 207 students in grade 8; additionally, students without an overall test status of complete were also excluded (i.e., students with an expired, invalidated, pending, or missing test status). This resulted in the exclusion of 72 students in grade 3 and 228 students in grade 8. The resulting samples used in the analyses included 82,795 students responding across 541<xref rid="fn0007" ref-type="fn"><sup>4</sup></xref> math items in grade 3 and 82,609 students responding across 429 math items in grade 8. Descriptive statistics for the grade 3 and grade 8 samples are presented in <xref rid="tab1" ref-type="table">Table 1</xref>.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Descriptive&#x002A; Statistics for Grades 3 and 8 on Mathematics on ILEARN.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top">Grade 3</th>
<th align="center" valign="top">Grade 8</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Total <italic>N</italic></td>
<td align="center" valign="top">82,795</td>
<td align="center" valign="top">82,609</td>
</tr>
<tr>
<td align="left" valign="top">Mean mathematics achievement (SD)</td>
<td align="center" valign="top">&#x2212;0.84 (1.01)&#x002A;&#x002A;&#x002A;</td>
<td align="center" valign="top">0.68 (1.44)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3">Gender (%)</td>
</tr>
<tr>
<td align="left" valign="top">Girls</td>
<td align="center" valign="top">48.76</td>
<td align="center" valign="top">48.92</td>
</tr>
<tr>
<td align="left" valign="top">Boys</td>
<td align="center" valign="top">51.24</td>
<td align="center" valign="top">51.08</td>
</tr>
<tr>
<td align="left" valign="top">Socioeconomic status&#x002A;&#x002A; (%)</td>
<td align="center" valign="top">47.43</td>
<td align="center" valign="top">52.70</td>
</tr>
<tr>
<td align="left" valign="top">English language learner (%)</td>
<td align="center" valign="top">9.43</td>
<td align="center" valign="top">3.34</td>
</tr>
<tr>
<td align="left" valign="top">Disability status (%)</td>
<td align="center" valign="top">2.22</td>
<td align="center" valign="top">2.69</td>
</tr>
<tr>
<td align="left" valign="top">Special education status (%)</td>
<td align="center" valign="top">16.42</td>
<td align="center" valign="top">14.33</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3">Ethnicity (%)</td>
</tr>
<tr>
<td align="left" valign="top">Asian</td>
<td align="center" valign="top">2.76</td>
<td align="center" valign="top">2.30</td>
</tr>
<tr>
<td align="left" valign="top">Black</td>
<td align="center" valign="top">12.59</td>
<td align="center" valign="top">11.67</td>
</tr>
<tr>
<td align="left" valign="top">Hispanic/Latino/a</td>
<td align="center" valign="top">13.05</td>
<td align="center" valign="top">12.38</td>
</tr>
<tr>
<td align="left" valign="top">Other</td>
<td align="center" valign="top">5.67</td>
<td align="center" valign="top">4.96</td>
</tr>
<tr>
<td align="left" valign="top">White, non-Hispanic</td>
<td align="center" valign="top">65.93</td>
<td align="center" valign="top">68.69</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x002A;Variable names reported here reflect the language and categories used on the assessment. &#x002A;&#x002A;Socioeconomic status variable name was used in the dataset to identify if a student qualified for a free or reduced price lunch (FRPL) or not. % in the table represent the % of students whose status was 1, indicating they qualified for FRPL. &#x002A;&#x002A;&#x002A;8 students in grade 3 and 5 students in grade 8 were excluded from reporting here due to missing value for their theta estimate. The mean mathematics achievement and its associated standard deviation are reported on the IRT scale. In calibration of item responses to obtain IRT theta scores, models are fixed to have a mean of 0 and variance of 1 (which allows us to think/interpret theta values here as standard <italic>z</italic>-scores, commonly used in educational research). From our results, we noted that in grade 3, students had lower average scores, about 0.8 standard deviation below the mean, while in grade 8, students mathematics performance was higher with an average score of approximately 2/3 standard deviation above the mean.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="sec9">
<title>Planned analysis</title>
<p>The normative threshold (NT) method for setting response time thresholds was used to study rapid guessing behavior. As a means of examining sensitivity of findings, we utilized three variants of the NT method (using different thresholds): NT10, NT20, and NT30. We offer next a brief description of the NT methods.</p>
</sec>
<sec id="sec10">
<title>Normative threshold methods</title>
<p>NT exploits item response times to identify rapid guesses. By setting a minimum response time threshold, the method differentiates rapid guessing and solution behavior. NT10 sets the response time threshold as 10% of the average time spent on an item by all students, with a maximum threshold value of 10&#x2009;s. Responses that have a shorter response time than the threshold are identified as rapid guesses while others are identified as solution/effortful behavior (see <xref ref-type="supplementary-material" rid="SM1">Appendix A</xref> for further computation explanation). Based on this identification, test engagement is represented by response time effort (RTE), or stated differently, by the proportion of effortful response. The maximum value of RTE is 1.00, indicating full engagement with the test. In the normative threshold approach, RTEs below 0.90 are defined as meaningful disengagement (<xref ref-type="bibr" rid="ref20">Wise, 2015</xref>; <xref ref-type="bibr" rid="ref23">Wise and Kingsbury, 2016</xref>; <xref ref-type="bibr" rid="ref21">Wise and Gao, 2017</xref>). The NT20 and NT30 approaches use a similar rule, setting the response threshold at 20% and 30%, respectively, of the average time spent on an item by all students, with the 10&#x2009;s rule preserved as in NT10 (<xref ref-type="bibr" rid="ref21">Wise and Gao, 2017</xref>; <xref ref-type="bibr" rid="ref22">Wise et al., 2021</xref>).</p>
</sec>
<sec id="sec11">
<title>Reported analysis</title>
<p>We conducted our analyses separately for each grade. In order to attend to our research aims, we first report results by examining RTE rates across grades, followed by examining results at the subgroup levels for typically reported and often policy relevant subgroups (i.e., those based on demographic variables including gender, free and reduced-price lunch (FRPL), special education status, and race/ethnicity). To understand student behavior more fully on ILEARN and to seek evidence for the validity of the NT method, we investigated the relationship between (low) effort, accuracy, and proficiency levels as well as conducted a sensitivity check across the three NT methods. All analyses were conducted in R (<xref ref-type="bibr" rid="ref7">R Core Team, 2022</xref>) using code written by the authors. Sample R code for the analyses is available upon request.</p>
</sec>
</sec>
<sec id="sec12" sec-type="results">
<title>Results</title>
<p>To attend to our first research question, we computed the RTE rates across grades and methods. Specifically, <xref rid="tab2" ref-type="table">Table 2</xref> reports the overall test engagement statistics using normative threshold approaches. Based on the NT10 approach, the RTEs were 0.991 and 0.980 in grade 3 and grade 8, respectively. NT20 and NT30 approaches revealed similar percentages of effortful responses. The RTEs based on NT20 and NT30 were 0.978 and 0.974 in grade 3. NT20 and NT30 reveal 0.971 and 0.967 RTE in grade 8. The percentage of RTEs below 0.90, which indicates meaningful disengagement, was smaller in grade 3 than grade 8 in all normative threshold approaches. For instance, 2.14% of the students had lower RTE than 0.90 in grade 3, while 5.74% of the students had lower RTE than 0.90 in grade 8. Based on NT10, the percentages of students who had RTE equal to 1, indicating full effortful response, were 86.28% and 76.85% in grades 3 and 8, respectively.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Overall test engagement statistics based on normative threshold (NT) approaches.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="middle" colspan="3">Grade 3</th>
<th align="center" valign="middle" colspan="3">Grade 8</th>
<th align="center" valign="middle">Mean RTE</th>
<th align="center" valign="middle">Percent of RTEs below 0.90</th>
<th align="center" valign="middle">Percent of RTEs&#x2009;=&#x2009;1</th>
<th align="center" valign="middle">Mean RTE</th>
<th align="center" valign="middle">Percent of RTEs below 0.90</th>
<th align="center" valign="middle">Percent of RTEs&#x2009;=&#x2009;1</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">NT10</td>
<td align="char" valign="middle" char=".">0.991</td>
<td align="char" valign="middle" char=".">2.14</td>
<td align="char" valign="middle" char=".">86.28</td>
<td align="char" valign="middle" char=".">0.980</td>
<td align="char" valign="middle" char=".">5.74</td>
<td align="char" valign="middle" char=".">76.85</td>
</tr>
<tr>
<td align="left" valign="middle">NT20</td>
<td align="char" valign="middle" char=".">0.978</td>
<td align="char" valign="middle" char=".">4.59</td>
<td align="char" valign="middle" char=".">62.11</td>
<td align="char" valign="middle" char=".">0.971</td>
<td align="char" valign="middle" char=".">7.83</td>
<td align="char" valign="middle" char=".">63.89</td>
</tr>
<tr>
<td align="left" valign="middle">NT30</td>
<td align="char" valign="middle" char=".">0.974</td>
<td align="char" valign="middle" char=".">5.22</td>
<td align="char" valign="middle" char=".">53.46</td>
<td align="char" valign="middle" char=".">0.967</td>
<td align="char" valign="middle" char=".">8.35</td>
<td align="char" valign="middle" char=".">55.51</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>RTE, response time effort.</p>
</table-wrap-foot>
</table-wrap>
<p><xref rid="tab3" ref-type="table">Table 3</xref> reports subgroup differences in rapid guessing rates for various demographic variables based on NT10 (attending to our second research question). At both grade levels, special education students, English language learners, male students, and relatively low achieving students had lower effortful responses than their peers. Special education students had the smallest effortful response among all subgroups at both grade levels. The mean RTEs for these students were 0.979 and 0.951 in grade 3 and grade 8, respectively. In addition, below proficiency students had lower RTEs than their relatively high achieving peers in both grades. The mean RTE of these students was 0.972 and 0.950 in grades 3 and 8, respectively. Similar results were obtained for NT20 and NT30 (see <xref ref-type="supplementary-material" rid="SM1">Appendix B, Tables B1, B2</xref>). As expected, the mean RTE rates slightly decreased under NT20 and NT30 methods for various subgroups, but the rates remained consistent across the subgroups (e.g., male students yielded lower RTEs than females across grades and methods). The results at subgroup levels were also consistent with results from <xref rid="tab2" ref-type="table">Table 2</xref> that reported at the grade levels across the three methods.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Subgroup (descriptive) differences in rapid guessing rates for various demographic variables (NT10).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th>Variable&#x002A;</th>
<th>Subgroup</th>
<th align="center" valign="middle" colspan="2">Mean RTE</th>
<th align="center" valign="middle">Grade 3</th>
<th align="center" valign="middle">Grade 8</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle" rowspan="2">Gender</td>
<td align="left" valign="middle">Female</td>
<td align="char" valign="middle" char=".">0.993</td>
<td align="char" valign="middle" char=".">0.987</td>
</tr>
<tr>
<td align="left" valign="middle">Male</td>
<td align="char" valign="middle" char=".">0.989</td>
<td align="char" valign="middle" char=".">0.973</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="2">Socioeconomic status</td>
<td align="left" valign="middle">FRLP</td>
<td align="char" valign="middle" char=".">0.995</td>
<td align="char" valign="middle" char=".">0.988</td>
</tr>
<tr>
<td align="left" valign="middle">Non FRLP</td>
<td align="char" valign="middle" char=".">0.988</td>
<td align="char" valign="middle" char=".">0.971</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="2">Special education</td>
<td align="left" valign="middle">Yes</td>
<td align="char" valign="middle" char=".">0.979</td>
<td align="char" valign="middle" char=".">0.951</td>
</tr>
<tr>
<td align="left" valign="middle">No</td>
<td align="char" valign="middle" char=".">0.994</td>
<td align="char" valign="middle" char=".">0.985</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="5">Ethnicity</td>
<td align="left" valign="middle">Asian</td>
<td align="char" valign="middle" char=".">0.994</td>
<td align="char" valign="middle" char=".">0.992</td>
</tr>
<tr>
<td align="left" valign="middle">Black</td>
<td align="char" valign="middle" char=".">0.983</td>
<td align="char" valign="middle" char=".">0.965</td>
</tr>
<tr>
<td align="left" valign="middle">Hispanic</td>
<td align="char" valign="middle" char=".">0.990</td>
<td align="char" valign="middle" char=".">0.979</td>
</tr>
<tr>
<td align="left" valign="middle">Other</td>
<td align="char" valign="middle" char=".">0.989</td>
<td align="char" valign="middle" char=".">0.970</td>
</tr>
<tr>
<td align="left" valign="middle">White</td>
<td align="char" valign="middle" char=".">0.993</td>
<td align="char" valign="middle" char=".">0.983</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="4">Performance level</td>
<td align="left" valign="middle">Below proficiency</td>
<td align="char" valign="middle" char=".">0.972</td>
<td align="char" valign="middle" char=".">0.950</td>
</tr>
<tr>
<td align="left" valign="middle">Approaching proficiency</td>
<td align="char" valign="middle" char=".">0.995</td>
<td align="char" valign="middle" char=".">0.993</td>
</tr>
<tr>
<td align="left" valign="middle">At proficiency</td>
<td align="char" valign="middle" char=".">0.998</td>
<td align="char" valign="middle" char=".">0.997</td>
</tr>
<tr>
<td align="left" valign="middle">Above proficiency</td>
<td align="char" valign="middle" char=".">0.998</td>
<td align="char" valign="middle" char=".">0.998</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x002A;Variable names reported here reflect the language and categories used on the assessment. &#x002A;&#x002A;Socioeconomic status variable name was used in the dataset to identify if a student qualified for a free or reduced price lunch (FRPL) or not.</p>
</table-wrap-foot>
</table-wrap>
<p>Additionally, the percentage of RTE below 0.90 for various groups was further examined based on NT10 (see <xref rid="tab4" ref-type="table">Table 4</xref>; NT20 and NT30 can be found in <xref ref-type="supplementary-material" rid="SM1">Appendix C, Tables C1, C2</xref>). Results were consistent in that RTE below 0.90 rates were higher in grade 8 than grade 3 for all studied subgroups, and in some groups, the rates were quite large. For example, in grade 8, below proficiency students and special education students were identified as the most disengaged studied groups with 15.52% and 15.18% of disengagement, respectively. In grade 3, the highest disengagement rates were associated with the same subgroups, however, at much lower rates of 8.49% and 6.44%, respectively. The groups with the lowest disengagement rates in grades 3 and 8 were those students at or above proficiency levels.</p>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption>
<p>Percent response time effort (RTE) below 0.90 for various demographic variables (NT10).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Variable&#x002A;</th>
<th align="center" valign="top">Subgroup</th>
<th align="center" valign="top">Grade 3</th>
<th align="center" valign="top">Grade 8</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle" rowspan="2">Gender</td>
<td align="left" valign="middle">Female</td>
<td align="char" valign="middle" char=".">1.54</td>
<td align="char" valign="middle" char=".">3.51</td>
</tr>
<tr>
<td align="left" valign="middle">Male</td>
<td align="char" valign="middle" char=".">2.71</td>
<td align="char" valign="middle" char=".">7.88</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="2">Socioeconomic status</td>
<td align="left" valign="middle">FRLP</td>
<td align="char" valign="middle" char=".">0.85</td>
<td align="char" valign="middle" char=".">3.20</td>
</tr>
<tr>
<td align="left" valign="middle">Non FRLP</td>
<td align="char" valign="middle" char=".">3.29</td>
<td align="char" valign="middle" char=".">8.58</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="2">Special education</td>
<td align="left" valign="middle">Yes</td>
<td align="char" valign="middle" char=".">6.44</td>
<td align="char" valign="middle" char=".">15.18</td>
</tr>
<tr>
<td align="left" valign="middle">No</td>
<td align="char" valign="middle" char=".">1.29</td>
<td align="char" valign="middle" char=".">4.16</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="5">Ethnicity</td>
<td align="left" valign="middle">Asian</td>
<td align="char" valign="middle" char=".">0.78</td>
<td align="char" valign="middle" char=".">1.84</td>
</tr>
<tr>
<td align="left" valign="middle">Black</td>
<td align="char" valign="middle" char=".">4.97</td>
<td align="char" valign="middle" char=".">10.68</td>
</tr>
<tr>
<td align="left" valign="middle">Hispanic</td>
<td align="char" valign="middle" char=".">2.58</td>
<td align="char" valign="middle" char=".">6.12</td>
</tr>
<tr>
<td align="left" valign="middle">Other</td>
<td align="char" valign="middle" char=".">3.27</td>
<td align="char" valign="middle" char=".">8.36</td>
</tr>
<tr>
<td align="left" valign="middle">White</td>
<td align="char" valign="middle" char=".">1.46</td>
<td align="char" valign="middle" char=".">4.78</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="4">Performance Level</td>
<td align="left" valign="middle">Below proficiency</td>
<td align="char" valign="middle" char=".">8.49</td>
<td align="char" valign="middle" char=".">15.52</td>
</tr>
<tr>
<td align="left" valign="middle">Approaching proficiency</td>
<td align="char" valign="middle" char=".">0.74</td>
<td align="char" valign="middle" char=".">1.18</td>
</tr>
<tr>
<td align="left" valign="middle">At proficiency</td>
<td align="char" valign="middle" char=".">0.09</td>
<td align="char" valign="middle" char=".">0.19</td>
</tr>
<tr>
<td align="left" valign="middle">Above proficiency</td>
<td align="char" valign="middle" char=".">0.03</td>
<td align="char" valign="middle" char=".">0.18</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x002A;Variable names reported here reflect the language and categories used on the assessment. &#x002A;&#x002A;Socioeconomic status variable name was used in the dataset to identify if a student qualified for a free or reduced price lunch (FRPL) or not.</p>
</table-wrap-foot>
</table-wrap>
<sec id="sec13">
<title>Accuracy by rapid guessing and estimated proficiency levels</title>
<p>To investigate our third research question, we evaluated the validity of the normative threshold approach by examining response accuracy by rapid guessing and proficiency level. It is hypothesized that a rapid response should have a substantively lower accuracy rate than solution (effortful) behavior (<xref ref-type="bibr" rid="ref26">Wise et al., 2019</xref>). Additionally, we posit that the identification of rapid guessing should be independent of examinees&#x2019; proficiency. Therefore, we anticipate that the accuracy rate of effortful response should increase as examinees&#x2019; proficiency increases, while the accuracy rate of rapid response should be similar across proficiency levels. For multiple-choice items, the accuracy rate of rapid guesses should approximate the chance rate &#x2013; in this case, 0.25 as the number of options on ILEARN multiple-choice mathematics items was four.</p>
<p><xref rid="fig1" ref-type="fig">Figure 1</xref> shows response accuracy (i.e., proportion correct response) results by rapid guess versus solution behavior. Namely, we divided proficiency levels into deciles (ten subgroups based on theta level) so that students who scored in approximately lowest 10% are grouped into the first decile, the second set of approximately 10% of students were grouped into the second decile, and so on. By doing so, we separated students by their performance into smaller groups. We then examined how accuracy of the response to an item differed for those students whose response to the item was classified as effortful vs. rapid guess responses. We replicated the same analysis for different normative thresholds across the grades and different levels of proficiency as represented by graphs 1.1 through 1.8 within <xref rid="fig1" ref-type="fig">Figure 1</xref>.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Response accuracy by proficiency (theta) decile and rapid guess.</p>
</caption>
<graphic xlink:href="feduc-08-1127644-g001.tif"/>
</fig>
<p>Specifically, as noted in graphs 1.1 and 1.2, NT10 was relatively robust to detecting rapid guessing by low achieving and moderate achieving examinees in grade 8 when compared to their grade 3 counterparts. Among the students in the seventh or lower proficiency deciles, the accuracy of rapid response was substantially lower than the accuracy of effortful response. Among these student groups in both grades, the accuracy rate of rapid guess response was lower than 0.30 across all proficiency levels, whereas the accuracy of effortful response gradually increased as proficiency increased. However, in the highest three achieving groups, the accuracy rate of rapid guessing was higher than expected, approximating 0.50 in the ninth decile. At the highest proficiency decile (the tenth decile), the accuracy rates were quite similar among rapid guesses and effortful responses.</p>
<p>In grade 3, NT10 revealed a weaker degree of validity evidence than grade 8. The accuracy rate of rapid guess increased as proficiency increased. It was observed that even in the low achieving group, the method failed to precisely identify rapid responses. And, we observed that the accuracy rate of rapid response was quite similar to rates of solution behavior among moderate and high achieving students (dark and light bars were of similar high suggesting similar rates of accuracy). These results suggested that these students responded quicker and more accurately than their peers, and that the model misidentified those responses as a rapid guess.</p>
<p>Graphs 1.3 through 1.6 of <xref rid="fig1" ref-type="fig">Figure 1</xref> present the accuracy rate across proficiency levels and rapid guess based on NT20 and NT30 approaches. The results showed that these approaches had weaker validity than NT10 across both tests. As proficiency increased, the accuracy rate of rapid response increased both in NT20 and NT30 approaches. In most of the proficiency groups, there were no substantial differences in the accuracy rate between rapid guessing and solution behavior, suggesting that these methods failed to precisely identify rapid guessing responses.</p>
<p>Furthermore, graphs 1.7 and 1.8 of <xref rid="fig1" ref-type="fig">Figure 1</xref> illustrate the response accuracy of rapid guessing behavior and solution behavior across proficiency levels only for multiple-choice items, allowing us to explore the extent to which the accuracy of rapid guessing deviated from the chance rate of 0.25. Results indicated that the NT10 approach was relatively more powerful than NT20 and NT30 to identify rapid guessing among low and moderate achieving students, especially in grade 8. While rapid guessing accuracy rates were closer to the chance rate among low achieving and moderate achieving grade 8 students in NT10, their accuracy rates were relatively higher in NT20 and NT30. For instance, the accuracy rate of rapid guessing was not substantially different from the chance rate among below-average students (i.e., those in the fifth or lower proficiency decile) in NT10, while their accuracy rate increased more rapidly across proficiency levels in NT20 and NT30. It is important to note, however, that for above-average students, even NT10 did not identify rapid responses when considering the accuracy rate of rapid guesses. For instance, among the two highest deciles, the accuracy rate of rapid response was higher than 0.50.</p>
<p>Similar to the all-item type comparisons, the normative threshold approaches for multiple-choice items have weaker validity in grade 3 than grade 8. NT10 results indicated that the accuracy rate of rapid responders increased more rapidly as proficiency increased in grade 3 than in grade 8. For instance, in grade 3, the accuracy rate of rapid response was close to the chance rate only in the first three proficiency deciles. In higher proficiency deciles, the accuracy rates were substantially higher than the chance rate of 0.25. Also, the accuracy rates of rapid response and solution behavior were quite similar among those students. NT20 and NT30 had poorer accuracy rates than NT10 in grade 3; except for the first proficiency decile, the accuracy rate of rapid guess was higher than the accuracy rate of solution behavior, suggesting that these approaches had serious limitations for detecting rapid guessing in the test.</p>
</sec>
<sec id="sec14">
<title>Percent of rapid responses across normative threshold approaches</title>
<p>One of the ways to examine the extent to which rapid guessing identification approaches revealed consistent results is to compare the percentage of rapid guess responses across normative threshold approaches. <xref rid="tab5" ref-type="table">Table 5</xref> reports the relationship between the percentage of item responses classified as non-effortful by the normative threshold methods. As observed, the relationship between normative threshold approaches was higher in grade 8 than in grade 3. For instance, the correlation between NT10 and NT20 was 0.909 in grade 8, whereas it was 0.590 in grade 3. Similarly, the relationships between NT10 and NT20 and between NT20 and NT30 were weaker in grade 3 than in grade 8.</p>
<table-wrap position="float" id="tab5">
<label>Table 5</label>
<caption>
<p>Pearson correlations across grades in mathematics for normative thresholds (NT) approaches.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top">NT10</th>
<th align="center" valign="top">NT20</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" colspan="4"><bold>Grade 3 (<italic>N</italic>&#x2009;=&#x2009;539 items)</bold></td>
</tr>
<tr>
<td align="left" valign="top">NT20</td>
<td align="char" valign="top" char=".">0.590</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">NT30</td>
<td align="char" valign="top" char=".">0.444</td>
<td align="char" valign="top" char=".">0.905</td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><bold>Grade 8 (<italic>N</italic>&#x2009;=&#x2009;429 items)</bold></td>
</tr>
<tr>
<td align="left" valign="top">NT20</td>
<td align="char" valign="top" char=".">0.909</td>
<td/>
</tr>
<tr>
<td align="left" valign="top">NT30</td>
<td align="char" valign="top" char=".">0.770</td>
<td align="char" valign="top" char=".">0.929</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>NT10, NT20, and NT30 represent the three variations of normative threshold method.</p>
</table-wrap-foot>
</table-wrap>
<p>Lastly, we visually examined the relationship between percentages of rapid guess responses across rapid guessing approaches (see <xref rid="fig2" ref-type="fig">Figure 2</xref>). The size of the points represents the mean response time of items, meaning that the larger points show longer mean response times. It was observed that consistency across normative approaches was shaped by mean response time across items. If the mean response time was larger than 10&#x2009;s, the normative threshold approaches yielded the same proportion of rapid response across items. This was due to all three normative threshold approaches setting 10&#x2009;s as the upper bound of the threshold. On the other hand, the smaller the mean response, the greater the deviance between the normative threshold approaches. Since smaller mean response times enabled NT20 and NT30 to capture larger proportions of rapid response, their deviances from NT10 were larger.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Percent of item responses classified as non-effortful by normative threshold methods. Note: point size represents the mean response time for an item, with larger points representing longer mean response times.</p>
</caption>
<graphic xlink:href="feduc-08-1127644-g002.tif"/>
</fig>
</sec>
</sec>
<sec id="sec15" sec-type="discussions">
<title>Discussion</title>
<p>Accurately detecting low effort among examinees is an important aspect of making valid interpretations and use of test scores, and the use of various detection methods has been studied in the literature. However, most of the recent studies examining test taking behavior have suffered from limitations related to modest sample sizes, few items (e.g., PISA-based studies), or have used fixed tests (i.e., not computer-adaptive tests; CATs). Only a few studies to our knowledge included adaptive test data with a large number of items and/or examinees, such as those conducted by <xref ref-type="bibr" rid="ref22">Wise et al. (2021)</xref> and <xref ref-type="bibr" rid="ref16">Soland et al. (2021)</xref>. Hence, the motivation of our study was rooted in gaining a better understanding of examinees&#x2019; test taking behavior on a large-scale state-wide standardized assessment where consequences to students are low or indirect, yet important for schools. This importance for schools is rooted in the fact that schools are held accountable by state and local authorities and poor performance on the assessment puts schools at risk for formal sanctions.</p>
<p>The most relevant literature with which to compare our results would be <xref ref-type="bibr" rid="ref22">Wise et al.&#x2019;s (2021)</xref> study. Namely, as in our current study, Wise et al.<xref rid="fn0008" ref-type="fn"><sup>5</sup></xref> examined test taking behavior on a summative assessment for 8th grade examinees. Our results were very consistent with Wise et al., who found the mean RTE rate using NT10 to be 0.979 (compared to our finding of 0.980). Further, the identified percentage of disengaged students was also very similar between the two studies: 5.50% in Wise et al., while we found 5.74% of disengagement. Similarly, approximately 75% of examinees of both of the assessments were deemed to be fully engaged (i.e., % of RTEs&#x2009;=&#x2009;1 were 75.70 and 76.85, respectively).</p>
<p>Despite reasonably high levels of engagement (on average), our results also suggested that for some groups of students, the RTE rates below 0.90 (which some have interpreted as disengagement) were quite a bit higher. While across all studied subgroups (on various demographics variables), the 8th graders were less engaged than were the 3<sup>rd</sup> graders, some subgroups showed quite a big difference in their engagement rates compared to others. Descriptively speaking, under the NT10 method, we observed that, males were less engaged than females (2.71% vs. 1.54% disengagement in grade 3; 7.88% vs. 3.51% in grade 8); those who did not report participating in free/reduced lunch price (FRLP) had RTE rates of 3.29 and 8.58 in grades 3 and 8, respectively (as compared to those in FRLP whose rates were lower at 0.85 and 3.20% in grades 3 and 8, respectively).</p>
<p>Large percentages of disengagement were also observed for those students in special education, in particular in grade 8 where 15.18% of students were classified as disengaged as opposed to 4.16% of their counterparts. In grade 3, across the reported ethnicities, students&#x2019; disengagement ranged from 0.78% (Asian) to 4.97% (Black), while in grade 8, larger ranges (and increases) in disengagements were observed. Specifically, while 1.84% of Asian students disengaged in grade 8 (lowest subgroup in terms of %), 10.68% of Black students reported disengagement (highest subgroup in terms of %). Lastly, when looking at the performance levels of students, those who were classified by their assessments scores as <italic>below proficiency</italic> were substantially less engaged in the assessment (8.49% and 15.52% for grades 3 and 8, respectively) compared to their peers in higher proficiency levels (approaching, at, and above proficiency) who yielded lower percentages of disengagement. We found that minimal disengagement was observed for those at or above proficiency level in either grade. We further noted that for NT20 and NT30 method, results provided similar patterns, although in some cases, the disengagement was even higher than under NT10 (see <xref ref-type="supplementary-material" rid="SM1">Appendix C, Tables C1, C2</xref>). One exception to the patterns between NT10 and other methods was found in the <italic>above proficiency</italic> performance levels which in 3<sup>rd</sup> grade were higher (1.07% and 1.63% for NT20 and NT30, respectively) than the rates for the <italic>at p</italic>roficiency subgroups.</p>
<sec id="sec16">
<title>Implications, strengths, limitations, and future directions</title>
<p>Conversation about the rapid guessing, student engagement, and ways to measure it, is unlikely to go away as long as we continue to assess students in schools. As <xref ref-type="bibr" rid="ref16">Soland et al. (2021)</xref> suggested, different methods to detect low effort have various strengths and weaknesses, and the tradeoffs may lay in the purpose and use of the assessment data, among other things (e.g., is the assessment CAT or fixed format). How should the low effort be treated operationally? One should ask whether or not the scores from students who showed low effort on a prespecified proportion of items be invalidated in order to preserve validity of the scores. Additionally, what is the intended use of the scores? For our current study, the use of ILEARN, as noted above, can be multifaced, and thus, had we found more meaningful low effort or large rates of rapid guessing, questions about inferences and validity of score interpretations would likely need to be weighted even more.</p>
<p>We believe our study holds several strengths, a primary one being the wealth of data at hand. Namely, we utilized a census-level dataset for grades 3 and 8 and had access to item level responses. With it being a computerized assessment, we were also afforded the opportunity to understand test taking behavior at a more nuanced level. While we did discard a few cases (due to missing data on variables of interest, see Methods), missing data rates were negligible.</p>
<p>One limitation of our study is the inclusion of only one subject (mathematics) as it is unknown if the findings would hold for other subject matter (e.g., science). A further limitation lies in our choice of the methods used to study test taking behavior. Specifically, our study employed NT as the method of choice, which is a common approach found in the literature. However, other methods exist, including, for example, the mixture log normal (MLN) method. Future research could triangulate efforts to describe test taking behavior from multiple methods/approaches to better understand how examinees engage with an assessment. However, having said that, we also recognize that some challenges may exist, as approaches have different strengths and limitations. For example, more complex approaches and models can be employed to detect rapid guessing behavior (e.g., <xref ref-type="bibr" rid="ref6">Lu et al., 2020</xref>&#x2019;s mixture model for responses and response times that incorporate hierarchical proficiency structure and information from other subsets on the assessment; or <xref ref-type="bibr" rid="ref19">Ulitzsch et al., 2020</xref>&#x2019;s model that incorporates a hierarchical latent model for joint investigation of engaged and disengaged responses). However, as Soland et al. point out, these models and approaches require large sample sizes at either item or response time levels, which for some contexts (such as in CAT), even with very large sample sizes (such as those reported in Soland, and in our study) may not be achievable.</p>
<p>Our study found behavior on the studied state assessment to be consistent with what has been found in the literature; however, future studies should further examine how test taking behavior manifests across different grades. Our observation of some differences between grades 3 and 8 suggests that, developmentally, students may take a different approach to engaging with test items even if, at the average, patterns of behaviors are similar. To build upon the results of the current study, it would also be beneficial to examine the impact of filtering rapid responses (e.g., <xref ref-type="bibr" rid="ref02">Rios et al., 2014</xref>, <xref ref-type="bibr" rid="ref01">2017</xref>) or incorporating measures of rapid guessing into proficiency estimation (S. L. <xref ref-type="bibr" rid="ref23">Wise and Kingsbury, 2016</xref>) with the aim of investigating the impact of such test taking behavior. As <xref ref-type="bibr" rid="ref9">Rios and Deng (2021)</xref> suggested, in doing so, we assume that rapid guessing can indeed be accurately identified, and this is still an open question. To that end, we also have not differentiated the preknowledge &#x2018;cheating&#x2019; from rapid guessing, which could suggest that once a student has foreknowledge of the item, the response time would also be fast (thus might be flagged as rapid guessing). While this is possible, in our current study, we were not as concerned about the foreknowledge. Reasons for that included the fact that ILEARN was recently revamped and is a state&#x2019;s standardized assessment with the purpose different from some high stakes admission tests, for example. Further, ILEARN was administered in CAT environment, to 3<sup>rd</sup> and 8th graders populations, and so taken altogether, we did not expect a large amount of preknowledge cheating occurring. However, as with any assessment, in particular those deemed to be high stakes, a potential issue of preknowledge is certainly present and ought to be considered in understanding student test taking behavior.</p>
<p>While our strengths were to examine test taking behavior using census-level data, the results of which told a consistent story, an open question remains whether these results would hold post pandemic. In other words, given the large disruption in education over the last two years due to COVID-19, it is not known if students&#x2019; engagement on assessments such as ILEARN would remain as high as found in the current study. Finally, future researchers should also examine whether item order (and other item-level characteristics) influences how test takers engage with items. Providing more nuanced understanding of test taking behavior can help strengthen our claims for valid test score use and interpretation.</p>
</sec>
</sec>
<sec id="sec17" sec-type="data-availability">
<title>Data availability statement</title>
<p>The data analyzed in this study is subject to the following licenses/restrictions: Data used in the study came from the State Department of Education in Indiana. Authors applied for restricted access to the data and obtained permission by the Institutional Review Board to conduct secondary analyses on the restricted data. Requests to access these datasets should be directed to <ext-link xlink:href="https://www.in.gov/doe/" ext-link-type="uri">https://www.in.gov/doe/</ext-link>.</p>
</sec>
<sec id="sec18">
<title>Author contributions</title>
<p>DSV led project conceptualization, methodology, writing, original draft preparation, and supervision. LR and DR contributed to conceptualization, review, and editing of the manuscript. YC performed the analyses and contributed to writing parts of the manuscript. SU contributed to reviewing and editing the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec19" sec-type="funding-information">
<title>Funding</title>
<p>This work was partially supported by two grants to the first author: the Maris M. Proffitt and Mary Higgins Proffitt Endowment Grant, Indiana University, and Indiana University Institute for Advanced Study, Indiana University&#x2014;Bloomington, IN.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="sec21" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/feduc.2023.1127644/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/feduc.2023.1127644/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.DOCX" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gneezy</surname> <given-names>U.</given-names></name> <name><surname>List</surname> <given-names>J. A.</given-names></name> <name><surname>Livingston</surname> <given-names>J. A.</given-names></name> <name><surname>Qin</surname> <given-names>X.</given-names></name> <name><surname>Sadoff</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Measuring success in education: the role of effort on the test itself</article-title>. <source>Am. Econ. Rev.</source> <volume>1</volume>, <fpage>291</fpage>&#x2013;<lpage>308</lpage>. doi: <pub-id pub-id-type="doi">10.1257/aeri.20180633</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Ercikan</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Differential rapid responding across language and cultural groups</article-title>. <source>Educ. Res. Eval.</source> <volume>26</volume>, <fpage>302</fpage>&#x2013;<lpage>327</lpage>. doi: <pub-id pub-id-type="doi">10.1080/13803611.2021.1963941</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hall</surname> <given-names>M.</given-names></name></person-group> (<year>n.d.</year>). <article-title>Volume 1 annual technical report</article-title>. <source>Technical Report</source>, <volume>1</volume>, <fpage>378</fpage>.</citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jensen</surname> <given-names>N.</given-names></name> <name><surname>Rice</surname> <given-names>A.</given-names></name> <name><surname>Soland</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>The influence of rapidly guessed item responses on teacher value-added estimates: implications for policy and practice</article-title>. <source>Educ. Eval. Policy Anal.</source> <volume>40</volume>, <fpage>267</fpage>&#x2013;<lpage>284</lpage>. doi: <pub-id pub-id-type="doi">10.3102/0162373718759600</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Jurich</surname> <given-names>D. P.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>A history of speededness: tracing the evolution of theory and practice</article-title>&#x201D; in <source>Integrating Timing Considerations to Improve Testing Practices</source>. eds. <person-group person-group-type="editor"><name><surname>Margolis</surname> <given-names>M. A.</given-names></name> <name><surname>Feinberg</surname> <given-names>R. A.</given-names></name></person-group> (<publisher-name>Routledge</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>118</lpage>.</citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Tao</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>A mixture model for responses and response times with a higher-order ability structure to detect rapid guessing behaviour</article-title>. <source>Br. J. Math. Stat. Psychol.</source> <volume>73</volume>, <fpage>261</fpage>&#x2013;<lpage>288</lpage>. doi: <pub-id pub-id-type="doi">10.1111/bmsp.12175</pub-id>, PMID: <pub-id pub-id-type="pmid">31385609</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="book"><person-group person-group-type="author"><collab id="coll1">R Core Team</collab></person-group> (<year>2022</year>). <article-title>R: a language and environment for statistical computing</article-title>. <source>R Foundation for Statistical Computing</source>, <publisher-loc>Vienna, Austria</publisher-loc>. <ext-link xlink:href="https://www.R-project.org/" ext-link-type="uri">https://www.R-project.org/</ext-link>.</citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rios</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Improving test-taking effort in low-stakes group-based educational testing: a meta-analysis of interventions</article-title>. <source>Appl. Meas. Educ.</source> <volume>34</volume>, <fpage>85</fpage>&#x2013;<lpage>106</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2021.1890741</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rios</surname> <given-names>J. A.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Does the choice of response time threshold procedure substantially affect inferences concerning the identification and exclusion of rapid guessing responses? A meta-analysis</article-title>. <source>Large-scale Assess. Educ.</source> <volume>9</volume>, <fpage>1</fpage>&#x2013;<lpage>25</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s40536-021-00110-8</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rios</surname> <given-names>J. A.</given-names></name> <name><surname>Guo</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Can culture be a salient predictor of test-taking engagement? An analysis of differential noneffortful responding on an international college-level assessment of critical thinking</article-title>. <source>Appl. Meas. Educ.</source> <volume>33</volume>, <fpage>263</fpage>&#x2013;<lpage>279</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2020.1789141</pub-id></citation></ref>
<ref id="ref01"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rios</surname> <given-names>J. A.</given-names></name> <name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Mao</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>O. L.</given-names></name></person-group> (<year>2017</year>). <article-title>Evaluating the impact of careless responding on aggregated-scores: To filter unmotivated examinees or not?</article-title> <source>Int. J. Test.</source> <volume>17</volume>, <fpage>74</fpage>&#x2013;<lpage>104</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305058.2016.1231193</pub-id></citation></ref>
<ref id="ref02"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rios</surname> <given-names>J. A.</given-names></name> <name><surname>Liu</surname> <given-names>O. L.</given-names></name> <name><surname>Bridgeman</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>Identifying low-effort examinees on student learning outcomes assessment: A comparison of two approaches: Identifying low-effort examinees on student learning outcomes assessment</article-title>. <source>New Dir. Inst. Res.</source> <volume>2014</volume>, <fpage>69</fpage>&#x2013;<lpage>82</lpage>. doi: <pub-id pub-id-type="doi">10.1002/ir.20068</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rutkowski</surname> <given-names>D.</given-names></name> <name><surname>Rutkowski</surname> <given-names>L.</given-names></name> <name><surname>Valdivia</surname> <given-names>D.</given-names></name> <name><surname>Canbolat</surname> <given-names>Y.</given-names></name> <name><surname>Underhill</surname> <given-names>S.</given-names></name></person-group> (<year>2023</year>). <article-title>A census-level, multi-grade analysis of the association between testing time, breaks, and achievement</article-title>. <source>Appl. Meas. Educ.</source> <volume>36</volume>, <fpage>14</fpage>&#x2013;<lpage>30</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2023.2172019</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soland</surname> <given-names>J.</given-names></name></person-group> (<year>2018a</year>). <article-title>The achievement gap or the engagement gap? Investigating the sensitivity of gaps estimates to test motivation</article-title>. <source>Appl. Meas. Educ.</source> <volume>31</volume>, <fpage>312</fpage>&#x2013;<lpage>323</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2018.1495213</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soland</surname> <given-names>J.</given-names></name></person-group> (<year>2018b</year>). <article-title>Are achievement gap estimates biased by differential student test effort? Putting an important policy metric to the test</article-title>. <source>Teach. Coll. Rec.</source> <volume>120</volume>, <fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi: <pub-id pub-id-type="doi">10.1177/016146811812001202</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soland</surname> <given-names>J.</given-names></name> <name><surname>Jensen</surname> <given-names>N.</given-names></name> <name><surname>Keys</surname> <given-names>T. D.</given-names></name> <name><surname>Bi</surname> <given-names>S. Z.</given-names></name> <name><surname>Wolk</surname> <given-names>E.</given-names></name></person-group> (<year>2019</year>). <article-title>Are test and academic disengagement related? Implications for measurement and practice</article-title>. <source>Educ. Assess.</source> <volume>24</volume>, <fpage>119</fpage>&#x2013;<lpage>134</lpage>. doi: <pub-id pub-id-type="doi">10.1080/10627197.2019.1575723</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soland</surname> <given-names>J.</given-names></name> <name><surname>Kuhfeld</surname> <given-names>M.</given-names></name> <name><surname>Rios</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Comparing different response time threshold setting methods to detect low effort on a large-scale assessment</article-title>. <source>Large-scale Assess. Educ.</source> <volume>9</volume>, <fpage>1</fpage>&#x2013;<lpage>21</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s40536-021-00100-w</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Stevenson</surname> <given-names>H.</given-names></name> <name><surname>Stigler</surname> <given-names>J. W.</given-names></name></person-group> (<year>1994</year>). <source>Learning Gap: Why Our Schools Are Failing and What We Can Learn From Japanese and Chinese Educ</source> <publisher-name>Simon and Schuster</publisher-name>.</citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ulitzsch</surname> <given-names>E.</given-names></name> <name><surname>von Davier</surname> <given-names>M.</given-names></name> <name><surname>Pohl</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>A hierarchical latent response model for inferences about examinee engagement in terms of guessing and item-level non-response</article-title>. <source>Br. J. Math. Stat. Psychol.</source> <volume>73</volume>, <fpage>83</fpage>&#x2013;<lpage>112</lpage>. doi: <pub-id pub-id-type="doi">10.1111/bmsp.12188</pub-id>, PMID: <pub-id pub-id-type="pmid">31709521</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name></person-group> (<year>2015</year>). <article-title>Effort analysis: individual score validation of achievement test data</article-title>. <source>Appl. Meas. Educ.</source> <volume>28</volume>, <fpage>237</fpage>&#x2013;<lpage>252</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2015.1042155</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name> <name><surname>Gao</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>A general approach to measuring test taking effort on computer-based tests</article-title>. <source>Appl. Meas. Educ.</source> <volume>30</volume>, <fpage>343</fpage>&#x2013;<lpage>354</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2017.1353992</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name> <name><surname>Im</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>The impact of disengaged test taking on a state&#x2019;s accountability test results</article-title>. <source>Educ. Assess.</source> <volume>26</volume>, <fpage>163</fpage>&#x2013;<lpage>174</lpage>. doi: <pub-id pub-id-type="doi">10.1080/10627197.2021.1956897</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name> <name><surname>Kingsbury</surname> <given-names>G. G.</given-names></name></person-group> (<year>2016</year>). <article-title>Modeling student test-taking motivation in the context of an adaptive achievement test</article-title>. <source>J. Educ. Meas.</source> <volume>53</volume>, <fpage>86</fpage>&#x2013;<lpage>105</lpage>. doi: <pub-id pub-id-type="doi">10.1111/jedm.12102</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name> <name><surname>Kong</surname> <given-names>X.</given-names></name></person-group> (<year>2005</year>). <article-title>Response time effort: a new measure of examinee motivation in computer-based tests</article-title>. <source>Appl. Meas. Educ.</source> <volume>18</volume>, <fpage>163</fpage>&#x2013;<lpage>183</lpage>. doi: <pub-id pub-id-type="doi">10.1207/s15324818ame1802_2</pub-id></citation></ref>
<ref id="ref03"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name> <name><surname>Kuhfeld</surname> <given-names>M. R.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>A cessation of measurement: Identifying test taker disengagement using response time</article-title>,&#x201D; in <source>Integrating Timing Considerations to Improve Testing Practices</source>. <publisher-name>Routledge</publisher-name>.</citation></ref>
<ref id="ref25"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>L.</given-names></name></person-group> (<year>2012</year>). <article-title>Setting response time thresholds for a CAT item pool: the normative threshold method</article-title>. <conf-name>Paper Presented at the Annual Meeting of the National Council on Measurement in Education</conf-name>, <conf-loc>Vancouver, Canada</conf-loc>.</citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wise</surname> <given-names>S. L.</given-names></name> <name><surname>Soland</surname> <given-names>J.</given-names></name> <name><surname>Bo</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>The (non)impact of differential test taker engagement on aggregated scores</article-title>. <source>Int. J. Test.</source> <volume>20</volume>, <fpage>57</fpage>&#x2013;<lpage>77</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305058.2019.1605999</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Woessmann</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>The importance of school systems: evidence from international differences in student achievement</article-title>. <source>J. Econ. Perspect.</source> <volume>30</volume>, <fpage>3</fpage>&#x2013;<lpage>32</lpage>. doi: <pub-id pub-id-type="doi">10.1257/jep.30.3.3</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn0004"><p><sup>1</sup>The <italic>being college ready</italic> indicator on ILEARN is connected to the performance level descriptors on the assessment such that students who achieve "At Proficiency" or "Above Proficiency" would be indicated as on track for being college ready. Students who received "Below Proficiency" or "Approaching Proficiency" would not be considered on track for college and career readiness based on their ILEARN results.</p></fn><fn id="fn0005"><p><sup>2</sup>Indiana uses student growth percentiles to measure growth more precisely.</p></fn>
<fn id="fn0006"><p><sup>3</sup>Institutional Review Board protocol to use data was filed and approved by the authors&#x2019; institution. Protocol type was not human subjects research because we used already collected and deidentified data.</p></fn>
<fn id="fn0007"><p><sup>4</sup>The starting numbers of math items were 551 and 454, respectively, but 10 and 25 items were removed from the analyses for being on a shared page (and, thus, not having a unique item-level response time).</p></fn>
<fn id="fn0008"><p><sup>5</sup>Wise et al. also examined English language arts and science as well as compared their results on summative assessment with another assessment, namely the MAP Growth. In our discussion, we focus only on direct possible comparisons between our respective results (i.e., grade 8 math).</p></fn>
</fn-group>
</back>
</article>