<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2024.1489054</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Confidence in mathematics is confounded by responses to reverse-coded items</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Antoniou</surname> <given-names>Faye</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1575723/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Alghamdi</surname> <given-names>Mohammed H.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1904281/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Educational Studies, National and Kapodistrian University of Athens</institution>, <addr-line>Athens</addr-line>, <country>Greece</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Self-Development Skills, King Saud University</institution>, <addr-line>Riyadh</addr-line>, <country>Saudi Arabia</country></aff>
<author-notes>
<fn id="fn0003" fn-type="edited-by"><p>Edited by: Dimitrios Stamovlasis, Aristotle University of Thessaloniki, Greece</p></fn>
<fn id="fn0004" fn-type="edited-by"><p>Reviewed by: Kosuke Kawai, University of California, Los Angeles, United States</p>
<p>Bo Zhang, Boston Children&#x2019;s Hospital, Harvard Medical School, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Faye Antoniou, <email>fayeantoniou@gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>24</day>
<month>10</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>15</volume>
<elocation-id>1489054</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>08</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>10</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2024 Antoniou and Alghamdi.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Antoniou and Alghamdi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>This study investigates the confounding effects of reverse-coded items on the measurement of confidence in mathematics using data from the 2019 Trends in International Mathematics and Science Study (TIMSS).</p>
</sec>
<sec>
<title>Methods</title>
<p>The sample came from the Saudi Arabian cohort of 8th graders in 2019 involving 4,515 students. Through mixture modeling, two subgroups responding in similar ways to reverse-coded items were identified representing approximately 9% of the sample.</p>
</sec>
<sec>
<title>Results</title>
<p>Their response to positively valenced and negatively valenced items showed inconsistency and the observed unexpected response patterns were further verified using Lz&#x002A;, U3, and the number of Guttman errors person fit indicators. Psychometric analyses on the full sample and the truncated sample after deleting the aberrant responders indicated significant improvements in both internal consistency reliability and factorial validity.</p>
</sec>
<sec>
<title>Discussion</title>
<p>It was concluded that reverse-coded items contribute to systematic measurement error that is associated with distorted item level parameters that compromised the scale&#x2019;s reliability and validity. The study underscores the need for reconsideration of reverse-coded items in survey design, particularly in contexts involving younger populations and low-achieving students.</p>
</sec>
</abstract>
<kwd-group>
<kwd>reverse coded items</kwd>
<kwd>aberrant responding</kwd>
<kwd>carelessness</kwd>
<kwd>random responding</kwd>
<kwd>person-fit indices</kwd>
<kwd>lz&#x002A;</kwd>
<kwd>Guttman errors</kwd>
<kwd>U3</kwd>
</kwd-group>
<contract-num rid="cn1">RSPD2024R601</contract-num>
<contract-sponsor id="cn1">King Saud University<named-content content-type="fundref-id">10.13039/501100002383</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="0"/>
<equation-count count="3"/>
<ref-count count="58"/>
<page-count count="10"/>
<word-count count="6281"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Quantitative Psychology and Measurement</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="sec1">
<label>1</label>
<title>What are the effects of reverse-coded items?</title>
<p>Several authors have pointed to the detrimental effects of reverse-coded items on the quality of the collected data (<xref ref-type="bibr" rid="ref39">Roszkowski and Soven, 2010</xref>; <xref ref-type="bibr" rid="ref53">Weems and Onwuegbuzie, 2001</xref>). For example, <xref ref-type="bibr" rid="ref6">Clauss and Bardeen (2020)</xref> reported that the presence of negatively worded items confounded the conclusions on the simple structure of the Attentional Control Scale (ACS), which, although should be unidimensional it was found to have a bifactor structure (see also <xref ref-type="bibr" rid="ref32">Pedersen et al., 2024</xref>; <xref ref-type="bibr" rid="ref36">Ponce et al., 2022</xref>). They concluded that reversing the negatively worded items resulted in producing a method factor rather than an attentional control factor, thus, complicating the conceptual factor structure, and producing an incoherent structure with unexplained item relations (<xref ref-type="bibr" rid="ref30">Merritt, 2012</xref>; <xref ref-type="bibr" rid="ref32">Pedersen et al., 2024</xref>). Concerns about internal consistency reliability and content differentiation have also been raised (<xref ref-type="bibr" rid="ref39">Roszkowski and Soven, 2010</xref>; <xref ref-type="bibr" rid="ref53">Weems and Onwuegbuzie, 2001</xref>). For example, <xref ref-type="bibr" rid="ref15">Jaensson and Nilsson (2017)</xref> reported that reversing negatively worded to positively worded items with the same content resulted in a poor to moderate agreement between the two, as was evident from intraclass correlation coefficients between 0.35 and 0.76. Thus, the obtained responses to the same content, using a different item format, resulted in discrepant responses, raising concerns about item scoring and item interpretation. <xref ref-type="bibr" rid="ref5">Bolt et al. (2020)</xref> reported that reverse-coded items may have deleterious effects on item comprehension. They reported that rates of confusion and misunderstanding of negatively worded items were as high as 53% for students in grades 3 through 5, representing a salient concern for these students as issues of identification, placement, and early intervention are most important at an early age. They further stated that skills and competencies were significantly underestimated in the presence of opposite format items. Similar distortions have been found with adolescents and adults such as teacher and student populations (<xref ref-type="bibr" rid="ref2">Barnette, 1996</xref>), although non-significant differences across age groups have also been reported (<xref ref-type="bibr" rid="ref44">Steinmann et al., 2024</xref>).</p>
<p>Interpretations as to the &#x201C;why&#x2019;s&#x201D; of differential responses to the same content as a function of item format have been traced to Personality traits such as agreeableness (<xref ref-type="bibr" rid="ref39">Roszkowski and Soven, 2010</xref>) or neuroticism (<xref ref-type="bibr" rid="ref22">Koutsogiorgi and Michaelides, 2022</xref>), the lack of motivation (<xref ref-type="bibr" rid="ref16">Kam, 2018</xref>), the presence of response styles such as acquiescence (<xref ref-type="bibr" rid="ref9">DiStefano and Motl, 2009</xref>), situational factors such as carelessness and inattention (<xref ref-type="bibr" rid="ref3">Baumgartner et al., 2018</xref>; <xref ref-type="bibr" rid="ref45">Steinmann et al., 2022</xref>; <xref ref-type="bibr" rid="ref47">Swain et al., 2008</xref>), differential interpretation of item content (<xref ref-type="bibr" rid="ref53">Weems and Onwuegbuzie, 2001</xref>), low achievement (<xref ref-type="bibr" rid="ref44">Steinmann et al., 2024</xref>; <xref ref-type="bibr" rid="ref54">Weems et al., 2003</xref>), mood (<xref ref-type="bibr" rid="ref22">Koutsogiorgi and Michaelides, 2022</xref>), and demographic variables such as gender with boys having higher levels of inconsistent responses (<xref ref-type="bibr" rid="ref27">Marsh and Grayson, 1995</xref>; <xref ref-type="bibr" rid="ref43">Steedle et al., 2019</xref>; <xref ref-type="bibr" rid="ref44">Steinmann et al., 2024</xref>). However, potential causes are beyond the scope of the present study and will not be discussed in more detail.</p>
<sec id="sec2">
<label>1.1</label>
<title>What are current recommendations for dealing with reverse-coded items?</title>
<p>At the analytical level, recommendations to deal with reverse-coded items include fitting multidimensional models such as the bifactor model to account for method variance likely attributed to the item format for negatively worded items. Identifying and deleting individuals who behave in unexpected ways as they provide invalid estimates for themselves and harm the psychometric qualities of the measured instrument (<xref ref-type="bibr" rid="ref31">Michaelides, 2019</xref>). Identifying sources of confusion for reverse-coded items and treating those items through refinement and revision to isolate the sources of confusion has also been recommended, especially for younger age groups (<xref ref-type="bibr" rid="ref5">Bolt et al., 2020</xref>; <xref ref-type="bibr" rid="ref12">Fukudome and Takeda, 2024</xref>; <xref ref-type="bibr" rid="ref44">Steinmann et al., 2024</xref>).</p>
</sec>
<sec id="sec3">
<label>1.2</label>
<title>Context and goals of the present study</title>
<p>We selected the examination of confidence ratings for students in Saudi Arabia using the Trends of International Mathematics Assessment (<xref ref-type="bibr" rid="ref9001">Mullis et al., 2020</xref>) international study for several reasons. First, Saudi students are classified among the lowest in mathematics achievement across TIMSS&#x2019;s participating countries, thus, examination of whether their assessments involve deficits in the psychological sphere of confidence is important. Second, recent data from PISA 2022 indicated that Saudi students had the highest ratings of &#x201C;straightlining&#x201D; responses in the survey instruments such as assertiveness, cooperation, etc. Thus, examination of aberrant responses in this population is important, because already approximately 4.5 to 5% of the participant responses are currently screened for and deleted from international databases due to straightlining. Thus, the goals of the present study were (a) to identify subgroups of participants who respond in the same manner in reverse-coded items using mixture modeling, (b) to validate the identification of aberrance mixtures using person fit indicators, (c) to test the presence of systematic error variance through supporting an additional &#x201C;method&#x201D; factor, and (d) evaluate the effects of aberrant responders on the psychometrics of a confidence scale related to mathematics achievement contrasting original and purified data (after deleting aberrant responders).</p>
</sec>
</sec>
<sec sec-type="methods" id="sec4">
<label>2</label>
<title>Methods</title>
<sec id="sec5">
<label>2.1</label>
<title>Participants and procedures</title>
<p>Participants were 5,680 Saudi 8<sup>th</sup>-grade students who took part in the 2019 TIMSS study. There were 2,884 males (50.8%) and 2,791 females (49.2%). Most students (91.7%) attended public schools with a small percentage (8.3%) attending international schools. The mean age was 13.926&#x2009;years (SD&#x2009;=&#x2009;0.679). In the TIMSS 2019 study, students in each country are selected for participation using multistage stratified random sampling to achieve representativeness to the population and specifically the characteristics of national student populations regarding geographic regions and school types. Sample sizes are more than 4,000 students and sampling engages at least 150 schools. The TIMSS guidelines also state clear guidelines to ensure high participation during the testing process, and to avoid biased samples that threaten generalization of the findings to the population. More details on the study and its methodology can be traced.<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref><sup>,</sup><xref ref-type="fn" rid="fn0002"><sup>2</sup></xref></p>
</sec>
<sec id="sec6">
<label>2.2</label>
<title>Measure</title>
<p>The &#x201C;Students Confidence in Mathematics scale&#x201D; was implemented which comprises nine items. Example items that are positively worded were &#x201C;I usually do well in mathematics,&#x201D; &#x201C;I learn things quickly in mathematics,&#x201D; and &#x201C;I am good at working out difficult mathematics problems.&#x201D; Sample negatively worded items were &#x201C;Mathematics makes me nervous&#x201D; and &#x201C;Mathematics makes me confused. Respondents indicated their agreement with each statement on a four-point Likert scale ranging from &#x201C;Agree a lot&#x201D; to &#x201C;Disagree a lot&#x201D; with no midpoint option. Higher scores on the confidence scale indicate greater confidence in mathematics. Internal consistency reliability was assessed using Cronbach&#x2019;s alpha coefficient and was 0.81. Based on TIMSS 2019, the scale is unidimensional, and scale scores are provided per country along with cutoff scores. Our country-based analysis using the Graded Response Model (GRM) showed marginal reliability estimates equal to 0.86 and adequate omnibus model fit via the Root Mean Squared Error of Approximation (RMSEA) that was equal to 0.08, after reversing the items that have the opposite meaning to reflect positive covariances across all items and post purification (i.e., after deleting aberrant responders). <xref ref-type="fig" rid="fig1">Figure 1</xref> displays the Test Information Function (TIF) and corresponding Conditional Standard Error of Measurement (CSEM) of the scale which shows a nice coverage of information across &#x00B1;2.5 theta scores and a center around the mean of zero as expected. Further analyses of category information curves are shown in the <xref ref-type="app" rid="app1">Appendix</xref> which support the used scaling system with no overlap or disordering.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption><p>Total information curve and conditional standard error of measurement for the mathematics confidence scale.</p></caption>
<graphic xlink:href="fpsyg-15-1489054-g001.tif"/>
</fig>
</sec>
<sec id="sec7">
<label>2.3</label>
<title>Data analyses</title>
<sec id="sec8">
<label>2.3.1</label>
<title>Criteria for classifying participants to groups</title>
<p>The following three criteria were utilized to evaluate model fit in the classification process. Lower values are indicative of better model fit. Although entropy refers to the estimation of the indices below, in a standalone form, entropy is not included in the process of concluding the most optimal latent class as earlier suggested (<xref ref-type="bibr" rid="ref9111">Masyn, 2013</xref>).</p>
<p>The Classification Likelihood Criterion (CLC) is a measure of model fit that considers both the log-likelihood of the model and the entropy of the classification. Entropy reflects the uncertainty of the in-class assignments, with higher entropy indicating more uncertainty (i.e., lower classification certainty). The goal of CLC is to find a balance between model fit (log-likelihood) and classification certainty (entropy). CLC is calculated as shown in <xref ref-type="disp-formula" rid="E1">Equation (1)</xref> below:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">C</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mtext>Entropy</mml:mtext></mml:mrow></mml:mfenced></mml:math></disp-formula>
<p>With LL being the log-likelihood of the model, and Entropy reflecting uncertainty in the classification of participants to subgroups.</p>
<p>The Akaike Weight of Evidence (AWE) is a model selection criterion that combines information about model fit (log-likelihood) and penalizes model complexity more strongly than the traditional Akaike Information Criterion (AIC). It favors simpler models unless the additional layers of complexity are justified based on omnibus model fit indicators. It is estimated as shown in <xref ref-type="disp-formula" rid="EQ2">Equation (2)</xref> below:</p>
<disp-formula id="EQ2"><label>(2)</label><mml:math id="M2"><mml:mtext>AWE</mml:mtext><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo></mml:msup><mml:msup><mml:mi>m</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mi>log</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>log</mml:mi><mml:mfenced open="(" close=")"><mml:mi>n</mml:mi></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mtext>Entropy</mml:mtext></mml:math></disp-formula>
<p>With LL being the log-likelihood of the model, <italic>m</italic> is the number of estimated parameters, and <italic>n is</italic> the sample size.</p>
<p>The Integrated Classification Likelihood-Bayesian Information Criterion (ICL-BIC) is a variant of the Bayesian Information Criterion (BIC) that also accounts for the quality of classification. Thus, the ICL-BIC criterion aims to select models that not only have a good fit (as indicated by the BIC) but also provide clear and distinct class memberships (as indicated by entropy). It is defined as shown in <xref ref-type="disp-formula" rid="EQ3">Equation (3)</xref> below:</p>
<disp-formula id="EQ3"><label>(3)</label><mml:math id="M3"><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">B</mml:mi><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">C</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">B</mml:mi><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">C</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mtext>Entropy</mml:mtext></mml:math></disp-formula>
<p>With BIC being the Bayesian Information Criterion, calculated as BIC&#x2009;=&#x2009;&#x2212;2&#x002A;LL&#x2009;+&#x2009;<italic>m</italic>&#x002A;log(<italic>n</italic>), and Entropy the classification uncertainty.</p>
</sec>
<sec id="sec9">
<label>2.3.2</label>
<title>Validating person aberrant behavior via analyzing response patterns</title>
<p>Three of the most prominent person fit indicators namely, U3, Guttman errors, and lz&#x002A; were selected to validate the results from the latent class analyses (<xref ref-type="bibr" rid="ref11">Emons, 2008</xref>; <xref ref-type="bibr" rid="ref8">Cui and Mousavi, 2015</xref>). Each of these indices has its advantages and limitations or is sensitive to specific patterns of aberrance. The U3 statistic provides a nuanced assessment of aberrant responding, being sensitive to subtle deviations from the Guttman pattern (<xref ref-type="bibr" rid="ref40">Schroeders et al., 2021</xref>). It is most efficacious in detecting inattention. Its disadvantage, however, is its sensitivity to test length, with brief measures jeopardizing its reliability. The number of Guttman errors offers a traditional measure of misfit by examining the type of response as a function of item difficulty (<xref ref-type="bibr" rid="ref29">Meijer, 1994</xref>; <xref ref-type="bibr" rid="ref49">Tendeiro and Meijer, 2014</xref>). It is easy to understand but it may not be sensitive to more subtle forms of misfit. The lz&#x002A; statistic is a powerful tool for detecting misfits in response patterns and contributes to a comprehensive analysis of response data. The lz&#x002A; statistic has been found effective in detecting various types of aberrant responding, such as fake good and random responding (<xref ref-type="bibr" rid="ref1">Av&#x015F;ar, 2022</xref>; <xref ref-type="bibr" rid="ref4">Beck et al., 2019</xref>; <xref ref-type="bibr" rid="ref21">Karabatsos, 2003</xref>). In terms of their direction, low values in lz&#x002A; (i.e., &#x003C;&#x2212;1.3) are indicative of aberrance and the opposite is true for the number of Guttman errors (G) and U3 for which larger values are indicative of aberrant responding. All person fit analyses were conducted using the Perfit package (<xref ref-type="bibr" rid="ref50">Tendeiro et al., 2016</xref>) in R (<xref ref-type="bibr" rid="ref48">Team R. C, 2015</xref>).</p>
</sec>
<sec id="sec10">
<label>2.3.3</label>
<title>Ancillary analyses involving confirmatory factor analyses (or the graded response model)</title>
<p>Several CFA models or the Graded Response Model (GRM) were fit to the data to estimate unidimensionality, the presence of a methods factor, item-level statistics, and omnibus model fit. A preferred index in all these tests was the Root Mean Squared Error of Approximation (RMSEA) for which values less than 0.08 signal acceptable model fit. In CFA descriptive fit indices such as the Comparative Fit Index (CFI) need to take on values greater than 0.900. To ensure that the sample size was adequate we conducted a Monte Carlo simulation positing a unidimensional construct with 9 items, factor loadings equal to 0.70 and residual variances equal to 0.51. Using either the full or truncated samples parameter recovery ranged between 95 and 96%, and power for the factor loadings was greater than 99.9%. The chi-square statistic was overpowered, which is why it was not relied upon when evaluating model fit.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="sec11">
<label>3</label>
<title>Results</title>
<sec id="sec12">
<label>3.1</label>
<title>Identification of aberrance using mixture modeling</title>
<p><xref ref-type="fig" rid="fig2">Figure 2</xref> displays an optimal solution judged by the above fit indices and a minimum sample size of <italic>n</italic>&#x2009;&#x003E;&#x2009;50 participants, selected so that sample representation would be greater than 1%. This latter criterion is justified because subgroups with n&#x2009;&#x003C;&#x2009;=50 may not represent true subgroups in the population but rather artifacts of the sampling process. Based on the above criteria, a 7-class solution was superior to an 8-class solution as indices of AWE and ICL-BIC were smaller for the 7-class model compared to the 8-class model but not so for the CLC (AWE<sub>7-class</sub>&#x2009;=&#x2009;96866.600; AWE<sub>8-class</sub>&#x2009;=&#x2009;97023.894; ICL-BIC<sub>7-class</sub>&#x2009;=&#x2009;95873.538; ICL-BIC<sub>8-class</sub>&#x2009;=&#x2009;95916.688; CLC<sub>7-class</sub>&#x2009;=&#x2009;95141.478; CLC<sub>8-class</sub>&#x2009;=&#x2009;95100.482). Similarly, the 7-class model had lower values compared to the 6-class model supporting its preference for the CLC and ICL-BIC only (AWE<sub>7-class</sub>&#x2009;=&#x2009;96866.600; AWE<sub>6-class</sub>&#x2009;=&#x2009;96848.480; ICL-BIC<sub>7-class</sub>&#x2009;=&#x2009;95873.538; ICL-BIC<sub>6-class</sub>&#x2009;=&#x2009;95969.564; CLC<sub>7-class</sub>&#x2009;=&#x2009;95141.478; CLC<sub>6-class</sub>&#x2009;=&#x2009;95321.648). As shown in the figure, classes 6 and 7 represent aberrant responders reflecting low and high confidence ratings, respectively. Given that the scores were reversed for the LCA analysis, the expectation is that the same direction/scoring of the items would reflect ratings that are similar across items, expecting horizontal lines that are parallel to the X-axis. Differences from a &#x201C;flat&#x201D; line would be indicative of content differences regarding confidence and should be expected. However, as seen in class 6, mean responses to the first four items (about being confident in math) were very low followed by very high levels in the items describing lack of confidence. However, as mentioned above, given that items were reverse coded for the analysis, the direction for all items was the same, and thus, the difference in ratings of that magnitude is likely indicative of inattention, carelessness, random responding, or other personal or situational factors reflecting some form of systematic error of measurement. In other words, participants provided the same rating (e.g., agreement) for items such as &#x201C;I am good at math&#x201D; and &#x201C;Mathematics is not my strength,&#x201D; which shows inconsistency and error. The same was true for latent class 7, for which confidence ratings were high in math followed by very low ratings which again, is incongruent given the same direction of items in this presentation (items were reverse-coded).</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption><p>Optimal latent class solution for the measurement of confidence for mathematics in Saudi Arabia.</p></caption>
<graphic xlink:href="fpsyg-15-1489054-g002.tif"/>
</fig>
</sec>
<sec id="sec13">
<label>3.2</label>
<title>Validating aberrance using person-fit indicators</title>
<p><xref ref-type="fig" rid="fig3">Figure 3</xref> displays densities and cutoff values using a cutoff threshold of 10% (or 90% depending on whether low or high scores were considered aberrant) for the three person-fit indices. As shown in the figure, the cutoff value for the lz&#x002A; indicator was &#x2212;1.348, for the number of Guttman errors 45.05, and for the U3 indicator 0.47. <xref ref-type="fig" rid="fig4">Figure 4</xref> provides densities for the three person-fit indicators by latent class. A means analysis using the Analysis of Variance (ANOVA) model was run to identify and contrast point estimates across the 7 classes. Results pointed to significant differences between groups using the omnibus F-test for Lz&#x002A; [<italic>F</italic>(6, 4,510)&#x2009;=&#x2009;456.926, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001], Guttman errors [<italic>F</italic>(6, 4,511)&#x2009;=&#x2009;594.943, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001] and the U3 index [<italic>F</italic>(6, 4,511)&#x2009;=&#x2009;273.023, <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001]. Using Tukey&#x2019;s <italic>post-hoc</italic> tests results indicated that the two aberrant classes (i.e., 6 and 7) were significantly more aberrant compared to all other classes. When contrasted with each other, class 7 estimates of aberrance were significantly elevated compared to class 6 estimates. Using the eta-squared effect size metric, differences between classes 6 and 7 and all other classes ranged between 0.27 and 0.44, reflecting larger-than-large effects (large eta squared =0.14; <xref ref-type="bibr" rid="ref7">Cohen, 1988</xref>; <xref ref-type="bibr" rid="ref24">Lakens, 2013</xref>).</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption><p>Cutoff values for indices of aberrant responding.</p></caption>
<graphic xlink:href="fpsyg-15-1489054-g003.tif"/>
</fig>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption><p>Densities of person-fit indices per latent class in the 7-class optimal solution (modal).</p></caption>
<graphic xlink:href="fpsyg-15-1489054-g004.tif"/>
</fig>
<p>When evaluating differences using the <xref ref-type="fig" rid="fig3">Figure 3</xref> cutoff values it is evident that the mean number of Guttman errors for classes 6 and 7 that were 74 and 49, respectively, exceed the critical values of 45 suggesting that most participants were aberrant responders emitting a significant number of Guttman errors. Regarding the lz&#x002A; and U3 indices, class 7 had mean estimates (MeanLz&#x002A;&#x2009;=&#x2009;2.925; MeanU3&#x2009;=&#x2009;0.633) much greater than the thresholds defining aberrant responding. For class 6, mean point estimates were close to the cutoff values but lower (MeanLz&#x002A;&#x2009;=&#x2009;&#x2212;0.449; MeanU3&#x2009;=&#x2009;0.395).</p>
</sec>
<sec id="sec14">
<label>3.3</label>
<title>Was the observed aberrance a function of a methods factor?</title>
<p>This hypothesis was tested by contrasting a unidimensional versus a two-factor model with the latter separating positively worded from negatively worded items. After fitting the data to a unidimensional measurement model using the Weighted Least Squares Mean and Variance adjusted (WLSMV) estimator that is appropriate for ordered data, results indicated that all items loaded significantly to their respective factor (at <italic>p</italic>&#x2009;&#x003C;&#x2009;0.001) but the global model fit was poor (CFI&#x2009;=&#x2009;0.824, RMSEA&#x2009;=&#x2009;0.210). On the contrary, separating the items based on valence into positively worded and negatively worded factors resulted in improved and acceptable model fit (CFI&#x2009;=&#x2009;0.974; RMSEA&#x2009;=&#x2009;0.078). Interestingly, the correlation between factors was <italic>r</italic>&#x2009;=&#x2009;0.492 using Pearson&#x2019;s r. These findings point to the existence of a methods factor that accounted for the intercorrelation between items due to item wording.</p>
</sec>
<sec id="sec15">
<label>3.4</label>
<title>Evaluating scale psychometrics using original and purified data</title>
<p>Using a purification procedure, analyses of internal consistency reliability and factorial validity were conducted contrasting the results from the full sample against those from a truncated sample that excluded the participants from classes 6 and 7. Regarding internal consistency reliability results indicated that the full sample estimate was 0.820 and was elevated to 0.837 for the truncated sample. When contrasting the two coefficients using a Fisher&#x2019;s z transformation Z-test (<xref ref-type="bibr" rid="ref14">Hinkle et al., 1988</xref>), results pointed to significant improvements in internal consistency reliability for the truncated sample (<italic>Z</italic>&#x2009;=&#x2009;2.542, <italic>p</italic>&#x2009;=&#x2009;0.011).</p>
<p>Using the Confirmatory Factor Analysis (CFA) model results indicated that model fit was improved using the truncated data compared to the full sample (e.g., CFI<sub>Full</sub>&#x2009;=&#x2009;0.824; CFI<sub>Truncated</sub>&#x2009;=&#x2009;0.890). Interestingly, when contrasting models using the RMSEA, the confidence intervals for the RMSEA using full data at 95% ranged between 0.205 and 0.214. The point estimate for the RMSEA using truncated data was 0.167 and its respective 95% confidence interval ranged between 0.162 and 0.171. Thus, the point estimate of the RMSEA using the truncated data was significantly different from the one using full data as the confidence interval for the full data did not include the point estimate of the RMSEA using truncated data.</p>
</sec>
</sec>
<sec sec-type="discussion" id="sec16">
<label>4</label>
<title>Discussion</title>
<p>The goals of the present study were (a) to identify subgroups of participants who respond in the same manner in reverse-coded items using mixture modeling, (b) to validate the identification of aberrance mixtures using person fit indicators, (c) to test the presence of systematic error variance through supporting an additional &#x201C;method&#x201D; factor, and (d) evaluate the effects of aberrant responders on the psychometrics of a confidence scale related to mathematics achievement.</p>
<p>One important finding was that two classes of individuals who responded in the same manner across reverse-coded items were identified using mixture modeling, one with low confidence and one with high confidence for mathematics. Furthermore, an analysis of these two groups using person fit indicators showed that mean levels of aberrance were significantly elevated in these two groups, compared to the remaining five subgroups. This finding provides further support for using mixture modeling to identify subgroups that potentially behave in aberrant ways. The concordance between person-fit indicators and the subgrouping produced via the LCA analysis was high, validating the use of person-based analyses. Furthermore, the combination of participants in classes 6 and 7 represented 9% of the sample. This magnitude is lower compared to unpublished data from Bandalos, Coleman, and Gerstner (cited in <xref ref-type="bibr" rid="ref38">Reise et al., 2016</xref>) who reported rates of inconsistent responses to positively worded and negatively worded items in the Rosenberg Self-Esteem scale (RSES) at rates between 12 and 17%.</p>
<p>Another important finding was that the inclusion of these two subgroups had important negative implications for the scale&#x2019;s psychometric analyses. An inferential statistical test indicated significant improvements in internal consistency reliability using the truncated sample compared to the full sample. Similarly, by employing the 95% confidence intervals of the RMSEA significant differences in model fit were present, with better model fit being associated with the truncated dataset. This finding agrees with earlier work in that negatively worded items were associated with poor psychometric characteristics such as enhanced item difficulty levels and lower discriminant ability compared to positively worded items, as per the IRT model (<xref ref-type="bibr" rid="ref17">Kam, 2023</xref>; <xref ref-type="bibr" rid="ref42">Sliter and Zickar, 2014</xref>). Thus, the idea that by reversing negatively worded items they function in equivalent ways with their positive counterparts simply does not hold.</p>
<p>A third finding was that the integration of person-based, and variable-based analysis contributed to our conclusion that positively worded and negatively worded items represent distinct facets due to item wording, representing method variance rather than distinct facets of the underlying confidence construct. In <xref ref-type="bibr" rid="ref28">Marsh et al. (2010)</xref> terms this systematic form of variance represents &#x201C;ephemeral&#x201D; variance. Using the factor model, results indicated that a 2-factor solution favored the unidimensional model with all positively and all negatively worded items loading on two distinct dimensions. This finding agrees with past studies that a method effects factor was identified (e.g., <xref ref-type="bibr" rid="ref20">Kam et al., 2021</xref>; <xref ref-type="bibr" rid="ref34">Podsakoff et al., 2003</xref>).</p>
<sec id="sec17">
<label>4.1</label>
<title>Recommendations limitations and future directions</title>
<p>Based on the empirical evidence described above and the current empirical findings it is suggested that reverse-coded items should be avoided as they may be associated with erroneous responses, carelessness, lack of understanding, or the tendency to satisfice (<xref ref-type="bibr" rid="ref23">Krosnick, 1991</xref>; <xref ref-type="bibr" rid="ref46">Su&#x00E1;rez-&#x00C1;lvarez et al., 2018</xref>). The empirical evidence has suggested that their inclusion likely contributes to spurious rather than substantive measurements (factors) due to item format (positive versus negative phrasing; <xref ref-type="bibr" rid="ref2">Barnette, 1996</xref>). Thus, the results from the present study add to previous recommendations that the practice of including negatively worded items in surveys and self-report instruments may introduce artificial methods effects and needs to be avoided as it may compromise both content and construct validity (<xref ref-type="bibr" rid="ref26">Marsh, 1996</xref>; <xref ref-type="bibr" rid="ref10">Dom&#x00ED;nguez-Salas et al., 2022</xref>). Specifically, effects of negatively worded items on cognitive fatigue or lack of cognitive reflection have been documented (<xref ref-type="bibr" rid="ref12">Fukudome and Takeda, 2024</xref>; <xref ref-type="bibr" rid="ref22">Koutsogiorgi and Michaelides, 2022</xref>; <xref ref-type="bibr" rid="ref30">Merritt, 2012</xref>) or have been linked to specific personality types (<xref ref-type="bibr" rid="ref54">Weems et al., 2003</xref>; <xref ref-type="bibr" rid="ref9">DiStefano and Motl, 2009</xref>) such as the behavioral inhibition system (BIS, <xref ref-type="bibr" rid="ref37">Quilty et al., 2006</xref>) or behavioral activation systems (BAS, <xref ref-type="bibr" rid="ref56">Weydmann et al., 2020</xref>). On the opposite side of this argument, however, <xref ref-type="bibr" rid="ref51">Vigil-Colet et al. (2020)</xref> suggested that reversed coded items could be valuable and could be used in survey measurement only if acquiescence and response biases are controlled for statistically. A series of novel methodologies are currently available in that regard (<xref ref-type="bibr" rid="ref33">Plieninger and Heck, 2018</xref>; <xref ref-type="bibr" rid="ref25">Machado et al., 2024</xref>).</p>
<p>The present study is limited for several reasons. First, there were no direct and observable indicators of aberrant responding; instead, aberrance was inferred from the person fit indices as they reflect deviations between observed and expected responses to items based on adherence to the Guttman pattern. The use of additional measurements such as eye-tracking, cognitive, or self-report measures could provide additional insight into the causes behind inconsistent responses in negatively valenced items.</p>
<p>In the future, it will be important to devise methodologies to both identify and correct estimates for the presence of aberrant responses due to reverse-coded items. <xref ref-type="bibr" rid="ref5">Bolt et al. (2020)</xref> presented an IRT mixture model that makes use of <xref ref-type="bibr" rid="ref9002">Samejima&#x2019;s (1969)</xref> Graded Response Model (GRM) to identify what they termed as &#x201C;confused&#x201D; classes. They presented models to identify full confusion, utilizing all items of a scale, or partial confusion utilizing half of the items. <xref ref-type="bibr" rid="ref16">Kam (2018)</xref> proposed the latent difference (LD) modeling approach from <xref ref-type="bibr" rid="ref35">Pohl et al. (2008)</xref> to identify method effects. Further combinations of mixture models that model separate ability and aberrance and adjust person-ability estimates may also be useful (<xref ref-type="bibr" rid="ref57">Yamamoto, 1989</xref>). <xref ref-type="bibr" rid="ref13">Garcia-Pardina et al. (2024)</xref> proposed a new model to estimate &#x201C;substantive dimensionality&#x201D; which accommodated variance due to wording effects which entailed exploratory graph analysis (EGA) and parallel analysis (PA). <xref ref-type="bibr" rid="ref18">Kam and Fan (2020)</xref> proposed multitrait-multimethod methodologies within the factor mixture models to capture heterogeneous responses and <xref ref-type="bibr" rid="ref19">Kam and Meyer (2022)</xref> suggested applying non-linear methodologies. Last, the inclusion of advanced technologies may also provide additional evidence on explanatory factors (<xref ref-type="bibr" rid="ref22">Koutsogiorgi and Michaelides, 2022</xref>).</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="sec18">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link xlink:href="https://timssandpirls.bc.edu/timss2019/" ext-link-type="uri">https://timssandpirls.bc.edu/timss2019/</ext-link>.</p>
</sec>
<sec sec-type="ethics-statement" id="sec19">
<title>Ethics statement</title>
<p>Ethical review and approval were not required for the study on human participants in accordance with the local legislation and institutional requirements. The studies were conducted in accordance with the local legislation and institutional requirements. This is a secondary data analysis. All ethical procedures are described at: <ext-link xlink:href="https://timssandpirls.bc.edu/timss2019/" ext-link-type="uri">https://timssandpirls.bc.edu/timss2019/</ext-link>. The studies were conducted in accordance with the local legislation and institutional requirements. Written informed consent for participation in this study was provided by the participants&#x2019; legal guardians/next of kin.</p>
</sec>
<sec sec-type="author-contributions" id="sec20">
<title>Author contributions</title>
<p>FA: Conceptualization, Formal analysis, Methodology, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. MA: Data curation, Funding acquisition, Software, Validation, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec sec-type="funding-information" id="sec21">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. We would like to thank Researchers Supporting Project number (RSPD2024R601), King Saud University, Riyadh, Saudi Arabia for funding this research work.</p>
</sec>
<sec sec-type="COI-statement" id="sec22">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
<p>The handling editor DS declared a past collaboration with the author FA.</p>
</sec>
<sec sec-type="disclaimer" id="sec23">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link xlink:href="https://www.iea.nl/studies/iea/timss/2019" ext-link-type="uri">https://www.iea.nl/studies/iea/timss/2019</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link xlink:href="https://timssandpirls.bc.edu/timss2019/" ext-link-type="uri">https://timssandpirls.bc.edu/timss2019/</ext-link></p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Av&#x015F;ar</surname> <given-names>A. &#x015E;.</given-names></name></person-group> (<year>2022</year>). <article-title>Aberrant individuals&#x2019; effects on fit indices both of confirmatory factor analysis and polytomous IRT models</article-title>. <source>Curr Psychol</source> <volume>41</volume>, <fpage>7427</fpage>&#x2013;<lpage>7440</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s12144-021-01563-4</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barnette</surname> <given-names>J. J.</given-names></name></person-group> (<year>1996</year>). <article-title>Responses that may indicate nonattending behaviors in three self-administered educational surveys</article-title>. <source>Res Sch</source> <volume>3</volume>, <fpage>49</fpage>&#x2013;<lpage>59</lpage>.</citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baumgartner</surname> <given-names>H.</given-names></name> <name><surname>Weijters</surname> <given-names>B.</given-names></name> <name><surname>Pieters</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>Misresponse to survey questions: a conceptual framework and empirical test of the effects of reversals, negations, and polar opposite core concepts</article-title>. <source>J Mark Res</source> <volume>55</volume>, <fpage>869</fpage>&#x2013;<lpage>883</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0022243718811848</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beck</surname> <given-names>M. F.</given-names></name> <name><surname>Albano</surname> <given-names>A. D.</given-names></name> <name><surname>Smith</surname> <given-names>W. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Person-fit as an index of inattentive responding: a comparison of methods using polytomous survey data</article-title>. <source>Appl Psychol Meas</source> <volume>43</volume>, <fpage>374</fpage>&#x2013;<lpage>387</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0146621618798666</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bolt</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>Y. C.</given-names></name> <name><surname>Meyer</surname> <given-names>R. H.</given-names></name> <name><surname>Pier</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>An IRT mixture model for rating scale confusion associated with negatively worded items in measures of social-emotional learning</article-title>. <source>Appl Meas Educ</source> <volume>33</volume>, <fpage>331</fpage>&#x2013;<lpage>348</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2020.1789140</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clauss</surname> <given-names>K.</given-names></name> <name><surname>Bardeen</surname> <given-names>J. R.</given-names></name></person-group> (<year>2020</year>). <article-title>Addressing psychometric limitations of the attentional control scale via bifactor modeling and item modification</article-title>. <source>J Pers Assess</source> <volume>102</volume>, <fpage>415</fpage>&#x2013;<lpage>427</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00223891.2018.1521417</pub-id>, PMID: <pub-id pub-id-type="pmid">30398371</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Cohen</surname> <given-names>J.</given-names></name></person-group> (<year>1988</year>). <source>Statistical power analysis for the behavioral sciences</source>. <edition>2nd</edition> Edn. <publisher-loc>New York</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates</publisher-name>.</citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cui</surname> <given-names>Y.</given-names></name> <name><surname>Mousavi</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Explore the usefulness of person-1t analysis on large-scale assessment</article-title>. <source>Int J Test</source> <volume>15</volume>, <fpage>23</fpage>&#x2013;<lpage>49</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305058.2014.977444</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiStefano</surname> <given-names>C.</given-names></name> <name><surname>Motl</surname> <given-names>R. W.</given-names></name></person-group> (<year>2009</year>). <article-title>Personality correlates of method effects due to negatively worded items on the Rosenberg self-esteem scale</article-title>. <source>Personal Individ Differ</source> <volume>46</volume>, <fpage>309</fpage>&#x2013;<lpage>313</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.paid.2008.10.020</pub-id>, PMID: <pub-id pub-id-type="pmid">37217890</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dom&#x00ED;nguez-Salas</surname> <given-names>S.</given-names></name> <name><surname>Andr&#x00E9;s-Villas</surname> <given-names>M.</given-names></name> <name><surname>Riera-Sampol</surname> <given-names>A.</given-names></name> <name><surname>Tauler</surname> <given-names>P.</given-names></name> <name><surname>Bennasar-Veny</surname> <given-names>M.</given-names></name> <name><surname>Aguilo</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Analysis of the psychometric properties of the sense of coherence scale (SOC-13) in patients with cardiovascular risk factors: a study of the method effects associated with negatively worded items</article-title>. <source>Health Qual Life Outcomes</source> <volume>20</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s12955-021-01914-6</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Emons</surname> <given-names>W.</given-names></name></person-group> (<year>2008</year>). <article-title>Nonparametric person-fit analysis of polytomous item scores</article-title>. <source>Appl Psychol Meas</source> <volume>32</volume>, <fpage>224</fpage>&#x2013;<lpage>247</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0146621607302479</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fukudome</surname> <given-names>K.</given-names></name> <name><surname>Takeda</surname> <given-names>T.</given-names></name></person-group> (<year>2024</year>). <article-title>The influence of cognitive reflection on consistency of responses between reversed and direct items</article-title>. <source>Personal Individ Differ</source> <volume>230</volume>:<fpage>112811</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.paid.2024.112811</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garcia-Pardina</surname> <given-names>A.</given-names></name> <name><surname>Abad</surname> <given-names>F. J.</given-names></name> <name><surname>Christensen</surname> <given-names>A. P.</given-names></name> <name><surname>Golino</surname> <given-names>H.</given-names></name> <name><surname>Garrido</surname> <given-names>L. E.</given-names></name></person-group> (<year>2024</year>). <article-title>Dimensionality assessment in the presence of wording effects: a network psychometric and factorial approach</article-title>. <source>Behav Res Methods</source> <volume>56</volume>, <fpage>6179</fpage>&#x2013;<lpage>6197</lpage>. doi: <pub-id pub-id-type="doi">10.3758/s13428-024-02348-w</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinkle</surname> <given-names>D. E.</given-names></name> <name><surname>Wiersma</surname> <given-names>W.</given-names></name> <name><surname>Jurs</surname> <given-names>S. G.</given-names></name></person-group> (<year>1988</year>). <source>Applied statistics for the behavioral sciences</source>. <edition>2nd</edition> Edn. <publisher-loc>Boston</publisher-loc>: <publisher-name>Houghton Mifflin Company</publisher-name>.</citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaensson</surname> <given-names>M.</given-names></name> <name><surname>Nilsson</surname> <given-names>U.</given-names></name></person-group> (<year>2017</year>). <article-title>Impact of changing positively worded items to negatively worded items in the Swedish web-version of the quality of recovery (SwQoR) questionnaire</article-title>. <source>J Eval Clin Pract</source> <volume>23</volume>, <fpage>502</fpage>&#x2013;<lpage>507</lpage>. doi: <pub-id pub-id-type="doi">10.1111/jep.12639</pub-id>, PMID: <pub-id pub-id-type="pmid">27650792</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kam</surname> <given-names>C. C. S.</given-names></name></person-group> (<year>2018</year>). <article-title>Novel insights into item keying/valence effect using latent difference modeling analysis</article-title>. <source>J Pers Assess</source> <volume>100</volume>, <fpage>389</fpage>&#x2013;<lpage>397</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00223891.2017.1369095</pub-id>, PMID: <pub-id pub-id-type="pmid">28980826</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kam</surname> <given-names>C. C. S.</given-names></name></person-group> (<year>2023</year>). <article-title>Why do regular and reversed items load on separate factors? Response difficulty vs. item extremity</article-title>. <source>Educ Psychol Meas</source> <volume>83</volume>, <fpage>1085</fpage>&#x2013;<lpage>1112</lpage>. doi: <pub-id pub-id-type="doi">10.1177/00131644221143972</pub-id>, PMID: <pub-id pub-id-type="pmid">37974659</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kam</surname> <given-names>C. C. S.</given-names></name> <name><surname>Fan</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Investigating response heterogeneity in the context of positively and negatively worded items by using factor mixture modeling</article-title>. <source>Organ Res Methods</source> <volume>23</volume>, <fpage>322</fpage>&#x2013;<lpage>341</lpage>. doi: <pub-id pub-id-type="doi">10.1177/1094428118790371</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kam</surname> <given-names>C. C. S.</given-names></name> <name><surname>Meyer</surname> <given-names>J. P.</given-names></name></person-group> (<year>2022</year>). <article-title>Testing the nonlinearity assumption underlying the use of reverse-keyed items: a logical response perspective</article-title>. <source>Assessment</source> <volume>30</volume>, <fpage>1569</fpage>&#x2013;<lpage>1589</lpage>. doi: <pub-id pub-id-type="doi">10.1177/10731911221106775</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kam</surname> <given-names>C. C. S.</given-names></name> <name><surname>Meyer</surname> <given-names>J. P.</given-names></name> <name><surname>Sun</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Why do people agree with both regular and reversed items?</article-title> <source>A logical response perspective Assessment</source> <volume>28</volume>, <fpage>1110</fpage>&#x2013;<lpage>1124</lpage>. doi: <pub-id pub-id-type="doi">10.1177/10731911211001931</pub-id>, PMID: <pub-id pub-id-type="pmid">33779309</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Karabatsos</surname> <given-names>G.</given-names></name></person-group> (<year>2003</year>). <article-title>Comparing the aberrant response detection performance of thirty-six person-fit statistics</article-title>. <source>Appl Meas Educ</source> <volume>16</volume>, <fpage>277</fpage>&#x2013;<lpage>298</lpage>. doi: <pub-id pub-id-type="doi">10.1207/S15324818AME1604_2</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koutsogiorgi</surname> <given-names>C. C.</given-names></name> <name><surname>Michaelides</surname> <given-names>M. P.</given-names></name></person-group> (<year>2022</year>). <article-title>Response tendencies due to item wording using eye-tracking methodology accounting for individual differences and item characteristics</article-title>. <source>Behav Res Methods</source> <volume>54</volume>, <fpage>2252</fpage>&#x2013;<lpage>2270</lpage>. doi: <pub-id pub-id-type="doi">10.3758/s13428-021-01719-x</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krosnick</surname> <given-names>J. A.</given-names></name></person-group> (<year>1991</year>). <article-title>Response strategies for coping with the cognitive demands of attitude measures in surveys</article-title>. <source>Appl Cogn Psychol</source> <volume>5</volume>, <fpage>213</fpage>&#x2013;<lpage>236</lpage>. doi: <pub-id pub-id-type="doi">10.1002/acp.2350050305</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lakens</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs</article-title>. <source>Front Psychol</source> <volume>4</volume>:<fpage>863</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2013.00863</pub-id>, PMID: <pub-id pub-id-type="pmid">24324449</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Machado</surname> <given-names>G. M.</given-names></name> <name><surname>Hauck-Filho</surname> <given-names>N.</given-names></name> <name><surname>Pallini</surname> <given-names>A. C.</given-names></name> <name><surname>Dias-Viana</surname> <given-names>J. L.</given-names></name> <name><surname>Chiappetta Santana</surname> <given-names>L. H. B.</given-names></name> <name><surname>Medeiros da Silva</surname> <given-names>C. A. N.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>Investigating the acquiescent responding impact in empathy measures</article-title>. <source>Int J Test</source>. <volume>24</volume>, <fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305058.2024.2364170</pub-id></citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>H. W.</given-names></name></person-group> (<year>1996</year>). <article-title>Positive and negative self-esteem: a substantively meaningful distinction or artifactors?</article-title> <source>J Pers Soc Psychol</source> <volume>70</volume>, <fpage>810</fpage>&#x2013;<lpage>819</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0022-3514.70.4.810</pub-id>, PMID: <pub-id pub-id-type="pmid">8636900</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>H. W.</given-names></name> <name><surname>Grayson</surname> <given-names>D.</given-names></name></person-group> (<year>1995</year>). &#x201C;<article-title>Latent variable models of multitrait-multimethod data</article-title>&#x201D; in <source>Structural equation modeling: Concept, issues, and applications</source>. ed. <person-group person-group-type="editor"><name><surname>Hoyle</surname> <given-names>R. H.</given-names></name></person-group> (<publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage</publisher-name>), <fpage>177</fpage>&#x2013;<lpage>198</lpage>.</citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>H. W.</given-names></name> <name><surname>Scalas</surname> <given-names>L. F.</given-names></name> <name><surname>Nagengast</surname> <given-names>B.</given-names></name></person-group> (<year>2010</year>). <article-title>Longitudinal tests of competing factor structures for the Rosenberg self-esteem scale: traits, ephemeral artifacts, and stable response styles</article-title>. <source>Psychol Assess</source> <volume>22</volume>, <fpage>366</fpage>&#x2013;<lpage>381</lpage>. doi: <pub-id pub-id-type="doi">10.1037/a0019225</pub-id>, PMID: <pub-id pub-id-type="pmid">20528064</pub-id></citation></ref>
<ref id="ref9111"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Masyn</surname> <given-names>K. E.</given-names></name></person-group> (<year>2013</year>). <article-title>Latent class analysis and finite mixture modeling</article-title>. In: <source>The Oxford handbook of quantitative methods: Statistical analysis</source>. Ed. <person-group person-group-type="editor"><name><surname>Little</surname> <given-names>T. D.</given-names></name></person-group> <publisher-name>(Oxford University Press)</publisher-name>, pp. <fpage>551</fpage>&#x2013;<lpage>611</lpage>.</citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meijer</surname> <given-names>R. R.</given-names></name></person-group> (<year>1994</year>). <article-title>The number of guttman errors as a simple and powerful person-1t statistic</article-title>. <source>Appl Psychol Meas</source> <volume>18</volume>, <fpage>311</fpage>&#x2013;<lpage>314</lpage>. doi: <pub-id pub-id-type="doi">10.1177/014662169401800402</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Merritt</surname> <given-names>S. M.</given-names></name></person-group> (<year>2012</year>). <article-title>The two-factor solution to Allen and Meyer&#x2019;s (1990) affective commitment scale: effects of negatively worded items</article-title>. <source>J Bus Psychol</source> <volume>27</volume>, <fpage>421</fpage>&#x2013;<lpage>436</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10869-011-9252-3</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Michaelides</surname> <given-names>M. P.</given-names></name></person-group> (<year>2019</year>). <article-title>Negative keying effects in the factor structure of TIMSS 2011 motivation scales and associations with reading achievement</article-title>. <source>Appl Meas Educ</source> <volume>32</volume>, <fpage>365</fpage>&#x2013;<lpage>378</lpage>. doi: <pub-id pub-id-type="doi">10.1080/08957347.2019.1660349</pub-id></citation></ref>
<ref id="ref9001"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Mullis</surname> <given-names>I. V. S.</given-names></name> <name><surname>Martin</surname> <given-names>M. O.</given-names></name> <name><surname>Foy</surname> <given-names>P.</given-names></name> <name><surname>Kelly</surname> <given-names>D. L.</given-names></name> <name><surname>Fishbein</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <source>TIMSS 2019 international results in mathematics and science</source>. <publisher-loc>International Association for the Evaluation of Educational Achievement (IEA)</publisher-loc>. Available at: <ext-link xlink:href="https://timssandpirls.bc.edu/timss2019/international-results/" ext-link-type="uri">https://timssandpirls.bc.edu/timss2019/international-results/</ext-link></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedersen</surname> <given-names>H. S.</given-names></name> <name><surname>Christensen</surname> <given-names>K. S.</given-names></name> <name><surname>Prior</surname> <given-names>A.</given-names></name> <name><surname>Christensen</surname> <given-names>K. B.</given-names></name></person-group> (<year>2024</year>). <article-title>The dimensionality of the perceived stress scale: the presence of opposing items is a source of measurement error</article-title>. <source>J Affect Disord</source> <volume>344</volume>, <fpage>485</fpage>&#x2013;<lpage>494</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jad.2023.10.109</pub-id>, PMID: <pub-id pub-id-type="pmid">37852582</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Plieninger</surname> <given-names>H.</given-names></name> <name><surname>Heck</surname> <given-names>D. W.</given-names></name></person-group> (<year>2018</year>). <article-title>A new model for acquiescence at the interface of psychometrics and cognitive psychology</article-title>. <source>Multivar Behav Res</source> <volume>53</volume>, <fpage>633</fpage>&#x2013;<lpage>654</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00273171.2018.1469966</pub-id>, PMID: <pub-id pub-id-type="pmid">29843531</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Podsakoff</surname> <given-names>P. M.</given-names></name> <name><surname>MacKenzie</surname> <given-names>S. B.</given-names></name> <name><surname>Lee</surname> <given-names>J. Y.</given-names></name> <name><surname>Podsakoff</surname> <given-names>N. P.</given-names></name></person-group> (<year>2003</year>). <article-title>Common method biases in behavioral research: a critical review of the literature and recommended remedies</article-title>. <source>J Appl Psychol</source> <volume>88</volume>, <fpage>879</fpage>&#x2013;<lpage>903</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0021-9010.88.5.879</pub-id>, PMID: <pub-id pub-id-type="pmid">14516251</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pohl</surname> <given-names>S.</given-names></name> <name><surname>Steyer</surname> <given-names>R.</given-names></name> <name><surname>Kraus</surname> <given-names>K.</given-names></name></person-group> (<year>2008</year>). <article-title>Modeling method effects as individual causal effects</article-title>. <source>J R Stat Soc Ser A</source> <volume>171</volume>, <fpage>41</fpage>&#x2013;<lpage>63</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-985X.2007.00517.x</pub-id>, PMID: <pub-id pub-id-type="pmid">39375710</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ponce</surname> <given-names>F. P.</given-names></name> <name><surname>Torres Irribarra</surname> <given-names>D.</given-names></name> <name><surname>Verg&#x00E9;s</surname> <given-names>&#x00C1;.</given-names></name> <name><surname>Arias</surname> <given-names>V. B.</given-names></name></person-group> (<year>2022</year>). <article-title>Wording effects in assessment: missing the trees for the forest</article-title>. <source>Multivar Behav Res</source> <volume>57</volume>, <fpage>718</fpage>&#x2013;<lpage>734</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00273171.2021.1925075</pub-id>, PMID: <pub-id pub-id-type="pmid">34048313</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quilty</surname> <given-names>L. C.</given-names></name> <name><surname>Oakman</surname> <given-names>J. M.</given-names></name> <name><surname>Risko</surname> <given-names>E.</given-names></name></person-group> (<year>2006</year>). <article-title>Correlates of the Rosenberg self-esteem scale method effects</article-title>. <source>Struct Equ Model</source> <volume>13</volume>, <fpage>99</fpage>&#x2013;<lpage>117</lpage>. doi: <pub-id pub-id-type="doi">10.1207/s15328007sem1301_5</pub-id>, PMID: <pub-id pub-id-type="pmid">39256755</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reise</surname> <given-names>S. P.</given-names></name> <name><surname>Kim</surname> <given-names>D. S.</given-names></name> <name><surname>Mansolf</surname> <given-names>M.</given-names></name> <name><surname>Widaman</surname> <given-names>K. F.</given-names></name></person-group> (<year>2016</year>). <article-title>Is the bifactor model a better model or is it just better at modeling implausible responses? Application of iteratively reweighted least squares to the Rosenberg self-esteem scale</article-title>. <source>Multivar Behav Res</source> <volume>51</volume>, <fpage>818</fpage>&#x2013;<lpage>838</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00273171.2016.1243461</pub-id>, PMID: <pub-id pub-id-type="pmid">27834509</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roszkowski</surname> <given-names>M. J.</given-names></name> <name><surname>Soven</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Shifting gears: consequences of including two negatively worded items in the middle of a positively worded questionnaire</article-title>. <source>Assess Eval High Educ</source> <volume>35</volume>, <fpage>117</fpage>&#x2013;<lpage>134</lpage>. doi: <pub-id pub-id-type="doi">10.1080/02602930802618344</pub-id></citation></ref>
<ref id="ref9002"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samejima</surname> <given-names>F.</given-names></name></person-group> (<year>1969</year>). <article-title>Estimation of latent ability using a response pattern of graded scores</article-title>. <source>Psychometrika Monograph Supplement</source>, <volume>34</volume>:<fpage>100</fpage>.</citation></ref>
<ref id="ref40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schroeders</surname> <given-names>U.</given-names></name> <name><surname>Schmidt</surname> <given-names>C.</given-names></name> <name><surname>Gnambs</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>Detecting careless responding in survey data using stochastic gradient boosting</article-title>. <source>Educ Psychol Meas</source> <volume>82</volume>, <fpage>29</fpage>&#x2013;<lpage>56</lpage>. doi: <pub-id pub-id-type="doi">10.1177/00131644211004708</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sliter</surname> <given-names>K. A.</given-names></name> <name><surname>Zickar</surname> <given-names>M. J.</given-names></name></person-group> (<year>2014</year>). <article-title>An IRT examination of the psychometric functioning of negatively worded personality items</article-title>. <source>Educ Psychol Meas</source> <volume>74</volume>, <fpage>214</fpage>&#x2013;<lpage>226</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0013164413504584</pub-id></citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steedle</surname> <given-names>J. T.</given-names></name> <name><surname>Hong</surname> <given-names>M.</given-names></name> <name><surname>Cheng</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>The effects of inattentive responding on construct validity evidence when measuring social&#x2013;emotional learning competencies</article-title>. <source>Educ Meas Issues Pract</source> <volume>38</volume>, <fpage>101</fpage>&#x2013;<lpage>111</lpage>. doi: <pub-id pub-id-type="doi">10.1111/emip.12256</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steinmann</surname> <given-names>I.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Braeken</surname> <given-names>J.</given-names></name></person-group> (<year>2024</year>). <article-title>Who responds inconsistently to mixed-worded scales? Differences by achievement, age group, and gender</article-title>. <source>Assess Educ Principles, Policy &#x0026; Practice</source> <volume>31</volume>, <fpage>5</fpage>&#x2013;<lpage>31</lpage>. doi: <pub-id pub-id-type="doi">10.1080/0969594X.2024.2318554</pub-id></citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steinmann</surname> <given-names>I.</given-names></name> <name><surname>Strietholt</surname> <given-names>R.</given-names></name> <name><surname>Braeken</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>A constrained factor mixture analysis model for consistent and inconsistent respondents to mixed-worded scales</article-title>. <source>Psychol Methods</source> <volume>27</volume>, <fpage>667</fpage>&#x2013;<lpage>702</lpage>. doi: <pub-id pub-id-type="doi">10.1037/met0000392</pub-id>, PMID: <pub-id pub-id-type="pmid">33829811</pub-id></citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Su&#x00E1;rez-&#x00C1;lvarez</surname> <given-names>J.</given-names></name> <name><surname>Pedrosa</surname> <given-names>I.</given-names></name> <name><surname>Lozano</surname> <given-names>L. M.</given-names></name> <name><surname>Garc&#x00ED;a-Cueto</surname> <given-names>E.</given-names></name> <name><surname>Cuesta</surname> <given-names>M.</given-names></name> <name><surname>Mu&#x00F1;iz</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Using reversed items in Likert scales: a questionable practice</article-title>. <source>Psicothema</source> <volume>30</volume>, <fpage>149</fpage>&#x2013;<lpage>158</lpage>. doi: <pub-id pub-id-type="doi">10.7334/psicothema2018.33</pub-id>, PMID: <pub-id pub-id-type="pmid">29694314</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swain</surname> <given-names>S. D.</given-names></name> <name><surname>Weathers</surname> <given-names>D.</given-names></name> <name><surname>Niedrich</surname> <given-names>R. W.</given-names></name></person-group> (<year>2008</year>). <article-title>Assessing three sources of misresponse to reversed Likert items</article-title>. <source>J Mark Res</source> <volume>45</volume>, <fpage>116</fpage>&#x2013;<lpage>131</lpage>. doi: <pub-id pub-id-type="doi">10.1509/jmkr.45.1.116</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="book"><person-group person-group-type="author"><collab id="coll1">Team R. C</collab></person-group> (<year>2015</year>). <source>R: A language and environment for statistical computing</source>. <publisher-loc>Vienna, Austria</publisher-loc>: <publisher-name>R Foundation for Statistical Computing, Vienna, Austria</publisher-name>.</citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tendeiro</surname> <given-names>J.</given-names></name> <name><surname>Meijer</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Detection of invalid test scores: the usefulness of simple nonparametric statistics</article-title>. <source>J Educ Meas</source> <volume>51</volume>, <fpage>239</fpage>&#x2013;<lpage>259</lpage>. doi: <pub-id pub-id-type="doi">10.1111/jedm.12046</pub-id></citation></ref>
<ref id="ref50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tendeiro</surname> <given-names>J. N.</given-names></name> <name><surname>Meijer</surname> <given-names>R. R.</given-names></name> <name><surname>Niessen</surname> <given-names>A. S. M.</given-names></name></person-group> (<year>2016</year>). <article-title>PerFit: an R package for person-fit analysis in IRT</article-title>. <source>J Stat Softw</source> <volume>74</volume>, <fpage>1</fpage>&#x2013;<lpage>27</lpage>. doi: <pub-id pub-id-type="doi">10.18637/jss.v074.i05</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vigil-Colet</surname> <given-names>A.</given-names></name> <name><surname>Navarro-Gonz&#x00E1;lez</surname> <given-names>D.</given-names></name> <name><surname>Morales-Vives</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>To reverse or to not reverse Likert-type items: that is the question</article-title>. <source>Psicothema</source> <volume>32</volume>, <fpage>108</fpage>&#x2013;<lpage>114</lpage>. doi: <pub-id pub-id-type="doi">10.7334/psicothema2019.286</pub-id>, PMID: <pub-id pub-id-type="pmid">31954423</pub-id></citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weems</surname> <given-names>G. H.</given-names></name> <name><surname>Onwuegbuzie</surname> <given-names>A. J.</given-names></name></person-group> (<year>2001</year>). <article-title>The impact of midpoint responses and reverse coding on survey data</article-title>. <source>Meas Eval Couns Dev</source> <volume>34</volume>:<fpage>166</fpage>. doi: <pub-id pub-id-type="doi">10.1080/07481756.2002.12069033</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weems</surname> <given-names>G. H.</given-names></name> <name><surname>Onwuegbuzie</surname> <given-names>A. J.</given-names></name> <name><surname>Lustig</surname> <given-names>D.</given-names></name></person-group> (<year>2003</year>). <article-title>Profiles of respondents who respond inconsistently to positively-and negatively-worded items on rating scales</article-title>. <source>Evaluation Res Educ</source> <volume>17</volume>, <fpage>45</fpage>&#x2013;<lpage>60</lpage>. doi: <pub-id pub-id-type="doi">10.1080/14664200308668290</pub-id></citation></ref>
<ref id="ref55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weems</surname> <given-names>G. H.</given-names></name> <name><surname>Onwuegbuzie</surname> <given-names>A. J.</given-names></name> <name><surname>Schreiber</surname> <given-names>J. B.</given-names></name> <name><surname>Eggers</surname> <given-names>S. J.</given-names></name></person-group> (<year>2003</year>). <article-title>Characteristics of respondents who respond differently to positively and negatively worded items on rating scales</article-title>. <source>Assess Eval High Educ</source> <volume>28</volume>, <fpage>587</fpage>&#x2013;<lpage>604</lpage>. doi: <pub-id pub-id-type="doi">10.1080/0260293032000130234</pub-id></citation></ref>
<ref id="ref56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weydmann</surname> <given-names>G.</given-names></name> <name><surname>Filho</surname> <given-names>N. H.</given-names></name> <name><surname>Bizarro</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>Acquiescent responding can distort the factor structure of the BIS/BAS scales</article-title>. <source>Personal Individ Differ</source> <volume>152</volume>:<fpage>109563</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.paid.2019.109563</pub-id></citation></ref>
<ref id="ref57"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Yamamoto</surname> <given-names>K. Y.</given-names></name></person-group> (<year>1989</year>). <source>HYBRID model of IRT and latent class models</source>. <publisher-loc>Princeton, NJ</publisher-loc>: <publisher-name>Educational Testing Service.</publisher-name></citation></ref>
</ref-list>
<app-group>
<app id="app1">
<title>Appendix</title>
<p>Category information curves as per the GRM model fitted to the confidence data.</p>
<p><inline-graphic xlink:href="fpsyg-15-1489054-g005.tif"/></p>
</app>
</app-group>
</back>
</article>
