<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Educ.</journal-id>
<journal-title>Frontiers in Education</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Educ.</abbrev-journal-title>
<issn pub-type="epub">2504-284X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/feduc.2025.1595658</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Education</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Detection of cultural and linguistic differential item functioning in reading assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Woo</surname>
<given-names>Yejin</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2993757/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/resources/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Choi</surname>
<given-names>Youn-Jeng</given-names>
</name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1408242/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>Department of Education, Ewha Womans University</institution>, <addr-line>Seoul</addr-line>, <country>Republic of Korea</country></aff>
<author-notes>
<fn fn-type="edited-by" id="fn0001">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2743301/overview">Ester Villalonga-Olives</ext-link>, University of Maryland, United States</p>
</fn>
<fn fn-type="edited-by" id="fn0002">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/664813/overview">Sebastian Weirich</ext-link>, Institute for Educational Quality Improvement (IQB), Germany</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/3034925/overview">Christopher Amissah</ext-link>, University of Maryland, United States</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Youn-Jeng Choi, <email>younjengchoi@ewha.ac.kr</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>09</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>10</volume>
<elocation-id>1595658</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>03</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>08</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2025 Woo and Choi.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Woo and Choi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>This study aims to determine whether differential item functioning (DIF) occurs in the PISA 2018 reading assessment and, if so, to explore which factors, such as linguistic elements, achievement goals, and perceived reading instructions as cultural elements, contribute most significantly to its occurrence. The United States was set as the reference group, and comparisons were made with Canada, Singapore, and South Korea. Item response theory-likelihood ratio (IRT-LR), logistic regression, and Rasch Tree analyses were utilized to identify DIF. Multiple methods consistently showed that item CR551Q06 exhibited DIF. The Rasch Tree analysis revealed that linguistic rather than cultural differences were the primary contributors to DIF. Interestingly, no DIF was detected using the Rasch Tree method in the comparisons between the United States and Canada and between the United States and Singapore, in contrast to the IRT-LR and logistic regression results. The analysis highlighted translation issues as a major source of bias, suggesting that careful adaptation of assessments is crucial to reducing DIF. These findings challenge assumptions about cultural differences in educational outcomes and emphasize the need for further research using varied DIF detection methods in different cultural contexts.</p>
</abstract>
<kwd-group>
<kwd>DIF</kwd>
<kwd>PISA</kwd>
<kwd>IRT-likelihood ratio</kwd>
<kwd>logistic regression</kwd>
<kwd>Rasch Tree</kwd>
</kwd-group>
<counts>
<fig-count count="3"/>
<table-count count="6"/>
<equation-count count="4"/>
<ref-count count="62"/>
<page-count count="16"/>
<word-count count="13387"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Assessment, Testing and Applied Measurement</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="sec1">
<label>1</label>
<title>Introduction</title>
<p>Assessing academic achievement serves not only to measure an individual&#x2019;s academic abilities but also to provide a basis for offering appropriate educational support to learners. Therefore, academic achievement assessments must employ valid tools that accurately measure the intended construct. A valid assessment tool should be aligned with the purpose of the test and free from any bias towards or against specific groups. Tests that favor or disadvantage a particular group prevent making fair conclusions about the test-takers based on their scores. This bias can manifest either across the entire test or within individual items, which is referred to as differential item functioning (DIF). Scholars have long argued that DIF can occur based on variables such as gender, race, or social status (<xref ref-type="bibr" rid="ref9">Coleman, 1968</xref>).</p>
<p>DIF occurs when test-takers with the same ability level perform differently on individual test items, which is closely related to test fairness and equity in large-scale assessments. The presence of DIF in an assessment can threaten its validity or lead to misinterpretation of item-level group differences. However, the presence of DIF does not necessarily undermine the validity of the entire test. Instead, it highlights how certain test items may function differently for various groups of test-takers. Thus, identifying and addressing DIF is crucial to ensuring fairness and equity in assessments.</p>
<p>Language differences are a key factor contributing to DIF. Test items have linguistic characteristics that make them more sensitive to translation errors compared to other types of texts (<xref ref-type="bibr" rid="ref49">Solano-Flores et al., 2009</xref>). Moreover, translation issues are among the primary causes of DIF in international comparative studies (<xref ref-type="bibr" rid="ref60">Yildirim and Berbero&#x011D;lu, 2009</xref>). Even when test items are translated properly, translation bias can still occur, as noted by <xref ref-type="bibr" rid="ref16">Grisay and Monseur (2007)</xref>. For example, translations may result in longer texts, altered word counts, or imprecise meaning, which can impact how test-takers comprehend the test. Consequently, even if the translated items ask the same questions as the source items, the test-takers&#x2019; understanding and perceived difficulty of the items may differ.</p>
<p>Cultural differences can also lead to DIF. Culture, shaped through interactions with the environment, reflects regional characteristics (<xref ref-type="bibr" rid="ref22">Jang, 2010</xref>). Eastern and Western cultures, for example, show significant differences: Eastern cultures tend to emphasize collectivism and community, while Western cultures prioritize individualism. These cultural distinctions can influence how test items are interpreted and processed by different groups, contributing to DIF. <xref ref-type="bibr" rid="ref42">Qian and Lau (2022)</xref> identified achievement goals and perceived reading instruction as cultural variables that may significantly affect DIF between Eastern and Western students.</p>
<p>Achievement goal theory, largely based on Western literature, may function differently in East Asian contexts due to the competitive education systems and emphasis on achievement (<xref ref-type="bibr" rid="ref29">Lau and Lee, 2008</xref>; <xref ref-type="bibr" rid="ref30">Lau and Nie, 2008</xref>). East Asian students exhibit unique patterns, with a strong positive correlation between mastery and performance goals (<xref ref-type="bibr" rid="ref19">Ho and Hau, 2008</xref>) and greater adoption of avoidance goals compared to Western students (<xref ref-type="bibr" rid="ref63">Zusho and Clayton, 2011</xref>). Traditional reading instruction in East Asia, characterized by teacher-centered, competitive, and exam-focused methods (<xref ref-type="bibr" rid="ref59">Watkins and Biggs, 2001</xref>), contrasts with mastery-oriented Western practices. However, recent reforms in some societies with Confucian heritage culture (CHC) have influenced traditional pedagogical practices (<xref ref-type="bibr" rid="ref28">Lau and Ho, 2015</xref>). In particular, the reforms in China promote student autonomy and cooperation, aligning more with mastery-oriented approaches (<xref ref-type="bibr" rid="ref56">The Ministry of Education and People&#x2019;s Republic of China, 2011</xref>). Examining how these aspects have changed after the educational reforms could provide meaningful insights.</p>
<p>Based on the cultural differences between East and West, <xref ref-type="bibr" rid="ref42">Qian and Lau (2022)</xref> utilized PISA 2018 reading achievement data to explore the impact of achievement goals and perceived reading instruction on the academic performance of Chinese students. The study showed that achievement goals and perceived reading instruction, particularly disciplinary climate, adaptive instruction, and teacher stimulation, influenced Chinese students&#x2019; reading performance. However, while the study examined variables likely to differ between Eastern and Western cultures, it focused solely on China, an Eastern country, and did not address potential differences in reading ability between students from Eastern and Western countries.</p>
<p>This study aimed to address this limitation in the literature by focusing on how linguistic and cultural factors contribute to DIF. It specifically examined whether cultural variables, such as achievement goals and perceived reading instruction, significantly impact differences in reading ability between students from Eastern and Western countries. It employed multiple DIF detection methods, including item response theory-likelihood ratio (IRT-LR), logistic regression (LR), and Rasch Tree (RT).</p>
<p>DIF detection techniques are grounded in two primary theories: Item response theory (IRT) and classical test theory. These techniques vary in their algorithms, synchronization criteria, and the cutoff points used to identify DIF. However, DIF detection methods are not entirely consistent with one another (<xref ref-type="bibr" rid="ref3">Bakan Kalayc&#x0131;o&#x011F;lu and Berbero&#x011F;lu, 2010</xref>). In response, simultaneously applying multiple DIF detection methods is often recommended (<xref ref-type="bibr" rid="ref17">Hambleton, 2006</xref>).</p>
<p>Traditional DIF detection methods include the Mantel&#x2013;Haenszel method, SIBTEST, and logistic regression, which are based on classical test theory, while methods such as the likelihood ratio test, Lord&#x2019;s method, and Raju&#x2019;s method are based on IRT. These traditional methods are statistically intuitive, relatively simple to interpret, and have been validated for reliability through several studies and practical evaluations. In this study, the likelihood ratio test was selected for its foundation in IRT (<xref ref-type="bibr" rid="ref6">Camilli and Shepard, 1994</xref>), and logistic regression was chosen for its ability to detect both uniform and non-uniform DIF (<xref ref-type="bibr" rid="ref61">Zumbo, 1999</xref>).</p>
<p>Although traditional DIF detection methods have these strengths, they often fall short in accounting for the diverse subgroups within the test-taker population because they require predefined distinctions between focal and reference groups. Several innovative approaches have been proposed to overcome this limitation. For example, the IRTree model (<xref ref-type="bibr" rid="ref5">B&#x00F6;ckenholt, 2012</xref>) models the test-taker&#x2019;s response process as a multi-step decision-making process, represented by a tree structure. Additionally, the Rasch Tree method (<xref ref-type="bibr" rid="ref53">Strobl et al., 2015</xref>) combines logistic regression with recursive partitioning to form subgroups based on various characteristics and response patterns, while other methods, such as the item-focused tree model (<xref ref-type="bibr" rid="ref58">Tutz and Berger, 2016</xref>) and the mixture IRT model combining latent class and IRT models, have also been introduced.</p>
<p>Among these, the Rasch Tree method has the advantage of using all explanatory variables in the data to define possible subgroups through recursive partitioning without the need for researchers to set arbitrary threshold values for group definitions (<xref ref-type="bibr" rid="ref23">Jang and Lee, 2023</xref>). For this reason, the present study employed the Rasch Tree method for analysis.</p>
<p>This study utilized data from the PISA 2018 student questionnaire and reading assessment. The United States was designated as the reference country, with the following countries included for comparison: (1) Canada, which shares the similar written language and culture as the U.S.; (2) Singapore, which shares the similar written language but differs culturally; and (3) South Korea, which differs from the U.S. in both written language and culture. The objective is to investigate the presence of DIF within the PISA 2018 reading assessment across these countries. Using IRT-LR, logistic regression, and Rasch Tree methods, the study aims to determine whether DIF occurs and, if so, identify common DIF items and examine their characteristics. To further explore the potential causes of DIF, IRT-LR and logistic regression are used to infer the indirect influences of linguistic and cultural factors based on predefined country groupings. Building on this, the Rasch Tree method is applied to identify specific factors that directly contribute to DIF without relying on predefined group structures. The specific research questions are as follows:<list list-type="order">
<list-item>
<p>Does DIF occur in the PISA 2018 Reading assessment when utilizing IRT-likelihood ratio (IRT-LR), logistic regression, and Rasch Tree methods across the four different countries?</p>
</list-item>
<list-item>
<p>If DIF occurs, what are the DIF items, and what characteristics do they exhibit?</p>
</list-item>
<list-item>
<p>If DIF occurs, what cultural and linguistic factors influence the DIF in the PISA 2018 Reading assessment, as identified by the Rasch Tree method?</p>
</list-item>
</list></p>
</sec>
<sec id="sec2">
<label>2</label>
<title>Theoretical frameworks</title>
<sec id="sec3">
<label>2.1</label>
<title>Test translation error and its impact</title>
<p>The theory of test translation error was developed to tackle the difficulties of accurately assessing diverse populations across multiple languages (<xref ref-type="bibr" rid="ref49">Solano-Flores et al., 2009</xref>). This theory characterizes translation error as the perceived discrepancies in content, structure, and meaning between the original and translated versions of test items. Accordingly, translation errors are not solely caused by poor-quality translations; even when translators perform exceptionally well, such errors remain. These errors occur because languages represent cultural experiences in distinct ways (<xref ref-type="bibr" rid="ref13">Greenfield, 1997</xref>) and utilize different semiotic systems (<xref ref-type="bibr" rid="ref4">Bezemer and Kress, 2008</xref>). For instance, certain languages employ classifiers with nouns modified by numbers (<xref ref-type="bibr" rid="ref1">Aikhenvald, 2003</xref>), with different classifiers required depending on the type of noun. These classifiers convey specific characteristics (such as shape or quantity) about the noun, leading to potential differences in the information provided by translated items compared to their source material. Similarly, translating quantifiers (like &#x201C;some&#x201D; or &#x201C;any&#x201D;) can be challenging when dealing with non-Indo-European languages because they do not utilize these elements in the same way as English or French (<xref ref-type="bibr" rid="ref14">Grisay, 2007</xref>).</p>
<p>The theory suggests that translation involves not just the translator&#x2019;s work but also various factors in both the translation process and the development of assessment tools. These factors can influence aspects such as content, vocabulary frequency, and grammatical complexity in the translated tests. Translation encompasses numerous features that may vary between language versions of the same test items, ranging from visual layout to the content volume, cognitive demands, linguistic requirements, and the cultural context assumed for test-takers. For example, even a slight change in the scale of a graph can make a curve appear steeper than in the original, and contextual details meant to make an item relatable may refer to scenarios that are unfamiliar to students in the target language. Although such errors are not the direct fault of the translator, they significantly influence the translated test&#x2019;s overall characteristics.</p>
<p>The theory posits that translation errors can be categorized based on various aspects, ranging from the design and format of the items to their linguistic features and the nature of the content. A key idea in the theory is that translation errors are multidimensional (<xref ref-type="bibr" rid="ref49">Solano-Flores et al., 2009</xref>). For instance, a punctuation mistake might not only be a style error but also affect the meaning, making it a semantic error. Similarly, word-for-word translations and the use of syntactic structures uncommon in the target language fall under grammar and syntax errors. Depending on the content being assessed, word-for-word translations may also cause errors in semantics or construct dimensions, as they can modify the intended meaning or alter the content being evaluated.</p>
<p>The concept of multidimensionality is further linked to the idea of a trade-off between error dimensions: correcting or minimizing an error in one dimension may inadvertently introduce errors in another. For example, avoiding the use of classifiers to prevent additional information about an object&#x2019;s characteristics (as mentioned in the previous example) may result in increased grammatical complexity.</p>
<p>The theory suggests that translation errors are unavoidable because languages encode cultural experiences differently (<xref ref-type="bibr" rid="ref13">Greenfield, 1997</xref>) and rely on distinct semiotic systems (<xref ref-type="bibr" rid="ref4">Bezemer and Kress, 2008</xref>). Additionally, the trade-offs between the error dimensions mentioned above make the complete elimination of translation errors impossible, even with high-quality translations. For instance, the amount of space required for printed text differs across languages, leading to variations in how much space the text occupies on a page and how much blank space is available for students to write their responses.</p>
<p>Although some translation errors are inevitable, many of them are insignificant&#x2014;they may go unnoticed or are unlikely to affect the constructs being measured or the difficulty of the items. However, studies (<xref ref-type="bibr" rid="ref50">Solano-Flores et al., 2005</xref>, <xref ref-type="bibr" rid="ref51">2013</xref>) have demonstrated a stronger correlation between translation errors and item difficulty in cases where numerous or severe errors occur&#x2014;those that are likely to change the constructs or meanings of the items&#x2014;compared to items with minimal or minor translation errors.</p>
<p>Among studies investigating cross-cultural measurement invariance in PISA assessments, <xref ref-type="bibr" rid="ref41">Oliden and Lizaso (2013)</xref> reported that the PISA 2009 reading skills test displayed metric invariance across different languages. However, <xref ref-type="bibr" rid="ref52">S&#x00F6;yler Ba&#x011F;du (2020)</xref> found that the PISA 2015 reading skills test did not exhibit measurement invariance between native English-speaking and non-native English-speaking countries.</p>
<p><xref ref-type="bibr" rid="ref7">Ceyhan (2019)</xref> investigated the measurement invariance of the PISA 2012 reading skills test, focusing on comparisons involving the same language with different cultures, as well as different languages with different cultures. The study showed that structural invariance was achieved for comparisons using the same language, while only weak invariance was observed for different languages. Analyses of the PISA 2000 reading skills items showed fewer items exhibiting DIF when comparing countries using the same language, as opposed to within-country comparisons using different languages (<xref ref-type="bibr" rid="ref15">Grisay et al., 2009</xref>; <xref ref-type="bibr" rid="ref16">Grisay and Monseur, 2007</xref>). This suggests that language&#x2014;in other words, translation&#x2014;plays a critical role in DIF. Similarly, in the present study, certain items were found to display DIF when comparing different languages and cultures.</p>
</sec>
<sec id="sec4">
<label>2.2</label>
<title>Cultural difference and its impact</title>
<p>Culture is shaped through interactions with the environment, reflecting regional characteristics (<xref ref-type="bibr" rid="ref22">Jang, 2010</xref>). As a result, Eastern and Western cultures that developed during the same historical period exhibit significant differences.</p>
<p>One key example is creativity, which is influenced by the cultural context in which individuals are situated. According to <xref ref-type="bibr" rid="ref54">Sung (2006)</xref>, culture and creativity are inseparably linked, with individualism and liberalism in Western cultures contrasting with collectivism and Confucianism in Eastern cultures, leading to differences in the development of creative traits and environments. Research suggests that Western individuals tend to excel in divergent and creative thinking, while Eastern individuals are generally more reflective and intuitive in their approaches. Furthermore, Eastern cultures often take a holistic view, perceiving objects as interconnected entities before analyzing their structure. In contrast, Western cultures typically adopt a more analytical worldview, focusing on individual elements first and then synthesizing them into a whole. These cultural distinctions align with collectivism and community-oriented values in the East and individualism in the West, shaped by their respective regional environments.</p>
<p>In the realm of education, these cultural differences are also reflected in distinct educational practices. Eastern education often emphasizes a strong foundation in cultural heritage, focusing on the acquisition and deep understanding of historical knowledge. In contrast, Western education prioritizes independent thinking, encouraging students to cultivate a broad range of skills through practice and experience, with an emphasis on learning how to approach problems critically and develop individual thought processes (<xref ref-type="bibr" rid="ref27">Ko, 2013</xref>).</p>
<p>This divergence is particularly pronounced in language education. For instance, in South Korea, the teaching of the Korean language is imbued with a nationalistic ideology that connects individual development with national progress. This connection is reflected in the curriculum, where grammar and literature are combined into a single comprehensive subject. By contrast, in the United States, literature is treated as one component of reading instruction, indicating a different pedagogical focus (<xref ref-type="bibr" rid="ref31">Lee, 2013</xref>). These differences highlight how culture shapes educational goals and practices, which in turn influence students&#x2019; learning experiences and outcomes.</p>
<p>From an educational achievement perspective, achievement goal theory, which originates from Western literature, operates differently in East Asian societies due to cultural influences (<xref ref-type="bibr" rid="ref29">Lau and Lee, 2008</xref>; <xref ref-type="bibr" rid="ref30">Lau and Nie, 2008</xref>). Studies have revealed unique patterns in the achievement goals of students from CHCs, where education systems are highly competitive, and achievement holds significant value. For example, East Asian students show a strong positive correlation between mastery and performance goals (<xref ref-type="bibr" rid="ref19">Ho and Hau, 2008</xref>). Moreover, performance goals may contribute positively to adaptive learning and academic achievement (<xref ref-type="bibr" rid="ref45">Salili and Lai, 2003</xref>). Another notable difference is that East Asian students tend to adopt higher levels of avoidance goals compared to their Western counterparts (<xref ref-type="bibr" rid="ref63">Zusho and Clayton, 2011</xref>).</p>
<p>The nature of reading instruction in East Asian societies also warrants attention, as it differs markedly from Western instructional practices. Traditional reading classrooms in East Asia, shaped by CHC, are typically teacher-centered, characterized by large class sizes, a competitive atmosphere, directive teaching methods, and an emphasis on examination performance (<xref ref-type="bibr" rid="ref59">Watkins and Biggs, 2001</xref>). These instructional methods often seem at odds with the mastery-oriented and student-centered approaches advocated in Western educational systems, which aim to foster adaptive achievement goals.</p>
<p>However, global educational reforms have altered some traditional teaching methods in these societies (<xref ref-type="bibr" rid="ref28">Lau and Ho, 2015</xref>). To examine whether these reforms have influenced reading instruction and achievement goals&#x2014;factors shaped by cultural contexts&#x2014;this study investigates DIF based on CHC. Using the United States as the reference country, the analysis compares traditionally CHC-influenced countries, such as South Korea and Singapore, with Canada, which, like the U.S., does not have a CHC background. This approach indirectly evaluates CHC&#x2019;s impact on educational environments and outcomes, particularly in reading instruction and achievement goals (<xref ref-type="bibr" rid="ref42">Qian and Lau, 2022</xref>).</p>
</sec>
</sec>
<sec sec-type="materials|methods" id="sec5">
<label>3</label>
<title>Materials and methods</title>
<sec id="sec6">
<label>3.1</label>
<title>Sample</title>
<p>The Programme for International Student Assessment (PISA), conducted by the Organization for Economic Co-operation and Development (OECD), was selected as the research focus. Since 2000, PISA has been administered every three years to assess the mathematics, science, and reading skills of 15-year-old students receiving formal education. Additionally, PISA collects data on various educational variables through student, teacher, school, and parent questionnaires. As a large-scale international comparative study, PISA serves as a valuable tool for countries to evaluate their educational systems and environments and to inform improvements in their education policies and practices.</p>
<p>While the most recent PISA cycle took place in 2021, this study utilizes data from the 2018 PISA cycle, where reading was designated as the core domain. In PISA, the core domain is assessed in greater detail, comprising approximately half of the total testing time (<xref ref-type="bibr" rid="ref40">OECD, 2019c</xref>). One core domain is assessed for all students, while the other domains are treated as minor and are not administered to all students (<xref ref-type="bibr" rid="ref40">OECD, 2019c</xref>). This emphasis provides more comprehensive coverage and a larger sample size, enabling deeper analysis and more reliable insights into student performance across countries (<xref ref-type="bibr" rid="ref40">OECD, 2019c</xref>). Thus, the 2018 PISA reading assessment data were selected for this study.</p>
<p>Given that this study aimed to examine potential bias in the reading items of PISA 2018, the analysis was conducted using published items. Specifically, the research focused on responses to seven items from the unit coded as &#x201C;Rapa Nui&#x201D;, which were uniquely both publicly released and implemented in the PISA 2018 reading main survey simultaneously (<xref ref-type="bibr" rid="ref38">OECD, 2019a</xref>). <xref ref-type="table" rid="tab1">Table 1</xref> presents the distribution of the items by code, item type, cognitive process being measured, text source, text organization and navigation, text format, text type, and item difficulty. As shown in <xref ref-type="table" rid="tab1">Table 1</xref>, these items measure a wide range of cognitive characteristics.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Item characteristics of &#x201C;Rapa Nui&#x201D; unit.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Item code</th>
<th align="left" valign="top">Item type</th>
<th align="left" valign="top">Cognitive process subscale</th>
<th align="left" valign="top">Cognitive process</th>
<th align="left" valign="top">Source</th>
<th align="left" valign="top">Text organization and navigation</th>
<th align="left" valign="top">Text format</th>
<th align="left" valign="top">Text type</th>
<th align="left" valign="top">Difficulty</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">CR551Q01</td>
<td align="left" valign="top">Simple multiple choice</td>
<td align="left" valign="top">Locate information</td>
<td align="left" valign="top">Access and retrieve information within a text</td>
<td align="left" valign="top">Single</td>
<td align="left" valign="top">Dynamic</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Narration</td>
<td align="left" valign="top">Level 4</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q05</td>
<td align="left" valign="top">Open Response</td>
<td align="left" valign="top">Understand</td>
<td align="left" valign="top">Represent literal meaning</td>
<td align="left" valign="top">Single</td>
<td align="left" valign="top">Dynamic</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Narration</td>
<td align="left" valign="top">Level 3</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q06</td>
<td align="left" valign="top">Complex multiple choice</td>
<td align="left" valign="top">Evaluate and reflect</td>
<td align="left" valign="top">Reflect on content and form</td>
<td align="left" valign="top">Single</td>
<td align="left" valign="top">Static</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Argument</td>
<td align="left" valign="top">Level 5</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q08</td>
<td align="left" valign="top">Simple multiple choice</td>
<td align="left" valign="top">Locate information</td>
<td align="left" valign="top">Access and retrieve information within a text</td>
<td align="left" valign="top">Single</td>
<td align="left" valign="top">Static</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Argument</td>
<td align="left" valign="top">Level 5</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q09</td>
<td align="left" valign="top">Simple multiple choice</td>
<td align="left" valign="top">Evaluate and reflect</td>
<td align="left" valign="top">Detect and handle conflict</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Level 4</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q10</td>
<td align="left" valign="top">Complex multiple choice</td>
<td align="left" valign="top">Understand</td>
<td align="left" valign="top">Integrate and generate inferences across multiple sources</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Level 5</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q11</td>
<td align="left" valign="top">Open Response</td>
<td align="left" valign="top">Evaluate and reflect</td>
<td align="left" valign="top">Detect and handle conflict</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Continuous</td>
<td align="left" valign="top">Multiple</td>
<td align="left" valign="top">Level 4</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The analysis included data from the United States as the reference group, with Canada (sharing the similar written language and culture), Singapore (sharing the similar written language but a different culture), and South Korea (with a different language and culture) as comparison groups. The United States was selected as the reference group due to the widespread use of English, the most commonly spoken language, and its frequent designation as a target country in prior studies (<xref ref-type="bibr" rid="ref26">Khorramdel et al., 2020</xref>; <xref ref-type="bibr" rid="ref36">Muench et al., 2022</xref>; <xref ref-type="bibr" rid="ref44">Sachse et al., 2016</xref>). Canada shares Western cultural traits with the United States, with both English and French as the languages of assessment. Canada was selected as a comparison country due to its cultural and linguistic similarities with the United States. To ensure consistency in the language of assessment, only the sample of students who took the test in English was included in the analysis. Singapore, while sharing the same language of assessment (English), represents a distinctly different Eastern cultural context. South Korea, as an East Asian country, represents a different cultural context and uses Korean as the primary language, which is linguistically distinct from English. This selection of comparison groups is expected to provide clearer insights into the effects of linguistic and cultural factors on DIF.</p>
<p>Participants with any missing values were excluded, resulting in the removal of approximately 15% of cases from the cognitive item data. For the Rasch Tree analysis, explanatory variables were categorized into linguistic and cultural constructs and merged with the data based on student IDs (see <xref ref-type="table" rid="tab2">Table 2</xref>). Specifically, the variable &#x201C;Language&#x201D; was used as the linguistic explanatory variable, while &#x201C;Perceived Reading Instruction&#x201D; and &#x201C;Achievement Goals&#x201D; were classified as cultural explanatory variables. All cases with missing values in these explanatory variables were removed using listwise deletion (missing rate; United States: 3.7%, Canada: 6.6%, Singapore: 1.8%, South Korea: 2.0%). In addition, for item &#x201C;CR551Q06,&#x201D; scores coded as &#x201C;2&#x201D; were recoded to &#x201C;1,&#x201D; while partial scores coded as &#x201C;1&#x201D; were recoded to &#x201C;0.&#x201D; The sample sizes from each country were as follows: United States (reference group) with 678 students, Canada with 2,898 students, Singapore with 1,222 students, and South Korea with 1,197 students.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Explanatory variables for Rasch Tree (<xref ref-type="bibr" rid="ref39">OECD, 2019b</xref>).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Aspect</th>
<th align="left" valign="top">Variable</th>
<th align="left" valign="top">Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">Language: language of assessment</td>
<td align="left" valign="middle">LANGTEST_COG</td>
<td align="left" valign="middle">Language of assessment is the language utilized during the actual administration of the test. English was encoded as &#x201C;313,&#x201D; French as &#x201C;493,&#x201D; and Korean as &#x201C;301,&#x201D; which are categorical variable.</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="6">Culture: perceived reading instruction</td>
<td align="left" valign="middle">DISCLIMA</td>
<td align="left" valign="middle">The disciplinary climate in the test language classroom is assessed through five items on a four-point Likert scale with the categories &#x201C;Every lesson,&#x201D; &#x201C;Most lessons,&#x201D; &#x201C;Some lessons,&#x201D; and &#x201C;Never or hardly ever.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">TEACHSUP</td>
<td align="left" valign="middle">Teacher support is measured through four items on a four-point Likert scale with the categories &#x201C;Every lesson,&#x201D; &#x201C;Most lessons,&#x201D; &#x201C;Some lessons,&#x201D; and &#x201C;Never or hardly ever.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">DIRINS</td>
<td align="left" valign="middle">Teacher-directed instruction is evaluated using four reverse-coded items on a four-point Likert scale with the categories &#x201C;Every lesson,&#x201D; &#x201C;Most lessons,&#x201D; &#x201C;Some lessons,&#x201D; and &#x201C;Never or hardly ever.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">PERFEED</td>
<td align="left" valign="middle">Perceived teacher feedback is assessed through three items on a four-point Likert scale with the categories &#x201C;Never or almost never,&#x201D; &#x201C;Some lessons,&#x201D; &#x201C;Many lessons,&#x201D; and &#x201C;Every lesson or almost every lesson.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">STIMREAD</td>
<td align="left" valign="middle">Teacher stimulation of reading and teaching strategies is measured through four items on a four-point Likert scale with the categories &#x201C;Never or hardly ever,&#x201D; &#x201C;In some lessons,&#x201D; &#x201C;In most lessons,&#x201D; and &#x201C;In all lessons.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">ADAPTIVITY</td>
<td align="left" valign="middle">Instruction adaptivity in test language lessons is evaluated through three items with a four-point Likert scale with the categories &#x201C;Never or almost never,&#x201D; &#x201C;Some lessons,&#x201D; &#x201C;Many lessons,&#x201D; and &#x201C;Every lesson or almost every lesson.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle" rowspan="3">Culture: achievement goals</td>
<td align="left" valign="middle">COMPETE</td>
<td align="left" valign="middle">Competitiveness is assessed through three items on a four-point Likert scale with the categories &#x201C;Strongly disagree,&#x201D; &#x201C;Disagree,&#x201D; &#x201C;Agree,&#x201D; and &#x201C;Strongly agree.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">WORKMAST</td>
<td align="left" valign="middle">Working motive and mastery achievement motive are measured with three items on a four-point Likert scale with the categories &#x201C;Strongly disagree,&#x201D; &#x201C;Disagree,&#x201D; &#x201C;Agree,&#x201D; and &#x201C;Strongly agree.&#x201D;</td>
</tr>
<tr>
<td align="left" valign="middle">GFOFAIL</td>
<td align="left" valign="middle">General fear of failure is evaluated using three items on a four-point Likert scale with the categories &#x201C;Strongly disagree,&#x201D; &#x201C;Disagree,&#x201D; &#x201C;Agree,&#x201D; and &#x201C;Strongly agree.&#x201D;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Because model fit indices can be influenced by sample size (<xref ref-type="bibr" rid="ref21">Hu and Bentler, 1995</xref>; <xref ref-type="bibr" rid="ref12">Fan et al., 1999</xref>; <xref ref-type="bibr" rid="ref32">Lei and Lomax, 2005</xref>; <xref ref-type="bibr" rid="ref11">Fan and Sivo, 2007</xref>; <xref ref-type="bibr" rid="ref35">Mahler, 2011</xref>), the study aimed to use an equal number of students from each country. Therefore, because the United States had the smallest sample size (678), the same number of students was randomly selected from the other countries. Consequently, data from a total of 2,712 students were analyzed.</p>
<p>For the Rasch Tree analysis, when comparing Canada to the United States, the variables &#x201C;TEACHSUP,&#x201D; &#x201C;DIRINS,&#x201D; and &#x201C;PERFEED&#x201D; were excluded, as these items were not administered in Canada. Additionally, the variable &#x201C;language of assessment&#x201D; was processed using one-hot encoding for comparisons involving different languages, such as Korea vs. the U.S. The variable &#x201C;LANGTEST_COG&#x201D; was dummy-coded to create new variables indicating the use of English (LANGTEST_COGEnglish), and Korean (LANGTEST_COGKorean). For example, when comparing Korea and the U.S., since the test languages are only Korean and English, the value for LANGTEST_COGEnglish is coded as 1, and LANGTEST_COGKorean is coded as 0 for those who took the test in English. However, for the United States&#x2013;Canada and United States&#x2013;Singapore comparisons, since all participants took the test in the same language (English), the variable was not included as it does not carry meaningful variance for analysis.</p>
</sec>
<sec id="sec7">
<label>3.2</label>
<title>Statistical analysis and interpretation</title>
<p>Using the complete dataset, IRT-LR was employed with the IRTLRDIF program (<xref ref-type="bibr" rid="ref57">Thissen, 2001</xref>) using the 3-parameter IRT model, and we used the &#x201C;difR&#x201D; package for logistic regression and the &#x201C;psychotree&#x201D; package for the Rasch Tree analyses in R.</p>
<p>IRT-LR, based on IRT, compares two different models: a compact model, which assumes no DIF, and an augmented model, which assumes DIF is possible in the item under study. For the IRT-LR analysis using the IRTLRDIF program, the likelihood ratio test statistic, <inline-formula>
<mml:math id="M1">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>, was employed. The formula for <inline-formula>
<mml:math id="M2">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>, as outlined by <xref ref-type="bibr" rid="ref57">Thissen (2001)</xref>, is as follows:<disp-formula id="E1">
<mml:math id="M3">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi mathvariant="italic">df</mml:mi>
<mml:mo stretchy="true">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>log</mml:mo>
<mml:mspace width="0.25em"/>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>log</mml:mo>
<mml:mspace width="0.25em"/>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo stretchy="true">)</mml:mo>
</mml:math>
</disp-formula></p>
<p>where <inline-formula>
<mml:math id="M4">
<mml:mi mathvariant="italic">df</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>p</mml:mi>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math id="M5">
<mml:mi>p</mml:mi>
</mml:math>
</inline-formula> represents the number of parameters. <inline-formula>
<mml:math id="M6">
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:math>
</inline-formula> is the compact model, and <inline-formula>
<mml:math id="M7">
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
</mml:math>
</inline-formula> is the augmented model. <inline-formula>
<mml:math id="M8">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi mathvariant="italic">df</mml:mi>
<mml:mo stretchy="true">)</mml:mo>
</mml:math>
</inline-formula> follows a chi-square distribution with degrees of freedom equal to the difference in the number of parameters between the augmented and compact models (<xref ref-type="bibr" rid="ref57">Thissen, 2001</xref>). Since the three-parameter model was applied, DIF is considered present when <inline-formula>
<mml:math id="M9">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula> exceeds 7.81, with 3 degrees of freedom (<inline-formula>
<mml:math id="M10">
<mml:mi mathvariant="italic">df</mml:mi>
</mml:math>
</inline-formula>) (<xref ref-type="bibr" rid="ref8">Choi et al., 2015</xref>). The formula for the three-parameter model is as follows:<disp-formula id="E2">
<mml:math id="M11">
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>&#x03B8;</mml:mi>
<mml:mo stretchy="true">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="true">)</mml:mo>
<mml:mo>&#x22C5;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mspace width="0.33em"/>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mi mathvariant="normal">e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x03B1;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>&#x03B8;</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x03B2;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="true">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:math>
</disp-formula></p>
<p>The second DIF technique used in this research was logistic regression (LR). LR assesses the effect of multiple independent variables on a binary outcome, determining which of two categories a subject belongs to. The logistic regression models can be described as follows:</p>
<p>Model 1 (full model): logit(<italic>p</italic>) = <inline-formula>
<mml:math id="M12">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mi>&#x039B;</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>&#x0393;</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>&#x039B;</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>&#x0393;</mml:mi>
<mml:mo stretchy="true">)</mml:mo>
</mml:math>
</inline-formula></p>
<p>Model 2 (1st reduced model): logit(<italic>p</italic>) = <inline-formula>
<mml:math id="M13">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mi>&#x039B;</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>&#x0393;</mml:mi>
</mml:math>
</inline-formula></p>
<p>Model 3 (2nd reduced model): logit(<italic>p</italic>) = <inline-formula>
<mml:math id="M14">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mi>&#x039B;</mml:mi>
</mml:math>
</inline-formula></p>
<p>where <italic>p</italic> represents the probability of answering the item correctly, <inline-formula>
<mml:math id="M15">
<mml:mi>&#x039B;</mml:mi>
</mml:math>
</inline-formula> is a measure of an individual&#x2019;s ability (e.g., IRT ability parameters (<inline-formula>
<mml:math id="M16">
<mml:mi>&#x03B8;</mml:mi>
</mml:math>
</inline-formula>) or total scores), <inline-formula>
<mml:math id="M17">
<mml:mi>&#x0393;</mml:mi>
</mml:math>
</inline-formula> is a categorical predictor variable indicating group membership for an individual (where <inline-formula>
<mml:math id="M18">
<mml:mi>&#x0393;</mml:mi>
</mml:math>
</inline-formula> = 1 for members of the focal group and <inline-formula>
<mml:math id="M19">
<mml:mi>&#x0393;</mml:mi>
</mml:math>
</inline-formula> = 0 for members of the reference group), <inline-formula>
<mml:math id="M20">
<mml:mo stretchy="true">(</mml:mo>
<mml:mi>&#x039B;</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>&#x0393;</mml:mi>
<mml:mo stretchy="true">)</mml:mo>
</mml:math>
</inline-formula> is the interaction of a person&#x2019;s ability and their group membership.</p>
<p>The term <inline-formula>
<mml:math id="M21">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> represents the main effect of a person&#x2019;s ability on their performance on the item. <inline-formula>
<mml:math id="M22">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> reflects the difference in intercepts between the focal and reference groups, which indicates uniform DIF when statistically significant. That is, uniform DIF occurs when one group consistently performs better or worse than the other across all levels of ability. <inline-formula>
<mml:math id="M23">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> captures the interaction between ability and group membership (i.e., whether the relationship between ability and item performance differs by group), and a significant <inline-formula>
<mml:math id="M24">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> implies the presence of non-uniform DIF. While <inline-formula>
<mml:math id="M25">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M26">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> are sometimes interpreted as approximating group-specific differences in intercepts (<inline-formula>
<mml:math id="M27">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> = <inline-formula>
<mml:math id="M28">
<mml:msub>
<mml:mi>&#x03B2;</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x03B2;</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>) and slopes (<inline-formula>
<mml:math id="M29">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>&#x03B2;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x03B2;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>), these are regression-based estimates and should not be directly equated with IRT-based item parameters. If the null hypothesis H<sub>0</sub>: <inline-formula>
<mml:math id="M30">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> = 0 (comparison between Model 1 and Model 2) is rejected, this suggests a significant interaction between ability and group, indicating non-uniform DIF; in this case, the DIF testing procedure concludes. If H<sub>0</sub>: <inline-formula>
<mml:math id="M31">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> = 0 is not rejected, the subsequent comparison between Model 2 and Model 3 tests H<sub>0</sub>: <inline-formula>
<mml:math id="M32">
<mml:msub>
<mml:mi>&#x03C4;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:math>
</inline-formula> = 0, and a significant result indicates uniform DIF (<xref ref-type="bibr" rid="ref55">Swaminathan and Rogers, 1990</xref>). This stepwise approach enables a clear distinction between the two types of DIF.</p>
<p>The statistical testing of LR DIF analysis was conducted by analyzing the difference in model fit between the two nested models using the chi-square statistic (<inline-formula>
<mml:math id="M33">
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>) (<xref ref-type="bibr" rid="ref47">Scott et al., 2010</xref>; <xref ref-type="bibr" rid="ref48">Sohn, 2010</xref>). The LRT (Likelihood Ratio Test) statistics evaluate DIF by comparing the fit of two nested models, while the Wald statistics assess model parameters using an appropriate contrast matrix (<xref ref-type="bibr" rid="ref25">Johnson and Wichern, 1998</xref>). Since LRT statistics is the default setting in the R package (<xref ref-type="bibr" rid="ref34">Magis et al., 2010</xref>), it was utilized in this study.</p>
<p>LR DIF analysis can also be viewed as a weighted least squares approach. The contributions of explanatory variables are reflected in the change in the coefficient of determination (<inline-formula>
<mml:math id="M34">
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>) between the augmented and compact models. This change is computed as:<disp-formula id="E3">
<mml:math id="M35">
<mml:mi>&#x0394;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mi>R</mml:mi>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:math>
</disp-formula>where <inline-formula>
<mml:math id="M36">
<mml:msubsup>
<mml:mi>R</mml:mi>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:math>
</inline-formula> represents the value for the augmented model and <inline-formula>
<mml:math id="M37">
<mml:msubsup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:math>
</inline-formula> represents the value for the compact model. The difference in <inline-formula>
<mml:math id="M38">
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula> illustrates the additional explanatory power provided by the variables in the augmented model. LR was selected for its ability to detect both uniform and non-uniform DIF (<xref ref-type="bibr" rid="ref61">Zumbo, 1999</xref>), making it a robust method.</p>
<p>The DIF level was decided according to the &#x0394;<italic>R</italic><sup>2</sup> value obtained from employing the LR technique. According to <xref ref-type="bibr" rid="ref24">Jodoin and Gierl (2001)</xref>, the &#x0394;<italic>R</italic><sup>2</sup> value is interpreted as follows: 0 &#x003C; &#x0394;<italic>R</italic><sup>2</sup> &#x003C; 0.035, no or negligible DIF; 0.035 &#x2264; &#x0394;<italic>R</italic><sup>2</sup> &#x003C;0.07, moderate DIF; &#x0394;<italic>R</italic><sup>2</sup> &#x2265; 0.07, high DIF. According to another source, &#x0394;<italic>R</italic><sup>2</sup> &#x003C; 0.13 indicates no or negligible DIF; 0.13 &#x2264; &#x0394;<italic>R</italic><sup>2</sup> &#x003C; 0.26, and &#x0394;<italic>R</italic><sup>2</sup> &#x2265; 0.26 indicate moderate and high DIF, respectively (<xref ref-type="bibr" rid="ref62">Zumbo and Thomas, 1996</xref>). In this study, <xref ref-type="bibr" rid="ref24">Jodoin and Gierl&#x2019;s (2001)</xref> criterion was to determine DIF. Items with a level of effect classified as B or higher in the logistic regression results were identified as exhibiting DIF.</p>
<p>To account for the increased risk of Type I errors due to multiple hypothesis testings, a Bonferroni correction was applied. This method adjusts the significance threshold by dividing the desired alpha level by the number of hypotheses conducted (<xref ref-type="bibr" rid="ref20">Holland and Thayer, 1986</xref>). In this study, seven items were analyzed, resulting in an adjusted significance level of 0.05/7&#x202F;=&#x202F;0.007 for each item. This correction was uniformly applied to both the IRT-LR and LR analyses to ensure consistency across methods. In the context of IRT-LR, the critical value for detecting DIF was determined as 12.11 based on the chi-square distribution with 3 degrees of freedom at the adjusted alpha level. For LR, which is based on a chi-square distribution with 1 degree of freedom, the corresponding critical value was approximately 7.27.</p>
<p>Nevertheless, Type I error can occur during the DIF detection process and may have significant impact on the analysis. However, according to the study by <xref ref-type="bibr" rid="ref2">Atar and Kamata (2011)</xref>, the simulation conditions considered three sample sizes (600, 1,200, 2,400) and two group sample size ratios (1:1 and 1:2). Regarding Type I error control, their findings indicated that the Type I error rates of both LRT (Likelihood Ratio Test) and LRP (Logistic Regression Procedure) were well controlled at clearly defined significance levels across all simulation conditions. However, previous studies have reported that when there is a difference in ability between groups, Type I error tends to increase (<xref ref-type="bibr" rid="ref10">DeMars, 2009</xref>; <xref ref-type="bibr" rid="ref33">Li et al., 2012</xref>; <xref ref-type="bibr" rid="ref37">Narayanan and Swaminathan, 1996</xref>). Nevertheless, as shown in <xref ref-type="table" rid="tab3">Table 3</xref>, all four countries belong to the highest or high-performing group in PISA reading. This suggests that the ability differences between groups are not substantial, indicating that Type I error is relatively well controlled in this study. Furthermore, using <inline-formula>
<mml:math id="M47">
<mml:mi>&#x0394;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula> along with the chi-square test is advantageous for controlling Type I error (<xref ref-type="bibr" rid="ref24">Jodoin and Gierl, 2001</xref>). Therefore, this study employed both <inline-formula>
<mml:math id="M48">
<mml:mi>&#x0394;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula> and the chi-square test to analyze DIF to reduce the likelihood of Type I error.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Average reading achievement among United States, Canada, Singapore, and South Korea (<xref ref-type="bibr" rid="ref38">OECD, 2019a</xref>).</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Countries</th>
<th align="center" valign="top">United States</th>
<th align="center" valign="top">Canada</th>
<th align="center" valign="top">Singapore</th>
<th align="center" valign="top">South Korea</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">Average reading score</td>
<td align="center" valign="middle">505</td>
<td align="center" valign="middle">520</td>
<td align="center" valign="middle">549</td>
<td align="center" valign="middle">514</td>
</tr>
<tr>
<td align="left" valign="middle">Rank among 81 countries (excluding Spain)</td>
<td align="center" valign="middle">13</td>
<td align="center" valign="middle">6</td>
<td align="center" valign="middle">2</td>
<td align="center" valign="middle">9</td>
</tr>
<tr>
<td align="left" valign="middle">Average reading score among OECD countries</td>
<td align="center" valign="middle" colspan="4">487</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, the Rasch Tree method (<xref ref-type="bibr" rid="ref53">Strobl et al., 2015</xref>) integrates logistic regression with recursive partitioning to identify subgroups based on response patterns and explanatory variables. This approach allows for the data-driven formation of subgroups without relying on arbitrary thresholds set by researchers (<xref ref-type="bibr" rid="ref23">Jang and Lee, 2023</xref>), making it a suitable method for DIF analysis. However, this method has certain limitations, particularly in terms of interpretative complexity due to its data-driven nature. Unlike traditional DIF detection methods that rely on predefined groups, Rasch Tree iteratively identifies subgroups based on statistical splits, which may lead to challenges in result interpretation and theoretical alignment (<xref ref-type="bibr" rid="ref53">Strobl et al., 2015</xref>). To address these limitations, this study incorporates traditional DIF detection methods, specifically the IRT-LR and LR, to enhance the robustness and interpretability of the analysis.</p>
<p>For interpreting the Rasch Tree results, item difficulty estimates were used. Higher item difficulty values indicate more difficult items, and groups with higher item difficulty estimates are considered disadvantaged in relation to those specific items. The Rasch Tree analysis examines differences in item difficulty parameters between groups under the null hypothesis that there is no difference in item difficulty parameters (<xref ref-type="bibr" rid="ref6">Camilli and Shepard, 1994</xref>). The threshold for defining DIF based on item difficulty was calculated using the Mantel&#x2013;Haenszel (MH) effect size from the Rasch Tree model (<xref ref-type="bibr" rid="ref18">Henninger et al., 2023</xref>; <xref ref-type="bibr" rid="ref20">Holland and Thayer, 1986</xref>; <xref ref-type="bibr" rid="ref43">Roussos et al., 1999</xref>). In this case, &#x0394;<italic>MH</italic> was used as an indicator using the MH odds ratio to evaluate whether items function differentially between groups.<disp-formula id="E4">
<mml:math id="M50">
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2.35</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mo stretchy="true">(</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi mathvariant="italic">iR</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mtext mathvariant="italic">iF</mml:mtext>
</mml:msub>
<mml:mo stretchy="true">)</mml:mo>
</mml:math>
</disp-formula></p>
<p><inline-formula>
<mml:math id="M51">
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi mathvariant="italic">iR</mml:mi>
</mml:msub>
</mml:math>
</inline-formula> = difficulty parameter for item i in the reference group; <inline-formula>
<mml:math id="M52">
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mtext mathvariant="italic">iF</mml:mtext>
</mml:msub>
</mml:math>
</inline-formula> = difficulty parameter for item i in the focal group.</p>
<p>Item difficulty estimates for each subgroup from the Rasch Tree results were used to calculate &#x0394;<italic>MH</italic>, and items were classified as A, B, or C based on the ETS classification system for the Mantel&#x2013;Haenszel effect size (<xref ref-type="bibr" rid="ref18">Henninger et al., 2023</xref>). Items classified as A, B, and C indicate no, medium, and large DIF, respectively. The stopping criterion for Rasch Tree analysis occurs when all items are classified as A or at least one item is classified as B. Specifically, if the absolute difference in item difficulty between groups exceeds 0.426(=1/2.35) but is less than 0.638(=1.5/2.35), the item is classified as B, indicating moderate DIF. If the absolute difference exceeds 0.638, the item is classified as C, indicating a large DIF effect. The item difficulty estimates for each comparison and the corresponding Mantel&#x2013;Haenszel effect size classifications are presented in <xref ref-type="table" rid="tab4">Table 4</xref>. Additionally, items with a level of effect classified as B or higher in the Rasch Tree results were identified as exhibiting DIF.</p>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption>
<p>ETS classification scheme for the Mantel&#x2013;Haenszel odds ratio in the <inline-formula>
<mml:math id="M54">
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
</mml:math>
</inline-formula>.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Class</th>
<th align="left" valign="top">Interpretation</th>
<th align="left" valign="top">Classification rule</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="middle">A</td>
<td align="left" valign="middle">Negligible DIF</td>
<td align="center" valign="middle">
<inline-formula>
<mml:math id="M55">
<mml:mn>0</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left" valign="middle">B</td>
<td align="left" valign="middle">Medium (moderate) DIF</td>
<td align="center" valign="middle">
<inline-formula>
<mml:math id="M56">
<mml:mn>1</mml:mn>
<mml:mo>&#x003C;</mml:mo>
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mo>&#x003C;</mml:mo>
<mml:mn>1.5</mml:mn>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left" valign="middle">C</td>
<td align="left" valign="middle">Large DIF</td>
<td align="center" valign="middle">
<inline-formula>
<mml:math id="M57">
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>1.5</mml:mn>
</mml:math>
</inline-formula>
</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="results" id="sec8">
<label>4</label>
<title>Results</title>
<sec id="sec9">
<label>4.1</label>
<title>DIF analysis with IRT-LR and LR</title>
<p>As shown in <xref ref-type="table" rid="tab5">Table 5</xref>, the comparison between the United States and Canada revealed that item CR551Q06 exhibited a DIF effect based on the IRT-LR and LR techniques. DIF had to be identified by at least one of the DIF detection methods for an item to be classified as a DIF item. In the case of the LR analyses, only items with at least a B level of effect were considered as exhibiting DIF. Item CR551Q06 showed negligible DIF based on LR techniques but were identified as displaying DIF in the IRT-LR analysis. Item CR551Q06 favored the United States in both the IRT-LR and LR analyses.</p>
<table-wrap position="float" id="tab5">
<label>Table 5</label>
<caption>
<p>A comparison between the USA, Canada, Singapore, and South Korea using IRT-LR &#x0026; logistic regression.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top" rowspan="3">Item code</th>
<th align="center" valign="top" colspan="6">Canada vs. USA</th>
<th align="center" valign="top" colspan="6">Singapore vs. USA</th>
<th align="center" valign="top" colspan="6">South Korea vs. USA</th>
</tr>
<tr>
<th align="center" valign="top" colspan="3">IRT-LR</th>
<th align="center" valign="top" colspan="3">Logistic regression</th>
<th align="center" valign="top" colspan="3">IRT-LR</th>
<th align="center" valign="top" colspan="3">Logistic regression</th>
<th align="center" valign="top" colspan="3">IRT-LR</th>
<th align="center" valign="top" colspan="3">Logistic regression</th>
</tr>
<tr>
<th align="center" valign="top">
<inline-formula>
<mml:math id="M58">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>
</th>
<th align="center" valign="top">DIF presence</th>
<th align="center" valign="top">Favored group</th>
<th align="center" valign="top">LRT <inline-formula>
<mml:math id="M59">
<mml:msup>
<mml:mi>&#x03C7;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula></th>
<th align="center" valign="top">Level of effect</th>
<th align="center" valign="top">Favored group</th>
<th align="center" valign="top">
<inline-formula>
<mml:math id="M60">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>
</th>
<th align="center" valign="top">DIF presence</th>
<th align="center" valign="top">Favored group</th>
<th align="center" valign="top">LRT <inline-formula>
<mml:math id="M61">
<mml:msup>
<mml:mi>&#x03C7;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula></th>
<th align="center" valign="top">Level of effect</th>
<th align="center" valign="top">Favored group</th>
<th align="center" valign="top">
<inline-formula>
<mml:math id="M62">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>
</th>
<th align="center" valign="top">DIF presence</th>
<th align="center" valign="top">Favored group</th>
<th align="center" valign="top">LRT <inline-formula>
<mml:math id="M63">
<mml:msup>
<mml:mi>&#x03C7;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula></th>
<th align="center" valign="top">Level of effect</th>
<th align="center" valign="top">Favored group</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">CR551Q01</td>
<td align="center" valign="top">10.4</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">10.84&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">Canada</td>
<td align="center" valign="top">1.7</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">5.07</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">55.0</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">Korea</td>
<td align="center" valign="top">8.82</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">-</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q05</td>
<td align="center" valign="top">2.9</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">2.17</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">4.0</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">2.11</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">26.0</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">Korea</td>
<td align="center" valign="top">119.37&#x002A;</td>
<td align="center" valign="top">C</td>
<td align="center" valign="top">Korea</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q06</td>
<td align="center" valign="top">21.5</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">23.13&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">77.1</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">105.82&#x002A;</td>
<td align="center" valign="top">B</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">466.5</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">616.56&#x002A;</td>
<td align="center" valign="top">C</td>
<td align="center" valign="top">USA</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q08</td>
<td align="center" valign="top">0.1</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">0.78</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">8.5</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">9.77&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">46.8</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">8.49</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">-</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q09</td>
<td align="center" valign="top">0.0</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">0.60</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">8.8</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">25.49&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">Singapore</td>
<td align="center" valign="top">14.6</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">11.61</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">-</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q10</td>
<td align="center" valign="top">0.5</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">0.0</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">13.92&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">Singapore</td>
<td align="center" valign="top">20.6</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">Korea&#x002A;&#x002A;</td>
<td align="center" valign="top">9.93&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">Korea</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q11</td>
<td align="center" valign="top">0.1</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">0.47</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">4.7</td>
<td align="center" valign="top">X</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">2.66</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">-</td>
<td align="center" valign="top">18.0</td>
<td align="center" valign="top">O</td>
<td align="center" valign="top">USA</td>
<td align="center" valign="top">18.89&#x002A;</td>
<td align="center" valign="top">A</td>
<td align="center" valign="top">Korea</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x002A;IRT-LR: DIF is present if <inline-formula>
<mml:math id="M64">
<mml:msup>
<mml:mi>G</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula> exceeds 12.11(<inline-formula>
<mml:math id="M65">
<mml:mi mathvariant="italic">df</mml:mi>
</mml:math>
</inline-formula> = 3) (Bonferroni correction) (<xref ref-type="bibr" rid="ref8">Choi et al., 2015</xref>), Logistic regression: &#x201C;A&#x201D;: negligible DIF (<inline-formula>
<mml:math id="M66">
<mml:mi>&#x0394;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x003C;</mml:mo>
<mml:mn>0.035</mml:mn>
</mml:math>
</inline-formula>), &#x201C;B&#x201D;: moderate DIF (<inline-formula>
<mml:math id="M67">
<mml:mn>0.035</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>&#x0394;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x003C;</mml:mo>
<mml:mn>.07</mml:mn>
</mml:math>
</inline-formula>), &#x201C;C&#x201D;: high DIF (<inline-formula>
<mml:math id="M68">
<mml:mn>0.07</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>&#x0394;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:math>
</inline-formula>) (<xref ref-type="bibr" rid="ref24">Jodoin and Gierl, 2001</xref>). Values marked with &#x002A; represent the Likelihood Ratio Test (LRT) chi-square values for uniform DIF, while unmarked values indicate non-uniform DIF. The IRTLRDIF program does not differentiate between uniform and nonuniform DIF; therefore, this study does not report them separately. Values marked with &#x002A;&#x002A; correspond to items for which the discrimination parameters were considerably higher in Korean group than in the U.S. group, according to the DIF analysis.</p>
</table-wrap-foot>
</table-wrap>
<p>In the comparison between the United States and Singapore, item CR551Q06 showed a DIF effect. Item CR551Q06 was detected as having DIF based on the IRT-LR technique and a moderate effect according to the LR technique. Conversely, items CR551Q08, CR551Q09, and CR551Q10 displayed negligible DIF based on the LR method and were not classified as DIF items according to the IRT-LR analysis. Item CR551Q06 favored the United States in both the IRT-LR and LR analyses and was classified as exhibiting uniform DIF in the LR analysis.</p>
<p>In the comparison between the United States and South Korea, all items were confirmed to exhibit a DIF effect. IRT-LR analysis classified all items as exhibiting DIF. LR analysis identified that items CR551Q05 and CR551Q06 demonstrated a C level of effect, indicating high DIF. In contrast, CR551Q01, CR551Q08, CR551Q09, CR551Q10, and CR551Q11 showed an A level of effect, suggesting negligible DIF. LR analysis identified CR551Q01, CR551Q08, and CR551Q09 as exhibiting non-uniform DIF, while CR551Q05, CR551Q06, CR551Q10, and CR551Q11 exhibited uniform DIF. According to the IRT-LR analysis, items CR551Q01 and CR551Q05 were found to favor Korea, while items CR551Q06, CR551Q08, CR551Q09, and CR551Q011 favored the United States. For item CR551Q010, the difficulty parameter (b) reported by the IRT-LR was 0.98 for both groups, making it difficult to determine which group the item favored. Based on LR analysis, among the items showing uniform DIF, items CR551Q05, CR551Q010, and CR551Q011 favored Korea, while item CR551Q06 favored the United States. In summary, both IRT-LR and LR analyses consistently indicated that item CR551Q05 favored Korea and item CR551Q06 favored the United States. Therefore, all items showed DIF when comparing these two countries with distinct written languages and cultures.</p>
<p>Based on the comparisons between the United States, Canada, Singapore, and South Korea, item CR551Q06 consistently displayed DIF across all comparisons. Additionally, item CR551Q06 showed moderate to high DIF (at least level B for LR analysis) in both the United States-Singapore and United States-South Korea. Furthermore, item CR551Q06 was consistently classified as a uniform DIF item favoring the United States across all three comparisons: United States&#x2013;Canada, United States&#x2013;Singapore, and United States&#x2013;South Korea. In contrast, item CR551Q05 was not identified as exhibiting DIF in either the IRT-LR or LR analyses for the United States&#x2013;Canada and United States&#x2013;Singapore comparisons. However, in the comparison with South Korea, it was detected as a DIF item in both methods. Notably, the LR analysis indicated a C level of effect, suggesting high DIF, and both IRT-LR and LR analyses consistently classified it as a uniform DIF item favoring Korea. The release of the Rapa Nui unit highlights the need to investigate which specific components of these items contribute to DIF.</p>
<p>Upon reviewing the characteristics of the items, no significant common features were immediately apparent. However, when the items were analyzed in terms of the cognitive processes required to solve them in each language, notable differences emerged. Specifically, item CR551Q06 requires &#x201C;reflecting on content and form,&#x201D; a relatively complex cognitive process. In contrast, item CR551Q05 involves &#x201C;representing literal meaning,&#x201D; which is considered a lower-level cognitive process. These findings highlight the need for a thorough review to determine whether differences in required cognitive processes contribute to the occurrence of DIF, and whether the importance or difficulty of these cognitive skills varies across the countries examined.</p>
<p>Based on the frequency and severity of DIF identified, it appears most significant in South Korea then Singapore, followed by Canada. This pattern corresponds to the degree of linguistic and cultural differences from the United States, which was used as the reference country. These findings suggest that linguistic and cultural disparities are positively associated with DIF.</p>
</sec>
<sec id="sec10">
<label>4.2</label>
<title>DIF analysis with Rasch Tree</title>
<p>The results of the DIF analysis for the seven items from the Rapa Nui unit of the PISA 2018 Reading Assessment, represented in a tree structure using Rasch Tree analysis, are shown in <xref ref-type="fig" rid="fig1">Figure 1</xref>. All linguistic and cultural variables listed in <xref ref-type="table" rid="tab2">Table 2</xref> were incorporated into the Rasch Tree model as candidate splitting variables. However, for the United States&#x2013;Canada and United States&#x2013;Singapore comparisons, language-related variables were excluded from the model because all participants took the test in English, and thus the language of assessment lacked variability, while cultural variables were retained.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Rasch Tree by comparison between the USA, Canada, Singapore, and South Korea. For the United States&#x2013;Canada and United States&#x2013;Singapore comparisons, the &#x201C;language of assessment&#x201D; variable was excluded due to the lack of variance, as all participants completed the assessment in English. Accordingly, only the two cultural variables&#x2014;Perceived Reading Instruction and Achievement Goals&#x2014;were included in these analyses, as shown in <xref ref-type="table" rid="tab2">Table 2</xref>. In contrast, the United States&#x2013;South Korea comparison included all linguistic and cultural variables listed in <xref ref-type="table" rid="tab2">Table 2</xref>.</p>
</caption>
<graphic xlink:href="feduc-10-1595658-g001.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Three line graphs compare data across Canada and USA, Singapore and USA, and South Korea and USA. Each has nodes with sample sizes, showing trends across the x-axis. South Korea and USA includes a decision tree splitting data at a significance level of p &#x003C; 0.001.</alt-text>
</graphic>
</fig>
<p>The nodes at the bottom of the figure represent the partitions within the decision tree model, each corresponding to a specific subgroup in the data. Initially, a Rasch Tree analysis was conducted using the academic performance data of students from the United States and Canada. All cultural variables listed in <xref ref-type="table" rid="tab2">Table 2</xref> were included in the Rasch Tree model as candidate splitting variables. However, no splits were observed, which is an expected outcome given the cultural similarities between the two countries. Moreover, this finding aligns with the IRT-LR and LR results, in which only item CR551Q06 was identified as exhibiting DIF.</p>
<p>Next, a Rasch Tree analysis was conducted for the United States and Singapore. The results showed no branches split, indicating the absence of significant DIF items and any meaningful influence from cultural factors. Although Singapore is culturally distinct from the United States, the lack of subgroup splits in the Rasch Tree analysis suggests that cultural differences are less likely to influence DIF in the context of international test translation.</p>
<p>Subsequently, we analyzed the reading assessment data of students from the United States and South Korea, which was the only comparison that resulted in a node split in the Rasch Tree analysis. Two subgroups were formed based on the language of assessment, as indicated by the &#x201C;LANGTEST_COGEnglish&#x201D; dummy variable: those using English and those using Korean. The left subgroup (node 2), where the &#x201C;LANGTEST_COGEnglish&#x201D; variable has a value of less than or equal to 0, represents the group assessed in Korean, while the right subgroup (node 3), where the variable has a value greater than 0, represents the group assessed in English. This reflects the fact that, while the language of assessment in the United States is exclusively English, in South Korea, the test is administered in Korean. Even though the subgroups were not predefined by country, the analysis revealed a clear division between the United States and South Korea, indicating a particularly strong DIF effect when comparing these two countries. Furthermore, despite including various cultural variables as background factors, the analysis confirmed that linguistic differences, rather than cultural differences, contribute to the presence of DIF in these items.</p>
<p>As shown in <xref ref-type="table" rid="tab6">Table 6</xref>, the effect size of the test was analyzed using <inline-formula>
<mml:math id="M69">
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
</mml:math>
</inline-formula>, and item CR551Q09 was classified as level B, indicating moderate DIF, while items CR551Q01, CR551Q05, CR551Q06, and CR551Q08 were classified as level C, indicating a large DIF. In particular, item CR551Q06 had an absolute <inline-formula>
<mml:math id="M70">
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
</mml:math>
</inline-formula> value of 7.8, and items CR551Q01 and CR551Q05 had absolute values of 3.3&#x202F;~&#x202F;3.8, showing a significant difference that led to their classification as DIF items. The fact that such items with large DIF values were identified in the subgroups divided by Korean and English suggests that a thorough review of the translation process from English to Korean is necessary. In this case, items CR551Q01, CR551Q05, CR551Q09, CR551Q10, and CR551Q11 favored the group that took the test in Korean (node 2) compared to the group that took the test in English (node 3). Conversely, items CR551Q06 and CR551Q08 favored the group that took the test in English (node 3) compared to the group that took the test in Korean (node 2).</p>
<table-wrap position="float" id="tab6">
<label>Table 6</label>
<caption>
<p>A comparison between the USA, Canada, Singapore, and South Korea using Rasch Tree.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Comparison</th>
<th align="center" valign="top">Canada and USA</th>
<th align="center" valign="top">Singapore and USA</th>
<th align="center" valign="top" colspan="6">South Korea and USA</th>
</tr>
<tr>
<th align="left" valign="middle" rowspan="2">Item code</th>
<th align="center" valign="middle">Item difficulty estimates</th>
<th align="center" valign="middle">Item difficulty estimates</th>
<th align="center" valign="top" colspan="6">Item difficulty estimates</th>
</tr>
<tr>
<th align="center" valign="middle">Node 1</th>
<th align="center" valign="middle">Node 1</th>
<th align="center" valign="middle">Node 2</th>
<th align="center" valign="middle">Node 3</th>
<th align="center" valign="top">Favored Group</th>
<th align="center" valign="middle">Difference</th>
<th align="center" valign="middle">
<inline-formula>
<mml:math id="M71">
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
</mml:math>
</inline-formula>
</th>
<th align="center" valign="top">Level of effect</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">CR551Q01</td>
<td align="center" valign="middle">0.1</td>
<td align="center" valign="middle">0.1</td>
<td align="center" valign="middle">&#x2212;1.1</td>
<td align="center" valign="middle">0.3</td>
<td align="center" valign="top"><bold>Korea</bold></td>
<td align="center" valign="middle"><bold>1.4</bold></td>
<td align="center" valign="middle"><bold>3.3</bold></td>
<td align="center" valign="top"><bold>C</bold></td>
</tr>
<tr>
<td align="left" valign="top">CR551Q05</td>
<td align="center" valign="middle">&#x2212;0.4</td>
<td align="center" valign="middle">&#x2212;0.4</td>
<td align="center" valign="middle">&#x2212;1.9</td>
<td align="center" valign="middle">&#x2212;0.3</td>
<td align="center" valign="top"><bold>Korea</bold></td>
<td align="center" valign="middle"><bold>1.6</bold></td>
<td align="center" valign="middle"><bold>3.8</bold></td>
<td align="center" valign="top"><bold>C</bold></td>
</tr>
<tr>
<td align="left" valign="top">CR551Q06</td>
<td align="center" valign="middle">&#x2212;0.5</td>
<td align="center" valign="middle">&#x2212;0.1</td>
<td align="center" valign="middle">2.6</td>
<td align="center" valign="middle">&#x2212;0.7</td>
<td align="center" valign="top"><bold>USA</bold></td>
<td align="center" valign="middle"><bold>3.3</bold></td>
<td align="center" valign="middle"><bold>7.8</bold></td>
<td align="center" valign="top"><bold>C</bold></td>
</tr>
<tr>
<td align="left" valign="top">CR551Q08</td>
<td align="center" valign="middle">&#x2212;0.2</td>
<td align="center" valign="middle">0.0</td>
<td align="center" valign="middle">0.6</td>
<td align="center" valign="middle">&#x2212;0.2</td>
<td align="center" valign="top"><bold>USA</bold></td>
<td align="center" valign="middle"><bold>0.8</bold></td>
<td align="center" valign="middle"><bold>1.9</bold></td>
<td align="center" valign="top"><bold>C</bold></td>
</tr>
<tr>
<td align="left" valign="top">CR551Q09</td>
<td align="center" valign="middle">0.2</td>
<td align="center" valign="middle">&#x2212;0.1</td>
<td align="center" valign="middle">&#x2212;0.3</td>
<td align="center" valign="middle">0.2</td>
<td align="center" valign="top"><bold>Korea</bold></td>
<td align="center" valign="middle"><bold>0.5</bold></td>
<td align="center" valign="middle"><bold>1.2</bold></td>
<td align="center" valign="top"><bold>B</bold></td>
</tr>
<tr>
<td align="left" valign="top">CR551Q10</td>
<td align="center" valign="middle">1.1</td>
<td align="center" valign="middle">0.8</td>
<td align="center" valign="middle">0.9</td>
<td align="center" valign="middle">1.1</td>
<td align="center" valign="top">Korea</td>
<td align="center" valign="middle">0.2</td>
<td align="center" valign="middle">0.5</td>
<td align="center" valign="top">A</td>
</tr>
<tr>
<td align="left" valign="top">CR551Q11</td>
<td align="center" valign="middle">&#x2212;0.3</td>
<td align="center" valign="middle">&#x2212;0.4</td>
<td align="center" valign="middle">&#x2212;0.8</td>
<td align="center" valign="middle">&#x2212;0.3</td>
<td align="center" valign="top"><bold>Korea</bold></td>
<td align="center" valign="middle"><bold>0.5</bold></td>
<td align="center" valign="middle"><bold>1.2</bold></td>
<td align="center" valign="top"><bold>B</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Items classified as B or C level, based on differences in item difficulty estimates between subgroups, are highlighted in bold. Rasch Tree: &#x201C;A&#x201D;: negligible DIF (<inline-formula>
<mml:math id="M72">
<mml:mn>0</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>), &#x201C;B&#x201D;: medium DIF (<inline-formula>
<mml:math id="M73">
<mml:mn>1</mml:mn>
<mml:mo>&#x003C;</mml:mo>
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mo>&#x003C;</mml:mo>
<mml:mn>1.5</mml:mn>
</mml:math>
</inline-formula>), &#x201C;C": large DIF (<inline-formula>
<mml:math id="M74">
<mml:mo>&#x2223;</mml:mo>
<mml:mi mathvariant="italic">&#x0394;MH</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>1.5</mml:mn>
</mml:math>
</inline-formula>) (<xref ref-type="bibr" rid="ref18">Henninger et al., 2023</xref>). For the United States&#x2013;Canada and United States&#x2013;Singapore comparisons, the &#x201C;language of assessment&#x201D; variable was excluded due to the lack of variance, as all participants completed the assessment in English. Accordingly, only the two cultural variables&#x2014;Perceived Reading Instruction and Achievement Goals&#x2014;were included in these analyses, as shown in <xref ref-type="table" rid="tab2">Table 2</xref>. In contrast, the United States&#x2013;South Korea comparison included all linguistic and cultural variables listed in <xref ref-type="table" rid="tab2">Table 2</xref>.</p>
</table-wrap-foot>
</table-wrap>
<p>In addition to the DIF results shown in <xref ref-type="table" rid="tab6">Table 6</xref>, we observed a notable pattern in the item difficulty estimates that warranted further exploration. While the main Rasch Tree analyses, as presented in <xref ref-type="table" rid="tab6">Table 6</xref>, did not include the country variable &#x201C;CNT&#x201D; and focused solely on linguistic and cultural covariates, no node splits were observed for the United States&#x2013;Canada and United States&#x2013;Singapore comparisons. As a result, item difficulty estimates in both cases were derived from the entire combined sample without country-level differentiation. Interestingly, the resulting item difficulty estimates from these two comparisons were similar. To supplement these findings and better understand the basis of this similarity, we conducted an additional Rasch Tree analysis that included the country variable &#x201C;CNT&#x201D; as a covariate. When the CNT variable was included, both the U.S.&#x2013;Canada and U.S.&#x2013;Singapore comparisons resulted in a binary split based solely on country, allowing for the estimation of item difficulties separately for each group. Although these estimates were not used for the main analysis, they are presented in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1</xref> to support a more comprehensive understanding of group-level patterns.</p>
</sec>
<sec id="sec11">
<label>4.3</label>
<title>Comparative analysis of IRT-LR, LR, and Rasch Tree results</title>
<p>The results of comparing and analyzing the outcomes of the IRT-LR, LR, and Rasch Tree analyses are as follows.</p>
<p>First, the comparison between the United States and Canada yielded divergent results across the three DIF detection methods. While both IRT-LR and LR identified item CR551Q06 as exhibiting DIF, the Rasch Tree analysis revealed no node splits, indicating no evidence of DIF. Given the strong cultural and linguistic alignment between the two countries, it is plausible that this discrepancy arises from the nature of subgroup specification. Traditional approaches like IRT-LR and LR rely on predefined groups&#x2014;typically based on nationality&#x2014;which may amplify even minor differences. In contrast, the Rasch Tree method identifies subgroups through a data-driven process without imposing prior group definitions. These findings imply that the DIF observed in item CR551Q06 may reflect the analytical structure of the traditional methods rather than genuine linguistic or cultural sources. Further investigation is warranted to clarify the origins of this DIF.</p>
<p>In addition, the supplementary analysis presented in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1</xref>&#x2014;conducted by including the country variable &#x201C;CNT&#x201D; in the Rasch Tree model&#x2014;produced results consistent with those in <xref ref-type="table" rid="tab5">Table 5</xref>, again identifying item CR551Q06 as exhibiting DIF. This finding confirms that when DIF is analyzed based on predefined country groups such as the United States and Canada, the same item tends to be detected as DIF regardless of the analytical method employed.</p>
<p>Next, a similar discrepancy was observed in the United States&#x2013;Singapore comparison. Item CR551Q06 was again identified as a DIF item by IRT-LR and LR, whereas the Rasch Tree analysis showed no evidence of subgroup splits. Unlike Canada, Singapore has a markedly different cultural background, despite sharing English as the test language. Although differences between Eastern and Western cultures exist, these differences are unlikely to be substantial enough to induce DIF. Traditional methods may have captured differences based on predefined national groups rather than cultural factors, such as perceived reading instruction and achievement goals, which were specifically examined in this study. This pattern reinforces the notion that DIF detection may be sensitive to the method&#x2019;s reliance on pre-established group structures.</p>
<p>In addition, the supplementary analysis presented in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1</xref>&#x2014;conducted by including the country variable &#x201C;CNT&#x201D; in the Rasch Tree model&#x2014;yielded results consistent with those in <xref ref-type="table" rid="tab5">Table 5</xref>, identifying items CR551Q06, CR551Q09, and CR551Q10 as exhibiting DIF. This consistency suggests that when DIF is analyzed based on predefined national groups such as the United States and Singapore, the same items tend to be flagged as DIF regardless of the analytical method employed.</p>
<p>Lastly, the comparison between the United States and South Korea shows partially consistent results. While all items were classified as DIF items using IRT-LR and LR, the Rasch Tree method identified CR551Q01, CR551Q05, CR551Q06, CR551Q08, CR551Q09, and CR551Q11 as DIF items. In the case of the Rasch Tree analysis, only items with at least a B level of effect were considered as exhibiting DIF. In the comparison between the United States and South Korea, all items except CR551Q10 were consistently identified as DIF items in at least two of the three analyses.</p>
<p>An analysis of the characteristics of the six items revealed the following: three items were simple multiple-choice, two was open-response, and one was complex multiple-choice item. In terms of cognitive processes, two items assessed students&#x2019; ability to access and retrieve information within a text, two items focused on detecting and handling conflict, one item assessed representing literal meaning, and one item assessed reflecting on content and form. Regarding the cognitive process subscale, three items focused on evaluating and reflecting, two items on locating information, and one on understanding. For text organization and navigation, two items were classified as dynamic, two as static, and two as multiple. All items had a continuous text format. The text types included two narrative items, two argumentative items, and two multiple type items. In terms of item difficulty, three items were classified as level 4, two as level 5, and one as level 3 according to the PISA proficiency scale, which ranges from below level 1 to level 6. Based on this classification, the items primarily represent high difficulty levels, with level 3 indicating moderate difficulty and levels 4 and 5 representing more challenging tasks. Although the six items were classified as DIF items in the analyses comparing the United States and South Korea, no distinct commonalities were revealed after analyzing their characteristics.</p>
<p>At least two of the three methods consistently identified item CR551Q06 (<xref ref-type="fig" rid="fig2">Figure 2</xref>) as DIF item when comparing the United States with two or more countries. All analyses indicated that item CR551Q06 favored the United States. In the comparison with South Korea, both items CR551Q05 and CR551Q06 exhibited large DIF effect sizes across all methods, with item 5 consistently favoring Korea (<xref ref-type="fig" rid="fig2">Figures 2</xref>, <xref ref-type="fig" rid="fig3">3</xref>). CR551Q06 is a complex multiple-choice item that targets the cognitive process of reflecting on content and form, categorized under the &#x201C;evaluate and reflect&#x201D; subscale. It relies on a single text source, features static organization and navigation, uses a continuous text format, and is of the argumentative type with a difficulty level of 5, which indicates a high level of complexity. Additionally, CR551Q05 is an open-response item that targets the cognitive process of representing literal meaning, categorized under the &#x201C;understand&#x201D; subscale. It relies on a single text source, features dynamic organization and navigation, uses a continuous text format, and is of the narrative type. The item has a difficulty level of 3, which indicates a moderate level of complexity. Further investigation, including a detailed review of the released items, is necessary to understand why this item consistently exhibited DIF across all analyses, and particularly in the United States&#x2013;South Korea comparison. Such analysis could help minimize DIF in future updates of international academic assessments.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Released item &#x201C;CR551Q06&#x201D; from Rapa Nui unit (Reproduced from <xref ref-type="bibr" rid="ref38">OECD, 2019a</xref>, &#x00A9; OECD 2019).</p>
</caption>
<graphic xlink:href="feduc-10-1595658-g002.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">A digital book review interface from "PISA 2018" displays a review of Jared Diamond's book "Collapse" on the right. The text highlights the environmental destruction on Rapa Nui and its consequences. On the left, a table asks users to categorize statements as fact or opinion, including points about the moai statues, European landings, and a recommendation of the book.</alt-text>
</graphic>
</fig>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption>
<p>Released item &#x201C;CR551Q05&#x201D; from Rapa Nui unit (Reproduced from <xref ref-type="bibr" rid="ref38">OECD, 2019a</xref>, &#x00A9; OECD 2019).</p>
</caption>
<graphic xlink:href="feduc-10-1595658-g003.tif" mimetype="image" mime-subtype="tiff">
<alt-text content-type="machine-generated">Interface of a digital classroom exercise from PISA 2018 related to Rapa Nui. The left panel has a question: "Refer to the Professor's Blog on the right. Type your answer to the question. In the last paragraph of the blog, the professor writes: 'Another mystery remained&#x2026;' To what mystery does she refer?" The right panel displays a professor's blog post from May 23, 11:22 a.m., describing the landscape of Rapa Nui, featuring grassy areas, blue skies, and volcanic backgrounds, accompanied by an image of moai statues.</alt-text>
</graphic>
</fig>
<p>Overall, there are notable similarities between the Rasch Tree results and those obtained using the IRT-LR and LR methods. Specifically, when considering both the presence and effect size of DIF in the IRT-LR and LR analyses, the frequency of DIF follows the order: South Korea &#x003E; Singapore &#x003E; Canada. This finding suggests that greater linguistic and cultural differences between countries are associated with a higher likelihood of DIF.</p>
<p>The Rasch Tree results further confirmed that among linguistic and cultural variables, the only significant factor influencing DIF occurrence was linguistic differences. This finding challenges the preconceived notion that perceived reading instruction and achievement goals differ significantly between Eastern and Western cultures. It may also imply that recent educational reforms in CHC cultures have substantially impacted educational culture, leading to outcomes that resemble those in the West.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="sec12">
<label>5</label>
<title>Conclusion</title>
<sec id="sec13">
<label>5.1</label>
<title>Discussion</title>
<p>This study applied traditional DIF detection methods, including IRT-LR and LR, as well as the newly emerging Rasch Tree method, to explore DIF and analyze its contributing factors. The analysis focused on the Rapa Nui Unit, consisting of seven items from the reading domain of PISA 2018. The reference group in the analysis was the United States, while Canada, Singapore, and South Korea were comparison groups.</p>
<p>This study applied traditional DIF detection methods, IRT-LR and LR, alongside the Rasch Tree method, to explore DIF across comparisons between the United States and Canada, Singapore, and South Korea. In the comparison between the United States and Canada, item CR551Q06 was identified as DIF item in the IRT-LR and LR analyses. However, the Rasch Tree analysis showed that no node splits, suggesting that the absence of substantial cultural differences between the two countries contributes to the low likelihood of DIF occurrence.</p>
<p>In the comparison between the United States and Singapore, a notable discrepancy was observed across methods. While IRT-LR and LR identified item CR551Q06 as exhibiting DIF, the Rasch Tree method found no significant subgroup splits or DIF items. The absence of node splits&#x2014;despite the inclusion of cultural background variables&#x2014;suggests that cultural differences alone are unlikely to lead to the occurrence of DIF.</p>
<p>This absence of node splits in the Rasch Tree analysis highlights the potential dependence of traditional methods on predefined subgroups. As supplementary evidence, when the Rasch Tree analysis was conducted with the country variable&#x2014;representing predefined national groups&#x2014;explicitly included, node splits emerged in both the U.S.&#x2013;Canada and U.S.&#x2013;Singapore comparisons. The resulting DIF patterns closely aligned with those identified by the IRT-LR and LR analyses. These findings suggest that when DIF is analyzed based on predefined country groups, the same items are flagged for DIF across methods.</p>
<p>When comparing the United States and South Korea, all items under consideration were classified as DIF by IRT-LR and LR, while the Rasch Tree method identified six items&#x2014;CR551Q01, CR551Q05, CR551Q06, CR551Q08, CR551Q09, and CR551Q11&#x2014;as exhibiting DIF. This was the only comparison where the Rasch Tree analysis revealed node splits, corresponding to differences in the language of assessment. These findings indicate that linguistic factors had a substantial impact on the observed DIF, partially consistent with the results from traditional methods.</p>
<p>Given the significant role of language in inducing DIF, greater attention should be paid to the accurate and culturally sensitive translation of test items into each country&#x2019;s target language. Translation effects have also been identified as a major source of bias, as prior studies have demonstrated their significant impact on DIF (<xref ref-type="bibr" rid="ref16">Grisay and Monseur, 2007</xref>; <xref ref-type="bibr" rid="ref15">Grisay et al., 2009</xref>; <xref ref-type="bibr" rid="ref41">Oliden and Lizaso, 2013</xref>; <xref ref-type="bibr" rid="ref50">Solano-Flores et al., 2005</xref>, <xref ref-type="bibr" rid="ref51">2013</xref>). These findings underscore the importance of careful and systematic adaptation procedures in cross-cultural assessments to reduce DIF. Potential sources of item bias must be considered during the item writing and adaptation phases. If necessary, specialized training for item writers and translators should be provided for this purpose. Based on the findings of this study, future research should be conducted using data from diverse cultural and linguistic contexts, employing multiple DIF detection techniques to validate and extend the current results.</p>
<p>These results also challenge the commonly held assumption that substantial differences in reading instruction and achievement goals exist between Eastern and Western cultures. They suggest that recent educational reforms in East Asia may have aligned instructional practices more closely with Western standards.</p>
<p>In addition, certain items repeatedly exhibited DIF across multiple comparisons. In particular, item CR551Q06 consistently showed DIF, with all analyses indicating that it favored the United States. In the comparison with South Korea, both items CR551Q05 and CR551Q06 demonstrated large DIF effect sizes across all detection methods, with CR551Q05 consistently favoring Korea. The consistent advantage of CR551Q06 for U.S. students across all comparisons warrants further investigation to verify and understand the underlying cause. Furthermore, in the U.S.&#x2013;South Korea comparison, the emergence of a relatively large number of DIF items and the magnitude of their effect sizes call for careful review and interpretation.</p>
<p>Moreover, future DIF research should incorporate both methods that require predefined groups and those that do not. In this study, IRT-LR and LR&#x2014;which necessitate prior group definitions&#x2014;were used alongside the Rasch Tree method, which operates without such assumptions. When IRT-LR and LR were applied using country-based groupings, DIF was detected in the comparisons between the United States and Canada, as well as between the United States and Singapore. In contrast, using the same dataset without country-based grouping, the Rasch Tree method did not identify any DIF. This contrast suggests that predefined groupings by country may lead to attributing DIF to prominent factors like language or culture, even though the true source of DIF may stem from other factors. Therefore, future research should incorporate methods like the Rasch Tree to more accurately uncover the underlying mechanisms that contribute to DIF, beyond predefined group structures.</p>
<p>Beyond methodological considerations, it is also important to reflect on the broader educational implications of how DIF is handled in international assessments. A related consideration concerns the inherent tension between cultural neutrality and task engagement. In efforts to develop culturally comparable assessments, item developers may be inclined to minimize or eliminate cultural references from tasks. However, doing so risks producing tasks that are overly generic, potentially diminishing students&#x2019; motivation and engagement. Tasks that completely avoid cultural context may fail to reflect authentic language use or meaningful scenarios, which are essential for assessing higher-order reading skills. This presents a dilemma for international assessment developers, as it may be more practical to accept a manageable level of DIF than to sacrifice the relevance and richness of the tasks. Acknowledging this trade-off is important when designing items that aim to be both culturally fair and pedagogically valuable.</p>
<p>Taken together, the findings from this study not only offer methodological insights but also highlight important practical considerations for international assessments. Finally, the present study contributes valuable insights into the potential cultural and linguistic sources of DIF, particularly by identifying items that may systematically favor or disadvantage specific groups. These insights can inform item development processes for future assessments. Specifically, the findings suggest that certain item features&#x2014;such as cultural references and linguistic complexity&#x2014;may introduce differential functioning that undermines cross-national comparability. Hence, test developers should carefully consider such factors during item construction, translation, and adaptation phases to ensure greater fairness and construct validity.</p>
</sec>
<sec id="sec14">
<label>5.2</label>
<title>Future research and limitations</title>
<p>Future research should explore how linguistic and cultural factors influence the occurrence of DIF items in the PISA 2022 Reading Core stage in comparison to the 2018 assessment. This study selected the 2018 data based on the fact that the core domain in PISA 2018 was reading. However, as linguistic and cultural influences evolve, it is crucial to examine which specific factors contribute to DIF detection in the modern context. Investigating these temporal changes will provide valuable insights into the shifting dynamics of language and culture and their impact on item functioning, helping to refine future assessments and reduce potential biases. However, because an individualized test application was used in PISA 2018, no students took identical tests, allowing for analysis across different items. For this reason, DIF and item bias studies should also be conducted across different item clusters.</p>
<p>While this study focused on a single reading unit (Rapa Nui) due to the constraints of PISA&#x2019;s multiple matrix sampling design, this narrow scope may limit the generalizability of the findings. In practice, large-scale assessments such as PISA aim to capture broad constructs using diverse item sets across multiple units. As such, results derived from a single unit should be interpreted with caution, particularly regarding their applicability to the entire reading construct. To address this limitation and enhance the practical utility of DIF analyses, future research should replicate this study&#x2019;s approach across a wider range of units and domains. Doing so will help validate the observed patterns and provide more robust evidence for improving the fairness and interpretability of international large-scale assessments.</p>
<p>While the Rasch Tree analysis found no cultural variables influencing DIF, this may be due to the exclusion of practical cultural factors as explanatory variables. This study did include all available background variables; rather, it focused on only achievement goals and perceived reading instruction, which have been reported to differ significantly between Eastern and Western cultures and are believed to influence reading achievement (<xref ref-type="bibr" rid="ref42">Qian and Lau, 2022</xref>). However, other cultural factors could also significantly impact DIF occurrence. Therefore, further research is needed to identify and incorporate additional practical cultural factors as explanatory variables to better assess their impact on DIF detection.</p>
<p>Finally, further analysis is needed for DIF items commonly identified across the comparisons of the United States, Canada, Singapore, and South Korea using IRT-LR, LR, and Rasch Tree analyses. In this study, DIF items were briefly analyzed based on item characteristics provided by the OECD. Beyond these basic characteristics, a detailed analysis of the released items is required. If the primary factor influencing DIF, as suggested by the Rasch Tree analysis, is linguistic, it is essential to assess whether the translation process for each country&#x2019;s language was appropriate and how the items were actually translated. Furthermore, consideration must be given to the characteristics of items that may introduce bias during translation.</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="sec15">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link xlink:href="https://www.oecd.org/en/data/datasets/pisa-2022-database.html#data" ext-link-type="uri">https://www.oecd.org/en/data/datasets/pisa-2022-database.html#data</ext-link>.</p>
</sec>
<sec sec-type="ethics-statement" id="sec16">
<title>Ethics statement</title>
<p>Ethical approval was not required for the study involving humans in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was not required from the participants or the participants&#x2019; legal guardians/next of kin in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec sec-type="author-contributions" id="sec17">
<title>Author contributions</title>
<p>YW: Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing &#x2013; original draft, Writing &#x2013; review &#x0026; editing. Y-JC: Conceptualization, Data curation, Methodology, Project administration, Software, Supervision, Validation, Writing &#x2013; review &#x0026; editing.</p>
</sec>
<sec sec-type="funding-information" id="sec18">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research and/or publication of this article.</p>
</sec>
<ack>
<p>This paper is an expanded version of the poster presented at the International Meeting of the Psychometric Society 2024.</p>
</ack>
<sec sec-type="COI-statement" id="sec19">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="sec20">
<title>Generative AI statement</title>
<p>The author(s) declare that Gen AI was used in the creation of this manuscript. The authors acknowledge the use of GPT-4o (OpenAI API, 2025 version) for language refinement and manuscript editing. The AI was not utilized for generating original scientific content, conducting data analysis, or formulating research interpretations. All AI-assisted edits were meticulously reviewed and manually validated by the authors to ensure factual accuracy, coherence, and adherence to ethical standards. Any necessary modifications were made accordingly. The initial and final prompts used in AI-assisted editing have been included in the supplementary materials for full transparency.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p>
</sec>
<sec sec-type="disclaimer" id="sec21">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="sec22">
<title>Supplementary material</title>
<p>The Supplementary material for this article can be found online at: <ext-link xlink:href="https://www.frontiersin.org/articles/10.3389/feduc.2025.1595658/full#supplementary-material" ext-link-type="uri">https://www.frontiersin.org/articles/10.3389/feduc.2025.1595658/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Aikhenvald</surname><given-names>A.</given-names></name></person-group> (<year>2003</year>). <source>Classifiers: A typology of noun categorization devices</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atar</surname><given-names>B.</given-names></name> <name><surname>Kamata</surname><given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Comparison of IRT likelihood ratio test and logistic regression DIF detection procedures</article-title>. <source>H. U. J. Educ.</source> <volume>41</volume>, <fpage>36</fpage>&#x2013;<lpage>47</lpage>.</citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bakan Kalayc&#x0131;o&#x011F;lu</surname><given-names>D.</given-names></name> <name><surname>Berbero&#x011F;lu</surname><given-names>G.</given-names></name></person-group> (<year>2010</year>). <article-title>Differential item functioning analysis of the science and mathematics items in the university entrance examinations in Turkey</article-title>. <source>J. Psychoeduc. Assess.</source> <volume>20</volume>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0734282910391623</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bezemer</surname><given-names>J.</given-names></name> <name><surname>Kress</surname><given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Writing in multimodal texts: a social semiotic account of designs for learning</article-title>. <source>Written Commun.</source> <volume>25</volume>, <fpage>166</fpage>&#x2013;<lpage>195</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0741088307313177</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>B&#x00F6;ckenholt</surname><given-names>U.</given-names></name></person-group> (<year>2012</year>). <article-title>Modeling multiple response processes in judgment and choice</article-title>. <source>Psychol. Methods</source> <volume>17</volume>, <fpage>665</fpage>&#x2013;<lpage>678</lpage>. doi: <pub-id pub-id-type="doi">10.1037/a0028111</pub-id>, PMID: <pub-id pub-id-type="pmid">22545594</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Camilli</surname><given-names>G.</given-names></name> <name><surname>Shepard</surname><given-names>L. A.</given-names></name></person-group> (<year>1994</year>). <source>Methods for identifying biased test items</source>, vol. <volume>4</volume>: <publisher-name>Sage</publisher-name>.</citation></ref>
<ref id="ref7"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ceyhan</surname><given-names>E.</given-names></name></person-group> (<year>2019</year>). <source>Assessing measurement invariance of PISA 2012 reading literacy scale among the countries determined in accordance with the language of application (master&#x2019;s thesis)</source>. <publisher-loc>Antalya</publisher-loc>: <publisher-name>Akdeniz University</publisher-name>.</citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Choi</surname><given-names>Y.-J.</given-names></name> <name><surname>Alexeev</surname><given-names>N.</given-names></name> <name><surname>Cohen</surname><given-names>A. S.</given-names></name></person-group> (<year>2015</year>). <article-title>DIF analysis using a mixture 3PL model with a covariate on the TIMSS 2007 mathematics test</article-title>. <source>Int. J. Test.</source> <volume>15</volume>, <fpage>239</fpage>&#x2013;<lpage>253</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305058.2015.1007241</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coleman</surname><given-names>J. S.</given-names></name></person-group> (<year>1968</year>). <article-title>Equality of educational opportunity</article-title>. <source>Equity Excell. Educ.</source> <volume>6</volume>, <fpage>19</fpage>&#x2013;<lpage>28</lpage>. doi: <pub-id pub-id-type="doi">10.1080/0020486680060504</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DeMars</surname><given-names>C. E.</given-names></name></person-group> (<year>2009</year>). <article-title>Modification of the mantel-haenszel and logistic regression DIF procedures to incorporate the SIBTEST regression correction</article-title>. <source>J. Educ. Behav. Stat.</source> <volume>34</volume>, <fpage>149</fpage>&#x2013;<lpage>170</lpage>. doi: <pub-id pub-id-type="doi">10.3102/1076998608329515</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname><given-names>X.</given-names></name> <name><surname>Sivo</surname><given-names>S. A.</given-names></name></person-group> (<year>2007</year>). <article-title>Sensitivity of fit indices to model misspecification and model types</article-title>. <source>Multivar. Behav. Res.</source> <volume>42</volume>, <fpage>509</fpage>&#x2013;<lpage>529</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00273170701382864</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname><given-names>X.</given-names></name> <name><surname>Thompson</surname><given-names>B.</given-names></name> <name><surname>Wang</surname><given-names>L.</given-names></name></person-group> (<year>1999</year>). <article-title>Effects of sample size, estimation methods, and model specification on structural equation modeling fit indexes</article-title>. <source>Struct. Equ. Modeling</source> <volume>6</volume>, <fpage>56</fpage>&#x2013;<lpage>83</lpage>. doi: <pub-id pub-id-type="doi">10.1080/10705519909540119</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Greenfield</surname><given-names>P. M.</given-names></name></person-group> (<year>1997</year>). <article-title>You can&#x2019;t take it with you: why ability assessments don&#x2019;t cross cultures</article-title>. <source>Am. Psychol.</source> <volume>52</volume>, <fpage>1115</fpage>&#x2013;<lpage>1124</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0003-066X.52.10.1115</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Grisay</surname><given-names>A.</given-names></name></person-group> (<year>2007</year>). The challenge of adapting PISA materials into non Indo-European languages: Some evidence from a brief of exploration of language issues in Chinese and Arabic. Available online at: <ext-link xlink:href="http://www.aspe.ulg.ac.be/grisay/fichiers/PISA07.pdf" ext-link-type="uri">http://www.aspe.ulg.ac.be/grisay/fichiers/PISA07.pdf</ext-link></citation></ref>
<ref id="ref15"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Grisay</surname><given-names>A.</given-names></name> <name><surname>Gonzalez</surname><given-names>E.</given-names></name> <name><surname>Monseur</surname><given-names>C.</given-names></name></person-group> (<year>2009</year>). &#x201C;<article-title>Equivalence of item difficulties across national versions of the PIRLS and PISA reading assessments</article-title>,&#x201D; in <source>IERI monograph series: Issues and methodologies in large-scale assessments</source>. eds. <person-group person-group-type="editor"><name><surname>Scheuermann</surname><given-names>F.</given-names></name> <name><surname>Bj&#x00F6;rnsson</surname><given-names>J.</given-names></name></person-group> (<publisher-loc>Hamburg, Germany</publisher-loc>: <publisher-name>IEA-ETS Research Institute</publisher-name>), <volume>2</volume>, <fpage>63</fpage>&#x2013;<lpage>83</lpage>.</citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grisay</surname><given-names>A.</given-names></name> <name><surname>Monseur</surname><given-names>C.</given-names></name></person-group> (<year>2007</year>). <article-title>Measuring the equivalence of item difficulty in the various versions of an international test</article-title>. <source>Stud. Educ. Eval.</source> <volume>33</volume>, <fpage>69</fpage>&#x2013;<lpage>86</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.stueduc.2007.01.006</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hambleton</surname><given-names>R. K.</given-names></name></person-group> (<year>2006</year>). <article-title>Good practices for identifying differential item functioning</article-title>. <source>Med. Care</source> <volume>44</volume>, <fpage>S182</fpage>&#x2013;<lpage>S188</lpage>. doi: <pub-id pub-id-type="doi">10.1097/01.mlr.0000245443.86671.c4</pub-id>, PMID: <pub-id pub-id-type="pmid">17060826</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henninger</surname><given-names>M.</given-names></name> <name><surname>Debelak</surname><given-names>R.</given-names></name> <name><surname>Strobl</surname><given-names>C.</given-names></name></person-group> (<year>2023</year>). <article-title>A new stopping criterion for Rasch trees based on the mantel-Haenszel effect size measure for differential item functioning</article-title>. <source>Educ. Psychol. Meas.</source> <volume>83</volume>, <fpage>181</fpage>&#x2013;<lpage>212</lpage>. doi: <pub-id pub-id-type="doi">10.1177/00131644221077135</pub-id>, PMID: <pub-id pub-id-type="pmid">36601252</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ho</surname><given-names>I. T.</given-names></name> <name><surname>Hau</surname><given-names>K. T.</given-names></name></person-group> (<year>2008</year>). <article-title>Academic achievement in the Chinese context: the role of goals, strategies, and effort</article-title>. <source>Int. J. Psychol.</source> <volume>43</volume>, <fpage>892</fpage>&#x2013;<lpage>897</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00207590701836323</pub-id>, PMID: <pub-id pub-id-type="pmid">22022794</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holland</surname><given-names>P. W.</given-names></name> <name><surname>Thayer</surname><given-names>D. T.</given-names></name></person-group> (<year>1986</year>). <article-title>Differential item functioning and the mantel-Haenszel procedure</article-title>. <source>ETS Res. Rep. Series</source> <volume>1986</volume>, <fpage>i</fpage>&#x2013;<lpage>24</lpage>. doi: <pub-id pub-id-type="doi">10.1002/j.2330-8516.1986.tb00186.x</pub-id>, PMID: <pub-id pub-id-type="pmid">40808747</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Hu</surname><given-names>L.</given-names></name> <name><surname>Bentler</surname><given-names>P.</given-names></name></person-group> (<year>1995</year>). &#x201C;<article-title>Evaluating model fit</article-title>&#x201D; in <source>Structural equation modeling: concepts, issues and application</source>. ed. <person-group person-group-type="editor"><name><surname>Hoyle</surname><given-names>R.</given-names></name></person-group> (<publisher-loc>Thousand Oasks</publisher-loc>: <publisher-name>Sage Publications</publisher-name>), <fpage>76</fpage>&#x2013;<lpage>99</lpage>.</citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jang</surname><given-names>K. B.</given-names></name></person-group> (<year>2010</year>). <article-title>A guideline for contents and method of cultural education based on the definition and characteristics of culture</article-title>. <source>Korean J. Cult. Arts Educ. Stud.</source> <volume>5</volume>, <fpage>19</fpage>&#x2013;<lpage>37</lpage>. doi: <pub-id pub-id-type="doi">10.15815/kjcaes.2010.5.2.19</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jang</surname><given-names>Y. S.</given-names></name> <name><surname>Lee</surname><given-names>J. Y.</given-names></name></person-group> (<year>2023</year>). <article-title>Exploring differential item functioning in PISA 2015 science test with the Rasch-tree</article-title>. <source>Korean Soc. Educ. Eval.</source> <volume>36</volume>, <fpage>83</fpage>&#x2013;<lpage>110</lpage>. doi: <pub-id pub-id-type="doi">10.31158/JEEV.2023.36.1.83</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jodoin</surname><given-names>M. G.</given-names></name> <name><surname>Gierl</surname><given-names>M. J.</given-names></name></person-group> (<year>2001</year>). <article-title>Evaluating type I error and power rates using an effect size measure with the logistic regression procedure for DIF detection</article-title>. <source>Appl. Meas. Educ.</source> <volume>14</volume>, <fpage>329</fpage>&#x2013;<lpage>349</lpage>. doi: <pub-id pub-id-type="doi">10.1207/S15324818AME1404_2</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Johnson</surname><given-names>R. A.</given-names></name> <name><surname>Wichern</surname><given-names>D. W.</given-names></name></person-group> (<year>1998</year>). <source>Applied multivariate statistical analysis</source> (4th ed.). <publisher-name>Upper Saddle River, NJ</publisher-name>: <publisher-name>Prentice-Hall</publisher-name>.</citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khorramdel</surname><given-names>L.</given-names></name> <name><surname>Pokropek</surname><given-names>A.</given-names></name> <name><surname>Joo</surname><given-names>S.-H.</given-names></name> <name><surname>Kirsch</surname><given-names>I.</given-names></name> <name><surname>Halderman</surname><given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>Examining gender DIF and gender differences in the PISA 2018 reading literacy scale: a partial invariance approach</article-title>. <source>Psychol. Test Assess. Model.</source> <volume>62</volume>, <fpage>179</fpage>&#x2013;<lpage>231</lpage>.</citation></ref>
<ref id="ref27"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ko</surname><given-names>Y. C.</given-names></name></person-group> (<year>2013</year>). <source>A study on cultural differences between the East and the West (Master&#x2019;s thesis)</source>. <publisher-loc>Chungju, South Korea</publisher-loc>: <publisher-name>Korea National University of Transportation, Graduate School of Humanities</publisher-name>.</citation></ref>
<ref id="ref28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lau</surname><given-names>K. L.</given-names></name> <name><surname>Ho</surname><given-names>E. S. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Reading performance and self-regulated learning of Hong Kong students: what we learnt from PISA 2009</article-title>. <source>Asia-Pac. Educ. Res.</source> <volume>25</volume>, <fpage>159</fpage>&#x2013;<lpage>171</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s40299-015-0246-1</pub-id></citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lau</surname><given-names>K. L.</given-names></name> <name><surname>Lee</surname><given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Examining Hong Kong students&#x2019; achievement goals and their relations with students&#x2019; perceived classroom environment and strategy use</article-title>. <source>Educ. Psychol.</source> <volume>28</volume>, <fpage>357</fpage>&#x2013;<lpage>372</lpage>. doi: <pub-id pub-id-type="doi">10.1080/01443410701612008</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lau</surname><given-names>S.</given-names></name> <name><surname>Nie</surname><given-names>Y.</given-names></name></person-group> (<year>2008</year>). <article-title>Interplay between personal goals and classroom goal structures in predicting student outcomes: a multilevel analysis of person-context interactions</article-title>. <source>J. Educ. Psychol.</source> <volume>100</volume>, <fpage>15</fpage>&#x2013;<lpage>29</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0022-0663.100.1.15</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname><given-names>K. K.</given-names></name></person-group> (<year>2013</year>). <article-title>A comparative study on national language curricula of Korea, China, and USA</article-title>. <source>Han-Geul</source> <volume>300</volume>, <fpage>183</fpage>&#x2013;<lpage>212</lpage>. doi: <pub-id pub-id-type="doi">10.22557/HG.2016.06.300.183</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lei</surname><given-names>M.</given-names></name> <name><surname>Lomax</surname><given-names>R. G.</given-names></name></person-group> (<year>2005</year>). <article-title>The effect of varying degrees of nonnormality in structural equation modeling</article-title>. <source>Struct. Equ. Model.</source> <volume>12</volume>, <fpage>1</fpage>&#x2013;<lpage>27</lpage>. doi: <pub-id pub-id-type="doi">10.1207/s15328007sem1201_1</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname><given-names>Y.</given-names></name> <name><surname>Brooks</surname><given-names>G. P.</given-names></name> <name><surname>Johanson</surname><given-names>G. A.</given-names></name></person-group> (<year>2012</year>). <article-title>Item discrimination and type I error in the detection of differential item functioning</article-title>. <source>Educ. Psychol. Meas.</source> <volume>72</volume>, <fpage>847</fpage>&#x2013;<lpage>861</lpage>. doi: <pub-id pub-id-type="doi">10.1177/0013164411426157</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Magis</surname><given-names>D.</given-names></name> <name><surname>Beland</surname><given-names>S.</given-names></name> <name><surname>Tuerlinckx</surname><given-names>F.</given-names></name> <name><surname>De Boeck</surname><given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>A general framework and an R package for the detection of dichotomous differential item functioning</article-title>. <source>Behav. Res. Methods</source> <volume>42</volume>, <fpage>847</fpage>&#x2013;<lpage>862</lpage>. doi: <pub-id pub-id-type="doi">10.3758/BRM.42.3.847</pub-id>, PMID: <pub-id pub-id-type="pmid">20805607</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Mahler</surname><given-names>C.</given-names></name></person-group> (<year>2011</year>). <source>The effects of misspecification type and nuisance variables on the behaviors of population fit indices used in structural equation modeling (Doctoral dissertation)</source>. <publisher-loc>Vancouver, Canada</publisher-loc>: <publisher-name>The University of British Columbia</publisher-name>.</citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Muench</surname><given-names>R.</given-names></name> <name><surname>Wieczorek</surname><given-names>O.</given-names></name> <name><surname>Gerl</surname><given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>Education regime and creativity: the eastern Confucian and the Western enlightenment types of learning in the PISA test</article-title>. <source>Cogent Educ.</source> <volume>9</volume>:<fpage>2144025</fpage>. doi: <pub-id pub-id-type="doi">10.1080/2331186X.2022.2144025</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Narayanan</surname><given-names>P.</given-names></name> <name><surname>Swaminathan</surname><given-names>H.</given-names></name></person-group> (<year>1996</year>). <article-title>Identification of items that show nonuniform DIF</article-title>. <source>Appl. Psychol. Meas.</source> <volume>20</volume>, <fpage>257</fpage>&#x2013;<lpage>274</lpage>. doi: <pub-id pub-id-type="doi">10.1177/014662169602000306</pub-id></citation></ref>
<ref id="ref38"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll1">OECD</collab></person-group> (<year>2019a</year>). PISA 2018 Results (Volume I): What Students Know and Can Do, PISA. Paris: OECD Publishing. doi: <pub-id pub-id-type="doi">10.1787/5f07c754-en</pub-id></citation></ref>
<ref id="ref39"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll2">OECD</collab></person-group>. (<year>2019b</year>). PISA 2018 technical report. Available online at: <ext-link xlink:href="https://www.oecd.org/pisa/data/pisa2018technicalreport/" ext-link-type="uri">https://www.oecd.org/pisa/data/pisa2018technicalreport/</ext-link></citation></ref>
<ref id="ref40"><citation citation-type="other"><person-group person-group-type="author"><collab id="coll3">OECD</collab></person-group>. (<year>2019c</year>). <source>PISA 2018 assessment and analytical framework</source>. <publisher-loc>Paris, France</publisher-loc>: <publisher-name>OECD Publishing</publisher-name>. doi: <pub-id pub-id-type="doi">10.1787/b25efab8-en</pub-id></citation></ref>
<ref id="ref41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oliden</surname><given-names>P. E.</given-names></name> <name><surname>Lizaso</surname><given-names>J. M.</given-names></name></person-group> (<year>2013</year>). <article-title>Invariance levels across language versions of the PISA 2009 reading comprehension tests in Spain</article-title>. <source>Psicothema</source> <volume>25</volume>, <fpage>390</fpage>&#x2013;<lpage>395</lpage>. doi: <pub-id pub-id-type="doi">10.7334/psicothema2013.46</pub-id></citation></ref>
<ref id="ref42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qian</surname><given-names>Q.</given-names></name> <name><surname>Lau</surname><given-names>K.-l.</given-names></name></person-group> (<year>2022</year>). <article-title>The effects of achievement goals and perceived reading instruction on Chinese student reading performance: evidence from PISA 2018</article-title>. <source>J. Res. Read.</source> <volume>45</volume>, <fpage>137</fpage>&#x2013;<lpage>156</lpage>. doi: <pub-id pub-id-type="doi">10.1111/1467-9817.12388</pub-id></citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roussos</surname><given-names>L. A.</given-names></name> <name><surname>Schnipke</surname><given-names>D. L.</given-names></name> <name><surname>Pashley</surname><given-names>P. J.</given-names></name></person-group> (<year>1999</year>). <article-title>A generalized formula for the mantel-Haenszel differential item functioning parameter</article-title>. <source>J. Educ. Behav. Stat.</source> <volume>24</volume>, <fpage>293</fpage>&#x2013;<lpage>322</lpage>. doi: <pub-id pub-id-type="doi">10.3102/10769986024003293</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sachse</surname><given-names>K. A.</given-names></name> <name><surname>Roppelt</surname><given-names>A.</given-names></name> <name><surname>Haag</surname><given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>A comparison of linking methods for estimating national trends in international comparative large-scale assessments in the presence of cross-national DIF</article-title>. <source>J. Educ. Meas.</source> <volume>53</volume>, <fpage>152</fpage>&#x2013;<lpage>171</lpage>. doi: <pub-id pub-id-type="doi">10.1111/jedm.12106</pub-id></citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salili</surname><given-names>F.</given-names></name> <name><surname>Lai</surname><given-names>M. K.</given-names></name></person-group> (<year>2003</year>). <article-title>Learning and motivation of Chinese students in Hong Kong: a longitudinal study of contextual influences on students&#x2019; achievement orientation and performance</article-title>. <source>Psychol. Schs.</source> <volume>40</volume>, <fpage>51</fpage>&#x2013;<lpage>70</lpage>. doi: <pub-id pub-id-type="doi">10.1002/pits.10069</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scott</surname><given-names>N. W.</given-names></name> <name><surname>Fayers</surname><given-names>P. M.</given-names></name> <name><surname>Aaronson</surname><given-names>N. K.</given-names></name> <name><surname>Bottomley</surname><given-names>A.</given-names></name> <name><surname>de Graeff</surname><given-names>A.</given-names></name> <name><surname>Groenvold</surname><given-names>M.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Differential item functioning (DIF) analyses of health-related quality of life instruments using logistic regression</article-title>. <source>Health Qual. Life Outcomes</source> <volume>8</volume>:<fpage>81</fpage>. doi: <pub-id pub-id-type="doi">10.1186/1477-7525-8-81</pub-id>, PMID: <pub-id pub-id-type="pmid">20684767</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sohn</surname><given-names>W.</given-names></name></person-group> (<year>2010</year>). <article-title>Exploring potential sources of DIF for PISA 2006 mathematics literacy items: application of logistic regression analysis</article-title>. <source>J. Educ. Eval.</source> <volume>23</volume>, <fpage>371</fpage>&#x2013;<lpage>390</lpage>.</citation></ref>
<ref id="ref49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Solano-Flores</surname><given-names>G.</given-names></name> <name><surname>Backhoff</surname><given-names>E.</given-names></name> <name><surname>Contreras-Ni&#x00F1;o</surname><given-names>L. &#x00C1;.</given-names></name></person-group> (<year>2009</year>). <article-title>Theory of test translation error</article-title>. <source>Int. J. Test.</source> <volume>9</volume>, <fpage>78</fpage>&#x2013;<lpage>91</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305050902880835</pub-id></citation></ref>
<ref id="ref50"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Solano-Flores</surname><given-names>G.</given-names></name> <name><surname>Contreras-Ni&#x00F1;o</surname><given-names>L. A.</given-names></name> <name><surname>Backhoff</surname><given-names>E.</given-names></name></person-group> (<year>2005</year>).The Mexican translation of TIMSS-95: test translation lessons from a post-mortem study. Paper presented at the annual meeting of the National Council on measurement in education, Montreal, Quebec, Canada.</citation></ref>
<ref id="ref51"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Solano-Flores</surname><given-names>G.</given-names></name> <name><surname>Contreras-Ni&#x00F1;o</surname><given-names>L. A.</given-names></name> <name><surname>Backhoff</surname><given-names>E.</given-names></name></person-group> (<year>2013</year>). &#x201C;<article-title>The measurement of translation error in PISA-2006 items: an application of the theory of test translation error</article-title>&#x201D; in <source>Research on PISA</source>. eds. <person-group person-group-type="editor"><name><surname>Prenzel</surname><given-names>M.</given-names></name> <name><surname>Kobarg</surname><given-names>M.</given-names></name> <name><surname>Sch&#x00F6;ps</surname><given-names>K.</given-names></name> <name><surname>R&#x00F6;nnebeck</surname><given-names>S.</given-names></name></person-group> (<publisher-loc>Dordrecht, The Netherlands</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>71</fpage>&#x2013;<lpage>85</lpage>.</citation></ref>
<ref id="ref52"><citation citation-type="book"><person-group person-group-type="author"><name><surname>S&#x00F6;yler Ba&#x011F;du</surname><given-names>P.</given-names></name></person-group> (<year>2020</year>). <source>Investigation of the measurement variability of PISA 2015 reading skills test according to the language variability (master&#x2019;s thesis)</source>. <publisher-loc>&#x0130;zmir</publisher-loc>: <publisher-name>Ege University, Institute of Educational Sciences</publisher-name>.</citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Strobl</surname><given-names>C.</given-names></name> <name><surname>Kopf</surname><given-names>J.</given-names></name> <name><surname>Zeileis</surname><given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Rasch trees: a new method for detecting differential item functioning in the Rasch model</article-title>. <source>Psychometrika</source> <volume>80</volume>, <fpage>289</fpage>&#x2013;<lpage>316</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11336-013-9388-3</pub-id>, PMID: <pub-id pub-id-type="pmid">24352514</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sung</surname><given-names>E. H.</given-names></name></person-group> (<year>2006</year>). <article-title>Creativity in the Est and the west</article-title>. <source>J. Korean Soc. Gift. Talent.</source> <volume>5</volume>, <fpage>67</fpage>&#x2013;<lpage>93</lpage>.</citation></ref>
<ref id="ref55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swaminathan</surname><given-names>H.</given-names></name> <name><surname>Rogers</surname><given-names>H. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Detecting differential item functioning using logistic regression procedures</article-title>. <source>J. Educ. Meas.</source> <volume>27</volume>, <fpage>361</fpage>&#x2013;<lpage>370</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1745-3984.1990.tb00754.x</pub-id></citation></ref>
<ref id="ref56"><citation citation-type="book"><person-group person-group-type="author"><collab id="coll4">The Ministry of Education and People&#x2019;s Republic of China</collab></person-group> (<year>2011</year>). <source>Curriculum standards of basic Chinese language education</source>. <publisher-loc>Beijing, China</publisher-loc>: <publisher-name>People&#x2019;s Education Press</publisher-name>.</citation></ref>
<ref id="ref57"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Thissen</surname><given-names>D.</given-names></name></person-group> (<year>2001</year>). <source>IRTLRDIF v2.0b&#x2014;Software for the computation of the statistics involved in item response theory likelihood-ratio tests for differential item functioning [computer software documentation]</source>. <publisher-loc>Chapel Hill</publisher-loc>: <publisher-name>L. L. Thurstone Psychometric Laboratory, University of North Carolina</publisher-name>.</citation></ref>
<ref id="ref58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tutz</surname><given-names>G.</given-names></name> <name><surname>Berger</surname><given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Item-focussed trees for the identification of items in differential item functioning</article-title>. <source>Psychometrika</source> <volume>81</volume>, <fpage>727</fpage>&#x2013;<lpage>750</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11336-015-9488-3</pub-id>, PMID: <pub-id pub-id-type="pmid">26596721</pub-id></citation></ref>
<ref id="ref59"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Watkins</surname><given-names>D. A.</given-names></name> <name><surname>Biggs</surname><given-names>J. B.</given-names></name></person-group> (<year>2001</year>). <source>Teaching the Chinese learner: Psychological and pedagogical perspectives</source>. <publisher-loc>Hong Kong</publisher-loc>: <publisher-name>Hong Kong University Press</publisher-name>.</citation></ref>
<ref id="ref60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yildirim</surname><given-names>H. H.</given-names></name> <name><surname>Berbero&#x011D;lu</surname><given-names>G.</given-names></name></person-group> (<year>2009</year>). <article-title>Judgmental and statistical DIF analyses of the PISA-2003 mathematics literacy items</article-title>. <source>Int. J. Test.</source> <volume>9</volume>, <fpage>108</fpage>&#x2013;<lpage>121</lpage>. doi: <pub-id pub-id-type="doi">10.1080/15305050902880736</pub-id></citation></ref>
<ref id="ref61"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Zumbo</surname><given-names>B. D.</given-names></name></person-group> (<year>1999</year>). <source>A handbook on the theory and methods of differential item functioning (DIF): Logistic regression modeling as a unitary framework forbinary and likert-type (ordinal) item scores</source>. <publisher-loc>Ottawa, ON</publisher-loc>: <publisher-name>Directorate of Human Resources Research and Evaluation, Department of National Defense</publisher-name>.</citation></ref>
<ref id="ref62"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Zumbo</surname><given-names>B. D.</given-names></name> <name><surname>Thomas</surname><given-names>D. R.</given-names></name></person-group> (<year>1996</year>). <source>A measure of DIF effect size using logistic regression procedures</source>. <publisher-loc>Philadelphia, PA</publisher-loc>: <publisher-name>Paper presented at National Board of Medical Examiners</publisher-name>.</citation></ref>
<ref id="ref63"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zusho</surname><given-names>A.</given-names></name> <name><surname>Clayton</surname><given-names>K.</given-names></name></person-group> (<year>2011</year>). <article-title>Culturalizing achievement goal theory and research</article-title>. <source>Educ. Psychol.</source> <volume>46</volume>, <fpage>239</fpage>&#x2013;<lpage>260</lpage>. doi: <pub-id pub-id-type="doi">10.1080/00461520.2011.614526</pub-id></citation></ref>
</ref-list>
</back>
</article>