<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2018.00325</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Using Anchoring Vignettes to Adjust Self-Reported Personality: A Comparison Between Countries</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Weiss</surname> <given-names>Selina</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/489121/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Roberts</surname> <given-names>Richard D.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Individual Differences and Psychological Assessment, Institute of Psychology and Education, Ulm University</institution>, <addr-line>Ulm</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>ProExam - an ACT Affiliated Company</institution>, <addr-line>New York, NY</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Martin S. Hagger, Curtin University, Australia</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Derwin King Chung Chan, University of Hong Kong, Hong Kong; Yu Yang, ShanghaiTech University, China; Katrin Rentzsch, University of Bamberg, Germany</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Selina Weiss <email>Selina.weiss&#x00040;uni-ulm.de</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Personality and Social Psychology, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>03</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>9</volume>
<elocation-id>325</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>10</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>02</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2018 Weiss and Roberts.</copyright-statement>
<copyright-year>2018</copyright-year>
<copyright-holder>Weiss and Roberts</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Data from self-report tools cannot be readily compared between cultures due to culturally specific ways of using a response scale. As such, anchoring vignettes have been proposed as a suitable methodology for correcting against this difference. We developed anchoring vignettes for the Big Five Inventory-44 (BFI-44) to supplement its Likert-type response options. Based on two samples (Rwanda: <italic>n</italic> &#x0003D; 423; Philippines: <italic>n</italic> &#x0003D; 143), we evaluated the psychometric properties of the measure both before and after applying the anchoring vignette adjustment. Results show that adjusted scores had better measurement properties, including improved reliability and a more orthogonal correlational structure, relative to scores based on the original Likert scale. Correlations of the Big Five Personality Factors with life satisfaction were essentially unchanged after the vignette-adjustment while correlations with counterproductive were noticeably lower. Overall, these changed findings suggest that the use of anchoring vignette methodology improves the cross-cultural comparability of self-reported personality, a finding of potential interest to the field of global workforce research and development as well as educational policymakers.</p>
</abstract>
<kwd-group>
<kwd>anchoring vignettes</kwd>
<kwd>personality scales and inventories</kwd>
<kwd>Big Five</kwd>
<kwd>differential item functioning</kwd>
<kwd>cross-cultural differences</kwd>
</kwd-group>
<contract-num rid="cn001">AID-OAA-LA-13-00008</contract-num>
<contract-sponsor id="cn001">United States Agency for International Development<named-content content-type="fundref-id">10.13039/100000200</named-content></contract-sponsor>
<counts>
<fig-count count="3"/>
<table-count count="7"/>
<equation-count count="0"/>
<ref-count count="100"/>
<page-count count="17"/>
<word-count count="13274"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Self-report questionnaires are a dominant assessment methodology in the social sciences. They are used to estimate important information about a participant&#x00027;s personality, attitudes, values, and beliefs. However, self-report questionnaires are prone to various biases that challenge the utility of this methodology including the validity of the data. These biases include cultural artifacts (e.g., measurement artifacts and differences in response sets due to differences in communication styles between cultures; Van de Vijver and Leung, <xref ref-type="bibr" rid="B90">1997</xref>; Fischer, <xref ref-type="bibr" rid="B24">2004</xref>), active deception (e.g., Ziegler et al., <xref ref-type="bibr" rid="B98">2011</xref>), and personal biases in response styles such as extreme responding, midpoint responding, and acquiescent (i.e., a tendency to agree with items and hence only using the upper half of the response option scale) and disacquiescent responding (i.e., a tendency to generally disagree with items and hence only use the bottom half of the response option scale; Van Vaerenbergh and Thomas, <xref ref-type="bibr" rid="B92">2013</xref>). Cross-cultural biases occur because participants compare themselves to the standards and values of their cultural group, also known as their reference group (Peng et al., <xref ref-type="bibr" rid="B70">1997</xref>; Heine et al., <xref ref-type="bibr" rid="B33">2002</xref>). The difference in item responses between the two groups is called <italic>Differential Item Functioning</italic> (DIF; Holland and Wainer, <xref ref-type="bibr" rid="B36">1993</xref>; Osterlind and Everson, <xref ref-type="bibr" rid="B68">2009</xref>). There are several reasons why an item, used in a cross-cultural context, can show DIF (Ellis et al., <xref ref-type="bibr" rid="B22">1993</xref>). An item can display DIF because of (1) mistakes in the item&#x00027;s translation, (2) because participants ascribe unique meanings to the item because of their culture, or (3) participants have different cultural knowledge (Johnson et al., <xref ref-type="bibr" rid="B43">2008</xref>). The anchoring vignette technique is a method that can detect DIF and adjust for some of these cross-cultural biases that lead to item DIF (King et al., <xref ref-type="bibr" rid="B49">2004</xref>; King and Wand, <xref ref-type="bibr" rid="B48">2007</xref>; Hopkins and King, <xref ref-type="bibr" rid="B37">2010</xref>). Understanding the impact of DIF is important for the development of new assessment tools and especially for their application. The anchoring vignettes provided in this study are a useful and easily applicable technique that can be applied to existing personality measures in cross-cultural research, which will help combat DIF.</p>
<p>To sufficiently support our hypotheses, the introduction is organized as follows. First, we provide an overview of traditional techniques for detecting DIF, followed by a summary of Anchoring Vignettes (<italic>AV</italic>s) and why they are superior over traditional methods. Second, we summarize how <italic>AV</italic>s have been applied both generally and specifically in personality research, specifically in regards to the Big Five Personality factor model. Finally, we demonstrate the utility of <italic>AV</italic>s for combating DIF in the assessment of the Big Five Personality factors, based on data from two countries: Rwanda and the Philippines.</p>
<p>The following passage describes several traditional techniques that are applied to detect DIF. Huang et al. (<xref ref-type="bibr" rid="B40">1997</xref>) used the five-factor personality model in two cultural contexts, the Philippines and America, and compared two DIF-statistics to examine measurement equivalence at the item level: (1) Item Response Theory (IRT) and (2) Mantel-Haenszel method (e.g., Ellis et al., <xref ref-type="bibr" rid="B22">1993</xref>; Huang et al., <xref ref-type="bibr" rid="B40">1997</xref>). DIF can be detected by classic IRT statistics, such as item discrimination and item difficulty (e.g., Camilli and Shepard, <xref ref-type="bibr" rid="B13">1994</xref>), or by the area between two cultures&#x00027; item characteristic curves (Thissen et al., <xref ref-type="bibr" rid="B86">1988</xref>). IRT item parameters are often assumed to be invariant over groups of participants. However, this is often not true (Rupp and Zumbo, <xref ref-type="bibr" rid="B79">2006</xref>). Instead, parameter invariance can be assessed to identify items that lack measurement equivalence, which means the item assesses the central construct differently for each group. However, it is argued that demonstrating factor congruence across cultures does not guarantee measurement equivalence (Bijnen et al., <xref ref-type="bibr" rid="B8">1986</xref>; Huang et al., <xref ref-type="bibr" rid="B40">1997</xref>).</p>
<p>The Mantel-Haenszel, a chi-square statistic comparing the actual and expected frequencies, can be used to detect DIF (Holland and Thayer, <xref ref-type="bibr" rid="B35">1988</xref>). This method has a lot of advantages including its simplicity and easy implementation. However, it does not detect <italic>non-uniform-</italic>DIF, which is an interaction between trait level and group membership so that mean differences in trait level between cultures would not be detected (Rogers and Swaminathan, <xref ref-type="bibr" rid="B76">1993</xref>).</p>
<p>Most historical DIF-statistics focused on binary scored items. Ordinal scaled items, such as Likert-scale items, require a different treatment, such as a logistic regression. Zumbo (<xref ref-type="bibr" rid="B100">1999</xref>) estimated a logistic regression to test DIF for ordinal scored items using the responses as a dependent variable with a grouping variable, total scale score, and an interaction of the group and the total scale score as independent variables (e.g., Crane et al., <xref ref-type="bibr" rid="B18">2006</xref>).</p>
<p>DIF can be also detected using multiple-group confirmatory factor analysis through establishing measurement invariance (Thissen et al., <xref ref-type="bibr" rid="B86">1988</xref>; Stark et al., <xref ref-type="bibr" rid="B82">2006</xref>; Teresi, <xref ref-type="bibr" rid="B85">2006</xref>). Configural invariance, the first step in testing measurement invariance, models the same factor structure across groups (Vandenberg and Lance, <xref ref-type="bibr" rid="B93">2000</xref>; Stark et al., <xref ref-type="bibr" rid="B82">2006</xref>). In the case where a sample does not demonstrate configural invariance across countries, it can be assumed that single items, or even the whole test, is affected by DIF (Teresi, <xref ref-type="bibr" rid="B85">2006</xref>). Likewise, DIF can be detected by comparing item factor loadings (e.g., Eysenck et al., <xref ref-type="bibr" rid="B23">1993</xref>).</p>
<p>The use of <italic>AVs</italic> (Thissen et al., <xref ref-type="bibr" rid="B87">1993</xref>) is perhaps the most promising approach that can be applied to detect DIF and the method offers the possibility to correct for it. The idea is to relate self-report answers with external benchmarks that measure the same concept but are more likely to be free of biases and therefore free of some DIF forms. Anchors specifically have been proposed as a useful tool to adjust the answers of different individuals to one underlying standard scale (King et al., <xref ref-type="bibr" rid="B49">2004</xref>). Anchors normally include descriptions (within vignettes) of one hypothetical person who, based on the theoretical description of the trait of interest, is described in a way to illustrate a certain trait level (Chevalier and Fielding, <xref ref-type="bibr" rid="B16">2011</xref>). Each participant then evaluates the behavior of this person on the same scale they used to answer the self-report questions. Because the anchors provide an external benchmark, <italic>AV</italic>s have a number of advantages over traditional DIF-detection procedures (M&#x000F5;ttus et al., <xref ref-type="bibr" rid="B59">2012</xref>). Primarily, while traditional DIF-statistics (described above) essentially plot single item scores against latent trait scores, with both types of scores derived from the same data, with <italic>AV</italic>s, there is independence between the scores (M&#x000F5;ttus et al., <xref ref-type="bibr" rid="B59">2012</xref>).</p>
<p>The following passage provides a short and general overview over the application of <italic>AV</italic>s in different research areas. <italic>AV</italic>s are not widely employed, though there are a few isolated instances of their use. They have been used in work related research (e.g., Kristensen and Johansson, <xref ref-type="bibr" rid="B51">2008</xref>), in research on life satisfaction (e.g., Kapteyn et al., <xref ref-type="bibr" rid="B47">2010</xref>) and quality of life (e.g., Crane et al., <xref ref-type="bibr" rid="B17">2016</xref>), in personality research (e.g, M&#x000F5;ttus et al., <xref ref-type="bibr" rid="B59">2012</xref>), and in an educational context (e.g., student-reported teachers&#x00027; classroom management; OECD, <xref ref-type="bibr" rid="B64">2012</xref>; Vonkova et al., <xref ref-type="bibr" rid="B94">2015</xref>). These applications mostly indicate that the use of <italic>AV</italic>s was beneficial and resulted in a more valid measure that offered cleaner comparisons between groups. Research on life satisfaction (Angelini et al., <xref ref-type="bibr" rid="B4">2014</xref>), for example, found that Danes and Italians report different levels of life satisfaction. But after adjusting self-report answers with <italic>AV</italic>s, these differences disappeared. Likewise, <italic>AV</italic>s helped improve measurement invariance in the Programme for International Student Assessment <italic>(PISA)</italic> (He and Van de Vijver, <xref ref-type="bibr" rid="B31">2016</xref>), and <italic>AV</italic>s were effectively used to identify and correct for DIF in a self-report physical health scale (Knott et al., <xref ref-type="bibr" rid="B50">2016</xref>).</p>
<p>Personality research, which is conducted in several countries, can especially benefit from the application of <italic>AV</italic>s. A long history of psychological research has shown that the Big Five Factor model of Personality represents the set of constructs that are most strongly differentiated, non-overlapping, and predictive across domains (Roberts et al., <xref ref-type="bibr" rid="B75">2015</xref>). Although they were first discovered in the English language, replication studies in other languages yielded the same five factors (see e.g., McCrae and Terracciano, <xref ref-type="bibr" rid="B56">2005</xref>; Schmitt et al., <xref ref-type="bibr" rid="B80">2007</xref>). But already Allport and Odbert (<xref ref-type="bibr" rid="B2">1936</xref>) noticed that culture and time period can influence responses. There are especially large differences in answering personality items when comparing Western and non-Western cultures (e.g., Mpofu and Nyanungo, <xref ref-type="bibr" rid="B61">1998</xref>; Byrne and Campbell, <xref ref-type="bibr" rid="B12">1999</xref>). DIF in personality items is known to appear because of inadequate translation, research, or development, sampling biases, and different response styles (e.g., Grimm and Church, <xref ref-type="bibr" rid="B26">1999</xref>; Van de Vijver and Leung, <xref ref-type="bibr" rid="B91">2000</xref>; Schmitt et al., <xref ref-type="bibr" rid="B80">2007</xref>).</p>
<p><italic>AV</italic>s have been shown to increase the reliability of scales assessing Conscientiousness and Openness in a representative study of 12th grade students in Brazil (<italic>N</italic> &#x0003D; 8,582) (Primi et al., <xref ref-type="bibr" rid="B71">2016</xref>). The study applied a set of three vignettes for Conscientiousness and Openness. Interestingly, they showed that the Openness vignettes were more frequently misordered relative to the Conscientiousness vignettes. In a study using <italic>AVs</italic> to compare facet-level measures of Conscientiousness across 21 countries, it was determined that, contrary to expectations, self-reported Conscientiousness was minimally affected by cultural differences (M&#x000F5;ttus et al., <xref ref-type="bibr" rid="B59">2012</xref>). Hence, the researchers concluded that it is not necessary to address comparability problems using <italic>AVs</italic> in personality. However, the sample size for each country was relatively small and some of the vignettes were abstract and most likely had differences in meaning due to the numerous translations. Hence, it is possible that the participants applied different standards in answering the <italic>AVs</italic> and the self-report personality questionnaires and hence violated the assumption of Response Consistency (discussed in detail below). Likewise, Primi et al. (<xref ref-type="bibr" rid="B71">2016</xref>) as well as M&#x000F5;ttus et al. (<xref ref-type="bibr" rid="B59">2012</xref>) do not report whether or not they tested any measurement assumptions (i.e., Response Consistency and Vignette Equivalence, which are described in detail below), which must be fulfilled in order to use <italic>AV</italic>s. He et al. (<xref ref-type="bibr" rid="B32">2017</xref>) compared different methods and procedures to improve the comparability between cultures including a vignette set with two levels of conscientiousness (<italic>N</italic> &#x0003D; 3,560 university students in 16 countries). They reported that the vignette sets showed a lack of invariance (arguably due to the characteristics of the vignettes) and hence were not free of DIF. However, they also found that the vignette technique was the only method which resulted in higher internal consistencies. Likewise, the use of <italic>AV</italic>s for the assessment of self-reported teamwork led to increased test information and item discrimination, and higher factor loadings and better model fit in a confirmatory factor analysis (Ham and Roberts, <xref ref-type="bibr" rid="B30">2015</xref>).</p>
<p>In our study, we evaluated the use of <italic>AVs</italic> in the assessment of personality in two countries: Rwanda and the Philippines. We selected these countries for several reasons. First, Rwanda is one the few African countries where the Big Five have not been replicated (Roberts et al., <xref ref-type="bibr" rid="B75">2015</xref>). Thus, the comparison of responses from a country where DIF in the Big Five has already been shown, specifically the Philippines, with another country where the Big Five factor structure have not been replicated, and has a different cultural and historical experience, is informative. Identifying a different structure to the Big Five Factor model for Rwanda would challenge the previously proclaimed generalizability of the Big Five Factor model.</p>
<p>Reviewing cross cultural personality research where <italic>AV</italic>-adjustment was not applied suggests a mean trait-level difference between the Philippines and Rwanda. Researchers who administered the BFI in 56 nations using 28 languages found significant differences in Openness to Experience in the geographical regions of South East Asia compared to other world regions (Schmitt et al., <xref ref-type="bibr" rid="B80">2007</xref>). Likewise, a comparison of the United States and the Philippines on the Big Five found significant mean differences between the groups. However, in that study, almost 40% of the items administered in the Philippines showed DIF, despite surveying both groups in English (Huang et al., <xref ref-type="bibr" rid="B40">1997</xref>).</p>
<p>There is also evidence suggesting a cultural difference in the approach toward a self-report questionnaire and the use of the response options. For example, Rwandans tend to place an extremely high value on authorities and people in a high status position, which can affect responding on self-report questionnaires (Staub et al., <xref ref-type="bibr" rid="B83">2005</xref>). Likewise, citizens of several African countries (i.e., Benin, South Africa, Senegal, and Burkina Faso) and Southeast Asia (e.g., the Philippines) showed the highest rates of extreme responding to self-reported personality (M&#x000F5;ttus et al., <xref ref-type="bibr" rid="B59">2012</xref>). In contrast, Germany shows the lowest rates of extreme responding, and most European nations and the United States are characterized by medium extreme responses. Thus, this suggests that individuals from the Philippines and African countries are characterized by a difference in understanding and interpretation of self-report questionnaire items, namely those assessing personality, which could explain the aforementioned mean trait-level differences. As such, it is important to test the extent to which there is DIF between the Philippines and Rwanda and whether this can be addressed through the use of <italic>AV</italic>s. As our study is the first to apply vignette sets (with three levels) on the BFI-44, comparison to previously published results with other countries is not possible.</p>
<p>The Big Five are linked to several important aspects of our life. Life satisfaction, a component of subjective well-being, is correlated to the Big Five, but the correlations are somewhat inconsistent between countries. In a representative Dutch sample, life satisfaction had small to medium positive correlations with all Big Five factors (M&#x000FC;ller, <xref ref-type="bibr" rid="B62">2014</xref>). However, in an Iranian sample, researchers found small negative correlations of life satisfaction with Conscientiousness, Openness, and Extraversion (<italic>r</italic> &#x0003D; &#x02212;0.24 to &#x02212;0.28) and non-significant correlations with Neuroticism and Agreeableness (Hosseinkhanzadeh and Taher, <xref ref-type="bibr" rid="B39">2013</xref>). In a Nigerian sample, correlations of life satisfaction with Neuroticism were negative but positive with the other Big Five factors (Onyishi et al., <xref ref-type="bibr" rid="B66">2012</xref>). The Big Five are also linked to work behavior. In a USA sample, all Big Five dimensions had small to medium negative correlations with counterproductive work behavior (Mount et al., <xref ref-type="bibr" rid="B60">2006</xref>). The extent to which this relation differs between cultures is unclear.</p>
<p>Overall, we hypothesize that using the <italic>AV</italic> technique will improve the psychometric characteristics and the cross-cultural comparability of self-reported personality. Here are our specific hypotheses:
<list list-type="simple">
<list-item><p><italic>Hypothesis 1</italic>: Using <italic>AVs</italic> to adjust self-report responses will improve the BFI-44 reliability, measured with omega &#x003C9; (an indicator of factor saturation; McDonald, <xref ref-type="bibr" rid="B58">1999</xref>), for each of the Big Five Personality factors, with estimates ranging from good to excellent. This hypothesis will be tested by comparing the overlap of the 95% confidence intervals of the omegas.</p></list-item>
<list-item><p><italic>Hypothesis 2: AV</italic>-adjusted scores fitted in a graded response model will show an increase in discriminant power, relative to the original scores, resulting in a wider range of thresholds and larger discrimination parameters. Furthermore, the increase in overall test information for the <italic>AV</italic>-adjusted scores will be indicated by a broader range of &#x003B8;-levels and small standard errors.</p></list-item>
<list-item><p><italic>Hypothesis 3:</italic> In a confirmatory factor analysis, models of the Big Five based on <italic>AV</italic>-adjusted scores will show acceptable fit to the data (CFI &#x02265; 0.90 and RMSEA &#x02264; 0.08) and a correlational structure such that Neuroticism is weakly negatively correlated with all of the other dimensions, and the other dimensions are weakly positively correlated with each other, supporting the theoretical structure of the Big Five model. We expect this result to hold for both samples.</p></list-item>
<list-item><p><italic>Hypothesis 4:</italic> Test-criterion relationships of life satisfaction and counterproductive work behavior with the <italic>AV</italic>-adjusted Big Five factor scores will be significantly stronger for the <italic>AV</italic>-adjusted scores.</p></list-item>
</list></p>
</sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<sec>
<title>Procedure and sample</title>
<p>The studies were conducted in Rwanda (Sample 1) and in the Philippines (Sample 2). Participants were recruited through the educational institute Akilah Institute for Woman of Akazi College in Africa, which is partially supported by the Educational Development Center (<italic>EDC</italic>) in Washington D.C. This study was carried out in accordance with the recommendations of the institutional review board <italic>(IRB</italic>; <italic>IRB</italic> Registration: IRB00000865) of the <italic>EDC</italic> Human Protections department. In addition, all participants provided written informed consent in accordance with the Declaration of Helsinki. In the Philippines, the parents of the participants were also informed of the study and the possible involvement of their child with a letter.</p>
<p>The newly-developed <italic>AV</italic>s were tested in psychology labs before the study was conducted. All items and translations were reviewed several times by all institutes involved. The studies were conducted by local interviewers who were employed and trained for data collection by the <italic>EDC</italic>. Each participant completed the questions in the same order, which were presented on a tablet provided by the <italic>EDC</italic>. Due to technical problems and power supply issues in both countries, some participants completed paper-pencil versions of the test material. The paper-pencil versions were entered into an electronic database by the local staff. Participation in these studies was voluntary and participants could withdraw from the study at any moment. The samples are summarized in Table <xref ref-type="table" rid="T1">1</xref>. The samples are based on adolescents and young adults who were either finishing school and/or applying for jobs. Several articles support this application of the Big Five in young adult and adolescent samples (see for example, Bratko and Maru&#x00161;i&#x00107;, <xref ref-type="bibr" rid="B10">1997</xref>; Digman, <xref ref-type="bibr" rid="B20">1997</xref>; Ehrler et al., <xref ref-type="bibr" rid="B21">1999</xref>; Rothbart et al., <xref ref-type="bibr" rid="B78">2000</xref>). This sampling procedure resulted in a relatively homogeneous sample regarding age and education, which were biases what we wanted to avoid. The participants completed the study during their college course time; therefore, they did not receive any financial compensation.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Descriptions of the Rwanda and Philippines samples.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold><italic>N</italic></bold></th>
<th valign="top" align="center"><bold>Age</bold></th>
<th valign="top" align="center"><bold>Sex</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Sample 1: Rwanda</td>
<td valign="top" align="center">423</td>
<td valign="top" align="center"><italic>M</italic> &#x0003D; 21.79 <italic>SD</italic> &#x0003D; 2.7</td>
<td valign="top" align="center">Female: <italic>N</italic> &#x0003D; 356 (84%)</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Range: 15&#x02013;33</td>
<td valign="top" align="center">Male: <italic>N</italic> &#x0003D; 67 (16%)</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Sample 2:Philippines</td>
<td valign="top" align="center">143</td>
<td valign="top" align="center"><italic>M</italic> &#x0003D; 15.5 <italic>SD</italic> &#x0003D; 0.83</td>
<td valign="top" align="center">Female: <italic>N</italic> &#x0003D; 99 (69%)</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="center">Range: 14&#x02013;19</td>
<td valign="top" align="center">Male: <italic>N</italic> &#x0003D; 44 (31%)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The Rwanda sample includes two different schools in different regions of Rwanda. The Philippines sample includes one school</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Measures</title>
<p>All measures were translated from English into either Kinyarwanda or Filipino. The translation was supervised by the <italic>EDC</italic> using backward and forward translation checks. The participants completed demographic questions first where they were asked to provide information about themselves, their family, and their home situation. These questions were also tailored for each country; for example, Filipinos were asked if they have a computer at home and Rwandans were asked if they have access to running water. Therefore, the demographics differed between both samples. As the studies were part of a larger project, additional measures were also included including 10 Situational Judgment Tests for Conscientiousness and the Conscientiousness Facet-Tool (MacCann et al., <xref ref-type="bibr" rid="B54">2009</xref>). Because of the focus of the article, these measures are not discussed further.</p>
<sec>
<title>Anchoring vignettes</title>
<p>The <italic>AVs</italic> included 15 hypothetical descriptions; three hypothetical descriptions for each personality dimension of males or females who embodied a certain level of the corresponding personality dimension. The <italic>AVs</italic> were developed by scientists from the Professional Examination Service in New York. Table <xref ref-type="table" rid="T2">2</xref> shows <italic>AVs</italic> for the Big Five, which show different levels of Conscientiousness, Agreeableness, Neuroticism, Openness, and Extraversion. In the first <italic>AV</italic> for conscientiousness, Sophia represents someone with a low level of Conscientiousness. In the second, Jacob shows a medium level of Conscientiousness, and in the third, Emma displays a high level of Conscientiousness. Participants were asked to rate the extent to which they agreed with the statement that Sophia, Jacob, and Emma are Conscientious. In this case, the suggested ratings for following the correct order would be &#x0201C;disagree strongly&#x0201D; or &#x0201C;disagree a little&#x0201D; for Sophia&#x00027;s statement, &#x0201C;neither agree nor disagree&#x0201D; for Jacob&#x00027;s statement, and &#x0201C;agree a little&#x0201D; or &#x0201C;agree strongly&#x0201D; for Emma&#x00027;s statement. Therefore, the person in Vignette 1 is rated as having lower Conscientiousness relative to the person in Vignette 2. Also, the person in Vignette 3 is rated as having the highest conscientiousness and is therefore higher than on conscientiousness the persons in described in Vignette 2 and 1.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><italic>AV</italic>s for Conscientiousness (C), Agreeableness (A), Neuroticism (N), Openness (O), and Extraversion (E).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>How much do you agree with this statement?</bold></th>
<th valign="top" align="center"><bold>Disagree strongly</bold></th>
<th valign="top" align="center"><bold>Disagree a little</bold></th>
<th valign="top" align="center"><bold>Neither agree nor disagree</bold></th>
<th valign="top" align="center"><bold>Agree a little</bold></th>
<th valign="top" align="center"><bold>Agree strongly</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">C1. Sophia tends to be somewhat careless. Other workers also comment that she is lazy. Sophia often also appears disorganized. Based on this information, to what extent do you agree with the statement &#x0201C;Sophia is conscientious/hard-working&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">C2. Jacob is a reliable worker and does all work with great efficiency, but he is easily distracted. Based on this information, to what extent do you agree with the statement &#x0201C;Jacob is conscientious/hard-working&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">C3. Emma always does a thorough job. She perseveres until all tasks are finished. Emma also makes plans and follows through with them. Based on this information, to what extent do you agree with the statement &#x0201C;Emma is conscientious/hard-working&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">A1. Jean tends to disagree with others, and as a result often starts quarrels. Indeed, many people consider Jean quite rude. Based on this information, to what extent do you agree with the statement &#x0201C;Jean is an agreeable person&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">A2. Even though Nicole is helpful and unselfish with others, some people find her cold and unfriendly. This does not matter so much, as she has a forgiving nature. Based on this information, to what extent do you agree with the statement &#x0201C;Nicole is an agreeable person&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">A3. Claude is considerate and kind to almost everyone. He is very trusting, and finds it easy to cooperate with others. Based on this information, to what extent do you agree with the statement &#x0201C;Claude is an agreeable person&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">N1. Carine frequently appears quite depressed to other people. She gets nervous easily. Based on this information, to what extent do you agree with the statement &#x0201C;Carine is emotionally stable&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">N2. Although in tense situations Paul remains calm, he can be quite moody. And he tends to worry quite a lot. Based on this information, to what extent do you agree with the statement &#x0201C;Paul is emotionally stable&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">N3. Aline always appears relaxed and to handle stress well. Indeed, she never comes across as upset. Aline remains calm in all situations. Based on this information, to what extent do you agree with the statement &#x0201C;Aline is emotionally stable&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">O1. Emmanuel has few artistic interests, and is not especially sophisticated either in music or literature. This has led some people to observe that Emmanuel does not appear especially curious about anything. Based on this information, to what extent do you agree with the statement &#x0201C;Emmanuel is open-minded&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">O2. Emma has an active imagination. This has led some people to calling her a deep thinker. Even so Emma prefers work that is routine. Based on this information, to what extent do you agree with the statement &#x0201C;Emma is open-minded&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">O3. Jean Bosco is original and always coming up with new ideas. This has led some people to calling him inventive. But beyond this, Jean Bosco values artistic, aesthetic experiences. Based on this information, to what extent do you agree with the statement &#x0201C;Jean Bosco is open-minded&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">E1. Claudine is very reserved. She tends to be quiet no matter what the circumstance. Indeed, people find her shy and inhibited. Based on this information, to what extent do you agree with the statement &#x0201C;Claudine is extraverted&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">E2. Emile is often talkative and generates a lot of enthusiasm in others. But on his day, Emile can be rather shy and inhibited. Based on this information, to what extent do you agree with the statement &#x0201C;Emile is extraverted&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">E3. Rosette has an assertive personality, and as a result appears outgoing and sociable. Indeed, people are always commenting on how full of energy Rosette is. Based on this information, to what extent do you agree with the statement &#x0201C;Rosette is extraverted&#x0201D;?</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
<td valign="top" align="center">O</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>The big five inventory (BFI-44)</title>
<p>The BFI-44 (John et al., <xref ref-type="bibr" rid="B42">1991</xref>; Benet-Mart&#x000ED;nez and John, <xref ref-type="bibr" rid="B5">1998</xref>) uses 44 items to measure the Big Five Personality factors: Extraversion (e.g., &#x0201C;<italic>I am someone who is talkative&#x0201D;)</italic>, Agreeableness (e.g., &#x0201C;<italic>I am someone who is helpful and unselfish with others&#x0201D;)</italic>, Conscientiousness (e.g., &#x0201C;<italic>I am someone who does a thorough job&#x0201D;)</italic>, Neuroticism (e.g., &#x0201C;<italic>I am someone who is depressed, blue&#x0201D;)</italic>, and Openness (e.g., &#x0201C;<italic>I am someone who is original, comes up with new ideas&#x0201D;)</italic>. The items are answered on a five-point Likert-scale with the poles &#x0201C;disagree strongly&#x0201D; and &#x0201C;agree strongly&#x0201D;. John and Srivastava (<xref ref-type="bibr" rid="B41">1999</xref>) established the validity and factor structure of this measurement. The reliabilities before and after the <italic>AV</italic>-adjustment are presented in the results section.</p>
</sec>
<sec>
<title>Satisfaction with life scale</title>
<p>This scale measures global life satisfaction with five items (e.g., &#x0201C;<italic>I am satisfied with my life&#x0201D;</italic>). It is known for good internal reliability and validity (Diener et al., <xref ref-type="bibr" rid="B19">1985</xref>). The reliability of the scale for the whole study was acceptable (&#x003C9; &#x0003D; 0.76).</p>
</sec>
<sec>
<title>Counterproductive behavior</title>
<p>This construct was measured with an adaption of the Interpersonal and Organizational Deviance items (Bennett and Robinson, <xref ref-type="bibr" rid="B6">2000</xref>) for a school and work context (e.g., &#x0201C;<italic>How often have you publicly embarrassed someone at school or work&#x0201D;</italic>). Respondents answered the items on a seven-point Likert-scale ranging from &#x0201C;never&#x0201D; to &#x0201C;daily.&#x0201D; The original instrument shows an acceptable fit in a confirmatory factor analysis and a two-factor structure. Our shorter adapted form has acceptable reliability (&#x003C9; &#x0003D; 0.79).</p>
</sec>
</sec>
<sec>
<title>Statistical analysis</title>
<sec>
<title>Data cleaning</title>
<p>To appropriately test the hypotheses and address cross-cultural comparability, we took several steps in terms of data cleaning and scoring prior to calculating the models. For data cleaning in both studies, we applied the following <italic>a priori</italic> standards to decrease noise in the data. Noise in the data can be due to inattentive participants or participants that are not willing to or fail to follow instructions (cf. Oppenheimer et al., <xref ref-type="bibr" rid="B67">2009</xref>; Maniaci and Rogge, <xref ref-type="bibr" rid="B53">2014</xref>). Noise can lead to low variance or indicate inappropriate response patterns in the data (e.g., participants always selecting the same response option), and it can lead to consistent order violations of the <italic>AVs</italic> due to the inattentive reading of instructions. Therefore, we removed participants with:
<list list-type="order">
<list-item><p>more than 10% missing entries in the data</p></list-item>
<list-item><p>low variance (&#x0003C;0.5) in answering the self-report questionnaires and <italic>AVs</italic></p></list-item>
<list-item><p>inappropriate response patterns in the <italic>AVs</italic></p></list-item>
<list-item><p>consistent order violations in the <italic>AVs</italic></p></list-item>
</list></p>
<p>18.7% of the original <italic>N</italic> &#x0003D; 520 participants in Sample 1 (Rwanda) and 28.5% of the original <italic>N</italic> &#x0003D; 200 participants in Sample 2 (Philippines) were removed in accordance to these criteria. Most participants were removed because of omissions in the data, which mostly originated from the paper-pencil versions.</p>
</sec>
</sec>
<sec>
<title>Analyzing the anchoring vignettes</title>
<p>In our study, we used a set of three vignettes varying in intensity to adjust self-report responses using a non-parametric approach. Specifically, responses to the self-report questionnaires were compared against the responses to the vignettes (see Table <xref ref-type="table" rid="T2">2</xref> for all examples relating a single self-report answer to a set of three vignettes). In this process, the responses of the original 5-point Likert-scales spread across a new 7-point Likert-scale. These adjusted answers are hypothetically free of some DIF forms and can thus be analyzed and interpreted like any other Likert-scale (King et al., <xref ref-type="bibr" rid="B49">2004</xref>; Wand, <xref ref-type="bibr" rid="B95">2013</xref>).</p>
<p>Table <xref ref-type="table" rid="T3">3</xref> shows the <italic>AV</italic>-adjusted scores if the participant rates the vignettes in the defined order. Of course, participants show individual differences in their ratings of these vignettes. For example, it is possible for participants to evaluate two vignettes equally if they decide that two hypothetical persons display the same intensity of a trait. Hence, they do not distinguish between two or even three vignettes (e.g., Vignette 1 &#x0003D; Vignette 2 &#x0003C; Vignette 3), which is referred to as tying one or more vignette pairs. Alternatively, participants can rate the vignettes as having a different intensity than originally defined, such as rating Vignette 2 as lower than Vignette 1 (Vignette 2 &#x0003C; Vignette 1 &#x0003C; Vignette 3) when the correct order is Vignette 1 &#x0003C; Vignette 2 &#x0003C; Vignette 3.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Rules for recoding self-report responses with three <italic>AVs</italic>.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Relative order ratings</bold></th>
<th valign="top" align="center"><bold>Adjusted score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Self &#x0003C; Vignette 1 &#x0003C; Vignette 2 &#x0003C; Vignette 3</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">Self &#x0003D; Vignette 1 &#x0003C; Vignette 2 &#x0003C; Vignette 3</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">Vignette 1 &#x0003C; Self &#x0003C; Vignette 2 &#x0003C; Vignette 3</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">Vignette 1 &#x0003C; Self &#x0003D; Vignette 2 &#x0003C; Vignette 3</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">Vignette 1 &#x0003C; Vignette 2 &#x0003C; Self &#x0003C; Vignette 3</td>
<td valign="top" align="center">5</td>
</tr>
<tr>
<td valign="top" align="left">Vignette 1 &#x0003C; Vignette 2 &#x0003C; Self &#x0003D; Vignette 3</td>
<td valign="top" align="center">6</td>
</tr>
<tr>
<td valign="top" align="left">Vignette 1 &#x0003C; Vignette 2 &#x0003C; Vignette 3 &#x0003C; Self</td>
<td valign="top" align="center">7</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x0201C;<italic>Self&#x0201D; represents a single self-report answer about one&#x00027;s own personal traits. Vignettes 1, 2, and 3 are the corresponding vignette set measuring the same trait, with the trait level highest in Vignette 3, followed by Vignette 2, with the trait level lowest in Vignette 1</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>These ties and order violations add complexity to analyses resulting in fragmentary information (Hopkins and King, <xref ref-type="bibr" rid="B37">2010</xref>). If the participant orders the vignettes in the correct order, the non-parametric adjustment through <italic>AV</italic>s eventuate in a single value (see Table <xref ref-type="table" rid="T2">2</xref>). Ties and order violations instead result in an interval solution and therefore in a set of values (King et al., <xref ref-type="bibr" rid="B49">2004</xref>). In our example, the interval can range from one to seven. The non-parametric approach has only a limited range of options for dealing with ties and order violations in <italic>AV</italic>s (Paccagnella, <xref ref-type="bibr" rid="B69">2013</xref>). Previous research has shown that choosing the lower bound of these intervals as an <italic>AV</italic>-adjusted answer leads to improved reliabilities (Kyllonen and Bertling, <xref ref-type="bibr" rid="B52">2014</xref>).</p>
<p>It is important that the assumptions of vignette equivalence and response consistency are met before evaluating <italic>AV</italic>-adjusted scores (King et al., <xref ref-type="bibr" rid="B49">2004</xref>). <italic>Vignette equivalence</italic> means that every participant perceives the <italic>AV</italic>s in the same way and therefore with the same ranking (Grol-Prokopczyk et al., <xref ref-type="bibr" rid="B28">2015</xref>), which should generally occur. This assumption would be violated if a large proportion of participants systematically interpret the vignettes in another way. Violation of this assumption leads to problems in non-parametric adjustments and incomparable thresholds in the parametric approach. In previous literature, this assumption was either assumed prima facie (Grol-Prokopczyk et al., <xref ref-type="bibr" rid="B28">2015</xref>) or considered to be met by simply looking at the consistencies when rank-ordering the vignettes (King et al., <xref ref-type="bibr" rid="B49">2004</xref>; Kristensen and Johansson, <xref ref-type="bibr" rid="B51">2008</xref>; Rice et al., <xref ref-type="bibr" rid="B74">2011</xref>). However, this assumption can be assessed by analyzing the amount of order violations within the <italic>AV</italic>s rank-order, with 10% or more indicating a significant amount of order violations. Generally, order violations are treated as measurement error. However, patterns in order violations can also have a diagnostic impact, providing information about the sample, the translation, or the quality of the vignette. For example, the World Health Organization (WHO) self-care vignettes show an order violation of 35.71% compared to an average 10% order violation for the other WHO vignettes. Systematic order violations can be also due to isolated &#x0201C;bad vignettes&#x0201D; (Grol-Prokopczyk et al., <xref ref-type="bibr" rid="B28">2015</xref>). Only if there are patterns in order violations is it problematic to analyze <italic>AV</italic>s based on the non-parametric approach (described later).</p>
<p>The assumption of <italic>response consistency</italic> tests the extent to which participants use the same thresholds for answering the self-report items and the <italic>AV</italic>s. Response consistency is violated if participants apply alternative standards to the self-report items and to the <italic>AVs</italic> or use varying standards in answering the <italic>AV</italic>s. Violations of this assumption lead to problems in adjusting self-report responses with the non-parametric approach (Grol-Prokopczyk et al., <xref ref-type="bibr" rid="B28">2015</xref>). There are options to test response consistency, such as comparing the thresholds of the <italic>AV</italic>s and the self-report items collected in different waves (Kapteyn et al., <xref ref-type="bibr" rid="B45">2011</xref>) or comparing thresholds of objective measures and self-report measures with responses to the <italic>AV</italic>s (Gupta et al., <xref ref-type="bibr" rid="B29">2010</xref>; Soest et al., <xref ref-type="bibr" rid="B81">2011</xref>; Hirve et al., <xref ref-type="bibr" rid="B34">2013</xref>). However, these options are often not available. Instead, another possibility is simply examining the IRT parameters, specifically the overlap of the threshold confidence intervals in a graded response model (i.e., a mathematical model for ordered polytomous categories) for the <italic>AVs</italic>. Mostly, this assumption is not examined but is instead assessed indirectly through interpreting the plausibility of the study results (King et al., <xref ref-type="bibr" rid="B49">2004</xref>; Grol-Prokopczyk, <xref ref-type="bibr" rid="B27">2014</xref>).</p>
<p>In our study, <italic>AVs</italic> were analyzed with the <italic>anchors</italic> package in R Studio version 3.3.2 (Wand et al., <xref ref-type="bibr" rid="B96">2011</xref>). Using this package, we assessed entropy (King and Wand, <xref ref-type="bibr" rid="B48">2007</xref>), which is an indicator of the informativeness of a given <italic>AV</italic> set. These statistics showed that all vignette sets, including the three vignettes in their defined order, were mostly informative. Next, we applied the non-parametric approach and calculated the <italic>AV</italic>-adjusted scores. Table <xref ref-type="supplementary-material" rid="SM1">S1</xref> in Supplemental Material shows an example of R-Code syntax used for the analysis of the <italic>AV</italic>s. Based on recommendations in the literature, we treated order violations as ties and chose the lower bound of the intervals (Kyllonen and Bertling, <xref ref-type="bibr" rid="B52">2014</xref>). Figures <xref ref-type="fig" rid="F1">1</xref>, <xref ref-type="fig" rid="F2">2</xref> show the means for the BFI-44 items for the original 5-point Likert-scale before the <italic>AV</italic>-adjustment and the means of the 7-point Likert-scale after the <italic>AV</italic>-adjustment, separated by country.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Scatterplots of the original and the <italic>AV</italic>-adjusted scales for Rwanda (<italic>N</italic> &#x0003D; 423) and the Philippines (<italic>N</italic> &#x0003D; 143).</p></caption>
<graphic xlink:href="fpsyg-09-00325-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Same as Figure <xref ref-type="fig" rid="F1">1</xref>.</p></caption>
<graphic xlink:href="fpsyg-09-00325-g0002.tif"/>
</fig>
</sec>
<sec>
<title>Computational approach</title>
<p>We assessed DIF through a multiple-group confirmatory factor analysis, establishing configural measurement invariance on the Big Five factor model between both samples. Then, several indices and programs were used to evaluate support for our hypotheses.</p>
<p>The first hypothesis was tested by applying McDonald&#x00027;s &#x003C9; (McDonald, <xref ref-type="bibr" rid="B58">1999</xref>), an estimate of general factor saturation that is considered a better indicator of reliability than Cronbach&#x00027;s &#x003B1; (Zinbarg et al., <xref ref-type="bibr" rid="B99">2005</xref>), and the confidence intervals of &#x003C9; were interpreted to evaluate whether reliability significantly improved after AV-adjustment.</p>
<p>For testing Hypothesis 2, we used a graded response model for the original and the <italic>AV</italic>-adjusted scores. The graded response model belongs to the polytomous item response theory models and can be applied for ordinal manifest variables. Reise and Waller (<xref ref-type="bibr" rid="B72">1990</xref>) showed that two-parameter logistic <italic>IRT</italic> models can be used for multidimensional data, which describes Personality questionnaire data. The graded response model was fitted with the multidimensional item response theory (full-information item factor analysis) <italic>MIRT</italic> package in R Studio version 3.3.2 (Chalmers, <xref ref-type="bibr" rid="B14">2012</xref>).</p>
<p>DIF and Hypothesis 3 were assessed based on a confirmatory factor analysis with the BFI items as manifest indicators of their respective Big Five Personality factor, which were allowed to correlate. Previous studies have reported problems in modeling personality self-report questionnaires in a confirmatory factor analysis (Marsh et al., <xref ref-type="bibr" rid="B55">2006</xref>). For example, the NEO-PI-R is known for encountering problems such as model misfit, negative item loadings, and high error correlations (e.g., Borkenau and Ostendorf, <xref ref-type="bibr" rid="B9">1990</xref>; McCrae et al., <xref ref-type="bibr" rid="B57">1996</xref>). A common technique for improving model fit in such a situation involves deleting items that are not loading high enough onto the corresponding factor (Aluja et al., <xref ref-type="bibr" rid="B3">2006</xref>; Tully et al., <xref ref-type="bibr" rid="B88">2011</xref>). Based on this procedure, items that were completely misfitting the model through small or even negative loadings were deleted from the original measure. In this study, this procedure resulted in a 36-item version of the BFI. To ensure comparability, all models are based on a 36-item solution. In the evaluation of Hypothesis 3, we modeled the 36-item solution of the BFI-44 for the unadjusted self-report scores [Model &#x00023;1 (Rwanda) and &#x00023;3 (Philippines)] and for the <italic>AV</italic>-adjusted scores [Model &#x00023;2 (Rwanda) and &#x00023;4 (Philippines)]. We then compared improvement in fit for each country (Model &#x00023;1 vs. Model &#x00023;2 and Model &#x00023;3 vs. Model &#x00023;4). A final model (Model &#x00023;5), which includes all participants and is based on the <italic>AV</italic>-adjusted scores, was estimated to evaluate whether the <italic>AV</italic>-adjusted scores provide stronger support for the Five Factor structure (i.e., orthogonal factorially-pure scales). As all models are based on the same factor structure based on the same 36 items, models can be compared by looking at improvement in the fit indices.</p>
<p>In a confirmatory factor analysis, several indices can be used to describe the fit between the theoretical model and the actual model. We used the criteria that a Comparative-Fit-Index (CFI) (Bentler, <xref ref-type="bibr" rid="B7">1990</xref>) greater than or equal to 0.90 and a Root Mean Square Error of Approximation (RMSEA) (Steiger, <xref ref-type="bibr" rid="B84">1990</xref>) less than or equal to 0.08 indicates acceptable fit (Steiger, <xref ref-type="bibr" rid="B84">1990</xref>). The confirmatory factor analyses were conducted with the either Mplus 7 (Muth&#x000E9;n and Muth&#x000E9;n, <xref ref-type="bibr" rid="B63">1998-2015</xref>) or the <italic>lavaan</italic> package in R Studio version 3.3.2 (Rosseel, <xref ref-type="bibr" rid="B77">2012</xref>). Basic statistics are based on the <italic>psych</italic> package (Revelle, <xref ref-type="bibr" rid="B73">2014</xref>) and the violin plots in Figure <xref ref-type="fig" rid="F3">3</xref> are based on the package <italic>vioplot</italic> (Adler, <xref ref-type="bibr" rid="B1">2005</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Violin plots of discrimination parameters of the original self-report questions and the <italic>AV</italic>-adjusted answers for Rwanda and Philippines.</p></caption>
<graphic xlink:href="fpsyg-09-00325-g0003.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<sec>
<title>Evaluating DIF</title>
<p>First, we tested the extent to which the original scores were affected by DIF. To this end, we conducted a multiple-group confirmatory factor analysis, with the above described fit standards for fit indices. In the first model, with all 44 items, imposing configural measurement invariance led to poor model fit (<italic>CFI</italic> &#x0003D; 0.60, <italic>RMSEA</italic> &#x0003D; 0.06). This indicates that the underlying five-factor model is not comparable between samples, and different constructs are represented. Likewise, there were large differences between the factor loadings, which ranged from &#x02212;0.55 to 0.78. Based on these results, we conclude that configural invariance between both studies on the full scale is not met and hence different constructs are assessed.</p>
<p>To compare improvement in fit from the original scores to the <italic>AV</italic>-adjusted scores, we conducted a second multiple-group confirmatory factor analysis based on the original self-report scores, but with the 36-item solution (Baseline Model: <italic>CFI</italic> &#x0003D; 0.67, <italic>RMSEA</italic> &#x0003D; 0.07). The poor model fit supports our initial conclusion that the original self-report items are affected by DIF and need an <italic>AV</italic>-adjustment to achieve cross-cultural comparability.</p>
</sec>
<sec>
<title>Vignette equivalence and response consistency</title>
<p>Next, before interpreting the AV-adjusted scores, we checked vignette equivalence and response consistency assumptions. The percentage of vignettes with the correct order ranged from 26 to 66% (Rwanda) and 37 to 72% (Philippines; see Table <xref ref-type="table" rid="T4">4</xref> for the percentage of all possible orders for each domain). In our study, the chance of randomly violating the correct order is 74% while the chance of randomly correctly ordering the <italic>AV</italic>s is 8%. Violations of 10% or lower can be treated as measurement error, however systematic violations (e.g., participants using answer pattern resulting in order violations) have to be excluded. Conscientiousness, Agreeableness, and Extraversion for both samples, as well as Openness for the Philippines sample, had order violations of 10% or less, indicating no problematic vignettes. However, violations for Neuroticism (Rwanda: 15%; Philippines: 18%) and Openness for the Rwanda sample (23%) were higher. For the latter, Openness displayed a higher amount of ties (51%) than correct orders (26%), in addition to the large percentage of order violations. Given there was only a partial violation of the vignette equivalence assumption, we felt comfortable continuing in the analyses. Possible reasons for these violations, as well as examples of other research where the same violations were found, is discussed in the Discussion section.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Percentage of correctly ordered vignettes, vignette ties, and order violations for each Big Five Personality factor, and the percentage of random chance ordering.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold>Correct order (1 &#x0003C; 2 &#x0003C; 3) random chance of correct order &#x0003D; 8%</bold></th>
<th valign="top" align="center"><bold>Ties (e.g., 1 &#x0003D; 2 &#x0003C; 3) random chance of ties &#x0003D; 18%</bold></th>
<th valign="top" align="center"><bold>Order violations (e.g., 2 &#x0003C; 1 &#x0003C; 3) random chance of order violation &#x0003D; 74%</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="4" style="background-color:#bbbdc0"><bold>CONSCIENTIOUSNESS</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda</td>
<td valign="top" align="center">66%</td>
<td valign="top" align="center">30%</td>
<td valign="top" align="center">4%</td>
</tr>
<tr>
<td valign="top" align="left">Philippines</td>
<td valign="top" align="center">72%</td>
<td valign="top" align="center">18%</td>
<td valign="top" align="center">10%</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4" style="background-color:#bbbdc0"><bold>AGREEABLENESS</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda</td>
<td valign="top" align="center">61%</td>
<td valign="top" align="center">33%</td>
<td valign="top" align="center">6%</td>
</tr>
<tr>
<td valign="top" align="left">Philippines</td>
<td valign="top" align="center">52%</td>
<td valign="top" align="center">43%</td>
<td valign="top" align="center">5%</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4" style="background-color:#bbbdc0"><bold>NEUROTICISM</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda</td>
<td valign="top" align="center">42%</td>
<td valign="top" align="center">43%</td>
<td valign="top" align="center">15%</td>
</tr>
<tr>
<td valign="top" align="left">Philippines</td>
<td valign="top" align="center">37%</td>
<td valign="top" align="center">45%</td>
<td valign="top" align="center">18%</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4" style="background-color:#bbbdc0"><bold>OPENNESS</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda</td>
<td valign="top" align="center">26%</td>
<td valign="top" align="center">51%</td>
<td valign="top" align="center">23%</td>
</tr>
<tr>
<td valign="top" align="left">Philippines</td>
<td valign="top" align="center">49%</td>
<td valign="top" align="center">41%</td>
<td valign="top" align="center">10%</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4" style="background-color:#bbbdc0"><bold>EXTRAVERSION</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda</td>
<td valign="top" align="center">47%</td>
<td valign="top" align="center">44%</td>
<td valign="top" align="center">8%</td>
</tr>
<tr>
<td valign="top" align="left">Philippines</td>
<td valign="top" align="center">56%</td>
<td valign="top" align="center">34%</td>
<td valign="top" align="center">10%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Response consistency was evaluated through an examination of whether or not respondents used the same thresholds while answering the <italic>AVs</italic> and the self-report questionnaire. We fitted a graded response model for the <italic>AVs</italic> and the self-report questions and compared the four thresholds of the <italic>AVs</italic> against the corresponding self-reports. The confidence intervals of most thresholds overlapped. Therefore, we consider the Response Consistency requirement met.</p>
</sec>
<sec>
<title>Hypothesis testing</title>
<p>Next, we applied the non-parametric approach (King et al., <xref ref-type="bibr" rid="B49">2004</xref>) on the data to produce <italic>AV</italic>-adjusted scores. These <italic>AV</italic>-adjusted scores are what was compared with the original self-report answers.</p>
<list list-type="simple">
<list-item><p><italic>Hypothesis 1:</italic> Reliability before and after AV-adjustment</p></list-item>
</list>
<p>We compared McDonald&#x00027;s &#x003C9; (McDonald, <xref ref-type="bibr" rid="B58">1999</xref>), a measure of reliability, for the original and <italic>AV</italic>-adjusted scores (see Table <xref ref-type="table" rid="T5">5</xref>). For the original scores, omega indicated poor reliability (&#x003C9; &#x0003D; 0.32 &#x02212;0.66) for all dimensions except conscientiousness (&#x003C9; &#x0003D; 0.74). Following the <italic>AV</italic>-adjustment, the omegas increased for every dimension in both samples to acceptable or good (&#x003C9; &#x0003D; 0.80 &#x02212;0.92). Looking at the 95% confidence interval, we see that the intervals of the omega estimates for original and the <italic>AV</italic>-adjusted scores do not overlap. In sum, the analysis shows that the <italic>AV</italic>-adjusted scales show better reliability than the original self-report scales, supporting hypothesis 1.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Scale reliability, indicated by McDonald&#x00027;s omega, and 95% Confidence Intervals for the original and <italic>AV</italic>-adjusted Big Five Personality factors.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Rwanda</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Philippines</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>Original</bold></th>
<th valign="top" align="center"><bold><italic>AV</italic>-adjusted</bold></th>
<th valign="top" align="center"><bold>Original</bold></th>
<th valign="top" align="center"><bold><italic>AV</italic>-adjusted</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Conscientiousness</td>
<td valign="top" align="center">0.74 [0.69 &#x02212;0.78]</td>
<td valign="top" align="center">0.92 [0.90 &#x02212;0.93]</td>
<td valign="top" align="center">0.76 [0.64 &#x02212;0.83]</td>
<td valign="top" align="center">0.89 [0.84 &#x02212;0.92]</td>
</tr>
<tr>
<td valign="top" align="left">Agreeableness</td>
<td valign="top" align="center">0.32 [0.21 &#x02212;0.42]</td>
<td valign="top" align="center">0.87 [0.83 &#x02212;0.90]</td>
<td valign="top" align="center">0.63 [0.45 &#x02212;0.80]</td>
<td valign="top" align="center">0.89 [0.85 &#x02212;0.92]</td>
</tr>
<tr>
<td valign="top" align="left">Neuroticism</td>
<td valign="top" align="center">0.66 [0.60 &#x02212;0.72]</td>
<td valign="top" align="center">0.80 [0.76 &#x02212;0.83]</td>
<td valign="top" align="center">0.63 [0.50 &#x02212;0.70]</td>
<td valign="top" align="center">0.82 [0.77 &#x02212;0.87]</td>
</tr>
<tr>
<td valign="top" align="left">Openness</td>
<td valign="top" align="center">0.66 [0.61 &#x02212;0.71]</td>
<td valign="top" align="center">0.91 [0.89 &#x02212;0.92]</td>
<td valign="top" align="center">0.57 [0.46 &#x02212;0.66]</td>
<td valign="top" align="center">0.88 [0.84 &#x02212;0.91]</td>
</tr>
<tr>
<td valign="top" align="left">Extraversion</td>
<td valign="top" align="center">0.51 [0.39 &#x02212;0.59]</td>
<td valign="top" align="center">0.81 [0.77 &#x02212;0.84]</td>
<td valign="top" align="center">0.62 [0.48 &#x02212;0.71]</td>
<td valign="top" align="center">0.82 [0.77 &#x02212;0.88]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>McDonald&#x00027;s omega and 95% confidence interval of omega for a model with five correlated factors. Original, original (5-point) self-report Likert-scale (before the AV-adjustment); AV-adjusted, 7-point Likert-scale (DIF-free, after the AV-adjustment)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<list list-type="simple">
<list-item><p><italic>Hypothesis 2:</italic> Item functioning before and after AV-adjustment</p></list-item>
</list>
<p>Next, we investigated the item and category function of the original and the <italic>AV</italic>-adjusted scores. Here, we assume that the DIF-free scores increase the discriminant power, as compared to the original self-report scores, because they are comparable between countries and more relevant to the measured trait in both studies. The violin plot in Figure <xref ref-type="fig" rid="F3">3</xref> shows that, for both studies, the discrimination parameters mostly increase as a result of the <italic>AV</italic>-adjustment.</p>
<p>The discriminant power also increases as the range of the threshold widens. The original self-report scores are based on a 5-point Likert-scale while the <italic>AV</italic>-adjusted scores use a 7-point Likert-scale. This results in a different number of thresholds: b<sub>1</sub> to b<sub>4</sub> (original 5-point Likert-scale) and b<sub>1</sub> to b<sub>6</sub> (7-point Likert-scale for the <italic>AV</italic>-adjusted answers). As b<sub>1</sub> to b<sub>6</sub> span a wider range of values, this indicates that the <italic>AV</italic>-adjusted scores differentiate higher levels of proficiency compared to the original self-report scores. In particular, for the original self-report scores, the last threshold was more effective in differentiation relative to the last category (b<sub>4</sub>). As we compared four to six thresholds, we examined correlations between the thresholds of the original and the <italic>AV</italic>-adjusted scores. This comparison shows that most thresholds are highly correlated with one another (<italic>r</italic>s &#x0003D; 0.47 to 0.94). Only the last threshold of the <italic>AV</italic>-adjusted scores (b<sub>6</sub>) shows negative correlations with the thresholds of the original self-report scores.</p>
<p>Next, we evaluated the overall test information (i.e., the degree of certainty of the proficiency estimates) for the original and <italic>AV</italic>-adjusted scores. The test information curve includes &#x003B8;-levels from &#x02212;6 to 6. For the original self-report scores, the test is mostly informative for &#x003B8; &#x0003C; 0. Above zero, however, the standard error increases greatly and the test information decreases. For the <italic>AV</italic>-adjusted scores, the test is less informative for &#x003B8; &#x0003C; 0, but it is still informative for &#x003B8; &#x0003E; 0, particularly between 2 and 4.</p>
<p>Overall, we conclude that hypothesis 2 was supported.</p>
<list list-type="simple">
<list-item><p><italic>Hypothesis 3:</italic> Big five factor structure before and after AV-adjustment</p></list-item>
</list>
<p>Next, we tested our hypothesis that the <italic>AV</italic>-adjusted scores will have improved model fit and a better loading pattern in a confirmatory factor analysis, relative to the original scores, thus showing better support for the Big Five factor structure.</p>
<p>Table <xref ref-type="table" rid="T6">6</xref> shows the results of the confirmatory factor analysis, based on the 36-item version of the BFI. As expected, the models based on the original self-report scores [Model &#x00023;1 (Rwanda) and &#x00023;3(Philippines)] have poor fit to the data, yielding a non-positive definite covariance matrix, with item loadings weak in magnitude, not significant, or even negative. These models are not clearly identified but are displayed here for purposes of comparison. In sum, Model &#x00023;1 and &#x00023;3 do not support the Big Five factor structure.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Confirmatory Factor Analysis model fit estimates based on the original and <italic>AV</italic>-adjusted scores.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>&#x00023;</bold></th>
<th valign="top" align="left"><bold>Model type</bold></th>
<th valign="top" align="center"><bold>CFI</bold></th>
<th valign="top" align="center"><bold>RMSEA</bold></th>
<th valign="top" align="center"><bold>&#x003C7;<sup>2</sup></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="left">Rwanda: Original<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">&#x003C7;<sub>2</sub>(584) &#x0003D; 1,479</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="left">Rwanda: <italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">&#x003C7;<sup>2</sup><sub>(584)</sub> &#x0003D; 1,423</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">Philippines: Original<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">&#x003C7;<sup>2</sup><sub>(584)</sub> &#x0003D; 1,165</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="left">Philippines: <italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">&#x003C7;<sup>2</sup><sub>(584)</sub> &#x0003D; 1,034</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">Rwanda and Philippines Combined: <italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center">&#x003C7;<sup>2</sup><sub>(584)</sub> &#x0003D; 1,640</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>All models are with the 36-item version. &#x0002B; These models have a non-positive definite covariance matrix, indicating they are not clearly identified; they are displayed her merely for comparison purposes</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>The models including the <italic>AV</italic>-adjusted scores (Model &#x00023;2, &#x00023;4, and &#x00023;5) show high item loadings and better model fit then Models &#x00023;1 and &#x00023;3. However, while Model &#x00023;2 (Rwanda) and &#x00023;5 (Rwanda and Philippines combined) show acceptable fit, the fit of Model &#x00023;4 (Philippines) is not acceptable. For Model &#x00023;5, the model fit improved over the original Baseline Model reported above, which was based on the original scores (<italic>CFI</italic> &#x0003D; 0.67, <italic>RMSEA</italic> &#x0003D; 0.07). Therefore, we assume that the studies are now comparable. In Model &#x00023;5, Neuroticism is negatively correlated with every other dimension (<italic>r</italic> &#x0003D; &#x02212;0.17 to &#x02212;0.31) and all other dimensions are weakly and positively correlated with one another (<italic>r</italic> &#x0003D; 0.13 to 0.31). Thus, the correlations of the <italic>AV</italic>-adjusted scores are much more in line with previous findings than the correlations for the original self-report scores. In sum, the confirmatory factor analysis shows that the <italic>AV</italic>-adjusted scores better support the original factor structure of the Big Five, supporting hypothesis 3.</p>
<list list-type="simple">
<list-item><p><italic>Hypothesis 4:</italic> Test-criterion relationships</p></list-item>
</list>
<p>Finally, we evaluated test-criterion relations with external outcome variables: satisfaction with life and counterproductive behavior (either at school or at work). The outcome variables were correlated with the Big Five Personality factors before and after the <italic>AV</italic>-adjustment (see Table <xref ref-type="table" rid="T7">7</xref>). The correlations of counterproductive behavior with Conscientiousness, Agreeableness and Neuroticism, decreased significantly after the <italic>AV</italic>-adjustment. However, for life satisfaction, with two exceptions, correlations were not significantly different after the <italic>AV</italic>-adjustment. The two exceptions are for the Rwanda sample where life satisfaction is negatively correlated to Openness and Conscientiousness after AV-adjustment, but unrelated before the adjustment. Overall, hypothesis 4, which stated correlations with satisfaction with life and counterproductive behavior should be stronger after <italic>AV</italic>-adjustment, was not supported.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Correlations (Spearman rho) of the Big Five Personality factors, based on the original and <italic>AV</italic>-adjusted scores, with two outcome variables.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold>Life satisfaction</bold></th>
<th valign="top" align="center"><bold>Counterproductive behavior</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="3" style="background-color:#bbbdc0"><bold>CONSCIENTIOUSNESS</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center"><italic>&#x02212;0.01/&#x02212;0.12</italic><xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.25<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref>/&#x02212;0.19<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref></td>
</tr>
<tr>
<td valign="top" align="left">Philippines: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.07/0.05</td>
<td valign="top" align="center"><italic>&#x02212;0.51<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref>/&#x02212;0.30</italic><xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref></td>
</tr>
<tr>
<td valign="top" align="left" colspan="3" style="background-color:#bbbdc0"><bold>AGREEABLENESS</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">&#x02212;0.03/&#x02212;0.04</td>
<td valign="top" align="center"><italic>0.22<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref>/&#x02212;0.13</italic><xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref></td>
</tr>
<tr>
<td valign="top" align="left">Philippines: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.04/&#x02212;0.02</td>
<td valign="top" align="center"><italic>&#x02212;0.50<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;</sup></xref>/&#x02212;0.19</italic><xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref></td>
</tr>
<tr>
<td valign="top" align="left" colspan="3" style="background-color:#bbbdc0"><bold>NEUROTICISM</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">&#x02212;0.01/&#x02212;0.01</td>
<td valign="top" align="center">0.11<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref>/0.07</td>
</tr>
<tr>
<td valign="top" align="left">Philippines: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.08/0.07</td>
<td valign="top" align="center"><italic>0.24</italic><xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref><italic>/0.12</italic></td>
</tr>
<tr>
<td valign="top" align="left" colspan="3" style="background-color:#bbbdc0"><bold>OPENNESS</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center"><italic>&#x02212;0.04/&#x02212;0.14</italic><xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.10<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref>/&#x02212;0.12<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref></td>
</tr>
<tr>
<td valign="top" align="left">Philippines: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.07/&#x02212;0.14</td>
<td valign="top" align="center">&#x02212;0.05/0.05</td>
</tr>
<tr>
<td valign="top" align="left" colspan="3" style="background-color:#bbbdc0"><bold>EXTRAVERSION</bold></td>
</tr>
<tr>
<td valign="top" align="left">Rwanda: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.03/&#x02212;0.05</td>
<td valign="top" align="center">0.01/0.01</td>
</tr>
<tr>
<td valign="top" align="left">Philippines: Original/<italic>AV</italic>-adjusted</td>
<td valign="top" align="center">0.15/0.13</td>
<td valign="top" align="center">&#x02212;0.02/0.03</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN1">
<label>&#x0002A;&#x0002A;</label>
<p><italic>p &#x0003C; 0.01</italic>;</p></fn>
<fn id="TN2">
<label>&#x0002A;</label>
<p><italic>p &#x0003C; 0.05</italic>.</p></fn>
<p><italic>Correlations presented in italics mean there is a significant difference in relations with that personality factor depending on whether the original or AV-adjusted scoring was used</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>Overall, the results suggest that the <italic>AV</italic> methodology is an appropriate tool for cross-cultural research, although the change in test-criterion relationships warrants further investigation.</p>
<sec>
<title>Summary and interpretation of the results</title>
<p>We demonstrated that not even the weakest degree of measurement invariance, configural invariance, was present across both countries. Hence, we show that the BFI-44 test and its items are affected by cross-cultural DIF. Because we performed several backward and forward translation checks, we presume our items showed no evidence of any translation problems, and that the observed DIF most likely occurred because participants from different countries displayed different probabilities of item endorsement. Thus, we utilized <italic>AVs</italic> to correct for DIF.</p>
<p>Before interpreting the AV-adjusted scores, we showed that the <italic>AV</italic>s mostly met the two basic assumptions of vignette equivalence and response consistency. On average, 89% of the participants ordered all vignettes correctly or rated them as ties. Thus, we can assume there was generally vignette equivalence, with most participants perceiving the vignettes in the same way and in the defined order. However, Neuroticism and Openness showed a non-negligible amount of order violations. A possible reason for the order violation on the Neuroticism <italic>AVs</italic> is that they were presented on a reversed scale. In previous studies, the use of a reversed scale also created confusion for participants (He et al., <xref ref-type="bibr" rid="B32">2017</xref>). The order violations within Openness are probably due to the conceptualization of the Openness factor and the corresponding <italic>AV</italic>-set. Primi et al. (<xref ref-type="bibr" rid="B71">2016</xref>) discovered similar results and suggested that this is because Openness is different from the other domains and is not as easy to rate as it mostly includes not observable behaviors compared to the other domains, such as Conscientiousness, which has observable behaviors. Likewise, given response consistency is traditionally difficult to confirm or disconfirm, we proposed a novel statistical solution and tested it with the present data set. We showed that the confidence intervals of most thresholds overlapped, implying that participants applied the same thresholds in answering <italic>AVs</italic> and the original self-report questions. Thus, we could conclude that the response consistency assumption was met.</p>
<p>We examined how the <italic>AV</italic>s influenced other psychometric characteristics by testing a series of hypotheses. In evaluation of the first hypothesis, we examined scale reliability, estimated through omega, for each of the Big Five factors before and after <italic>AV</italic>-adjustment. We found support for this hypothesis such that there was a higher internal consistency for the <italic>AV</italic>-adjusted scores indicating better measurement properties.</p>
<p>Hypothesis 2 was also confirmed: the discrimination parameters based on the AV-adjusted scores were larger, while the thresholds spanned a wider range. In sum, the <italic>AV</italic>-adjustment increased the overall test information, adding power and precision to the test. This finding again suggests that the <italic>AVs</italic> are very beneficial from a psychometric perspective.</p>
<p>To test Hypothesis 3, we assumed that the <italic>AV</italic>-adjusted scores provided clearer support for the Big Five Factor structure of personality, compared to the original scores, which were DIF-affected items and failed to show measurement invariance. Overall, we found the <italic>AV</italic>-adjusted scores better predict the estimated level of the latent factor, including more reliable and factorially-pure scales aligned with the Big Five Factor structure. They are therefore more in line with previous findings (mostly based on exploratory factor analysis) regarding the Big Five factor structure (notably obtained with samples that are more homogenous culturally than the two chosen for the present investigation). It has to be noted that Model &#x00023;4, based on the <italic>AV</italic>-adjusted scores for the Filipino sample, did not have acceptable fit. However, with personality data, a confirmatory factor analysis is not always desirable due to specific characteristics of the data (e.g., model complexity; Hopwood and Donnellan, <xref ref-type="bibr" rid="B38">2010</xref>; Fischer, <xref ref-type="bibr" rid="B25">2014</xref>). In general, the confirmation of the third hypothesis is also in line with the findings of the graded response model (Hypothesis 2). In conclusion, all psychometric characteristics are improved following the <italic>AV</italic>-adjustment: <italic>AV</italic>-adjusted items that are DIF-free appear to improve comparability across countries.</p>
<p>For the final hypothesis (Hypothesis 4), we evaluated test-criterion relationships by examining correlations between the original and <italic>AV</italic>-adjusted scores with two external outcome measures. In our results, Conscientiousness, Agreeableness, and Openness, based on both the original scores and the <italic>AV</italic>-adjusted scores, correlated negatively with counterproductive work behavior, as expected. Likewise, Neuroticism had a slightly positive correlation and Extraversion was unrelated before the <italic>AV</italic>-adjustment. However, the magnitude of these correlations was significantly lower when using the <italic>AV</italic>-adjusted scores. Both studies showed weak non-significant correlations of the Big Five with satisfaction with life. Looking at these correlations, we see that relations of satisfaction with life with Openness and Conscientiousness are significantly different depending on whether <italic>AV</italic>-adjustment is used or not. Agreeableness, Neuroticism, and Extraversion showed no significant variation before and after the <italic>AV</italic>-adjustment. Especially unexpected is the negative correlation of satisfaction with life with Conscientiousness after the <italic>AV</italic>-adjustment. However, these results are in line with previous findings in Iranian samples (Hosseinkhanzadeh and Taher, <xref ref-type="bibr" rid="B39">2013</xref>). Notably, the unexpected drop in correlation magnitude of the Big Five with both outcome variables is consistent with the findings of Ham and Roberts (<xref ref-type="bibr" rid="B30">2015</xref>) who found a similar reduction in correlation with outcome values after applying the <italic>AV</italic> methodology.</p>
<p>One possible explanation for a fairly systematic reduction in these correlations is as follows. <italic>AVs</italic> were only applied on the personality items and not on the outcome measures <italic>per se</italic>; that is, DIF was only corrected for in personality. Thus, this method triggered a decrease in the covariance as the comparison is made between DIF-free items and items that are still DIF-affected. A further explanation is that the individual&#x00027;s ranking of personality substantially moved after the <italic>AV</italic>-adjustment, since the correlation between the original self-report dimensions and the <italic>AV</italic>-adjusted dimensions were around <italic>r</italic> &#x0003D; 0.38 to 0.64. Moreover, the stability of the fairly small correlations is somewhat questionable. Some correlations drop to a non-significant level if countries are examined separately. Thus, adjusting both the predictor and outcome variables is worth considering.</p>
</sec>
<sec>
<title>Limitations</title>
<p><italic>AVs</italic> are a strong theoretical and practical tool for accommodating DIF. However, this tool faces some general limitations worth mentioning. First, Buckley (<xref ref-type="bibr" rid="B11">2008</xref>) demonstrates that context effects can bias the vignette response, as do the order of the vignettes relative to the self-report. We did not test for order effects in the current study (<italic>AV</italic>s came before the BFI-44) largely because of concerns by the local administration of having multiple forms. Nevertheless, this concern is worthy of consideration.</p>
<p>J&#x000FC;rges and Winter (<xref ref-type="bibr" rid="B44">2013</xref>) showed the importance of vignette equivalence by highlighting that vignette ratings are somewhat sensitive to the sex and age of the hypothetical person described in the vignette. Likewise, participants may apply different thresholds for male and female hypothetical scenarios affecting response consistency (Kapteyn et al., <xref ref-type="bibr" rid="B46">2007</xref>). Randomization can be used as a technique to neutralize violations of response consistency (Chan et al., <xref ref-type="bibr" rid="B15">2015</xref>). In our study, we randomized the sex of the possible descriptions, as well as the names, which could also show some relation to age groups. However, we were unable to randomize the order of the AVs, which would have allowed us to address contextual effects.</p>
<p>Our study was limited in the number of countries assessed, the regions surveyed, and the sample size within each country. In particular, this limitation prevented us from using the parametric approach (King et al., <xref ref-type="bibr" rid="B49">2004</xref>) to analyze the <italic>AVs</italic>; hence, we applied the non-parametric approach for both samples. The non-parametric approach shows limitations when dealing with order violations: inconsistencies are grouped and the non-parametric solution can only deal with scalar values, resulting in a loss of information (Paccagnella, <xref ref-type="bibr" rid="B69">2013</xref>). Hence, all non-systematic order violations in our study were treated as ties, leading to a loss of information. Most studies experience some degree of order violation. He et al. (<xref ref-type="bibr" rid="B32">2017</xref>) reported order violations ranging from 3 to 13% for facets of Conscientiousness, with the exception of the facet industriousness, which had an order violation of 30%. They concluded that all <italic>AV</italic>-sets worked well, even though presenting vignettes with two levels, instead of vignettes with three levels, expect the <italic>AV</italic>-set for the Conscientiousness facet of industriousness, which was hence excluded from the analysis (He et al., <xref ref-type="bibr" rid="B32">2017</xref>). In our study, at least two Big Five factors showed more than 10% order violations. It has been argued that it is not problematic for later interpretation when vignettes are ordered incorrectly because a participant experienced the <italic>AV</italic>s differently based on their circumstances (Wand, <xref ref-type="bibr" rid="B95">2013</xref>). Any kind of disagreement on the actual vignette order should be explored as a possible design problem and as an indication of a poor vignette. Hence, the quality of our Openness vignettes, where order violations ranged from 10% (Philippines) to 23% (Rwanda), can be improved and a revision of this vignette set, as well as the neuroticism vignettes (15% order violation in Rwanda and 18% on the Philippines), should be pursued in future studies.</p>
<p>The representativeness of the results for both countries may be restricted to the specific regions of the countries where the data was collected and may be slightly more representative of females, as both samples include a high percentage of women.</p>
</sec>
<sec>
<title>Future considerations</title>
<p>Based on these considerations, future studies should include larger samples, more countries, and more regions within countries, as this will allow the researcher to use the parametric approach. Those results could then be compared to results using a non-parametric approach. Using <italic>AV</italic>-adjustment in countries where the Big Five Factor structure is replicated and well-established would allow for an interpretation of mean-shifts in trait scores after the <italic>AV</italic>-adjustment. Also, it is important to note that the Big Five Factor structure was replicated quite well in Rwanda after the <italic>AV</italic>-adjustment. However, this result should be replicated in future research. Future studies should also further explore test-criterion relationships after <italic>AV</italic>-adjustment, particularly with a wider range of criterion variables, given our work in this area was limited.</p>
<p>Our findings are especially relevant for researchers interested in alternative or competing methods to measure personality. Our study provides insights concerning the robustness and the universality of the Big Five Personality factors. Influential articles describe the Big Five as a psychometrically sound measure that can be applied in different countries and cultures (e.g., McCrae and Terracciano, <xref ref-type="bibr" rid="B56">2005</xref>; Schmitt et al., <xref ref-type="bibr" rid="B80">2007</xref>). Self-report measures of the Big Five, which have their origin in a lexical approach, are based on principal component analysis with a varimax rotation, meaning the five factors are kept orthogonal (e.g., Tupes and Christal, <xref ref-type="bibr" rid="B89">1961</xref>). However, if the fit of the model is assessed with confirmatory factor analysis or item response theory, this structure often has a poor fit to the data and insufficient psychometric properties (Olaru et al., <xref ref-type="bibr" rid="B65">2015</xref>). The confirmatory models based on the original self-report scores in our study supports these concerns. The <italic>AV</italic>-adjustments leading to DIF-free scores show a promising solution toward a psychometrically sound measurement with interpretable, reliable, and factorially-pure scales.</p>
<p>AV-adjustment is especially relevant today given personality research is facing a debate on the comparability of results based Likert-scale response options, ranging from issues with cross-cultural comparison (He et al., <xref ref-type="bibr" rid="B32">2017</xref>) to comparability between genders (Weisberg et al., <xref ref-type="bibr" rid="B97">2011</xref>). The application of an external benchmark, like the <italic>AVs</italic>, for all Big Five dimensions is not only of interest for correcting cross-cultural bias, but rather for any kind of bias between different groups (e.g., men and women).</p>
<p>The <italic>AV</italic>s provided in Table <xref ref-type="table" rid="T2">2</xref> can be applied not only in a research context (e.g., global workforce, developmental and educational research and policymakers), but also in occupational context. For example, large international companies that base their application and selection process on assessed cognitive and non-cognitive skills can apply vignettes in order to minimize cross-cultural bias in the assessment of non-cognitive skills.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>Conclusion</title>
<p>This study is one of the first to use <italic>AV</italic> methodology to adjust all Big Five dimensions in more than one country (cf. Primi et al., <xref ref-type="bibr" rid="B71">2016</xref>). In order to use <italic>AVs</italic>, we tested and showed that we met the basic measurement assumptions. The literature regarding the utility of <italic>AV</italic>-adjustment for the assessment of personality is mixed, with researchers finding either the method is not necessary or very beneficial. We showed that personality self-reports in Rwanda and the Philippines are affected by DIF and improved with an <italic>AV</italic>-adjustment. Even though the trait-level means for the original and <italic>AV</italic>-adjusted scores were not drastically different, several psychometric characteristics were improved when <italic>AV</italic>-adjusted scores were used. In the end, the DIF-free scores led to more reliable, powerful, and precise scales that are in line with the Big Five Factor structure. However, test-criterion relations were somewhat reduced after the <italic>AV</italic>-adjustment&#x02014;a finding discussed above. This finding notwithstanding, we argue that the <italic>AV</italic>-adjustment makes personality across countries more comparable and offers a possible solution to cross-cultural comparison problems.</p>
<p>In sum, this study and its results act as an important step toward explaining and handling cross-cultural comparability problems. Overall, <italic>AVs</italic> provide a useful external benchmark for adjusting self-report scores when measuring personality. Future studies should consider implementing similar adjustments to the assessment of other psychological constructs.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>SW contributed to the conception of the study, the data analysis, and the interpretation of data for the work. SW drafted the article and finalized the version for publication. RR contributed to the conception and design of the study, and the acquisition and interpretation of the data. RR also edited the final manuscript. SW and RR agreed to be held accountable for all aspects of the work and ensure that questions related to the accuracy or integrity of any part of the work will be appropriately investigated and resolved.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>The data for this paper was made available by a project implemented by the Education Development Center, Professional Examination Services, and the Akilah Institute for Women that was funded through the Workforce Connections grant, led by FHI360, from the United States Agency for International Development (USAID). Funding for the publication costs of this paper was provided by Ulm University. We thank Sally Olderbak for English edits and helpful feedback on previous versions of this article.</p>
</ack>
<sec sec-type="supplementary-material" id="s7">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyg.2018.00325/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpsyg.2018.00325/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Adler</surname> <given-names>D.</given-names></name></person-group> (<year>2005</year>). <source>vioplot: Violin plot</source>. R package version 0.2. Available online at: <ext-link ext-link-type="uri" xlink:href="http://CRAN.R-project.org/package=vioplot">http://CRAN.R-project.org/package=vioplot</ext-link></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Allport</surname> <given-names>G. W.</given-names></name> <name><surname>Odbert</surname> <given-names>H. S.</given-names></name></person-group> (<year>1936</year>). <article-title>Trait-names: a psycho-lexical study</article-title>. <source>Psychol. Monogr.</source> <volume>47</volume>, <fpage>1</fpage>&#x02013;<lpage>178</lpage>. <pub-id pub-id-type="doi">10.2307/452250</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aluja</surname> <given-names>A.</given-names></name> <name><surname>Rossier</surname> <given-names>J.</given-names></name> <name><surname>Garc&#x000ED;a</surname> <given-names>L. F.</given-names></name> <name><surname>Angleitner</surname> <given-names>A.</given-names></name> <name><surname>Kuhlman</surname> <given-names>M.</given-names></name> <name><surname>Zuckerman</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>A cross-cultural shortened form of the ZKPQ (ZKPQ-50-cc) adapted to English, French, German, and Spanish languages</article-title>. <source>Pers. Individ. Dif.</source> <volume>41</volume>, <fpage>619</fpage>&#x02013;<lpage>628</lpage>. <pub-id pub-id-type="doi">10.1016/j.paid.2006.03.001</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Angelini</surname> <given-names>V.</given-names></name> <name><surname>Cavapozzi</surname> <given-names>D.</given-names></name> <name><surname>Corazzini</surname> <given-names>L.</given-names></name> <name><surname>Paccagnella</surname> <given-names>O.</given-names></name></person-group> (<year>2014</year>). <article-title>Do Danes and Italians rate life satisfaction in the same way? Using Vignettes to correct for individual-specific scale biases</article-title>. <source>Oxf. Bull. Econ. Stat.</source> <volume>76</volume>, <fpage>643</fpage>&#x02013;<lpage>666</lpage>. <pub-id pub-id-type="doi">10.1111/obes.12039</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benet-Mart&#x000ED;nez</surname> <given-names>V.</given-names></name> <name><surname>John</surname> <given-names>O. P.</given-names></name></person-group> (<year>1998</year>). <article-title>Los Cinco Grandes: across cultures and ethnic groups: multitrait-multimethod analyses of the Big Five in Spanish and English</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>75</volume>, <fpage>729</fpage>&#x02013;<lpage>750</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.75.3.729</pub-id><pub-id pub-id-type="pmid">9781409</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bennett</surname> <given-names>R. J.</given-names></name> <name><surname>Robinson</surname> <given-names>S. L.</given-names></name></person-group> (<year>2000</year>). <article-title>Development of a measure of workplace deviance</article-title>. <source>J. Appl. Psychol.</source> <volume>85</volume>, <fpage>349</fpage>&#x02013;<lpage>360</lpage>. <pub-id pub-id-type="doi">10.1037/0021-9010.85.3.349</pub-id><pub-id pub-id-type="pmid">10900810</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bentler</surname> <given-names>P. M.</given-names></name></person-group> (<year>1990</year>). <article-title>Comparative fit indexes in structural models</article-title>. <source>Psychol. Bull.</source> <volume>107</volume>, <fpage>238</fpage>&#x02013;<lpage>246</lpage>. <pub-id pub-id-type="doi">10.1037/0033-2909.107.2.238</pub-id><pub-id pub-id-type="pmid">2320703</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bijnen</surname> <given-names>E. J.</given-names></name> <name><surname>Van Der Net</surname> <given-names>T. Z.</given-names></name> <name><surname>Poortinga</surname> <given-names>Y. H.</given-names></name></person-group> (<year>1986</year>). <article-title>On cross-cultural comparative studies with the Eysenck Personality Questionnaire</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>17</volume>, <fpage>3</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1177/0022002186017001001</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borkenau</surname> <given-names>P.</given-names></name> <name><surname>Ostendorf</surname> <given-names>F.</given-names></name></person-group> (<year>1990</year>). <article-title>Comparing exploratory and confirmatory factor analysis: a study on the 5-factor model of personality</article-title>. <source>Pers. Individ. Dif.</source> <volume>11</volume>, <fpage>515</fpage>&#x02013;<lpage>524</lpage>. <pub-id pub-id-type="doi">10.1016/0191-8869(90)90065-y</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bratko</surname> <given-names>D.</given-names></name> <name><surname>Maru&#x00161;i&#x00107;</surname> <given-names>I.</given-names></name></person-group> (<year>1997</year>). <article-title>Family study of the big five personality dimensions</article-title>. <source>Pers. Individ. Dif.</source> <volume>23</volume>, <fpage>365</fpage>&#x02013;<lpage>369</lpage>. <pub-id pub-id-type="doi">10.1016/s0191-8869(97)00081-0</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Buckley</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <source>Survey Context Effects in Anchoring Vignettes</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.456.9907&#x00026;rep=rep1&#x00026;type=pdf">http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.456.9907&#x00026;rep=rep1&#x00026;type=pdf</ext-link></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Byrne</surname> <given-names>B. M.</given-names></name> <name><surname>Campbell</surname> <given-names>T. L.</given-names></name></person-group> (<year>1999</year>). <article-title>Cross-cultural comparisons and the presumption of equivalent measurement and theoretical structure: a look beneath the surface</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>30</volume>, <fpage>555</fpage>&#x02013;<lpage>574</lpage>. <pub-id pub-id-type="doi">10.1177/0022022199030005001</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Camilli</surname> <given-names>G.</given-names></name> <name><surname>Shepard</surname> <given-names>L. A.</given-names></name></person-group> (<year>1994</year>). <source>Methods for Identifying Biased Test Items</source>. <publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chalmers</surname> <given-names>R. P.</given-names></name></person-group> (<year>2012</year>). <article-title>Mirt: a multidimensional item response theory package for the R environment</article-title>. <source>J. Stat. Softw.</source> <volume>48</volume>, <fpage>1</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v048.i06</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>D. K. C.</given-names></name> <name><surname>Ivarsson</surname> <given-names>A.</given-names></name> <name><surname>Stenling</surname> <given-names>A.</given-names></name> <name><surname>Yang</surname> <given-names>X. S.</given-names></name> <name><surname>Chatzisarantis</surname> <given-names>N. L.</given-names></name> <name><surname>Hagger</surname> <given-names>M. S.</given-names></name></person-group> (<year>2015</year>). <article-title>Response-order effects in survey methods: a randomized controlled crossover study in the context of sport injury prevention</article-title>. <source>J. Sport Exerc. Psychol.</source> <volume>37</volume>, <fpage>666</fpage>&#x02013;<lpage>673</lpage>. <pub-id pub-id-type="doi">10.1123/jsep.2015-0045</pub-id><pub-id pub-id-type="pmid">26866774</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chevalier</surname> <given-names>A.</given-names></name> <name><surname>Fielding</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>An introduction to anchoring vignettes</article-title>. <source>J. R. Stat. Soc. Ser. A</source> <volume>174</volume>, <fpage>569</fpage>&#x02013;<lpage>574</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-985x.2011.00703.x</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crane</surname> <given-names>M.</given-names></name> <name><surname>Rissel</surname> <given-names>C.</given-names></name> <name><surname>Greaves</surname> <given-names>S.</given-names></name> <name><surname>Gebel</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Correcting bias in self-rated quality of life: an application of anchoring vignettes and ordinal regression models to better understand QoL differences across commuting modes</article-title>. <source>Qual. Life Res.</source> <volume>25</volume>, <fpage>257</fpage>&#x02013;<lpage>266</lpage>. <pub-id pub-id-type="doi">10.1007/s11136-015-1090-8</pub-id><pub-id pub-id-type="pmid">26254800</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crane</surname> <given-names>P. K.</given-names></name> <name><surname>Gibbons</surname> <given-names>L. E.</given-names></name> <name><surname>Jolley</surname> <given-names>L.</given-names></name> <name><surname>van Belle</surname> <given-names>G.</given-names></name></person-group> (<year>2006</year>). <article-title>Differential item functioning analysis with ordinal logistic regression techniques: DIFdetect and difwithpar</article-title>. <source>Med. Care</source> <volume>44</volume>, <fpage>S115</fpage>&#x02013;<lpage>S123</lpage>. <pub-id pub-id-type="doi">10.1097/01.mlr.0000245183.28384.ed</pub-id><pub-id pub-id-type="pmid">17060818</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diener</surname> <given-names>E. D.</given-names></name> <name><surname>Emmons</surname> <given-names>R. A.</given-names></name> <name><surname>Larsen</surname> <given-names>R. J.</given-names></name> <name><surname>Griffin</surname> <given-names>S.</given-names></name></person-group> (<year>1985</year>). <article-title>The satisfaction with life scale</article-title>. <source>J. Pers. Assess.</source> <volume>49</volume>, <fpage>71</fpage>&#x02013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1207/s15327752jpa4901_13</pub-id><pub-id pub-id-type="pmid">16367493</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Digman</surname> <given-names>J. M.</given-names></name></person-group> (<year>1997</year>). <article-title>Higher-order factors of the Big Five</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>73</volume>, <fpage>1246</fpage>&#x02013;<lpage>1256</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514-73.6.1246</pub-id><pub-id pub-id-type="pmid">9418278</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ehrler</surname> <given-names>D. J.</given-names></name> <name><surname>Evans</surname> <given-names>J. G.</given-names></name> <name><surname>McGhee</surname> <given-names>R. L.</given-names></name></person-group> (<year>1999</year>). <article-title>Extending Big-Five theory into childhood: a preliminary investigation into the relationship between Big-Five personality traits and behavior problems in children</article-title>. <source>Psychol. Sch.</source> <volume>36</volume>, <fpage>451</fpage>&#x02013;<lpage>458</lpage>. <pub-id pub-id-type="doi">10.1002/(sici)1520-6807(199911)</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ellis</surname> <given-names>B. B.</given-names></name> <name><surname>Becker</surname> <given-names>P.</given-names></name> <name><surname>Kimmel</surname> <given-names>H. D.</given-names></name></person-group> (<year>1993</year>). <article-title>An item response theory evaluation of an English version of the Trier Personality Inventory (TPI)</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>24</volume>, <fpage>133</fpage>&#x02013;<lpage>148</lpage>. <pub-id pub-id-type="doi">10.1177/0022022193242001</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eysenck</surname> <given-names>S. B.</given-names></name> <name><surname>Barrett</surname> <given-names>P. T.</given-names></name> <name><surname>Barnes</surname> <given-names>G. E.</given-names></name></person-group> (<year>1993</year>). <article-title>A cross-cultural study of personality: Canada and England</article-title>. <source>Pers. Individ. Dif.</source> <volume>14</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/0191-8869(93)90168-3</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fischer</surname> <given-names>R.</given-names></name></person-group> (<year>2004</year>). <article-title>Standardization to account for cross-cultural response bias a classification of Score Adjustment Procedures and Review of Research in JCCP</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>35</volume>, <fpage>263</fpage>&#x02013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1177/0022022104264122</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fischer</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>What values can (and cannot) tell us about individuals, society and culture</article-title>, in <source>Advances in Culture and Psychology</source>, eds <person-group person-group-type="editor"><name><surname>Gelfand</surname> <given-names>M. J.</given-names></name> <name><surname>Chiu</surname> <given-names>C.</given-names></name> <name><surname>Hong</surname> <given-names>Y.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>), <fpage>218</fpage>&#x02013;<lpage>272</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grimm</surname> <given-names>S. D.</given-names></name> <name><surname>Church</surname> <given-names>A. T.</given-names></name></person-group> (<year>1999</year>). <article-title>A cross-cultural study of response biases in personality measures</article-title>. <source>J. Res. Pers.</source> <volume>33</volume>, <fpage>415</fpage>&#x02013;<lpage>441</lpage>. <pub-id pub-id-type="doi">10.1006/jrpe.1999.2256</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grol-Prokopczyk</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>Age and sex effects in anchoring vignette studies: methodological and empirical contributions</article-title>. <source>Surv. Res. Methods</source> <volume>8</volume>, <fpage>1</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="pmid">25621079</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grol-Prokopczyk</surname> <given-names>H.</given-names></name> <name><surname>Verdes-Tennant</surname> <given-names>E.</given-names></name> <name><surname>McEniry</surname> <given-names>M.</given-names></name> <name><surname>Isp&#x000E1;ny</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Promises and pitfalls of anchoring vignettes in health survey research</article-title>. <source>Demography</source> <volume>52</volume>, <fpage>1703</fpage>&#x02013;<lpage>1728</lpage>. <pub-id pub-id-type="doi">10.1007/s13524-015-0422-1</pub-id><pub-id pub-id-type="pmid">26335547</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gupta</surname> <given-names>N. D.</given-names></name> <name><surname>Kristensen</surname> <given-names>N.</given-names></name> <name><surname>Pozzoli</surname> <given-names>D.</given-names></name></person-group> (<year>2010</year>). <article-title>External validation of the use of vignettes in cross-country health studies</article-title>. <source>Econ. Model.</source> <volume>27</volume>, <fpage>854</fpage>&#x02013;<lpage>865</lpage>. <pub-id pub-id-type="doi">10.1016/j.econmod.2009.11.007</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ham</surname> <given-names>E. H.</given-names></name> <name><surname>Roberts</surname> <given-names>R. D.</given-names></name></person-group> (<year>2015</year>). <article-title>An application of anchoring vignettes for improving interpersonal comparability of student self-reported teamwork scores</article-title>. <volume>28</volume>, <fpage>1107</fpage>&#x02013;<lpage>1128</lpage>. <pub-id pub-id-type="doi">10.1037/e552422014-001</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>Van de Vijver</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>Correcting for scale usage differences among Latin American Countries, Portugal, and Spain in PISA</article-title>. <source>Electron. J. Educ. Res. Assess. Eval.</source> <volume>22</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.7203/relieve.22.1.8282</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>Van de Vijver</surname> <given-names>F. J.</given-names></name> <name><surname>Fetvadjiev</surname> <given-names>V. H.</given-names></name> <name><surname>Carmen Dominguez Espinosa</surname> <given-names>A.</given-names></name> <name><surname>Adams</surname> <given-names>B.</given-names></name> <name><surname>Alonso-Arbiol</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>On enhancing the cross-cultural comparability of likert-scale personality and value measures: a comparison of common procedures</article-title>. <source>Eur. J. Pers.</source> <volume>31</volume>, <fpage>642</fpage>&#x02013;<lpage>657</lpage>. <pub-id pub-id-type="doi">10.1002/per.2132</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heine</surname> <given-names>S. J.</given-names></name> <name><surname>Lehman</surname> <given-names>D. R.</given-names></name> <name><surname>Peng</surname> <given-names>K.</given-names></name> <name><surname>Greenholtz</surname> <given-names>J.</given-names></name></person-group> (<year>2002</year>). <article-title>What&#x00027;s wrong with cross-cultural comparisons of subjective Likert scales?: The reference-group effect</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>82</volume>, <fpage>903</fpage>&#x02013;<lpage>918</lpage>. <pub-id pub-id-type="doi">10.1037//0022-3514.82.6.903</pub-id><pub-id pub-id-type="pmid">12051579</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hirve</surname> <given-names>S.</given-names></name> <name><surname>G&#x000F3;mez-Oliv&#x000E9;</surname> <given-names>X.</given-names></name> <name><surname>Oti</surname> <given-names>S.</given-names></name> <name><surname>Debpuur</surname> <given-names>C.</given-names></name> <name><surname>Juvekar</surname> <given-names>S.</given-names></name> <name><surname>Tollman</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Use of anchoring vignettes to evaluate health reporting behavior amongst adults aged 50 years and above in Africa and Asia-testing assumptions</article-title>. <source>Glob. Health Action</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.3402/gha.v6i0.21064</pub-id><pub-id pub-id-type="pmid">24011254</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Holland</surname> <given-names>P. W.</given-names></name> <name><surname>Thayer</surname> <given-names>D.</given-names></name></person-group> (<year>1988</year>). <article-title>Differential item performance and the Mantel-Haenszel procedure</article-title>, in <source>Test Vailidity</source>, eds <person-group person-group-type="editor"><name><surname>Wainer</surname> <given-names>H.</given-names></name> <name><surname>Braun</surname> <given-names>H. I.</given-names></name></person-group> (<publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates Publishers</publisher-name>), <fpage>129</fpage>&#x02013;<lpage>145</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Holland</surname> <given-names>P. W.</given-names></name> <name><surname>Wainer</surname> <given-names>H.</given-names></name></person-group> (eds.). (<year>1993</year>). <source>Differential Item Functioning.</source> <publisher-loc>New York, NY; London</publisher-loc>: <publisher-name>Routledge</publisher-name>.</citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopkins</surname> <given-names>D. J.</given-names></name> <name><surname>King</surname> <given-names>G.</given-names></name></person-group> (<year>2010</year>). <article-title>Improving anchoring vignettes designing surveys to correct interpersonal incomparability</article-title>. <source>Public Opin. Q.</source> <volume>74</volume>, <fpage>201</fpage>&#x02013;<lpage>222</lpage>. <pub-id pub-id-type="doi">10.1093/poq/nfq011</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hopwood</surname> <given-names>C. J.</given-names></name> <name><surname>Donnellan</surname> <given-names>M. B.</given-names></name></person-group> (<year>2010</year>). <article-title>How should the internal structure of personality inventories be evaluated?</article-title> <source>Pers. Soc. Psychol. Rev.</source> <volume>14</volume>, <fpage>332</fpage>&#x02013;<lpage>346</lpage> <pub-id pub-id-type="doi">10.1177/1088868310361240</pub-id><pub-id pub-id-type="pmid">20435808</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hosseinkhanzadeh</surname> <given-names>A. A.</given-names></name> <name><surname>Taher</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>The relationship between personality traits with life satisfaction</article-title>. <source>Sociol. Mind</source> <volume>3</volume>, <fpage>99</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.4236/sm.2013.31015</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>C. D.</given-names></name> <name><surname>Church</surname> <given-names>A. T.</given-names></name> <name><surname>Katigbak</surname> <given-names>M. S.</given-names></name></person-group> (<year>1997</year>). <article-title>Identifying cultural differences in items and traits differential item functioning in the NEO personality inventory</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>28</volume>, <fpage>192</fpage>&#x02013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1177/0022022197282004</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>John</surname> <given-names>O. P.</given-names></name> <name><surname>Srivastava</surname> <given-names>S.</given-names></name></person-group> (<year>1999</year>). <article-title>The Big Five trait taxonomy: History, measurement, and theoretical perspectives</article-title>, in <source>Handbook of Personality: Theory and Research, 2nd Edn</source>, eds <person-group person-group-type="editor"><name><surname>Pervin</surname> <given-names>L. A.</given-names></name> <name><surname>John</surname> <given-names>O. P.</given-names></name></person-group> (<publisher-loc>New York, NY; London</publisher-loc>: <publisher-name>The Guilford Press</publisher-name>), <fpage>102</fpage>&#x02013;<lpage>138</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>John</surname> <given-names>O. P.</given-names></name> <name><surname>Donahue</surname> <given-names>E. M.</given-names></name> <name><surname>Kentle</surname> <given-names>R. L.</given-names></name></person-group> (<year>1991</year>). <source>The Big Five Inventory: Versions 4a and 54</source>. <publisher-loc>Berkeley, CA</publisher-loc>: <publisher-name>University of California, Berkeley, Institute of Personality and Social Research</publisher-name>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>W.</given-names></name> <name><surname>Spinath</surname> <given-names>F.</given-names></name> <name><surname>Krueger</surname> <given-names>R. F.</given-names></name> <name><surname>Angleitner</surname> <given-names>A.</given-names></name> <name><surname>Riemann</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Personality in Germany and Minnesota: an IRT-Based Comparison of MPQ Self-Reports</article-title>. <source>J. Pers.</source> <volume>76</volume>, <fpage>665</fpage>&#x02013;<lpage>706</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-6494.2008.00500.x</pub-id><pub-id pub-id-type="pmid">18399949</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>J&#x000FC;rges</surname> <given-names>H.</given-names></name> <name><surname>Winter</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Are anchoring vignettes ratings sensitive to vignette age and sex?</article-title>. <source>Health Econ.</source> <volume>22</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1002/hec.1806</pub-id><pub-id pub-id-type="pmid">22083845</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kapteyn</surname> <given-names>A.</given-names></name> <name><surname>Smith</surname> <given-names>J. P.</given-names></name> <name><surname>Soest</surname> <given-names>A. V.</given-names></name> <name><surname>Vonkova</surname> <given-names>H.</given-names></name></person-group> (<year>2011</year>). <article-title>Anchoring vignettes and response consistency</article-title>. <source>Labor Popul.</source> WR-<volume>840</volume>, <fpage>1</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.2139/ssrn.1799563</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kapteyn</surname> <given-names>A.</given-names></name> <name><surname>Smith</surname> <given-names>J.</given-names></name> <name><surname>Soest</surname> <given-names>A. V.</given-names></name></person-group> (<year>2007</year>). <article-title>Vignettes and self-reported work disability in the US and the Netherlands</article-title>. <source>Am. Econ. Rev.</source> <volume>97</volume>, <fpage>461</fpage>&#x02013;<lpage>473</lpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kapteyn</surname> <given-names>A.</given-names></name> <name><surname>Smith</surname> <given-names>J.</given-names></name> <name><surname>Soest</surname> <given-names>A. V.</given-names></name></person-group> (<year>2010</year>). <article-title>Life satisfaction</article-title>, in <source>International Differences in Subjective Well-Being</source>, eds <person-group person-group-type="editor"><name><surname>Diener</surname> <given-names>E.</given-names></name> <name><surname>Helliwell</surname> <given-names>J.</given-names></name> <name><surname>Kahneman</surname> <given-names>D.</given-names></name></person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>), <fpage>70</fpage>&#x02013;<lpage>104</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>King</surname> <given-names>G.</given-names></name> <name><surname>Wand</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Comparing incomparable survey responses: evaluating and selecting anchoring vignettes</article-title>. <source>Polit. Anal.</source> <volume>15</volume>, <fpage>46</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1093/pan/mpl011</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>King</surname> <given-names>G.</given-names></name> <name><surname>Murray</surname> <given-names>C. J.</given-names></name> <name><surname>Salomon</surname> <given-names>J. A.</given-names></name> <name><surname>Tandon</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Enhancing the validity and cross-cultural comparability of measurement in survey research</article-title>. <source>Am. Polit. Sci. Rev.</source> <volume>98</volume>, <fpage>191</fpage>&#x02013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-531-91826-6_16</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Knott</surname> <given-names>R.</given-names></name> <name><surname>Lorgelly</surname> <given-names>P.</given-names></name> <name><surname>Black</surname> <given-names>N.</given-names></name> <name><surname>Hollingsworth</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <source>Differential Item Functioning in the EQ-5D: An Exploratory Analysis Using Anchoring Vignettes (No. 16/14).</source> HEDG, c/o Department of Economics, <publisher-name>University of York</publisher-name>.</citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kristensen</surname> <given-names>N.</given-names></name> <name><surname>Johansson</surname> <given-names>E.</given-names></name></person-group> (<year>2008</year>). <article-title>New evidence on cross-country differences in job satisfaction using anchoring vignettes</article-title>. <source>Labour Econ.</source> <volume>15</volume>, <fpage>96</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.labeco.2006.11.001</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kyllonen</surname> <given-names>P. C.</given-names></name> <name><surname>Bertling</surname> <given-names>J. P.</given-names></name></person-group> (<year>2014</year>). <source>Anchoring Vignettes reduce Bias in Noncognitive Rating Scale Responses.</source> <publisher-loc>Princeton, NJ</publisher-loc>: <publisher-name>ETS/OECD</publisher-name>.</citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maniaci</surname> <given-names>M. R.</given-names></name> <name><surname>Rogge</surname> <given-names>R. D.</given-names></name></person-group> (<year>2014</year>). <article-title>Caring about carelessness: participant inattention and its effects on research</article-title>. <source>J. Res. Pers.</source> <volume>48</volume>, <fpage>61</fpage>&#x02013;<lpage>83</lpage>. <pub-id pub-id-type="doi">10.1016/j.jrp.2013.09.008</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>MacCann</surname> <given-names>C.</given-names></name> <name><surname>Duckworth</surname> <given-names>A. L.</given-names></name> <name><surname>Roberts</surname> <given-names>R. D.</given-names></name></person-group> (<year>2009</year>). <article-title>Empirical identification of the major facets of conscientiousness</article-title>. <source>Learn. Individ. Differ.</source> <volume>19</volume>, <fpage>451</fpage>&#x02013;<lpage>458</lpage>. <pub-id pub-id-type="doi">10.1016/j.lindif.2009.03.007</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>H. W.</given-names></name> <name><surname>Trautwein</surname> <given-names>U.</given-names></name> <name><surname>L&#x000FC;dtke</surname> <given-names>O.</given-names></name> <name><surname>K&#x000F6;ller</surname> <given-names>O.</given-names></name> <name><surname>Baumert</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Integration of multidimensional self-concept and core personality constructs: construct validation and relations to well-being and achievement</article-title>. <source>J. Pers.</source> <volume>74</volume>, <fpage>403</fpage>&#x02013;<lpage>456</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-6494.2005.00380</pub-id><pub-id pub-id-type="pmid">16529582</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCrae</surname> <given-names>R. R.</given-names></name> <name><surname>Terracciano</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>Universal features of personality traits from the observer&#x00027;s perspective: data from 50 cultures</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>88</volume>, <fpage>547</fpage>&#x02013;<lpage>582</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.88.3.547</pub-id><pub-id pub-id-type="pmid">15740445</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCrae</surname> <given-names>R. R.</given-names></name> <name><surname>Zonderman</surname> <given-names>A. B.</given-names></name> <name><surname>Costa</surname> <given-names>P. T.</given-names> <suffix>Jr.</suffix></name> <name><surname>Bond</surname> <given-names>M. H.</given-names></name> <name><surname>Paunonen</surname> <given-names>S. V.</given-names></name></person-group> (<year>1996</year>). <article-title>Evaluating replicability of factors in the Revised NEO Personality Inventory: confirmatory factor analysis versus Procrustes rotation</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>70</volume>, <fpage>552</fpage>&#x02013;<lpage>566</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.70.3.552</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>McDonald</surname> <given-names>R. P.</given-names></name></person-group> (<year>1999</year>). <source>Test Theory: A Unified Treatment.</source> <publisher-loc>Mahwah, NJ</publisher-loc>: <publisher-name>Erlbaum</publisher-name>.</citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>M&#x000F5;ttus</surname> <given-names>R.</given-names></name> <name><surname>Allik</surname> <given-names>J.</given-names></name> <name><surname>Realo</surname> <given-names>A.</given-names></name> <name><surname>Rossier</surname> <given-names>J.</given-names></name> <name><surname>Zecca</surname> <given-names>G.</given-names></name> <name><surname>Ah-Kion</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>The effect of response style on self-reported conscientiousness across 20 countries</article-title>. <source>Pers. Soc. Psychol. Bull.</source> <volume>38</volume>, <fpage>1423</fpage>&#x02013;<lpage>1436</lpage>. <pub-id pub-id-type="doi">10.1177/0146167212451275</pub-id><pub-id pub-id-type="pmid">22745332</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mount</surname> <given-names>M.</given-names></name> <name><surname>Ilies</surname> <given-names>R.</given-names></name> <name><surname>Johnson</surname> <given-names>E.</given-names></name></person-group> (<year>2006</year>). <article-title>Relationship of personality traits and counterproductive work behaviors: the mediating effects of job satisfaction</article-title>. <source>Pers. Psychol.</source> <volume>59</volume>, <fpage>591</fpage>&#x02013;<lpage>622</lpage>. <pub-id pub-id-type="doi">10.1111/j.1744-6570.2006.00048.x</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mpofu</surname> <given-names>E.</given-names></name> <name><surname>Nyanungo</surname> <given-names>K. R.</given-names></name></person-group> (<year>1998</year>). <article-title>Educational and psychological testing in Zimbabwean schools: past, present and future</article-title>. <source>Eur. J. Psychol. Assess.</source> <volume>14</volume>, <fpage>71</fpage>&#x02013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1027/1015-5759.14.1.71</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="thesis"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>M. L.</given-names></name></person-group> (<year>2014</year>). <source>The Development of Life Satisfaction: Does Personality Matter? A Five Year Longitudinal Study</source>. Master Thesis, <publisher-name>University of Twente</publisher-name>, <publisher-loc>Enschede</publisher-loc>.</citation>
</ref>
<ref id="B63">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Muth&#x000E9;n</surname> <given-names>L. K.</given-names></name> <name><surname>Muth&#x000E9;n</surname> <given-names>B. O.</given-names></name></person-group> (<year>1998-2015</year>). <source>Mplus Version 7 User&#x00027;s Guide. Statistical Analysis with Latent Variables</source>. <publisher-loc>Los Angeles, CA</publisher-loc>: <publisher-name>Muth&#x000E9;n &#x00026; Muth&#x000E9;n</publisher-name>.</citation>
</ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><collab>OECD</collab></person-group> (<year>2012</year>). <source>PISA 2012 - Technical Report</source>. <publisher-loc>Paris</publisher-loc>: <publisher-name>OECD</publisher-name>.</citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olaru</surname> <given-names>G.</given-names></name> <name><surname>Witth&#x000F6;ft</surname> <given-names>M.</given-names></name> <name><surname>Wilhelm</surname> <given-names>O.</given-names></name></person-group> (<year>2015</year>). <article-title>Methods matter: testing competing models for designing short-scale big-five assessments</article-title>. <source>J. Res. Pers.</source> <volume>59</volume>, <fpage>56</fpage>&#x02013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1016/j.jrp.2015.09.001</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Onyishi</surname> <given-names>I. E.</given-names></name> <name><surname>Okongwu</surname> <given-names>O. E.</given-names></name> <name><surname>Ugwu</surname> <given-names>F. O.</given-names></name></person-group> (<year>2012</year>). <article-title>Personality and social support as predictors of life satisfaction of Nigerian prisons officers</article-title>. <source>Eur. Sci. J.</source> <volume>8</volume>, <fpage>110</fpage>&#x02013;<lpage>125</lpage>.</citation>
</ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oppenheimer</surname> <given-names>D. M.</given-names></name> <name><surname>Meyvis</surname> <given-names>T.</given-names></name> <name><surname>Davidenko</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <article-title>Instructional manipulation checks: Detecting satisficing to increase statistical power</article-title>. <source>J. Exp. Soc. Psychol.</source> <volume>45</volume>, <fpage>867</fpage>&#x02013;<lpage>872</lpage>. <pub-id pub-id-type="doi">10.1016/j.jesp.2009.03.009</pub-id></citation>
</ref>
<ref id="B68">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Osterlind</surname> <given-names>S. J.</given-names></name> <name><surname>Everson</surname> <given-names>H. T.</given-names></name></person-group> (<year>2009</year>). <source>Differential Item Functioning, Vol. 161</source>. <publisher-loc>Thousand Oaks, CA</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paccagnella</surname> <given-names>O.</given-names></name></person-group> (<year>2013</year>). <article-title>Modelling individual heterogeneity in ordered choice models: anchoring vignettes and the Chopit Model</article-title>. <source>J. Methodol. Appl. Stat.</source> <volume>15</volume>, <fpage>69</fpage>&#x02013;<lpage>94</lpage>.</citation>
</ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>K.</given-names></name> <name><surname>Nisbett</surname> <given-names>R. E.</given-names></name> <name><surname>Wong</surname> <given-names>N. Y.</given-names></name></person-group> (<year>1997</year>). <article-title>Validity problems comparing values across cultures and possible solutions</article-title>. <source>Psychol. Methods</source> <volume>2</volume>, <fpage>329</fpage>&#x02013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1037/1082-989x.2.4.329</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Primi</surname> <given-names>R.</given-names></name> <name><surname>Zanon</surname> <given-names>C.</given-names></name> <name><surname>Santos</surname> <given-names>D.</given-names></name> <name><surname>De Fruyt</surname> <given-names>F.</given-names></name> <name><surname>John</surname> <given-names>O. P.</given-names></name></person-group> (<year>2016</year>). <article-title>Anchoring Vignettes: can they make adolescent self-reports of social-emotional skills more reliable, discriminant, and criterion-valid?</article-title>. <source>Eur. J. Psychol. Assess.</source> <volume>32</volume>, <fpage>39</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1027/1015-5759/a000336</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reise</surname> <given-names>S. P.</given-names></name> <name><surname>Waller</surname> <given-names>N. G.</given-names></name></person-group> (<year>1990</year>). <article-title>Fitting the two-parameter model to personality data</article-title>. <source>Appl. Psychol. Meas.</source> <volume>14</volume>, <fpage>45</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1177/014662169001400105</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Revelle</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <source>psych: Procedures for Personality and Psychological Research.</source> <publisher-loc>R package version, 1(1), Evanston, IL</publisher-loc>: <publisher-name>Northwestern University</publisher-name>.</citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rice</surname> <given-names>N.</given-names></name> <name><surname>Robone</surname> <given-names>S.</given-names></name> <name><surname>Smith</surname> <given-names>P.</given-names></name></person-group> (<year>2011</year>). <article-title>Analysis of the validity of the vignette approach to correct for heterogeneity in reporting health system responsiveness</article-title>. <source>Eur. J. Health Econ.</source> <volume>12</volume>, <fpage>141</fpage>&#x02013;<lpage>162</lpage>. <pub-id pub-id-type="doi">10.1007/s10198-010-0235-5</pub-id><pub-id pub-id-type="pmid">20349262</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Roberts</surname> <given-names>R. D.</given-names></name> <name><surname>Martin</surname> <given-names>J.</given-names></name> <name><surname>Olaru</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <source>A Rosetta Stone for Noncognitive Skills: Understanding, Assessing, and Enhancing Noncognitive Skills in Primary and Secondary Education</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Asia Society and ProExam</publisher-name>.</citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogers</surname> <given-names>H. J.</given-names></name> <name><surname>Swaminathan</surname> <given-names>H.</given-names></name></person-group> (<year>1993</year>). <article-title>A comparison of logistic regression and Mantel-Haenszel procedures for detecting differential item functioning</article-title>. <source>Appl. Psychol. Meas.</source> <volume>17</volume>, <fpage>105</fpage>&#x02013;<lpage>116</lpage>. <pub-id pub-id-type="doi">10.1177/014662169301700201</pub-id></citation>
</ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosseel</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>Lavaan: an R package for structural equation modeling</article-title>. <source>J. Stat. Softw.</source> <volume>48</volume>, <fpage>1</fpage>&#x02013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v048.i02</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rothbart</surname> <given-names>M. K.</given-names></name> <name><surname>Ahadi</surname> <given-names>S. A.</given-names></name> <name><surname>Evans</surname> <given-names>D. E.</given-names></name></person-group> (<year>2000</year>). <article-title>Temperament and personality: origins and outcomes</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>78</volume>, <fpage>122</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.78.1.122</pub-id><pub-id pub-id-type="pmid">10653510</pub-id></citation>
</ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rupp</surname> <given-names>A. A.</given-names></name> <name><surname>Zumbo</surname> <given-names>B. D.</given-names></name></person-group> (<year>2006</year>). <article-title>Understanding parameter invariance in unidimensional IRT models</article-title>. <source>Educ. Psychol. Meas.</source> <volume>66</volume>, <fpage>63</fpage>&#x02013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1177/0013164404273942</pub-id></citation>
</ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmitt</surname> <given-names>D.</given-names></name> <name><surname>Allik</surname> <given-names>J.</given-names></name> <name><surname>McCrae</surname> <given-names>R. R.</given-names></name> <name><surname>Benet-Martinez</surname> <given-names>V.</given-names></name></person-group> (<year>2007</year>). <article-title>The geographic distribution of Big Five personality traits: patterns and profiles of human self&#x02013;description across 56 nations</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>38</volume>, <fpage>173</fpage>&#x02013;<lpage>212</lpage>. <pub-id pub-id-type="doi">10.1177/0022022106297299</pub-id></citation>
</ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soest</surname> <given-names>A. V.</given-names></name> <name><surname>Delaney</surname> <given-names>L.</given-names></name> <name><surname>Harmon</surname> <given-names>C.</given-names></name> <name><surname>Kapteyn</surname> <given-names>A.</given-names></name> <name><surname>Smith</surname> <given-names>J. P.</given-names></name></person-group> (<year>2011</year>). <article-title>Validating the use of anchoring vignettes for the correction of response scale differences in subjective questions</article-title>. <source>J. R. Stat. Soc. Ser. A</source> <volume>174</volume>, <fpage>575</fpage>&#x02013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-985x.2011.00694.x</pub-id></citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stark</surname> <given-names>S.</given-names></name> <name><surname>Chernyshenko</surname> <given-names>O. S.</given-names></name> <name><surname>Drasgow</surname> <given-names>F.</given-names></name></person-group> (<year>2006</year>). <article-title>Detecting differential item functioning with confirmatory factor analysis and item response theory: toward a unified strategy</article-title>. <source>J. Appl. Psychol.</source> <volume>91</volume>, <fpage>1292</fpage>&#x02013;<lpage>1306</lpage>. <pub-id pub-id-type="doi">10.1037/0021-9010.91.6.1292</pub-id><pub-id pub-id-type="pmid">17100485</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Staub</surname> <given-names>E.</given-names></name> <name><surname>Pearlman</surname> <given-names>L. A.</given-names></name> <name><surname>Gubin</surname> <given-names>A.</given-names></name> <name><surname>Hagengimana</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>Healing, reconciliation, forgiving and the prevention of violence after genocide or mass killing: an intervention and its experimental evaluation in Rwanda</article-title>. <source>J. Soc. Clin. Psychol.</source> <volume>24</volume>, <fpage>297</fpage>&#x02013;<lpage>334</lpage>. <pub-id pub-id-type="doi">10.1521/jscp.24.3.297.65617</pub-id></citation>
</ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steiger</surname> <given-names>J. H.</given-names></name></person-group> (<year>1990</year>). <article-title>Structural model evaluation and modification: an interval estimation approach</article-title>. <source>Multivariate Behav. Res.</source> <volume>25</volume>, <fpage>173</fpage>&#x02013;<lpage>180</lpage>. <pub-id pub-id-type="doi">10.1207/s15327906mbr2502_4</pub-id><pub-id pub-id-type="pmid">26794479</pub-id></citation>
</ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teresi</surname> <given-names>J. A.</given-names></name></person-group> (<year>2006</year>). <article-title>Overview of quantitative measurement methods: equivalence, invariance, and differential item functioning in health applications</article-title>. <source>Med. Care</source> <volume>44</volume>, <fpage>39</fpage>&#x02013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1097/01.mlr.0000245452.48613.45</pub-id><pub-id pub-id-type="pmid">17060834</pub-id></citation>
</ref>
<ref id="B86">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thissen</surname> <given-names>D.</given-names></name> <name><surname>Steinberg</surname> <given-names>L.</given-names></name> <name><surname>Wainer</surname> <given-names>H.</given-names></name></person-group> (<year>1988</year>). <article-title>Use of item response theory in the study of group differences in trace lines</article-title>, in <source>Test Validity</source>, eds <person-group person-group-type="editor"><name><surname>Wainer</surname> <given-names>H.</given-names></name> <name><surname>Braun</surname> <given-names>H. I.</given-names></name></person-group> (<publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum</publisher-name>), <fpage>147</fpage>&#x02013;<lpage>169</lpage>.</citation>
</ref>
<ref id="B87">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thissen</surname> <given-names>D.</given-names></name> <name><surname>Steinberg</surname> <given-names>L.</given-names></name> <name><surname>Wainer</surname> <given-names>H.</given-names></name></person-group> (<year>1993</year>). <article-title>Detection of differential item functioning using the parameters of item response models</article-title>, in <source>Differential Item Functioning</source>, eds <person-group person-group-type="editor"><name><surname>Holland</surname> <given-names>P. W.</given-names></name> <name><surname>Wainer</surname> <given-names>H.</given-names></name></person-group> (<publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates</publisher-name>), <fpage>67</fpage>&#x02013;<lpage>113</lpage>.</citation>
</ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tully</surname> <given-names>P. J.</given-names></name> <name><surname>Winefield</surname> <given-names>H. R.</given-names></name> <name><surname>Baker</surname> <given-names>R. A.</given-names></name> <name><surname>Turnbull</surname> <given-names>D. A.</given-names></name> <name><surname>de Jonge</surname> <given-names>P.</given-names></name></person-group> (<year>2011</year>). <article-title>Confirmatory factor analysis of the Beck Depression Inventory-II and the association with cardiac morbidity and mortality after coronary revascularization</article-title>. <source>J. Health Psychol.</source> <volume>16</volume>, <fpage>584</fpage>&#x02013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1177/1359105310383604</pub-id><pub-id pub-id-type="pmid">21346014</pub-id></citation>
</ref>
<ref id="B89">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tupes</surname> <given-names>E. C.</given-names></name> <name><surname>Christal</surname> <given-names>R. E.</given-names></name></person-group> (<year>1961</year>). <source>Recurrent personality factors based on trait ratings (No. ASD-TR-61-97).</source> <publisher-loc>Lackland, TX</publisher-loc>: <publisher-name>Personnel Research Lab</publisher-name>.</citation>
</ref>
<ref id="B90">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Van de Vijver</surname> <given-names>F. J.</given-names></name> <name><surname>Leung</surname> <given-names>K.</given-names></name></person-group> (<year>1997</year>). <source>Methods and Data Analysis for Cross-Cultural Research, Vol. 1</source>. <publisher-name>Sage</publisher-name>.</citation>
</ref>
<ref id="B91">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van de Vijver</surname> <given-names>F. J.</given-names></name> <name><surname>Leung</surname> <given-names>K.</given-names></name></person-group> (<year>2000</year>). <article-title>Methodological issues in psychological research on culture</article-title>. <source>J. Cross Cult. Psychol.</source> <volume>31</volume>, <fpage>33</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1177/0022022100031001004</pub-id></citation>
</ref>
<ref id="B92">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Vaerenbergh</surname> <given-names>Y.</given-names></name> <name><surname>Thomas</surname> <given-names>T. D.</given-names></name></person-group> (<year>2013</year>). <article-title>Response styles in survey research: a literature review of antecedents, consequences, and remedies</article-title>. <source>Int. J. Public Opin. Res.</source> <volume>25</volume>, <fpage>1</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1093/ijpor/eds021</pub-id></citation>
</ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vandenberg</surname> <given-names>R. J.</given-names></name> <name><surname>Lance</surname> <given-names>C. E.</given-names></name></person-group> (<year>2000</year>). <article-title>A review and synthesis of the measurement invariance literature: suggestions, practices, and recommendations for organizational research</article-title>. <source>Organ. Res. Methods</source> <volume>3</volume>, <fpage>4</fpage>&#x02013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1177/109442810031002</pub-id></citation>
</ref>
<ref id="B94">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Vonkova</surname> <given-names>H.</given-names></name> <name><surname>Zamarro</surname> <given-names>G.</given-names></name> <name><surname>Deberg</surname> <given-names>V.</given-names></name> <name><surname>Hitt</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <source>Comparisons of Student Perceptions of Teacher&#x00027;s Performance in the Classroom: Using Parametric Anchoring Vignette Methods for Improving Comparability</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://aefpweb.org/sites/default/files/webform/aefp40/PISApaper_AEFP.pdf">https://aefpweb.org/sites/default/files/webform/aefp40/PISApaper_AEFP.pdf</ext-link></citation>
</ref>
<ref id="B95">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wand</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Credible comparisons using interpersonally incomparable data: nonparametric scales with anchoring vignettes</article-title>. <source>Am. J. Pol. Sci.</source> <volume>57</volume>, <fpage>249</fpage>&#x02013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1111/j.1540-5907.2012.00597.x</pub-id></citation>
</ref>
<ref id="B96">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wand</surname> <given-names>J.</given-names></name> <name><surname>King</surname> <given-names>G.</given-names></name> <name><surname>Lau</surname> <given-names>O.</given-names></name></person-group> (<year>2011</year>). <article-title>Anchors: software for anchoring vignettes data</article-title>. <source>J. Stat. Softw.</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v042.i03</pub-id></citation>
</ref>
<ref id="B97">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weisberg</surname> <given-names>Y. J.</given-names></name> <name><surname>DeYoung</surname> <given-names>C. G.</given-names></name> <name><surname>Hirsh</surname> <given-names>J. B.</given-names></name></person-group> (<year>2011</year>). <article-title>Gender differences in personality across the ten aspects of the Big Five</article-title>. <source>Front. Psychol.</source> <volume>2</volume>:<fpage>178</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2011.00178</pub-id><pub-id pub-id-type="pmid">21866227</pub-id></citation>
</ref>
<ref id="B98">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Ziegler</surname> <given-names>M.</given-names></name> <name><surname>MacCann</surname> <given-names>C.</given-names></name> <name><surname>Roberts</surname> <given-names>R. D.</given-names></name></person-group> (eds.) (<year>2011</year>). <source>New Perspectives on Faking in Personality Assessment</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B99">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zinbarg</surname> <given-names>R. E.</given-names></name> <name><surname>Revelle</surname> <given-names>W.</given-names></name> <name><surname>Yovel</surname> <given-names>I.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2005</year>). <article-title>Cronbach&#x00027;s &#x003B1;, Revelle&#x00027;s &#x003B2;, and McDonald&#x00027;s &#x003C9; H: their relations with each other and two alternative conceptualizations of reliability</article-title>. <source>Psychometrika</source> <volume>70</volume>, <fpage>123</fpage>&#x02013;<lpage>133</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-003-0974-7</pub-id></citation>
</ref>
<ref id="B100">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zumbo</surname> <given-names>B. D.</given-names></name></person-group> (<year>1999</year>). <source>A Handbook on the Theory and Methods of Differential Item Functioning (DIF)</source>. <publisher-loc>Ottawa, ON</publisher-loc>: <publisher-name>National Defense Headquarters</publisher-name>.</citation>
</ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> The financial support for this study is based on the Workforce Connections award (Grant Number: AID-OAA-LA-13-00008) given to FHI360 through United States Agency for International Development (USAID). Financial support did not influence what is written in the submitted work. The Educational Development Center Washington (EDC) owns the copyrights of the final anchoring vignette questionnaire.</p>
</fn>
</fn-group>
</back>
</article>