<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2018.00097</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Establishing a Longitudinal Comparable Scale of Chinese Children&#x00027;s Cognitive Development through Calibrated Projection Linking</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Ouyang</surname> <given-names>Xiangzi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/471685/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Qiusi</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/479843/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Xin</surname> <given-names>Tao</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/497520/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Fu</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute of Developmental Psychology, Beijing Normal University</institution>, <addr-line>Beijing</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of English, Purdue University</institution>, <addr-line>West Lafayette, IN</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Collaborative Innovation Center of Assessment toward Basic Education Quality, Beijing Normal University</institution>, <addr-line>Beijing</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>Shenzhen Seaskyland Educational Evaluation Co. Ltd</institution>, <addr-line>Shenzhen</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Ioannis Tsaousis, University of Crete, Greece</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Brooke Magnus, Marquette University, United States; Mi kyoung Yim, Korea Health Personnel Licensing Examination Institute, South Korea</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Tao Xin <email>xintao&#x00040;bnu.edu.cn</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Quantitative Psychology and Measurement, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>02</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>9</volume>
<elocation-id>97</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>09</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>01</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2018 Ouyang, Zhang, Xin and Liu.</copyright-statement>
<copyright-year>2018</copyright-year>
<copyright-holder>Ouyang, Zhang, Xin and Liu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>In the past decades, the longitudinal approach has been remarkably and increasingly used in the investigations of children&#x00027;s cognitive development. Recently, many researchers have started to realize the importance and necessity of examining measurement invariance for any further longitudinal analysis. However, there are few empirical studies demonstrating how to conduct further analysis when the assumption of measurement invariance of an instrument is violated. The primary purpose of this study is to explore how a newly-developed calibrated projection method can be applied to reduce the impact of lack of parameter invariance in a longitudinal study of preschool children&#x00027;s cognitive development. The sample consisted of 882 children from China who participated in two waves of the cognitive tests when they were 4 and 5 years old. Before this study was conducted, the IRT method was used to examine the measurement invariance of the instrument. The results showed that five items presented difficulty parameter drift and three items presented discrimination/slope parameter drift. In the study, the invariant items were treated as &#x0201C;common items&#x0201D; and calibrated projection linking was used to establish a comparable scale across two time points. Then the linking method was evaluated by three properties: grade-to-grade growth, grade-to-grade variability, and the separation of distributions. The results showed that the grade-to-grade growth across two waves was larger and exhibited a larger effect size; the grade-to-grade variability showed less scale shrinkage, which indicated a smaller measurement error; the separation of distributions showed a larger growth as well.</p></abstract>
<kwd-group>
<kwd>children cognition</kwd>
<kwd>measurement invariance</kwd>
<kwd>longitudinal study</kwd>
<kwd>multidimensional IRT</kwd>
<kwd>calibrated projection</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="4"/>
<equation-count count="5"/>
<ref-count count="43"/>
<page-count count="10"/>
<word-count count="6649"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>When the trend of children&#x00027;s cognitive development is assessed, the longitudinal approach is important because it facilitates the understanding of the dynamic processes of developmental change in children&#x00027;s cognition. As opposed to describing cognitive skills at different ages (Ornstein and Haden, <xref ref-type="bibr" rid="B27">2001</xref>), longitudinal studies place an emphasis on developmental change and can elucidate developmental trajectories of skill acquisition (Grammer et al., <xref ref-type="bibr" rid="B12">2013</xref>). For example, some longitudinal studies showed a systematic transition from relatively passive to more active remembering across elementary school years (e.g., Schneider and Sodian, <xref ref-type="bibr" rid="B34">1997</xref>; Sodian and Schneider, <xref ref-type="bibr" rid="B36">1999</xref>). The age-related trends revealed a picture of gradual development throughout childhood. In addition, the longitudinal method enables an examination of the mechanisms that may underlie the developmental changes as well as the skills associated with the changes over time. For example, Grammer et al. (<xref ref-type="bibr" rid="B13">2011</xref>) used the latent curve model to estimate the trajectories of children&#x00027;s strategy use and metamemory, which showed that the use of subsequent strategy is predictable by the metamemory at earlier time points.</p>
<p>Traditionally, the studies of children&#x00027;s cognitive development often relied on comparisons of manifest scale scores over time. For every child, the item scores of each scale would be averaged at each wave. The means were then compared using either paired-samples <italic>t-</italic>tests, when there were two measurement waves, or repeated measures ANOVA or other latent growth models, when more than two measurement waves were involved.</p>
<p>However, such a simple comparison of manifest scale scores over time may yield inaccurate results when the measurement of the underlying scale is not equivalent over time. That is because the manifest scale scores for the children&#x00027;s cognition scale depends not only on the latent true cognition score at each wave, but on the whole underlying measurement model (Steinmetz et al., <xref ref-type="bibr" rid="B37">2009</xref>). As the children&#x00027;s cognitive ability develops fast in the preschool period, a unified instrument of a cognitive test is most likely inappropriate across different ages (e.g., some items are too hard or too easy for different ages). Therefore, the measurement invariance of the scale should always be ensured in a longitudinal comparison (Marsh and Grayson, <xref ref-type="bibr" rid="B22">1994</xref>; Wu et al., <xref ref-type="bibr" rid="B41">2010</xref>). Otherwise, it would be difficult to explain whether the changes in the manifest scale scores are due to the actual cognitive development (changes in the latent means) or merely the changes of the measurement (Vaillancourt et al., <xref ref-type="bibr" rid="B39">2003</xref>). Thus, if the scale of the longitudinal measurement is not stable, conclusions derived from comparisons of manifest scale scores over time will be untrustworthy (Shadish et al., <xref ref-type="bibr" rid="B35">2002</xref>).</p>
<sec>
<title>Measurement invariance</title>
<p>Measurement invariance is defined as the stable property of psychometric features of an instrument across different situations or time periods (Mellenbergh, <xref ref-type="bibr" rid="B24">1989</xref>; Meredith and Millsap, <xref ref-type="bibr" rid="B25">1992</xref>).</p>
<p>Establishing measurement invariance is a critical requirement for making inferences about treatment effects and changes in constructs over time. Ensuring that the structure of the measures remains stable over time can reduce measurement error and maximize the interpretability of the findings (Pitts et al., <xref ref-type="bibr" rid="B29">1996</xref>). Therefore, longitudinal measurement invariance should be guaranteed before any further longitudinal analysis. Willoughby et al. (<xref ref-type="bibr" rid="B40">2012</xref>) investigated the longitudinal measurement invariance of Executive Function task battery before further longitudinal analysis, and found that two tasks exhibited partial measurement non-invariance, although the performance on the entire battery was stable over time.</p>
<p>Both the confirmatory factor analysis (CFA) method and the item response theory (IRT) method can be used to investigate measurement invariance. In CFA framework, a series of tests are required to investigate the measurement invariance, including tests for variance-covariance matrices, configural invariance, factor loadings invariance, intercept invariance, etc. (Schmitt and Kuljanin, <xref ref-type="bibr" rid="B33">2008</xref>). Unlike the CFA approach, which often examines the measurement invariance at a test level, IRT is conducted at both the overall test level and item level. The examination of measurement invariance in the IRT framework can provide information on whether the discrimination/slope (<italic>a</italic> parameters) or difficulty (<italic>b</italic> parameters) of each item has changed across different time periods or situations, which is beneficial to the revision of items. Meade et al. (<xref ref-type="bibr" rid="B23">2005</xref>) compared the CFA and IRT methods in establishing measurement invariance. By utilizing a longitudinal assessment of job satisfaction as an example, they demonstrated that the differences in items&#x00027; difficulty parameters over time could be effectively detected by IRT rather than CFA.</p>
<p>In many previous studies, researchers have attempted to examine measurement invariance before conducting longitudinal analysis and reported partial measurement invariance when some items showed drifted parameters (Willoughby et al., <xref ref-type="bibr" rid="B40">2012</xref>; Hakulinen et al., <xref ref-type="bibr" rid="B14">2014</xref>). For example, Meade et al. (<xref ref-type="bibr" rid="B23">2005</xref>) examined the measurement invariance of the instrument of job satisfaction using IRT, and the results indicated that three items functioned differently at Time 1 (T1) and Time 2 (T2). However, there has been a lack of discussions in literature about the solution to such a problem in longitudinal studies.</p>
<p>The solution proposed in the present study is &#x0201C;calibrated projection linking.&#x0201D; This is a newly-developed method, which was previously used in linking parallel tests (Thissen et al., <xref ref-type="bibr" rid="B38">2011</xref>; Cai, <xref ref-type="bibr" rid="B8">2015</xref>). Calibrated projection involves a two-tier IRT model (Cai, <xref ref-type="bibr" rid="B7">2010</xref>) to link two measures, which is distinct from the conventional calibration that requires the two measures to be of the same construct. It, therefore, allows the lack of measurement invariance of the instrument. In linking parallel tests without common items, the nearly identical item pairs in the two instruments were set to be common items to link scores on the PedsQL Symptoms Scale to the IRT metric of the PROMIS pediatric asthma impact scale (PAIS) (Thissen et al., <xref ref-type="bibr" rid="B38">2011</xref>).</p>
<p>The present study aims to explore the applications of calibrated projection to establish a longitudinal comparable scale of a 4- to 5-year-old children cognitive development test, of which some items lacked measurement invariance. The cognitive ability growth of the children from ages 4 to 5 is first described. The procedure of applying calibrated projection linking in the longitudinal studies is then illustrated with an example of a cognitive development test. Last, the performance of this method is presented and discussed. Overall, this study can be of interest to both substantive and methodological researchers.</p>
</sec>
</sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<sec>
<title>Measures</title>
<p>The instrument in this study is part of a series of instruments in a project on the Chinese national 3- to 6-year-old children&#x00027;s learning and development. The instrument was designed with heavy reference to the Chinese version of the Binet test and WISC-IV and was then refined after 30 psychologists expertized in the children&#x00027;s cognitive development were interviewed. The refined cognitive test consists of nine items: comparing quantity, orientation, addition and subtraction, jigsaw, classification, sorting, patterning, measurement, and fetching objects. Each item consists of four tasks at different levels, ranging from easy to difficult. The children started their test at different levels according to their ages. For instance, the 4-year-olds started at level 1, and 5-year-olds level 2. Only when they accomplished one task could they move on to the next level. At last, their performances were scored according to how many tasks they completed. Each task was worth 1 point, so the score of each item ranged from 0 to 4. Every child was tested by an experimenter who received professional training.</p>
</sec>
<sec>
<title>Samples</title>
<p>The data of this study were derived from part of the project mentioned above, which was conducted by UNICEF and Ministry of Education in China. The sample consists of 882 children from different provinces across China, including Inner Mongolia, Sichuan, Heilongjiang, Hebei, Jiangsu, and Fujian. Of all the 882 children, 422 (48%) were male and 460 (52%) were female. In addition, 460 (52%) were from urban areas and 422(48%) were from rural areas. All of the procedures conducted in the study were approved by the Institutional Review Board (IRB) and the participants&#x00027; parents.</p>
</sec>
<sec>
<title>Analysis</title>
<p>In the study, the two-tier IRT model (Cai, <xref ref-type="bibr" rid="B7">2010</xref>) was used to link the cognitive tests across two time points. The two-tier model describes the probability of each item response as a function of a set of item parameters and the latent variables measured by the scale. The rationale of adopting the two-tier model rests on the following facts. Firstly, from a substantive view, young children tend to develop an understanding of mathematical concepts, which are reflected by their informal ideas of more and less, taking away, shape, size, location, time, pattern, and position (Baroody et al., <xref ref-type="bibr" rid="B5">2006</xref>; Clements and Sarama, <xref ref-type="bibr" rid="B11">2009</xref>; Lee et al., <xref ref-type="bibr" rid="B21">2009</xref>). The two-tier model can model both the general factor (mathematic ability) and specific dimensions (e.g., &#x0201C;classification&#x0201D;) at the same time. Secondly, from a methodological view, the two-tier model is suitable for investigating a longitudinal study, since it takes into account the time effects of the general dimension, which represents the mathematical abilities at ages 4 and 5 (&#x003B8;<sub>1</sub> and &#x003B8;<sub>2</sub> in <bold>Figure 2</bold>). The two-tier model for graded response (Samejima, <xref ref-type="bibr" rid="B30">1969</xref>, <xref ref-type="bibr" rid="B31">1997</xref>; Cai, <xref ref-type="bibr" rid="B7">2010</xref>) is denoted as</p>
<disp-formula id="E1"><mml:math id="M1"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mtext>k</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mtext>k</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mtext>k</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>a</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mo>&#x003B8;</mml:mo></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x003B6;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;for&#x000A0;k</mml:mtext><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>5</mml:mn><mml:mo>&#x02026;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>Response&#x000A0;categories</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where subscript <italic>k</italic> represents the response category; <italic>j</italic> represents the items; subscript <italic>a</italic> represents the general dimension; subscript <italic>s</italic> represents specific dimensions; <bold>a</bold><sub><italic>j</italic></sub> is the vector of slope parameters; <bold>&#x003B8;</bold><sub><italic>j</italic></sub> is the vector of abilities at different time points; &#x003B2;<sub><italic>jk</italic></sub>is the intercept parameter; <italic><bold>a</bold></italic><sub><italic>a</italic></sub> is the slope parameter of general dimension; <italic>a</italic><sub>s</sub> is the slope parameter of specific dimension; &#x003B6;<sub>s</sub> is the ability of specific dimension.</p>
</sec>
<sec>
<title>Calibrated projection</title>
<p>Calibrated projection is a new statistical procedure that exploits the two-tier IRT model to link two measures (Thissen et al., <xref ref-type="bibr" rid="B38">2011</xref>). With calibrated projection, the item responses from two time points were fitted to the model presented above: &#x003B8;<sub>1</sub> denotes the underlying cognitive ability of the 4-year-olds, and &#x003B8;<sub>2</sub> denotes the cognitive ability of the 5-year-olds. Prior to this study, Ouyang et al. (<xref ref-type="bibr" rid="B28">2016</xref>) conducted a study that examined measurement invariance using the IRT model on the same instrument. The results of the study indicated that three items exhibited <italic>a</italic> parameter drift: &#x0201C;comparing quantity&#x0201D; (<inline-formula><mml:math id="M2"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>6</mml:mn><mml:mo>.</mml:mo><mml:mn>87</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>), &#x0201C;addition and subtraction&#x0201D; (<inline-formula><mml:math id="M3"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>14</mml:mn><mml:mo>.</mml:mo><mml:mn>17</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>), and &#x0201C;measurement&#x0201D; (<inline-formula><mml:math id="M4"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>6</mml:mn><mml:mo>.</mml:mo><mml:mn>86</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>). Five items exhibited category intercept <italic>(d)</italic> parameter drift: &#x0201C;comparing quantity&#x0201D; (<inline-formula><mml:math id="M5"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>50</mml:mn><mml:mo>.</mml:mo><mml:mn>71</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>), &#x0201C;addition and subtraction&#x0201D; (<inline-formula><mml:math id="M6"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>28</mml:mn><mml:mo>.</mml:mo><mml:mn>67</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>), &#x0201C;orientation&#x0201D; (<inline-formula><mml:math id="M7"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>140</mml:mn><mml:mo>.</mml:mo><mml:mn>34</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>), &#x0201C;jigsaw&#x0201D; (<inline-formula><mml:math id="M8"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>27</mml:mn><mml:mo>.</mml:mo><mml:mn>65</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>), and &#x0201C;classification&#x0201D; (<inline-formula><mml:math id="M9"><mml:mo>&#x00394;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>26</mml:mn><mml:mo>.</mml:mo><mml:mn>06</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>01</mml:mn></mml:math></inline-formula>). In order to place the items at two time points on the same scale, this study treated the invariant items as &#x0201C;common items,&#x0201D; which is shown in Figures <xref ref-type="fig" rid="F1">1</xref>, <xref ref-type="fig" rid="F2">2</xref>, and then the <italic>a</italic> parameters of each common item at T1 were set equal to the <italic>a</italic> parameters for its counterpart at T2. The common items for linking <italic>a</italic> parameters are &#x0201C;Orientation,&#x0201D; &#x0201C;Jigsaw,&#x0201D; &#x0201C;Classification,&#x0201D; &#x0201C;Sorting,&#x0201D; &#x0201C;Patterning,&#x0201D; and &#x0201C;Fetching Objects,&#x0201D; which are the bold lines in Figure <xref ref-type="fig" rid="F2">2</xref>. The category intercept parameters of each of the common items were also consistently set equal across the two time points. The common items for linking &#x003B2; parameters are: &#x0201C;Sorting,&#x0201D; &#x0201C;Patterning,&#x0201D; &#x0201C;Measurement,&#x0201D; and &#x0201C;Fetching Objects.&#x0201D; Then the IRT scale ability scores were estimated at two time points, and then transformed to T-scores for the convenience of comparison.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Linking design of 4&#x02013;5 years old children cognitive test.</p></caption>
<graphic xlink:href="fpsyg-09-00097-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Two-tier model for linking. Q, Comparing quantity; O, Orientation; AS, Addition and subtraction; J, Jigsaw; C, Classification; S, Sorting; P, Patterning; ME, Measurement; F, Fetching.</p></caption>
<graphic xlink:href="fpsyg-09-00097-g0002.tif"/>
</fig>
</sec>
<sec>
<title>Evaluation</title>
<p>In the study, as calibrated projection was applied in real samples rather than simulated samples, very few properties could be employed to evaluate the performance of the proposed approach. Therefore, the properties often applied in evaluating vertical scaling in longitudinal studies were adopted, which are grade-to-grade growth, grade-to-grade variability, and the separation of grade distributions (Kolen and Brennan, <xref ref-type="bibr" rid="B20">2004</xref>; Kim, <xref ref-type="bibr" rid="B19">2007</xref>). The method without calibrated projection was used as the baseline in the study. Thus, the performance of calibrated projection was evaluated by comparing the properties with the ones of the baseline method.</p>
<p>Grade-to-grade growth is defined as &#x0201C;the change from one grade to the next over the content taught in a particular grade&#x0201D; (Kolen and Brennan, <xref ref-type="bibr" rid="B20">2004</xref>, p. 377). The indicator of grade-to-grade growth is the mean difference between consecutive grades. The mean estimates are expected to increase with age regardless of content areas or the proficiency estimators used (Kim, <xref ref-type="bibr" rid="B19">2007</xref>). By examining the mean value at each age, questions like the following can be answered: how much do students grow, on average, from 1 year to the next? Are growth patterns different at different ages?</p>
<p>Grade-to-grade variability refers to the pattern of within-grade variability at different ages. Hoover (<xref ref-type="bibr" rid="B17">1984</xref>) argued that within-grade variability should increase with age because in young ages low-achieving students are expected to grow at a slower rate than high-achieving students. The indicator of grade-to-grade variability is the difference among the standard deviations (SDs) of each age on their own scale. Any dramatic change over grades, say 10 times larger or smaller, would indicate that the scale might not be functioning well.</p>
<p>The separation of grade distributions is the degree of overlap between scale score distributions of consecutive grades. One index of the separation of grade distributions is the horizontal distances between the distributions of consecutive grades (Holland, <xref ref-type="bibr" rid="B16">2002</xref>), which is based upon the difference between the two distributions at selected percentile points (Holland, <xref ref-type="bibr" rid="B16">2002</xref>; Kim, <xref ref-type="bibr" rid="B19">2007</xref>). To compute horizontal distances, certain percentile points of the score distributions must be selected. If <italic>p</italic> denotes a percentile point, then the p-percentile, <italic>X(p)</italic>, of the cumulative distribution function (CDF), <italic>F</italic>, is defined as</p>
<disp-formula id="E2"><label>(1)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><italic>X</italic>(<italic>p</italic>) is usually referred to as &#x0201C;the <italic>p</italic>th percentile&#x0201D; of <italic>F</italic>, and <italic>F</italic><sup>&#x02212;1</sup>(<italic>p</italic>) denotes the inverse function of <italic>F</italic>. Likewise, the percentiles of another CDF of G can be denoted by <italic>Y(p)</italic>. Then, using Equation (1),</p>
<disp-formula id="E3"><label>(2)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Then, the horizontal distance between two distributions of <italic>F</italic> and <italic>G, HD(p)</italic>, can be defined as:</p>
<disp-formula id="E4"><mml:math id="M12"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><italic>HD</italic>(<italic>p</italic>)represents the difference between the <italic>p</italic>th percentiles of the two distributions. For example, for the distribution of 4-year-old children, <italic>F</italic>, the percentile rank of a &#x003B8; of 1, is 50. For the distribution of 5-year-old children, G, the percentile rank of a &#x003B8; of 1.3, is 50. By Equations (1&#x02013;2), the horizontal distance between the two distributions of <italic>F</italic> and <italic>G</italic> at the 50th percentile is 0.3.</p>
<p>Horizontal distances were computed to examine the gaps at selected locations throughout the entire distributions: 5th, 10th, 25th, 50th, 75th, 90th, and 95th. In the present study, horizontal distances of the scale with calibration projection applied were compared with the ones of the baseline method. Any dramatic change over grades or percentile points, say 10 times larger or smaller, would indicate that the scale might not be functioning well (Kim, <xref ref-type="bibr" rid="B19">2007</xref>).</p>
<p>Calibrated projection process was conducted using IRTPRO 2.1, and the outputs from IRTPRO 2.1 were then analyzed using SPSS20.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<sec>
<title>Descriptive statistics</title>
<p>Table <xref ref-type="table" rid="T1">1</xref> shows the reliability of the children&#x00027;s cognitive test at two time points. In psychological tests, alpha coefficient 0.7 is the cut-off value for being acceptable (Santos, <xref ref-type="bibr" rid="B32">1999</xref>). As our cognitive tests only included 9 items, the reliability of the test is acceptable.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Cronbach&#x00027;s alpha coefficient.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Time points</bold></th>
<th valign="top" align="center"><bold>Cronbach&#x00027;s &#x003B1; coefficient</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Wave 1</td>
<td valign="top" align="center">0.70</td>
</tr>
<tr>
<td valign="top" align="left">Wave 2</td>
<td valign="top" align="center">0.72</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Linking the children&#x00027;s cognitive longitudinal test</title>
<p>The present study used the two-tier IRT model (Cai, <xref ref-type="bibr" rid="B7">2010</xref>) to link the test across two waves. Table <xref ref-type="table" rid="T2">2</xref> shows the item parameters of the cognitive test at two time points. The invariant item parameters across two time periods were bold in the table. The 3rd&#x02212;7th columns show the slope (<italic>a</italic>) parameters on the general dimension representing the time effects of the mathematics test and the category intercept (&#x003B2;) parameters that were freely estimated in the two-tier model. In the 9th&#x02212;13th columns are the item parameters obtained after the common items were constrained to be equal. The 8th and 14th columns show the item slope (<italic>s</italic>) parameters on specific dimensions of mathematics, such as &#x0201C;comparing quantity,&#x0201D; &#x0201C;addition and subtraction,&#x0201D; and &#x0201C;orientation,&#x0201D; etc., which were fixed because the contents of the nine items did not change in two waves. The correlation between cognitive abilities of the 4- and 5-year-olds is 0.86.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Item parameters with and without linking.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Items</bold></th>
<th valign="top" align="left"><bold>Time points</bold></th>
<th valign="top" align="center" colspan="6" style="border-bottom: thin solid #000000;"><bold>Without linking</bold></th>
<th valign="top" align="center" colspan="6" style="border-bottom: thin solid #000000;"><bold>With linking</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="center"><bold><italic>a</italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>1</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>2</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>3</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>4</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>s</italic></bold></th>
<th valign="top" align="center"><bold><italic>a</italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>1</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>2</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>3</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>&#x003B2;<sub>4</sub></italic></bold></th>
<th valign="top" align="center"><bold><italic>s</italic></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Q</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">2.09</td>
<td valign="top" align="center">&#x02212;0.45</td>
<td valign="top" align="center">&#x02212;1.26</td>
<td valign="top" align="center">&#x02212;2.14</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">2.16</td>
<td valign="top" align="center">&#x02212;0.37</td>
<td valign="top" align="center">&#x02212;1.19</td>
<td valign="top" align="center">&#x02212;2.07</td>
<td valign="top" align="center">0.50</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.06</td>
<td valign="top" align="center">2.16</td>
<td valign="top" align="center">&#x02212;0.55</td>
<td valign="top" align="center">&#x02212;1.23</td>
<td valign="top" align="center">&#x02212;1.51</td>
<td/>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">1.54</td>
<td valign="top" align="center">&#x02212;1.17</td>
<td valign="top" align="center">&#x02212;1.86</td>
<td valign="top" align="center">&#x02212;2.14</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">O</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.15</td>
<td valign="top" align="center">2.57</td>
<td valign="top" align="center">&#x02212;0.13</td>
<td valign="top" align="center">&#x02212;2.90</td>
<td valign="top" align="center">&#x02212;6.99</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center"><bold>1.03</bold></td>
<td valign="top" align="center">2.58</td>
<td valign="top" align="center">&#x02212;0.04</td>
<td valign="top" align="center">&#x02212;2.75</td>
<td valign="top" align="center">&#x02212;6.80</td>
<td valign="top" align="center">0.64</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.10</td>
<td valign="top" align="center">4.38</td>
<td valign="top" align="center">1.65</td>
<td valign="top" align="center">&#x02212;0.82</td>
<td valign="top" align="center">&#x02212;2.71</td>
<td/>
<td valign="top" align="center"><bold>1.03</bold></td>
<td valign="top" align="center">3.74</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">&#x02212;1.55</td>
<td valign="top" align="center">&#x02212;3.47</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">AS</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.20</td>
<td valign="top" align="center">0.18</td>
<td valign="top" align="center">&#x02212;0.13</td>
<td valign="top" align="center">&#x02212;1.31</td>
<td valign="top" align="center">&#x02212;3.04</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">1.19</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">&#x02212;1.16</td>
<td valign="top" align="center">&#x02212;2.88</td>
<td valign="top" align="center">0.49</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">3.43</td>
<td valign="top" align="center">2.83</td>
<td valign="top" align="center">&#x02212;0.20</td>
<td valign="top" align="center">&#x02212;2.29</td>
<td/>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">2.93</td>
<td valign="top" align="center">2.33</td>
<td valign="top" align="center">&#x02212;0.71</td>
<td valign="top" align="center">&#x02212;2.79</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">J</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.02</td>
<td valign="top" align="center">1.41</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">&#x02212;2.79</td>
<td valign="top" align="center">&#x02212;4.54</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center"><bold>1.10</bold></td>
<td valign="top" align="center">1.61</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">&#x02212;2.66</td>
<td valign="top" align="center">&#x02212;4.42</td>
<td valign="top" align="center">0.93</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.38</td>
<td valign="top" align="center">2.96</td>
<td valign="top" align="center">2.53</td>
<td valign="top" align="center">&#x02212;1.94</td>
<td valign="top" align="center">&#x02212;2.99</td>
<td/>
<td valign="top" align="center"><bold>1.10</bold></td>
<td valign="top" align="center">2.17</td>
<td valign="top" align="center">1.75</td>
<td valign="top" align="center">&#x02212;2.64</td>
<td valign="top" align="center">&#x02212;3.68</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">C</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.03</td>
<td valign="top" align="center">1.77</td>
<td valign="top" align="center">&#x02212;0.97</td>
<td valign="top" align="center">&#x02212;4.36</td>
<td valign="top" align="center">&#x02212;5.88</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center"><bold>1.00</bold></td>
<td valign="top" align="center">1.88</td>
<td valign="top" align="center">&#x02212;0.85</td>
<td valign="top" align="center">&#x02212;4.23</td>
<td valign="top" align="center">&#x02212;5.74</td>
<td valign="top" align="center">0.72</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.16</td>
<td valign="top" align="center">2.33</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">&#x02212;2.40</td>
<td valign="top" align="center">&#x02212;3.71</td>
<td/>
<td valign="top" align="center"><bold>1.00</bold></td>
<td valign="top" align="center">1.65</td>
<td valign="top" align="center">&#x02212;0.39</td>
<td valign="top" align="center">&#x02212;3.10</td>
<td valign="top" align="center">&#x02212;4.42</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">C</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.60</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">&#x02212;0.49</td>
<td valign="top" align="center">&#x02212;1.49</td>
<td valign="top" align="center">&#x02212;2.47</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center"><bold>1.66</bold></td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.32</bold></td>
<td valign="top" align="center"><bold>&#x02212;1.24</bold></td>
<td valign="top" align="center"><bold>&#x02212;2.23</bold></td>
<td valign="top" align="center">0.71</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">2.26</td>
<td valign="top" align="center">1.87</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">&#x02212;0.13</td>
<td valign="top" align="center">&#x02212;1.19</td>
<td/>
<td valign="top" align="center"><bold>1.66</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.85</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.32</bold></td>
<td valign="top" align="center"><bold>&#x02212;1.24</bold></td>
<td valign="top" align="center"><bold>&#x02212;2.23</bold></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">P</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.34</td>
<td valign="top" align="center">2.89</td>
<td valign="top" align="center">0.16</td>
<td valign="top" align="center">&#x02212;0.64</td>
<td valign="top" align="center">&#x02212;1.40</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center"><bold>1.23</bold></td>
<td valign="top" align="center"><bold>3.00</bold></td>
<td valign="top" align="center"><bold>0.31</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.58</bold></td>
<td valign="top" align="center"><bold>&#x02212;1.32</bold></td>
<td valign="top" align="center">0.60</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.32</td>
<td valign="top" align="center">4.01</td>
<td valign="top" align="center">1.20</td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">&#x02212;0.48</td>
<td/>
<td valign="top" align="center"><bold>1.23</bold></td>
<td valign="top" align="center"><bold>3.00</bold></td>
<td valign="top" align="center"><bold>0.31</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.58</bold></td>
<td valign="top" align="center"><bold>&#x02212;1.32</bold></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">ME</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">&#x02212;1.08</td>
<td valign="top" align="center">&#x02212;2.35</td>
<td valign="top" align="center">&#x02212;3.40</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center"><bold>0.68</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.90</bold></td>
<td valign="top" align="center"><bold>&#x02212;2.04</bold></td>
<td valign="top" align="center"><bold>&#x02212;3.20</bold></td>
<td valign="top" align="center">0.76</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.36</td>
<td valign="top" align="center">1.49</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">&#x02212;1.05</td>
<td valign="top" align="center">&#x02212;2.25</td>
<td/>
<td valign="top" align="center">1.25</td>
<td valign="top" align="center"><bold>0.68</bold></td>
<td valign="top" align="center"><bold>&#x02212;0.90</bold></td>
<td valign="top" align="center"><bold>&#x02212;2.04</bold></td>
<td valign="top" align="center"><bold>&#x02212;3.20</bold></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">F</td>
<td valign="top" align="left">T1</td>
<td valign="top" align="center">1.17</td>
<td valign="top" align="center">4.15</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">&#x02212;2.26</td>
<td valign="top" align="center">&#x02212;3.74</td>
<td valign="top" align="center">1.54</td>
<td valign="top" align="center"><bold>1.29</bold></td>
<td valign="top" align="center"><bold>4.35</bold></td>
<td valign="top" align="center"><bold>1.28</bold></td>
<td valign="top" align="center"><bold>&#x02212;2.02</bold></td>
<td valign="top" align="center"><bold>&#x02212;3.54</bold></td>
<td valign="top" align="center">1.52</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">T2</td>
<td valign="top" align="center">1.44</td>
<td valign="top" align="center">4.98</td>
<td valign="top" align="center">2.29</td>
<td valign="top" align="center">&#x02212;1.09</td>
<td valign="top" align="center">&#x02212;2.63</td>
<td/>
<td valign="top" align="center"><bold>1.29</bold></td>
<td valign="top" align="center"><bold>4.35</bold></td>
<td valign="top" align="center"><bold>1.28</bold></td>
<td valign="top" align="center"><bold>&#x02212;2.02</bold></td>
<td valign="top" align="center"><bold>&#x02212;3.54</bold></td>
<td/>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Q, Comparing quantity; O, Orientation; AS, Addition and subtraction; J, Jigsaw; C, Classification; S, Sorting; P, Patterning; ME, Measurement; F, Fetching Objects</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>The slope parameters represent the discrimination of the items in the IRT framework, and 0.64 or greater is considered as moderate or high discrimination (Baker, <xref ref-type="bibr" rid="B4">2001</xref>). Thus, all items, except &#x0201C;Comparing quantity,&#x0201D; were highly discriminating. &#x003B2; parameter represents the category intercept parameter, which is opposite to difficulty parameter. The higher the &#x003B2; parameter is, the easier the task level. For example, in Table <xref ref-type="table" rid="T2">2</xref> &#x0201C;Measurement&#x0201D; is more difficult than &#x0201C;Patterning&#x0201D; at all task levels. By comparing the category intercepts of &#x0201C;addition and subtraction,&#x0201D; it can be seen that &#x003B2;<sub>1</sub> and &#x003B2;<sub>2</sub> drifted severely across two waves (0.34, 0.22 at the first wave and 2.93, 2.33 at the second wave), which indicates that the first and second task levels may have been too easy for 5-year-old children. Furthermore, the slope parameters can be compared between general and specific dimensions. For example, the slope parameters of &#x0201C;addition and subtraction&#x0201D; on the general dimension are 1.20 and 0.87, and the one on the specific dimension is 0.49. This indicates that &#x0201C;addition and subtraction&#x0201D; explained more variability of the entire mathematics test. In contrast, the slope parameters of &#x0201C;fetching objects&#x0201D; on the general dimension are 1.17 and 1.44, and the one on the specific dimension is 1.54. This indicates that this item is highly related to both specific and general dimensions.</p>
</sec>
<sec>
<title>Evaluation</title>
<p>After the calibrated scale was established, the ability parameters were estimated, and then transformed to T-score for convenience, which is</p>
<disp-formula id="E5"><mml:math id="M13"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>T</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>50</mml:mn><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The ability distributions without linking and with linking are both shown below. Figure <xref ref-type="fig" rid="F3">3</xref> is the histogram and normalized ability distribution of the baseline method, which means all slope parameters on the general dimension of the nine items in the two-tier model were freely estimated. In this figure, the upper graph shows the cognitive ability distribution of the 4-year-old children, and the lower graph shows that of the 5-year-old children. The red lines in both graphs denote the means of the two distributions. The mean of 4-year-old children&#x00027;s cognitive ability is 44.36, and that of the 5-year-olds is 51.96. Figure <xref ref-type="fig" rid="F4">4</xref> shows the normalized cognitive ability distributions of 4- to 5-year-old children after calibrated projection was applied. The mean of the 4-year-olds is 43.08, and that of the 5-year-olds is 59.12. By comparing Figures <xref ref-type="fig" rid="F3">3</xref>, <xref ref-type="fig" rid="F4">4</xref>, it can be seen that with calibrated projection, the ability distributions presented a larger growth across two waves, which is aligned with the findings of some previous studies that the preschool is a key period in which children&#x00027;s cognitive ability grows rapidly (Chang, <xref ref-type="bibr" rid="B10">2009</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Ability distribution without calibrated projection.</p></caption>
<graphic xlink:href="fpsyg-09-00097-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Ability distribution after calibrated projection.</p></caption>
<graphic xlink:href="fpsyg-09-00097-g0004.tif"/>
</fig>
</sec>
<sec>
<title>Grade-to-grade growth</title>
<p>In the present study, grade-to-grade growth means the average ability growth between 4- and 5-year-old children. In Figure <xref ref-type="fig" rid="F5">5</xref>, the average ability score is increased by 7.59 without linking and 16.04 with linking. This difference indicates that, with calibrated projection, there was a larger growth in the cognitive ability of children from ages 4 to 5.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparison between 4 and 5 years old ability means without and with linking.</p></caption>
<graphic xlink:href="fpsyg-09-00097-g0005.tif"/>
</fig>
<p>Furthermore, by paired sample <italic>T</italic>-test, the significance and effect size of children&#x00027;s ability growth with and without linking were compared. Table <xref ref-type="table" rid="T3">3</xref> shows that regardless of whether the linking was applied or not, the ability growths of children from ages 4 to 5 are both significant. However, the effect size with linking is almost twice as large as the one of the baseline method.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Paired sample <italic>T</italic>-test for score difference between 4- and 5-year old children without and with linking.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold><italic>SD</italic></bold></th>
<th valign="top" align="center"><bold>SE</bold></th>
<th valign="top" align="center"><bold><italic>t</italic></bold></th>
<th valign="top" align="center"><bold>Cohen&#x00027;s d</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Without linking</td>
<td valign="top" align="center">&#x02212;7.60</td>
<td valign="top" align="center">2.82</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">&#x02212;80.12<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.10</td>
</tr>
<tr>
<td valign="top" align="left">With linking</td>
<td valign="top" align="center">&#x02212;16.04</td>
<td valign="top" align="center">3.07</td>
<td valign="top" align="center">0.10</td>
<td valign="top" align="center">&#x02212;155.03<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.19</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN1"><label>&#x0002A;&#x0002A;&#x0002A;</label><p><italic>p &#x0003C; 0.001</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Grade-to-grade variability</title>
<p>Figure <xref ref-type="fig" rid="F6">6</xref> shows the differences in the standard deviations between the two scales obtained with and without linking. With linking, the SDs at two time points are 8.79 and 9.51, respectively. Without linking, however, the SDs at two time points are 8.80 and 8.10, respectively, showing a &#x0201C;scale shrinkage&#x0201D; problem (Hoover, <xref ref-type="bibr" rid="B17">1984</xref>, <xref ref-type="bibr" rid="B18">1988</xref>). It implies a decrease in variability of the score with age.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Comparison between 4 and 5 years old ability standard deviation without and with linking.</p></caption>
<graphic xlink:href="fpsyg-09-00097-g0006.tif"/>
</fig>
</sec>
<sec>
<title>Separation of grade distributions</title>
<p>The separation of grade distributions is mainly demonstrated by horizontal distance (HD, Holland, <xref ref-type="bibr" rid="B16">2002</xref>). To examine gaps between distributions of consecutive grades, the HDs were computed at the following selected percentile points: 5th, 10th, 25th, 50th, 75th, 90th, and 95th, and then averaged. As is shown in Table <xref ref-type="table" rid="T4">4</xref>, the average HD is 7.56 without linking and 16.04 with linking. This result indicates that the difference of ability distributions between two time points is larger after calibrated projection was applied, which is consistent with the result of grade-to-grade growth.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Comparison between 4 and 5 years old ability average HD without and with linking.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Percentile HD</bold></th>
<th valign="top" align="center"><bold>Without linking</bold></th>
<th valign="top" align="center"><bold>With linking</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">5th</td>
<td valign="top" align="center">8.78</td>
<td valign="top" align="center">15.12</td>
</tr>
<tr>
<td valign="top" align="left">10th</td>
<td valign="top" align="center">8.38</td>
<td valign="top" align="center">14.96</td>
</tr>
<tr>
<td valign="top" align="left">25th</td>
<td valign="top" align="center">8.01</td>
<td valign="top" align="center">15.37</td>
</tr>
<tr>
<td valign="top" align="left">50th</td>
<td valign="top" align="center">7.90</td>
<td valign="top" align="center">16.20</td>
</tr>
<tr>
<td valign="top" align="left">75th</td>
<td valign="top" align="center">6.67</td>
<td valign="top" align="center">16.12</td>
</tr>
<tr>
<td valign="top" align="left">90th</td>
<td valign="top" align="center">6.69</td>
<td valign="top" align="center">16.94</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left">95th</td>
<td valign="top" align="center">6.51</td>
<td valign="top" align="center">17.56</td>
</tr> <tr>
<td valign="top" align="left">Average HD</td>
<td valign="top" align="center">7.56</td>
<td valign="top" align="center">16.04</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>In longitudinal studies, measurement invariance is a significant property that needs to be established before any further analysis is conducted. In the field of children&#x00027;s cognitive development, as the children&#x00027;s cognitive abilities grow very fast during the childhood, the instrument can be aptly drifted across different ages. The previous study of examining the longitudinal measurement invariance (Ouyang et al., <xref ref-type="bibr" rid="B28">2016</xref>) showed that among nine items, 3 <italic>a</italic> parameters and 5 category intercept (&#x003B2;) parameters presented a drift across two waves, although the construct of the test over two waves remained stable by reference to the high correlation of 0.86. As so many item parameters were drifted over time, the reliability or predictive validity of the test could have been compromised (e.g., Alvares and Hulin, <xref ref-type="bibr" rid="B1">1972</xref>; Henry and Hulin, <xref ref-type="bibr" rid="B15">1987</xref>). In order to achieve a more accurate measurement of children&#x00027;s cognitive developing trajectory from 4- to 5-year-old, calibrated projection was applied to establish a comparable scale in this longitudinal test.</p>
<p>Calibrated projection was mostly applied to link parallel tests in previous studies (e.g., Thissen et al., <xref ref-type="bibr" rid="B38">2011</xref>; Monroe et al., <xref ref-type="bibr" rid="B26">2014</xref>). The present study extended the method to reduce the impact of lack of measurement invariance in longitudinal tests. Calibrated projection is based on the two-tier IRT model (Cai, <xref ref-type="bibr" rid="B7">2010</xref>), of which each item loads on both general dimensions and a specific dimension. In the previous studies of linking parallel tests, the item parameters that loaded on the specific dimension representing the same content of the items across two tests were set equal so as to play roles of &#x0201C;common items.&#x0201D; However, in the present longitudinal study the common item parameters that load on the general dimension representing the time effect were set equal during the process of estimation. Furthermore, the category intercept (&#x003B2;) parameters of common items were also set to be equal to reduce the item difficulty parameters drift. The IRT method of examining measurement invariance can provide more information about the performance of different items. For example, Table <xref ref-type="table" rid="T2">2</xref> shows that the first two tasks of &#x0201C;addition and subtraction&#x0201D; may be too easier for 5 years old children, which need revision in future.</p>
<p>In order to compare the ability scales established by the proposed method and the baseline method, three evaluation criteria that have been used in vertical scaling were adopted in the study. They are grade-to-grade growth, grade-to-grade variability, and separation of grade distributions (Kim, <xref ref-type="bibr" rid="B19">2007</xref>). These criteria were represented by mean difference, standard deviation (SD), and average horizontal distance (HD), respectively (Hoover, <xref ref-type="bibr" rid="B17">1984</xref>; Camilli, <xref ref-type="bibr" rid="B9">1988</xref>; Kim, <xref ref-type="bibr" rid="B19">2007</xref>).</p>
<p>First of all, with calibrated projection, the mean difference shows a larger growth of children&#x00027;s cognitive ability from ages 4 to 5. Furthermore, the statistical test shows that the mean differences are significant in both cases, but the effect size was larger with linking. The growth pattern with calibration project applied in this study provides strong supports for many previous studies about the rapid development of children&#x00027;s cognition from ages 4 to 5 (Chang, <xref ref-type="bibr" rid="B10">2009</xref>; Zhao, <xref ref-type="bibr" rid="B43">2009</xref>). For example, Yang (<xref ref-type="bibr" rid="B42">2009</xref>) suggested that Chinese children younger than 4 can only accomplish the task of sorting 4 items, while children at 5 can accomplish 10 items.</p>
<p>Secondly, the result of the grade-to-grade variability shows that within-grade SD with calibrated projection applied increases with age, which supports Hoover&#x00027;s (<xref ref-type="bibr" rid="B17">1984</xref>) study. Hoover explained the reason of the result, based on the expectation of a slower growth rate of low-achieving students than high-achieving students at young ages. This growth pattern of mathematic ability across preschool years is also supported by other studies (Bast and Reitsma, <xref ref-type="bibr" rid="B6">1997</xref>; Aunola et al., <xref ref-type="bibr" rid="B3">2002</xref>, <xref ref-type="bibr" rid="B2">2004</xref>). In addition, the grade-to-grade variability shows a decrease with age with the baseline method, which indicates &#x0201C;scale shrinkage&#x0201D; (Kim, <xref ref-type="bibr" rid="B19">2007</xref>). According to Hoover (<xref ref-type="bibr" rid="B17">1984</xref>, <xref ref-type="bibr" rid="B18">1988</xref>), scale shrinkage is not very common in real data. Camilli (<xref ref-type="bibr" rid="B9">1988</xref>) indicated that scale shrinkage may be caused by systematic estimation error or measurement error, drawing on the findings of some simulation studies that if the variability of item parameters is set to be different, then more scale shrinkage problems would occur. Therefore, scale shrinkage problem that occurred in the cognitive ability distributions with the baseline method suggests that there might exist some systematic estimation error.</p>
<p>Thirdly, the separation of grade distributions represents the difference between the ability distributions in two waves. In the present study, the comparison of average HD indicates that with calibrated projection, the difference between ability distributions from age 4 to 5 is larger, which is similar to the result of grade-to-grade growth. The consistency between the HD results and the results of grade-to-grade growth is also indicated in some previous studies (Kim, <xref ref-type="bibr" rid="B19">2007</xref>).</p>
</sec>
<sec id="s5">
<title>Limitations and conclusion</title>
<p>Despite the strengths of this study, there also exist a few limitations. Firstly, the sample only covered 4- to 5-year-old children, and the 1-year range was limited for further analysis. In some vertical scaling studies, the age range is often 3&#x02013;4 years or more, which would yield more information about the scaling method by comparing evaluation properties across every 2 consecutive years. Thus, in future, the age range should be extended to 3&#x02013;6 years old, the whole preschool stage for Chinese children, so as to investigate whether calibrated projection will performance in the same way when used to evaluate the growth across other consecutive ages. Secondly, as this study was conducted with a real sample, it constrained the use of the evaluation criteria. Because there is no standard value or true value for the grade-to-grade growth, variability, and the separation of grade distributions for comparisons. The evaluation in this study was mostly based on the results of previous studies about their performances in different situations. In the future, more simulation studies are needed to evaluate the performance of calibrated projection, so that the results can be compared with other linking or vertical scaling methods in longitudinal studies.</p>
<p>Despite these limitations, the current study demonstrates ways of applying calibrated projection method to link longitudinal tests when there occur item parameter drifts in the instrument across different waves. This is critical, as changes in the psychometric properties of a test over time could sacrifice its reliability or predictive validity (e.g., Alvares and Hulin, <xref ref-type="bibr" rid="B1">1972</xref>; Henry and Hulin, <xref ref-type="bibr" rid="B15">1987</xref>). Furthermore, the results of this study indicated that with linking, grade-to-grade growth and its effect size are larger. The result of grade-to-grade variability after linking is aligned with the result of the study of Hoover (<xref ref-type="bibr" rid="B17">1984</xref>) and shows less scale shrinkage, which indicates a smaller measurement error. Moreover, the conspicuous separation of grade distributions supports the result of grade-to-grade growth. In summary, comparisons of the three properties showed a possible consequence of ignoring the measurement invariance in a longitudinal analysis as well as the performance of calibrated projection from a practical view.</p>
</sec>
<sec id="s6">
<title>Ethics statement</title>
<p>The study was approved by the Institutional Review Board (IRB) of Beijing Normal University. All the parents of participants provided written informed consent.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>XO wrote the first draft of the manuscript and assisted study design, and data analyses. QZ revised it critically for important intellectual content. TX was the principal investigator of the study and the data provider. All of the authors participated in the final approval of the version to be published and agreed to be accountable for all aspects of the work.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alvares</surname> <given-names>K. M.</given-names></name> <name><surname>Hulin</surname> <given-names>C. L.</given-names></name></person-group> (<year>1972</year>). <article-title>Two explanations of temporal changes in ability-skill relationships: a literature review and theoretical analysis</article-title>. <source>Hum. Factors</source> <volume>14</volume>, <fpage>295</fpage>&#x02013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1177/001872087201400402</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aunola</surname> <given-names>K.</given-names></name> <name><surname>Leskinen</surname> <given-names>E.</given-names></name> <name><surname>Lerkkanen</surname> <given-names>M. K.</given-names></name> <name><surname>Nurmi</surname> <given-names>J. E.</given-names></name></person-group> (<year>2004</year>). <article-title>Developmental dynamics of math performance from preschool to grade 2</article-title>. <source>J. Educ. Psychol.</source> <volume>96</volume>:<fpage>699</fpage>. <pub-id pub-id-type="doi">10.1037/0022-0663.96.4.699</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aunola</surname> <given-names>K.</given-names></name> <name><surname>Leskinen</surname> <given-names>E.</given-names></name> <name><surname>Onatsu-Arvilommi</surname> <given-names>T.</given-names></name> <name><surname>Nurmi</surname> <given-names>J. E.</given-names></name></person-group> (<year>2002</year>). <article-title>Three methods for studying developmental change: a case of reading skills and self-concept</article-title>. <source>Br. J. Educ. Psychol.</source> <volume>72</volume>, <fpage>343</fpage>&#x02013;<lpage>364</lpage>. <pub-id pub-id-type="doi">10.1348/000709902320634447</pub-id><pub-id pub-id-type="pmid">12396310</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baker</surname> <given-names>F. B.</given-names></name></person-group> (<year>2001</year>). <source>The Basics of Item Response Theory</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>ERIC Publications</publisher-name>.</citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baroody</surname> <given-names>A. J.</given-names></name> <name><surname>Lai</surname> <given-names>M.-l.</given-names></name> <name><surname>Mix</surname> <given-names>K. S.</given-names></name></person-group> (<year>2006</year>). <article-title>The development of young children&#x00027;s early number and operation sense and its implications for early childhood education</article-title>, in <source>Handbook of Research on the Education of Young Children</source>, eds <person-group person-group-type="editor"><name><surname>Spodek</surname> <given-names>B.</given-names></name> <name><surname>Saracho</surname> <given-names>O. N.</given-names></name></person-group> (<publisher-loc>Mahwah, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates Publishers</publisher-name>), <fpage>187</fpage>&#x02013;<lpage>221</lpage>.</citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bast</surname> <given-names>J.</given-names></name> <name><surname>Reitsma</surname> <given-names>P.</given-names></name></person-group> (<year>1997</year>). <article-title>Mathew effects in reading: a comparison of latent growth curve models and simplex models with structured means</article-title>. <source>Multivariate Behav. Res.</source> <volume>32</volume>, <fpage>135</fpage>&#x02013;<lpage>167</lpage>. <pub-id pub-id-type="doi">10.1207/s15327906mbr3202_3</pub-id><pub-id pub-id-type="pmid">26788756</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>L.</given-names></name></person-group> (<year>2010</year>). <article-title>A two-tier full-information item factor analysis model with applications</article-title>. <source>Psychometrika</source> <volume>75</volume>, <fpage>581</fpage>&#x02013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-010-9178-0</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Lord&#x02013;Wingersky algorithm version 2.0 for hierarchical item factor models with applications in test scoring, scale alignment, and model fit testing</article-title>. <source>Psychometrika</source> <volume>80</volume>, <fpage>535</fpage>&#x02013;<lpage>559</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-014-9411-3</pub-id><pub-id pub-id-type="pmid">25233839</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Camilli</surname> <given-names>G.</given-names></name></person-group> (<year>1988</year>). <article-title>Scale shrinkage and the estimation of latent distribution parameters</article-title>. <source>J. Educ. Stat.</source> <volume>13</volume>, <fpage>227</fpage>&#x02013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.3102/10769986013003227</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <source>The Study on the Development of Composition of 2-Dimensional Geometric Figures in Young Children Aged 3-6 (in Chinese).</source> Master dissertation, <publisher-name>East China Normal University</publisher-name>.</citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Clements</surname> <given-names>D. H.</given-names></name> <name><surname>Sarama</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>Learning trajectories in early mathematics&#x02013;sequences of acquisition and teaching</article-title>, in <source>Encyclopedia of Language and Literacy Development</source>, eds <person-group person-group-type="editor"><name><surname>New</surname> <given-names>R. S.</given-names></name> <name><surname>Cochran</surname> <given-names>M.</given-names></name></person-group> (<publisher-loc>London, ON</publisher-loc>: <publisher-name>Canadian Language and Literacy Research Network</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>7</lpage>.</citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grammer</surname> <given-names>J. K.</given-names></name> <name><surname>Coffman</surname> <given-names>J. L.</given-names></name> <name><surname>Ornstein</surname> <given-names>P. A.</given-names></name> <name><surname>Morrison</surname> <given-names>F. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Change over time: conducting longitudinal studies of children&#x00027;s cognitive development</article-title>. <source>J. Cogn. Dev.</source> <volume>14</volume>, <fpage>515</fpage>&#x02013;<lpage>528</lpage>. <pub-id pub-id-type="doi">10.1080/15248372.2013.833925</pub-id><pub-id pub-id-type="pmid">24955035</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grammer</surname> <given-names>J. K.</given-names></name> <name><surname>Purtell</surname> <given-names>K. M.</given-names></name> <name><surname>Coffman</surname> <given-names>J. L.</given-names></name> <name><surname>Ornstein</surname> <given-names>P. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Relations between children&#x00027;s metamemory and strategic performance: time-varying covariates in early elementary school</article-title>. <source>J. Exp. Child Psychol.</source> <volume>108</volume>, <fpage>139</fpage>&#x02013;<lpage>155</lpage>. <pub-id pub-id-type="doi">10.1016/j.jecp.2010.08.001</pub-id><pub-id pub-id-type="pmid">20863515</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hakulinen</surname> <given-names>C.</given-names></name> <name><surname>Jokela</surname> <given-names>M.</given-names></name> <name><surname>Keltikangas-J&#x000E4;rvinen</surname> <given-names>L.</given-names></name> <name><surname>Merjonen</surname> <given-names>P.</given-names></name> <name><surname>Raitakari</surname> <given-names>O. T.</given-names></name> <name><surname>Hintsanen</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Longitudinal measurement invariance, stability and change of anger and cynicism</article-title>. <source>J. Behav. Med.</source> <volume>37</volume>, <fpage>434</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1007/s10865-013-9501-1</pub-id><pub-id pub-id-type="pmid">23479114</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Henry</surname> <given-names>R. A.</given-names></name> <name><surname>Hulin</surname> <given-names>C. L.</given-names></name></person-group> (<year>1987</year>). <article-title>Stability of skilled performance across time: some generalizations and limitations on utilities</article-title>. <source>J. Appl. Psychol.</source> <volume>72</volume>:<fpage>457</fpage>. <pub-id pub-id-type="doi">10.1037/0021-9010.72.3.457</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holland</surname> <given-names>P. W.</given-names></name></person-group> (<year>2002</year>). <article-title>Two measures of change in the gaps between the CDFs of test-score distributions</article-title>. <source>J. Educ. Behav. Stat.</source> <volume>27</volume>, <fpage>3</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.3102/10769986027001003</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoover</surname> <given-names>H. D.</given-names></name></person-group> (<year>1984</year>). <article-title>The most appropriate scores for measuring educational development in the elementary schools: GE&#x00027;s</article-title>. <source>Educ. Meas.</source> <volume>3</volume>, <fpage>8</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1111/j.1745-3992.1984.tb00768.x</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoover</surname> <given-names>H. D.</given-names></name></person-group> (<year>1988</year>). <article-title>Growth expectations for low-achieving students: a reply to yen</article-title>. <source>Educ. Meas.</source> <volume>7</volume>, <fpage>21</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1111/j.1745-3992.1988.tb00841.x</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <source>A Comparison of Calibration Methods and Proficiency Estimators for Creating IRT Vertical Scales.</source> Doctoral dissertation, <publisher-name>The University of Iowa</publisher-name>.</citation></ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kolen</surname> <given-names>M. J.</given-names></name> <name><surname>Brennan</surname> <given-names>R. L.</given-names></name></person-group> (<year>2004</year>). <source>Test Equating, Scaling, and Linking.</source> <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Ng</surname> <given-names>E. L.</given-names></name> <name><surname>Ng</surname> <given-names>S. F.</given-names></name></person-group> (<year>2009</year>). <article-title>The contributions of working memory and executive functioning to problem representation and solution generation in algebraic word problems</article-title>. <source>J. Educ. Psychol.</source> <volume>101</volume>, <fpage>373</fpage>. <pub-id pub-id-type="doi">10.1037/a0013843</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsh</surname> <given-names>H. W.</given-names></name> <name><surname>Grayson</surname> <given-names>D.</given-names></name></person-group> (<year>1994</year>). <article-title>Longitudinal stability of latent means and individual differences: a unified approach</article-title>. <source>Struct. Equat. Model. Multidisc. J.</source> <volume>1</volume>, <fpage>317</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1080/10705519409539984</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meade</surname> <given-names>A. W.</given-names></name> <name><surname>Lautenschlager</surname> <given-names>G. J.</given-names></name> <name><surname>Hecht</surname> <given-names>J. E.</given-names></name></person-group> (<year>2005</year>). <article-title>Establishing measurement equivalence and invariance in longitudinal data with item response theory</article-title>. <source>Int. J. Test.</source> <volume>5</volume>, <fpage>279</fpage>&#x02013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.1207/s15327574ijt0503_6</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mellenbergh</surname> <given-names>G. J.</given-names></name></person-group> (<year>1989</year>). <article-title>Item bias and item response theory</article-title>. <source>Int. J. Educ. Res.</source> <volume>13</volume>, <fpage>127</fpage>&#x02013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.1016/0883-0355(89)90002-5</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meredith</surname> <given-names>W.</given-names></name> <name><surname>Millsap</surname> <given-names>R. E.</given-names></name></person-group> (<year>1992</year>). <article-title>On the misuse of manifest variables in the detection of measurement bias</article-title>. <source>Psychometrika</source> <volume>57</volume>, <fpage>289</fpage>&#x02013;<lpage>311</lpage>. <pub-id pub-id-type="doi">10.1007/BF02294510</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Monroe</surname> <given-names>S.</given-names></name> <name><surname>Cai</surname> <given-names>L.</given-names></name> <name><surname>Choi</surname> <given-names>K.</given-names></name></person-group> (<year>2014</year>). <source>Student Growth Percentiles Based on MIRT: Implications of Calibrated Projection</source>. CRESST Report 842, <publisher-name>National Center for Research on Evaluation, Standards, and Student Testing (CRESST)</publisher-name>.</citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ornstein</surname> <given-names>P. A.</given-names></name> <name><surname>Haden</surname> <given-names>C. A.</given-names></name></person-group> (<year>2001</year>). <article-title>Memory development or the development of memory?</article-title> <source>Curr. Dir. Psychol. Sci.</source> <volume>10</volume>, <fpage>202</fpage>&#x02013;<lpage>205</lpage>. <pub-id pub-id-type="doi">10.1111/1467-8721.00149</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ouyang</surname> <given-names>X. Z.</given-names></name> <name><surname>Tian</surname> <given-names>W.</given-names></name> <name><surname>Xin</surname> <given-names>T.</given-names></name> <name><surname>Zhan</surname> <given-names>P. D.</given-names></name></person-group> (<year>2016</year>). <article-title>Use IRT to analyze longitudinal data measurement invariance &#x02013; the case of 4-5 Children&#x00027;s cognitive ability test</article-title>. <source>Psychol. Sci.</source> <volume>39</volume>, <fpage>606</fpage>&#x02013;<lpage>613</lpage>. <pub-id pub-id-type="doi">10.16719/j.cnki.1671-6981.20160315</pub-id> [in Chinese].</citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pitts</surname> <given-names>S. C.</given-names></name> <name><surname>West</surname> <given-names>S. G.</given-names></name> <name><surname>Tein</surname> <given-names>J. Y.</given-names></name></person-group> (<year>1996</year>). <article-title>Longitudinal measurement models in evaluation research: examining stability and change</article-title>. <source>Eval. Program Plann.</source> <volume>19</volume>, <fpage>333</fpage>&#x02013;<lpage>350</lpage>. <pub-id pub-id-type="doi">10.1016/S0149-7189(96)00027-4</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samejima</surname> <given-names>F.</given-names></name></person-group> (<year>1969</year>). <article-title>Estimation of latent ability using a response pattern of graded scores</article-title>. <source>Psychometr. Monogr.</source> <volume>34</volume>, <fpage>1</fpage>&#x02013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1007/BF03372160</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Samejima</surname> <given-names>F.</given-names></name></person-group> (<year>1997</year>). <article-title>Graded response model</article-title>, in <source>Handbook of Modern Item Response Theory</source>, eds <person-group person-group-type="editor"><name><surname>van der Linden</surname> <given-names>W. J.</given-names></name> <name><surname>Hambleton</surname> <given-names>R. K.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>85</fpage>&#x02013;<lpage>100</lpage>.</citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santos</surname> <given-names>J. R. A.</given-names></name></person-group> (<year>1999</year>). <article-title>Cronbach&#x00027;s alpha: a tool for assessing the reliability of scales</article-title>. <source>J. Extens.</source> <volume>37</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmitt</surname> <given-names>N.</given-names></name> <name><surname>Kuljanin</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Measurement invariance: review of practice and implications</article-title>. <source>Hum. Resour. Manage. Rev.</source> <volume>18</volume>, <fpage>210</fpage>&#x02013;<lpage>222</lpage>. <pub-id pub-id-type="doi">10.1016/j.hrmr.2008.03.003</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schneider</surname> <given-names>W.</given-names></name> <name><surname>Sodian</surname> <given-names>B.</given-names></name></person-group> (<year>1997</year>). <article-title>Memory strategy development: lessons from longitudinal research</article-title>. <source>Dev. Rev.</source> <volume>17</volume>, <fpage>442</fpage>&#x02013;<lpage>461</lpage>. <pub-id pub-id-type="doi">10.1006/drev.1997.0441</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shadish</surname> <given-names>W. R.</given-names></name> <name><surname>Cook</surname> <given-names>T. D.</given-names></name> <name><surname>Campbell</surname> <given-names>D. T.</given-names></name></person-group> (<year>2002</year>). <article-title>Construct validity and external validity</article-title>, in <source>Experimental and Quasi-Experimental Designs for Generalized Causal Inference</source>, eds <person-group person-group-type="editor"><name><surname>Shadish</surname> <given-names>W. R.</given-names></name> <name><surname>Cook</surname> <given-names>T. D.</given-names></name> <name><surname>Campbell</surname> <given-names>D. T.</given-names></name></person-group> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Houghton Mifflin</publisher-name>), <fpage>64</fpage>&#x02013;<lpage>102</lpage>.</citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sodian</surname> <given-names>B.</given-names></name> <name><surname>Schneider</surname> <given-names>W.</given-names></name></person-group> (<year>1999</year>). <article-title>Memory strategy development-gradual increase, sudden insight, or roller coaster?</article-title>, in <source>Individual development from 3 to 12: Findings from the Munich Longitudinal Study</source>, eds <person-group person-group-type="editor"><name><surname>Weinert</surname> <given-names>F. E.</given-names></name> <name><surname>Schneider</surname> <given-names>W.</given-names></name></person-group> (<publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>), <fpage>61</fpage>&#x02013;<lpage>77</lpage>.</citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steinmetz</surname> <given-names>H.</given-names></name> <name><surname>Schmidt</surname> <given-names>P.</given-names></name> <name><surname>Tina-Booh</surname> <given-names>A.</given-names></name> <name><surname>Wieczorek</surname> <given-names>S.</given-names></name> <name><surname>Schwartz</surname> <given-names>S. H.</given-names></name></person-group> (<year>2009</year>). <article-title>Testing measurement invariance using multigroup CFA: differences between educational groups in human values measurement</article-title>. <source>Qual. Quant.</source> <volume>43</volume>, <fpage>599</fpage>&#x02013;<lpage>616</lpage>. <pub-id pub-id-type="doi">10.1007/s11135-007-9143-x</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thissen</surname> <given-names>D.</given-names></name> <name><surname>Varni</surname> <given-names>J. W.</given-names></name> <name><surname>Stucky</surname> <given-names>B. D.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Irwin</surname> <given-names>D. E.</given-names></name> <name><surname>DeWalt</surname> <given-names>D. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Using the PedsQL&#x02122; 3.0 asthma module to obtain scores comparable with those of the PROMIS pediatric asthma impact scale (PAIS)</article-title>. <source>Qual. Life Res.</source> <volume>20</volume>, <fpage>1497</fpage>&#x02013;<lpage>1505</lpage>. <pub-id pub-id-type="doi">10.1007/s11136-011-9874-y</pub-id><pub-id pub-id-type="pmid">21384264</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vaillancourt</surname> <given-names>T.</given-names></name> <name><surname>Brendgen</surname> <given-names>M.</given-names></name> <name><surname>Boivin</surname> <given-names>M.</given-names></name> <name><surname>Tremblay</surname> <given-names>R. E.</given-names></name></person-group> (<year>2003</year>). <article-title>A longitudinal confirmatory factor analysis of indirect and physical aggression: evidence of two factors over time?</article-title>. <source>Child Dev.</source> <volume>74</volume>, <fpage>1628</fpage>&#x02013;<lpage>1638</lpage>. <pub-id pub-id-type="doi">10.1046/j.1467-8624.2003.00628.x</pub-id><pub-id pub-id-type="pmid">14669886</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Willoughby</surname> <given-names>M. T.</given-names></name> <name><surname>Blair</surname> <given-names>C. B.</given-names></name> <name><surname>Wirth</surname> <given-names>R. J.</given-names></name> <name><surname>Greenberg</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>The measurement of executive function at age 5: psychometric properties and relationship to academic achievement</article-title>. <source>Psychol. Assess.</source> <volume>24</volume>, <fpage>226</fpage>. <pub-id pub-id-type="doi">10.1037/a0025361</pub-id><pub-id pub-id-type="pmid">21966934</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>A. D.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Gadermann</surname> <given-names>A. M.</given-names></name> <name><surname>Zumbo</surname> <given-names>B. D.</given-names></name></person-group> (<year>2010</year>). <article-title>Multiple-indicator multilevel growth model: a solution to multiple methodological challenges in longitudinal studies</article-title>. <source>Soc. Indic. Res.</source> <volume>97</volume>, <fpage>123</fpage>&#x02013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1007/s11205-009-9496-8</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z. Y.</given-names></name></person-group> (<year>2009</year>). <source>The Study on the Development of Seriation and Event Seriation in Young Children Aged 3-6 (in Chinese)</source>. Master dissertation, <publisher-name>East China Normal University</publisher-name>.</citation></ref>
<ref id="B43">
<citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Z. G.</given-names></name></person-group> (<year>2009</year>). <source>The Study on the Relationship Among Numerosity Estimation, Counting Ability and Visual-Spatial Cognitive Ability in Young Children Aged 3-6 (in Chinese).</source> Doctoral dissertation, <publisher-name>East China Normal University</publisher-name>.</citation></ref>
</ref-list> 
</back>
</article>
