<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Educ.</journal-id>
<journal-title>Frontiers in Education</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Educ.</abbrev-journal-title>
<issn pub-type="epub">2504-284X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/feduc.2022.854378</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Education</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Text Mining to Alleviate the Cold-Start Problem of Adaptive Comparative Judgments</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>De Vrindt</surname> <given-names>Michiel</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1632412/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Van den Noortgate</surname> <given-names>Wim</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/571860/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Debeer</surname> <given-names>Dries</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/538058/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Imec Research Group ITEC, KU Leuven</institution>, <addr-line>Kortrijk</addr-line>, <country>Belgium</country></aff>
<aff id="aff2"><sup>2</sup><institution>Faculty of Psychology and Educational Sciences, KU Leuven</institution>, <addr-line>Leuven</addr-line>, <country>Belgium</country></aff>
<aff id="aff3"><sup>3</sup><institution>Faculty of Psychology and Educational Sciences, Ghent University</institution>, <addr-line>Ghent</addr-line>, <country>Belgium</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Sven De Maeyer, University of Antwerp, Belgium</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Jinnie Shin, University of Florida, United States; Elise Crompvoets, Tilburg University, Netherlands</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Michiel De Vrindt <email>michiel.de.vrindt&#x00040;gmail.com</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Assessment, Testing and Applied Measurement, a section of the journal Frontiers in Education</p></fn></author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>07</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>7</volume>
<elocation-id>854378</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 De Vrindt, Van den Noortgate and Debeer.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>De Vrindt, Van den Noortgate and Debeer</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Comparative judgments permit the assessment of open-ended student works by constructing a latent quality scale through repeated pairwise comparisons (i.e., which works &#x0201C;win&#x0201D; or &#x0201C;lose&#x0201D;). Adaptive comparative judgments speed up the judgment process by maximizing the Fisher information of the next comparison. However, at the start of a judgment process, such an adaptive algorithm will not perform well. In order to reliably approximate the Fisher Information of possible pairs well, multiple comparisons are needed. In addition, adaptive comparative judgments have been shown to inflate the scale separation coefficient, which is a reliability estimator for the quality estimates. Current methods to solve the inflation issue increase the number of required comparisons. The goal of this study is to alleviate the cold-start problem of adaptive comparative judgments for essays or other textual assignments, but also to minimize the bias of the scale separation coefficient. By using text-mining techniques, which can be performed before the first judgment, essays can be adaptively compared from the start. More specifically, we propose a selection rule that is based both on a high (1) cosine similarity of the vector representations and (2) Fisher Information of essay pairs. At the start of the judgment process, the cosine similarity has the highest weight in the selection rule. With more judgments, this weight decreases progressively, whereas the weight of the Fisher Information increases. Using simulated data, the proposed strategy is compared with existing approaches. The results indicate that the proposed selection rule can mitigate both the cold-start. That is, fewer judgments are needed to obtain accurate and reliable quality estimates. In addition, the selection rule was found to reduce the inflation of the scale separation reliability.</p></abstract>
<kwd-group>
<kwd>text mining</kwd>
<kwd>natural language processing</kwd>
<kwd>comparative judgments</kwd>
<kwd>educational assessment</kwd>
<kwd>computational linguistics</kwd>
<kwd>psychometrics</kwd>
<kwd>educational technology</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="2"/>
<equation-count count="11"/>
<ref-count count="28"/>
<page-count count="16"/>
<word-count count="10457"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>For rubric marking of students&#x00027; works, assessors are required to isolate and accurately evaluate the criteria of the works. Grades or marks follow from how well certain criteria or the so-called &#x02018;grade-descriptors&#x00027; are satisfied (Pollitt, <xref ref-type="bibr" rid="B21">2004</xref>). Especially when the students&#x00027; works are open-ended (e.g., essay text, portfolios, and mathematical proofs), rubric marking can be a difficult task for assessors (Jones and Inglis, <xref ref-type="bibr" rid="B15">2015</xref>; Jones et al., <xref ref-type="bibr" rid="B14">2019</xref>). Even when assessors are well-experienced, their assessments are likely to be influenced by earlier assessments, inevitably making the given grades relative to some extent. The method of comparative judgments (CJ), as introduced by Thurstone (<xref ref-type="bibr" rid="B26">1927</xref>), directly exploits the relative aspect of assessing open ended works. In CJ, rather than assessing individual works, pairs of works are holistically and repeatedly compared. That is, assessors (or judges) are not required to assign a grade on a specific (or multiple) grading scale(s); they only need to select the better work (i.e., the winner) of each pair that was assigned to them. Consequently, differences in rater severity (i.e., assessors that systemically score more severe or more lenient) and differences in perceived qualities between assessors become negligible (Pollitt, <xref ref-type="bibr" rid="B22">2012</xref>). Based on the win-lose judgments of the comparisons, quality estimates of the students&#x00027; works are obtained. As such, CJ allows a reliable and valid assessment of open-ended works that require subjective judgments. In addition to the capability of creating a valid and reliable quality scale, the process of CJ has proven to decrease the cognitive load that is required for the assessment process and develops the assessor&#x00027;s assessment skills (Coenen et al., <xref ref-type="bibr" rid="B8">2018</xref>). From the students&#x00027; perspective, CJ can include quantitative and qualitative feedback. Quantitative feedback is directly available from the final rank-order of essays, whereas quantitative feedback can be incorporated by including assessors&#x00027; remarks (e.g., strong and weak points of essays). Hence, CJ can be used for both summative and formative assessments.</p>
<p>The original CJ algorithm pairs students&#x00027; works randomly. A drawback of random pairings is that it typically requires many comparisons to obtain sufficiently reliable quality estimates. Consequently, the assessors&#x00027; workload can be high. Several strategies have been proposed to minimize the number of pairwise comparisons while maintaining the reliability of the quality estimates and the final ranking of the works. Generally, these strategies try to make the repeated selection of pairs as optimal as possible (Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>; Bramley and Vitello, <xref ref-type="bibr" rid="B6">2019</xref>; Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>). For instance, Pollitt (<xref ref-type="bibr" rid="B22">2012</xref>) proposed a selection rule that speeds up the &#x0201C;scale-building&#x0201D; process by repeatedly selecting the pair for which a comparison would add the most information to the estimated qualities. More specifically, pairs are selected so that the expected Fisher Information of each next comparison is maximized based on the current quality estimates (refer to below). Because the quality estimates are repeatedly updated during the process, and because, based on the updated estimates, the most informative pair is repeatedly selected, this selection algorithm will be referred to as &#x0201C;adaptive comparative judgments&#x0201D; (ACJ).</p>
<p>Adaptive comparative judgments has two important shortcomings. First, at the start of the judgment process, the pairings cannot be made adaptively because quality estimates are only available after a minimal number of comparisons. This issue is typically referred to as the &#x0201C;cold-start problem.&#x0201D; Current implementations of ACJ generally select the initial pairs randomly, where the adaptive selection starts only after these initial random pairings. Yet the first adaptive pairings are highly determined by the outcomes of the initial comparisons and judgments. Thus, if by chance low-quality works are paired with other low-quality works, it is possible that a low-quality work &#x0201C;wins&#x0201D; multiple initial comparisons, resulting in a high first quality estimate. When a low-quality work with a high first quality estimate is subsequently paired with a high-quality work (which is likely in the first ACJ-based pairings), the obtained judgment will have a limited contribution to the final quality estimate and ranking. Moreover, it may take multiple additional comparisons before the quality estimate of the low-quality work is properly adjusted and ACJ can have its beneficial impact. To prevent this behavior, Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) proposed a selection rule that introduces randomness in the selection of initial pairs while Rangel-Smith and Lynch (<xref ref-type="bibr" rid="B23">2018</xref>) selected initial pairs with more different initial quality estimates. Yet, although these selection rules may reduce the probability of strong distortions in ACJ, it also reduces the efficiency of the judgment process.</p>
<p>Second, the adaptive selection of pairs based on the maximum Fisher Information typically pairs work with similar true qualities. Therefore, low-quality works are often compared with other low-quality works and high-quality works with other high-quality works. These adaptive comparisons not only increase the reliability of the quality estimates (i.e., they lower the standard errors), but they also tend to inflate the estimated quality scale when the number of comparisons is still small (i.e., the estimated qualities are more extreme than the true qualities) (Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>). The combination of lower standard errors and an inflated latent scale can cause inflation of the scale separation reliability (SSR), which is a commonly used estimator for the reliability of the obtained quality estimates (Bramley, <xref ref-type="bibr" rid="B5">2015</xref>; Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>; Bramley and Vitello, <xref ref-type="bibr" rid="B6">2019</xref>; Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>). This is problematic because the SSR is typically used to decide when to stop the ACJ process. That is, the ACJ process is typically stopped when predefined reliability, as estimated by the SSR, is reached. When the SSR is overestimated due to the adaptive selection algorithm, there is a risk that the ACJ process is stopped prematurely. Indeed, Bramley (<xref ref-type="bibr" rid="B5">2015</xref>) reported that for true reliability of 0.70, an SSR of 0.95 may be expected when using ACJ. Moreover, Bramley and Vitello (<xref ref-type="bibr" rid="B6">2019</xref>) compared the quality estimates of the works that were obtained using ACJ and a limited number of comparisons per work, with quality estimates obtained by comparing every work to every other work (&#x0201C;all-by-all&#x0201D; design). The SD of the ACJ-obtained scale was 0.391 times larger than the SD of the all-by-all-obtained scale.</p>
<p>The issue of the SSR inflation in ACJ is widely known and some solutions have been proposed. These solutions consist of modifying the assessment design in order to increase the number of comparisons, add randomness to the adaptive selection algorithm or impose a minimal difference between quality parameter estimates to be selected (Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>; Bramley and Vitello, <xref ref-type="bibr" rid="B6">2019</xref>; Crompvoets et al., <xref ref-type="bibr" rid="B10">2021</xref>). Yet all strategies decrease the efficiency of the judgment process (i.e., more comparisons are required). Therefore, in this study, we explore a new strategy to alleviate the cold-start problem and reduce the SSR inflation in ACJ. We focus on the application of ACJ to assess textual works and propose the use of text-mining techniques to obtain numerical representations of the texts that capture semantic and syntactical information. Based on these numerical representations, the semantic and syntactical similarities of the texts can be computed. Because both the text mining techniques and the computation of the similarities can be performed before the start of the ACJ process, the initial pairings can be based on the similarities of the texts, rather than randomly pairing texts. As such, the cold-start problem and the SSR inflation may be mitigated. We explore different text mining techniques and evaluate our strategy using two sets of textual works.</p>
<p>In the remainder of this article, we first introduce the Bradley-Terry-Luce model (Bradley and Terry, <xref ref-type="bibr" rid="B4">1952</xref>) for comparative judgment data and discuss the ACJ process in more detail (Pollitt, <xref ref-type="bibr" rid="B22">2012</xref>). After presenting the SSR reliability estimator, the proposed text-mining strategy is explained, including the necessary text-pre-processing for extracting textual information. Different representation techniques are considered: term frequency-inverse document frequency (Aizawa, <xref ref-type="bibr" rid="B2">2003</xref>), averaged word embeddings (Mikolov et al., <xref ref-type="bibr" rid="B19">2013</xref>), and document embeddings (Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). Subsequently, we explain how the textual representations can be used to select initial pairs of essays by computing the similarity between the texts. More specifically, we propose a new progressive selection rule, in which the adaptive selection rule gradually becomes more important. We illustrate the proposed strategy using two real essay sets. Moreover, using simulated data the performance of the new progressive selection rule and the different text representation techniques is evaluated. The impact on the SSR inflation and the precision of the quality estimates is compared across conditions. After discussing the results, limitations and future research opportunities are discussed.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>2. Methods</title>
<sec>
<title>2.1. Comparative Judgements-Design</title>
<sec>
<title>2.1.1. Bradley-Terry-Luce Model</title>
<p>Let there be a set <italic>S</italic> of <italic>N</italic> works that should be assessed. Consider work <italic>i</italic> and work <italic>j</italic> with <italic>j</italic> and <italic>i</italic> in <italic>S</italic>. According to the Bradley-Terry-Luce model (BTL), the probability that work <italic>i</italic> wins over work <italic>j</italic> in a comparison, Pr(<italic>x</italic><sub><italic>ij</italic></sub> &#x0003D; 1), depends on the quality parameters &#x003B8;<sub><italic>i</italic></sub> and &#x003B8;<sub><italic>j</italic></sub> of work <italic>i</italic> and <italic>j</italic>, respectively (Bradley and Terry, <xref ref-type="bibr" rid="B4">1952</xref>):</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">Pr</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">where</mml:mtext><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0007E;</mml:mo><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Based on the win-lose (i.e., 0, 1) data of many comparisons, the vector of all quality parameters <italic><bold>&#x003B8;</bold></italic><sub>1&#x000D7;<italic>N</italic></sub> can then be estimated by applying maximum-likelihood based methods to the BTL (Hunter, <xref ref-type="bibr" rid="B13">2004</xref>).</p>
</sec>
<sec>
<title>2.1.2. Adaptive Comparative Judgement</title>
<p>When &#x003B8;<sub><italic>i</italic></sub> &#x0003D; &#x003B8;<sub><italic>j</italic></sub> (i.e., the works <italic>i</italic> and <italic>j</italic> have equal quality parameters), then following Equation (1), the probability that work <italic>i</italic> wins over work <italic>j</italic> in a comparison is equal to Pr(<italic>x</italic><sub><italic>ij</italic></sub> &#x0003D; 1|&#x003B8;<sub><italic>i</italic></sub>, &#x003B8;<sub><italic>j</italic></sub>) &#x0003D; 0.5. Moreover, the outcome for comparisons with Pr(<italic>x</italic><sub><italic>ij</italic></sub> &#x0003D; 1|&#x003B8;<sub><italic>i</italic></sub>, &#x003B8;<sub><italic>j</italic></sub>) &#x0003D; 0.5 has the highest possible variance <inline-formula><mml:math id="M3"><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>25</mml:mn></mml:math></inline-formula>, and the expected Fisher information will be maximal. Therefore, the outcome of such a comparison will add the maximal amount of information to the estimation for the quality parameters (Pollitt, <xref ref-type="bibr" rid="B21">2004</xref>). For ACJ as in Pollitt (<xref ref-type="bibr" rid="B22">2012</xref>), the works with the smallest difference in estimated quality parameters will be paired together, because the computed Fisher information is highest for these pairs.</p>
<p>Although the BTL allows multiple comparisons between pairs of works, CJ and ACJ typically restrict the number of comparisons per pair (by a single rater) to be maximally one: <italic>x</italic><sub><italic>ij</italic></sub> &#x0003D; {0, 1} (<italic>i</italic> &#x02260; <italic>j</italic>). For <italic>N</italic> works, there are <inline-formula><mml:math id="M4"><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula> unique comparisons. We denote this set of unique comparisons as <italic>B</italic>. In addition, let <italic>B</italic><sub><italic>m</italic></sub> be the set of unique pairs that is not yet compared after the <italic>m</italic>th judgment. Hence, generally in ACJ, the pair that will be selected for the <italic>m</italic>&#x0002B;1th comparison is the pair with the highest expected Fisher information <inline-formula><mml:math id="M5"><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> (i.e., with the smallest distance between the quality estimates <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M7"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>) in <italic>B</italic><sub><italic>m</italic></sub>.</p>
<p>Which pair has the highest Fisher information changes through the ACJ process because the quality estimates are continuously updated. Originally, Pollitt (<xref ref-type="bibr" rid="B22">2012</xref>) proposed to update all quality estimates <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> simultaneously after &#x02018;a round of comparisons&#x00027;in which all works were compared once. However, because updating and re-estimating <inline-formula><mml:math id="M9"><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:math></inline-formula> only after a certain number of comparisons results in a selection of pairs that are not based on the most up-to-date quality estimates (Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>), <inline-formula><mml:math id="M10"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> is updated after every single comparison <italic>m</italic> in this study.</p>
<p>To repeatedly estimate the quality parameters after each comparison <italic>m</italic>, an expectation maximization algorithm is used (Hunter, <xref ref-type="bibr" rid="B13">2004</xref>). Formally, for comparison <italic>m</italic> &#x0002B; 1 all qualities <inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:math></inline-formula> for work <italic>i</italic>, &#x02026;, <italic>N</italic> are estimated using:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M13"><mml:mrow><mml:msubsup><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:msubsup><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mstyle></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>where <italic>n</italic><sub><italic>ij</italic></sub> is an indicator variable indicating whether work <italic>i</italic> and <italic>j</italic> are compared yet and <italic>x</italic><sub><italic>i</italic></sub> is the total number of wins of work <italic>i</italic>.</p>
<p>After updating every <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, all quality parameters are centered so that the mean of the quality estimates will be zero (Equation 4). If the work has not been compared yet or it loses every comparison, its quality estimate is unidentifiable. To make the quality parameters identifiable, a small quantity is added to <italic>x</italic><sub><italic>ij</italic></sub> (i.e., 10<sup>&#x02212;3</sup>) (Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>).</p>
</sec>
<sec>
<title>2.1.3. Stochastic Adaptive Comparative Judgments</title>
<p>In the original ACJ algorithm by Pollitt (<xref ref-type="bibr" rid="B22">2012</xref>), only the point estimates of the quality parameters are considered in the selection algorithm. However, the uncertainty of these point estimates can be large, especially at the beginning of the ACJ process when there are few judgments per work. In order to also consider the uncertainty of the quality estimates, Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) included the standard error of the quality estimate in the selection algorithm. That is, for comparison <italic>m</italic> &#x0002B; 1, first the work <italic>i</italic> with the largest standard error of the quality estimate <inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> is selected. Then, rather than selecting the work <italic>j</italic> for which <inline-formula><mml:math id="M16"><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is maximized (with the comparison of <italic>i</italic> and <italic>j</italic> still in <italic>B</italic><sub><italic>m</italic></sub>), the work <italic>j</italic> is randomly selected from all candidates left in <italic>B</italic><sub><italic>m</italic></sub> with a probability that is a function of the distance between <inline-formula><mml:math id="M17"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M19"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula>. More specifically, the selection probabilities are proportional to the densities of the <inline-formula><mml:math id="M20"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> in a normal distribution with mean <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and variance <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> (Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>).</p>
<p>This adaptive selection rule is stochastic and introduces randomness to the algorithm. If few comparisons have been made with work <italic>i</italic>, the normal distribution of the quality parameter will have wider tails, which causes the selection rule to be more random. As more comparisons are made, the normal distribution will become more peaked and student works with similar quality parameters will be selected with a higher probability. A drawback of this algorithm is that only <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> is considered. <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> is not taken into account.</p>
<p>To compute the standard error of a quality parameter estimate <inline-formula><mml:math id="M25"><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> after each comparison, the observed Fisher Information function with respect to <inline-formula><mml:math id="M26"><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> given all the judgment outcomes <bold>x</bold> is used:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x02202;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>&#x02113;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle><mml:mo>|</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E6"><label>(6)</label><mml:math id="M28"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>x</italic><sub><italic>ij</italic></sub> is 1 when work <italic>i</italic> wins the comparison over <italic>j</italic> (<italic>x</italic><sub><italic>ij</italic></sub> &#x0003D; 1 &#x02212; <italic>x</italic><sub><italic>ji</italic></sub>). In Equation (6), superscript (<italic>m</italic>) is dropped for the ease of reading. In this article, the &#x02018;stochastic ACJ&#x00027; selection rule by Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) is used for all ACJ.</p>
</sec>
<sec>
<title>2.1.4. SSR as Reliability Estimator</title>
<p>If the true quality parameters <italic><bold>&#x003B8;</bold></italic> of a set of works are known, the reliability of the estimated qualities <inline-formula><mml:math id="M29"><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:math></inline-formula> can be obtained from the squared Pearson correlation of the true quality and estimated parameters <inline-formula><mml:math id="M30"><mml:msubsup><mml:mrow><mml:mi>&#x003C1;</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>. This corresponds to the ratio of the variance of the true quality levels and the variance of the estimated quality parameters. The more similar the variances are, the higher the reliability will be. In practice, the reliability of the assessment is an important criterion. Often, a minimum value for reliability is required. In real assessment situations, however, the true quality parameters are not available, which makes it impossible to compute the reliability as <inline-formula><mml:math id="M31"><mml:msubsup><mml:mrow><mml:mi>&#x003C1;</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B8;</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>.</p>
<p>An estimator for the reliability that can be computed without the true quality parameters is the Scale Separation Reliability (SSR), which is based on the estimated quality parameters and their uncertainty (Brennan, <xref ref-type="bibr" rid="B7">2010</xref>). To compute the SSR, the unknown true variance of the quality parameters, denoted &#x003C3;<sup>2</sup>, is approximated by the difference between the variance of the quality estimates, denoted <inline-formula><mml:math id="M32"><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and the mean squared error of the standard errors of the quality estimates, <inline-formula><mml:math id="M33"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>. The SSR is defined as:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M34"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mtext class="textrm" mathvariant="normal">MSE</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E8"><label>(8)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">with MSE</mml:mtext><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle class="math"><mml:mtext class="textrm" mathvariant="normal">E</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Equation (7) indicates that a higher variance of the quality estimates and smaller standard errors of the estimates will lead to a higher SSR. For the full derivation of the SSR, refer to Verhavert et al. (<xref ref-type="bibr" rid="B28">2018</xref>). For the SSR to be estimable, <inline-formula><mml:math id="M36"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>&#x0003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M37"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>&#x02265;</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">E</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> must hold.</p>
</sec>
<sec>
<title>2.1.5. Vector Representations of Essays</title>
<p>Numerical representations of texts should capture the most important features of the texts, both with respect to syntax and semantics. Statistical language modeling allows the mapping of natural unstructured text to a vector of numeric values. We consider three representation techniques to represent essay tests as numerical vectors: term frequency-inverse document frequency (&#x0201C;tf-idf&#x0201D;) (Aizawa, <xref ref-type="bibr" rid="B2">2003</xref>), averaged word embeddings (Mikolov et al., <xref ref-type="bibr" rid="B19">2013</xref>), and document embeddings (Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). A brief explanation of the construction of the three representation techniques will be given.</p>
<p>First, tf-idf representations are constructed based on word frequencies: the relative frequency of words in a document is offset by how often words appear across documents (Aizawa, <xref ref-type="bibr" rid="B2">2003</xref>). A word that occurs frequently in a document but that doesn&#x00027;t occur often in other documents, receives a higher weight. However, because it only considers word frequencies, tf-idf is limited in terms of extracting syntactical meanings. One way to extract syntactical information is by grouping sequences of words that often occur together, called &#x0201C;n-grams.&#x0201D; Yet even in the case of n-grams, tf-idf representations do not incorporate the syntactical meaning of texts apart from relations between n-grams. In addition, because every word (or n-gram) across the documents corresponds to one dimension, tf-idf representations are typically highly dimensional.</p>
<p>Second, average word embeddings are a more complex representation technique that incorporates syntactical information and that is not highly dimensional. Average word embeddings are distributional representations based on the so-call &#x0201C;skip-gram word embeddings&#x0201D; neural network architecture. In the skip-gram architecture, a shallow neural network is constructed with a word as input and its surrounding words as output (Mikolov et al., <xref ref-type="bibr" rid="B19">2013</xref>). The &#x0201C;embeddings&#x0201D; are the weights of the hidden layer in the neural network, which are obtained from predicting the set of surrounding words for each input word. The predicted surrounding words are the words that have the largest probability on average as given by the sigmoid function of the dot product of the embeddings of each surrounding word with the input word. However, iterating over all possible combinations of surrounding words and calculating probabilities is computationally intensive. As an alternative, the objective function is minimized by correctly distinguishing between surrounding words and sampled non-surrounding words (i.e., &#x0201C;negative sampling&#x0201D;). Ultimately, essay representations are obtained from the average pooling of the word embeddings of all the words in each essay. A disadvantage of averaged word embeddings is that it does not account for the dependence of the meaning of words coming from the document (or essay) they are part of.</p>
<p>Finally, document embeddings are an extension of word embeddings that allow for this document-dependence (Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). Instead of learning embeddings on the level of words and aggregating it to embeddings of documents, document embeddings can be learned directly. The distributed continuous-bag-of-words architecture are neural networks that predict whether words occur in a given document. The words are those with the highest probability on average as given by the sigmoid function of the dot product of a document embedding and word embeddings. Negative sampling is also possible by sampling words that do not occur in a given document. The distributed bag-of-words architecture for document embeddings can be initialized by a pre-trained set of word embeddings (Tulkens et al., <xref ref-type="bibr" rid="B27">2016</xref>). The pre-trained model consists of embeddings of words that are trained on a very large corpus of texts. The reason for using a large corpus is that words can be learned from or &#x02018;embedded&#x00027; in many different contexts. Pre-trained models are often used in natural language processing as sample corpora are often not large enough. If the contexts in which words are learned are very different from those in the essay texts, then the pre-trained word embeddings would not be fit. However, this possibility is only small as pre-trained corpora are very large.</p>
<p>The main differences between the three representations are three-folded. First, the dimensions of the vector representation can have either an explicit interpretation based on term frequencies (tf-idf) or an implicit interpretation (averaged word embeddings and document embeddings). Second, the length of the vector can be variable (tf-idf) or fixed (averaged word embeddings and document embeddings). Finally, the representations can be sparse with many zero dimensions (tf-idf) or dense with few zero dimensions (averaged word embeddings and document embeddings).</p>
<p>When comparing average word embeddings with document embeddings, document embeddings have a clear advantage, which is apparent from the clustering of the embeddings in vector space. Document embeddings tend to be located close to the embeddings of the keywords of the document (Lau and Baldwin, <xref ref-type="bibr" rid="B16">2016</xref>). Average word embeddings, on the other hand, tend to be located at the centroid of the word embeddings of all the words in a document. However, document embeddings are not free of issues. Ai et al. (<xref ref-type="bibr" rid="B1">2016</xref>) pointed out that shorter documents can be overfitted and often show too much similarity; the sampling distribution used in the document embeddings is improper in that frequent words can be penalized too rigidly; and sometimes document embeddings do not detect synonyms of words in different documents even though the context is alike. Despite these issues, document embeddings showed better results for various tasks when compared to tf-idf or averaged word embeddings (Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). Therefore, we expect that use document embeddings to represent essays and select pairs of essays based on these representations to outperform the tf-id and average word embeddings.</p>
</sec>
</sec>
<sec>
<title>2.2. Progressive Selection Rule Based on Vector Similarities</title>
<p>In this manuscript, we propose a progressive selection rule that combines the stochastic ACJ selection of Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) with a similarity component based on the cosine similarities of the vector representations of essays. Initially, the progressive selection rule selects pairs based on the similarity of their representations (i.e., how close they are to each other in vector space). As more judgment outcomes become available, the weight of the &#x02018;stochastic adaptivity&#x00027; component increases so that pairs are increasingly selected based on the quality parameter estimates of the works.</p>
<p>To quantify the similarity between the vector representations, the cosine similarity is chosen over the Euclidean distance and Jaccard similarity. First, unlike the Euclidean distance, the cosine similarity is a normalized measure (with range [&#x02212;1, 1]). Second, although also normalized, the Jaccard similarity tends to not work well for detecting similarities between texts when there are many overlapping words between essays (Singh and Singh, <xref ref-type="bibr" rid="B25">2021</xref>). The cosine similarity between two works <italic>i</italic> and <italic>j</italic> is the cosine of the angle of their corresponding vector representations <italic><bold>&#x003B3;</bold></italic><sub><italic>i</italic></sub> and <italic><bold>&#x003B3;</bold></italic><sub><italic>j</italic></sub>:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M38"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B3;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B3;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B3;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B3;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B3;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B3;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Note that <italic><bold>&#x003B3;</bold></italic><sub><italic>i</italic></sub> is of variable-length for tf-idf representations and fixed-length for averaged word embeddings and document embeddings. The fixed length is determined by the dimensionality of the pre-trained word embeddings which in this case is 320 (Tulkens et al., <xref ref-type="bibr" rid="B27">2016</xref>).</p>
<p>For the similarity component in the progressive selection rule, the cosine similarities of all works <italic>j</italic> with respect to work <italic>i</italic> are non-linearly transformed so that higher similarities are up-weighted and lower similarities are down-weighted. This can be achieved by assigning the probability mass of the CDF of a normal distribution to all cosine similarity values of works <italic>j</italic> with respect to work <italic>i</italic>. To encourage the selection of pairs with very high similarities (which can be rare) an upper quantile of the cosine similarities is chosen as the mean of the normal CDF. The quantile will function as a (soft) threshold parameter. So the probability to select works <italic>j</italic> with any lower similarity value than the quantile will be close to 0. A second component is the stochastic adaptive selection rule as in Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) (refer to above). As such, the parameter uncertainty of work <italic>i</italic> can be taken into account for the selection of work <italic>j</italic>.</p>
<p>The cosine similarity also measures dissimilarities (i.e., negative values). However, dissimilarities are uninformative for the pairing of essays, and negative values cannot be used as probabilities in the progressive selection rule. Hence, the cosine similarities are truncated at 0.</p>
<p>The two components are combined in the progressive selection rule as follows: a pair {<italic>i, j</italic>} is selected from <italic>B</italic><sub><italic>m</italic></sub> so that work <italic>i</italic> has the minimum number of comparisons out of all the works, and work <italic>j</italic> is sampled with a probability given by the weighted sum of the similarity and the adaptivity component. The weights depend on the number of times work <italic>i</italic> has been compared. Formally, at the <italic>m</italic><sub><italic>i</italic></sub> &#x0002B; 1-th comparison of work <italic>i</italic> it is paired with work <italic>j</italic> given the probability mass function:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M39"><mml:mrow><mml:mtext>Pr</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy='false'>&#x0007C;</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003A6;</mml:mi><mml:mo stretchy='true'>(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003B3;</mml:mi></mml:mstyle><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003B3;</mml:mi></mml:mstyle><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>S</mml:mi></mml:mstyle><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='true'>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>&#x0007D;</mml:mo><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mi>&#x003A6;</mml:mi><mml:mo stretchy='true'>(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003B3;</mml:mi></mml:mstyle><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003B3;</mml:mi></mml:mstyle><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>S</mml:mi></mml:mstyle><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='true'>)</mml:mo></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003D5;</mml:mi><mml:mo stretchy='true'>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003C3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mover accent='true'><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo stretchy='true'>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>&#x0007D;</mml:mo><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mi>&#x003D5;</mml:mi><mml:mo stretchy='true'>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>&#x003C3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mover accent='true'><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo stretchy='true'>)</mml:mo></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>where &#x003A6; is the CDF of a standard normal distribution with as mean the p-th quantile of all cosine similarities with the essay <italic>i</italic> except itself, <italic>Q</italic><sub><italic><bold>S</bold></italic><sub><italic>i</italic></sub></sub>(<italic>p</italic>) with <italic><bold>S</bold></italic><sub><italic>i</italic></sub> &#x0003D; (<italic>S</italic>(<italic><bold>&#x003B3;</bold></italic><sub><italic>i</italic></sub>, <italic><bold>&#x003B3;</bold></italic><sub><italic>j</italic></sub>), &#x02026;, <italic>S</italic>(<italic><bold>&#x003B3;</bold></italic><sub><italic>i</italic></sub>, <italic><bold>&#x003B3;</bold></italic><sub><italic>N</italic>&#x02212;1</sub>)). For the adaptive component, the density values of all <inline-formula><mml:math id="M40"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for the normal distribution with mean <inline-formula><mml:math id="M41"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and standard error <inline-formula><mml:math id="M42"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:msub></mml:math></inline-formula> are taken. The weight <italic>w</italic><sub><italic>i</italic></sub> &#x02208; [0, 1] of work <italic>i</italic> depends on <italic>m</italic><sub><italic>i</italic></sub> (this is the number of times work <italic>i</italic> has been compared) and on <italic>m</italic><sub><italic>d</italic></sub> (this is the minimal desired number of comparisons for each work) with <italic>m</italic><sub><italic>i</italic></sub> &#x02264; <italic>m</italic><sub><italic>d</italic></sub> and decay parameter <italic>t</italic> (<italic>t</italic> &#x0003E; 0) as follows:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M43"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#x000A0;</mml:mtext><mml:msub><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>m</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>t</mml:mi></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>otherwise</mml:mtext><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>If <italic>m</italic><sub><italic>i</italic></sub> &#x0003D; 0, work <italic>i</italic> is compared for the first time and will be allocated only based on the similarity component. Moreover, one needs to determine the speed at which the weight of the similarity selection rule decays in favor of the adaptive component by setting the parameter <italic>t</italic>. In computerized adaptive testing, where progressive selection rules with a random component have been proposed, <italic>t</italic> &#x0003D; 1 is often chosen, which corresponds with a linear decrease of the weight of the random component (Revuelta and Ponsoda, <xref ref-type="bibr" rid="B24">1998</xref>; Barrada et al., <xref ref-type="bibr" rid="B3">2010</xref>). In this study, however, we tune the decay parameter <italic>t</italic> to find the optimal progressive rule. A higher <italic>t</italic> leads to a slower decrease in the similarity component, whereas a smaller <italic>t</italic> leads to a faster decrease of the similarity component. For <italic>t</italic> &#x0003D; 0, the progressive rule reduces to the stochastic ACJ selection rule.</p>
</sec>
<sec>
<title>2.3. Experiment</title>
<sec>
<title>2.3.1. Datasets: Essay Sets</title>
<p>The proposed selection rule will be tested on two different essay sets. The essay sets along with quality scores were provided by the company Comproved. The qualities of these essays were estimated from CJ-assessments and are centered around zero. For this study, these are assumed to be the true quality levels, which is a reasonable assumption given that each essay was compared up to 20 times with random CJ. Both essay sets are of a similar size although the length of the essays in essay set 1 is more variable than those in essay set 2 (refer to <xref ref-type="table" rid="T1">Table 1</xref>). The quality levels show a symmetric distribution around zero. For essay set 1, 16-year-old students were asked to write a two-page research proposal on a topic of their choice. For essay set 2, 16-year-old students needed to write a two-page argumentative essay about the conservation of wildlife. For both essay sets, the true quality levels show only a small spread. This corresponds to assessment situations where it would be hard for the assessors to discriminate between the quality levels of essays (Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Description of the contents of two essay sets.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>Essay set 1</bold></th>
<th valign="top" align="center"><bold>Essay set 2</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Assignment</td>
<td valign="top" align="center">Research proposal</td>
<td valign="top" align="center">Argumentation</td>
</tr>
<tr>
<td valign="top" align="left"><italic>N</italic></td>
<td valign="top" align="center">141</td>
<td valign="top" align="center">150</td>
</tr>
<tr>
<td valign="top" align="left">SD of qualities</td>
<td valign="top" align="center">1.66</td>
<td valign="top" align="center">1.13</td>
</tr>
<tr>
<td valign="top" align="left">Range of qualities</td>
<td valign="top" align="center">&#x02212;5.42, 4.92</td>
<td valign="top" align="center">&#x02212;3.62, 2.10</td>
</tr>
<tr>
<td valign="top" align="left">Proportion qualities &#x02264; 0</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">0.45</td>
</tr>
<tr>
<td valign="top" align="left">Proportion qualities &#x0003E;0</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.55</td>
</tr>
<tr>
<td valign="top" align="left">Total &#x00023; of words</td>
<td valign="top" align="center">67340</td>
<td valign="top" align="center">58037</td>
</tr>
<tr>
<td valign="top" align="left">Avg. length essays</td>
<td valign="top" align="center">474</td>
<td valign="top" align="center">386</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>2.3.2. Preprocessing of Essay Texts</title>
<p>The initial preprocessing steps on the essay texts are common for every representation technique and they are in accordance with the steps performed on the pre-trained SoNaR corpus (Oostdijk et al., <xref ref-type="bibr" rid="B20">2013</xref>; Tulkens et al., <xref ref-type="bibr" rid="B27">2016</xref>). This involves lowercasing, removing punctuations, removing numbers, removing single letter words, and decoding utf-8 encoding. The only single letter word that is included is &#x0201C;u&#x0201D; which is a Dutch formal pronoun. In contrast to Tulkens et al. (<xref ref-type="bibr" rid="B27">2016</xref>), we chose to also include sentences shorter than 5 words. The reason being that the essay set is short (1 or 2 pages) so every sentence may be meaningful (<xref ref-type="table" rid="T1">Table 1</xref>).</p>
<p>Some additional preprocessing steps on the texts depend on the representation technique. For the tf-idf representation of the essays, the essay texts will be normalized to a higher extent. This is necessary as the size of the essay sets is relatively small and no pre-trained corpus can be used with tf-idf. Extended normalization will decrease the length of the vocabulary, and hence, increase the similarities between essays. However, there may be a loss of information as well. A first additional step is the lemmatization of the words so that they are simplified to their root word, which is an existing word&#x02014;unlike with stemming. In addition, for tf-idf the syntactical structures will be represented to some extent by allowing bi-grams of word pairs that often occur together. Including <italic>n</italic>-grams also decreases the high dimensionality of the vector representation because the vocabulary size decreases. Note that for the tf-idf representations, the idf-term is smoothed in order to prevent zero division (Aizawa, <xref ref-type="bibr" rid="B2">2003</xref>).</p>
<p>For the representation of essays based on averaged word embeddings and document embeddings, the pre-trained SoNaR corpus with embeddings of Dutch words is used (Tulkens et al., <xref ref-type="bibr" rid="B27">2016</xref>). The pre-trained corpus consists of 28.1 million sentences and 398.2 million words from various media outlets (news stories, magazines, auto-cues, legal texts, Wikipedia, etc.) (Oostdijk et al., <xref ref-type="bibr" rid="B20">2013</xref>). The embeddings were learned using a skip-gram architecture with negative sampling (Mikolov et al., <xref ref-type="bibr" rid="B19">2013</xref>). The embeddings have 320 dimensions. The pre-trained SoNaR corpus showed excellent results for training word embeddings in Tulkens et al. (<xref ref-type="bibr" rid="B27">2016</xref>). Note that this pre-trained corpus only contains correctly spelled words. This implies that misspelled words in the essays will not be represented, which may decrease their usability for making pairs. Also, grammatical mistakes can have an influence on the essay embeddings because word embeddings and document embeddings are sensitive to word order as it used for their training (Mikolov et al., <xref ref-type="bibr" rid="B19">2013</xref>; Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). Preprocessing techniques like lemmatization or stemming are not performed for these representations to keep the essays closest to their original semantical and syntactical meaning. This is feasible given that almost all words can be found in the large pre-trained SoNaR corpus (Oostdijk et al., <xref ref-type="bibr" rid="B20">2013</xref>).</p>
</sec>
<sec>
<title>2.3.3. Baseline Selection Rules and Simulation Design</title>
<p>Three baseline selection rules will be tested: the random CJ, the stochastic ACJ as in Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>), and a progressive selection rule with a random component for the initial comparisons. For the progressive rule with a similarity component (Equation 10), three essay representation techniques will be considered (i.e., tf-idf, averaged word embeddings, and document embeddings) and a progressive rule with a random component instead of a similarity component. The progressive selection rule with a random component is constructed to evaluate whether the similarity component in the progressive selection rule is more informative for the initial pairing of works than random pairs.</p>
<p>The performance of each selection rule will be assessed based on the SSR, the true reliability, and the SSR bias (their difference) for a given number of comparisons per work on average. Next, differences in SSR between the proposed progressive rule and the baseline selection rules will be evaluated based on the two components that determine the SSR, namely the spread of the quality parameter estimates and their standard errors with respect to the ranking of essays (Equation 7). For brevity, not all representation techniques will be compared to the baseline selection rules here, only the one that performs the best in terms of SSR.</p>
<p>To simulate the judgment process the probability that work <italic>i</italic> wins as obtained from BTL-model (Equation 1) is compared to a random number drawn from a continuous uniform distribution between 0 and 1 (Davey et al., <xref ref-type="bibr" rid="B11">1997</xref>; Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>). If the probability is higher than the random value, work <italic>i</italic> wins the comparison. If the sampled value is smaller, work <italic>j</italic> wins the comparison. As such, one can imitate the stochastic process of judging. For each of the selection rules, the judgment process will be simulated 100 times (Matteucci and Veldkamp, <xref ref-type="bibr" rid="B18">2013</xref>; Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>). A minimum of 40 work comparisons for all works is defined as a stopping rule (<italic>m</italic><sub><italic>d</italic></sub> &#x0003D; 40). This can show the asymptotic behavior of the SSR estimator for the different selection rules. For a minimum of 40 work comparisons per work, at least 50% of the possible pairings are compared given that <italic>N</italic><sub>1</sub> &#x0003D; 141 and <italic>N</italic><sub>2</sub> &#x0003D; 150. Preliminary simulations are conducted to tune the decay parameter <italic>t</italic> and the quantile <italic>p</italic> of the cosine similarities (Equation 10). That is, the true reliability and the SSR bias are evaluated for a grid of every parameter combinations for <italic>p</italic> &#x0003D; {0.50, 0.70, 0.80, 0.90, 0.95} and <italic>t</italic> &#x0003D; {0.20, 0.40, 0.60, 0.80, 1.00, 2.00}. For each condition (5 &#x000D7; 5) 50 simulations are conducted.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Tuning of the Decay Parameter and the Cosine Similarity Quantile</title>
<p>The preliminary simulations showed that a decay parameter (<italic>t</italic>) of 0.4 and a cosine similarity quantile (<italic>p</italic>) from 70 to 90% result in the highest SSR with a small bias (below 0.05). The 80% upper quantile of the cosine similarities was chosen. The cosine similarity corresponding to the 80% quantile is the smallest for tf-idf (0.15 and 0.24 for essay set 1 and 2, respectively) and the largest for averaged word embeddings (0.34 and 0.42 for essay set 1 and 2, respectively) (<xref ref-type="table" rid="T2">Table 2</xref>). The 80% quantile of the cosine similarities using document embeddings is 0.27 and 0.26 for essay set 1 and 2, respectively.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Quantiles of the cosine similarities between essays using different essay representation techniques for two essay sets.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Essay representation</bold></th>
<th valign="top" align="center"><bold>Quantile (%)</bold></th>
<th valign="top" align="center"><bold>Essay set 1</bold></th>
<th valign="top" align="center"><bold>Essay set 2</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Tf-idf</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">0.12</td>
<td valign="top" align="center">0.21</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">70</td>
<td valign="top" align="center">0.14</td>
<td valign="top" align="center">0.23</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">80</td>
<td valign="top" align="center">0.15</td>
<td valign="top" align="center">0.24</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">90</td>
<td valign="top" align="center">0.17</td>
<td valign="top" align="center">0.26</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Averaged word emb.</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">0.3</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">70</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.37</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">80</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.42</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">90</td>
<td valign="top" align="center">0.40</td>
<td valign="top" align="center">0.51</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Document emb.</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">0.22</td>
<td valign="top" align="center">0.22</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">70</td>
<td valign="top" align="center">0.24</td>
<td valign="top" align="center">0.24</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">80</td>
<td valign="top" align="center">0.27</td>
<td valign="top" align="center">0.26</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">90</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.28</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.2. Performance of SSR Estimator</title>
<p>We will first describe the performance of the SSR estimator for the proposed progressive selection rule with different essay representation techniques. Subsequently, we will compare the progressive selection rule with the best performing representation technique to the CJ and ACJ baseline selection rules.</p>
<sec>
<title>3.2.1. Performance of SSR for the Progressive Selection Rules</title>
<p>The performance of the SSR for the progressive rule with a similarity component is highly dependent on the chosen essay representation technique. For essay set 1, a similarity component based on averaged word embeddings and document embeddings seems to perform equally well in terms of reliability and SSR bias (<xref ref-type="fig" rid="F1">Figures 1A</xref>, <xref ref-type="fig" rid="F2">2A</xref>). The progressive selection rule with a similarity component based on tf-idf representations results in small true reliability similar to the progressive selection rule with a random component. This indicates that the similarities based on the tf-idf representations of essay set 1 are close to being random. However, for essay set 2 the progressive selection rule with a similarity component based on tf-idf performs better than with a random component, and unexpectedly, better than with a similarity component based on averaged word embeddings (<xref ref-type="fig" rid="F1">Figures 1B</xref>, <xref ref-type="fig" rid="F2">2B</xref>). For both essay sets, the progressive rule with a similarity component based on document embeddings performs at least as good as the progressive rule based on tf-idf or averaged word embeddings, and is always better than the progressive rule with a random component. This indicates that initial pairings based on the large cosine similarities of document embeddings can be beneficial.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>For essay set 1 (<bold>A</bold>, <italic>N</italic><sub>1</sub> &#x0003D; 141) and 2 (<bold>B</bold>, <italic>N</italic><sub>2</sub> &#x0003D; 150), the true reliability resulting from the progressive selection with a random component, and a similarity component using tf-idf, averaged word embeddings and document embeddings (100 simulations). The solid lines indicate the mean values and the transparent bands indicate the 95% point-wise confidence intervals.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>For essay set 1 (<bold>A</bold>, <italic>N</italic><sub>1</sub> &#x0003D; 141) and 2 (<bold>B</bold>, <italic>N</italic><sub>2</sub> &#x0003D; 150), SSR bias resulting from the progressive selection with a random component, and a similarity component using tf-idf, averaged word embeddings and document embeddings (100 simulations). The solid lines indicate the mean values and the transparent bands indicate the 95% point-wise confidence intervals.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0002.tif"/>
</fig>
<p>The progressive rule with a similarity component produces higher true reliability than random CJ (<xref ref-type="fig" rid="F1">Figure 1</xref>). When the similarity component is computed based on document embeddings, true reliability is reached that is 0.02&#x02013;0.03 higher than for random CJ. The true reliability under the progressive rule with a similarity component is close to the high reliability under ACJ. Compared to ACJ, however, the progressive rule with a similarity component has an SSR bias that converges faster to below 0.05. For essay set 1, the SSR bias is even smaller than for random CJ (<xref ref-type="fig" rid="F2">Figure 2A</xref>). For essay set 2, the SSR bias is more persistent than for random CJ which may be due to the smaller spread of its true quality levels (<xref ref-type="fig" rid="F2">Figure 2B</xref> and <xref ref-type="table" rid="T1">Table 1</xref>).</p>
</sec>
<sec>
<title>3.2.2. Performance of SSR for the Baseline Selection Rules</title>
<p>The performance of the CJ and ACJ baseline selection rules in terms of the SSR is as expected given the average number of comparisons. The random CJ can result in an SSR that can both under- and over-estimate the true reliability at the start of the CJ process (<xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref>). Crompvoets et al. (<xref ref-type="bibr" rid="B10">2021</xref>) also reported positive SSR bias for the random CJ selection rule. The SSR bias for random CJ converges to &#x0003C;0.05 after on average 5 comparisons per work. In other words, up to 355 and 375 comparisons were needed for essay set 1 and 2, respectively. ACJ on the other hand results in an SSR that clearly overestimates the true reliability. After on average 10 comparisons per work, the SSR is 25% larger than the true reliability for essay set 1 (<xref ref-type="fig" rid="F3">Figure 3A</xref>), and 52% for essay set 2 (<xref ref-type="fig" rid="F3">Figure 3B</xref>). For ACJ, the SSR bias is only negligible (below 0.05) after on average 20 comparisons per work for both essay sets (<xref ref-type="fig" rid="F4">Figure 4</xref>). Both baseline selection rules show evidence that their SSR is asymptotically unbiased&#x02014;although the rate at which the bias reduces is the highest for random CJ. Note that for all selection rules, the SSR bias is negative until on average 5 comparisons per work are made. Even though ACJ produces inflated SSR estimates, it can produce true reliability that is 0.02&#x02013;0.03 higher than for random CJ (<xref ref-type="fig" rid="F3">Figure 3</xref>). This is already observed for more than 5 comparisons per work on average. The performance of the SSR for random CJ and ACJ is similar to in Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) and Rangel-Smith and Lynch (<xref ref-type="bibr" rid="B23">2018</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>For essay set 1 (<bold>A</bold>, <italic>N</italic><sub>1</sub> &#x0003D; 141) and 2 (<bold>B</bold>, <italic>N</italic><sub>2</sub> &#x0003D; 150), the true reliability resulting from the selection rules: for random (CJ), adaptive (ACJ), and a progressive selection with a random and a similarity component using document embeddings of essays (100 simulations). The solid lines indicate the mean values and the transparent bands indicate the 95% point-wise confidence intervals.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>For essay set 1 (<bold>A</bold>, <italic>N</italic><sub>1</sub> &#x0003D; 141) and 2 (<bold>B</bold>, <italic>N</italic><sub>2</sub> &#x0003D; 150), SSR bias resulting from the selection rules: random (CJ), adaptive (ACJ) and a progressive selection with a random and a similarity component using document embeddings (100 simulations). The solid lines indicate the mean values and the transparent bands indicate the 95% point-wise confidence intervals.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0004.tif"/>
</fig>
<p>The results for the true reliability and the SSR produced by the progressive rule with a random component are inconsistent between essay sets. For essay set 1, the progressive rule with a random component results in quality parameter estimates that have the lowest true reliability out of all the selection rules (<xref ref-type="fig" rid="F3">Figure 3A</xref>). For essay set 2, the progressive rule with a random component results in true reliability that is higher than for the random CJ and ACJ (<xref ref-type="fig" rid="F3">Figure 3B</xref>). For both essay sets, the SSR bias for the progressive rule with a random component is smaller than for ACJ but larger than for random CJ (<xref ref-type="fig" rid="F4">Figure 4</xref>).</p>
<p>The progressive selection rule with a similarity component based on document embeddings requires fewer judgments per work to reach the desired reliability (for instance, 0.70 or 0.80). For essay set 1, this progressive selection rule can reach reliability of 0.80 in 14 comparisons per work, while 16 comparisons on average are required for random CJ (<xref ref-type="fig" rid="F3">Figures 3A</xref>, <xref ref-type="fig" rid="F4">4A</xref>). In total, with the proposed selection rule 141 fewer comparisons are needed to reach true reliability of 0.80. For essay set 2, with the proposed selection rule on average 3 comparisons per work less are required as compared to random CJ (<xref ref-type="fig" rid="F3">Figures 3B</xref>, <xref ref-type="fig" rid="F4">4B</xref>). Then, 225 fewer comparisons are needed. Note that the gain in true reliability of the novel progressive selection rule is only moderate with respect to random CJ (0.02-0.03). This can be explained by the relatively large essay sets and the small standard deviations of the true quality levels (<xref ref-type="table" rid="T1">Table 1</xref>; Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>; Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>).</p>
</sec>
</sec>
<sec>
<title>3.3. Evaluation of the Quality Parameter Estimates</title>
<p>To investigate the performance of the SSR estimator we focus on the spread of the quality estimates on the scale and their precision (i.e., uncertainty) (Equation 7). Only the results obtained using the document embeddings as the text representation technique are considered because the SSR results (refer to above) were best for both essay sets. Again random CJ, ACJ, and the progressive selection rule with a random component serve as baselines for comparison.</p>
<sec>
<title>3.3.1. Spread of the Quality Estimates</title>
<p>Because the absolute differences in quality estimates can vary, the cumulative ranking of the estimates is evaluated, for different average numbers of comparisons.</p>
<p>For five comparisons per work on average, all selection rules result in equivalent estimated quality parameters given their ranking (<xref ref-type="fig" rid="F5">Figures 5A</xref>, <xref ref-type="fig" rid="F6">6A</xref>). For 10 comparisons per work on average, the differences in estimated quality parameters between ACJ and the other selection rules become noticeable (<xref ref-type="fig" rid="F5">Figures 5B</xref>, <xref ref-type="fig" rid="F6">6B</xref>). ACJ tends to produce quality estimates that are more spread out than the other selection rules. For ACJ &#x0007E;20% of the highest and lowest ranking works will have estimated qualities greater than &#x000B1;3. For the other selection rules, this is only the case for 5% of the most extreme quality parameter values. The inflated spread of the quality parameter estimates can explain the inflation of the SSR for ACJ (Equation 7). The higher the inflation of the spread of the quality parameter estimates, the more biased the estimates can be. Moreover, when comparing the results of set 1 (<xref ref-type="fig" rid="F5">Figure 5</xref>) with the results of set 2 (<xref ref-type="fig" rid="F6">Figure 6</xref>), there seems to be an inverse relation between the spread of true quality levels (<xref ref-type="table" rid="T1">Table 1</xref>) and the spread of the estimated quality parameters for ACJ. Namely, the smaller the spread of the true quality levels, the larger the inflation of the spread of the quality parameter estimates, and therefore, the larger the SSR bias for ACJ will be.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>For essay set 1 (N1 = 141), quality parameter estimates with respect to their cumulative ranking for 5 <bold>(A)</bold> and 10 <bold>(B)</bold> Comparisons per work on average. This is assessed for different selection rules: random (CJ), adaptive (ACJ), and a progressive selection with a random component and with a similarity component using document embeddings (100 simulations).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>For essay set 2 (<italic>N</italic><sub>2</sub> &#x0003D; 150), quality parameter estimates with respect to their cumulative ranking for 5 <bold>(A)</bold> and 10 <bold>(B)</bold> comparisons per work on average. This is assessed for different selection rules: random (CJ), adaptive (ACJ), and a progressive selection with a random component and with a similarity component using document embeddings (100 simulations).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0006.tif"/>
</fig>
</sec>
<sec>
<title>3.3.2. The Precision of the Quality Estimates</title>
<p>As the spread of the quality estimates differs between selection rules (<xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>), the parameter uncertainty is assessed with respect to the cumulative rank order. It can be seen that quality parameters are estimated most precisely for middle-ranked essays (<xref ref-type="fig" rid="F7">Figures 7</xref>, <xref ref-type="fig" rid="F8">8</xref>). This can be explained by the fact that most essay parameters are located around the median. On the other hand, the highest and lowest ranking essay qualities are estimated with less precision. The precision difference between extreme and middle-ranked essays reduces as the average number of comparisons per work increases. This decrease is stronger for essay set 1 (<xref ref-type="fig" rid="F7">Figure 7</xref>), which has a larger SD of the true qualities than essay set 2 (<xref ref-type="table" rid="T1">Table 1</xref>). However, for both essay sets ACJ results in more precise quality parameter estimates for 10 or more comparisons per work on average (<xref ref-type="fig" rid="F7">Figures 7B</xref>, <xref ref-type="fig" rid="F8">8B</xref>). The smaller standard errors for ACJ can inflate the SSR (Equation 7). Note that the increase in precision in ACJ is in itself a desired property; it is its high bias in quality parameter estimates that is undesirable. As the average number of comparisons per work increases, the parameter uncertainty becomes similar for all selection rules (<xref ref-type="fig" rid="F7">Figures 7</xref>, <xref ref-type="fig" rid="F8">8</xref>). But even then, random CJ results in more uncertain parameter estimates than ACJ.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>For essay set 1 (<italic>N</italic><sub>1</sub> &#x0003D; 141), the precision of quality parameter estimates with respect to their cumulative ranking for different averages of work comparisons: 5 <bold>(A)</bold>, 10 <bold>(B)</bold>, 20 <bold>(C)</bold>, and 30 <bold>(D)</bold>, respectively. This is assessed for different selection rules: random (CJ), adaptive (ACJ), and a progressive selection with a random component and with a similarity component using document embeddings (100 simulations). The solid lines indicate the mean values and the transparent bands indicate the 95% point-wise CIs.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0007.tif"/>
</fig>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>For essay set 2 (<italic>N</italic><sub>2</sub> &#x0003D; 150), standard error of quality parameter estimates with respect to their cumulative ranking for different averages of work comparisons: 5 <bold>(A)</bold>, 10 <bold>(B)</bold>, 20 <bold>(C)</bold>, and 30 <bold>(D)</bold>, respectively. This is assessed for different selection rules: random (CJ), adaptive (ACJ), and a progressive selection with a random component and with a similarity component using document embeddings (100 simulations). The solid lines indicate the mean values and the transparent bands indicate the 95% point-wise CIs.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-07-854378-g0008.tif"/>
</fig>
<p>The progressive selection rule with a similarity component based on document embeddings can show improvements upon random CJ in terms of the precision of the quality parameter estimates. Namely, for essay set 1 a lower uncertainty for high and lower ranked works is obtained after 10 comparisons on average (<xref ref-type="fig" rid="F7">Figure 7B</xref>). With respect to the progressive rule with a random component, there is a visible gain in precision for the estimation of quality parameters. For essay set 2, the differences in uncertainty are small (<xref ref-type="fig" rid="F8">Figure 8</xref>). This may be explained by the smaller spread of the true quality levels of essay set 2 (<xref ref-type="table" rid="T1">Table 1</xref>).</p>
<p>In sum, the new progressive rule with a similarity component (based on document embeddings), unlike ACJ, does not show inflation of the spread of the quality estimates (<xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>). This is also observed for the progressive rule with a random component. However, the progressive rule with a similarity component can result in more precise quality parameter estimates than with a random component (<xref ref-type="fig" rid="F7">Figures 7</xref>, <xref ref-type="fig" rid="F8">8</xref>). This is most notably the case for essay set 1 where the spread of the true quality levels is larger (<xref ref-type="table" rid="T1">Table 1</xref>). For true quality levels that are more spread out a high, unbiased SSR can be obtained with the progressive selection rule based on document embeddings (<xref ref-type="fig" rid="F4">Figures 4A</xref>, <xref ref-type="fig" rid="F3">3A</xref>) without inflating the spread of the scale of quality parameter estimates (<xref ref-type="fig" rid="F5">Figure 5</xref>) and while increasing the precision of the quality parameter estimates (<xref ref-type="fig" rid="F7">Figure 7</xref>).</p>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>With the proposed selection rule, the essays were initially paired based on the cosine similarities of their vector representations. After the initial phase, the ACJ selection criterion progressively weighted higher in the selection rule (Equation 10). Even though the gain in SSR and true reliability was small, an improvement in terms of SSR estimates and its bias were observed when compared to CJ, ACJ, and a progressive selection rule with a random component. Hence, the proposed selection rule reduced the number of comparisons needed to obtain reliable quality estimates for the essays. The progressive selection rule with a similarity component based on document embeddings performed consistently better than any other selection rule for the two different essay sets. Most importantly, this progressive rule with a similarity component resulted in higher true reliability than the progressive rule with a random component while still reducing the SSR bias quickly. Thus, there is not only evidence that one can alleviate the cold-start by using a progressive selection rule based on the cosine similarities, but also that one can improve the true reliability and the SSR with this selection rule. However, the results indicate the importance of selecting the most appropriate essay representation technique, which was found to be the document embeddings (Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). The document embeddings were initialized by a pre-trained corpus of word embeddings. A limitation of the simulation design is that in practice multiple raters can compare the same pair while in our design the restriction of one comparison per pair was held. We do not except that by elevating this restriction the results of the proposed progressive selection rule relative to the baseline selection rules would be very different.</p>
<p>Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) selected essays to be judged that have parameter estimates with the largest standard errors. It was observed that when selecting essays to be judged (work <italic>i</italic>) that way, a large discrepancy occurs in the number of comparisons per essay. Essays with extremer parameter estimates would consistently be selected as the essay qualities are almost always more uncertain. Instead in this study, it was opted to select the essay to be judged based on the minimal number of times it has been judged. Note that the number of comparisons is also related to the standard errors of the parameter estimates: the standard errors decrease with the number of comparisons (Equation 6). Our approach reflects more practical assessment situations where having an equal amount of comparisons for all works may be preferred. It can be seen as unfair by assessors and students if one essay would be compared more often than another. From a statistical point of view, however, targeting the essays to be judged based on the maximal uncertainty of the parameter estimates may increase the precision of the quality estimates and the SSR even further. Therefore, the selection rule proposed in this study may be improved upon by selecting every essay to be judged (work <italic>i</italic>) based on the maximal standard error of its parameter estimate. Future research is required with respect to the effects of selecting the essays to be judged based on a combination of the number of times it has been compared and their parameter uncertainty. By doing so, one can prevent too large discrepancies in the number of comparisons per essay while still improving the SSR.</p>
<p>It is expected that for smaller essay sets, the benefits of the progressive selection rule with a similarity component over random CJ will become more apparent. Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>) and Bramley and Vitello (<xref ref-type="bibr" rid="B6">2019</xref>) observed that for smaller samples, ACJ can result in a higher gain in the precision of quality parameters and the reliability than random CJ. Furthermore, ACJ can perform well when there is more spread in the true quality levels of works (&#x003C3; &#x0003E; 2) (Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>). The current results showed that the novel selection rule can produce high true reliability without an increase in SSR bias. Given these results, it is expected that with the proposed selection rule a higher SSR with a small bias can be obtained when it is tested on smaller sample sizes than in the current study. Such cases would represent small classroom assessment situations. Note that document embeddings can be used for smaller essay sets as they can be initiated by a pre-trained corpus of word embeddings (Oostdijk et al., <xref ref-type="bibr" rid="B20">2013</xref>). It is also expected that the benefits would be greater for essay sets that show more high similarities or similarities with more variance. Then more informative initial pairs could be selected. For this study, the essay representations showed rather low similarities (refer to <xref ref-type="table" rid="T2">Table 2</xref>).</p>
<p>As opposed to alleviating the cold-start of ACJ, one can also improve the ACJ-algorithm itself. The proposed progressive selection rules implement the stochastic approach of ACJ from Crompvoets et al. (<xref ref-type="bibr" rid="B9">2020</xref>). For an essay to be paired with another, an essay will be selected based on its density value for the distribution of the essay quality estimate that is to be compared (work <italic>i</italic>). That way, the uncertainty of the quality estimate of the essay that is compared is taken into account. However, this assumes that all other essay quality estimates (every work <italic>j</italic>) are deterministic. In order to take the uncertainty of all essay quality estimates into account, a different approach of adaptive pairing is required. A Bayesian adaptive selection rule as proposed in Crompvoets et al. (<xref ref-type="bibr" rid="B10">2021</xref>) takes the parameter uncertainty of both work <italic>i</italic> and <italic>j</italic> into account. Every work <italic>i</italic> and <italic>j</italic> are sampled from the conditional posterior distribution of their quality parameter. In the context of item response theory, Barrada et al. (<xref ref-type="bibr" rid="B3">2010</xref>) have summarized multiple selection rules that integrate over the weighted likelihood function of an ability parameter: e.g., the Fisher information weighted by the likelihood function or the Kullback-Leibler function weighted by the likelihood function. It is expected that the progressive selection rule with a similarity component would benefit from such a redefined ACJ selection rule.</p>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>The objective of this study was to alleviate the cold-start problem of adaptive comparative judgments, while simultaneously minimizing the bias of the scale separation coefficient that can occur (Bramley, <xref ref-type="bibr" rid="B5">2015</xref>; Rangel-Smith and Lynch, <xref ref-type="bibr" rid="B23">2018</xref>; Bramley and Vitello, <xref ref-type="bibr" rid="B6">2019</xref>; Crompvoets et al., <xref ref-type="bibr" rid="B9">2020</xref>). We proposed the use of text mining as it is possible to extract essay representations before the judgment process has started. A variety of essay representation techniques were considered: term frequency-inverse document frequency, averaged word embeddings, and document embeddings (Aizawa, <xref ref-type="bibr" rid="B2">2003</xref>; Mikolov et al., <xref ref-type="bibr" rid="B19">2013</xref>; Le and Mikolov, <xref ref-type="bibr" rid="B17">2014</xref>). Subsequently, the representations of essays were used to select initial pairs of essays that have high cosine similarities between their representations. Progressively, the selection rule will be more determined by the closeness of the quality estimates given the parameter uncertainty. The simulation results showed that the progressive selection rule can minimize the bias of the scale separation coefficient while still resulting in high true reliability. Out of all representation techniques, the document embeddings of the essays (as initialized by pre-trained word embeddings) consistently showed the best results in terms of scale separation reliability. Moreover, the proposed progressive rule prevents the inflation of the variability of the quality estimates, and it can reduce the uncertainty of the quality estimates&#x02014;especially for low and high quality essays when the variability of the true quality levels is high. Although the gain in reliability and parameter precision was moderate, it is expected that this gain will be larger for smaller essay sets that show more variability in the true essay qualities and for essays that show more high similarities. A practical example would be its use in classroom assessment contexts.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The data analyzed in this study is subject to the following licenses/restrictions: the data was previously used for commercial purposes. Requests to access these datasets should be directed to <email>info&#x00040;comproved.com</email>.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>MD and DD: conceptualization and presentation of the problem and design of the simulation study. MD: execution and analysis of the simulation and writing&#x02014;original draft preparation. MD, DD, and WV: writing&#x02014;review and editing. DD and WV: supervision. This article originated from the Master thesis MD wrote under the supervision of WV and DD (De Vrindt, <xref ref-type="bibr" rid="B12">2021</xref>). All the authors approved the final version of the manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack><p>The company Comproved was thanked for allowing this study by providing essay texts and scores.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ai</surname> <given-names>Q.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Guo</surname> <given-names>J.</given-names></name> <name><surname>Croft</surname> <given-names>W. B.</given-names></name></person-group> (<year>2016</year>). <source>Analysis of the Paragraph Vector Model for Information Retrieval</source>. <publisher-loc>New York, NY</publisher-loc>. <pub-id pub-id-type="doi">10.1145/2970398.2970409</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aizawa</surname> <given-names>A.</given-names></name></person-group> (<year>2003</year>). <article-title>An information-theoretic perspective of tf&#x02013;idf measures</article-title>. <source>Inform. Process. Manage</source>. <volume>39</volume>, <fpage>45</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1016/S0306-4573(02)00021-3</pub-id><pub-id pub-id-type="pmid">18350930</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barrada</surname> <given-names>J. R.</given-names></name> <name><surname>Olea</surname> <given-names>J.</given-names></name> <name><surname>Ponsoda</surname> <given-names>V.</given-names></name> <name><surname>Abad</surname> <given-names>F. J.</given-names></name></person-group> (<year>2010</year>). <article-title>A method for the comparison of item selection rules in computerized adaptive testing</article-title>. <source>Appl. Psychol. Measure</source>. <volume>34</volume>, <fpage>438</fpage>&#x02013;<lpage>452</lpage>. <pub-id pub-id-type="doi">10.1177/0146621610370152</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bradley</surname> <given-names>R. A.</given-names></name> <name><surname>Terry</surname> <given-names>M. E.</given-names></name></person-group> (<year>1952</year>). <article-title>Rank analysis of incomplete block designs: I. the method of paired comparisons</article-title>. <source>Biometrika</source> <volume>39</volume>, <fpage>324</fpage>&#x02013;<lpage>345</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/39.3-4.324</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bramley</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <source>Investigating the Reliability of Adaptive Comparative Judgment</source>. Tech. rep., <publisher-name>Cambridge Assessment</publisher-name>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bramley</surname> <given-names>T.</given-names></name> <name><surname>Vitello</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>The effect of adaptivity on the reliability coefficient in adaptive comparative judgement</article-title>. <source>Assess. Educ</source>. <volume>26</volume>, <fpage>43</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1080/0969594X.2017.1418734</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brennan</surname> <given-names>R. L.</given-names></name></person-group> (<year>2010</year>). <article-title>Generalizability theory and classical test theory</article-title>. <source>Appl. Measure. Educ</source>. <volume>24</volume>, <fpage>1</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1080/08957347.2011.532417</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coenen</surname> <given-names>T.</given-names></name> <name><surname>Coertjens</surname> <given-names>L.</given-names></name> <name><surname>Vlerick</surname> <given-names>P.</given-names></name> <name><surname>Lesterhuis</surname> <given-names>M.</given-names></name> <name><surname>Mortier</surname> <given-names>A. V.</given-names></name> <name><surname>Donche</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>An information system design theory for the comparative judgement of competences</article-title>. <source>Eur. J. Inform. Syst</source>. <volume>27</volume>, <fpage>248</fpage>&#x02013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1080/0960085X.2018.1445461</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crompvoets</surname> <given-names>E. A.</given-names></name> <name><surname>B&#x000E9;guin</surname> <given-names>A. A.</given-names></name> <name><surname>Sijtsma</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Adaptive pairwise comparison for educational measurement</article-title>. <source>J. Educ. Behav. Stat</source>. <volume>45</volume>, <fpage>316</fpage>&#x02013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.3102/1076998619890589</pub-id><pub-id pub-id-type="pmid">27363500</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crompvoets</surname> <given-names>E. A. V.</given-names></name> <name><surname>Beguin</surname> <given-names>A.</given-names></name> <name><surname>Sijtsma</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <source>Pairwise Comparison Using a Bayesian Selection Algorithm: Efficient Holistic Measurement</source>. <pub-id pub-id-type="doi">10.31234/osf.io/32nhp</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Davey</surname> <given-names>T.</given-names></name> <name><surname>Nering</surname> <given-names>M. L.</given-names></name> <name><surname>Thompson</surname> <given-names>T.</given-names></name></person-group> (<year>1997</year>). <source>Realistic Simulation of Item Response Data, Vol. 97</source>. <publisher-loc>Iowa City, IA</publisher-loc>: <publisher-name>ERIC</publisher-name>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>De Vrindt</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <source>Text mining to alleviate the cold-start problem</source> (Master&#x00027;s thesis). <publisher-loc>KU Leuven, Leuven, Belgium</publisher-loc>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hunter</surname> <given-names>D. R.</given-names></name></person-group> (<year>2004</year>). <article-title>MM Algorithms for generalized Bradley Terry models</article-title>. <source>Ann. Stat</source>. <volume>32</volume>, <fpage>384</fpage>&#x02013;<lpage>406</lpage>. <pub-id pub-id-type="doi">10.1214/aos/1079120141</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jones</surname> <given-names>I.</given-names></name> <name><surname>Bisson</surname> <given-names>M.</given-names></name> <name><surname>Gilmore</surname> <given-names>C.</given-names></name> <name><surname>Inglis</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Measuring conceptual understanding in randomised controlled trials: can comparative judgement help?</article-title> <source>Br. Educ. Res. J</source>. <volume>45</volume>, <fpage>662</fpage>&#x02013;<lpage>680</lpage>. <pub-id pub-id-type="doi">10.1002/berj.3519</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jones</surname> <given-names>I.</given-names></name> <name><surname>Inglis</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>The problem of assessing problem solving: can comparative judgement help?</article-title> <source>Educ. Stud. Math</source>. <volume>89</volume>, <fpage>337</fpage>&#x02013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1007/s10649-015-9607-1</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lau</surname> <given-names>J. H.</given-names></name> <name><surname>Baldwin</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>An empirical evaluation of doc2vec with practical insights into document embedding generation</article-title>, in <source>Proceedings of the 1st Workshop on Representation Learning for NLP</source> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>78</fpage>&#x02013;<lpage>86</lpage>. <pub-id pub-id-type="doi">10.18653/v1/W16-1609</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>Q. V.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Distributed representations of sentences and documents</article-title>. <source>arXiv [Preprint]</source>. <volume>arXiv</volume>: <fpage>1405.4053</fpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1405.4053</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matteucci</surname> <given-names>M.</given-names></name> <name><surname>Veldkamp</surname> <given-names>B. P.</given-names></name></person-group> (<year>2013</year>). <article-title>On the use of MCMC computerized adaptive testing with empirical prior information to improve efficiency</article-title>. <source>Stat. Methods Appl</source>. <volume>22</volume>, <fpage>243</fpage>&#x02013;<lpage>267</lpage>. <pub-id pub-id-type="doi">10.1007/s10260-012-0216-1</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Efficient estimation of word representations in vector space</article-title>. <source>arXiv [Preprint]</source>. <volume>arXiv</volume>: <fpage>1301.3781</fpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1301.3781</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oostdijk</surname> <given-names>N.</given-names></name> <name><surname>Reynaert</surname> <given-names>M.</given-names></name> <name><surname>Hoste</surname> <given-names>V.</given-names></name> <name><surname>Schuurman</surname> <given-names>I.</given-names></name></person-group> (<year>2013</year>). <article-title>The construction of a 500-million-word reference corpus of contemporary written Dutch</article-title>, in <source>Essential Speech and Language Technology for Dutch</source>, eds <person-group person-group-type="editor"><name><surname>Spijns</surname> <given-names>P.</given-names></name> <name><surname>Odijk</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>219</fpage>&#x02013;<lpage>247</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-30910-6_13</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pollitt</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Let&#x00027;s stop marking exams</article-title>, in <source>IAEA Conference</source> (<publisher-loc>Philadelphia, PA</publisher-loc>).</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pollitt</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>The method of adaptive comparative judgement</article-title>. <source>Assess. Educ</source>. <volume>19</volume>, <fpage>281</fpage>&#x02013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.1080/0969594X.2012.665354</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rangel-Smith</surname> <given-names>C.</given-names></name> <name><surname>Lynch</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>Addressing the issue of bias in the measurement of reliability in the method of adaptive comparative judgment</article-title>, in <source>36th Pupils&#x00027; Attitudes towards Technology Conference</source> (<publisher-loc>Athlone</publisher-loc>), <fpage>378</fpage>&#x02013;<lpage>387</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Revuelta</surname> <given-names>J.</given-names></name> <name><surname>Ponsoda</surname> <given-names>V.</given-names></name></person-group> (<year>1998</year>). <article-title>A comparison of item exposure control methods in computerized adaptive testing</article-title>. <source>J. Educ. Meas</source>. <volume>35</volume>, <fpage>311</fpage>&#x02013;<lpage>327</lpage>. <pub-id pub-id-type="doi">10.1111/j.1745-3984.1998.tb00541.x</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>R.</given-names></name> <name><surname>Singh</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Text similarity measures in news articles by vector space model using NLP</article-title>. <source>J. Instit. Eng. Ser. B</source> <volume>102</volume>, <fpage>329</fpage>&#x02013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1007/s40031-020-00501-5</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thurstone</surname> <given-names>L. L.</given-names></name></person-group> (<year>1927</year>). <article-title>The method of paired comparisons for social values</article-title>. <source>J. Abnorm. Soc. Psychol</source>. <volume>21</volume>:<fpage>384</fpage>. <pub-id pub-id-type="doi">10.1037/h0065439</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tulkens</surname> <given-names>S.</given-names></name> <name><surname>Emmery</surname> <given-names>C.</given-names></name> <name><surname>Daelemans</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Evaluating unsupervised Dutch Word embeddings as a linguistic resource</article-title>. <source>arXiv [Preprint]</source>. <volume>arXiv</volume>: <fpage>1607.00225</fpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1607.00225</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verhavert</surname> <given-names>S.</given-names></name> <name><surname>De Maeyer</surname> <given-names>S.</given-names></name> <name><surname>Donche</surname> <given-names>V.</given-names></name> <name><surname>Coertjens</surname> <given-names>L.</given-names></name></person-group> (<year>2018</year>). <article-title>Scale separation reliability: what does it mean in the context of comparative judgment?</article-title> <source>Appl. Psychol. Meas</source>. <volume>42</volume>, <fpage>428</fpage>&#x02013;<lpage>445</lpage>. <pub-id pub-id-type="doi">10.1177/0146621617748321</pub-id><pub-id pub-id-type="pmid">30787486</pub-id></citation></ref>
</ref-list>
</back>
</article>