<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Educ.</journal-id>
<journal-title>Frontiers in Education</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Educ.</abbrev-journal-title>
<issn pub-type="epub">2504-284X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/feduc.2019.00028</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Education</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The Influence of Variance in Learner Answers on Automatic Content Scoring</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Horbach</surname> <given-names>Andrea</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/654195/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zesch</surname> <given-names>Torsten</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/565160/overview"/>
</contrib>
</contrib-group>
<aff><institution>Language Technology Lab, University Duisburg-Essen</institution>, <addr-line>Duisburg</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Ronny Scherer, Department of Teacher Education and School Research, Faculty of Educational Sciences, University of Oslo, Norway</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Mark Gierl, University of Alberta, Canada; Dirk Ifenthaler, Universit&#x000E4;t Mannheim, Germany</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Andrea Horbach <email>andrea.horbach&#x00040;uni-due.de</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Educational Psychology, a section of the journal Frontiers in Education</p></fn></author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>04</month>
<year>2019</year>
</pub-date>
<pub-date pub-type="collection">
<year>2019</year>
</pub-date>
<volume>4</volume>
<elocation-id>28</elocation-id>
<history>
<date date-type="received">
<day>07</day>
<month>12</month>
<year>2018</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>03</month>
<year>2019</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2019 Horbach and Zesch.</copyright-statement>
<copyright-year>2019</copyright-year>
<copyright-holder>Horbach and Zesch</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Automatic content scoring is an important application in the area of automatic educational assessment. Short texts written by learners are scored based on their content while spelling and grammar mistakes are usually ignored. The difficulty of automatically scoring such texts varies according to the variance within the learner answers. In this paper, we first discuss factors that influence variance in learner answers, so that practitioners can better estimate if automatic scoring might be applicable to their usage scenario. We then compare the two main paradigms in content scoring: (i) similarity-based and (ii) instance-based methods, and discuss how well they can deal with each of the variance-inducing factors described before.</p></abstract>
<kwd-group>
<kwd>automatic content scoring</kwd>
<kwd>short-answer questions</kwd>
<kwd>natural language processing</kwd>
<kwd>linguistic variance</kwd>
<kwd>machine learning</kwd>
</kwd-group>
<contract-num rid="cn001">01PL16075</contract-num>
<contract-sponsor id="cn001">Bundesministerium f&#x000FC;r Bildung und Forschung<named-content content-type="fundref-id">10.13039/501100002347</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="49"/>
<page-count count="14"/>
<word-count count="10130"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Automatic content scoring is a task from the field of educational natural language processing (NLP). In this task, a free-text answer written by students should be automatically assigned a score or correctness label in the same way as a human teacher would do. Content scoring tasks have been a popular exercise type for a variety of subjects and educational scenarios, such as listening or reading comprehension (in language learning) or definition questions (in science education). In a traditional classroom-setting, answers to such exercises are manually scored by a teacher, but in recent years, their automatic scoring has received growing attention as well (for an overview, see e.g., Ziai et al., <xref ref-type="bibr" rid="B49">2012</xref> and Burrows et al. (<xref ref-type="bibr" rid="B8">2014</xref>)). Automatic content scoring may decrease the manual scoring workload (Burstein et al., <xref ref-type="bibr" rid="B9">2001</xref>) as well as offer more consistency in scoring (Haley et al., <xref ref-type="bibr" rid="B21">2007</xref>). Additionally, automatic scoring provides the advantage that evaluation can happen in the absence of a teacher so that students may receive feedback immediately without having to wait for human scoring. With the increasing popularity of MOOCS and other online learning platforms, automatic scoring has become a topic of growing importance for educators in general.</p>
<p>In this paper, we restrict ourselves to short-answer questions as one instance of free-form assessment. While other test types, such as multiple choice items, are much easier to score, free-text items have an advantage from a testing perspective. They require active formulation instead of just selecting the correct answer from a set of alternatives, i.e., they test production instead of recognition.</p>
<p>Answers to short-answer questions have a typical length between a single phrase and two to three sentences. This places them in length between gap-filling exercises, which often ask for single words, and essays, where learners write longer texts. We do not cover automatic essay scoring in this article, even if it is related to short-answer scoring, and to some extent even the same methods might be applied. The main reason is that scoring essays usually takes into consideration the form of the essay (style, grammar, spelling, etc.) in addition to content (Burstein et al., <xref ref-type="bibr" rid="B10">2013</xref>), which introduces many additional factors of influence that are beyond our scope.</p>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> shows examples from three different content scoring datasets (<sc>Asap</sc>, <sc>Powergrading</sc> and <sc>SemEval</sc>) and highlights the main components of a content scoring scenario: a prompt, a set of learner answers with scoring labels, and (one or several) reference answers.</p>
<list list-type="bullet">
<list-item><p>A <bold>prompt</bold> consists of a particular question and optionally some textual or graphical material the question is about (this additional material is omitted in <xref ref-type="table" rid="T1">Table 1</xref> for space reasons).</p></list-item>
<list-item><p>A set of <bold>learner answers</bold> that are given in response to that prompt. The learner answers in our example have different length ranging from short phrasal answers in <sc>Powergrading</sc> to short paragraphs in <sc>Asap</sc>. They may also contain spelling or grammatical errors. As discussed above, these errors should not be taken into consideration when scoring an answer.</p></list-item>
<list-item><p>The task of automatic scoring is to assign a <bold>scoring label</bold> to a learner answer. If we want to learn such an assignment mechanism, we typically need some scored examples, i.e., learner answers with a gold-standard scoring label assigned by a human. As we can see in the example, the kind of label varies between datasets and can be either numeric or categorial, depending on the nature of the task and also of the purpose of the automatic scoring.</p>
<p>Numeric or binary scoring labels, as we see in <sc>Asap</sc> and <sc>Powergrading</sc>, can be easily summed up and compared. They are thus often used in summative feedback, where the goal is to inform teachers, e.g., about the performance of students in a homework assignment. For formative feedback, which is directed toward the learner, in contrast, a more informative categorical label might be preferable, e.g., to inform a student of their learning progress. The <sc>SemEval</sc> data is an example for scoring labels aiming into that direction.</p></list-item>
<list-item><p>In addition to learner answers, datasets often include teacher-specified <bold>reference answers</bold> for each label. A reference answer showcases a representative answer for a given score and can be used for (human or automatic) comparison with a learner answer. Alternatively, scoring guidelines describing properties of answers with a certain score can be provided. This is often the case when answers are so complex that just providing a small number of reference answers does not nearly cover the conceptual range of possible correct answers and misconceptions. This is for example the case for the <sc>Asap</sc> dataset. When reference answers are given, many datasets only provide reference answers for correct answers and not for incorrect ones, e.g., <sc>Powergrading</sc> and <sc>SemEval</sc>.</p></list-item>
</list>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Exemplary content scoring prompts from three different datasets with reference answers (if available) as well as several learner answers with their scoring labels.</p></caption>
<graphic xlink:href="feduc-04-00028-g0001.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Dataset statistics.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Corpus</bold></th>
<th valign="top" align="center"><bold>&#x00023; Answers</bold></th>
<th valign="top" align="center"><bold>&#x00023; Prompts</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>&#x02205; <italic><sup>tokens</sup></italic>/<italic><sub>answer</sub></italic></bold></th>
</tr>
<tr>
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>min</bold></th>
<th valign="top" align="center"><bold>med</bold></th>
<th valign="top" align="center"><bold>max</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ASAP</td>
<td valign="top" align="center">33,320</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">26.5</td>
<td valign="top" align="center">48.5</td>
<td valign="top" align="center">66.2</td>
</tr>
<tr>
<td valign="top" align="left">ASAP-DE</td>
<td valign="top" align="center">903</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">24.6</td>
<td valign="top" align="center">33.0</td>
<td valign="top" align="center">33.9</td>
</tr>
<tr>
<td valign="top" align="left">CREE</td>
<td valign="top" align="center">566</td>
<td valign="top" align="center">62</td>
<td valign="top" align="center">5.7</td>
<td valign="top" align="center">21.6</td>
<td valign="top" align="center">68.1</td>
</tr>
<tr>
<td valign="top" align="left">CREG</td>
<td valign="top" align="center">1,032</td>
<td valign="top" align="center">177</td>
<td valign="top" align="center">5.0</td>
<td valign="top" align="center">9.7</td>
<td valign="top" align="center">45.8</td>
</tr>
<tr>
<td valign="top" align="left">CS</td>
<td valign="top" align="center">630</td>
<td valign="top" align="center">21</td>
<td valign="top" align="center">6.2</td>
<td valign="top" align="center">20.6</td>
<td valign="top" align="center">36.0</td>
</tr>
<tr>
<td valign="top" align="left">CSSAG</td>
<td valign="top" align="center">1,840</td>
<td valign="top" align="center">31</td>
<td valign="top" align="center">10.9</td>
<td valign="top" align="center">23.5</td>
<td valign="top" align="center">42.6</td>
</tr>
<tr>
<td valign="top" align="left">Powergrading</td>
<td valign="top" align="center">6,980</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">1.9</td>
<td valign="top" align="center">3.4</td>
<td valign="top" align="center">8.4</td>
</tr>
<tr>
<td valign="top" align="left">PT_ASAG</td>
<td valign="top" align="center">3,675</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">9.5</td>
<td valign="top" align="center">14.3</td>
<td valign="top" align="center">40.8</td>
</tr>
<tr>
<td valign="top" align="left">SRA</td>
<td valign="top" align="center">5,239</td>
<td valign="top" align="center">182</td>
<td valign="top" align="center">3.4</td>
<td valign="top" align="center">11.7</td>
<td valign="top" align="center">44.3</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Tokens per answer are counted individually across all answers for one prompt and the minimum, median<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>, and maximum of these values reported. i.e., the prompt with the shortest answers in ASAP has on average 26.5 tokens</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>The content scoring scenario with its interrelated textual components &#x02013; a prompt, learner answers, and a reference answer &#x02013; render automatic content scoring a challenging application of Natural Language Processing which bears strong resemblances to various core NLP fields like paraphrasing (Bhagat and Hovy, <xref ref-type="bibr" rid="B7">2013</xref>), textual entailment (Dagan et al., <xref ref-type="bibr" rid="B12">2013</xref>), and textual similarity (B&#x000E4;r et al., <xref ref-type="bibr" rid="B4">2012</xref>). In all those fields, the semantic relation between two texts is assessed, a method that directly transfers to the comparison between learner and reference answers, as we will see later.</p>
<p>During recent years, many approaches for automatic content scoring have been published on various datasets (see Burrows et al. (<xref ref-type="bibr" rid="B8">2014</xref>) for an overview). A practitioner who is considering using automatic scoring for their own educational data might easily feel overwhelmed. They might find it hard to compare approaches and draw conclusions for their applicability on their specific scoring scenario. In particular, approaches often apply various machine learning methods with a variety of features and are trained and evaluated using different datasets. Thus, comparing any two approaches from the literature can be difficult.</p>
<p>This paper aims to shed light on the individual factors influencing automatic content scoring and identifies the variance in the answers as one key factor that makes scoring difficult. We start in section 2 by discussing the nature of this variance, followed by a discussion of datasets and their parameters that influence variance. We discuss in section 3 properties of automatic scoring methods and review existing approaches, especially with respect to whether they score answers based on features extracted from the answers themselves or based on a comparison with a reference answer. We then discuss in section 4 how these factors can be isolated in scoring experiments. We either provide own experiments, discuss relevant studies from the literature, or formulate requirements for datasets that would make currently infeasible experiments possible.</p>
</sec>
<sec id="s2">
<title>2. Variance in Learner Answers</title>
<p>Variance is the reasons why automatic scoring has to go beyond simply matching learner answers to reference answers. The more variance we find in the learner answers, the more complex the scoring model has to be and therefore the harder is the content scoring task (Pad&#x000F3;, <xref ref-type="bibr" rid="B35">2016</xref>). Thus, in this section, we discuss <italic>why</italic> variance increases the difficulty of automatic scoring and analyze publicly available datasets with respect to the variance-inducing properties.</p>
<sec>
<title>2.1. Sources of Variance</title>
<p>From an NLP perspective, automating content scoring of free-text prompts is a challenging task, mainly due to the textual variance of answers given by the learner. Variance can occur on several levels, as highlighted in <xref ref-type="fig" rid="F2">Figure 2</xref>. It can occur both on the conceptual level as well as on the realization level, whereas variance in realization can mean variance of the linguistic expression as well as orthographic variance.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Sources of variance in content scoring.</p></caption>
<graphic xlink:href="feduc-04-00028-g0002.tif"/>
</fig>
<sec>
<title>2.1.1. Conceptual Variance</title>
<p>Conceptual variance occurs when a prompt asks for multiple aspects or has more than one correct solution. For example, in the prompt <italic>Name one state that borders Mexico</italic> from the <sc>Powergrading</sc> dataset, there are four different correct solutions: <italic>California, Arizona, New Mexico</italic>, and <italic>Texas</italic>. A scoring method needs to take all of them into account. However, conceptually different correct solutions are not the main problem, as their number is usually rather small. The much bigger problem is variance within incorrect answers, as there are usually many ways for a learner to get an answer wrong so that incorrect answers often correspond to several misconceptions. For the Powergrading example prompt in <xref ref-type="table" rid="T1">Table 1</xref> (asking <italic>What is the economic system in the United States?</italic>), frequent misconceptions center around <italic>democracy</italic> or <italic>US dollar</italic>, but there also is a long tail of infrequent other misconceptions.</p>
</sec>
<sec>
<title>2.1.2. Variance in Realization</title>
<p>In contrast to the conceptual variance we have just discussed, which covers different ways of conceptually answering a question, variance in realization means different ways of formulating the same conceptual answer. We consider variance in linguistic expression as well as variance on the orthographic level.</p>
<sec>
<title>2.1.2.1. Variance of Linguistic Expression</title>
<p>This refers to the fact that natural language provides many possibilities to express roughly the same meaning (Meecham and Rees-Miller, <xref ref-type="bibr" rid="B29">2005</xref>; Bhagat and Hovy, <xref ref-type="bibr" rid="B7">2013</xref>). This variance of expression makes it in most cases impossible to preemptively enumerate all correct solutions to a prompt and score new learner answers by string comparison alone. For example consider the following three sentences. They all come from the <sc>SemEval</sc> prompt in <xref ref-type="fig" rid="F1">Figure 1</xref>. The first is a reference answer, while the other two are learner answers.</p>
<list list-type="bullet">
<list-item><p><bold>R</bold> <italic>Voltage is the difference in electrical states between two terminals</italic></p></list-item>
<list-item><p><italic>L</italic><sub>1</sub> <italic>[Voltage] is the difference in electrial stat between terminals</italic></p></list-item>
<list-item><p><italic>L</italic><sub>2</sub> <italic>[Voltage is] the measurement between the electrical states of the positive and negative terminals of a battery</italic>.</p></list-item>
</list>
<p>While the first learner answer in the example above shares many words with the reference answer, the second learner answer has much lower overlap. The term <italic>difference</italic> is replaced by the related term <italic>measurement</italic>. For such cases of lexical variance, we need some form of external knowledge to decide that <italic>difference</italic> and <italic>measurement</italic> are similar.</p>
</sec>
<sec>
<title>2.1.2.2. Orthographic Variance</title>
<p>A property of (especially non-native) learner data that also contributes toward high realization variance in the data is the orthographic variability and occurrence of linguistic deviations from the standard (Ellis and Barkhuizen, <xref ref-type="bibr" rid="B16">2005</xref>), which can also make it hard for humans to understand what was intended (Reznicek et al., <xref ref-type="bibr" rid="B40">2013</xref>). For example in the learner answer <italic>L</italic><sub>1</sub> above, the learner misspelled <italic>electrical state</italic> as <italic>electrial stat</italic>. The number of spelling errors &#x02013; and thus how pronounced this deviation is &#x02013; depends on a number of factors, such as whether answers have been written by language learners or native speakers or whether answers refer to a text visually available to the learner at the time of writing the answer or not.</p>
</sec>
</sec>
</sec>
<sec>
<title>2.2. Content Scoring Datasets</title>
<p>In the following, we introduce publicly available datasets for content scoring. Afterwards, we categorize all datasets in <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref> according to various factors that influence variance. The datasets come from different research contexts, we present them here in alphabetical order:
<list list-type="bullet">
<list-item><p>The <sc><bold>Asap</bold></sc> <bold>dataset</bold><xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> has been released for the purpose of a scoring competition and contains answers collected at US high schools for 10 different ppts from various subjects. The main distinguishing features for this dataset are the large number of answers per individual prompt as well as relative high length of answers. A German version of the dataset, <sc><bold>Asap-De</bold></sc>, addressing three of the science prompts, has been collected by Horbach et al. (<xref ref-type="bibr" rid="B26">2018</xref>).</p></list-item>
<list-item><p>The <bold>CREE dataset</bold> (Bailey and Meurers, <xref ref-type="bibr" rid="B3">2008</xref>) contains answers given by learners of English as a foreign language for reading comprehension questions. The number of answers per prompt as well as the overall number of learner answers in this dataset is comparably low.</p></list-item>
<list-item><p>The <bold>CREG dataset</bold> (Meurers et al., <xref ref-type="bibr" rid="B30">2011a</xref>) is similar to CREE in that it targets reading comprehension questions for foreign language learners, but here the data is in German, so it is an instance of a non-English dataset. Answers were given by beginning and intermediate German-as-a-foreign-language learners at two US universities and respond to reading comprehension questions.</p></list-item>
<list-item><p>The <bold>CS dataset</bold> (Mohler and Mihalcea, <xref ref-type="bibr" rid="B34">2009</xref>) contains answers to computer science questions given by participants of a university course. In this dataset, the questions stand alone and do not address additional material, such as reading texts or experiment descriptions.</p></list-item>
<list-item><p>The <bold>CSSAG dataset</bold> (Pado and Kiefer, <xref ref-type="bibr" rid="B37">2015</xref>) contains computer science questions collected from participants of a university-level computer-science class in German.</p></list-item>
<list-item><p>The <bold>Powergrading dataset</bold> (Basu et al., <xref ref-type="bibr" rid="B5">2013</xref>) addresses questions from US immigration exams and learner answers have been crowd-sourced. It is unclear what the language proficiency of the writers is, including whether they are native speakers or not. The dataset contains the shortest learner answers of all datasets.</p></list-item>
<list-item><p>The Portuguese <bold>PT_ASAG dataset</bold> (Galhardi et al., <xref ref-type="bibr" rid="B20">2018</xref>) contains learner answers collected in biology classes at schools in Brazil using a web system. Apart from reference answers for each question, the dataset also contains keywords specifying aspects of a good question.</p></list-item>
<list-item><p>The <bold>Student Response Analysis (SRA) dataset</bold> (Dzikovska et al., <xref ref-type="bibr" rid="B15">2013</xref>) was used in SemEval-2013 shared task. It consists of two subsets, both dealing with science questions: The Beetle subset covers student interactions with a tutoring system, while the SciEntsBank subset contains answers to assessment questions. A special feature of this dataset is that learner answers are annotated with three different types of labels: (i) binary correct/incorrect decisions, (ii) with categories used for recognizing textual entailment such as whether an answer entails or contradicts the reference answer), as well as (iii) formative assessment labels, informing students, e.g., that an answer is partially correct, but incomplete.</p></list-item>
</list></p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Overview of content scoring datasets.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Corpus</bold></th>
<th valign="top" align="left"><bold>Prompt type</bold></th>
<th valign="top" align="left"><bold>Language</bold></th>
<th valign="top" align="left"><bold>Learner population</bold></th>
<th valign="top" align="left"><bold>Scoring labels</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ASAP</td>
<td valign="top" align="left">Sciences, biology, reading comprehension</td>
<td valign="top" align="left">English</td>
<td valign="top" align="left">High school students</td>
<td valign="top" align="left">Numeric [0, 1, 2, (3)]</td>
</tr>
<tr>
<td valign="top" align="left">ASAP-DE</td>
<td valign="top" align="left">Sciences</td>
<td valign="top" align="left">German</td>
<td valign="top" align="left">Crowdworkers</td>
<td valign="top" align="left">Numeric [0, 1, 2, (3)]</td>
</tr>
<tr>
<td valign="top" align="left">CREE</td>
<td valign="top" align="left">Reading comprehension for language learning</td>
<td valign="top" align="left">English</td>
<td valign="top" align="left">university students learning English</td>
<td valign="top" align="left">Binary &#x00026; diagnostic</td>
</tr>
<tr>
<td valign="top" align="left">CREG</td>
<td valign="top" align="left">Reading comprehension for language learning</td>
<td valign="top" align="left">German</td>
<td valign="top" align="left">US university students learning German</td>
<td valign="top" align="left">Binary &#x00026; diagnostic</td>
</tr>
<tr>
<td valign="top" align="left">CS</td>
<td valign="top" align="left">Computer science questions</td>
<td valign="top" align="left">English</td>
<td valign="top" align="left">university students</td>
<td valign="top" align="left">Numeric [0, 0.5, &#x02026;, 5]</td>
</tr>
<tr>
<td valign="top" align="left">CSSAG</td>
<td valign="top" align="left">Computer science</td>
<td valign="top" align="left">German</td>
<td valign="top" align="left">University students</td>
<td valign="top" align="left">Numeric [0, 0.5, &#x02026;, 2]</td>
</tr>
<tr>
<td valign="top" align="left">Powergrading</td>
<td valign="top" align="left">Immigration exams</td>
<td valign="top" align="left">English</td>
<td valign="top" align="left">Unknown (crowdworkers)</td>
<td valign="top" align="left">Binary</td>
</tr>
<tr>
<td valign="top" align="left">PT_ASAG</td>
<td valign="top" align="left">Biology</td>
<td valign="top" align="left">Portuguese</td>
<td valign="top" align="left">8th &#x00026; 9th grade students</td>
<td valign="top" align="left">Numeric [0, 1, 2, (3)]</td>
</tr>
<tr>
<td valign="top" align="left">SRA</td>
<td valign="top" align="left">Science questions</td>
<td valign="top" align="left">English</td>
<td valign="top" align="left">High school students</td>
<td valign="top" align="left">Entailment labels (binary &#x00026; diagnostic)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>2.3. Dataset Properties Influencing Variance</title>
<p>We now discuss dataset-inherent properties that can help us to estimate the amount of variance to be expected in data.</p>
<sec>
<title>2.3.1. Prompt Type</title>
<p>The type of prompt has a strong influence on the expected answer variance. Imagine, for example, a factual question like <italic>Where was Mozart born?</italic> and a reading comprehension question such as <italic>What conclusion can you draw from the text?</italic> For the first question, there is no variance in the correct answers (<italic>Salzburg</italic>) and probably only little variance in the misconceptions (<italic>Vienna</italic>). For the second question, a very high variance is to be expected. In general, the more open-ended a question is, the harder it will be to automatize its scoring.</p>
<p>Different answer taxonomies have been proposed to classify questions in the classroom according to the cognitive processes involved for the student and they provide also clues about ease of automatic scoring. Anderson et al. (<xref ref-type="bibr" rid="B1">2001</xref>) provide a classification scheme according to the cognitive skills that are involved in solving an exercise: remembering, understanding, applying, analyzing, evaluating, and creating in ascending order of difficulty for the student. This taxonomy could of course also be applied to content scoring prompts. Pad&#x000F3; (<xref ref-type="bibr" rid="B36">2017</xref>) annotates questions in the CSSAG dataset according to this taxonomy and finds that questions from the lower categories are not only easier for students, but produce also less variance and need less elaborate methods for automatic scoring. She also finds that the instructional context of a question needs to be considered when assigning a level (e.g., to differentiate between a real analyzing question and one that is actually a remembering question because the analysis has been explicitly made in the course). Therefore it is hard to apply such a taxonomy to a dataset where the instructional context is unknown.</p>
<p>A taxonomy specifically for reading comprehension questions has been developed by Day and Park (<xref ref-type="bibr" rid="B14">2005</xref>). It classifies questions by comprehension as literal, reorganization, inference, prediction, evaluation, and personal response (again ordered from easy to hard). Literal questions are the easiest because their answers can be found verbatim in the text. Such questions tend to have lower variance, especially when given to low-proficiency learners, as they often lift their answers from the text. Also for this taxonomy, it has been found that reading comprehension prompts for language learners focus on the lower comprehension types (Meurers et al., <xref ref-type="bibr" rid="B31">2011b</xref>) and that among these literal questions are easier to score than reorganization and inference questions. We argue that questions with comprehension types higher in the taxonomy contain so much variance that they are difficult to handle automatically. An example for a personal response question from Day and Park (<xref ref-type="bibr" rid="B14">2005</xref>) is <italic>What do you like or dislike about this article?</italic> We argue that answers to such questions go beyond content-based evaluation and rather touch the area of essay scoring, as how an opinion is expressed it might be more important than its actual content.</p>
<p>The modality of a prompt also plays a role. By modality, we mean whether a question refers to a written or a spoken text. Especially for non-native speakers, listening comprehension exercises will yield a much higher variance as learners cannot copy material from the text based on the written form, but mostly write what they think they understood auditorily. This leads especially to a high orthographic variance and makes scoring harder compared to a similar prompt administered as reading comprehension exercise.</p>
<p><xref ref-type="table" rid="T2">Table 2</xref> shows that existing datasets cover very diverse prompts from reading comprehension for language learning over science question to biology and literature questions, but that they do not nearly cover all possible prompt types.</p>
</sec>
<sec>
<title>2.3.2. Answer Length</title>
<p>Answer length of course is strongly related to the type of question asked. <italic>Where</italic> or <italic>when</italic> questions usually require only a phrasal answer, whereas <italic>why</italic> questions are often answered with complete sentences. Shorter answers consisting of only a few words often correspond only to a single concept mentioned in the answer (see the example from the <sc>Powergrading</sc> dataset in <xref ref-type="table" rid="T1">Table 1</xref>), whereas longer answers (as we saw in the <sc>Asap</sc> example) tend to be also conceptually more complex. It seems intuitive that this conceptual complexity is accompanied by a higher variance in the data. In a longer answer, there are more options how to phrase and order ideas in different ways.</p>
<p>Answer length is a measure that can be easily determined for a new dataset once the learner answers are collected, so it can serve as a quick indicator for the ease of scoring. In general, shorter answers can be scored better than longer answers. Of course, also datasets with answers of the same length can display different types of complexity and variance. Nevertheless, we consider answer length as a good and at the same time cheap indicator.</p>
<p><xref ref-type="table" rid="T1">Table 1</xref> presents some core answer length statistics for each dataset. A dataset usually consists of several individual prompts and different prompts in a dataset might differ more or less from each other. To characterize the variance between prompts in a dataset better we give the average answer length in tokens, as well as the minimum, median, and maximum value across the different prompts. <xref ref-type="fig" rid="F3">Figure 3</xref> visualizes for each dataset the distribution of the average answer length per prompt. We see that the individual datasets span a wide range of lengths from very short phrasal answers in <sc>Powergrading</sc> to long answers almost resembling short essays in ASAP. We also see that the number of different prompts and individual learner answers and thus also the number of learner answers for each prompt varies considerably, from datasets with only a very restricted number of answers for each question, such as in CREE and CREG, to several thousand answers per prompt in ASAP.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Average lengths of answers per dataset.</p></caption>
<graphic xlink:href="feduc-04-00028-g0003.tif"/>
</fig>
</sec>
<sec>
<title>2.3.3. Language</title>
<p>The language that is used to answer a prompt, such as English, German, or Chinese, is also an important factor influencing the answer variance. Methods that work well for one language may not be directly transferable to other languages. This is due both to the linguistic properties of individual languages as well as to the availability of language-specific NLP resources used for scoring. By linguistic properties we mean especially the morphological richness of a language and the restrictiveness of word order. If an answer given in English talks about a red apple, it might be sufficient to look for the term <italic>red apple</italic>, while in German, depending on the grammatical context, terms such as <italic>(ein) roter Apfel, (der) rote Apfel, (einen) roten Apfel</italic>, or <italic>(des) roten Apfels</italic> might occur. Thus, a scoring approach based on token n-grams usually needs fewer training instances in English compared to German, as an English n-gram often corresponds to several German n-grams. For morphologically-richer languages such as Finnish or Turkish, approaches developed for English might completely fail.</p>
<p>Freeness of word order is related to morphological richness. Highly inflected languages, such as German, have usually a less restricted word order than English. Thus, n-gram models work well for the mainly linear grammatical structures in English, but less so for German with freer word-order and more long-distance dependencies (Andresen and Zinsmeister, <xref ref-type="bibr" rid="B2">2017</xref>).</p>
<p>As for language resources used in content scoring methods, there are two main areas which have to be considered: linguistic processing tools as well as external resources. Many scoring methods rely on some sort of linguistic processing. The automatic detection of word and sentence boundaries (tokenization) is a minimal requirement necessary for almost all approaches, while some methods additionally use for example lemmatization (detecting the base form of a word), part-of-speech-tagging (labeling words as nouns, verbs or adjectives), or parsing sentences into syntax trees, which represent the internal linguistic structure of a sentence. External resources can be, for example, dictionaries used for spellchecking, but also resources providing information about the similarity between words in a language. Coming back to the example above, to know that <italic>measurement</italic> and <italic>difference</italic> are related, one would either need an ontology crafted by an expert, such as WordNet (Fellbaum, <xref ref-type="bibr" rid="B17">1998</xref>), or would need similarity information derived from large corpora, based on the core observation in distributional semantics that words are similar if they often appear in similar contexts (Firth, <xref ref-type="bibr" rid="B18">1957</xref>). The availability of such tools and resources has to be taken into consideration when planning automatic scoring for a new language.</p>
</sec>
<sec>
<title>2.3.4. Learner Population and Language Proficiency</title>
<p>The learner population is another important factor to consider, as it defines the language proficiency of the learners, i.e., whether they are beginning foreign language learners or highly proficient native speakers. Language proficiency can have two, at first glance contradicting, effects: A low language proficiency might lead to a high variance in terms of orthography, because beginners are more likely to make spelling or grammatical errors. At the same time, being a low-proficiency learner, can equally reduce variance, but on the lexical and syntactic level. This is because such a learner will have a more restricted vocabulary and has acquired fewer grammatical constructions than a native speaker. Moreover, low-proficiency learners might stay closer to the formulations in the prompt, especially when dealing with reading comprehension exercises, where the process of re-using material from the text for an answer is known as &#x0201C;lifting.&#x0201D;</p>
<p>Beginning language learners and fully proficient students are of course only the far points of the scale, while students from different grades in school would rank somewhere in between. <xref ref-type="table" rid="T2">Table 2</xref> shows that the discussed datasets indeed cover a wide range of language proficiencies.</p>
<p>Also the homogeneity of the learner population plays a role: Learners from a homogeneous population can be expected to produce more homogeneous answers. It has, for example, been shown that the native language of a language learner influences the errors a learner makes (Ringbom and Jarvis, <xref ref-type="bibr" rid="B41">2009</xref>). A German learner of English might be more inclined to misspell the word <italic>marmalade</italic> as <italic>marmelade</italic> because of the German cognate <italic>Marmelade</italic> (Beinborn et al., <xref ref-type="bibr" rid="B6">2016</xref>). An automatic scoring engine trained on learner answers given by German learners might thus encounter the misspelling <italic>marmelade</italic> often enough to learn that an answer containing this word is as good as an answer containing the right spelling. However, a model trained on answers by learners from many different countries might not be able to learn (partially overlapping) error patterns for each individual first language of the learners. In a slightly different way, this also applies to native speakers. Consider e.g., answers by students from one university which all attended the same lecture and used the same slides and textbooks for studying (low variance) vs. answers by students from different universities using different learning materials (high variance).</p>
</sec>
<sec>
<title>2.3.5. Other Factors</title>
<p>The following factors do not directly influence the variance found in the data, but are other data-inherent factors that influence the difficulty of automatic scoring.</p>
<sec>
<title>2.3.5.1. Dataset Size</title>
<p>When using machine learning models to perform content scoring, as do all the approaches we discuss in this article, the availability of already-scored answers from which the scoring method can learn is an important parameter (Heilman and Madnani, <xref ref-type="bibr" rid="B22">2015</xref>). The more answers there are to learn from, the better we can usually model what a correct or incorrect answer looks like. The range of available answers covered varies between less than 10 answers for a prompt (as for example in the <sc>Creg</sc> dataset where a model across individual questions is learnt by most approaches dealing with this dataset) and over 3,000 answers per prompt in the <sc>Asap</sc> dataset.</p>
<p>In many practical settings, only a small part of the available data is manually scored and used for training. It has been shown that the choice of training data heavily influences scoring performance and that the variance within the instances selected for training is a major influencing factor (Zesch et al., <xref ref-type="bibr" rid="B46">2015a</xref>; Horbach and Palmer, <xref ref-type="bibr" rid="B25">2016</xref>).</p>
</sec>
<sec>
<title>2.3.5.2. Label Set</title>
<p>Different label sets have been proposed for different content scoring datasets. The educational purpose of the scoring scenario is the main determining factor for this choice. Some datasets such as <sc>CREG</sc> and <sc>SRA</sc> have even more than one label set so that different usage scenarios can be addressed. This purpose can either be to generate <italic>summative</italic> or <italic>formative feedback</italic> (Scriven, <xref ref-type="bibr" rid="B43">1967</xref>). The recipient of summative feedback is the teacher who wants to get an overview of the performance of a number of learners, for example in a placement test or exam situation. In this case, it is important that scores are comparable and can be aggregated so that there is an overall result for a test consisting of several prompts. Binary or numeric scores fit this purpose well. Formative feedback in contrast, as given through the categorical labels in <sc>SRA</sc>, <sc>CREG</sc>, and <sc>CREE</sc>, is directed toward the learner and meant to inform learners about their progress and the problems they might have had with answering a question. This type of feedback in content scoring is, for example, used in automatic tutoring systems. For a learner, the information that she scored 3.5 out of 5 points might be not as informative as a more meaningful feedback message stating that she missed an important concept required in a correct answer. Thus, datasets meant for formative feedback often use categorical labels rather than numeric ones.</p>
<p>The kind of label that is to be predicted obviously influences the scoring difficulty. In general, the more fine-grained the labels, the harder they are to predict given the same overall amount of training data. Also the conceptual spread covered by the labels can make the task more or less difficult. If the labels intend to make very subtle distinctions between similar concepts, the task is more complex than a scoring scheme that differentiates between coarser categories and considers everything as correct that is somewhat related to the correct answer.</p>
</sec>
<sec>
<title>2.3.5.3. Difficulty of the Scoring Task for Humans</title>
<p>All machine learning algorithms learn from a gold-standard produced by having human experts (such as teachers) label the data. If the scoring task is difficult, humans will make errors and label data inconsistently. This noise in the data impedes performance of a machine learning algorithm. If the gold-standard dataset is constructed from two trained human annotators, the inter-annotator-agreement between these two is considered to be an upper bound of the performance that can be expected from a machine. If two teachers agree only in 90% of the scores they assign for the same task, 90% agreement with the gold-standard is also considered the best possible result obtainable by automatic scoring (Gale et al., <xref ref-type="bibr" rid="B19">1992</xref>; Resnik and Lin, <xref ref-type="bibr" rid="B39">2010</xref>). The same argument can be applied for self-consistency. If a teacher labels the same data twice and can reproduce his own cores only for 90% of all answers, we can consider this 90% an upper bound for machine learning. This influence parameter obviously depends on most of the others and cannot be considered in isolation, but it helps to estimate which level of performance is to be expected for a particular prompt.</p>
</sec>
</sec>
</sec>
<sec>
<title>2.4. Summary</title>
<p>In this section, we have discussed several factors that are influencing the variance to be found in learner answers: the prompt type, answer length, language and learner population. We also introduced dataset size, the label set and the scoring difficulty for human scorers as additional parameters that influence the suitability of a dataset for human scoring. In the next section, we first give an overview of content scoring methods and then present a set of experiments that show the influence of some of the discussed factors on content scoring.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Automatic Content Scoring</title>
<p>As explained in the introduction, the overall aim of content scoring is to mimic a teacher&#x00027;s scoring behavior by assigning labels to a learner answers indicating how good the answer is content-wise.</p>
<p>A very large number of automatic content scoring methods have been proposed (see Burrows et al., <xref ref-type="bibr" rid="B8">2014</xref> for an overview), but we argue that most existing methods can be categorized into two main paradigms: similarity-based and instance-based scoring. Hence, instead of analyzing the properties of single scoring methods, we can draw interesting conclusions by comparing the two paradigms.</p>
<sec>
<title>3.1. Similarity-Based Approaches</title>
<p><xref ref-type="fig" rid="F4">Figure 4</xref> gives a schematic overview of similarity-based scoring. The learner answer is compared with a reference answer (or a high-scoring learner answer) based on a similarity metric. If the similarity surpasses a certain threshold (exemplified by 0.7 in <xref ref-type="fig" rid="F4">Figure 4</xref>), the learner answer is considered as correct. Note that reference answers are always examples for correct answers. In the datasets discussed in section 2.2, there are no samples for incorrect answers, although we have seen earlier that also incorrect answers might form groups of answers expressing the same content.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Schematic overview of similarity-based scoring.</p></caption>
<graphic xlink:href="feduc-04-00028-g0004.tif"/>
</fig>
<p>An important factor in the performance of such similarity-based approaches is how the similarity between answers is computed. In the simplest form, it can be computed based on surface overlap, such as token overlap, where the amount of words or characters shared between answers is measured or edit distance, where the number of editing steps necessary to transform one answer into another is counted. These methods work well when different correct answers can be expected to mainly employ the same lexical material. However, when paraphrases are expected to be lexically diverse, surface-based methods might not be optimal. Consider the hypothetical sentence pair <italic>Paul presented his mother with a book - Mary received a novel from her son as a gift</italic>. In such a case the overlap between the two sentences on the surface is low, while it is clear to human readers that the two sentences convey a very similar meaning. To retrieve the information that <italic>present</italic> and <italic>gift</italic> from the above example are highly similar, semantic similarity methods make use of ontologies like WordNet Fellbaum (<xref ref-type="bibr" rid="B17">1998</xref>) or large background corpora [e.g., latent semantic analysis (Landauer and Dumais, <xref ref-type="bibr" rid="B28">1997</xref>)].</p>
<p>In the content scoring literature, all these kinds of similarities are used. While Meurers et al. (<xref ref-type="bibr" rid="B32">2011c</xref>) mainly rely on similarity on the surface level for different linguistic units (tokens, chunks, dependency triples), methods such as Mohler and Mihalcea (<xref ref-type="bibr" rid="B34">2009</xref>) rely on external knowledge about semantic similarity between words.</p>
</sec>
<sec>
<title>3.2. Instance-Based Approaches</title>
<p>In instance-based approaches, lexical properties of correct answers (words, phrases, or even parts of words) are learned from other learner answers labeled as correct, while commonalities between incorrect answers inform the classifier about common misconceptions in learner answers. One would, for example, as depicted in <xref ref-type="fig" rid="F5">Figure 5</xref>, learn that certain n-grams, such as <italic>electrical states</italic>, are indicators for correct answers while others, such as <italic>battery</italic>, are indicators for incorrect answers. For the scoring process, learner answers are then represented as feature vectors where each feature represents the occurrence of one such n-gram. The information about good n-grams is prompt-specific. For a different prompt, such as one asking for the power source in a certain experiment, <italic>battery</italic> might indicate a good answer, while answers containing the bigram <italic>electrical states</italic> would likely be wrong.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Schematic overview of instance-based scoring.</p></caption>
<graphic xlink:href="feduc-04-00028-g0005.tif"/>
</fig>
<p>As the knowledge used for classification usually comes from the dataset itself and, in many approaches, no external knowledge is used in the scoring process (in contrast to similarity-based scoring), instance-based methods tend to need more training data and do not generalize as well across prompts. Instance-based methods have been used, for example, for various work on the <sc>Asap</sc> dataset (Higgins et al., <xref ref-type="bibr" rid="B23">2014</xref>; Zesch et al., <xref ref-type="bibr" rid="B47">2015b</xref>), including all the top-performing systems from the ASAP scoring competition (Conort, <xref ref-type="bibr" rid="B11">2012</xref>; Jesensky, <xref ref-type="bibr" rid="B27">2012</xref>; Tandalla, <xref ref-type="bibr" rid="B44">2012</xref>; Zbontar, <xref ref-type="bibr" rid="B45">2012</xref>), as well as in commercially used systems.</p>
</sec>
<sec>
<title>3.3. Comparison</title>
<p>We presented two conceptually different ways of content scoring, one relying on the similarity with a reference answer (similarity-based) and the other on information about lexical material in the learner answers (instance-based). While we have presented the paradigmatic case for each side, there are of course less clear-cut cases. For example, an instance-based k-nearest-neighbor classifier scores new unlabeled answers by assigning them the label of the closest labeled learner answer. By doing so the classier inherently exploits similarities between answers.</p>
<sec>
<title>3.3.1. Associated Machine Learning Approaches</title>
<p>Classical supervised machine learning approaches have been associated with both types of scoring paradigms. Instance-based approaches often work on feature vectors representing lexical items, while similarity-based approaches (Meurers et al., <xref ref-type="bibr" rid="B32">2011c</xref>; Mohler et al., <xref ref-type="bibr" rid="B33">2011</xref>) use various overlap measures as features or rely on just one similarity metric (Mohler and Mihalcea, <xref ref-type="bibr" rid="B34">2009</xref>). Deep learning methods have been applied for instance-based scoring Riordan et al. (<xref ref-type="bibr" rid="B42">2017</xref>) as well as similarity-based scoring Patil and Agrawal (<xref ref-type="bibr" rid="B38">2018</xref>). As content scoring datasets are often rather small, the performance gain by using deep learning methods has far not been as in other NLP areas, if there was a reported gain at all.</p>
</sec>
<sec>
<title>3.3.2. Source of Knowledge</title>
<p>In general, instance-based approaches mainly use lexical material present in the answers while similarity-based methods often leverage external knowledge resources like WordNet or distributional semantics to bridge the vocabulary gap between differently phrased answers. Deep learning approaches usually also make use of external knowledge in the form of embeddings that also encode similarity between words.</p>
</sec>
<sec>
<title>3.3.3. Prompt Transfer</title>
<p>Another aspect to consider when comparing scoring paradigms is the transferability of models to new prompts. As similarity-methods learn about a relation between two texts rather than the occurrence of certain words or word combinations, such a model can also be transferred to new prompts for which it has not been trained. For instance-based approaches, a particular word combination indicating a good answer for one prompt might not have the same importance for another prompt. We can therefore generally expect that similarity-based models transfer more easily to new prompts.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments and Discussion</title>
<p>In the previous sections, we have introduced (i) the factors influencing the variance of learner answers and the overall difficulty of the scoring task, and (ii) the two major paradigms in automatic content scoring: similarity-based and instance-based scoring. In this section, we bring both together. In the few cases where empirical evidence already exists, we direct the reader to experiments in the literature that address these influences. We design and conduct a set of experiments to explore those sources of variance that have been experimentally examined yet. However, for some dimensions of variance we have no empirical basis as evaluation datasets are sparse and do not cover the full range of necessary properties. In these cases, we instead describe desiderata for datasets that would be needed to investigate such influences. The discussion in this section is aimed at providing guidance for matching paradigms with use-cases in order to allow a practitioner to choose a setup according to the needs of their automatic scoring scenario.</p>
<sec>
<title>4.1. Experimental Setup</title>
<p>Our experiments (instance-based as well as similarity-based) build on the Escrito scoring toolkit (Zesch and Horbach, <xref ref-type="bibr" rid="B48">2018</xref>) (in version 0.0.1) that is implemented based on DKPro TC (Daxenberger et al., <xref ref-type="bibr" rid="B13">2014</xref>) (in version 1.0.1). For preprocessing, we use DKPro Core.<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> We apply sentence splitting, tokenization, POS-tagging and lemmatization. We did not spellcheck the data, as Horbach et al. (<xref ref-type="bibr" rid="B24">2017</xref>) found that the amount of spelling errors in the <sc>Asap</sc> data did not impede scoring performance in an experimental setup similar to ours.</p>
<p>We use a standard machine learning setup, variants of which have been used widely. We extract token and lemma n-gram features, using uni- to trigrams for tokens and bi- to four-grams for characters. We train a support vector machine using the Weka SVM classifier with SMO optimization in its standard configuration, i.e., without standard parameter tuning.</p>
<sec>
<title>4.1.1. Datasets</title>
<p>We select datasets from those discussed above (see section 2.2). The main selection criterion is, that a dataset contains a high number of learner answers per prompt, so that we can investigate the influence of training data size in prompt-specific models. To meet this criterion we use <sc>Powergrading</sc>, <sc>Asap</sc>, and <sc>SemEval</sc>.</p>
</sec>
<sec>
<title>4.1.2. Evaluation Metric</title>
<p>One common type of evaluation measure applicable for all label sets in short answer scoring is accuracy, i.e., the percentage of correctly classified items. This often goes together with a per-class evaluation of precision, recall, and F-score. Kappa values, taking into account the chance agreement between the machine learning outcome and the gold standard also are quite popular. This holds especially for Quadratically Weighted Kappa (QWK) for numeric scores, as it not only considers whether an answer is correctly classified or not, but also how far of an incorrect answer is. As QWK became a quasi-standard through its usage in the Kaggle ASAP challenge, we use it for our experiments as well.</p>
</sec>
<sec>
<title>4.1.3. Learning Curves</title>
<p>We listed the amount of available training data as one important influence factor for scoring performance. We can simulate datasets of different sizes by using random subsamples of a dataset. By doing this iteratively several times and for several amounts of training data, we obtain a learning curve. If a classifier learns from more data results usually improve until the learning curve approximates a flat line. When we provide learning curve experiments, we always sample 100 times for each amount of training data and average over the results.</p>
</sec>
</sec>
<sec>
<title>4.2. Answer Length</title>
<p>As to our knowledge answer length has not been examined as an influencing factor so far, we test the hypothesis that shorter answers are easier to score, as they should have less variance in general. For this purpose, we conduct experiments with increasing amounts of training data and plot the resulting learning curves. Prompts from datasets with shorter answers should converge faster and at a higher kappa than prompts with longer answers. Note that we restrict ourselves to instance-based experiments here, as there is an insufficient number of datasets providing the necessary reference answers. However, we expect the general results to also hold similarity-based methods, as the similarity of longer answers is harder to compute than for shorter answers.</p>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> shows the results for instance-based scoring for a number of prompts covering a wide variety of different average lengths, selected from <sc>Powergrading</sc> (short answers), <sc>Sra</sc> (medium length answers), and <sc>Asap</sc> (long answers, split in two prompts with on-average about 25 tokens per answer as well as eight prompts with more than 45 tokens per answer). We observe that (as expected) shorter prompts are easier to score, but the results between individual prompts (thin lines) within a dataset vary considerably. Thus, we also present the average over all prompts from the dataset (thick line), that clearly support the hypothesis.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Instance-based learning curves for datasets with different average lengths. Thin lines are individual prompts, while the thick line is the average for this dataset. <bold>(A)</bold> very short (POWERGRADING). <bold>(B)</bold> short (SRA). <bold>(C)</bold> medium (ASAP short). <bold>(D)</bold> long (ASAP long).</p></caption>
<graphic xlink:href="feduc-04-00028-g0006.tif"/>
</fig>
<p>These experiments also tell us something about the influence of the number of training data. An obvious finding is that more data yields, for most prompts, better results. A more interesting observation is that the curves for the <sc>Sra</sc> answers level off earlier than for the <sc>Asap</sc> and <sc>Powergrading</sc> datasets. This means we could not learn much more given the current machine learning algorithm, parameter settings and feature set even if we had more training data. The <sc>Asap</sc> and <sc>Powergrading</sc> curves, in contrast, are still raising: if we had more training data available, we could expect a better scoring performance.</p>
</sec>
<sec>
<title>4.3. Prompt Type</title>
<p>In our experiments regarding answer length, we cannot fully isolate effects originating from the length of the answers from other effects like the prompt type (as some prompts require longer answers than others) and learner population (as certain prompts are suitable only for a certain learner population). Therefore, we now try to isolate the effect of the prompt type by choosing prompts with answers of the same length and coming from the same dataset, thus from the same learner population and language.</p>
<p>We select four different prompts from <sc>Powergrading</sc> with a mean length between 3.3 and 4.8 tokens per answers and three different prompts from the <sc>Asap</sc> dataset with an average length between 45 and 53 tokens and show the resulting learning curves for an instance-based setup in <xref ref-type="fig" rid="F7">Figure 7</xref>. We observe that these prompts behave very differently despite a comparable length of the answers. Especially for the <sc>Powergrading</sc> data, performance with very few training data instances varies considerably showing other factors than length contribute to the performance. We assume that for these prompts (with often repetitive answers) the label distribution plays a role, as performance with few training instances suffers because chances are high that only members of the majority class are selected for scoring. For the ASAP prompts, those differences are less pronounced.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Instance-based learning curves for <sc>Powergrading</sc> and <sc>Asap</sc> prompts with comparable lengths.</p></caption>
<graphic xlink:href="feduc-04-00028-g0007.tif"/>
</fig>
<p>With the currently available data, we cannot make any claims about the influence of the prompt type itself, e.g., regarding domain (like <italic>biology prompts are easier than literature prompts</italic>) or modality of the prompt (as this would require having comparable prompts for example as listening and reading comprehension).</p>
</sec>
<sec>
<title>4.4. Language</title>
<p>In order to compare approaches solely based on the language involved, one would need the same prompts administered to comparable learner population but in different languages. The only such available datasets we know about are <sc>Asap</sc> and <sc>Asap-De</sc>. <sc>Asap-De</sc> uses a subset of the prompts of <sc>Asap</sc> translated to German and provides answers from German-speaking crowdworkers (Horbach et al., <xref ref-type="bibr" rid="B26">2018</xref>). These answers were annotated according to the same annotation guidelines. So, while trying to be as comparable as possible, the datasets still differ in the learner population, in addition to the language. Horbach et al. (<xref ref-type="bibr" rid="B26">2018</xref>) compared instance-based automatic scoring on the two datasets and found results to be in a similar range with a slight performance benefit for the German data. However, they also reported differences in the nature of the data &#x02013; resulting potentially from the different learner populations &#x02013;, such as a different label distribution and considerably shorter answers for German, which they attribute to crowdworkers being potentially less motivated then school students in an assessment situation. Therefore, it is unclear whether any of those differences can be blamed on the language difference or the difference in learner population. More controlled data collections would be possible to get results that are specific to the language difference only. One such data collection with answers from students from different countries and thus various language backgrounds is the data from the PISA studies.<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref> Such data would be an ideal testbed to compare learner populations with different native languages on the same prompt administered in various languages.</p>
</sec>
<sec>
<title>4.5. Learner Population</title>
<p>The results mentioned above for the different languages might equally be used as a potential example for the influence of different learner populations. In order to fully isolate the effect of learner population, one would need to collect the same dataset from two different learner groups such as native speakers vs. language learners or high-school vs. university students. To the best of our knowledge, such data is currently not available.</p>
<p>However, one aspect of different learner population is their tendency to make spelling errors. In experiments on the <sc>Asap</sc> dataset, Horbach et al. (<xref ref-type="bibr" rid="B24">2017</xref>) found that the amount of spelling errors present in the data did not negatively influence content scoring performance. Only if the amount of spelling errors per answer was artificially increased, scoring performance decreased, especially, if errors followed a random pattern (unlikely to occur in real data) and if scoring methods relied on the occurrence of certain words and ignored sub-word information (i.e., certain character combinations).</p>
</sec>
<sec>
<title>4.6. Label Set</title>
<p>When discussion influence factors, we assumed that a dataset with more individual labels is harder to score than a dataset with binary labels. The influence of different label sets was already tested in previous work, especially in the SemEval Shared Task &#x0201C;The Joint Student Response Analysis and 8th Recognizing Textual Entailment Challenge&#x0201D; (Dzikovska et al., <xref ref-type="bibr" rid="B15">2013</xref>). The <sc>Sra</sc> dataset used for this challenge is annotated with three label sets of different granularity: two, three or five labels providing increasing levels of feedback to the learner. The two-way task just informs learners whether their answer was correct or not. The 3-way task additionally distinguishes between contradictory answers (contradicting the learner answers) and other incorrect answers. In the 5-way task, answers classified as incorrect in the 3-way task are classified in an even more fine-grained manner as &#x0201C;partially correct, but incomplete,&#x0201D; &#x0201C;irrelevant for the question,&#x0201D; or &#x0201C;not in the domain&#x0201D; (such as <italic>I don&#x00027;t know</italic>.).</p>
<p>Seven out of nine systems participating the SemEval Shared Task reported results for each of these label sets. For all of them performance was best for the 2-way task (with a mean weighted F-Score of .720 for the best performing system) and worst for the 5-way task (0.547 mean weighted F-Score, again for the best performing system, which was a different one then for the 2-way result). This clearly shows that the expected effect of more fine-grained label sets being more difficult to score automatically.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Conclusion and Future Work</title>
<p>In this paper, we discussed the different influence factors that determine how much variance we see in the learner answers toward a specific prompt and how this variance influences automatic scoring performance. These factors include the type of prompt, the language in the data, the average length of answers as well as the number of training instances that are available. Of course, these factors are interdependent and influence each other. It is thus hard to decide based on purely theoretical speculations whether, for example, medium length answers to a factoid question given by German native speakers annotated with binary scoring labels and with a large number of training instances are easier or harder to score than shorter answers in non-native English with numeric labels and a smaller set of training instances. Such questions can only be answered empirically, but the available datasets do not nearly cover the available parameter space exhaustively, so that such experiments are not possible in a straightforward manner. That makes it hard to compare different approaches in the literature and it is also a challenge to estimate the performance on new data. Therefore, we presented experiments that show the influence of some of the discussed factors on content scoring.</p>
<p>Our findings give researchers as well as educational practitioners hints about whether content scoring might work for a certain new dataset. At the same time, our paper also highlights the demand for more systematic research, both in terms of dataset creation and automatic scoring. For a number of influence factors, we were not able to clearly assess their influence because data that would allow to investigate a single influence parameter in isolation does not exist. It would thus be desirable for the automatic scoring community to systematically collect new datasets varying only in specific dimensions, such as to ask the same prompt to different learner populations and in different languages in order to further broaden our knowledge about the full contribution of these factors.</p>
</sec>
<sec id="s6">
<title>Data Availability</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: https: www.kaggle.com/c/asap-sas, <ext-link ext-link-type="uri" xlink:href="https://www.microsoft.com/en-us/download/details.aspx?id=52397">https://www.microsoft.com/en-us/download/details.aspx?id=52397</ext-link> and <ext-link ext-link-type="uri" xlink:href="https://www.cs.york.ac.uk/semeval-2013/task7/index.php%3Fid=data.html">https://www.cs.york.ac.uk/semeval-2013/task7/index.php%3Fid=data.html</ext-link>.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>All authors listed have made a substantial, direct and intellectual contribution to the work, and approved it for publication.</p>
<sec>
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anderson</surname> <given-names>L.</given-names></name> <name><surname>Krathwohl</surname> <given-names>D.</given-names></name> <name><surname>Bloom</surname> <given-names>B.</given-names></name></person-group> (<year>2001</year>). <source>A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom&#x00027;s Taxonomy of Educational Objectives</source>. Longman.</citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andresen</surname> <given-names>M.</given-names></name> <name><surname>Zinsmeister</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). &#x0201C;The benefit of syntactic vs. linear n-grams for linguistic description,&#x0201D; in <source>Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017)</source>, <fpage>4</fpage>&#x02013;<lpage>14</lpage>.</citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bailey</surname> <given-names>S.</given-names></name> <name><surname>Meurers</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>Diagnosing meaning errors in short answers to reading comprehension questions</article-title>, in <source>Proceedings of the Third Workshop on Innovative Use of NLP for Building Educational Applications</source> (<publisher-loc>Columbus</publisher-loc>), <fpage>107</fpage>&#x02013;<lpage>115</lpage>.</citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>B&#x000E4;r</surname> <given-names>D.</given-names></name> <name><surname>Biemann</surname> <given-names>C.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name> <name><surname>Zesch</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>UKP: computing semantic textual similarity by combining multiple content similarity measures</article-title>, in <source>Proceedings of the 6th International Workshop on Semantic Evaluation, Held in Conjunction With the 1st Joint Conference on Lexical and Computational Semantics</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>435</fpage>&#x02013;<lpage>440</lpage>.</citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Basu</surname> <given-names>S.</given-names></name> <name><surname>Jacobs</surname> <given-names>C.</given-names></name> <name><surname>Vanderwende</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>Powergrading: a clustering approach to amplify human effort for short answer grading</article-title>. <source>Trans. Assoc. Comput. Linguist.</source> <volume>1</volume>, <fpage>391</fpage>&#x02013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1162/tacla00236</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Beinborn</surname> <given-names>L.</given-names></name> <name><surname>Zesch</surname> <given-names>T.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name></person-group> (<year>2016</year>). <article-title>Predicting the spelling difficulty of words for language learners</article-title>, in <source>Proceedings of the Building Educational Applications Workshop at NAACL</source> (<publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>ACL</publisher-name>), <fpage>73</fpage>&#x02013;<lpage>83</lpage>.</citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bhagat</surname> <given-names>R.</given-names></name> <name><surname>Hovy</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). What is a paraphrase? <source>Comput. Linguist.</source> <volume>39</volume>, <fpage>463</fpage>&#x02013;<lpage>472</lpage>. <pub-id pub-id-type="doi">10.1162/COLI_a_00166</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burrows</surname> <given-names>S.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name> <name><surname>Stein</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>The eras and trends of automatic short answer grading</article-title>. <source>Int. J. Art. Intell. Educ.</source> <volume>25</volume>, <fpage>60</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1007/s40593-014-0026-8</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burstein</surname> <given-names>J.</given-names></name> <name><surname>Leacock</surname> <given-names>C.</given-names></name> <name><surname>Swartz</surname> <given-names>R.</given-names></name></person-group> (<year>2001</year>). <source>Automated evaluation of essays and short answers</source>. Abingdon-on-Thames: Taylor &#x00026; Francis Group.</citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burstein</surname> <given-names>J.</given-names></name> <name><surname>Tetreault</surname> <given-names>J.</given-names></name> <name><surname>Madnani</surname> <given-names>N.</given-names></name></person-group> (<year>2013</year>). <article-title>The e-rater automated essay scoring system</article-title>. <source>Handbook of Automated Essay Evaluation: Current Applications and New Directions</source>, eds M. D. Shermis and J. Burstein (Routledge), <fpage>55</fpage>&#x02013;<lpage>67</lpage>.</citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conort</surname> <given-names>X.</given-names></name></person-group> (<year>2012</year>). <article-title>Short answer scoring: explanation of Gxav solution</article-title>, in <source>ASAP Short Answer Scoring Competition System Description</source>.</citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dagan</surname> <given-names>I.</given-names></name> <name><surname>Roth</surname> <given-names>D.</given-names></name> <name><surname>Sammons</surname> <given-names>M.</given-names></name> <name><surname>Zanzotto</surname> <given-names>F. M.</given-names></name></person-group> (<year>2013</year>). <source>Recognizing Textual Entail-ment: Models and Applications</source>. Morgan &#x00026; Claypool Publishers.</citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Daxenberger</surname> <given-names>J.</given-names></name> <name><surname>Ferschke</surname> <given-names>O.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name> <name><surname>Zesch</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Dkpro tc: a java-based framework for supervised learning experiments on textual data</article-title>, in <source>Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source> (<publisher-loc>Baltimore, MD</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>61</fpage>&#x02013;<lpage>66</lpage>.</citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Day</surname> <given-names>R. R.</given-names></name> <name><surname>Park</surname> <given-names>J. S.</given-names></name></person-group> (<year>2005</year>). <article-title>Developing reading comprehension questions</article-title>. <source>Reading Foreign Lang.</source> <volume>17</volume>, <fpage>60</fpage>&#x02013;<lpage>73</lpage>.</citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dzikovska</surname> <given-names>M. O.</given-names></name> <name><surname>Nielsen</surname> <given-names>R.</given-names></name> <name><surname>Brew</surname> <given-names>C.</given-names></name> <name><surname>Leacock</surname> <given-names>C.</given-names></name> <name><surname>Giampiccolo</surname> <given-names>D.</given-names></name> <name><surname>Bentivogli</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Semeval-2013 task 7: The joint student response analysis and 8th recognizing textual entailment challenge</article-title>, <source>*SEM 2013: The First Joint Conference on Lexical and Computational Semantics.</source></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ellis</surname> <given-names>R.</given-names></name> <name><surname>Barkhuizen</surname> <given-names>G. P.</given-names></name></person-group> (<year>2005</year>). <source>Analysing Learner Language</source>. Oxford University Press Oxford.</citation></ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Fellbaum</surname> <given-names>C.</given-names></name></person-group> (ed.). (<year>1998</year>). <source>WordNet: An Electronic Lexical Database</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Language, Speech, and Communication</publisher-name>. MIT Press.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Firth</surname> <given-names>J. R.</given-names></name></person-group> (<year>1957</year>). <source>A Synopsis of Linguistic Theory 1930-55.</source> <publisher-loc>London</publisher-loc>: <publisher-name>Longmans</publisher-name>.</citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gale</surname> <given-names>W.</given-names></name> <name><surname>Church</surname> <given-names>K. W.</given-names></name> <name><surname>Yarowsky</surname> <given-names>D.</given-names></name></person-group> (<year>1992</year>). <article-title>Estimating upper and lower bounds on the performance of word-sense disambiguation programs</article-title>, in <source>Proceedings of the 30th Annual Meeting on Association for Computational Linguistics</source> (<publisher-loc>Association for Computational Linguistics</publisher-loc>), <fpage>249</fpage>&#x02013;<lpage>256</lpage>.</citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galhardi</surname> <given-names>L.</given-names></name> <name><surname>Barbosa</surname> <given-names>C. R.</given-names></name> <name><surname>de Souza</surname> <given-names>R. C. T.</given-names></name> <name><surname>Brancher</surname> <given-names>J. D.</given-names></name></person-group> (<year>2018</year>). <article-title>Portuguese automatic short answer grading</article-title>, in <source>Brazilian Symposium on Computers in Education (Simp&#x000F3;sio Brasileiro de Inform&#x000E1;tica na Educa&#x000E7;&#x000E3;o-SBIE)</source>, <fpage>1373</fpage>.</citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haley</surname> <given-names>D. T.</given-names></name> <name><surname>Thomas</surname> <given-names>P.</given-names></name> <name><surname>De Roeck</surname> <given-names>A.</given-names></name> <name><surname>Petre</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <article-title>Measuring improvement in latent semantic analysis-based marking systems: using a computer to mark questions about html</article-title>, in <source>Proceedings of the Ninth Australasian Conference on Computing Education - Volume 66</source>, ACE &#x00027;07 (<publisher-loc>Darlinghurst, NSW</publisher-loc>: <publisher-name>Australian Computer Society, Inc.</publisher-name>), <fpage>35</fpage>&#x02013;<lpage>42</lpage></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heilman</surname> <given-names>M.</given-names></name> <name><surname>Madnani</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). &#x0201C;The impact of training data on automated short answer scoring performance,&#x0201D; in <source>Proceedings of the Tenth Workshop on Innovative Use of NLP for Building Educational Applications</source>, <fpage>81</fpage>&#x02013;<lpage>85</lpage>.</citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Higgins</surname> <given-names>D.</given-names></name> <name><surname>Brew</surname> <given-names>C.</given-names></name> <name><surname>Heilman</surname> <given-names>M.</given-names></name> <name><surname>Ziai</surname> <given-names>R.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Cahill</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Is getting the right answer just about choosing the right words? the role of syntactically-informed features in short answer scoring</article-title>. <source>arXiv [preprint]. arXiv:1403.0801</source>.</citation></ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Horbach</surname> <given-names>A.</given-names></name> <name><surname>Ding</surname> <given-names>Y.</given-names></name> <name><surname>Zesch</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>The influence of spelling error on content scoring performance</article-title>, in <source>Proceedings of the 4th Workshop on Natural Language Processing Techniques for Educational Applications</source> (<publisher-loc>Taipei</publisher-loc>: <publisher-loc>AFNLP</publisher-loc>), <fpage>45</fpage>&#x02013;<lpage>53</lpage>.</citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Horbach</surname> <given-names>A.</given-names></name> <name><surname>Palmer</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Investigating active learning for short-answer scoring</article-title>, in <source>Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>, <fpage>301</fpage>&#x02013;<lpage>311</lpage>.</citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Horbach</surname> <given-names>A.</given-names></name> <name><surname>Stennmanns</surname> <given-names>S.</given-names></name> <name><surname>Zesch</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>Cross-lingual content scoring</article-title>, in <source>Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications</source> (<publisher-loc>New Orleans, LA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>410</fpage>&#x02013;<lpage>419</lpage>.</citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jesensky</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Team JJJ technical methods paper</article-title>, in <source>ASAP Short Answer Scoring Competition System Description</source>.</citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Landauer</surname> <given-names>T. K.</given-names></name> <name><surname>Dumais</surname> <given-names>S. T.</given-names></name></person-group> (<year>1997</year>). <article-title>A solution to plato&#x00027;s problem: the latent semantic analysis theory of acquisition, induction, and representation of knowledge</article-title>. <source>Psychol. Rev.</source> <volume>104</volume>:<fpage>211</fpage>.</citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Meecham</surname> <given-names>M.</given-names></name> <name><surname>Rees-Miller</surname> <given-names>J.</given-names></name></person-group> (<year>2005</year>). <article-title>Language in social contexts</article-title>. <source>Contemporary Linguistics</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Bedford/St. Martin&#x00027;s</publisher-name>), <fpage>537</fpage>&#x02013;<lpage>590</lpage>.</citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meurers</surname> <given-names>D.</given-names></name> <name><surname>Ziai</surname> <given-names>R.</given-names></name> <name><surname>Ott</surname> <given-names>N.</given-names></name> <name><surname>Kopp</surname> <given-names>J.</given-names></name></person-group> (<year>2011a</year>). <article-title>Corpus of reading comprehension exercises in German</article-title>. <source>CREG-1032, SFB 833: Bedeutungskonstitution - Dynamik und Adaptivit&#x000E4;t sprachlicher Strukturen, Project A4, Universit&#x000E4;t T&#x000FC;bingen.</source></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Meurers</surname> <given-names>D.</given-names></name> <name><surname>Ziai</surname> <given-names>R.</given-names></name> <name><surname>Ott</surname> <given-names>N.</given-names></name> <name><surname>Kopp</surname> <given-names>J.</given-names></name></person-group> (<year>2011b</year>). &#x0201C;Evaluating answers to reading comprehension questions in context: results for german and the role of information structure,&#x0201D; in <source>Proceedings of the TextInfer 2011 Workshop on Textual Entailment</source>, TIWTE &#x00027;11 (<publisher-loc>Stroudsburg, PA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Meurers</surname> <given-names>D.</given-names></name> <name><surname>Ziai</surname> <given-names>R.</given-names></name> <name><surname>Ott</surname> <given-names>N.</given-names></name> <name><surname>Kopp</surname> <given-names>J.</given-names></name></person-group> (<year>2011c</year>). <article-title>Evaluating answers to reading comprehension questions in context: results for german and the role of information structure</article-title>, in <source>Proceedings of the TextInfer 2011 Workshop on Textual Entailment</source> (<publisher-loc>Edinburgh</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohler</surname> <given-names>M.</given-names></name> <name><surname>Bunescu</surname> <given-names>R. C.</given-names></name> <name><surname>Mihalcea</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Learning to grade short answer questions using semantic similarity measures and dependency graph alignments</article-title>, in <source>ACL</source>, <fpage>752</fpage>&#x02013;<lpage>762</lpage>.</citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mohler</surname> <given-names>M.</given-names></name> <name><surname>Mihalcea</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>Text-to-text semantic similarity for automatic short answer grading</article-title>, in <source>Proceedings of the 12th Conference of the European Chapter of the Association for Computational Linguistics</source> (<publisher-loc>Association for Computational Linguistics</publisher-loc>), <fpage>567</fpage>&#x02013;<lpage>575</lpage>.</citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pad&#x000F3;</surname> <given-names>U.</given-names></name></person-group> (<year>2016</year>). <article-title>Get semantic with me! the usefulness of different feature types for short-answer grading</article-title>, in <source>Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers</source>, <fpage>2186</fpage>&#x02013;<lpage>2195</lpage>.</citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pad&#x000F3;</surname> <given-names>U.</given-names></name></person-group> (<year>2017</year>). <article-title>Question difficulty&#x02013;how to estimate without norming, how to use for automated grading</article-title>, in <source>Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications</source>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>.</citation></ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pado</surname> <given-names>U.</given-names></name> <name><surname>Kiefer</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Short answer grading: when sorting helps and when it doesn&#x00027;t</article-title>, in <source>Proceedings of the 4th workshop on NLP for Computer Assisted Language Learning at NODALIDA 2015, Vilnius, 11th May, 2015</source> (<publisher-loc>Link&#x000F6;ping University Electronic Press</publisher-loc>), <fpage>42</fpage>&#x02013;<lpage>50</lpage>.</citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Patil</surname> <given-names>P.</given-names></name> <name><surname>Agrawal</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <source>Auto Grader for Short Answer Questions.</source> Stanford University, CS229.</citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Resnik</surname> <given-names>P.</given-names></name> <name><surname>Lin</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Evaluation of nlp systems</article-title>, in <source>The Handbook of Computational Linguistics and Natural Language Processing</source>.</citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reznicek</surname> <given-names>M.</given-names></name> <name><surname>Ludeling</surname> <given-names>A.</given-names></name> <name><surname>Hirschmann</surname> <given-names>H.</given-names></name></person-group> (<year>2013</year>). <article-title>Competing target hypotheses in the falko corpus</article-title>, in <source>Automatic Treatment and Analysis of Learner Corpus Data</source>, Vol 59, eds A. D&#x000ED;az-Negrillo, N. Ballier, and P. Thompson (John Benjamins Publishing Company), <fpage>101</fpage>&#x02013;<lpage>123</lpage>.</citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ringbom</surname> <given-names>H.</given-names></name> <name><surname>Jarvis</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <source>Chapter 7: The Importance of Cross-Linguistic Similarity in Foreign Language Learning</source>. John Wiley &#x00026; Sons, Ltd.</citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Riordan</surname> <given-names>B.</given-names></name> <name><surname>Horbach</surname> <given-names>A.</given-names></name> <name><surname>Cahill</surname> <given-names>A.</given-names></name> <name><surname>Zesch</surname> <given-names>T.</given-names></name> <name><surname>Lee</surname> <given-names>C. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Investigating neural architectures for short answer scoring</article-title>, in <source>Proceedings of the Building Educational Applications Workshop at EMNLP</source> (<publisher-loc>Copenhagen</publisher-loc>), <fpage>159</fpage>&#x02013;<lpage>168</lpage>.</citation></ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Scriven</surname> <given-names>M.</given-names></name></person-group> (<year>1967</year>). <article-title>The methodology of evaluation</article-title>, in <source>Perspectives of Curriculum Evaluation, AERA Monograph Series on Curriculum Evaluation</source>, Volume 1, eds R. Tyler, R. Gagn&#x000E9;, and M. Scriven (<publisher-loc>Chicago, IL</publisher-loc>: <publisher-name>Rand McNally</publisher-name>), <fpage>39</fpage>&#x02013;<lpage>83</lpage>.</citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tandalla</surname> <given-names>L.</given-names></name></person-group> (<year>2012</year>). <article-title>Scoring short answer essays</article-title>. <source>ASAP Short Answer Scoring Competition System Description</source>. The Hewlett Foundation.</citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zbontar</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Short answer scoring by stacking</article-title>. <source>ASAP Short Answer Scoring Competition System Description. Retrieved July</source>.</citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zesch</surname> <given-names>T.</given-names></name> <name><surname>Heilman</surname> <given-names>M.</given-names></name> <name><surname>Cahill</surname> <given-names>A.</given-names></name></person-group> (<year>2015a</year>). <article-title>Reducing annotation efforts in supervised short answer scoring</article-title>, in <source>Proceedings of the Tenth Workshop on Innovative Use of NLP for Building Educational Applications</source>, <fpage>124</fpage>&#x02013;<lpage>132</lpage>.</citation></ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zesch</surname> <given-names>T.</given-names></name> <name><surname>Heilman</surname> <given-names>M.</given-names></name> <name><surname>Cahill</surname> <given-names>A.</given-names></name></person-group> (<year>2015b</year>). <article-title>Reducing annotation efforts in supervised short answer scoring</article-title>, in <source>Proceedings of the Building Educational Applications Workshop at NAACL</source> (<publisher-loc>Denver, CO</publisher-loc>), <fpage>124</fpage>&#x02013;<lpage>132</lpage>.</citation></ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zesch</surname> <given-names>T.</given-names></name> <name><surname>Horbach</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>ESCRITO - An NLP-Enhanced Educational Scoring Toolkit</article-title>, in <source>Proceedings of the Language Resources and Evaluation Conference (LREC)</source> (<publisher-loc>Miyazaki</publisher-loc>: <publisher-name>European Language Resources Association (ELRA)</publisher-name>).</citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ziai</surname> <given-names>R.</given-names></name> <name><surname>Ott</surname> <given-names>N.</given-names></name> <name><surname>Meurers</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>Short answer assessment: establishing links between research strands</article-title>, in <source>Proceedings of the Seventh Workshop on Building Educational Applications Using NLP</source>, NAACL HLT &#x00027;12 (<publisher-loc>Stroudsburg, PA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>190</fpage>&#x02013;<lpage>200</lpage>.</citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>For the median, we report the lower median if there is an even number of items, so that the value corresponds to the average number of tokens per answer of a specific prompt.</p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/c/asap-sas">https://www.kaggle.com/c/asap-sas</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup><ext-link ext-link-type="uri" xlink:href="https://dkpro.github.io/dkpro-core/">https://dkpro.github.io/dkpro-core/</ext-link></p></fn>
<fn id="fn0004"><p><sup>4</sup><ext-link ext-link-type="uri" xlink:href="http://www.oecd.org/pisa/">http://www.oecd.org/pisa/</ext-link></p></fn>
</fn-group>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This work has been funded by the German Federal Ministry of Education and Research under grant no. FKZ 01PL16075.</p></fn>
</fn-group>
</back>
</article>