<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2023.1112365</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A study on surprisal and semantic relatedness for eye-tracking data prediction</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Salicchi</surname> <given-names>Lavinia</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2118801/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chersoni</surname> <given-names>Emmanuele</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2130569/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Lenci</surname> <given-names>Alessandro</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/472987/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Chinese and Bilingual Studies, The Hong Kong Polytechnic University, Kowloon</institution>, <addr-line>Hong Kong SAR</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Computational Linguistics Laboratory (CoLing Lab), University of Pisa</institution>, <addr-line>Pisa</addr-line>, <country>Italy</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Nora Hollenstein, University of Copenhagen, Denmark</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Joseph Marvin Imperial, University of Bath, United Kingdom; Yohei Oseki, The University of Tokyo, Japan</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Lavinia Salicchi &#x02709; <email>lavinia.salicchi&#x00040;connect.polyu.hk</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Language Sciences, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1112365</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Salicchi, Chersoni and Lenci.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Salicchi, Chersoni and Lenci</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Previous research in computational linguistics dedicated a lot of effort to using language modeling and/or distributional semantic models to predict metrics extracted from eye-tracking data. However, it is not clear whether the two components have a distinct contribution, with recent studies claiming that surprisal scores estimated with large-scale, deep learning-based language models subsume the semantic relatedness component. In our study, we propose a regression experiment for estimating different eye-tracking metrics on two English corpora, contrasting the quality of the predictions with and without the surprisal and the relatedness components. Different types of relatedness scores derived from both static and contextual models have also been tested. Our results suggest that both components play a role in the prediction, with semantic relatedness surprisingly contributing also to the prediction of function words. Moreover, they show that when the metric is computed with the contextual embeddings of the BERT model, it is able to explain a higher amount of variance.</p></abstract>
<kwd-group>
<kwd>cognitive modeling</kwd>
<kwd>surprisal</kwd>
<kwd>semantic relatedness</kwd>
<kwd>cosine similarity</kwd>
<kwd>language models</kwd>
<kwd>distributional semantics</kwd>
<kwd>eye-tracking</kwd>
</kwd-group>
<counts>
<fig-count count="1"/>
<table-count count="8"/>
<equation-count count="1"/>
<ref-count count="85"/>
<page-count count="12"/>
<word-count count="10099"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Eye-tracking data recorded during reading provide important evidence about the factors influencing language comprehension (Rayner et al., <xref ref-type="bibr" rid="B64">1989</xref>; Rayner, <xref ref-type="bibr" rid="B62">1998</xref>). In the investigation of potential predictors of human reading patterns, cognitive studies have focused their attention on two specific factors, among the others: (i) the semantic coherence of a word with the rest of the sentence (Ehrlich and Rayner, <xref ref-type="bibr" rid="B15">1981</xref>; Pynte et al., <xref ref-type="bibr" rid="B58">2008</xref>; Mitchell et al., <xref ref-type="bibr" rid="B51">2010</xref>), which is typically assessed <italic>via semantic relatedness</italic> metrics (usually the <italic>cosine</italic>) computed with <italic>distributional word embeddings</italic>, and (ii) the predictability of the word from its previous context, as measured by <italic>surprisal</italic> (Hale, <xref ref-type="bibr" rid="B25">2001</xref>; Levy, <xref ref-type="bibr" rid="B44">2008</xref>). Initially, the two factors were considered separately, and the general idea was that words having low semantic coherence and low in-context predictability (i.e., high surprisal) induce longer reading times. This hypothesis was instead questioned by Frank (<xref ref-type="bibr" rid="B18">2017</xref>), who argued that previous findings had to be attributed to a confound between semantic relatedness and word predictability and that the effect of the former disappeared once surprisal was factored out.</p>
<p>Our work aims at providing further evidence about the complex interplay between semantic relatedness and surprisal as predictors of eye-tracking data. For example, it is unclear whether the fact that no independent effect of relatedness has been found depends on the specific word embedding model being used for measuring it. In fact, there is a large variety of Distributional Semantic Models (DSMs) that are trained with different objectives, and they have been shown to perform differently depending on the task (Lenci et al., <xref ref-type="bibr" rid="B43">2022</xref>). Moreover, the recent introduction of contextual embedding models such as ELMo (Peters et al., <xref ref-type="bibr" rid="B56">2018</xref>) and BERT (Devlin et al., <xref ref-type="bibr" rid="B14">2019</xref>) has also radically changed the way semantic relatedness can be assessed. In particular, contextual embeddings now make it possible to compare the semantic representations of <italic>words in specific contexts</italic> (<italic>token-level representations</italic>), and not just type-level representations that tend to conflate multiple senses of the same word.</p>
<p>The goals of this paper can thus be summarized as follows:</p>
<list list-type="order">
<list-item><p>Investigating whether distributional measures of semantic relatedness between a word and its previous contexts are indeed made redundant by surprisal, or have instead an autonomous explanatory role to model eye-tracking data;</p></list-item>
<list-item><p>Looking into different types of word embeddings, to check whether &#x0201C;classical&#x0201D; static models and contextual ones interact differently or not with surprisal.</p></list-item>
</list>
<p>To explore these issues, we implemented four different linear models to predict three eye-tracking features on two eye-tracking corpora: i) a baseline with word-level features, ii) a model with baseline features and the surprisal between target word and context, iii) a model with baseline features and the relatedness between the vector representing the target word and the vector representing the context, and iv) a model with all the above-mentioned regression features. While surprisal has been consistently computed using a state-of-the-art neural language model GPT2-xl (Radford et al., <xref ref-type="bibr" rid="B61">2019</xref>), the vectors employed in the cosine similarity calculation were obtained using either SGNS (Mikolov et al., <xref ref-type="bibr" rid="B50">2013</xref>) or BERT (Devlin et al., <xref ref-type="bibr" rid="B14">2019</xref>), to compare static and contextual word embedding models.</p>
<p>Our results show that the models including both relatedness and surprisal perform better than the other three, suggesting that, despite the overlap between the two, they contribute differently in explaining the variance in the data. Furthermore, when comparing the models using only relatedness, we noticed that BERT vectors outperform SGNS ones, confirming the added value of contextual embeddings when modeling the relatedness of words in contexts. Finally, we investigated how our models predict eye-tracking feature values for different parts of speech, and we found that while surprisal helps on content words, semantic relatedness contributes to improving the predictions on both function and content words.</p>
</sec>
<sec id="s2">
<title>2. Computational models of human reading times: Surprisal and semantic relatedness</title>
<p>Since the cognitive processes of meaning construction involve the integration of individual word meanings into the syntactic and semantic context, the literature in natural language processing and cognitive science got interested in how such contextual effects on word fixations could be modeled. A first class of computational models has relied on distributional semantics to assess the relatedness of a word with its wider semantic context (Section 2.1); another class of models has explored the connection between the logarithmic probabilities of words in context and their processing difficulty (Section 2.2).</p>
<sec>
<title>2.1. Computational measures for semantic coherence</title>
<p>A fruitful line of research has been investigating the usage of cosine similarity between word embeddings for predicting reading times. The employment of word vectors for modeling reading times originated from classical DSMs (Lenci and Sahlgren, <xref ref-type="bibr" rid="B42">2023</xref>). Pynte et al. (<xref ref-type="bibr" rid="B58">2008</xref>) and Mitchell et al. (<xref ref-type="bibr" rid="B51">2010</xref>) used the semantic distance between a target word and the context as a predictor, measured as 1 min the traditional cosine similarity metric (Turney and Pantel, <xref ref-type="bibr" rid="B79">2010</xref>; Lenci, <xref ref-type="bibr" rid="B41">2018</xref>). The context was in turn modeled as the sum of the distributional vectors representing the words before the target. These studies found strong correlations between semantic distance and reading times: The more semantically related the words, the shorter the fixation durations.</p>
<p>Originally, vector spaces were obtained from the extraction and counting (hence the name of <italic>count models</italic>) of the co-occurrences between the target words and the relevant linguistic contexts. Raw co-occurrences were usually weighted <italic>via</italic> different types of statistical association measures [e.g., Mutual Information, log-likelihood; see Evert (<xref ref-type="bibr" rid="B16">2005</xref>) for an overview] and then the vector space was optionally transformed with some algebraic operation for dimensionality reduction, such as Singular Value Decomposition (Landauer and Dumais, <xref ref-type="bibr" rid="B40">1997</xref>; Bullinaria and Levy, <xref ref-type="bibr" rid="B8">2012</xref>). The contexts could consist either in the words occurring within a window surrounding the target (Lund and Burgess, <xref ref-type="bibr" rid="B47">1996</xref>; Sahlgren, <xref ref-type="bibr" rid="B67">2008</xref>), or in the words linked to the target by syntactic (Pad&#x000F3; and Lapata, <xref ref-type="bibr" rid="B54">2007</xref>; Baroni and Lenci, <xref ref-type="bibr" rid="B4">2010</xref>) or semantic relations (Sayeed et al., <xref ref-type="bibr" rid="B73">2015</xref>).</p>
<p>Later, with the increasing success of deep learning techniques in Natural Language Processing, the so-called <italic>predict models</italic> established themselves as a new standard (Mikolov et al., <xref ref-type="bibr" rid="B50">2013</xref>; Bojanowski et al., <xref ref-type="bibr" rid="B5">2017</xref>). In such models, the learning of word vectors is based on neural network training and framed as a self-supervised language modeling task. One of the most popular predict DSMs is Word2Vec (Mikolov et al., <xref ref-type="bibr" rid="B50">2013</xref>), which includes two main architectures: CBOW, trained for predicting a target word given the context surrounding it, and Skip-Gram, whose learning objective is to predict the surrounding context given a target word. The most common implementation of Skip-Gram makes use of negative sampling (SGNS), whose objective is to discriminate between word sequences that are actually occurring in the data (positive samples) and &#x00022;corrupted&#x00022; samples, which are obtained by randomly replacing a word in a true sequence from the corpus (negative samples).</p>
<p>One of the main limitations of &#x0201C;traditional&#x0201D; word embeddings, both count and predict ones, is that they provide <italic>static</italic> representations of the semantics of a word. They assign a single embedding to each word type, thereby conflating the possible senses of a lexeme and hampering the possibility to address the pervasive phenomena of polysemy and homography. For example, <italic>bank</italic> as a financial agency will have the same vector representation of <italic>bank</italic> as the bank of the river. This way, lexical semantic representations are built at the <italic>type</italic> level only, and the embedding will be a sort of distributional summary of all the instances of a word, no matter how different their senses might be (and probably, the most frequent senses would obscure the minority ones).</p>
<p>The most recent generation of DSMs is said to be <italic>contextual</italic> because they produce a distinct vector for each word instance in context, that is a <italic>token</italic> level representation (Peters et al., <xref ref-type="bibr" rid="B56">2018</xref>; Devlin et al., <xref ref-type="bibr" rid="B14">2019</xref>; Liu et al., <xref ref-type="bibr" rid="B45">2019</xref>). Contextual DSMs generally rely on a multi-encoder network and the word vectors are learned as a function of the internal states, so that a word appearing in different sentence contexts determines different activation states and, as a consequence, is represented by a different vector.</p>
<p>Most contextual DSMs are based on <italic>Transformers</italic> (Vaswani et al., <xref ref-type="bibr" rid="B81">2017</xref>), which use a self-attention mechanism (Bahdanau et al., <xref ref-type="bibr" rid="B2">2014</xref>) for getting the most salient elements in a sentence context and assign them higher weights. BERT (Devlin et al., <xref ref-type="bibr" rid="B14">2019</xref>) is probably the most popular model for generating contextual word representations. BERT is trained on a masked language modeling objective function: random words in the input sentences are replaced by a <monospace>&#x02018;[MASK]&#x00027;</monospace> token and the model attempts to predict the masked word based on the surrounding context. Simultaneously, BERT is optimized on a next sentence prediction task, as the model receives sentence pairs in input and has to predict whether the second sentence is subsequent to the first one in the training data. It should be noticed that BERT is defined as <italic>deeply bidirectional</italic> as, in fact, it takes into account the left-hand and the right-hand context of a word to predict the word filling the masked token. The contextual embeddings produced by BERT have been shown to improve the state-of-the-art performance in several Natural Language Processing tasks (Devlin et al., <xref ref-type="bibr" rid="B14">2019</xref>) and it has been reported that its multilingual versions (i.e., Multilingual BERT, XLM) are able to predict human fixations in multiple languages (Hollenstein et al., <xref ref-type="bibr" rid="B31">2021</xref>, <xref ref-type="bibr" rid="B29">2022a</xref>,<xref ref-type="bibr" rid="B30">b</xref>). Significantly, it was shown that it is possible to extract semantic representations at the type level from BERT just by averaging token vectors of randomly-sampled sentences, and those can achieve a performance close to traditional word embeddings on word similarity tasks (Bommasani et al., <xref ref-type="bibr" rid="B6">2020</xref>; Chronis and Erk, <xref ref-type="bibr" rid="B10">2020</xref>; Lenci et al., <xref ref-type="bibr" rid="B43">2022</xref>) and on word association modeling (Rodriguez and Merlo, <xref ref-type="bibr" rid="B66">2020</xref>).</p>
</sec>
<sec>
<title>2.2. Computational measures for word predictability</title>
<p>A significant part of the psycholinguistic and computational studies modeled naturalistic reading data by means of language model probabilities, being inspired by <italic>surprisal theory</italic> (Hale, <xref ref-type="bibr" rid="B25">2001</xref>, <xref ref-type="bibr" rid="B26">2016</xref>), with the idea that the predictability of a word is the main factor determining the reading times. More specifically, the processing difficulty of a word is considered to be proportional to its <italic>surprisal</italic>, that is, the negative logarithm of the probability of the word given the context. Several studies based on language models adopted surprisal theory as a reference framework for the prediction of eye-tracking data (Demberg and Keller, <xref ref-type="bibr" rid="B13">2008</xref>; Frank and Bod, <xref ref-type="bibr" rid="B19">2011</xref>; Fossum and Levy, <xref ref-type="bibr" rid="B17">2012</xref>; Monsalve et al., <xref ref-type="bibr" rid="B52">2012</xref>; Smith and Levy, <xref ref-type="bibr" rid="B76">2013</xref>). The predictions were typically evaluated on the Dundee Corpus (Kennedy et al., <xref ref-type="bibr" rid="B37">2003</xref>), as one of the earliest corpora with gold standard annotations of eye-tracking measures.</p>
<p>Later research has focused on the quality of the language model to estimate conditional probabilities, finding that models with lower perplexity are a better fit to human reading times (Goodkind and Bicknell, <xref ref-type="bibr" rid="B22">2018</xref>). Following studies confirmed the model perplexity as a significant determinant, making use of more and more advanced neural architectures, such as LSTM (van Schijndel and Linzen, <xref ref-type="bibr" rid="B80">2018</xref>), GRU (Aurnhammer and Frank, <xref ref-type="bibr" rid="B1">2019</xref>), Transformers (Merkx and Frank, <xref ref-type="bibr" rid="B48">2021</xref>), GPT-2 (Wilcox et al., <xref ref-type="bibr" rid="B82">2020</xref>).</p>
<p>Is contextual predictability, that is surprisal, all we need to model human reading behavior? Some recent results suggest that this may not be the case. Goodkind and Bicknell (<xref ref-type="bibr" rid="B23">2021</xref>), for example, investigated the role played on local word statistics, such as word bigram and trigram probability, in sentence processing, and consequently their impact on reading times, finding that they affect processing independently of surprisal. Moreover, Hofmann et al. (<xref ref-type="bibr" rid="B28">2021</xref>) compared different models for computing surprisal as predictors of eye-tracking fixations and found that they explain different and independent proportions of variance in the viewing parameters. For example, classical n-gram-based language models are better at predicting metrics related to short-range access, while RNN models better predict the early preprocessing of the next word.</p>
<p>The models of the GPT family are based on Transformer architectures (Radford et al., <xref ref-type="bibr" rid="B60">2018</xref>, <xref ref-type="bibr" rid="B61">2019</xref>; Brown et al., <xref ref-type="bibr" rid="B7">2020</xref>). Differently from BERT, GPT is a uni-directional, autoregressive Transformer language model, which means that the training objective is to predict the next word, given all of the previous words. GPT-2, in particular, has been commonly used in eye-tracking studies, as the surprisal scores computed by this language model have been proved to be strong predictors of reading times and eye fixations in English (Hao et al., <xref ref-type="bibr" rid="B27">2020</xref>; Wilcox et al., <xref ref-type="bibr" rid="B82">2020</xref>; Merkx and Frank, <xref ref-type="bibr" rid="B48">2021</xref>) and in other languages (e.g., Dutch, German, Hindi, Chinese, Russian) (Salicchi et al., <xref ref-type="bibr" rid="B69">2022</xref>).</p>
<p>The research work on semantic relatedness and surprisal led Frank (<xref ref-type="bibr" rid="B18">2017</xref>) to ask whether these two factors have actually independent effects in the modeling of reading times. The question was motivated by the fact that not all the studies on reading times found effects associated with semantic relatedness (e.g., Traxler et al., <xref ref-type="bibr" rid="B78">2000</xref>; Gordon et al., <xref ref-type="bibr" rid="B24">2006</xref>), although vector space metrics clearly proved to be useful for modeling other types of experimental data on naturalistic reading, such as the N400 amplitude in EEG recordings (Frank and Willems, <xref ref-type="bibr" rid="B20">2017</xref>). Frank suggested that, since DSMs like Word2Vec (Mikolov et al., <xref ref-type="bibr" rid="B50">2013</xref>) are based on word co-occurrence and are optimized for predicting words in context, previous results were due to a confound between semantic relatedness and word predictability. Indeed, when surprisal was factored out, the author showed that the semantic distance effects disappeared. Moreover, the different results obtained in modeling the N400 component in the EEG data were attributed to differences in the stimuli presentation method: while in eye-tracking participants read the text naturally, in many EEG studies words are presented one at a time with unnaturally long durations. Following the findings of Wlotko and Federmeier (<xref ref-type="bibr" rid="B83">2015</xref>) and Frank (<xref ref-type="bibr" rid="B18">2017</xref>) pointed out that, the more natural the presentation rates of the words in the experimental setting in EEG, the smaller the semantic relatedness effects on N400 data tend to be, with no effects at all for behavioral metrics on naturalistic reading. Is distributional semantic relatedness really made redundant by surprisal, or were the results by Frank (<xref ref-type="bibr" rid="B18">2017</xref>) also conditioned by the specific type of embeddings used in the experiments? The analyzes in Sections 3, 4 aim at clarifying this issue.</p>
</sec>
</sec>
<sec sec-type="materials and methods" id="s3">
<title>3. Materials and methods</title>
<sec>
<title>3.1. Definition of eye-tracking metrics in psycholinguistic studies</title>
<p>Several metrics have been defined to describe eye movement features (Rayner, <xref ref-type="bibr" rid="B62">1998</xref>). In this work, we focus on first fixation duration, number of fixations and total reading time. The first fixation duration (FFD), that is the time spent fixing a word for the first time, is typically associated with lexical information processing, like lexical access (Inhoff, <xref ref-type="bibr" rid="B32">1984</xref>), which is heavily affected by word frequency (Balota and Chumbley, <xref ref-type="bibr" rid="B3">1984</xref>). Fast word recognition is obtained when a word can be recognized with a single glance. In this sense, a short FFD reflects a quick and successful lexical access (Hofmann et al., <xref ref-type="bibr" rid="B28">2021</xref>).</p>
<p>However, several words may not be accessed immediately. Words may receive multiple fixations before the eyes move to the next word, and this is reflected by the number of fixations (NF), depending on the integration of the word within the sentence semantics or syntax (Frazier and Rayner, <xref ref-type="bibr" rid="B21">1982</xref>). An alternative metric for this &#x0201C;delayed&#x0201D; lexical access is known as <italic>gaze duration</italic>, which computes directly the sum of the duration of individual fixations before moving to the next word (Inhoff and Radach, <xref ref-type="bibr" rid="B33">1998</xref>; Rayner, <xref ref-type="bibr" rid="B62">1998</xref>).</p>
<p>Finally, the total reading time (TRT), as the sum of all fixation durations on the word, including regressions, is affected by both lexical and sentence-level processing. The TRT is likely to indicate the time required for the full semantic integration of the word in the sentence context (Radach and Kennedy, <xref ref-type="bibr" rid="B59">2013</xref>).</p>
<p>What are the factors affecting word fixations during reading? There is a general consensus that word position, word length, and the number of syllables within the word affect language processing and, consequently, reading behavior and fixations (Just and Carpenter, <xref ref-type="bibr" rid="B35">1980</xref>). It has also been observed that low-frequency words tend to have longer gaze durations and, additionally, they lead to longer gaze on the immediately following words, a phenomenon typically referred to as <italic>spillover effect</italic> (Rayner and Duffy, <xref ref-type="bibr" rid="B63">1986</xref>; Rayner et al., <xref ref-type="bibr" rid="B64">1989</xref>; Remington et al., <xref ref-type="bibr" rid="B65">2018</xref>). A common explanation is that rare and longer words have a higher cognitive load, as they require more time for the semantic integration in the sentence context (Pollatsek et al., <xref ref-type="bibr" rid="B57">2008</xref>), and therefore they may influence the processing of the following words.</p>
</sec>
<sec>
<title>3.2. Eye-tracking corpora</title>
<p>Traditional corpora annotated with eye-tracking data consist of short isolated sentences (or even single words) with particular structures or lexemes, in order to investigate specific syntactic and semantic phenomena. In the present work, we use GECO (Cop et al., <xref ref-type="bibr" rid="B11">2017</xref>) and Provo (Luke and Christianson, <xref ref-type="bibr" rid="B46">2018</xref>), two eye-tracking corpora containing long, complete, and coherent texts.</p>
<p><bold>GECO</bold> is a bilingual corpus in English and Dutch composed of the entire Agatha Christie&#x00027;s novel <italic>The Mysterious Affair at Styles</italic>. The corpus is freely downloadable with a related dataset containing eye-tracking data of 33 subjects (19 of them bilingual, 14 English monolingual) reading the full novel text, presented paragraph-by-paragraph on a screen<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>. In total, GECO is composed of 54,364 tokens.</p>
<p><bold>Provo</bold> contains 55 short English texts about various topics, with 2.5 sentences and 50 words on average, for a total of 2, 689 tokens, and a vocabulary of 1,197 words. These texts were read by 84 native English speakers and their eye-tracking measures were collected and made publicly available online<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>.</p>
<p>GECO and Provo are particularly interesting for our goals because they are corpora of naturalistic reading since data have been recorded from subjects reading real texts, instead of short stimuli created <italic>in vitro</italic>. For every word in the corpora, we extracted the mean total reading time, mean first fixation duration, and mean number of fixations. Mean values were obtained by averaging over the subjects. The choice of modeling mean eye-tracking measures is justified by the high inter-subject consistency of the recorded data.</p>
</sec>
<sec>
<title>3.3. Method</title>
<p>We implemented and compared four main types of linear models (see <xref ref-type="table" rid="T1">Table 1</xref>):</p>
<list list-type="order">
<list-item><p>A baseline model with word-related statistics that are known to influence sentence and word processing (i.e., word frequency, word length, word position within the sentence, previous word frequency, previous word length, and whether or not the previous word was fixated);</p></list-item>
<list-item><p>Two models combining baseline features and cosine similarity, one using Skip-Gram vectors (SGNS), one using BERT vectors;</p></list-item>
<list-item><p>One model with baseline features &#x0002B; surprisal computed using GPT2-xl;</p></list-item>
<list-item><p>Two models with baseline features &#x0002B; surprisal computed using GPT2-xl &#x0002B; cosine similarity, one using SGNS vectors, one using BERT vectors.</p></list-item>
</list>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of the linear models implemented for the experiments.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th valign="top" align="left"><bold>Model name</bold></th>
<th valign="top" align="left"><bold>Features</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="left">Word frequency<break/> Word length<break/> Word position within the sentence<break/> Previous word frequency<break/> Previous word length<break/> Whether or not the previous word was fixated</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="left">Baseline features (same as BL)<break/> Cosine similarity (BERT vectors)</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Baseline features (same as BL)<break/> Cosine similarity (SGNS vectors)</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="left">Baseline features (same as BL)<break/> Surprisal (GPT2-xl)</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-cos</td>
<td valign="top" align="left">Baseline features (same as BL)<break/> Surprisal (GPT2-xl)<break/> Cosine similarity (SGNS vectors)</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">Baseline features (same as BL)<break/> Surprisal (GPT2-xl)<break/> Cosine similarity (BERT vectors)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Recent works have cast doubts on the application of cosine in similarity task while employing contextual vector models. In fact, in contextual embeddings a small number of dimensions (e.g., 3-5) tend to dominate the similarity metric, accounting for most of the data variance (Timkey and van Schijndel, <xref ref-type="bibr" rid="B77">2021</xref>). Moreover, it has been shown that the removal of the outlier dimensions leads to drastic performance drops both in language modeling and in downstream tasks (Kovaleva et al., <xref ref-type="bibr" rid="B39">2021</xref>).</p>
<p>To address this issue, for similarity tasks it has been suggested to correct the comparisons by discounting the &#x0201C;rogue&#x0201D; dimensions or to adopt metrics based on the rank of the dimensions themselves, rather than on their absolute values (Timkey and van Schijndel, <xref ref-type="bibr" rid="B77">2021</xref>). In order to take into account the potential effect of rogue dimensions on computing cosine similarity with BERT, we followed the latter suggestion and we also implemented two further models, in which we use Spearman correlation instead of cosine similarity.</p>
<p>Rank-based metrics have been reported to outperform vector cosine in semantic relatedness tasks (Santus et al., <xref ref-type="bibr" rid="B70">2016a</xref>,<xref ref-type="bibr" rid="B71">b</xref>, <xref ref-type="bibr" rid="B72">2018</xref>; Zhelezniak et al., <xref ref-type="bibr" rid="B85">2019</xref>), and it has been shown that Spearman itself is more correlated with human judgments than cosine (Timkey and van Schijndel, <xref ref-type="bibr" rid="B77">2021</xref>). For each of the resulting eight models, the values to be predicted were first fixation duration (FFD), number of fixations (NF) and total reading time (TRT). We predicted those metrics on both GECO and Provo corpus. We also experimented with models with and without interactions between the features. The models were implemented using the generalized linear models available in R, which have also been used for the statistical analysis.</p>
<p>After we fitted the data of the eye-tracking features with each model, we compared them using the corrected Akaike Information Criterion (AICc) in order to determine the extent to which the goodness of fit improves with the addition of semantic relatedness and surprisal as predictors. Additionally, we also analyzed i) the correlations between linear model errors (as Mean Absolute Error, MAE) and word features, and ii) which parts of speech are easier or harder for each model to predict.</p>
</sec>
<sec>
<title>3.4. Regression features</title>
<sec>
<title>3.4.1. Baseline features</title>
<p>The baseline model includes the following word features: i) the target word and previous word length, computed as the number of letters within the word to be predicted; ii) the target word and previous word frequency, whose values are extracted from Wikipedia;<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> iii) the target word position, as the index of the word within the current sentence; iv) a Boolean value corresponding to 1 if the word preceding the target word was fixated, 0 otherwise. The baseline features are the same used by Frank (<xref ref-type="bibr" rid="B18">2017</xref>).</p>
</sec>
<sec>
<title>3.4.2. Metrics of semantic relatedness</title>
<p>To compute the semantic relatedness between the context and the target word, we extracted vectors for each word, represented the sentence context with a vector, and finally computed, alternatively, the <italic>cosine similarity</italic> or the <italic>Spearman correlation</italic> between the context and the target vectors (the latter metric was used only with the BERT vectors only).</p>
<p>With <bold>SGNS</bold> embeddings, we extracted the pre-trained vectors for each word, and we computed the context vector using an additive model: We summed the vectors of all the words preceding the target and took this as the context representation. For example, given the sentence <italic>The dog chases the cat</italic>, if the target word is <italic>chases</italic>, the context vector will be <inline-formula><mml:math id="M1"><mml:mover class="overrightarrow"><mml:mrow><mml:mi>T</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mover class="overrightarrow"><mml:mrow><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover></mml:math></inline-formula>, while if the target word is <italic>cat</italic>, the context vector will be <inline-formula><mml:math id="M2"><mml:mover class="overrightarrow"><mml:mrow><mml:mi>T</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mover class="overrightarrow"><mml:mrow><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mover class="overrightarrow"><mml:mrow><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mover class="overrightarrow"><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover></mml:math></inline-formula>.</p>
<p>On the other hand, given the bidirectional nature of the <bold>BERT</bold> language model, the input to extract the embeddings from this model required a special preprocessing, since we wanted to avoid the model to &#x0201C;see the future,&#x0201D; by having the target word vector including information also from the right-hand context. Therefore, we fed BERT with sub-sentences. For instance, given the sentence <italic>The dog chases the cat</italic>, we generated the following sub-sentences:</p>
<list list-type="simple">
<list-item><p>S[0] = [<italic>The</italic>]</p></list-item>
<list-item><p>S[1] = [<italic>The dog</italic>]</p></list-item>
<list-item><p>S[2] = [<italic>The dog chases</italic>]</p></list-item>
<list-item><p>S[3] = [<italic>The dog chases the</italic>]</p></list-item>
<list-item><p>S[4] = [<italic>The dog chases the cat</italic>]</p></list-item>
</list>
<p>For each target word, we extracted its vector, when the lexeme occurs at the end of a sub-sentence (e.g., <italic>The</italic> will be extracted in S[0], <italic>dog</italic> in S[1], <italic>chases</italic> in S[2], and so on).</p>
<p>Regarding the context, we used the vector of the special token [CLS], which is created by BERT as a global representation of the input sentence, taking into account how salient each word is for the sentence&#x00027;s meaning. Again, to avoid a representation of the target word itself within the [CLS] vector, we computed the cosine similarity and the Spearman correlation between the target word embedding, and the [CLS] vector of the previous sub-sentence. For example, if <italic>cat</italic> is the target word, we computed the cosine similarity between <inline-formula><mml:math id="M3"><mml:mover class="overrightarrow"><mml:mrow><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover></mml:math></inline-formula> from S[4] and <inline-formula><mml:math id="M4"><mml:mover class="overrightarrow"><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>3</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x020D7;</mml:mo></mml:mover></mml:math></inline-formula>. In order to find the optimal layer for the computation of the similarity scores, we extracted vectors from all the 24 layers of BERT Large and computed the Spearman correlations with each one of the target features.</p>
<p>The results can be seen in <xref ref-type="fig" rid="F1">Figure 1</xref>. Consistently with the findings of Salicchi et al. (<xref ref-type="bibr" rid="B68">2021</xref>), the layers with the highest absolute correlation values are the ones immediately before the last one. We chose layer 22 as the one with the highest inverse correlation to our data.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Spearman correlations between TRT (dot line), FFD (square), and NF (triangle) and the cosine similarity using vectors produced by different layers of BERT Large, on GECO <bold>(left)</bold> and Provo <bold>(right</bold>). Layer 24, whose values are systematically higher than the average, is intentionally left out for plot reading purposes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-14-1112365-g0001.tif"/>
</fig></sec>
<sec>
<title>3.4.3. Surprisal</title>
<p>To model the influence of word predictability on eye-tracking measures, we included in the regression models the surprisal of the target words given their previous context. For each target word we computed the surprisal as the negative logarithm of its probability given all the words preceding the target:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mo class="qopname">log</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The probability <italic>P</italic> is computed by GPT2-xl, the largest publicly available version of GPT-2. Similarly to the original model, GPT2-xl was also trained on the WebText corpus (40 GB of text data), but it has a larger architecture (48 layers, for a total of 1542M parameters) and was shown to have the lowest perplexity on the evaluation corpora of Radford et al. (<xref ref-type="bibr" rid="B61">2019</xref>).</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Results and discussion</title>
<sec>
<title>4.1. General analysis</title>
<sec>
<title>4.1.1. Cosine similarity vs. Spearman correlation</title>
<p>We first checked whether Spearman correlation was a better similarity metric than cosine with BERT contextual embeddings. Therefore, we compared BL-cos and BL-Spearman, namely models with baseline features and the similarity metric only, and we compared BL-sur-cos and BL-sur-Spearman, which are the models using baseline features, surprisal, and the similarity metric. The AICc values reported in <xref ref-type="table" rid="T2">Tables 2</xref>&#x02013;<xref ref-type="table" rid="T5">5</xref> clearly show that cosine similarity is a better predictor of eye-tracking features than Spearman correlation: on GECO, the difference between BL-cos and BL-Spearman is 1,279, and between BL-sur-cos and BL-sur-Spearman is 901; on Provo the differences are 333 and 318, respectively. Given these results, we henceforth focus our analyzes only on cosine similarity and its relationship with surprisal. Our findings suggest that, within the linear models we propose, BERT embeddings anisotropy does not affect the eye movements modeling, and therefore, cosine similarity is a suitable feature to be used for this eye tracking feature prediction task.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Average AICc, and AICc for TRT, FFD, and NF on GECO with SGNS vectors.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th/>
<th valign="top" align="center" colspan="2"><bold>Avg</bold></th>
<th valign="top" align="center" colspan="2"><bold>TRT</bold></th>
<th valign="top" align="center" colspan="2"><bold>FFD</bold></th>
<th valign="top" align="center" colspan="2"><bold>NF</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td valign="top" align="left"><bold>Model</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-cos</td>
<td valign="top" align="center">60,286</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">88,611</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">80,296</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">11,951</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">60,492</td>
<td valign="top" align="center">206</td>
<td valign="top" align="center">88,835</td>
<td valign="top" align="center">224</td>
<td valign="top" align="center">80,576</td>
<td valign="top" align="center">280</td>
<td valign="top" align="center">12,065</td>
<td valign="top" align="center">115</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">60,982</td>
<td valign="top" align="center">696</td>
<td valign="top" align="center">89,409</td>
<td valign="top" align="center">798</td>
<td valign="top" align="center">80,903</td>
<td valign="top" align="center">607</td>
<td valign="top" align="center">12,634</td>
<td valign="top" align="center">683</td>
</tr>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">61,466</td>
<td valign="top" align="center">1,180</td>
<td valign="top" align="center">89,948</td>
<td valign="top" align="center">1,337</td>
<td valign="top" align="center">81,483</td>
<td valign="top" align="center">1,186</td>
<td valign="top" align="center">12,969</td>
<td valign="top" align="center">1,018</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Average AICc, and AICc for TRT, FFD, and NF on GECO with BERT vectors.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th/>
<th valign="top" align="center" colspan="2"><bold>Avg</bold></th>
<th valign="top" align="center" colspan="2"><bold>TRT</bold></th>
<th valign="top" align="center" colspan="2"><bold>FFD</bold></th>
<th valign="top" align="center" colspan="2"><bold>NF</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td valign="top" align="left"><bold>Model</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-cos</td>
<td valign="top" align="center">59,566</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">87,758</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">79,232</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">11,709</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">60,151</td>
<td valign="top" align="center">585</td>
<td valign="top" align="center">88,413</td>
<td valign="top" align="center">654</td>
<td valign="top" align="center">79,697</td>
<td valign="top" align="center">465</td>
<td valign="top" align="center">12,346</td>
<td valign="top" align="center">637</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-Spearman</td>
<td valign="top" align="center">60,467</td>
<td valign="top" align="center">901</td>
<td valign="top" align="center">88,803</td>
<td valign="top" align="center">1,045</td>
<td valign="top" align="center">80,538</td>
<td valign="top" align="center">1,307</td>
<td valign="top" align="center">12,060</td>
<td valign="top" align="center">350</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">60,492</td>
<td valign="top" align="center">926</td>
<td valign="top" align="center">88,835</td>
<td valign="top" align="center">1,077</td>
<td valign="top" align="center">80,576</td>
<td valign="top" align="center">1,345</td>
<td valign="top" align="center">12,065</td>
<td valign="top" align="center">356</td>
</tr>
<tr>
<td valign="top" align="left">BL-Spearman</td>
<td valign="top" align="center">61,430</td>
<td valign="top" align="center">1,864</td>
<td valign="top" align="center">89,902</td>
<td valign="top" align="center">2,145</td>
<td valign="top" align="center">81,432</td>
<td valign="top" align="center">2,200</td>
<td valign="top" align="center">12,957</td>
<td valign="top" align="center">1,247</td>
</tr>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">61,466</td>
<td valign="top" align="center">1,900</td>
<td valign="top" align="center">89,948</td>
<td valign="top" align="center">2,190</td>
<td valign="top" align="center">81,483</td>
<td valign="top" align="center">2,251</td>
<td valign="top" align="center">12,969</td>
<td valign="top" align="center">1,259</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Average AICc, and AICc for TRT, FFD, and NF on Provo with SGNS vectors.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th/>
<th valign="top" align="center" colspan="2"><bold>Avg</bold></th>
<th valign="top" align="center" colspan="2"><bold>TRT</bold></th>
<th valign="top" align="center" colspan="2"><bold>FFD</bold></th>
<th valign="top" align="center" colspan="2"><bold>NF</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td valign="top" align="left"><bold>Model</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-cos</td>
<td valign="top" align="center">279</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1,309</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">288</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02013;762</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">391</td>
<td valign="top" align="center">112</td>
<td valign="top" align="center">1,436</td>
<td valign="top" align="center">127</td>
<td valign="top" align="center">441</td>
<td valign="top" align="center">153</td>
<td valign="top" align="center">&#x02013;704</td>
<td valign="top" align="center">58</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">437</td>
<td valign="top" align="center">158</td>
<td valign="top" align="center">1,468</td>
<td valign="top" align="center">159</td>
<td valign="top" align="center">406</td>
<td valign="top" align="center">118</td>
<td valign="top" align="center">&#x02013;594</td>
<td valign="top" align="center">168</td>
</tr>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">619</td>
<td valign="top" align="center">340</td>
<td valign="top" align="center">1,683</td>
<td valign="top" align="center">374</td>
<td valign="top" align="center">643</td>
<td valign="top" align="center">354</td>
<td valign="top" align="center">&#x02013;470</td>
<td valign="top" align="center">292</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Average AICc, and AICc for TRT, FFD, and NF on Provo with BERT vectors.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th/>
<th valign="top" align="center" colspan="2"><bold>Avg</bold></th>
<th valign="top" align="center" colspan="2"><bold>TRT</bold></th>
<th valign="top" align="center" colspan="2"><bold>FFD</bold></th>
<th valign="top" align="center" colspan="2"><bold>NF</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td valign="top" align="left"><bold>Model</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
<td valign="top" align="center"><bold>AICc</bold></td>
<td valign="top" align="center"><bold>Delta</bold></td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-cos</td>
<td valign="top" align="center">67</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1,081</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02013;88</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">&#x02013;791</td>
<td valign="top" align="center">0</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">196</td>
<td valign="top" align="center">129</td>
<td valign="top" align="center">1,216</td>
<td valign="top" align="center">135</td>
<td valign="top" align="center">&#x02013;0.26</td>
<td valign="top" align="center">87</td>
<td valign="top" align="center">&#x02013;627</td>
<td valign="top" align="center">165</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur-Spearman</td>
<td valign="top" align="center">385</td>
<td valign="top" align="center">318</td>
<td valign="top" align="center">1,429</td>
<td valign="top" align="center">348</td>
<td valign="top" align="center">434</td>
<td valign="top" align="center">521</td>
<td valign="top" align="center">&#x02013;707</td>
<td valign="top" align="center">85</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">391</td>
<td valign="top" align="center">324</td>
<td valign="top" align="center">1,436</td>
<td valign="top" align="center">355</td>
<td valign="top" align="center">441</td>
<td valign="top" align="center">529</td>
<td valign="top" align="center">&#x02013;704</td>
<td valign="top" align="center">88</td>
</tr>
<tr>
<td valign="top" align="left">BL-Spearman</td>
<td valign="top" align="center">529</td>
<td valign="top" align="center">462</td>
<td valign="top" align="center">1,674</td>
<td valign="top" align="center">593</td>
<td valign="top" align="center">633</td>
<td valign="top" align="center">721</td>
<td valign="top" align="center">&#x02013;474</td>
<td valign="top" align="center">315</td>
</tr>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">619</td>
<td valign="top" align="center">552</td>
<td valign="top" align="center">1,683</td>
<td valign="top" align="center">602</td>
<td valign="top" align="center">643</td>
<td valign="top" align="center">730</td>
<td valign="top" align="center">&#x02013;470</td>
<td valign="top" align="center">321</td>
</tr>
</tbody>
</table>
</table-wrap></sec>
<sec>
<title>4.1.2. Linear models comparison</title>
<p>For each implemented model, we used AICc values to determine which one was the best fit for the data. On both corpora, we notice that the best predictor of eye-tracking features is BL-sur-cos, including the interactions between baseline features, but with no interactions between cosine and surprisal. The fact that the regression model using both surprisal and cosine consistently performs better than the ones using only one of the two is strong evidence that they are both explanatory factors of reading times. Furthermore, while comparing BL-cos-sur with SGNS embeddings, and BL-cos-sur with BERT embeddings, it is possible to notice how the usage of the latter set of vectors improves the model (AICc values on GECO: 60,286 with SGNS-59,566 with BERT; AICc values on Provo: 279 with SGNS-67 with BERT).</p>
<p>Looking at the <italic>p</italic>-values of the regression features of our BL-sur-cos model, we observe that both cosine similarity and surprisal are statistically highly significant at <italic>p</italic> &#x0003C; 0.001 (for a complete analysis of regression features significance scores see <xref ref-type="supplementary-material" rid="SM1">Appendix 1</xref>). Although the combination of both cosine similarity and surprisal is the best performing model on both corpora, it is useful to focus also on the performances of BL-cos, and BL-sur while employing different vector models for BL-cos, to get further insights on the different contributions of surprisal and cosine similarity. We performed nested model comparisons with the R <italic>anova</italic> function using BL-sur-cos and three partial models: one excluding the cosine similarity (BL-sur), and the other two excluding surprisal (BL-cos with BERT vectors and BL-cos with SGNS vectors), in order to check whether the two features make independent contributions. We obtained strongly significant <italic>p</italic>-values (<italic>p</italic> &#x0003C; 0.001) on both corpora, regardless of vector type and for all the eye-tracking features, indicating that both semantic relatedness and surprisal provide an independent and significant contribution.</p>
<p>Focusing now on BL-cos and BL-sur, the performance on <bold>GECO</bold> is reported in <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref>. BL-cos with BERT vectors: Delta cosine similarity is 585, Delta surprisal is 926 (surprisal: &#x0002B;341) (<xref ref-type="table" rid="T3">Table 3</xref>); BL-cos with SGNS vectors: Delta surprisal is 206, Delta cosine similarity is 696 (surprisal: &#x02212;490) (<xref ref-type="table" rid="T2">Table 2</xref>); On <bold>Provo</bold> instead BL-cos with BERT vectors: Delta cosine similarity is 129, Delta surprisal is 324 (surprisal: &#x0002B;195) (<xref ref-type="table" rid="T5">Table 5</xref>); BL-cos with SGNS vectors: Delta surprisal is 112, Delta cosine is 158 (surprisal: &#x02212;46) (<xref ref-type="table" rid="T4">Table 4</xref>). This first analysis shows that BL-cos and BL-sur have <italic>quantitatively</italic> similar behavior, suggesting that cosine and surprisal help to predict eye-tracking values to the same extent. A difference in the salience of the two features is instead highlighted by the Part-of-Speech analysis (see the related subsection below).</p>
<p>It is also clear that models using SGNS vectors have poorer performances than the ones relying on BERT. Not only, as already mentioned, the usage of BERT embeddings improves the performances of the BL-cos-sur model, but while comparing the BL-cos models and the BL-sur model, the first shows better performances than the latter only when BERT vectors are involved. This difference in the capability of BL-cos models in predicting eye-tracking features suggests that the findings in Frank (<xref ref-type="bibr" rid="B18">2017</xref>) might be influenced by the specific type of embedding model used for the experiments (SGNS).</p>
<p>Once confirmed that the model including both surprisal and cosine similarity is the one performing better, we performed further analysis focused on BL, BL-sur, and BL-cos only, in order to understand the individual contribution of the two computational metrics.</p>
</sec>
<sec>
<title>4.1.3. Error analysis</title>
<p>In order to have a more fine-grained view of the performance differences between models BL-cos and BL-sur, we also analyzed the correlation between the Mean Absolute Error (MAE) of the models and word-level features. We tested the following features: target and previous word length, target and previous word frequency, target word length, target word position, fixation of the previous word (a boolean feature), and the reading complexity of the sentence from the beginning to the target word, which we computed using the Dale-Chall readability formula (Dale and Chall, <xref ref-type="bibr" rid="B12">1948</xref>).</p>
<p>After we averaged the correlations among all the eye-tracking features to be predicted (see <xref ref-type="supplementary-material" rid="SM1">Appendix 2</xref>) we noticed that almost all the values are negative, suggesting that: (i) longer and more frequent words are easier to be predicted; (ii) words at the beginning of the sentence are harder to predict for our models, plausibly because a wider and richer context benefits both cosine similarity and surprisal; (iii) sentences with higher readability make better predictions possible. Even so, the correlations between MAE and these features are generally low, ranging from 0.002 for previous word length to 0.1 for target word length. However, it is possible to use these values for a comparison between models BL-cos and BL-sur. We notice that surprisal seems to be more sensitive to target word frequency and previous word fixation if compared to cosine similarity, while the latter shows slightly higher correlations with target word length and position within the sentence.</p>
</sec>
<sec>
<title>4.1.4. POS analysis</title>
<p>Both GECO and Provo provide information regarding the part of speech (POS) of each word in the corpora. We used this information to check the performances of BL-cos and BL-sur on different POS. We first checked the average MAE of BL, BL-cos, and BL-sur for function words (pronouns, conjunctions, determiners, numeral, existential there&#x00027;s, prepositions, interjections) and content words (nouns, verbs, adverbs, adjectives) for each eye-tracking feature (<xref ref-type="table" rid="T6">Table 6</xref>). Then for a more detailed analysis, we ranked the words following the MAE values, and finally, we focused on the 10, 100, 500, and 1,000 words with the highest MAE.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Average MAE on Provo and GECO content and function words from models BL, BL-cos, and BL-sur for the three eye-tracking features and their mean.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th valign="top" align="left"><bold>Feature</bold></th>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center" colspan="4"><bold>Word type</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td/>
<td/>
<td valign="top" align="center" colspan="2"><bold>Content</bold></td>
<td valign="top" align="center" colspan="2"><bold>Function</bold></td>
</tr>
<tr style="background-color:#919497; color:#ffffff;">
<td/>
<td/>
<td valign="top" align="center"><bold>Provo</bold></td>
<td valign="top" align="center"><bold>GECO</bold></td>
<td valign="top" align="center"><bold>Provo</bold></td>
<td valign="top" align="center"><bold>GECO</bold></td>
</tr>
<tr>
<td valign="top" align="left">TRT</td>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">0.228</td>
<td valign="top" align="center">0.337</td>
<td valign="top" align="center">0.290</td>
<td valign="top" align="center">0.457</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">0.217</td>
<td valign="top" align="center">0.333</td>
<td valign="top" align="center">0.281</td>
<td valign="top" align="center">0.457</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">0.215</td>
<td valign="top" align="center">0.330</td>
<td valign="top" align="center">0.275</td>
<td valign="top" align="center">0.454</td>
</tr>
<tr>
<td valign="top" align="left">FFD</td>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">0.180</td>
<td valign="top" align="center">0.295</td>
<td valign="top" align="center">0.246</td>
<td valign="top" align="center">0.425</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">0.159</td>
<td valign="top" align="center">0.281</td>
<td valign="top" align="center">0.216</td>
<td valign="top" align="center">0.422</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">0.172</td>
<td valign="top" align="center">0.289</td>
<td valign="top" align="center">0.236</td>
<td valign="top" align="center">0.423</td>
</tr>
<tr>
<td valign="top" align="left">NF</td>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">0.178</td>
<td valign="top" align="center">0.228</td>
<td valign="top" align="center">0.147</td>
<td valign="top" align="center">0.187</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">0.177</td>
<td valign="top" align="center">0.228</td>
<td valign="top" align="center">0.132</td>
<td valign="top" align="center">0.184</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">0.170</td>
<td valign="top" align="center">0.226</td>
<td valign="top" align="center">0.140</td>
<td valign="top" align="center">0.185</td>
</tr>
<tr>
<td valign="top" align="left">Avg</td>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">0.195</td>
<td valign="top" align="center">0.287</td>
<td valign="top" align="center">0.228</td>
<td valign="top" align="center">0.356</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center"><bold>0.185</bold></td>
<td valign="top" align="center"><bold>0.281</bold></td>
<td valign="top" align="center"><bold>0.210</bold></td>
<td valign="top" align="center">0.354</td>
</tr>
 <tr>
<td/>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center">0.186</td>
<td valign="top" align="center">0.282</td>
<td valign="top" align="center">0.217</td>
<td valign="top" align="center">0.354</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The bold formatting indicates the lowest MAE averaged over the 3 eye tracking features.</p>
</table-wrap-foot>
</table-wrap>
<p>We found that for all three models function words are harder to be predicted than content words, especially coordinating conjunctions and pronouns. Noticeably, previous research had already found that the semantics of function words is difficult to model even for Transformers (Kim et al., <xref ref-type="bibr" rid="B38">2019</xref>), and that fine-tuned multilingual Transformer model struggle the most with the prediction of their fixation metrics (Hollenstein et al., <xref ref-type="bibr" rid="B30">2022b</xref>). Regarding the performances of BL-cos and BL-sur, even if both cosine similarity and surprisal help in lowering the average MAE, if compared to the baseline, cosine similarity employment improves slightly more the performance of the model for both content words and function words.</p>
</sec>
</sec>
<sec>
<title>4.2. Eye-tracking features analysis</title>
<p>While comparing the different models, it was clear that some performance differences were due to the eye-tracking feature the models had to predict. For example, the data showed in the Avg column of <xref ref-type="table" rid="T2">Tables 2</xref>&#x02013;<xref ref-type="table" rid="T5">5</xref> are mean values computed using the AICc scores of TRT, FFD, and NF, but if we focus on the performances of models BL-cos and BL-sur, depending on the target eye-tracking features, we notice some interesting and substantial differences: on TRT cosine similarity-only and surprisal-only models follow the general tendency we described in Section 4.1 (i.e., surprisal better than cosine similarity when BL-cos makes use of SGNS vectors to compute cosine), but with cosine similarity performing generally slightly better than surprisal; on FFD the model using baseline regression features and cosine similarity only performs consistently better, except when using SGNS on GECO (but not on Provo), while on NF model BL-sur outperforms BL-cos on both corpora, even when using BERT vectors in BL-cos.</p>
<p>In the analysis of the correlations between models MAE and word features, we found that for TRT and FFD, the highest correlation (especially on GECO) is the one between MAE and the word length. Since it is a negative correlation, we can conclude that shorter words induce higher MAE: The shorter the word, the harder for the model to predict the feature value. On the other hand, with NF, word length has the highest, but <italic>positive</italic>, correlation with the MAE, thus suggesting that for this eye-tracking feature shorter words are easier to be predicted. Finally, for all the eye-tracking features on both corpora, word frequency is negatively correlated. As expected, prediction is more difficult for the rarest words.</p>
<p>When we checked the contribution of BL-cos and BL-sur in comparison to the baseline for different parts of speech, we noticed that for FFD cosine similarity generally decreases the MAE, while for TRT surprisal gives a generally higher contribution, except for verbs and adjectives (<xref ref-type="table" rid="T7">Tables 7</xref>, <xref ref-type="table" rid="T8">8</xref>). Regarding NF, cosine similarity lowers the MAE for function words, while surprisal has a major impact on content words. However, for the NF feature content words are less easily predicted.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Average MAE on Provo content words.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center" colspan="4"><bold>TRT</bold></th>
<th valign="top" align="center" colspan="4"><bold>FFD</bold></th>
<th valign="top" align="center" colspan="4"><bold>NF</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td/>
<td valign="top" align="center"><bold>N</bold></td>
<td valign="top" align="center"><bold>RB</bold></td>
<td valign="top" align="center"><bold>V</bold></td>
<td valign="top" align="center"><bold>J</bold></td>
<td valign="top" align="center"><bold>N</bold></td>
<td valign="top" align="center"><bold>RB</bold></td>
<td valign="top" align="center"><bold>V</bold></td>
<td valign="top" align="center"><bold>J</bold></td>
<td valign="top" align="center"><bold>N</bold></td>
<td valign="top" align="center"><bold>RB</bold></td>
<td valign="top" align="center"><bold>V</bold></td>
<td valign="top" align="center"><bold>J</bold></td>
</tr>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">0.243</td>
<td valign="top" align="center">0.242</td>
<td valign="top" align="center">0.210</td>
<td valign="top" align="center">0.209</td>
<td valign="top" align="center">0.190</td>
<td valign="top" align="center">0.184</td>
<td valign="top" align="center">0.171</td>
<td valign="top" align="center">0.166</td>
<td valign="top" align="center">0.195</td>
<td valign="top" align="center">0.180</td>
<td valign="top" align="center">0.152</td>
<td valign="top" align="center">0.180</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">0.231</td>
<td valign="top" align="center">0.243</td>
<td valign="top" align="center"><bold>0.198</bold></td>
<td valign="top" align="center"><bold>0.199</bold></td>
<td valign="top" align="center"><bold>0.178</bold></td>
<td valign="top" align="center">0.184</td>
<td valign="top" align="center"><bold>0.157</bold></td>
<td valign="top" align="center"><bold>0.154</bold></td>
<td valign="top" align="center">0.192</td>
<td valign="top" align="center">0.178</td>
<td valign="top" align="center"><bold>0.149</bold></td>
<td valign="top" align="center">0.179</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center"><bold>0.228</bold></td>
<td valign="top" align="center"><bold>0.221</bold></td>
<td valign="top" align="center">0.200</td>
<td valign="top" align="center">0.205</td>
<td valign="top" align="center">0.181</td>
<td valign="top" align="center"><bold>0.176</bold></td>
<td valign="top" align="center">0.163</td>
<td valign="top" align="center">0.163</td>
<td valign="top" align="center"><bold>0.183</bold></td>
<td valign="top" align="center"><bold>0.171</bold></td>
<td valign="top" align="center">0.150</td>
<td valign="top" align="center"><bold>0.178</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>N, nouns; RB, adverbs; V, verbs; J, adjectives-for the three eye-tracking features. The bold formatting indicates the values with the lowest MAE of each POS within each eye-tracking feature.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>Average MAE on GECO content words.</p></caption>
<table frame="box" rules="all">
<thead><tr style="background-color:#919497; color:#ffffff;">
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center" colspan="4"><bold>TRT</bold></th>
<th valign="top" align="center" colspan="4"><bold>FFD</bold></th>
<th valign="top" align="center" colspan="4"><bold>NF</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919497; color:#ffffff;">
<td/>
<td valign="top" align="center"><bold>N</bold></td>
<td valign="top" align="center"><bold>RB</bold></td>
<td valign="top" align="center"><bold>V</bold></td>
<td valign="top" align="center"><bold>J</bold></td>
<td valign="top" align="center"><bold>N</bold></td>
<td valign="top" align="center"><bold>RB</bold></td>
<td valign="top" align="center"><bold>V</bold></td>
<td valign="top" align="center"><bold>J</bold></td>
<td valign="top" align="center"><bold>N</bold></td>
<td valign="top" align="center"><bold>RB</bold></td>
<td valign="top" align="center"><bold>V</bold></td>
<td valign="top" align="center"><bold>J</bold></td>
</tr>
<tr>
<td valign="top" align="left">BL</td>
<td valign="top" align="center">0.335</td>
<td valign="top" align="center">0.365</td>
<td valign="top" align="center">0.334</td>
<td valign="top" align="center">0.309</td>
<td valign="top" align="center">0.289</td>
<td valign="top" align="center">0.322</td>
<td valign="top" align="center">0.294</td>
<td valign="top" align="center">0.273</td>
<td valign="top" align="center">0.238</td>
<td valign="top" align="center">0.226</td>
<td valign="top" align="center">0.217</td>
<td valign="top" align="center">0.242</td>
</tr>
<tr>
<td valign="top" align="left">BL-cos</td>
<td valign="top" align="center">0.328</td>
<td valign="top" align="center">0.367</td>
<td valign="top" align="center">0.332</td>
<td valign="top" align="center">0.301</td>
<td valign="top" align="center">0.280</td>
<td valign="top" align="center">0.323</td>
<td valign="top" align="center">0.292</td>
<td valign="top" align="center">0.264</td>
<td valign="top" align="center">0.237</td>
<td valign="top" align="center">0.226</td>
<td valign="top" align="center">0.217</td>
<td valign="top" align="center">0.241</td>
</tr>
<tr>
<td valign="top" align="left">BL-sur</td>
<td valign="top" align="center"><bold>0.323</bold></td>
<td valign="top" align="center"><bold>0.360</bold></td>
<td valign="top" align="center">0.332</td>
<td valign="top" align="center"><bold>0.299</bold></td>
<td valign="top" align="center">0.280</td>
<td valign="top" align="center"><bold>0.320</bold></td>
<td valign="top" align="center">0.292</td>
<td valign="top" align="center"><bold>0.262</bold></td>
<td valign="top" align="center"><bold>0.234</bold></td>
<td valign="top" align="center"><bold>0.225</bold></td>
<td valign="top" align="center"><bold>0.216</bold></td>
<td valign="top" align="center"><bold>0.240</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>N, nouns; RB, adverbs; V, verbs; J, adjectives-for the three eye-tracking features. The bold formatting indicates the values with the lowest MAE of each POS within each eye-tracking feature.</p>
</table-wrap-foot>
</table-wrap>
<p>We surmise that the different performances of BL-sur and BL-cos in predicting these three eye-tracking features might be explained by taking into account the reading process stage each feature is related to. On one hand, since FFD is typically associated with early stages of reading, such as lexical information process, it is not surprising that the model relying on semantic relatedness between the context and the target word performs better. On the other hand, the performances of BL-cos and BL-sur on TRT and NF, features that reflect later stages of the reading process, including information-structural integration, may suggest that predictability is a key factor in handling syntagmatic relations and integrating semantic and syntactic information.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>In this paper, we implemented four different kinds of regression models to predict three eye-tracking features of two corpora collecting eye movements data, with the aim of investigating the role and interplay between distributional measures of target-context semantic relatedness, and target surprisal, as computed with a state-of-the-art neural language model. The main research question was whether semantic relatedness is indeed made redundant by surprisal, as argued by Frank (<xref ref-type="bibr" rid="B18">2017</xref>), or instead plays an independent role in explaining eye-tracking data. The models include: (i) a baseline with word-level features, (ii) the same baseline with cosine similarity, (iii) the baseline with surprisal, iv) the baseline with both cosine similarity and surprisal.</p>
<p>Our results show that the complete model systematically outperforms the others for every eye-tracking feature and that both semantic relatedness and surprisal benefit the prediction of eye-tracking features, given the performance drop while factoring one of them out. Surprisal and distributional semantic relatedness clearly overlap, especially since the latter is nowadays commonly computed using word embeddings produced by DSMs trained with a prediction objective, like the one that surprisal formalizes. Yet, they capture different linguistic dimensions. Surprisal models the <italic>syntagmatic</italic> predictability of a word, given the preceding ones. On the other hand, both static and contextual DSMs use prediction as a distributional signal to form internal representations of lexical meaning that capture information more directly pertaining to the <italic>paradigmatic</italic> dimension, such as belonging to the same semantic classes and domains or sharing similar features. For instance, the words <italic>pie</italic> and <italic>cake</italic> are paradigmatically related because they share several salient attributes, such as being edible, sweet, etc. (Chersoni et al., <xref ref-type="bibr" rid="B9">2021</xref>) showed that word embeddings encode a vast range of linguistically and cognitively relevant semantic features. Therefore, the results of our analyzes suggest that, despite their overlap, corpus-based semantic relatedness and surprisal capture different dimensions that play an autonomous role during reading. While surprisal reflects how predictable the target word is from the previous context, semantic relatedness models how coherent the meaning of the target is with respect to the context one (e.g., they belong to the same semantic field or describe a prototypical situation). Frank and Willems (<xref ref-type="bibr" rid="B20">2017</xref>) found that syntagmatic surprisal and paradigmatic semantic relatedness can have neurally distinguishable effects during language comprehension. Our analyzes show that their independent effect can be detected in eye-tracking data too.</p>
<p>We also analyzed whether the relatedness and surprisal have a differential effect depending on the target part-of-speech. Comparing the average MAE of our models, we noticed that surprisal mainly helps to improve the model&#x00027;s performances on content words, while the contribution of semantic relatedness includes function words as well. Finally, we investigated whether the interplay between surprisal and relatedness is affected by the type of word embeddings used to compute the latter, in particular considering the difference between static DSMs (SGNS) and contextual ones (BERT). The experiments show that when using BERT vectors, which are inherently able to account for context-dependent meaning shifts and carry out an implicit form of word-sense disambiguation, the model <bold>BL-cos</bold> performs better than <bold>BL-sur</bold>, while static vectors make the latter outrank the model using semantic relatedness only. Overall, our findings suggest that the kind of word embedding employed for computing vector distances has a significant impact, which may explain the differences from the findings by Frank (<xref ref-type="bibr" rid="B18">2017</xref>).</p>
<p>The present work admittedly has some limitations. For example, we employed and compared a restricted pool of language models and word embedding models, and a possible future direction could be testing other, more recent models (e.g., XLNet Yang et al. <xref ref-type="bibr" rid="B84">2019</xref>, among others, RoBERTa Liu et al., <xref ref-type="bibr" rid="B45">2019</xref>), or different static embedding models (e.g., GloVe Pennington et al., <xref ref-type="bibr" rid="B55">2014</xref>, FastText Bojanowski et al., <xref ref-type="bibr" rid="B5">2017</xref>). A particularly interesting issue, raised by some recent works, is the relationship between the size of a language model and its capacity to model human behavioral data (Oh and Schuler, <xref ref-type="bibr" rid="B53">2022</xref>; Shain et al., <xref ref-type="bibr" rid="B74">2022</xref>). In particular, Oh and Schuler (<xref ref-type="bibr" rid="B53">2022</xref>) found that larger language models are worse at predicting human reading times: larger models tend to be less surprised by open-class words because they have been trained on many more word sequences than those available to humans. Moreover, phenomena of inverse scaling have also been reported for language modeling of negations (Jang et al., <xref ref-type="bibr" rid="B34">2022</xref>) and quantifiers (Kalouli et al., <xref ref-type="bibr" rid="B36">2022</xref>; Michaelov and Bergen, <xref ref-type="bibr" rid="B49">2022</xref>). It might be worth testing whether this increasing lack of alignment with human performance as scale increases can be observed also at the level of similarity estimation with the embeddings, or it is an effect limited to language model predictions. With this purpose, it could be interesting to compare embedding models of different size with BERT, and see if there are differences in modeling open class vs. function words.</p>
<p>Another limitation is due to the fact that we used English materials only, and this leaves open the question whether our results would apply to other languages. An interesting research path to pursue is to compare models with cosine similarity and surprisal using multilingual data. In fact, we plan to extend our analyzes to the recently-published MECO corpus (Siegelman et al., <xref ref-type="bibr" rid="B75">2022</xref>), which provides eye-tracking data on comparable texts for 13 different languages.</p>
<p>Finally, if the importance and independence of surprisal and semantic relatedness are clear, given the results shown in the present paper, a preliminary feature importance analysis using a random forest regression model (see <xref ref-type="supplementary-material" rid="SM1">Appendix 3</xref>) revealed how target and previous word lengths are the features with the higher impact, and most importantly, surprisal systematically seems to have a larger effect on the model compared to cosine similarity. These preliminary results suggest one further possible research direction: the employment and comparison of different models and a consequent feature importance analysis, in order to find even more generalizable insights regarding the role of semantic relatedness and predictability in the reading process.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>AL and EC contributed to the conception and design of the study. LS was responsible for the coding part, the data analysis, and the creation of the first draft of the manuscript. AL, EC, and LS contributed equally to the final form of the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This project was supported by the COnversational BRAins (CoBra) European Training Network (H-ZG9X). EC was supported by the Startup Fund (1-BD8S) by the Hong Kong Polytechnic University.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1112365/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1112365/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://expsy.ugent.be/downloads/geco/">https://expsy.ugent.be/downloads/geco/</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://osf.io/sjefs/">https://osf.io/sjefs/</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup>The Wikipedia frequencies were extracted from <ext-link ext-link-type="uri" xlink:href="https://github.com/IlyaSemenov/wikipedia-word-frequency">https://github.com/IlyaSemenov/wikipedia-word-frequency</ext-link></p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aurnhammer</surname> <given-names>C.</given-names></name> <name><surname>Frank</surname> <given-names>S. L.</given-names></name></person-group> (<year>2019</year>). <article-title>Evaluating information-theoretic measures of word prediction in naturalistic sentence reading</article-title>. <source>Neuropsychologia</source>. <volume>134</volume>, <fpage>107198</fpage> <pub-id pub-id-type="doi">10.1016/j.neuropsychologia.2019.107198</pub-id><pub-id pub-id-type="pmid">31553896</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bahdanau</surname> <given-names>D.</given-names></name> <name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>Neural Machine Translation by Jointly Learning to Align and Translate</article-title>. <source>arXiv preprint arXiv</source>:1409.0473.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balota</surname> <given-names>D. A.</given-names></name> <name><surname>Chumbley</surname> <given-names>J. I.</given-names></name></person-group> (<year>1984</year>). <article-title>Are lexical decisions a good measure of lexical access? The role of word frequency in the neglected decision stage</article-title>. <source>J. Exp. Psychol</source>. <volume>10</volume>, <fpage>340</fpage>. <pub-id pub-id-type="doi">10.1037/0096-1523.10.3.340</pub-id><pub-id pub-id-type="pmid">6242411</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baroni</surname> <given-names>M.</given-names></name> <name><surname>Lenci</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Distributional memory: a general framework for corpus-based semantics</article-title>. <source>Comput. Linguist</source>. <volume>36</volume>, <fpage>673</fpage>&#x02013;<lpage>721</lpage>. <pub-id pub-id-type="doi">10.1162/coli_a_00016</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <name><surname>Grave</surname> <given-names>E.</given-names></name> <name><surname>Joulin</surname> <given-names>A.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Enriching word vectors with subword information</article-title>. <source>Trans. Assoc. Computat. Linguist</source>. <volume>5</volume>, <fpage>135</fpage>&#x02013;<lpage>146</lpage>. <pub-id pub-id-type="doi">10.1162/tacl_a_00051</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bommasani</surname> <given-names>R.</given-names></name> <name><surname>Davis</surname> <given-names>K.</given-names></name> <name><surname>Cardie</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Interpreting pretrained contextualized representations via reductions to static embeddings,&#x0201D;</article-title> in <source>Proceedings of ACL</source>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>T.</given-names></name> <name><surname>Mann</surname> <given-names>B.</given-names></name> <name><surname>Ryder</surname> <given-names>N.</given-names></name> <name><surname>Subbiah</surname> <given-names>M.</given-names></name> <name><surname>Kaplan</surname> <given-names>J. D.</given-names></name> <name><surname>Dhariwal</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;Language models are few-shot learners,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 33</source>, <fpage>1877</fpage>&#x02013;<lpage>1901</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bullinaria</surname> <given-names>J. A.</given-names></name> <name><surname>Levy</surname> <given-names>J. P.</given-names></name></person-group> (<year>2012</year>). <article-title>Extracting semantic representations from word co-occurrence statistics: stop-lists, stemming, and SVD</article-title>. <source>Behav. Res. Methods</source> <volume>44</volume>, <fpage>890</fpage>&#x02013;<lpage>907</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-011-0183-8</pub-id><pub-id pub-id-type="pmid">22258891</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chersoni</surname> <given-names>E.</given-names></name> <name><surname>Santus</surname> <given-names>E.</given-names></name> <name><surname>Huang</surname> <given-names>C.-R.</given-names></name> <name><surname>Lenci</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Decoding word embeddings with brain-based semantic features</article-title>. <source>Comput. Linguist</source>. <volume>47</volume>, <fpage>663</fpage>&#x02013;<lpage>698</lpage>. <pub-id pub-id-type="doi">10.1162/coli_a_00412</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chronis</surname> <given-names>G.</given-names></name> <name><surname>Erk</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;When is a bishop not like a rook? When it&#x00027;s like a rabbi! multi-prototype BERT embeddings for estimating semantic relationships,&#x0201D;</article-title> in <source>Proceedings of CONLL</source>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cop</surname> <given-names>U.</given-names></name> <name><surname>Dirix</surname> <given-names>N.</given-names></name> <name><surname>Drieghe</surname> <given-names>D.</given-names></name> <name><surname>Duyck</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>Presenting GECO: an eye-tracking corpus of monolingual and bilingual sentence reading</article-title>. <source>Behav. Re. Methods</source> <volume>49</volume>, <fpage>602</fpage>&#x02013;<lpage>615</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-016-0734-0</pub-id><pub-id pub-id-type="pmid">27193157</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dale</surname> <given-names>E.</given-names></name> <name><surname>Chall</surname> <given-names>J. S.</given-names></name></person-group> (<year>1948</year>). <article-title>A formula for predicting readability: instructions</article-title>. <source>Educ. Res. Bull</source>. <volume>27</volume>, <fpage>37</fpage>&#x02013;<lpage>54</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Demberg</surname> <given-names>V.</given-names></name> <name><surname>Keller</surname> <given-names>F.</given-names></name></person-group> (<year>2008</year>). <article-title>Data from eye-tracking corpora as evidence for theories of syntactic processing complexity</article-title>. <source>Cognition</source> <volume>109</volume>, <fpage>193</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2008.07.008</pub-id><pub-id pub-id-type="pmid">18930455</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;BERT: pre-training of deep bidirectional transformers for language understanding,&#x0201D;</article-title> in <source>Proceedings of NAACL</source> (<publisher-loc>Minneapolis, MN</publisher-loc>).</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ehrlich</surname> <given-names>S. E.</given-names></name> <name><surname>Rayner</surname> <given-names>K.</given-names></name></person-group> (<year>1981</year>). <article-title>Contextual effects on word perception and eye movements during reading</article-title>. <source>J. Verbal Learn. Verbal Behav</source>. <volume>20</volume>, <fpage>641</fpage>&#x02013;<lpage>665</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-5371(81)90220-6</pub-id><pub-id pub-id-type="pmid">28333501</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Evert</surname> <given-names>S.</given-names></name></person-group> (<year>2005</year>). <source>The Statistics of Word Cooccurrences: Word Pairs and Collocations</source> (Ph.D. thesis). University of Stuttgart.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fossum</surname> <given-names>V.</given-names></name> <name><surname>Levy</surname> <given-names>R.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Sequential vs. hierarchical syntactic models of human incremental sentence processing,&#x0201D;</article-title> in <source>Proceedings of the NAACL Workshop on Cognitive Modeling and Computational Linguistics</source> (<publisher-loc>Montreal, QC</publisher-loc>).</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Frank</surname> <given-names>S. L.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Word embedding distance does not predict word reading time,&#x0201D;</article-title> in <source>Proceedings of CogSci</source> (<publisher-loc>London</publisher-loc>).</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frank</surname> <given-names>S. L.</given-names></name> <name><surname>Bod</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Insensitivity of the human sentence-processing system to hierarchical structure</article-title>. <source>Psychol. Sci</source>. <volume>22</volume>, <fpage>829</fpage>&#x02013;<lpage>834</lpage>. <pub-id pub-id-type="doi">10.1177/0956797611409589</pub-id><pub-id pub-id-type="pmid">21586764</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frank</surname> <given-names>S. L.</given-names></name> <name><surname>Willems</surname> <given-names>R. M.</given-names></name></person-group> (<year>2017</year>). <article-title>Word predictability and semantic similarity show distinct patterns of brain activity during language comprehension</article-title>. <source>Lang. Cogn. Neurosci</source>. <volume>32</volume>, <fpage>1192</fpage>&#x02013;<lpage>1203</lpage>. <pub-id pub-id-type="doi">10.1080/23273798.2017.1323109</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frazier</surname> <given-names>L.</given-names></name> <name><surname>Rayner</surname> <given-names>K.</given-names></name></person-group> (<year>1982</year>). <article-title>Making and correcting errors during sentence comprehension: eye movements in the analysis of structurally ambiguous sentences</article-title>. <source>Cogn. Psychol</source>. <volume>14</volume>, <fpage>178</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1016/0010-0285(82)90008-1</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Goodkind</surname> <given-names>A.</given-names></name> <name><surname>Bicknell</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Predictive power of word surprisal for reading times is a linear function of language model quality,&#x0201D;</article-title> in <source>Proceedings of the LSA Workshop on Cognitive Modeling and Computational Linguistics</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>).</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goodkind</surname> <given-names>A.</given-names></name> <name><surname>Bicknell</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>Local word statistics affect reading times independently of surprisal</article-title>. <source>arXiv preprint</source> arXiv:2103.04469. <pub-id pub-id-type="doi">10.48550/arXiv.2103.04469</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gordon</surname> <given-names>P. C.</given-names></name> <name><surname>Hendrick</surname> <given-names>R.</given-names></name> <name><surname>Johnson</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>Y.</given-names></name></person-group> (<year>2006</year>). <article-title>Similarity-based interference during language comprehension: evidence from eye tracking during reading</article-title>. <source>J. Exp. Psychol. Learn. Mem. Cogn</source>. <volume>32</volume>, <fpage>1304</fpage>. <pub-id pub-id-type="doi">10.1037/0278-7393.32.6.1304</pub-id><pub-id pub-id-type="pmid">17087585</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hale</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>&#x0201C;A probabilistic earley parser as a psycholinguistic model,&#x0201D;</article-title> in <source>Proceedings of NAACL</source> (<publisher-loc>Pittsburgh, PA</publisher-loc>).<pub-id pub-id-type="pmid">17662975</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hale</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Information-theoretical complexity metrics</article-title>. <source>Lang. Linguist. Compass</source>. <volume>10</volume>, <fpage>397</fpage>&#x02013;<lpage>412</lpage>. <pub-id pub-id-type="doi">10.1111/lnc3.12196</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hao</surname> <given-names>Y.</given-names></name> <name><surname>Mendelsohn</surname> <given-names>S.</given-names></name> <name><surname>Sterneck</surname> <given-names>R.</given-names></name> <name><surname>Martinez</surname> <given-names>R.</given-names></name> <name><surname>Frank</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Probabilistic predictions of people perusing: evaluating metrics of language model performance for psycholinguistic modeling,&#x0201D;</article-title> in <source>Proceedings of the EMNLP Workshop on Cognitive Modeling and Computational Linguistics</source>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hofmann</surname> <given-names>M. J.</given-names></name> <name><surname>Remus</surname> <given-names>S.</given-names></name> <name><surname>Biemann</surname> <given-names>C.</given-names></name> <name><surname>Radach</surname> <given-names>R.</given-names></name> <name><surname>Kuchinke</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Language models explain word reading times better than empirical predictability</article-title>. <source>Fronti. Artif. Intell</source>. <volume>4</volume>, <fpage>730570</fpage>. <pub-id pub-id-type="doi">10.3389/frai.2021.730570</pub-id><pub-id pub-id-type="pmid">35187472</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hollenstein</surname> <given-names>N.</given-names></name> <name><surname>Chersoni</surname> <given-names>E.</given-names></name> <name><surname>Jacobs</surname> <given-names>C. L.</given-names></name> <name><surname>Oseki</surname> <given-names>Y.</given-names></name> <name><surname>Pr&#x000E9;vot</surname> <given-names>L.</given-names></name> <name><surname>Santus</surname> <given-names>E.</given-names></name></person-group> (<year>2022a</year>). <article-title>&#x0201C;CMCL 2022 shared task on multilingual and crosslingual prediction of human reading behavior,&#x0201D;</article-title> in <source>Proceedings of the ACL Workshop on Cognitive Modeling and Computational Linguistics</source> (<publisher-loc>Dublin</publisher-loc>).</citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hollenstein</surname> <given-names>N.</given-names></name> <name><surname>Gonzalez-Dios</surname> <given-names>I.</given-names></name> <name><surname>Beinborn</surname> <given-names>L.</given-names></name> <name><surname>Jaeger</surname> <given-names>L.</given-names></name></person-group> (<year>2022b</year>). <article-title>&#x0201C;Patterns of text readability in human and predicted eye movements,&#x0201D;</article-title> in <source>Proceedings of the AACL Workshop on Cognitive Aspects of the Lexicon</source> (<publisher-loc>Taipei</publisher-loc>).<pub-id pub-id-type="pmid">28151950</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollenstein</surname> <given-names>N.</given-names></name> <name><surname>Pirovano</surname> <given-names>F.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>J&#x000E4;ger</surname> <given-names>L.</given-names></name> <name><surname>Beinborn</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Multilingual language models predict human reading behavior,&#x0201D;</article-title> in <source>Proceedings of NAACL</source>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Inhoff</surname> <given-names>A. W.</given-names></name></person-group> (<year>1984</year>). <article-title>Two stages of word processing during eye fixations in the reading of prose</article-title>. <source>J. Verbal Learn. Verbal Behav</source>. <volume>23</volume>, <fpage>612</fpage>&#x02013;<lpage>624</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-5371(84)90382-7</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Inhoff</surname> <given-names>A. W.</given-names></name> <name><surname>Radach</surname> <given-names>R.</given-names></name></person-group> (<year>1998</year>). <article-title>&#x0201C;Definition and computation of oculomotor measures in the study of cognitive processes,&#x0201D;</article-title> in <source>Eye Guidance in Reading and Scene Perception</source>, <fpage>29</fpage>&#x02013;<lpage>53</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jang</surname> <given-names>J.</given-names></name> <name><surname>Ye</surname> <given-names>S.</given-names></name> <name><surname>Seo</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Can large language models truly understand prompts? a case study with negated prompts</article-title>. <source>arXiv preprint</source> arXiv:2209.12711. <pub-id pub-id-type="doi">10.48550/arXiv.2209.12711</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Just</surname> <given-names>M. A.</given-names></name> <name><surname>Carpenter</surname> <given-names>P. A.</given-names></name></person-group> (<year>1980</year>). <article-title>A theory of reading: From eye fixations to comprehension</article-title>. <source>Psychol. Rev</source>. <volume>87</volume>, <fpage>329</fpage>&#x02013;<lpage>354</lpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.87.4.329</pub-id><pub-id pub-id-type="pmid">7413885</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kalouli</surname> <given-names>A.-L.</given-names></name> <name><surname>Sevastjanova</surname> <given-names>R.</given-names></name> <name><surname>Beck</surname> <given-names>C.</given-names></name> <name><surname>Romero</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Negation, coordination, and quantifiers in contextualized language models,&#x0201D;</article-title> in <source>Proceedings of COLING</source> (<publisher-loc>Gyeongju</publisher-loc>).</citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kennedy</surname> <given-names>A.</given-names></name> <name><surname>Hill</surname> <given-names>R.</given-names></name> <name><surname>Pynte</surname> <given-names>J.</given-names></name></person-group> (<year>2003</year>). <article-title>&#x0201C;The dundee corpus,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Eye Movement</source> (<publisher-loc>Dundee</publisher-loc>).</citation>
</ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>N.</given-names></name> <name><surname>Patel</surname> <given-names>R.</given-names></name> <name><surname>Poliak</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>A.</given-names></name> <name><surname>Xia</surname> <given-names>P.</given-names></name> <name><surname>McCoy</surname> <given-names>R. T.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Probing what different NLP tasks teach machines about function word comprehension,&#x0201D;</article-title> in <source>Proceedings of &#x0002A;SEM</source> (<publisher-loc>Minneapolis, MN</publisher-loc>).</citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kovaleva</surname> <given-names>O.</given-names></name> <name><surname>Kulshreshtha</surname> <given-names>S.</given-names></name> <name><surname>Rogers</surname> <given-names>A.</given-names></name> <name><surname>Rumshisky</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;BERT busters: outlier dimensions that disrupt transformers,&#x0201D;</article-title> in <source>Findings of ACL</source>.</citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Landauer</surname> <given-names>T. K.</given-names></name> <name><surname>Dumais</surname> <given-names>S. T.</given-names></name></person-group> (<year>1997</year>). <article-title>A solution to plato&#x00027;s problem: the latent semantic analysis theory of acquisition, induction, and representation of knowledge</article-title>. <source>Psychol. Rev</source>. <volume>104</volume>, <fpage>211</fpage>. <pub-id pub-id-type="doi">10.1037/0033-295X.104.2.211</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lenci</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Distributional models of word meaning</article-title>. <source>Ann. Rev. Linguist</source>. <volume>4</volume>, <fpage>151</fpage>&#x02013;<lpage>171</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-linguistics-030514-125254</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lenci</surname> <given-names>A.</given-names></name> <name><surname>Sahlgren</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <source>Distributional Semantics</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lenci</surname> <given-names>A.</given-names></name> <name><surname>Sahlgren</surname> <given-names>M.</given-names></name> <name><surname>Jeuniaux</surname> <given-names>P.</given-names></name> <name><surname>Gyllensten</surname> <given-names>A. C.</given-names></name> <name><surname>Miliani</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>A comprehensive comparative evaluation and analysis of distributional semantic models</article-title>. <source>Lang. Resour. Evaluat</source>. <volume>56</volume>, <fpage>1269</fpage>&#x02013;<lpage>1313</lpage>. <pub-id pub-id-type="doi">10.1007/s10579-021-09575-z</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levy</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Expectation-based syntactic comprehension</article-title>. <source>Cognition</source> <volume>106</volume>, <fpage>1126</fpage>&#x02013;<lpage>1177</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2007.05.006</pub-id><pub-id pub-id-type="pmid">17662975</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Ott</surname> <given-names>M.</given-names></name> <name><surname>Goyal</surname> <given-names>N.</given-names></name> <name><surname>Du</surname> <given-names>J.</given-names></name> <name><surname>Joshi</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>RoBERTa: a robustly optimized BERT pretraining approach</article-title>. <source>arXiv preprint</source> arXiv:1907.11692. <pub-id pub-id-type="doi">10.48550/arXiv.1907.11692</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luke</surname> <given-names>S. G.</given-names></name> <name><surname>Christianson</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>The provo corpus: a large eye-tracking corpus with predictability norms</article-title>. <source>Behav. Res. Methods</source> <volume>50</volume>, <fpage>826</fpage>&#x02013;<lpage>833</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-017-0908-4</pub-id><pub-id pub-id-type="pmid">28523601</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lund</surname> <given-names>K.</given-names></name> <name><surname>Burgess</surname> <given-names>C.</given-names></name></person-group> (<year>1996</year>). <article-title>Producing high-dimensional semantic spaces from lexical co-occurrence</article-title>. <source>Behav. Res. Methods Instruments Comput</source>. <volume>28</volume>, <fpage>203</fpage>&#x02013;<lpage>208</lpage>. <pub-id pub-id-type="doi">10.3758/BF03204766</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Merkx</surname> <given-names>D.</given-names></name> <name><surname>Frank</surname> <given-names>S. L.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Human sentence processing: recurrence or attention?&#x0201D;</article-title> in <source>Proceedings of the NAACL Workshop on Cognitive Modeling and Computational Linguistics</source>.</citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Michaelov</surname> <given-names>J. A.</given-names></name> <name><surname>Bergen</surname> <given-names>B. K.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x00027;Rarely&#x00027; a problem? language models exhibit inverse scaling in their predictions following &#x00027;few&#x00027;-type quantifiers</article-title>. <source>arXiv preprint</source> arXiv:2212.08700. <pub-id pub-id-type="doi">10.48550/arXiv.2212.08700</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Efficient estimation of word representations in vector space</article-title>. <source>arXiv preprint</source> arXiv:1301.3781. <pub-id pub-id-type="doi">10.48550/arXiv.1301.3781</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mitchell</surname> <given-names>J.</given-names></name> <name><surname>Lapata</surname> <given-names>M.</given-names></name> <name><surname>Demberg</surname> <given-names>V.</given-names></name> <name><surname>Keller</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Syntactic and semantic factors in processing difficulty: an integrated measure,&#x0201D;</article-title> in <source>Proceedings of ACL</source> (<publisher-loc>Uppsala</publisher-loc>).</citation>
</ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Monsalve</surname> <given-names>I. F.</given-names></name> <name><surname>Frank</surname> <given-names>S. L.</given-names></name> <name><surname>Vigliocco</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Lexical surprisal as a general predictor of reading time,&#x0201D;</article-title> in <source>Proceedings of EACL</source> (<publisher-loc>Avignon</publisher-loc>).</citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oh</surname> <given-names>B.-D.</given-names></name> <name><surname>Schuler</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Entropy-and distance-based predictors from GPT-2 attention patterns predict reading times over and above GPT-2 surprisal,&#x0201D;</article-title> in <source>Proceedings of EMNLP</source> (<publisher-loc>Abu Dhabi</publisher-loc>).</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pad&#x000F3;</surname> <given-names>S.</given-names></name> <name><surname>Lapata</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <article-title>Dependency-based construction of semantic space models</article-title>. <source>Comput. Linguist</source>. <volume>33</volume>, <fpage>161</fpage>&#x02013;<lpage>199</lpage>. <pub-id pub-id-type="doi">10.1162/coli.2007.33.2.161</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pennington</surname> <given-names>J.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Manning</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Glove: global vectors for word representation,&#x0201D;</article-title> in <source>Proceedings of EMNLP</source> (<publisher-loc>Doha</publisher-loc>).</citation>
</ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Peters</surname> <given-names>M. E.</given-names></name> <name><surname>Neumann</surname> <given-names>M.</given-names></name> <name><surname>Iyyer</surname> <given-names>M.</given-names></name> <name><surname>Gardner</surname> <given-names>M.</given-names></name> <name><surname>Clark</surname> <given-names>C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Deep contextualized word representations,&#x0201D;</article-title> in <source>Proceedings of NAACL</source> (<publisher-loc>New Orleans, LA</publisher-loc>).<pub-id pub-id-type="pmid">34343877</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pollatsek</surname> <given-names>A.</given-names></name> <name><surname>Juhasz</surname> <given-names>B. J.</given-names></name> <name><surname>Reichle</surname> <given-names>E. D.</given-names></name> <name><surname>Machacek</surname> <given-names>D.</given-names></name> <name><surname>Rayner</surname> <given-names>K.</given-names></name></person-group> (<year>2008</year>). <article-title>Immediate and delayed effects of word frequency and word length on eye movements in reading: a reversed delayed effect of word length</article-title>. <source>J. Exp. Psychol</source>. <volume>34</volume>, <fpage>726</fpage>. <pub-id pub-id-type="doi">10.1037/0096-1523.34.3.726</pub-id><pub-id pub-id-type="pmid">18505334</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pynte</surname> <given-names>J.</given-names></name> <name><surname>New</surname> <given-names>B.</given-names></name> <name><surname>Kennedy</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>On-line contextual influences during reading normal text: a multiple-regression analysis</article-title>. <source>Vision Res</source>. <volume>48</volume>, <fpage>2172</fpage>&#x02013;<lpage>2183</lpage>. <pub-id pub-id-type="doi">10.1016/j.visres.2008.02.004</pub-id><pub-id pub-id-type="pmid">18701125</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Radach</surname> <given-names>R.</given-names></name> <name><surname>Kennedy</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Eye movements in reading: some theoretical context</article-title>. <source>Q. J. Exp. Psychol</source>. <volume>66</volume>, <fpage>429</fpage>&#x02013;<lpage>452</lpage>. <pub-id pub-id-type="doi">10.1080/17470218.2012.750676</pub-id><pub-id pub-id-type="pmid">23289943</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Narasimhan</surname> <given-names>K.</given-names></name> <name><surname>Salimans</surname> <given-names>T.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2018</year>). <source>Improving Language Understanding by Generative Pre-training</source>. Open-AI Blog.</citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Child</surname> <given-names>R.</given-names></name> <name><surname>Luan</surname> <given-names>D.</given-names></name> <name><surname>Amodei</surname> <given-names>D.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2019</year>). <source>Language Models are Unsupervised Multitask Learners</source>. In Open-AI Blog.<pub-id pub-id-type="pmid">35637722</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rayner</surname> <given-names>K.</given-names></name></person-group> (<year>1998</year>). <article-title>Eye movements in reading and information processing: 20 years of research</article-title>. <source>Psychol. Bull</source>. <volume>124</volume>, <fpage>372</fpage>&#x02013;<lpage>422</lpage>. <pub-id pub-id-type="doi">10.1037/0033-2909.124.3.372</pub-id><pub-id pub-id-type="pmid">9849112</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rayner</surname> <given-names>K.</given-names></name> <name><surname>Duffy</surname> <given-names>S. A.</given-names></name></person-group> (<year>1986</year>). <article-title>Lexical complexity and fixation times in reading: effects of word frequency, verb complexity, and lexical ambiguity</article-title>. <source>Mem. Cogn</source>. <volume>14</volume>, <fpage>191</fpage>&#x02013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.3758/BF03197692</pub-id><pub-id pub-id-type="pmid">3736392</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rayner</surname> <given-names>K.</given-names></name> <name><surname>Sereno</surname> <given-names>S. C.</given-names></name> <name><surname>Morris</surname> <given-names>R. K.</given-names></name> <name><surname>Schmauder</surname> <given-names>A. R.</given-names></name> <name><surname>Clifton Jr</surname> <given-names>C.</given-names></name></person-group> (<year>1989</year>). <article-title>Eye movements and on-line language comprehension processes</article-title>. <source>Lang. Cogn. Process</source>. <volume>4</volume>, <fpage>SI21</fpage>-<lpage>SI49</lpage>. <pub-id pub-id-type="doi">10.1080/01690968908406362</pub-id></citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Remington</surname> <given-names>R. W.</given-names></name> <name><surname>Burt</surname> <given-names>J. S.</given-names></name> <name><surname>Becker</surname> <given-names>S. I.</given-names></name></person-group> (<year>2018</year>). <article-title>The curious case of spillover: does it tell us much about saccade timing in reading?</article-title> <source>Attent. Percept. Psychophys</source>. <volume>80</volume>, <fpage>1683</fpage>&#x02013;<lpage>1690</lpage>. <pub-id pub-id-type="doi">10.3758/s13414-018-1544-5</pub-id><pub-id pub-id-type="pmid">29968083</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>M. A.</given-names></name> <name><surname>Merlo</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Word associations and the distance properties of context-aware word embeddings,&#x0201D;</article-title> in <source>Proceedings of CONLL</source>.</citation>
</ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahlgren</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>The distributional hypothesis</article-title>. <source>Italian J. Comput. Linguist</source>. <volume>20</volume>, <fpage>33</fpage>&#x02013;<lpage>53</lpage>.</citation>
</ref>
<ref id="B68">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Salicchi</surname> <given-names>L.</given-names></name> <name><surname>Lenci</surname> <given-names>A.</given-names></name> <name><surname>Chersoni</surname> <given-names>E.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Looking for a role for word embeddings in eye-tracking features prediction: does semantic similarity help?&#x0201D;</article-title> in <source>Proceedings of IWCS</source> (<publisher-loc>Dublin</publisher-loc>).</citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salicchi</surname> <given-names>L.</given-names></name> <name><surname>Xiang</surname> <given-names>R.</given-names></name> <name><surname>Hsu</surname> <given-names>Y.-Y.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;HkAmsters at CMCL 2022 shared task: predicting eye-tracking data from a gradient boosting framework with linguistic features,&#x0201D;</article-title> in <source>Proceedings of the ACL Workshop on Cognitive Modeling and Computational Linguistics</source>.</citation>
</ref>
<ref id="B70">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Santus</surname> <given-names>E.</given-names></name> <name><surname>Chersoni</surname> <given-names>E.</given-names></name> <name><surname>Lenci</surname> <given-names>A.</given-names></name> <name><surname>Huang</surname> <given-names>C.-R.</given-names></name> <name><surname>Blache</surname> <given-names>P</given-names></name></person-group>. (<year>2016a</year>). <article-title>&#x0201C;Testing APsyn against Vector cosine on similarity estimation,&#x0201D;</article-title> in <source>Proceedings of PACLIC</source> (<publisher-loc>Seoul</publisher-loc>).</citation>
</ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santus</surname> <given-names>E.</given-names></name> <name><surname>Chiu</surname> <given-names>T.-S.</given-names></name> <name><surname>Lu</surname> <given-names>Q.</given-names></name> <name><surname>Lenci</surname> <given-names>A.</given-names></name> <name><surname>Huang</surname> <given-names>C.-R.</given-names></name></person-group> (<year>2016b</year>). <article-title>&#x0201C;What a Nerd! beating students and vector cosine in the ESL and TOEFL datasets,&#x0201D;</article-title> in Proceedings <italic>of LREC</italic> (Portoro&#x0017E;).</citation>
</ref>
<ref id="B72">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Santus</surname> <given-names>E.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Chersoni</surname> <given-names>E.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;A rank-based similarity metric for word embeddings,&#x0201D;</article-title> in <source>Proceedings of ACL</source> (<publisher-loc>Melbourne, VIC</publisher-loc>).</citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sayeed</surname> <given-names>A.</given-names></name> <name><surname>Shkadzko</surname> <given-names>P.</given-names></name> <name><surname>Demberg</surname> <given-names>V.</given-names></name></person-group> (<year>2015</year>). <article-title>An exploration of semantic features in an unsupervised thematic fit evaluation framework</article-title>. <source>Italian J. Comput. Linguist</source>. <volume>1</volume>, <fpage>31</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.4000/ijcol.298</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shain</surname> <given-names>C.</given-names></name> <name><surname>Meister</surname> <given-names>C.</given-names></name> <name><surname>Pimentel</surname> <given-names>T.</given-names></name> <name><surname>Cotterell</surname> <given-names>R.</given-names></name> <name><surname>Levy</surname> <given-names>R. P.</given-names></name></person-group> (<year>2022</year>). <article-title>Large-scale evidence for logarithmic effects of word predictability on reading time</article-title>. <source>PsyArXiv</source>. <pub-id pub-id-type="doi">10.31234/osf.io/4hyna</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siegelman</surname> <given-names>N.</given-names></name> <name><surname>Schroeder</surname> <given-names>S.</given-names></name> <name><surname>Acart&#x000FC;rk</surname> <given-names>C.</given-names></name> <name><surname>Ahn</surname> <given-names>H.-D.</given-names></name> <name><surname>Alexeeva</surname> <given-names>S.</given-names></name> <name><surname>Amenta</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Expanding horizons of cross-linguistic research on reading: the multilingual eye-movement corpus (meco)</article-title>. <source>Behav. Res. Methods</source> <volume>2022</volume>, <fpage>1</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-021-01772-6</pub-id><pub-id pub-id-type="pmid">35112286</pub-id></citation></ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>N. J.</given-names></name> <name><surname>Levy</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <article-title>The effect of word predictability on reading time is logarithmic</article-title>. <source>Cognition</source> <volume>128</volume>, <fpage>302</fpage>&#x02013;<lpage>319</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2013.02.013</pub-id><pub-id pub-id-type="pmid">23747651</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Timkey</surname> <given-names>W.</given-names></name> <name><surname>van Schijndel</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;All bark and no bite: rogue dimensions in transformer language models obscure representational quality,&#x0201D;</article-title> in <source>Proceedings of EMNLP</source> (<publisher-loc>Punta Cana</publisher-loc>).</citation>
</ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Traxler</surname> <given-names>M. J.</given-names></name> <name><surname>Foss</surname> <given-names>D. J.</given-names></name> <name><surname>Seely</surname> <given-names>R. E.</given-names></name> <name><surname>Kaup</surname> <given-names>B.</given-names></name> <name><surname>Morris</surname> <given-names>R. K.</given-names></name></person-group> (<year>2000</year>). <article-title>Priming in sentence processing: intralexical spreading activation, schemas, and situation models</article-title>. <source>J. Psycholinguist. Res</source>. <volume>29</volume>, <fpage>581</fpage>&#x02013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1023/A:1026416225168</pub-id><pub-id pub-id-type="pmid">11196064</pub-id></citation></ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turney</surname> <given-names>P. D.</given-names></name> <name><surname>Pantel</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>From frequency to meaning: vector space models of semantics</article-title>. <source>J. Artif. Intell. Res</source>. <volume>37</volume>, <fpage>141</fpage>&#x02013;<lpage>188</lpage>. <pub-id pub-id-type="doi">10.1613/jair.2934</pub-id><pub-id pub-id-type="pmid">27138012</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>van Schijndel</surname> <given-names>M.</given-names></name> <name><surname>Linzen</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;A neural model of adaptation in reading,&#x0201D;</article-title> in <source>Proceedings of EMNLP</source> (<publisher-loc>Brussels</publisher-loc>).</citation>
</ref>
<ref id="B81">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Attention is all you need,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>5998</fpage>&#x02013;<lpage>6008</lpage>.</citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wilcox</surname> <given-names>E. G.</given-names></name> <name><surname>Gauthier</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Qian</surname> <given-names>P.</given-names></name> <name><surname>Levy</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;On the predictive power of neural language models for human real-time comprehension behavior,&#x0201D;</article-title> in <source>Proceedings of CogSci</source>.</citation>
</ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wlotko</surname> <given-names>E. W.</given-names></name> <name><surname>Federmeier</surname> <given-names>K. D.</given-names></name></person-group> (<year>2015</year>). <article-title>Time for prediction? the effect of presentation rate on predictive sentence comprehension during word-by-word reading</article-title>. <source>Cortex</source> <volume>68</volume>, <fpage>20</fpage>&#x02013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1016/j.cortex.2015.03.014</pub-id><pub-id pub-id-type="pmid">25987437</pub-id></citation></ref>
<ref id="B84">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Dai</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Carbonell</surname> <given-names>J.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;XlNet: generalized autoregressive pretraining for language understanding,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 32</source> (<publisher-loc>Vancouver, BC</publisher-loc>).</citation>
</ref>
<ref id="B85">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhelezniak</surname> <given-names>V.</given-names></name> <name><surname>Savkov</surname> <given-names>A.</given-names></name> <name><surname>Shen</surname> <given-names>A.</given-names></name> <name><surname>Hammerla</surname> <given-names>N.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Correlation coefficients and semantic textual similarity,&#x0201D;</article-title> in <source>Proceedings of NAACL</source> (<publisher-loc>Minneapolis, MN</publisher-loc>).<pub-id pub-id-type="pmid">34842531</pub-id></citation></ref>
</ref-list>
</back>
</article>