<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3-mathml3.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" dtd-version="1.3" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Sci.</journal-id>
<journal-title-group>
<journal-title>Frontiers in Computer Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Sci.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2624-9898</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcomp.2025.1634228</article-id>
<article-version article-version-type="Version of Record" vocab="NISO-RP-8-2008"/>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Original Research</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Beyond speech: leveraging mouse movements for information adaptation in voice interfaces</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Kontogiorgos</surname> <given-names>Dimosthenis</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Methodology" vocab-term-identifier="https://credit.niso.org/contributor-roles/methodology/">Methodology</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Investigation" vocab-term-identifier="https://credit.niso.org/contributor-roles/investigation/">Investigation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Software" vocab-term-identifier="https://credit.niso.org/contributor-roles/software/">Software</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Visualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/visualization/">Visualization</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Formal analysis" vocab-term-identifier="https://credit.niso.org/contributor-roles/formal-analysis/">Formal analysis</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; original draft" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-original-draft/">Writing &#x2013; original draft</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Validation" vocab-term-identifier="https://credit.niso.org/contributor-roles/validation/">Validation</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Data curation" vocab-term-identifier="https://credit.niso.org/contributor-roles/data-curation/">Data curation</role>
<uri xlink:href="https://loop.frontiersin.org/people/681827"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Schlangen</surname> <given-names>David</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Funding acquisition" vocab-term-identifier="https://credit.niso.org/contributor-roles/funding-acquisition/">Funding acquisition</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Writing &#x2013; review &amp; editing" vocab-term-identifier="https://credit.niso.org/contributor-roles/writing-review-editing/">Writing &#x2013; review &#x00026; editing</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Supervision" vocab-term-identifier="https://credit.niso.org/contributor-roles/supervision/">Supervision</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Project administration" vocab-term-identifier="https://credit.niso.org/contributor-roles/project-administration/">Project administration</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Resources" vocab-term-identifier="https://credit.niso.org/contributor-roles/resources/">Resources</role>
<role vocab="credit" vocab-identifier="https://credit.niso.org/" vocab-term="Conceptualization" vocab-term-identifier="https://credit.niso.org/contributor-roles/conceptualization/">Conceptualization</role>
<uri xlink:href="https://loop.frontiersin.org/people/1315607"/>
</contrib>
</contrib-group>
<aff id="aff1"><label>1</label><institution>Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology</institution>, <city>Cambridge, MA</city>, <country country="us">United States</country></aff>
<aff id="aff2"><label>2</label><institution>Department of Linguistics, University of Potsdam</institution>, <city>Potsdam</city>, <country country="de">Germany</country></aff>
<author-notes>
<corresp id="c001"><label>&#x0002A;</label>Correspondence: Dimosthenis Kontogiorgos, <email xlink:href="mailto:dimos@csail.mit.edu">dimos@csail.mit.edu</email></corresp>
</author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2025-12-12">
<day>12</day>
<month>12</month>
<year>2025</year>
</pub-date>
<pub-date publication-format="electronic" date-type="collection">
<year>2025</year>
</pub-date>
<volume>7</volume>
<elocation-id>1634228</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>05</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>11</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Kontogiorgos and Schlangen.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Kontogiorgos and Schlangen</copyright-holder>
<license>
<ali:license_ref start_date="2025-12-12">https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</license-p>
</license>
</permissions>
<abstract>
<p>As human speakers naturally adapt their linguistic styles to one another, voice user interfaces that prompt similar linguistic adaptations can augment human-like interaction. In this study, we leverage a corpus of human instructions to model the effectiveness of incremental instruction generation in artificial agents. Participants interacted with agents that guided them in selecting virtual puzzle pieces, varying the amount of information provided in each instruction. Through an empirical examination of the Gricean maxims in utterance construction, our initial perception study highlighted the significance of adaptive instruction generation. By employing mouse movements as a proxy for user understanding, we developed computational models that enabled agents to detect user uncertainty and refine instructions incrementally. Comparing speaker-based and listener-based models, we found that agents encouraging linguistic adaptations were preferred by users. Our findings offer new insights into the value of mouse movements as indicators of user comprehension and introduce a methodological framework for developing adaptive interactive systems that generate instructions dynamically.</p></abstract>
<kwd-group>
<kwd>instructions</kwd>
<kwd>language production</kwd>
<kwd>mouse tracking</kwd>
<kwd>common ground</kwd>
<kwd>adaptive systems</kwd>
<kwd>incremental instruction</kwd>
<kwd>voice user interfaces</kwd>
<kwd>conversational agents</kwd>
</kwd-group>
<funding-group>
<award-group id="gs1">
<funding-source id="sp1">
<institution-wrap>
<institution>Deutsche Forschungsgemeinschaft</institution>
<institution-id institution-id-type="doi" vocab="open-funder-registry" vocab-identifier="10.13039/open_funder_registry">10.13039/501100001659</institution-id>
</institution-wrap>
</funding-source>
</award-group>
<award-group id="gs2">
<funding-source id="sp2">
<institution-wrap>
<institution>Knut och Alice Wallenbergs Stiftelse</institution>
<institution-id institution-id-type="doi" vocab="open-funder-registry" vocab-identifier="10.13039/open_funder_registry">10.13039/501100004063</institution-id>
</institution-wrap>
</funding-source>
</award-group>
<funding-statement>The author(s) declare that financial support was received for the research and/or publication of this article. This research was partially supported by the German Research Foundation (DFG) project RECOLAGE (423217434) and a PostDoctoral Research Fellowship by the Knut and Alice Wallenberg Foundation.</funding-statement>
</funding-group>
<counts>
<fig-count count="9"/>
<table-count count="5"/>
<equation-count count="0"/>
<ref-count count="130"/>
<page-count count="19"/>
<word-count count="13170"/>
</counts>
<custom-meta-group>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Human-Media Interaction</meta-value>
</custom-meta>
</custom-meta-group>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<label>1</label>
<title>Introduction</title>
<p>&#x0201C;<italic>Pass me the Vernier, please,&#x0201D;</italic> said Jakob. &#x0201C;<italic>The what?&#x0201D;</italic> asked Frida. &#x0201C;<italic>The calliper&#x0201D;</italic>, Jakob replied, observing Frida&#x00027;s confusion. &#x0201C;<italic>The long metal thing... right in front of you&#x0201D;</italic>, continued Jakob. Although this dialogue is only represented in text, it illustrates the richness of an interactional setting and how embodied actions can prompt the reformulation of speech. In <italic>task-oriented interactions</italic>, speakers collaboratively build <italic>common ground</italic>&#x02014;a mutual understanding of shared goals (<xref ref-type="bibr" rid="B27">Clark and Marshall, 1981</xref>; <xref ref-type="bibr" rid="B28">Clark and Wilkes-Gibbs, 1986</xref>; <xref ref-type="bibr" rid="B45">Fussell and Krauss, 1992</xref>; <xref ref-type="bibr" rid="B30">Dafoe et al., 2021</xref>). Utterances are constructed in a cooperative manner (<xref ref-type="bibr" rid="B24">Clark, 1996</xref>), following a principle known as <italic>audience design</italic> (<xref ref-type="bibr" rid="B13">Bell, 1984</xref>; <xref ref-type="bibr" rid="B17">Brennan and Hanna, 2009</xref>).</p>
<sec>
<label>1.1</label>
<title>Grice&#x00027;s cooperative principle</title>
<p>This process is part of the <italic>cooperative principle</italic> that Grice defined as the <italic>maxim of quantity</italic> (<xref ref-type="bibr" rid="B51">Grice, 1975</xref>, <xref ref-type="bibr" rid="B52">1989</xref>), which represents <italic>efficiency</italic> in communication by conveying the <italic>most</italic> accurate information with the <italic>least</italic> effort required. Speakers minimize collaborative effort by always seeking <italic>positive evidence of understanding</italic>, a concept known as the <italic>grounding process</italic> (<xref ref-type="bibr" rid="B25">Clark and Brennan, 1991</xref>; <xref ref-type="bibr" rid="B16">Brennan and Clark, 1996</xref>). They construct utterances in an <italic>opportunistic approach</italic>, incrementally gathering evidence that the criteria for mutual understanding are met (<xref ref-type="bibr" rid="B96">Sacks et al., 1978</xref>; <xref ref-type="bibr" rid="B103">Schober and Clark, 1989</xref>; <xref ref-type="bibr" rid="B49">Gonsior et al., 2010</xref>). Difficult descriptions are often conveyed through episodic utterances or <italic>incremental units</italic>, where the communicative act itself, rather than the information conveyed, holds the most significance (<xref ref-type="bibr" rid="B76">Krahmer and Van Deemter, 2012</xref>). This type of adaptation poses a challenge for interfaces, as it requires them to have robust representations of the interaction state and effectively interpret the user&#x00027;s signals.</p>
</sec>
<sec>
<label>1.2</label>
<title>Approach</title>
<p>In this paper, we examine these interactional phenomena by replicating incremental utterance construction in voice user interfaces. We begin with the assumption that there is an alignment between <italic>the complexity of instructions and the level of assistance required from the system</italic>. Some users may rely less on system cues and are more likely to achieve their goals with systems that adapt to their individual needs (<xref ref-type="bibr" rid="B117">Torrey et al., 2006</xref>). While low system effort may lead to misunderstandings, excessive effort could overwhelm the listener (<xref ref-type="bibr" rid="B115">Torrey et al., 2013</xref>; <xref ref-type="bibr" rid="B22">Chai et al., 2014</xref>; <xref ref-type="bibr" rid="B72">Kontogiorgos and Gustafson, 2021</xref>).</p>
<p>We examine this adaptation process from the perspective of <italic>common ground</italic>. Our approach begins by analyzing a corpus in which humans instruct each other in a task-oriented setting. The annotated instructions were synthesized through a Text-To-Speech (TTS) system and evaluated in the first study to explore the balance of information necessary for task completion. Participants&#x00027; mouse movements were collected, analyzed, and used as a proxy for utterance comprehension. In a second study, employing Machine-Learning methods, the system automatically assessed participants&#x00027; understanding based on their mouse movements. This adaptive instruction generation method enabled the system to predict in real-time whether users would successfully complete the task and to adjust the construction of instructions incrementally.</p>
</sec>
<sec>
<label>1.3</label>
<title>Research questions</title>
<p>The corpus analysis and two studies address the following research questions:</p>
<list list-type="simple">
<list-item><p><bold>RQ1:</bold> How do human speakers produce instructions in incremental units, and what are their attributes (e.g., timing, duration)?</p></list-item>
<list-item><p><bold>RQ2:</bold> What is the optimal granularity of information that an interface should provide at each incremental unit of a goal-oriented task? Specifically, how does varying the amount and detail of information affect task performance and user understanding?</p></list-item>
<list-item><p><bold>RQ3:</bold> How can the interface dynamically adapt its communication strategy when the user&#x00027;s attention or behavior deviates from the expected interaction pattern?</p></list-item>
</list>
</sec>
<sec>
<label>1.4</label>
<title>Contributions of this article</title>
<p>Our findings indicate that voice user interfaces that utilize incremental instructions can effectively minimize collaborative effort with users. We demonstrate that mouse movements serve as a <italic>reliable proxy for utterance comprehension</italic> and influence instruction behavior in incremental units. The goal of this paper is to estimate how well this coordination is maintained by predicting, throughout the interaction, whether the user&#x00027;s goal is uncertain (<xref ref-type="bibr" rid="B74">Kontogiorgos et al., 2019</xref>). We design a system that adapts to the user&#x00027;s information needs in a virtual puzzle task by providing referential information incrementally (<xref ref-type="bibr" rid="B41">Engonopoulos et al., 2013</xref>; <xref ref-type="bibr" rid="B129">Zarrie&#x000DF; and Schlangen, 2016</xref>) (<xref ref-type="fig" rid="F1">Figure 1</xref>).</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>Research design of this study: (1) Extraction of instructions from a dataset of human instructors (<xref ref-type="bibr" rid="B128">Zarrie&#x000DF; et al., 2016</xref>). (2) Modeling and synthesizing the extracted instructions, followed by evaluation in an online study (Study 1). (3) Utilizing mouse movement data from Study 1 to train adaptive instruction models, with subsequent evaluation of these models in a perception study (Study 2). <bold>(a)</bold> Human instructors dataset. <bold>(b)</bold> Study 1: ML model training. <bold>(c)</bold> Study 2: human evaluation.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0001.tif">
<alt-text content-type="machine-generated">Diagram illustrating a research workflow in three parts: (a) Human instructors dataset showing two individuals with images and audio data labeled with &#x0201C;there&#x02019;s a piece that&#x02019;s like an L shape,&#x0201D; (b) Study 1 depicting machine learning model training with mouse features, BERT, and multimodal fusion, (c) Study 2 showing a human evaluating on a computer, repeating the L-shape description.</alt-text>
</graphic>
</fig>
<p>This line of work contributes to empirical findings in linguistic alignment research (<xref ref-type="bibr" rid="B15">Branigan et al., 2010</xref>). Specifically, we demonstrate how an interface can exhibit adaptive and collaborative behavior by providing <italic>as much information as needed by users</italic>. We evaluate three approaches to eliciting adaptive behavior by comparing two interaction strategies: (a) a <italic>speaker-based model</italic> and (b) a <italic>listener-based model</italic>, against (c) a control condition where users explicitly request the information they need. Differences are assessed in terms of users&#x00027; task performance, user behavior, and perceptions of the three agents, providing valuable insights for designing future voice user interfaces that deliver personalized instructions.</p>
</sec>
<sec>
<label>1.5</label>
<title>Background and related work</title>
<sec>
<label>1.5.1</label>
<title>Mutual understanding with voice user interfaces</title>
<p>Incremental language construction behaviors demonstrate the <italic>affordances</italic> of interaction, showing that mutual understanding can be shaped as a continuous, participatory, and collaborative process (<xref ref-type="bibr" rid="B24">Clark, 1996</xref>; <xref ref-type="bibr" rid="B26">Clark and Krych, 2004</xref>; <xref ref-type="bibr" rid="B10">Baumann et al., 2013</xref>). Like human speakers, voice user interfaces should employ data-driven approaches and adapt their instruction strategies to match users&#x00027; evolving levels of understanding (<xref ref-type="bibr" rid="B89">Pelikan and Broth, 2016</xref>; <xref ref-type="bibr" rid="B73">Kontogiorgos and Pelikan, 2020</xref>; <xref ref-type="bibr" rid="B12">Behnke et al., 2020</xref>). While much of the HCI research has concentrated on preventing miscommunication with users, it often overlooks that human dialogue is grounded in the <italic>cooperative principle</italic> (<xref ref-type="bibr" rid="B52">Grice, 1989</xref>), characterized by variability in interactional phenomena (e.g., disfluencies, repairs, hesitations) (<xref ref-type="bibr" rid="B75">Kousidis et al., 2014</xref>; <xref ref-type="bibr" rid="B20">Buschmeier et al., 2012</xref>; <xref ref-type="bibr" rid="B120">Wagner et al., 2015</xref>; <xref ref-type="bibr" rid="B55">Haake et al., 2019</xref>). Given the inherently social nature of human communication, user interfaces must incrementally monitor users for social cues that signal mutual understanding (<xref ref-type="bibr" rid="B71">Kontogiorgos, 2022</xref>), a human-like capability that voice user interfaces currently lack.</p></sec>
<sec>
<label>1.5.2</label>
<title>Adaptation in cooperative AI</title>
<p>In this paper, we focus on computer adaptation, which is further demonstrated in the two studies presented. We examine adaptation within the domain of referential language,<xref ref-type="fn" rid="fn0003"><sup>1</sup></xref> particularly in the instructional use of language that describes objects within the shared space of attention between the user and the computer (<xref ref-type="bibr" rid="B7">Axelsson and Skantze, 2020</xref>). The user&#x00027;s visual attention is considered in the form of mouse movements to guide subsequent instructions.<xref ref-type="fn" rid="fn0004"><sup>2</sup></xref></p>
<p>A significant body of research on referring expression (RE) generation in HCI has focused on producing REs as the shortest possible expressions with minimal ambiguity (<xref ref-type="bibr" rid="B31">Dale and Reiter, 1995</xref>; <xref ref-type="bibr" rid="B125">Williams and Scheutz, 2017</xref>). However, this approach does not fully align with how humans naturally communicate. Humans depend on the cooperation of their conversational partner to resolve ambiguities, constructing descriptions in an <italic>opportunistic</italic> manner that often results in non-optimal, yet adequate, utterances. These utterances can be repaired and adjusted according to the listener&#x00027;s understanding. This process is inherently collaborative, particularly in interactive settings, where the RE is tailored to the specific listener, and a sequence of utterances is iteratively refined until mutual comprehension is achieved. <italic>The objective of this process is to minimize the joint effort, thereby producing REs with the least collaborative effort</italic> (<xref ref-type="bibr" rid="B28">Clark and Wilkes-Gibbs, 1986</xref>; <xref ref-type="bibr" rid="B25">Clark and Brennan, 1991</xref>).</p>
<p>Some HCI research has investigated <italic>incremental</italic> descriptions or ambiguous instructions across various contexts, including computational studies on situated dialogue among humans (<xref ref-type="bibr" rid="B65">Kelleher and Kruijff, 2006</xref>; <xref ref-type="bibr" rid="B32">Dethlefs et al., 2011</xref>; <xref ref-type="bibr" rid="B68">Kirk and Fraser, 2017</xref>; <xref ref-type="bibr" rid="B82">Magassouba et al., 2018</xref>), and visual search (<xref ref-type="bibr" rid="B78">Kraut et al., 2003</xref>; <xref ref-type="bibr" rid="B129">Zarrie&#x000DF; and Schlangen, 2016</xref>; <xref ref-type="bibr" rid="B79">Li et al., 2020</xref>; <xref ref-type="bibr" rid="B94">Rojowiec et al., 2020</xref>). This work often involves incremental units (<xref ref-type="bibr" rid="B105">Skantze and Hjalmarsson, 2010</xref>; <xref ref-type="bibr" rid="B11">Baumann and Schlangen, 2012</xref>; <xref ref-type="bibr" rid="B66">Kennington and Schlangen, 2017</xref>; <xref ref-type="bibr" rid="B63">Jensen et al., 2020</xref>) or leverages the incremental algorithm (<xref ref-type="bibr" rid="B34">DeVault et al., 2005</xref>). Researchers have also incorporated signals like users&#x00027; eye-gaze (<xref ref-type="bibr" rid="B70">Koller et al., 2012</xref>; <xref ref-type="bibr" rid="B107">Staudte et al., 2012</xref>; <xref ref-type="bibr" rid="B83">Mitev et al., 2018</xref>) and employed paradigms in virtual environments (<xref ref-type="bibr" rid="B108">Stoia et al., 2006</xref>; <xref ref-type="bibr" rid="B109">Striegnitz et al., 2012</xref>; <xref ref-type="bibr" rid="B46">Garoufi and Koller, 2014</xref>). Additionally, instructions have been examined within Human-Robot Interaction (HRI) through human-robot collaborative instruction tasks (<xref ref-type="bibr" rid="B42">Fang et al., 2015</xref>; <xref ref-type="bibr" rid="B121">Wallbridge et al., 2019</xref>; <xref ref-type="bibr" rid="B36">Do&#x0011F;an et al., 2020</xref>; <xref ref-type="bibr" rid="B123">Weerakoon et al., 2020</xref>; <xref ref-type="bibr" rid="B122">Wallbridge et al., 2021</xref>; <xref ref-type="bibr" rid="B37">Do&#x0011F;an and Leite, 2021</xref>) and in robotic navigation (<xref ref-type="bibr" rid="B112">Tellex et al., 2011</xref>).</p>
<p>What information interfaces disclose in each incremental unit or which instructional strategies may be most effective has been sparsely explored, primarily within HRI research (<xref ref-type="bibr" rid="B117">Torrey et al., 2006</xref>, <xref ref-type="bibr" rid="B116">2007</xref>, <xref ref-type="bibr" rid="B115">2013</xref>; <xref ref-type="bibr" rid="B99">Saupp and Mutlu, 2014</xref>; <xref ref-type="bibr" rid="B100">Saupp&#x000E9; and Mutlu, 2015</xref>). Collaborative human-AI utterance generation has also been approached through abstraction matching, generating grounded utterances that align with user intent in Python code generation (<xref ref-type="bibr" rid="B81">Liu et al., 2023</xref>). Providing explanations as repair strategies has been demonstrated to be effective in conversational interactions (<xref ref-type="bibr" rid="B6">Ashktorab et al., 2019</xref>). While this article focuses on the social dynamics of generating utterances incrementally, some research has raised questions about building relationships and bonds in conversational agent communication, favoring a focus on transactional and utilitarian aspects without directly mimicking human-to-human conversation (<xref ref-type="bibr" rid="B29">Clark et al., 2019</xref>).</p></sec>
<sec>
<label>1.5.3</label>
<title>Adaptation in the form of incremental units</title>
<p>Instructions in situated interactions are often composed of multiple fragmentary utterance units, described as a <italic>series of corrections</italic> (<xref ref-type="bibr" rid="B80">Lindwall and Ekstr&#x000F6;m, 2012</xref>). These instructions depend on mutual understanding, from their initial formulation to the iterative process of being reformulated based on the actions of the person receiving the instruction. This adaptation process poses a significant challenge for computers, as they must continuously monitor the user and adjust instructions in real-time. Thus, the generation of instructions should not be viewed solely as <italic>information exchange</italic>,<xref ref-type="fn" rid="fn0005"><sup>3</sup></xref> but also as an opportunity to demonstrate socially intelligent behavior.</p>
<p>Such behavior is likely influenced by the speaker&#x00027;s ability to adhere to the cooperative principle and the <italic>Gricean maxim of quantity</italic>, providing as much information as necessary with as few utterances as possible (<xref ref-type="bibr" rid="B48">Gigliobianco et al., 2024</xref>). This approach is also likely intended to minimize collaborative effort (<xref ref-type="bibr" rid="B42">Fang et al., 2015</xref>; <xref ref-type="bibr" rid="B72">Kontogiorgos and Gustafson, 2021</xref>), with difficult instructions being presented incrementally until common ground is established. One reason for this behavior could be that it is easier for speakers to plan utterances incrementally rather than constructing a single, unambiguous instruction; another reason is the flexibility it provides in adapting to the listener&#x00027;s understanding. This joint orientation of incremental turns is often facilitated through pausing and forming intonational phrases, allowing the speaker to adapt and reformulate instructions as multi-utterance contributions to the conversation (<xref ref-type="bibr" rid="B28">Clark and Wilkes-Gibbs, 1986</xref>), in synchrony with the listener&#x00027;s signals of understanding&#x02014;a challenging task for voice user interfaces.</p></sec></sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Modeling instructions</title>
<sec>
<label>2.1</label>
<title>Human instructor corpus</title>
<p>We used a corpus of human instructors from (<xref ref-type="bibr" rid="B128">Zarrie&#x000DF; et al., 2016</xref>). The corpus consists of 11 dialogue pairs of native English speakers interacting through video. Their task was to solve a virtual pentomino puzzle, forming shapes such as the elephant shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The role of the &#x0201C;Instructor&#x0201D; was to guide the &#x0201C;User&#x0201D; participant on how to solve the puzzle.</p>
<fig position="float" id="F2">
<label>Figure 2</label>
<caption><p>The collaborative task in the corpus (<xref ref-type="bibr" rid="B128">Zarrie&#x000DF; et al., 2016</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0002.tif">
<alt-text content-type="machine-generated">Four-panel comic featuring blue Tetris-like shapes on a grid. Panel 1: A cursor points at a blue shape resembling the letter &#x0201C;V&#x0201D; with text &#x0201C;Okay so next one is the one that looks like the letter V.&#x0201D; Panel 2: A red arrow indicates a different shape with text &#x0201C;It is like...3 by 3 blocks?&#x0201D; Panel 3: A shape resembling an arrow is highlighted with the text &#x0201C;Looks a bit like an arrow.&#x0201D; Panel 4: The correct shape is highlighted in green in the left corner with the text &#x0201C;Yeah, right there on the left corner!&#x0201D; </alt-text>
</graphic>
</fig>
</sec>
<sec>
<label>2.2</label>
<title>Instructions in incremental units</title>
<p>The corpus included utterance-level annotations, increment transcripts, and the types of referring strategies used to identify the pentomino pieces (<xref ref-type="bibr" rid="B102">Schlangen and Fern&#x000E1;ndez, 2008</xref>). A transcript of a fragment of an interaction is shown below, with the incremental units highlighted in <bold>bold</bold>:</p>
<list list-type="simple">
<list-item><p>INSTRUCTOR: <italic>letter</italic> <bold>There&#x00027;s a piece like an L shape</bold>.</p></list-item>
<list-item><p>USER: Mhm, yeah!</p></list-item>
<list-item><p>INSTRUCTOR: <italic>geometrical shape</italic> <bold>Where you know one piece is longer than the other</bold>.</p></list-item>
<list-item><p>INSTRUCTOR: <italic>blocks</italic> <bold>It&#x00027;s about four units by two units</bold>.</p></list-item>
<list-item><p>USER: This one?</p></list-item>
<list-item><p>USER: So, you, you can see it when I&#x00027;m moving it here?</p></list-item>
<list-item><p>INSTRUCTOR: No. I, I just see the solution, yeah?</p></list-item>
<list-item><p>INSTRUCTOR: I&#x00027;m looking at an elephant, believe it or not.</p></list-item>
<list-item><p>INSTRUCTOR: Okay, so eh.</p></list-item>
<list-item><p>USER: -laughter-</p></list-item>
<list-item><p>INSTRUCTOR: <italic>elephant solution</italic> <bold>It&#x00027;s like the back leg</bold>.</p></list-item>
<list-item><p>INSTRUCTOR: <italic>location</italic> <bold>The bottom right of the ... grid</bold>.</p></list-item>
</list>
<p>An important aspect of using human instruction utterances is that they are generated in a collaborative manner. We used the incremental units as templates to generate computer instructions. Instructors typically altered the referring strategy [type of referent attribute (<xref ref-type="bibr" rid="B91">Reigeluth et al., 1980</xref>)] with each new increment. We extracted and modeled the incremental units from a total of 3,174 referring expressions.</p>
</sec>
<sec>
<label>2.3</label>
<title>Automatic instruction generation</title>
<p>We analyzed the instructions by examining their turn-taking characteristics. The average number of incremental units per instruction was 2.0 &#x000B1; 1.5, with a minimum of 1 and a maximum of 5 incremental units. The average incremental unit duration was 2.4<italic>s</italic>&#x000B1;1.6<italic>s</italic>, with an average of 6.6 &#x000B1; 6.3 words per unit and a pause duration of 1.4<italic>s</italic>&#x000B1;2.2<italic>s</italic> between units. To utilize these utterances in the studies, we filtered the data by selecting the five most frequently used strategies (corresponding to the maximum number of incremental in this corpus). We also removed outliers from the data (e.g., very long pauses, very long utterances) by filtering data points more than two standard deviations away from the mean, resulting in a total of 1,588 utterances.</p>
<p>We then cleaned the utterances for disfluencies or prosodic information included in the annotations. A random set of utterances was selected and checked for coherence. Next, we grouped the utterances and filtered out the referent object to use as templates. Aside from the referent-specific information, the remaining utterance attributes, including the syntactic structure, were preserved as originally spoken by the instructors. The agent only needed to select a human utterance and adjust it to the current referent target.</p></sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Study 1: The maxim of quantity</title>
<p>We conducted a study with an interface instructing humans to evaluate the effectiveness of the incremental units observed in the corpus. Participants were exposed to two types of stimuli: (i) <italic>visual</italic> (the pentomino pieces), and (ii) <italic>auditory</italic> (task instructions).</p>
<sec>
<label>3.1</label>
<title>Materials and methods</title>
<sec>
<label>3.1.1</label>
<title>Implementation</title>
<p>We implemented a web version of the Pentomino task, where each scene included the referent target object among a set of distractor objects. Instructions were generated for each Pentomino and referring strategy using Amazon Polly Text-to-Speech (TTS), and we created the agent <italic>Matthew</italic>. All participants were exposed to the same stimuli. Task boards were generated with Pentomino pieces randomly positioned (see <xref ref-type="fig" rid="F3">Figure 3</xref>). Matthew delivered an incremental unit for each of the five referring strategies observed in the human corpus (transcript in Section 2.2).</p>
<fig position="float" id="F3">
<label>Figure 3</label>
<caption><p>Illustration of a voice user interface incrementally constructing instructions based on the user&#x00027;s signals.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0003.tif">
<alt-text content-type="machine-generated">Grid of Tetris-like blue and green shapes at the top. In the lower section, there's a pixelated grid with a yellow area. It includes headshots of two people labeled &#x0201C;Instructor&#x0201D; and &#x0201C;User,&#x0201D; their eyes obscured.</alt-text>
</graphic>
</fig>
<p>Although current off-the-shelf TTS services provide limited control over prosodic variation (<xref ref-type="bibr" rid="B110">Sz&#x000E9;kely et al., 2019</xref>), we manipulated the end of each incremental unit by introducing a rising pitch to suggest that the agent might continue with an additional utterance (<xref ref-type="bibr" rid="B118">Traum and Hinkelman, 1992</xref>; <xref ref-type="bibr" rid="B18">Brennan and Schober, 2001</xref>). After each unit, the agent paused. To determine the duration of these pauses (<xref ref-type="bibr" rid="B130">Zellner, 1994</xref>), we used the timing data (mean and standard deviation) from the human instructor corpus to decide how long to wait before delivering the next incremental unit. While most of the pauses were silent, in some instances, Matthew generated filled pauses (e.g., &#x0201C;<italic>uhm,&#x0201D; &#x0201C;uh&#x0201D;</italic>), based on their occurrence rate in the human corpus (22% of the pauses). The interface monitored participants&#x00027; visual attention (<xref ref-type="bibr" rid="B87">M&#x000FC;ller and Krummenacher, 2006</xref>; <xref ref-type="bibr" rid="B38">Eckstein, 2011</xref>) through their mouse movements.</p></sec>
<sec>
<label>3.1.2</label>
<title>Independent variables</title>
<p>In this study, we investigated the amount of information that needs to be conveyed in instructions. We hypothesized that even when subjects are exposed to the same amount of information, some individuals might be more dependent on the adaptation to their specific information needs. The aim was to estimate the minimal amount of information required to achieve high accuracy and to determine whether additional information is always beneficial or potentially detrimental. We expected to observe a trade-off between the spoken effort exerted by the interface and the users&#x00027; accuracy in the task.</p>
<p>We manipulated the amount of information provided to subjects by using the minimum (1) and maximum (5) number of incremental units employed by human instructors and tested five variations of instructions. To control for order effects, we also tested each of the five referring strategies appearing first, resulting in a combination of 25 (5 incremental units &#x000D7; 5 referring strategies) instructions for each of the Pentominoes. In total, 300 instructions were evaluated (25 &#x000D7; 12 Pentominoes) using a balanced Latin Square design.</p></sec>
<sec>
<label>3.1.3</label>
<title>Dependent variables</title>
<sec>
<label>3.1.3.1</label>
<title>Behavioral measures</title>
<p><bold>User actions:</bold> We measured users&#x00027; overall <bold>accuracy</bold> in the task (percentage of correctly identified objects). For each instruction, we also recorded users&#x00027; <bold>response time</bold>, indicating the effort spent on visual search, as well as <bold>idle time</bold>, which represents the time taken by users to initiate a mouse movement, and whether the user&#x00027;s <bold>mouse had moved</bold> as a binary feature. <bold>Mouse uncertainty:</bold> For each incremental unit, we also calculated a set of features representing the user&#x00027;s mouse movement uncertainty (see <xref ref-type="table" rid="T1">Table 1</xref>). These features allowed us to estimate the degree of unpredictability in the mouse movements as a proxy for the user&#x00027;s attention (<xref ref-type="bibr" rid="B44">Fitts, 1954</xref>) (see <xref ref-type="fig" rid="F4">Figure 4</xref>). The continuous signal of mouse movements was extracted every 200ms and concatenated to represent the mouse movement during the incremental unit, while also preserving the temporal dynamics; moving away from or toward the target piece can be interpreted as a user&#x00027;s display of understanding. Mouse movements have been utilized in cognitive psychology (<xref ref-type="bibr" rid="B119">Wachsmuth et al., 2008</xref>; <xref ref-type="bibr" rid="B114">Tomlinson Jr and Bott, 2013</xref>; <xref ref-type="bibr" rid="B113">Tomlinson Jr and Assimakopoulos, 2013</xref>; <xref ref-type="bibr" rid="B126">Xiao and Yamauchi, 2014</xref>; <xref ref-type="bibr" rid="B21">Calcagn&#x00300;&#x00131; et al., 2017</xref>; <xref ref-type="bibr" rid="B93">Rheem et al., 2018</xref>; <xref ref-type="bibr" rid="B58">Horwitz et al., 2020</xref>; <xref ref-type="bibr" rid="B104">Schoemann et al., 2021</xref>) to assess cognitive load, as well as in HCI (<xref ref-type="bibr" rid="B124">Whisenand and Emurian, 1999</xref>; <xref ref-type="bibr" rid="B86">Mueller and Lockerd, 2001</xref>; <xref ref-type="bibr" rid="B5">Ashdown et al., 2005</xref>; <xref ref-type="bibr" rid="B4">Arroyo et al., 2006</xref>; <xref ref-type="bibr" rid="B54">Guo and Agichtein, 2010</xref>; <xref ref-type="bibr" rid="B35">Diaz et al., 2013</xref>; <xref ref-type="bibr" rid="B84">Monaro et al., 2017</xref>; <xref ref-type="bibr" rid="B67">Kieslich et al., 2019</xref>; <xref ref-type="bibr" rid="B77">Krassanakis and Kesidis, 2020</xref>) and information retrieval (<xref ref-type="bibr" rid="B53">Guo and Agichtein, 2008</xref>; <xref ref-type="bibr" rid="B61">Huang et al., 2011</xref>, <xref ref-type="bibr" rid="B60">2012b</xref>; <xref ref-type="bibr" rid="B19">Br&#x000FC;ckner et al., 2021</xref>) to identify user attention and engagement (<xref ref-type="bibr" rid="B64">Johnson et al., 2012</xref>; <xref ref-type="bibr" rid="B106">Smucker et al., 2014</xref>; <xref ref-type="bibr" rid="B1">Arapakis and Leiva, 2016</xref>, <xref ref-type="bibr" rid="B2">2020</xref>; <xref ref-type="bibr" rid="B3">Arapakis et al., 2020</xref>; <xref ref-type="bibr" rid="B69">Kirsh, 2020</xref>), often showing a correlation with eye movements (<xref ref-type="bibr" rid="B23">Chen et al., 2001</xref>; <xref ref-type="bibr" rid="B59">Huang et al., 2012a</xref>; <xref ref-type="bibr" rid="B90">Qvarfordt, 2017</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Behavioral and subjective measures.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center" colspan="2"><bold>Behavioral: user actions</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Accuracy</td>
<td valign="top" align="left">Percentage of correct pieces selected</td>
</tr>
<tr>
<td valign="top" align="left">Response time</td>
<td valign="top" align="left">Time it took subjects to select a piece</td>
</tr>
<tr>
<td valign="top" align="left">Idle time</td>
<td valign="top" align="left">Time it took subjects to initiate a mouse movement</td>
</tr>
<tr>
<td valign="top" align="left">Mouse Moved</td>
<td valign="top" align="left">Binary feature indicating the mouse has moved</td>
</tr>
<tr>
<td valign="top" align="left" colspan="2"><bold>Behavioral: mouse features</bold></td>
</tr>
<tr>
<td valign="top" align="left">Distance to target (mean)</td>
<td valign="top" align="left">Average linear distance to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Distance to target (std)</td>
<td valign="top" align="left">Standard deviation of distance to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Distance to target (min)</td>
<td valign="top" align="left">Minimum distance to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Distance to target (max)</td>
<td valign="top" align="left">Maximum distance to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Distance to target (range)</td>
<td valign="top" align="left">Range of distance to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Distance to target (slope)</td>
<td valign="top" align="left">Slope of distance to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Distance traveled</td>
<td valign="top" align="left">Total mouse distance traveled</td>
</tr>
<tr>
<td valign="top" align="left">Velocity (mean)</td>
<td valign="top" align="left">Velocity of mouse movement</td>
</tr>
<tr>
<td valign="top" align="left">Direction change X</td>
<td valign="top" align="left">Number of changes in trajectory direction (X-axis)</td>
</tr>
<tr>
<td valign="top" align="left">Direction change Y</td>
<td valign="top" align="left">Number of changes in trajectory direction (Y-axis)</td>
</tr>
<tr>
<td valign="top" align="left">Area under the curve (AUC)</td>
<td valign="top" align="left">AUC from observed mouse trajectory to direct path to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Mean absolute deviation</td>
<td valign="top" align="left">Mean absolute deviation to direct path to target piece</td>
</tr>
<tr>
<td valign="top" align="left">Max absolute deviation</td>
<td valign="top" align="left">Max absolute deviation to direct path to target piece</td>
</tr>
<tr>
<td valign="top" align="left" colspan="2"><bold>Subjective measures</bold></td>
</tr>
<tr>
<td valign="top" align="left">Ambiguousness</td>
<td valign="top" align="left">Matthew&#x00027;s instruction was [<italic>unambiguous / ambiguous</italic>]</td>
</tr>
<tr>
<td valign="top" align="left">Human-likeness</td>
<td valign="top" align="left">Matthew&#x00027;s instruction was [<italic>machine-like / human-like</italic>]</td>
</tr>
<tr>
<td valign="top" align="left">Information</td>
<td valign="top" align="left">Matthew&#x00027;s instruction had [<italic>too little / too much</italic>] information</td>
</tr>
<tr>
<td valign="top" align="left">Effort</td>
<td valign="top" align="left">Matthew put [<italic>too little / too much</italic>] effort in this instruction</td>
</tr></tbody>
</table>
</table-wrap>
<fig position="float" id="F4">
<label>Figure 4</label>
<caption><p>Area under the curve (AUC) of the observed mouse trajectory compared to the direct path toward the referent.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0004.tif">
<alt-text content-type="machine-generated">Diagram illustrating a maze with blue blocks, showing a mouse's start position. A red dashed line indicates the observed trajectory, while a black dashed line shows the direct path. A blue arrow pointing to &#x0201C;AUC&#x0201D; suggests an alternative route.</alt-text>
</graphic>
</fig>
</sec>
<sec>
<label>3.1.3.2</label>
<title>Subjective measures</title>
<p>Users were asked to respond to four <italic>7-point Likert-scale questions</italic> (see <xref ref-type="table" rid="T1">Table 1</xref>) regarding the <bold>Instruction Appropriateness</bold>.</p></sec></sec>
<sec>
<label>3.1.4</label>
<title>Statistical analyses</title>
<p>We conducted statistical analyses in R (<xref ref-type="bibr" rid="B111">Team, 2009</xref>). Using the <italic>lme4, lmerTest, glmmTMB</italic> packages (<xref ref-type="bibr" rid="B9">Bates et al., 2014</xref>), we fitted linear mixed-effects models (LMMs) and generalized linear mixed-effects models (GLMMs) to examine the relationship between <italic>the number of incremental units</italic> uttered (fixed factor with five levels), <italic>instruction correctness</italic> (fixed factor with two levels), and the dependent variables. As random effects, we included intercepts for the <italic>participants</italic>, the <italic>pentomino pieces</italic>, the <italic>referring strategies</italic>, and the <italic>type of mouse device</italic> used. Continuous dependent variables were analyzed with LMMs after log transformation, where appropriate. Binary outcomes were analyzed with GLMMs using a binomial link. We opted for LMMs due to their ability to model variance in the data, such as the variability in mouse behavior across users, using the following notation: <italic>DV</italic> &#x0007E; <italic>IncrementalUnits * InstructionCorrectness &#x0002B; (1|Participant) &#x0002B; (1|PentominoOrder) &#x0002B; (1|ReferringStrategy) &#x0002B; (1|MouseType)</italic>. Model assumptions were validated using the DHARMa package, and effect sizes were reported as standardized &#x003B2; (LMMs) or odds ratios (GLMMs) with 95% confidence intervals. Participants were not restricted on when they could select pieces. For a small subset of the data (11%), participants selected a pentomino before hearing the complete instruction; therefore, we included <italic>interruption</italic> as a confounding factor to account for this variance in the model. Maximum likelihood estimation tests were used to determine the chi-square and p-values, comparing the null models to the full models.</p></sec>
<sec>
<label>3.1.5</label>
<title>Procedure and data collection</title>
<p>At the beginning of the task, Matthew asked participants for their informed consent. After each instruction, Matthew indicated whether the instruction was correct and placed the piece on the elephant structure accordingly. At the end of the interaction, Matthew asked participants to provide their demographic data. Participants were debriefed on the study&#x00027;s purpose and the experimental manipulations.</p>
<p>Eighty participants were recruited online (<xref ref-type="bibr" rid="B39">Eerola et al., 2021</xref>). Five participants did not fully complete the task or experienced technical issues and were excluded, resulting in a total of 75 participants. The participants evaluated a total of 900 instructions, with 2,700 incremental units used as data points in the statistical and machine-learning models. The mean age of the participants was 31.8 (&#x000B1;6.9) years, with 30 identifying as female and 45 as male. Their self-reported English fluency was 6.3 (&#x000B1;0.9) on a scale of 1 to 7. The task took, on average, 18.0 (&#x000B1;8.2) minutes to complete, and each instruction lasted, on average, 11.8 (&#x000B1;5.2) seconds. Forty-four participants used a mouse, while 31 participants used a trackpad.</p></sec>
<sec>
<label>3.1.6</label>
<title>Manipulation check</title>
<p>Since the stimuli used were synthesized human instructions, we had limited control over the amount of information conveyed in each incremental unit. We used the number of words spoken as a proxy for the information transmitted to determine whether the stimulus was consistent across incremental units. Fitting linear mixed-effects models indicated that there were no significant differences in the number of words spoken per incremental unit (6.9 &#x000B1; 3.5): &#x003C7;<sup>2</sup> &#x0003D; 1.42, <italic>p</italic>&#x0003E;0.05, with a marginal <italic>R</italic><sup>2</sup> of 0.001, suggesting that each incremental unit carried approximately the same amount of information (as indicated by the number of words). To further assess the informational similarity between utterances, we computed the average pairwise semantic similarity using Sentence-BERT embeddings (all-MiniLM-L6-v2) and cosine similarity, which yielded a mean value of 0.29, indicating that the utterances were moderately similar in meaning; not identical, yet sharing some overlap in informational content across incremental units. However, it remains subjective as to what information is considered ambiguous in this paradigm, which we aim to evaluate in this study.</p>
</sec>
</sec>
<sec>
<label>3.2</label>
<title>Results</title>
<sec>
<label>3.2.1</label>
<title>Behavioral measures</title>
<p><bold>Effects of user actions: Accuracy</bold>. On average, participants had an accuracy of 70.8%, with 8.5 (&#x000B1;1.4) out of 12 referents correctly identified. Fitting generalized linear mixed models revealed a statistically significant difference in the number of correct pieces identified per incremental unit (see <xref ref-type="fig" rid="F5">Figure 5</xref>), with the mean values presented in <xref ref-type="table" rid="T2">Table 2</xref>, indicating that more information presented led to better performance, however without a clear indication of additional information being perceived as overwhelming. We also tested the effect of referring strategies order, which showed a statistically significant impact: &#x003C7;<sup>2</sup> &#x0003D; 27.87, <italic>p</italic> &#x0003C; 0.001. <bold>Idle &#x00026; response time</bold>. The mean idle time was 9.09 (&#x000B1;10.2) seconds, and the mean response time was 14.7 (&#x000B1;11.4) seconds. LMMs (with log transformation) indicated that both idle and response times were statistically different across incremental units, with a rising trend in time (see <xref ref-type="table" rid="T2">Table 2</xref> and <xref ref-type="fig" rid="F6">Figure 6</xref>). The models also showed that users were faster at selecting pieces when they selected the correct piece. Through GLMMs, the <bold>Mouse moved</bold> measure was found to be statistically different across incremental units, with users&#x00027; mouse movements starting early during the instruction (see <xref ref-type="table" rid="T2">Table 2</xref>).</p>
<fig position="float" id="F5">
<label>Figure 5</label>
<caption><p>Users&#x00027; accuracy in identifying the correct pentomino pieces per incremental unit. Based on the number of incremental units spoken, the interface can estimate the probability of the user correctly identifying the referent.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0005.tif">
<alt-text content-type="machine-generated">Bar chart showing the percentage of correct pentominoes picked across different incremental units. The units range from one to five, each represented by a different color: red, blue, green, purple, and orange. Percentages increase with more incremental units, ranging from approximately 56% to 78%, with error bars displayed for each unit.</alt-text>
</graphic>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Behavioral and subjective measures for each incremental unit, with unit means (1-5) comparing the null model to the full model.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><bold>Predictor</bold></th>
<th valign="top" align="center"><bold>I1</bold></th>
<th valign="top" align="center"><bold>I2</bold></th>
<th valign="top" align="center"><bold>I3</bold></th>
<th valign="top" align="center"><bold>I4</bold></th>
<th valign="top" align="center"><bold>I5</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
<th valign="top" align="center"><bold>Chi-square</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Accuracy (%)</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.029</td>
<td valign="top" align="center">21.209</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Idle Time</td>
<td valign="top" align="center">8.38</td>
<td valign="top" align="center">8.59</td>
<td valign="top" align="center">9.70</td>
<td valign="top" align="center">8.97</td>
<td valign="top" align="center">10.33</td>
<td valign="top" align="center">0.035</td>
<td valign="top" align="center">35.083</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Response Time</td>
<td valign="top" align="center">12.6</td>
<td valign="top" align="center">13.1</td>
<td valign="top" align="center">15.1</td>
<td valign="top" align="center">15.7</td>
<td valign="top" align="center">17.3</td>
<td valign="top" align="center">0.291</td>
<td valign="top" align="center">232.71</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Moved (%)</td>
<td valign="top" align="center">0.99</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.861</td>
<td valign="top" align="center">215.39</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (mean)</td>
<td valign="top" align="center">852</td>
<td valign="top" align="center">898</td>
<td valign="top" align="center">852</td>
<td valign="top" align="center">799</td>
<td valign="top" align="center">784</td>
<td valign="top" align="center">0.009</td>
<td valign="top" align="center">23.291</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (std)</td>
<td valign="top" align="center">386</td>
<td valign="top" align="center">203</td>
<td valign="top" align="center">159</td>
<td valign="top" align="center">119</td>
<td valign="top" align="center">112</td>
<td valign="top" align="center">0.112</td>
<td valign="top" align="center">280.87</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (min)</td>
<td valign="top" align="center">126</td>
<td valign="top" align="center">551</td>
<td valign="top" align="center">622</td>
<td valign="top" align="center">618</td>
<td valign="top" align="center">621</td>
<td valign="top" align="center">0.056</td>
<td valign="top" align="center">142.65</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (max)</td>
<td valign="top" align="center">1,163</td>
<td valign="top" align="center">1,086</td>
<td valign="top" align="center">1,027</td>
<td valign="top" align="center">930</td>
<td valign="top" align="center">912</td>
<td valign="top" align="center">0.033</td>
<td valign="top" align="center">87.929</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (range)</td>
<td valign="top" align="center">1030</td>
<td valign="top" align="center">539</td>
<td valign="top" align="center">407</td>
<td valign="top" align="center">314</td>
<td valign="top" align="center">290</td>
<td valign="top" align="center">0.133</td>
<td valign="top" align="center">339.02</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (slope)</td>
<td valign="top" align="center">&#x02013;24.8</td>
<td valign="top" align="center">&#x02013;16.8</td>
<td valign="top" align="center">&#x02013;15.2</td>
<td valign="top" align="center">&#x02013;14.3</td>
<td valign="top" align="center">&#x02013;12.2</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">31.331</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance traveled</td>
<td valign="top" align="center">1430</td>
<td valign="top" align="center">723</td>
<td valign="top" align="center">531</td>
<td valign="top" align="center">399</td>
<td valign="top" align="center">370</td>
<td valign="top" align="center">0.142</td>
<td valign="top" align="center">364.27</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Velocity</td>
<td valign="top" align="center">32.9</td>
<td valign="top" align="center">23.1</td>
<td valign="top" align="center">21.0</td>
<td valign="top" align="center">18.4</td>
<td valign="top" align="center">17.5</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">36.204</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Dir. change (x)</td>
<td valign="top" align="center">2.655</td>
<td valign="top" align="center">1.328</td>
<td valign="top" align="center">1.045</td>
<td valign="top" align="center">0.899</td>
<td valign="top" align="center">0.865</td>
<td valign="top" align="center">0.070</td>
<td valign="top" align="center">167.97</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Dir. change (y)</td>
<td valign="top" align="center">2.915</td>
<td valign="top" align="center">1.473</td>
<td valign="top" align="center">1.177</td>
<td valign="top" align="center">0.981</td>
<td valign="top" align="center">0.974</td>
<td valign="top" align="center">0.081</td>
<td valign="top" align="center">195.15</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">AUC</td>
<td valign="top" align="center">33,415</td>
<td valign="top" align="center">228,324</td>
<td valign="top" align="center">259,276</td>
<td valign="top" align="center">254,554</td>
<td valign="top" align="center">257,854</td>
<td valign="top" align="center">0.050</td>
<td valign="top" align="center">136.72</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Mean abs. deviation</td>
<td valign="top" align="center">86.8</td>
<td valign="top" align="center">95.5</td>
<td valign="top" align="center">90.1</td>
<td valign="top" align="center">90.7</td>
<td valign="top" align="center">90.5</td>
<td valign="top" align="center">0.067</td>
<td valign="top" align="center">146.35</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Max abs. deviation</td>
<td valign="top" align="center">139</td>
<td valign="top" align="center">143</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">122</td>
<td valign="top" align="center">121</td>
<td valign="top" align="center">0.055</td>
<td valign="top" align="center">108.89</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Ambiguousness</td>
<td valign="top" align="center">4.74</td>
<td valign="top" align="center">4.15</td>
<td valign="top" align="center">4.03</td>
<td valign="top" align="center">3.71</td>
<td valign="top" align="center">3.90</td>
<td valign="top" align="center">0.255</td>
<td valign="top" align="center">192.48</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Human-likeness</td>
<td valign="top" align="center">4.08</td>
<td valign="top" align="center">4.23</td>
<td valign="top" align="center">4.32</td>
<td valign="top" align="center">4.40</td>
<td valign="top" align="center">4.09</td>
<td valign="top" align="center">0.047</td>
<td valign="top" align="center">36.224</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Information</td>
<td valign="top" align="center">2.52</td>
<td valign="top" align="center">3.17</td>
<td valign="top" align="center">3.57</td>
<td valign="top" align="center">3.95</td>
<td valign="top" align="center">4.15</td>
<td valign="top" align="center">0.330</td>
<td valign="top" align="center">330.16</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Effort</td>
<td valign="top" align="center">2.60</td>
<td valign="top" align="center">3.37</td>
<td valign="top" align="center">3.69</td>
<td valign="top" align="center">3.99</td>
<td valign="top" align="center">4.13</td>
<td valign="top" align="center">0.329</td>
<td valign="top" align="center">287.27</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p><italic>P</italic>-value indicators: &#x02212;<italic>p</italic>&#x0003E;0.05, <sup>&#x0002A;</sup><italic>p</italic> &#x02264; 0.05, <sup>&#x0002A;&#x0002A;</sup><italic>p</italic> &#x02264; 0.01, <sup>&#x0002A;&#x0002A;&#x0002A;</sup><italic>p</italic> &#x02264; 0.001.</p>
</table-wrap-foot>
</table-wrap>
<fig position="float" id="F6">
<label>Figure 6</label>
<caption><p><bold>(a)</bold> Perceived information amount separated by correctness. <bold>(b)</bold> Mean distance to target, indicating the probability of the mouse position being on or close to the target. <bold>(c)</bold> Response time separated by correctness.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0006.tif">
<alt-text content-type="machine-generated">Chart A shows information increasing with the number of units heard, with correct picks (blue circles) higher than incorrect (red triangles). Chart B displays mean distance to target decreasing, correct choices generally closer. Chart C illustrates response time comparisons between correct (brown) and incorrect (blue) picks, showing variability across units. Each chart includes error bars.</alt-text>
</graphic>
</fig>
<p><bold>Effects of mouse uncertainty</bold>. Linear mixed-effects models revealed statistically significant differences in how participants utilized mouse movements and how uncertainty was expressed when selecting pieces during each incremental unit (see <xref ref-type="table" rid="T2">Table 2</xref> and <xref ref-type="fig" rid="F6">Figure 6</xref>). Consistent with users&#x00027; accuracy in the task, mouse movements indicated that less uncertainty was associated with a higher number of incremental units spoken by the interface.</p></sec>
<sec>
<label>3.2.2</label>
<title>Subjective measures</title>
<p>Fitting linear mixed-effects models on <bold>Instruction Appropriateness</bold> revealed statistically significant effects on how incremental units were perceived (see <xref ref-type="table" rid="T2">Table 2</xref>). The findings indicated that there were differences in how the amount of information and ambiguity were perceived, with selection accuracy influencing whether or not participants believed additional information was necessary. Perceived information amount was strongly affected by whether a user made a correct selection (see <xref ref-type="fig" rid="F6">Figure 6</xref>); however, the actual amount of information provided remained constant, regardless of the user&#x00027;s performance.</p>
</sec>
</sec>
<sec>
<label>3.3</label>
<title>Estimating user uncertainty</title>
<p>To estimate whether each instruction was ambiguous, we trained two Random Forest (RF) classifiers using the Scikit-Learn framework (<xref ref-type="bibr" rid="B88">Pedregosa et al., 2011</xref>). We selected RFs for their robustness against overfitting and their interpretability in identifying the most informative features. Using mouse features, we were able to estimate user uncertainty and predict whether users were likely to succeed. The first model was a <italic>speaker-based</italic> model, where the system &#x02018;<italic>looks back&#x00027;</italic> at what it has said and predicts whether the reference will be successful. We employed Sentence-BERT (<xref ref-type="bibr" rid="B92">Reimers and Gurevych, 2019</xref>) to convert each instruction into a 384-dimensional vector. The second model was a <italic>listener-based</italic> model that incorporated both the instruction embeddings and the user&#x00027;s mouse movements (see <xref ref-type="table" rid="T1">Table 1</xref>). Both models utilized these features to estimate the user&#x00027;s confidence in their selection; an adaptive interface with this knowledge can incrementally determine whether the user requires additional information.</p>
<p>To better understand the models, we calculated the most informative features in the classification task. Both classifiers were evaluated using subject-independent 10-fold cross-validation. For each incremental unit, we utilized the users&#x00027; piece selections as ground truth in the models. The underlying assumption is that either the user&#x00027;s attention or the adequacy of the incremental unit contains information that leads to correct actions, which a machine learning model can leverage.</p>
<p>A total of 900 instructions were used. The models extracted features within the sampling window between incremental units. Since the classification classes were imbalanced, we applied the Scikit-Learn <italic>resampling</italic> method (<xref ref-type="bibr" rid="B88">Pedregosa et al., 2011</xref>) to re-sample the majority-class segments, balancing the dataset. This process resulted in 1,266 data points, with an average sampling window of 5.6 (&#x000B1;7.5) seconds (between turns) and a total duration of about 2 h of mouse-movement data. For evaluation metrics, we report the average Accuracy of the models, as well as Precision, Recall, and F1 scores.</p>
<sec>
<label>3.3.1</label>
<title>Speaker-based model</title>
<p>Using the semantic representations, the speaker-based model simulated a human speaker self-repairing their utterance (<xref ref-type="bibr" rid="B74">Kontogiorgos et al., 2019</xref>), essentially evaluating whether it was a good instruction based on its linguistic features. Hyperparameters for the Random Forest model were optimized using grid search: [max-depth = 110, max-features = &#x0201C;auto,&#x0201D; min-samples-leaf = 4, min-samples-split = 10, n-estimators = 100]. The results, presented in <xref ref-type="table" rid="T3">Table 3</xref>, show better-than-chance accuracy, although relatively low.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Summary of the performance of the machine learning models.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="left"><bold>Features</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="left">[random chance]</td>
<td valign="top" align="center">0.5</td>
<td valign="top" align="center">0.5</td>
<td valign="top" align="center">0.5</td>
<td valign="top" align="center">0.5</td>
</tr>
<tr>
<td valign="top" align="left">Speaker-based</td>
<td valign="top" align="left">BERT x384</td>
<td valign="top" align="center">0.61 (&#x000B1;0.05)</td>
<td valign="top" align="center">0.60 (&#x000B1;0.05)</td>
<td valign="top" align="center">0.66 (&#x000B1;0.05)</td>
<td valign="top" align="center">0.63 (&#x000B1;0.05)</td>
</tr>
<tr>
<td valign="top" align="left">Listener-based</td>
<td valign="top" align="left">BERT x384 &#x00026; Mouse Movements</td>
<td valign="top" align="center"><bold>0.87</bold> (&#x000B1;0.04)</td>
<td valign="top" align="center"><bold>0.81</bold> (&#x000B1;0.05)</td>
<td valign="top" align="center"><bold>0.96</bold> (&#x000B1;0.03)</td>
<td valign="top" align="center"><bold>0.88</bold> (&#x000B1;0.03)</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Highest performance indicated in bold.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<label>3.3.2</label>
<title>Listener-based model</title>
<p>The hyperparameters for the listener-based model were optimized using grid search: [max-depth = 100, max-features = &#x0201C;auto,&#x0201D; min-samples-leaf = 3, min-samples-split = 12, n-estimators = 100]. This model yielded better results (see <xref ref-type="table" rid="T3">Table 3</xref>), indicating that paying attention to the listener may better simulate human speaker behavior. A <italic>post-hoc</italic> examination of the model&#x00027;s features revealed that the mouse features combined had an F-score of 0.67, compared to 0.33 for the linguistic features.</p>
</sec>
</sec>
<sec>
<label>3.4</label>
<title>Discussion</title>
<p>The focus of this study was to investigate the <italic>maxim of quantity</italic> (<xref ref-type="bibr" rid="B51">Grice, 1975</xref>), specifically the amount of relevant information presented to participants to successfully disambiguate referring expressions. The findings indicated that information adaptation is crucial, as users require utterances adapted to their information needs. The fact that users rated the information received differently is a significant insight for information adaptation, suggesting that they assess the amount of perceived information based on their task performance rather than the actual information received.</p>
<p>While we initially expected that providing more information might overwhelm users, the results did not clearly support that assumption, as users&#x00027; accuracy did not consistently reflect this effect. We also observed differences in idle and response times; participants with higher accuracy responded more quickly, indicating that slower response times were associated with higher cognitive effort.</p>
<p>We also observed statistically significant differences in how participants rated the interface. We had predicted that users would not always achieve high accuracy, as the instructions provided were sometimes incomplete. However, focusing solely on accuracy could lead to designing an agent that continuously provides information until all ambiguity is resolved, regardless of how overwhelming this might be for users. This highlights a challenge that socially intelligent agents must address: maximizing accuracy while minimizing collaborative effort (investigated in Study 2).</p>
<p>Finally, the two prediction models estimated user uncertainty, demonstrating strong performance in identifying which types of features (linguistic vs. behavioral) an interface should focus on to deliver incremental units where turn transitions are &#x0201C;<italic>relevant&#x0201D;</italic> (<xref ref-type="bibr" rid="B96">Sacks et al., 1978</xref>).</p></sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Study 2: The principle of least collaborative effort</title>
<p>Study 2 investigated how to elicit adaptive behavior through instructions using the models trained in Study 1. A significant difference between the two Studies was that Study 2 examined adaptive instruction strategies, while Study 1 utilized predetermined structures of instructions evaluated by users.</p>
<sec>
<label>4.1</label>
<title>Materials and methods</title>
<sec>
<label>4.1.1</label>
<title>System design</title>
<p>We utilized the same web application, modified to incorporate the ML models. We took a new sample of the human instructions and synthesized them using the same TTS method described in Study 1. We created three agents (<italic>Kevin, David</italic>, and <italic>Peter</italic>), each corresponding to a variation of the three models evaluated. Each agent began with a single incremental unit and then monitored the user to determine whether additional information was necessary. The order of referring strategies was based on their frequency of usage in the human corpus. We deployed the machine-learning models on the interface using the <italic>Sklearn-Porter</italic> open-source framework (<xref ref-type="bibr" rid="B85">Morawiec, 2021</xref>).</p></sec>
<sec>
<label>4.1.2</label>
<title>Independent variables</title>
<p>We evaluated three separate models for planning spoken instructions. We hypothesized that a model that monitors the user would have an advantage over a model that only monitors what is being spoken. We compared these two models to a control condition in which users actively indicated whether a new instruction unit should be spoken (Tell-Me-More). A total of 60 CUI instruction units (same for each condition) were evaluated (5 incremental units &#x000D7; 12 Pentominoes). This study aimed to examine the optimal model for providing instructions, investigating whether adaptive models that follow the principles of least collaborative effort offer a benefit in the interaction.</p>
<sec>
<label>4.1.2.1</label>
<title>Adaptive baseline (Tell-Me-More)</title>
<p>This interface only responded to the need for additional information when manually prompted by the user. Using the same pausing behavior, the interface displayed a Tell-Me-More button at the end of the pause, waiting for the user to indicate if a new instruction unit should be spoken. Pressing the button can be seen as the user&#x00027;s continuous attempt to establish common ground (<xref ref-type="bibr" rid="B47">Garoufi et al., 2016</xref>). Users received a higher payment for correct answers across all conditions; however, in this baseline, they were informed that each press of the button would reduce their bonus payment. Through this process, users aimed to achieve as many correct answers as possible with the minimal amount of information required.</p></sec>
<sec>
<label>4.1.2.2</label>
<title>Speaker-based model</title>
<p>We extracted and employed the speaker-based model from Study 1. As long as the model predicted that the participant was unlikely to be successful, it would respond with phrases like &#x0201C;<italic>no&#x0201D;</italic>, &#x0201C;<italic>not this one&#x0201D;</italic>, or &#x0201C;<italic>hm&#x0201D;</italic> before proceeding to the next instruction unit (<xref ref-type="bibr" rid="B95">Rookhuiszen et al., 2009</xref>; <xref ref-type="bibr" rid="B83">Mitev et al., 2018</xref>). When the model predicted that the user would be successful, it would respond with &#x0201C;<italic>yeah&#x0201D;</italic>, &#x0201C;<italic>yes&#x0201D;</italic>, or &#x0201C;<italic>yup&#x0201D;</italic>, followed by the next instruction unit. We anticipated that the low accuracy of this model would induce additional uncertainty in user behavior.</p></sec>
<sec>
<label>4.1.2.3</label>
<title>Listener-based model</title>
<p>This model not only evaluated its own utterances as the speaker-based model but also monitored users&#x00027; mouse movements using the same features as in Study 1 as input. It utilized the same feedback behavior as the speaker-based model based on its predictions. As this model was adaptive to the user, we predicted that it would result in less uncertainty in user actions, as indicated by mouse movement behavior.</p></sec>
</sec>
<sec>
<label>4.1.3</label>
<title>Dependent variables</title>
<sec>
<label>4.1.3.1</label>
<title>Behavioral measures</title>
<p>For each model, we measured <bold>user actions</bold> and <bold>mouse features</bold> as outlined in <xref ref-type="table" rid="T1">Table 1</xref>. We also compared <bold>system behavior</bold>, including the <bold>number of incremental units</bold> uttered per model, as well as <bold>model predictions</bold> in relation to the <bold>number of times the Tell-Me-More button was pressed</bold> by users.</p></sec>
<sec>
<label>4.1.3.2</label>
<title>Subjective measures</title>
<p>Users rated each instruction for <bold>appropriateness</bold> using two questions from <xref ref-type="table" rid="T1">Table 1</xref>: <bold>Ambiguousness</bold> and <bold>Information</bold>. <bold>System perception</bold>: At the end of the interaction, users evaluated the agent, focusing on <bold>Instruction Comprehension</bold> and whether the instructions were perceived as <italic>Understood, Complete, Helpful</italic>, and <italic>Collaborative</italic>. We also measured the <bold>Agent Rating</bold> using two items from the Godspeed questionnaire (<xref ref-type="bibr" rid="B8">Bartneck et al., 2008</xref>) related to <italic>Likeability</italic> and <italic>Intelligence</italic>. Finally, we added an <bold>adaptivity</bold> item to assess <bold>how well each model was perceived to adapt to users</bold>.</p></sec>
<sec>
<label>4.1.3.3</label>
<title>Procedure and data collection</title>
<p>The agents followed the same procedure as in Study 1. A total of 71 participants were recruited online. Eight participants were excluded due to technical issues or failure to adhere to study requirements, resulting in 63 participants (21 in each model, using a between-subjects design). The mean age of the participants was 26.9 (&#x000B1;5.9); 39 identified as female, 23 as male, and 1 preferred not to answer. The self-reported English language fluency was 5.7 (&#x000B1;0.9). The task took, on average, 11.4 (&#x000B1;5.0) minutes to complete; 33 participants used a computer mouse, and 30 participants used a mouse trackpad.</p>
</sec>
</sec>
</sec>
<sec>
<label>4.2</label>
<title>Results</title>
<p>As in Study 1, linear mixed-effects models were utilized, incorporating the same fixed and random factors.</p>
<sec>
<label>4.2.1</label>
<title>Behavioral measures</title>
<sec>
<label>4.2.1.1</label>
<title>Effects of user actions</title>
<p><bold>Accuracy</bold>. Participants had an overall accuracy of 64.9%. GLMMs did not show a significant difference in user accuracy among conditions (see <xref ref-type="table" rid="T4">Table 4</xref>). <bold>Idle &#x00026; response time</bold>. The mean idle time was 6.6 (&#x000B1;10.3) seconds, and the mean response time was 12.5 (&#x000B1;10.3) seconds. Log transformations for both idle and response times were statistically different across conditions, with users acting faster when they provided correct answers and also faster when interacting with the listener-based model. Bonferroni corrected pairwise comparisons showed that both the speaker-based and listener-based models led to faster responses by users compared to the adaptive baseline (see <xref ref-type="table" rid="T4">Table 4</xref> and <xref ref-type="fig" rid="F7">Figure 7</xref>).</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Behavioral and subjective measures for each condition (c1: speaker-based model, c2: listener-based model, c3: adaptive baseline).</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><bold>Predictor</bold></th>
<th valign="top" align="center"><bold>C1</bold></th>
<th valign="top" align="center"><bold>C2</bold></th>
<th valign="top" align="center"><bold>C3</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
<th valign="top" align="center"><bold>Chi-square</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Incremental units</td>
<td valign="top" align="center">2.8</td>
<td valign="top" align="center">3.0</td>
<td valign="top" align="center">1.2</td>
<td valign="top" align="center">0.444</td>
<td valign="top" align="center">89.09</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Accuracy (%)</td>
<td valign="top" align="center">0.631</td>
<td valign="top" align="center">0.655</td>
<td valign="top" align="center">0.663</td>
<td valign="top" align="center">0.002</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Idle time</td>
<td valign="top" align="center">5.79</td>
<td valign="top" align="center">5.77</td>
<td valign="top" align="center">10.81</td>
<td valign="top" align="center">0.026</td>
<td valign="top" align="center">17.353</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Response time</td>
<td valign="top" align="center">13.3</td>
<td valign="top" align="center">11.2</td>
<td valign="top" align="center">13.8</td>
<td valign="top" align="center">0.040</td>
<td valign="top" align="center">34.986</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Moved (%)</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.049</td>
<td valign="top" align="center">25.324</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (mean)</td>
<td valign="top" align="center">835</td>
<td valign="top" align="center">876</td>
<td valign="top" align="center">999</td>
<td valign="top" align="center">0.031</td>
<td valign="top" align="center">21.293</td>
<td valign="top" align="center">***</td>
</tr>
<tr>
<td valign="top" align="left">Distance (std)</td>
<td valign="top" align="center">106</td>
<td valign="top" align="center">59</td>
<td valign="top" align="center">158</td>
<td valign="top" align="center">0.054</td>
<td valign="top" align="center">36.494</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (min)</td>
<td valign="top" align="center">686</td>
<td valign="top" align="center">781</td>
<td valign="top" align="center">752</td>
<td valign="top" align="center">0.025</td>
<td valign="top" align="center">23.832</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (max)</td>
<td valign="top" align="center">950</td>
<td valign="top" align="center">937</td>
<td valign="top" align="center">1137</td>
<td valign="top" align="center">0.034</td>
<td valign="top" align="center">15.513</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (range)</td>
<td valign="top" align="center">263</td>
<td valign="top" align="center">155</td>
<td valign="top" align="center">385</td>
<td valign="top" align="center">0.052</td>
<td valign="top" align="center">34.376</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance (slope)</td>
<td valign="top" align="center">&#x02013;13.8</td>
<td valign="top" align="center">&#x02013;5.16</td>
<td valign="top" align="center">&#x02013;25.3</td>
<td valign="top" align="center">0.039</td>
<td valign="top" align="center">30.654</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Distance traveled</td>
<td valign="top" align="center">279</td>
<td valign="top" align="center">228</td>
<td valign="top" align="center">457</td>
<td valign="top" align="center">0.045</td>
<td valign="top" align="center">33.265</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Velocity</td>
<td valign="top" align="center">17.7</td>
<td valign="top" align="center">9.9</td>
<td valign="top" align="center">30.4</td>
<td valign="top" align="center">0.053</td>
<td valign="top" align="center">31.681</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Dir. change (x)</td>
<td valign="top" align="center">0.635</td>
<td valign="top" align="center">&#x02013;0.337</td>
<td valign="top" align="center">0.705</td>
<td valign="top" align="center">0.150</td>
<td valign="top" align="center">88.371</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Dir. change (y)</td>
<td valign="top" align="center">0.681</td>
<td valign="top" align="center">&#x02013;0.320</td>
<td valign="top" align="center">0.875</td>
<td valign="top" align="center">0.157</td>
<td valign="top" align="center">81.378</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">AUC</td>
<td valign="top" align="center">297,209</td>
<td valign="top" align="center">&#x02013;8,227</td>
<td valign="top" align="center">341,591</td>
<td valign="top" align="center">0.341</td>
<td valign="top" align="center">104.34</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Mean abs. deviation</td>
<td valign="top" align="center">91.7</td>
<td valign="top" align="center">39.2</td>
<td valign="top" align="center">85.2</td>
<td valign="top" align="center">0.088</td>
<td valign="top" align="center">33.367</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Max abs. deviation</td>
<td valign="top" align="center">119</td>
<td valign="top" align="center">44</td>
<td valign="top" align="center">132</td>
<td valign="top" align="center">0.191</td>
<td valign="top" align="center">55.846</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Ambiguousness</td>
<td valign="top" align="center">3.69</td>
<td valign="top" align="center">3.77</td>
<td valign="top" align="center">3.90</td>
<td valign="top" align="center">0.010</td>
<td valign="top" align="center">4.49</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Information</td>
<td valign="top" align="center">3.83</td>
<td valign="top" align="center">3.92</td>
<td valign="top" align="center">3.24</td>
<td valign="top" align="center">0.199</td>
<td valign="top" align="center">97.90</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;&#x0002A;</sup></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p><italic>P</italic>-value indicators: &#x02212;<italic>p</italic>&#x0003E;0.05, <sup>&#x0002A;</sup><italic>p</italic> &#x0003D; 0.05, <sup>&#x0002A;&#x0002A;</sup><italic>p</italic> &#x0003D; 0.01, <sup>&#x0002A;&#x0002A;&#x0002A;</sup><italic>p</italic> &#x0003D; 0.001.</p>
</table-wrap-foot>
</table-wrap>
<fig position="float" id="F7">
<label>Figure 7</label>
<caption><p>Response time and instruction appropriateness across different conditions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0007.tif">
<alt-text content-type="machine-generated">Two line graphs compare response time and information across conditions. The left graph shows response time, with incorrect responses in red above correct responses in blue for Speaker, Listener, and Baseline conditions. The right graph shows information, where incorrect responses in red drop significantly in the Baseline, and correct responses in blue decline slightly from Speaker-Based to Baseline.</alt-text>
</graphic>
</fig>
</sec>
<sec>
<label>4.2.1.2</label>
<title>Effects of user uncertainty</title>
<p>LMMs revealed statistically significant differences in how participants utilized mouse movements across conditions (see <xref ref-type="table" rid="T4">Table 4</xref> and <xref ref-type="fig" rid="F8">Figure 8</xref>). The results indicate that mouse movement uncertainty is lower in the listener-based model.</p>
<fig position="float" id="F8">
<label>Figure 8</label>
<caption><p>Mouse trajectory distance traveled and mean absolute deviation to target, separated by accuracy.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0008.tif">
<alt-text content-type="machine-generated">Bar charts compare correct and incorrect data for three methods: Speaker-Based, Listener-Based, and Baseline. The left chart shows distance traveled, with Baseline having the highest distance for correct cases. The right chart illustrates mean absolute deviation, with Speaker-Based having the highest for incorrect cases. Error bars are present.</alt-text>
</graphic>
</fig>
</sec>
<sec>
<label>4.2.1.3</label>
<title>Effects of system behavior</title>
<p>LMMs revealed a statistically significant difference in the number of incremental units per condition, also considering the user&#x00027;s correctness, as shown in <xref ref-type="table" rid="T4">Table 4</xref>. Bonferroni corrected pair-wise comparisons indicated that a significantly lower number of incremental units were uttered with the baseline condition. We calculated the model accuracies aggregated by user (0.623 for the Speaker-Based Model and 0.560 for the Listener-Based Model), which indicated that the Speaker-Based Model had a better prediction match with actual user accuracy. In the adaptive baseline, users requested additional information 25.4% of the time, resulting in significantly fewer spoken installments compared to the two models.</p></sec></sec>
<sec>
<label>4.2.2</label>
<title>Subjective measures</title>
<sec>
<label>4.2.2.1</label>
<title>Instruction appropriateness</title>
<p>LMMs partially revealed statistically significant effects on how the incremental units were perceived, using condition and correctness as fixed factors (see <xref ref-type="table" rid="T4">Table 4</xref> and <xref ref-type="fig" rid="F7">Figure 7</xref>). We observed differences in the perceived amount of information, with the baseline condition having the lowest perceived information overall, according to Bonferroni-corrected pairwise comparisons; user accuracy also influenced whether they felt additional information was necessary.</p></sec>
<sec>
<label>4.2.2.2</label>
<title>System perception</title>
<p>LMMs partially revealed significant effects on how each agent was perceived (see <xref ref-type="table" rid="T5">Table 5</xref> and <xref ref-type="fig" rid="F9">Figure 9</xref>). Bonferroni-corrected pairwise comparisons showed that the instructions provided by the listener-based model were perceived as the most complete, while those from the adaptive baseline were perceived as the least complete.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>System perception measures for each condition.</p></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><bold>Predictor</bold></th>
<th valign="top" align="center"><bold>C1</bold></th>
<th valign="top" align="center"><bold>C2</bold></th>
<th valign="top" align="center"><bold>C3</bold></th>
<th valign="top" align="center"><bold><italic>R</italic><sup>2</sup></bold></th>
<th valign="top" align="center"><bold>Chi-square</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Comply to agent (before)</td>
<td valign="top" align="center">6.76</td>
<td valign="top" align="center">6.48</td>
<td valign="top" align="center">6.52</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">1.092</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Comply to agent (after)</td>
<td valign="top" align="center">5.19</td>
<td valign="top" align="center">5.57</td>
<td valign="top" align="center">5.86</td>
<td valign="top" align="center">0.038</td>
<td valign="top" align="center">2.4145</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Likeability</td>
<td valign="top" align="center">4.92</td>
<td valign="top" align="center">5.34</td>
<td valign="top" align="center">4.91</td>
<td valign="top" align="center">0.023</td>
<td valign="top" align="center">1.4472</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Intelligence</td>
<td valign="top" align="center">4.69</td>
<td valign="top" align="center">5.02</td>
<td valign="top" align="center">4.60</td>
<td valign="top" align="center">0.028</td>
<td valign="top" align="center">1.7614</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Understanding</td>
<td valign="top" align="center">4.48</td>
<td valign="top" align="center">4.81</td>
<td valign="top" align="center">4.90</td>
<td valign="top" align="center">0.017</td>
<td valign="top" align="center">1.0864</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Completeness</td>
<td valign="top" align="center">4.62</td>
<td valign="top" align="center">4.95</td>
<td valign="top" align="center">3.62</td>
<td valign="top" align="center">0.152</td>
<td valign="top" align="center">10.217</td>
<td valign="top" align="center"><sup>&#x0002A;&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">Helpful</td>
<td valign="top" align="center">3.90</td>
<td valign="top" align="center">4.52</td>
<td valign="top" align="center">4.71</td>
<td valign="top" align="center">0.054</td>
<td valign="top" align="center">3.4583</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Collaborative</td>
<td valign="top" align="center">4.29</td>
<td valign="top" align="center">4.95</td>
<td valign="top" align="center">4.76</td>
<td valign="top" align="center">0.034</td>
<td valign="top" align="center">2.1442</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">Adaptive</td>
<td valign="top" align="center">4.48</td>
<td valign="top" align="center">4.38</td>
<td valign="top" align="center">3.62</td>
<td valign="top" align="center">0.078</td>
<td valign="top" align="center">5.0468</td>
<td valign="top" align="center"> &#x02264; 0.1</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p><italic>P</italic>-value indicators: &#x02212;<italic>p</italic>&#x0003E;0.05, <sup>&#x0002A;</sup><italic>p</italic> &#x02264; 0.05, <sup>&#x0002A;&#x0002A;</sup><italic>p</italic> &#x02264; 0.01, <sup>&#x0002A;&#x0002A;&#x0002A;</sup><italic>p</italic> &#x02264; 0.001.</p>
</table-wrap-foot>
</table-wrap>
<fig position="float" id="F9">
<label>Figure 9</label>
<caption><p>Perceived likeability and intelligence across different conditions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-07-1634228-g0009.tif">
<alt-text content-type="machine-generated">Two bar charts compare Speaker-Based, Listener-Based, and Baseline conditions in terms of Likeability and Intelligence. The Listener-Based condition scores highest in both metrics, followed by Speaker-Based and Baseline.</alt-text>
</graphic>
</fig>
</sec>
</sec>
</sec>
<sec>
<label>4.3</label>
<title>Discussion</title>
<p>The findings in this study indicated that adaptation is necessary, with behavioral and subjective preferences leaning toward the <italic>listener-based model</italic>, even though no significant differences were observed in task accuracy compared to the <italic>speaker-based model</italic>. Adaptivity in this context may imply not just improving accuracy with more data but also the ability to dynamically adjust data usage based on user need. Both models were preferred over the <italic>adaptive baseline</italic>. We observed statistically significant differences in how participants rated the amount of information they received; the baseline condition was rated with the lowest amount of information, while the listener-based model was rated with the highest amount, corroborated by the highest number of incremental units overall. Users rated the listener-based model&#x00027;s instructions as the most complete, indicating favorable outcomes in user adaptation; however, no significant differences were found in likeability and intelligence.</p>
<p>Less uncertainty was observed in the listener-based model, as well as when users provided correct answers. The idle and response times also suggested lower cognitive load with the listener-based model, with both models overall performing better than the baseline, which required more effort from the user. Aligning with the listener-based model, users&#x00027; mouse behavior indicated less uncertainty, even though they were not aware of which model was actually considering their mouse movements. The comparison between a speech-only model and a speech-and-mouse model is valuable for analytical purposes. These models are not in direct competition; rather, the comparison helps us understand the relevance of mouse movements in constructing incremental speech. Even though user accuracy was consistent across conditions, the speaker-based model was somewhat better at predicting whether the user would answer correctly.</p></sec>
</sec>
<sec id="s5">
<label>5</label>
<title>General discussion</title>
<sec>
<label>5.1</label>
<title>Key findings</title>
<sec>
<label>5.1.1</label>
<title>RQ1: How do human speakers produce instructions in incremental units?</title>
<p>We observed in the human instructor corpus that instructions are constructed collaboratively and are often incomplete, including errors in production and variations in pauses. The main outcome of the corpus analysis was that speakers consistently adapt their instructions based on their listeners&#x00027; signals of understanding, and the main goal is to train models of user uncertainty based on mouse movements. Through the corpus, we also identified the fundamental attributes of instructions, such as timing and pauses. The analysis further revealed the presentation of speech through continuing contributions, as well as &#x0201C;meta-communicative acts,&#x0201D; such as the user&#x00027;s public display of understanding, which we define in our studies through mouse behavior. These incremental units represent the joint project between the interface and the user to establish common ground. We can conceptualize instructions as the intention of the instructor to refer to a specific part of the assembly, with incremental units serving as the continuing contributions that achieve that goal.</p></sec>
<sec>
<label>5.1.2</label>
<title>RQ2: How much information should an interface convey in each incremental unit?</title>
<p>Study 1 investigated information as the main variable, specifically exploring whether replicating the behavior of humans adjusting their instructions to their listeners&#x00027; information needs has an impact when implemented in a machine. The findings indicated that information plays a significant role in users&#x00027; accuracy in the task, as well as in their displayed uncertainty. Classification accuracy was used as a proxy for the quality of each model; however, it did not fully capture the model&#x00027;s effectiveness or how it was perceived by users, which was addressed in RQ3.</p></sec>
<sec>
<label>5.1.3</label>
<title>RQ3: How should the interface adapt its instructions if the user&#x00027;s attention does not meet the expected behavior?</title>
<p>RQ3 was examined through a user study that compared three adaptive models of constructing instructions. We predicted that differences in users&#x00027; accuracy in the task would be observed if the appropriate model adapted to their information needs; however, task accuracy did not appear to differ by model. Nevertheless, we did observe differences in user behavior, with the listener-based model prompting less uncertainty. Instructions from this model were perceived as more complete, even though the incremental units were identical across all conditions. Comparing the three models showed that &#x02018;observing user signals recurrently gives interfaces the advantage of planning utterances collaboratively, with the user being part of the process&#x00027; (<xref ref-type="bibr" rid="B71">Kontogiorgos, 2022</xref>).</p>
</sec>
</sec>
<sec>
<label>5.2</label>
<title>Implications for adaptive user interfaces</title>
<p>We use voice rather than text as the medium of communication because instructions in incremental units have primarily been observed as a conversational phenomenon. Voice also allows the user to focus on visually scanning for the referring objects while receiving information incrementally. Therefore, these findings have implications primarily for conversational user interfaces in task collaboration settings, as well as teaching and instructional interfaces utilizing mouse movements. While not tested in this study, such instructional behaviors are important for embodied interfaces that observe the users&#x00027; embodied signals when uttering instructions. Instructions in incremental units provide an opportunity to convey social behavior, which may be expected by human users, even when the interlocutor is a computer. Similar to the use of discourse markers, <italic>displaying information incrementally may help to mitigate directness</italic>, balancing between brevity and information exchange as a politeness strategy (<xref ref-type="bibr" rid="B50">Goodman and Stuhlm&#x000FC;ller, 2013</xref>; <xref ref-type="bibr" rid="B127">Yoon et al., 2016</xref>).</p>
<p>However, presenting information incrementally may not always be preferred by users, depending on the interface&#x00027;s utility and the changes in the user&#x00027;s state (e.g., during emergencies). Additionally, different users may interpret incrementality in various ways, meaning that the interface must also consider users&#x00027; personality traits and what they perceive as efficient vs. polite communication. Recognizing that a turn unit is more flexible than &#x0201C;<italic>push-to-talk&#x0201D;</italic> interactions (<xref ref-type="bibr" rid="B43">Fern&#x000E1;ndez et al., 2007</xref>) enables the possibility to co-construct instructions with the user (<xref ref-type="bibr" rid="B72">Kontogiorgos and Gustafson, 2021</xref>).</p>
</sec>
<sec>
<label>5.3</label>
<title>Limitations and future work</title>
<p>In this paper, we used a set of puzzle pieces to study incremental utterance production. Similar to the Tangram puzzles used in psycholinguistics, the Tetris-like Pentomino shapes lack the appearance of common objects, making them a suitable paradigm for examining linguistic alignment when people collaboratively develop new terms to describe objects. Each step in the puzzle is grounded incrementally, making it ideal for investigating computer-generated incremental instructions. This constrained nature of the task offers an advantage in examining instructions and provides a level of control over how conversational phenomena evolve. While our findings provide novel insights into utterance construction, they should be interpreted with caution. The collaborative nature of the task may limit generalization to other forms of conversation, such as &#x0201C;<italic>open-world dialogues&#x0201D;</italic> (<xref ref-type="bibr" rid="B14">Bohus and Horvitz, 2009</xref>), which are not object-focused and may not involve collaboration. Nonetheless, the parallel to real-world tasks can be drawn to any machine-guided assembly, whether it involves building IKEA furniture or receiving instructions through a visual interface.</p>
<p>In this article, we used a speaker-and-listener modeling approach to facilitate mutual understanding. However, a much simpler model could use delays in task progress as a proxy for a lack of grounding; when a user does not respond to an instruction, the system can assume that the instruction was either not heard or not understood. Since common ground is a &#x0201C;<italic>feeling&#x0201D;</italic> among speakers, it can be challenging to methodologically establish a ground truth for what is understood by users (<xref ref-type="bibr" rid="B33">DeVault, 2008</xref>). While we can confirm that each incremental unit is heard, we cannot ensure that it is also understood (as shown in the lack of significant findings on accuracy in Study 2).</p>
<p>An important limitation of this work has been the use of prosody. We employed standard TTS services that are not designed for co-constructed speech; instructions in incremental units are a conversational phenomenon where appropriate intonation carries pragmatic information, such as signaling that information may be incomplete or inviting the listener to participate in its construction. Current TTS services lack this flexibility, which may have influenced how users perceived the agents&#x00027; adaptation to their behavior.</p>
<p>Additionally, the utterances were originally spoken by human instructors and constructed collaboratively; it is inherently subjective what information is ambiguous, as all human instructions are, to some degree, ambiguous and incomplete. User uncertainty was treated as the user&#x00027;s attempt to express clarification requests, which often leads to utterance reformulation rather than the provision of new information (<xref ref-type="bibr" rid="B101">Schlangen and Fern&#x000E1;ndez, 2007</xref>). In this work, however, we chose to always present new information. Future research should investigate how to repair utterances when user uncertainty is detected and explore sequential learning of linguistic strategies based on the state of the user and the environment (<xref ref-type="bibr" rid="B40">Ekstedt and Skantze, 2020</xref>; <xref ref-type="bibr" rid="B97">Sadler et al., 2023</xref>; <xref ref-type="bibr" rid="B98">Sadler and Schlangen, 2023</xref>), as well as alternative architectures to speaker-based and listener-based models. Future work should also consider the impact of such proactive interfaces that may have implications for the user&#x00027;s task workflow interruption, as well as approach incremental unit construction through the principles of mixed-initiative user interfaces (<xref ref-type="bibr" rid="B57">Horvitz, 1999</xref>).</p>
<p>Finally, participants were mainly young adults, fluent English speakers, which limits the generalisability of our findings. Assistive technologies intended for older adults may face different interaction needs, communication styles, and attentional patterns. Future work should therefore validate the proposed approach with older adult populations to assess its applicability to age-related assistive settings.</p>
</sec>
<sec>
<label>5.4</label>
<title>Real-world applications</title>
<p>The approach of using mouse movements as implicit signals of user understanding offers significant potential for real-world applications in domains requiring multitasking interactions. In assistive technologies, adaptive voice interfaces could monitor subtle interaction cues (e.g., cursor hesitation) to adjust the timing, complexity, or repetition of instructions. Outside the mouse-movement domain, in smart home environments, where users may interact with devices while engaged in physical tasks, such interfaces could reduce cognitive effort by incrementally delivering guidance and monitoring non-verbal cues such as eye gaze, hand gestures, or interaction delays.</p>
<p>Beyond individual tasks, this research can be applied in collaborative or instructional settings, such as remote education, collaborative design platforms, or training simulators. In these settings, systems that detect user uncertainty through behavioral signals can better support novice users by tailoring information delivery to their comprehension level. For example, in remote technical support, systems can detect whether the user is struggling with a step and proactively offer clarification without requiring explicit feedback. In human&#x02014;robot collaboration, detecting hesitation or misalignment in operator behavior can help robots adjust their verbal instructions or actions in real time.</p>
<p>As the puzzle task used in this study offers experimental control but may limit ecological validity, future work should investigate how these findings transfer to more naturalistic environments and more diverse input modalities. Particularly, integrating implicit cues beyond mouse trajectories, such as gaze behavior, body orientation, or hesitation in speech, may improve robustness and generalisability in real-world applications.</p></sec></sec>
<sec sec-type="conclusions" id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In summary, this article presented: (i) an analysis of the attributes of human incremental instruction, (ii) empirical evidence demonstrating the benefits of adapting the delivery of information to user behavior, and (iii) a user study showing that mouse movements are a reliable implicit indicator of uncertainty. To the best of our knowledge, this work is the first to utilize real-time turn-taking decisions based solely on users&#x00027; mouse movements. We also showed that users&#x00027; movement patterns reveal potential ambiguities in instructions, which a voice user interface can leverage to adjust its guidance.</p>
<p>Taken together, our findings demonstrate that <italic>process adaptivity</italic>, rather than outcome differences, improves the interaction dynamics of incremental guidance systems. While overall task accuracy did not differ significantly between models, systems that adapted their timing and information granularity in response to users&#x00027; behavior led to smoother interaction, reduced idle time, and more favorable user perceptions. This highlights the importance of monitoring ongoing behavioral cues to maintain mutual understanding during action execution.</p>
<p>The analysis of human incremental instruction provides insights into how humans structure assistance and how such strategies can be operationalised in intelligent interfaces. These results have practical implications for the automatic generation of human-like, responsive instructions in assistive and collaborative systems. More broadly, this article contributes to a central challenge in HCI: how to assess and maintain common ground incrementally as interaction unfolds.</p></sec>
</body>
<back>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec sec-type="ethics-statement" id="s8">
<title>Ethics statement</title>
<p>The studies involving humans were approved by the German Research Foundation (Deutsche Forschungsgemeinschaft). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.</p>
</sec>
<sec sec-type="author-contributions" id="s9">
<title>Author contributions</title>
<p>DK: Conceptualization, Writing &#x02013; review &#x00026; editing, Methodology, Investigation, Software, Visualization, Formal analysis, Writing &#x02013; original draft, Validation, Data curation. DS: Funding acquisition, Writing &#x02013; review &#x00026; editing, Supervision, Project administration, Resources, Conceptualization.</p>
</sec>
<ack><title>Acknowledgments</title><p>The authors would like to thank Jana G&#x000F6;tze, Karla Friedrichs, Hannah Pelikan, and Joakim Gustafson for contributing to the discussions on earlier versions of this study, as well as the reviewers for constructive feedback on earlier versions of the paper.</p></ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s11">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p>
<p>Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.</p></sec>
<sec sec-type="disclaimer" id="s12">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s13">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fcomp.2025.1634228/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fcomp.2025.1634228/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Video_1.mp4" id="SM1" mimetype="video/mp4" xmlns:xlink="http://www.w3.org/1999/xlink"/></sec>
<ref-list>
<title>References</title>
<ref id="B1">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Arapakis</surname> <given-names>I.</given-names></name> <name><surname>Leiva</surname> <given-names>L. A.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Predicting user engagement with direct displays using mouse cursor information,&#x0201D;</article-title> in <source>Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval</source>, <fpage>599</fpage>&#x02013;<lpage>608</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2911451.2911505</pub-id></mixed-citation>
</ref>
<ref id="B2">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Arapakis</surname> <given-names>I.</given-names></name> <name><surname>Leiva</surname> <given-names>L. A.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Learning efficient representations of mouse movements to predict user attention,&#x0201D;</article-title> in <source>Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>, <fpage>1309</fpage>&#x02013;<lpage>1318</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3397271.3401031</pub-id></mixed-citation>
</ref>
<ref id="B3">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Arapakis</surname> <given-names>I.</given-names></name> <name><surname>Penta</surname> <given-names>A.</given-names></name> <name><surname>Joho</surname> <given-names>H.</given-names></name> <name><surname>Leiva</surname> <given-names>L. A.</given-names></name></person-group> (<year>2020</year>). <article-title>A price-per-attention auction scheme using mouse cursor information</article-title>. <source>ACM Trans. Inf. Syst</source>. <volume>38</volume>, <fpage>1</fpage>&#x02013;<lpage>30</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3374210</pub-id></mixed-citation>
</ref>
<ref id="B4">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Arroyo</surname> <given-names>E.</given-names></name> <name><surname>Selker</surname> <given-names>T.</given-names></name> <name><surname>Wei</surname> <given-names>W.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x0201C;Usability tool for analysis of web designs using mouse tracks,&#x0201D;</article-title> in <source>CHI&#x00027;06 Extended Abstracts on Human Factors in Computing Systems</source>, <fpage>484</fpage>&#x02013;<lpage>489</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1125451.1125557</pub-id></mixed-citation>
</ref>
<ref id="B5">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ashdown</surname> <given-names>M.</given-names></name> <name><surname>Oka</surname> <given-names>K.</given-names></name> <name><surname>Sato</surname> <given-names>Y.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;Combining head tracking and mouse input for a gui on multiple monitors,&#x0201D;</article-title> in <source>CHI&#x00027;05 Extended Abstracts on Human Factors in Computing Systems</source>, <fpage>1188</fpage>&#x02013;<lpage>1191</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1056808.1056873</pub-id></mixed-citation>
</ref>
<ref id="B6">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ashktorab</surname> <given-names>Z.</given-names></name> <name><surname>Jain</surname> <given-names>M.</given-names></name> <name><surname>Liao</surname> <given-names>Q. V.</given-names></name> <name><surname>Weisz</surname> <given-names>J. D.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Resilient chatbots: repair strategy preferences for conversational breakdowns,&#x0201D;</article-title> in <source>Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems</source>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3290605.3300484</pub-id></mixed-citation>
</ref>
<ref id="B7">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Axelsson</surname> <given-names>N.</given-names></name> <name><surname>Skantze</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Using knowledge graphs and behaviour trees for feedback-aware presentation agents,&#x0201D;</article-title> in <source>Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents</source>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3383652.3423884</pub-id></mixed-citation>
</ref>
<ref id="B8">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bartneck</surname> <given-names>C.</given-names></name> <name><surname>Croft</surname> <given-names>E.</given-names></name> <name><surname>Kulic</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Measuring the anthropomorphism, animacy, likeability, perceived intelligence and perceived safety of robots,&#x0201D;</article-title> in <source>Metrics for HRI Workshop, Technical Report</source>, <fpage>37</fpage>&#x02013;<lpage>44</lpage>.</mixed-citation>
</ref>
<ref id="B9">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bates</surname> <given-names>D.</given-names></name> <name><surname>M&#x000E4;chler</surname> <given-names>M.</given-names></name> <name><surname>Bolker</surname> <given-names>B.</given-names></name> <name><surname>Walker</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Fitting linear mixed-effects models using lme4</article-title>. <source>arXiv preprint arXiv:1406.5823</source>.</mixed-citation>
</ref>
<ref id="B10">
<mixed-citation publication-type="web"><person-group person-group-type="author"><name><surname>Baumann</surname> <given-names>T.</given-names></name> <name><surname>Paetzel</surname> <given-names>M.</given-names></name> <name><surname>Schlesinger</surname> <given-names>P.</given-names></name> <name><surname>Menzel</surname> <given-names>W.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Using Affordances to shape the interaction in a hybrid spoken dialogue system,&#x0201D;</article-title> in <source>Studientexte zur Sprachkommunikation: Elektronische Sprachsignalverarbeitung 2013</source> (<publisher-loc>Dresden</publisher-loc>: <publisher-name>TUD Press</publisher-name>), <fpage>1219</fpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.essv.de/pdf/2013_12_19.pdf">https://www.essv.de/pdf/2013_12_19.pdf</ext-link></mixed-citation>
</ref>
<ref id="B11">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Baumann</surname> <given-names>T.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Inpro_iss: a component for just-in-time incremental speech synthesis,&#x0201D;</article-title> in <source>Proceedings of the ACL 2012 System Demonstrations</source>, <fpage>103</fpage>&#x02013;<lpage>108</lpage>.</mixed-citation>
</ref>
<ref id="B12">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Behnke</surname> <given-names>G.</given-names></name> <name><surname>Bercher</surname> <given-names>P.</given-names></name> <name><surname>Kraus</surname> <given-names>M.</given-names></name> <name><surname>Schiller</surname> <given-names>M.</given-names></name> <name><surname>Mickeleit</surname> <given-names>K.</given-names></name> <name><surname>H&#x000E4;ge</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;New developments for robert-assisting novice users even better in diy projects,&#x0201D;</article-title> in <source>Proceedings of the International Conference on Automated Planning and Scheduling</source>, <fpage>343</fpage>&#x02013;<lpage>347</lpage>. doi: <pub-id pub-id-type="doi">10.1609/icaps.v30i1.6679</pub-id></mixed-citation>
</ref>
<ref id="B13">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bell</surname> <given-names>A.</given-names></name></person-group> (<year>1984</year>). <article-title>Language style as audience design</article-title>. <source>Lang. Soc</source>. <volume>13</volume>, <fpage>145</fpage>&#x02013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1017/S004740450001037X</pub-id></mixed-citation>
</ref>
<ref id="B14">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bohus</surname> <given-names>D.</given-names></name> <name><surname>Horvitz</surname> <given-names>E.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Open-world dialog: challenges, directions, and prototype,&#x0201D;</article-title> in <source>6th IJCAI Workshop on Knowledge and Reasoning in Practical Dialogue Systems</source>, <fpage>34</fpage>.</mixed-citation>
</ref>
<ref id="B15">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Branigan</surname> <given-names>H. P.</given-names></name> <name><surname>Pickering</surname> <given-names>M. J.</given-names></name> <name><surname>Pearson</surname> <given-names>J.</given-names></name> <name><surname>McLean</surname> <given-names>J. F.</given-names></name></person-group> (<year>2010</year>). <article-title>Linguistic alignment between people and computers</article-title>. <source>J. Pragmat</source>. <volume>42</volume>, <fpage>2355</fpage>&#x02013;<lpage>2368</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.pragma.2009.12.012</pub-id></mixed-citation>
</ref>
<ref id="B16">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Brennan</surname> <given-names>S. E.</given-names></name> <name><surname>Clark</surname> <given-names>H. H.</given-names></name></person-group> (<year>1996</year>). <article-title>Conceptual pacts and lexical choice in conversation</article-title>. <source>J. Exper. Psychol</source>. <volume>22</volume>:<fpage>1482</fpage>. doi: <pub-id pub-id-type="doi">10.1037//0278-7393.22.6.1482</pub-id><pub-id pub-id-type="pmid">8921603</pub-id></mixed-citation>
</ref>
<ref id="B17">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Brennan</surname> <given-names>S. E.</given-names></name> <name><surname>Hanna</surname> <given-names>J. E.</given-names></name></person-group> (<year>2009</year>). <article-title>Partner-specific adaptation in dialog</article-title>. <source>Top. Cogn. Sci</source>. <volume>1</volume>, <fpage>274</fpage>&#x02013;<lpage>291</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1756-8765.2009.01019.x</pub-id><pub-id pub-id-type="pmid">25164933</pub-id></mixed-citation>
</ref>
<ref id="B18">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Brennan</surname> <given-names>S. E.</given-names></name> <name><surname>Schober</surname> <given-names>M. F.</given-names></name></person-group> (<year>2001</year>). <article-title>How listeners compensate for disfluencies in spontaneous speech</article-title>. <source>J. Mem. Lang</source>. <volume>44</volume>, <fpage>274</fpage>&#x02013;<lpage>296</lpage>. doi: <pub-id pub-id-type="doi">10.1006/jmla.2000.2753</pub-id></mixed-citation>
</ref>
<ref id="B19">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Br&#x000FC;ckner</surname> <given-names>L.</given-names></name> <name><surname>Arapakis</surname> <given-names>I.</given-names></name> <name><surname>Leiva</surname> <given-names>L. A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;When choice happens: a systematic examination of mouse movement length for decision making in web search,&#x0201D;</article-title> in <source>Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>, <fpage>2318</fpage>&#x02013;<lpage>2322</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3404835.3463055</pub-id></mixed-citation>
</ref>
<ref id="B20">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Buschmeier</surname> <given-names>H.</given-names></name> <name><surname>Baumann</surname> <given-names>T.</given-names></name> <name><surname>Dosch</surname> <given-names>B.</given-names></name> <name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Combining incremental language generation and incremental speech synthesis for adaptive information presentation,&#x0201D;</article-title> in <source>Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue</source>, <fpage>295</fpage>&#x02013;<lpage>303</lpage>.</mixed-citation>
</ref>
<ref id="B21">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Calcagn&#x000ED;</surname> <given-names>A.</given-names></name> <name><surname>Lombardi</surname> <given-names>L.</given-names></name> <name><surname>Sulpizio</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Analyzing spatial data from mouse tracker methodology: an entropic approach</article-title>. <source>Behav. Res. Methods</source> <volume>49</volume>, <fpage>2012</fpage>&#x02013;<lpage>2030</lpage>. doi: <pub-id pub-id-type="doi">10.3758/s13428-016-0839-5</pub-id><pub-id pub-id-type="pmid">28078571</pub-id></mixed-citation>
</ref>
<ref id="B22">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Chai</surname> <given-names>J. Y.</given-names></name> <name><surname>She</surname> <given-names>L.</given-names></name> <name><surname>Fang</surname> <given-names>R.</given-names></name> <name><surname>Ottarson</surname> <given-names>S.</given-names></name> <name><surname>Littley</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>&#x0201C;Collaborative effort towards common ground in situated human-robot dialogue,&#x0201D;</article-title> in <source>2014 9th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>33</fpage>&#x02013;<lpage>40</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2559636.2559677</pub-id></mixed-citation>
</ref>
<ref id="B23">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>M. C.</given-names></name> <name><surname>Anderson</surname> <given-names>J. R.</given-names></name> <name><surname>Sohn</surname> <given-names>M. H.</given-names></name></person-group> (<year>2001</year>). <article-title>&#x0201C;What can a mouse cursor tell us more? Correlation of eye/mouse movements on web browsing,&#x0201D;</article-title> in <source>CHI&#x00027;01 Extended Abstracts on Human Factors in Computing Systems</source>, <fpage>281</fpage>&#x02013;<lpage>282</lpage>. doi: <pub-id pub-id-type="doi">10.1145/634067.634234</pub-id></mixed-citation>
</ref>
<ref id="B24">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>H. H.</given-names></name></person-group> (<year>1996</year>). <source>Using Language</source>. Cambridge: Cambridge University Press. doi: <pub-id pub-id-type="doi">10.1017/CBO9780511620539</pub-id></mixed-citation>
</ref>
<ref id="B25">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>H. H.</given-names></name> <name><surname>Brennan</surname> <given-names>S. E.</given-names></name></person-group> (<year>1991</year>). <article-title>&#x0201C;Grounding in communication,&#x0201D;</article-title> in <source>Perspectives on socially shared cognition</source>, eds. L. B. Resnick, J. M. Levine, and S. D. Teasley (New York: American Psychological Association), <fpage>127</fpage>&#x02013;<lpage>149</lpage>. doi: <pub-id pub-id-type="doi">10.1037/10096-006</pub-id></mixed-citation>
</ref>
<ref id="B26">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>H. H.</given-names></name> <name><surname>Krych</surname> <given-names>M. A.</given-names></name></person-group> (<year>2004</year>). <article-title>Speaking while monitoring addressees for understanding</article-title>. <source>J. Mem. Lang</source>. <volume>50</volume>, <fpage>62</fpage>&#x02013;<lpage>81</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jml.2003.08.004</pub-id></mixed-citation>
</ref>
<ref id="B27">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>H. H.</given-names></name> <name><surname>Marshall</surname> <given-names>C. R.</given-names></name></person-group> (<year>1981</year>). <article-title>&#x0201C;Definite knowledge and mutual knowledge,&#x0201D;</article-title> in <source>Elements of Discourse Understanding</source>, eds. A. K. Joshi, B. L. Webber, and I. A. Sag (Cambridge University Press), <fpage>10</fpage>&#x02013;<lpage>63</lpage>.</mixed-citation>
</ref>
<ref id="B28">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>H. H.</given-names></name> <name><surname>Wilkes-Gibbs</surname> <given-names>D.</given-names></name></person-group> (<year>1986</year>). <article-title>Referring as a collaborative process</article-title>. <source>Cognition</source> <volume>22</volume>, <fpage>1</fpage>&#x02013;<lpage>39</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0010-0277(86)90010-7</pub-id><pub-id pub-id-type="pmid">3709088</pub-id></mixed-citation>
</ref>
<ref id="B29">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>L.</given-names></name> <name><surname>Pantidi</surname> <given-names>N.</given-names></name> <name><surname>Cooney</surname> <given-names>O.</given-names></name> <name><surname>Doyle</surname> <given-names>P.</given-names></name> <name><surname>Garaialde</surname> <given-names>D.</given-names></name> <name><surname>Edwards</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;What makes a good conversation? Challenges in designing truly conversational agents,&#x0201D;</article-title> in <source>Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems</source>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3290605.3300705</pub-id></mixed-citation>
</ref>
<ref id="B30">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Dafoe</surname> <given-names>A.</given-names></name> <name><surname>Bachrach</surname> <given-names>Y.</given-names></name> <name><surname>Hadfield</surname> <given-names>G.</given-names></name> <name><surname>Horvitz</surname> <given-names>E.</given-names></name> <name><surname>Larson</surname> <given-names>K.</given-names></name> <name><surname>Graepel</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>Cooperative ai: machines must learn to find common ground</article-title>. <source>Nature</source> <volume>593</volume>, <fpage>33</fpage>&#x02013;<lpage>36</lpage>. doi: <pub-id pub-id-type="doi">10.1038/d41586-021-01170-0</pub-id><pub-id pub-id-type="pmid">33947992</pub-id></mixed-citation>
</ref>
<ref id="B31">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Dale</surname> <given-names>R.</given-names></name> <name><surname>Reiter</surname> <given-names>E.</given-names></name></person-group> (<year>1995</year>). <article-title>Computational interpretations of the gricean maxims in the generation of referring expressions</article-title>. <source>Cogn. Sci</source>. <volume>19</volume>, <fpage>233</fpage>&#x02013;<lpage>263</lpage>. doi: <pub-id pub-id-type="doi">10.1207/s15516709cog1902_3</pub-id></mixed-citation>
</ref>
<ref id="B32">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Dethlefs</surname> <given-names>N.</given-names></name> <name><surname>Cuay&#x000E1;huitl</surname> <given-names>H.</given-names></name> <name><surname>Viethen</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Optimising natural language generation decision making for situated dialogue,&#x0201D;</article-title> in <source>Proceedings of the SIGDIAL 2011 Conference</source>, <fpage>78</fpage>&#x02013;<lpage>87</lpage>.</mixed-citation>
</ref>
<ref id="B33">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>DeVault</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <source>Contribution Tracking: Participating in Task-Oriented Dialogue Under Uncertainty</source>. <publisher-loc>New Jersey</publisher-loc>: <publisher-name>Rutgers The State University of New Jersey-New Brunswick</publisher-name>.</mixed-citation>
</ref>
<ref id="B34">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>DeVault</surname> <given-names>D.</given-names></name> <name><surname>Rutgers</surname> <given-names>N. K.</given-names></name> <name><surname>Kothari</surname> <given-names>A.</given-names></name> <name><surname>Oved</surname> <given-names>I.</given-names></name> <name><surname>Stone</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;An information-state approach to collaborative reference,&#x0201D;</article-title> in <source>Proceedings of the ACL Interactive Poster and Demonstration Sessions</source>, <fpage>1</fpage>&#x02013;<lpage>4</lpage>. doi: <pub-id pub-id-type="doi">10.3115/1225753.1225754</pub-id></mixed-citation>
</ref>
<ref id="B35">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Diaz</surname> <given-names>F.</given-names></name> <name><surname>White</surname> <given-names>R.</given-names></name> <name><surname>Buscher</surname> <given-names>G.</given-names></name> <name><surname>Liebling</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Robust models of mouse movement on dynamic web search results pages,&#x0201D;</article-title> in <source>Proceedings of the 22nd ACM international conference on Information &#x00026;Knowledge Management</source>, <fpage>1451</fpage>&#x02013;<lpage>1460</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2505515.2505717</pub-id></mixed-citation>
</ref>
<ref id="B36">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Do&#x0011F;an</surname> <given-names>F. I.</given-names></name> <name><surname>Gillet</surname> <given-names>S.</given-names></name> <name><surname>Carter</surname> <given-names>E. J.</given-names></name> <name><surname>Leite</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>The impact of adding perspective-taking to spatial referencing during human-robot interaction</article-title>. <source>Rob. Auton. Syst</source>. <volume>134</volume>:<fpage>103654</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.robot.2020.103654</pub-id></mixed-citation>
</ref>
<ref id="B37">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Do&#x0011F;an</surname> <given-names>F. I.</given-names></name> <name><surname>Leite</surname> <given-names>I.</given-names></name></person-group> (<year>2021</year>). <article-title>Open challenges on generating referring expressions for human-robot interaction</article-title>. <source>arXiv preprint arXiv:2104.09193</source>.</mixed-citation>
</ref>
<ref id="B38">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Eckstein</surname> <given-names>M. P.</given-names></name></person-group> (<year>2011</year>). <article-title>Visual search: a retrospective</article-title>. <source>J. Vis</source>. <volume>11</volume>, <fpage>14</fpage>&#x02013;<lpage>14</lpage>. doi: <pub-id pub-id-type="doi">10.1167/11.5.14</pub-id></mixed-citation>
</ref>
<ref id="B39">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Eerola</surname> <given-names>T.</given-names></name> <name><surname>Armitage</surname> <given-names>J.</given-names></name> <name><surname>Lavan</surname> <given-names>N.</given-names></name> <name><surname>Knight</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Online data collection in auditory perception and cognition research: recruitment, testing, data quality and ethical considerations</article-title>. <source>Audit. Percept. Cogn</source>. <volume>4</volume>, <fpage>251</fpage>&#x02013;<lpage>280</lpage>. doi: <pub-id pub-id-type="doi">10.1080/25742442.2021.2007718</pub-id></mixed-citation>
</ref>
<ref id="B40">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ekstedt</surname> <given-names>E.</given-names></name> <name><surname>Skantze</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog,&#x0201D;</article-title> in <source>Findings of the Association for Computational Linguistics: EMNLP 2020</source>, <fpage>2981</fpage>&#x02013;<lpage>2990</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/2020.findings-emnlp.268</pub-id></mixed-citation>
</ref>
<ref id="B41">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Engonopoulos</surname> <given-names>N.</given-names></name> <name><surname>Villalba</surname> <given-names>M.</given-names></name> <name><surname>Titov</surname> <given-names>I.</given-names></name> <name><surname>Koller</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Predicting the resolution of referring expressions from user behavior,&#x0201D;</article-title> in <source>Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing</source>, <fpage>1354</fpage>&#x02013;<lpage>1359</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/D13-1134</pub-id></mixed-citation>
</ref>
<ref id="B42">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>R.</given-names></name> <name><surname>Doering</surname> <given-names>M.</given-names></name> <name><surname>Chai</surname> <given-names>J. Y.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Embodied collaborative referring expression generation in situated human-robot interaction,&#x0201D;</article-title> in <source>Proceedings of the Tenth Annual ACM/IEEE International Conference on Human-Robot Interaction</source>, <fpage>271</fpage>&#x02013;<lpage>278</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2696454.2696467</pub-id></mixed-citation>
</ref>
<ref id="B43">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fern&#x000E1;ndez</surname> <given-names>R.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name> <name><surname>Lucht</surname> <given-names>T.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Push-to-talk ain&#x00027;t always bad! Comparing different interactivity settings in task-oriented dialogue,&#x0201D;</article-title> in <source>Proceedings of the DECALOG 2007</source>, <fpage>25</fpage>.</mixed-citation>
</ref>
<ref id="B44">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fitts</surname> <given-names>P. M.</given-names></name></person-group> (<year>1954</year>). <article-title>The information capacity of the human motor system in controlling the amplitude of movement</article-title>. <source>J. Exp. Psychol</source>. <volume>47</volume>:<fpage>381</fpage>. doi: <pub-id pub-id-type="doi">10.1037/h0055392</pub-id><pub-id pub-id-type="pmid">13174710</pub-id></mixed-citation>
</ref>
<ref id="B45">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fussell</surname> <given-names>S. R.</given-names></name> <name><surname>Krauss</surname> <given-names>R. M.</given-names></name></person-group> (<year>1992</year>). <article-title>Coordination of knowledge in communication: effects of speakers&#x00027; assumptions about what others know</article-title>. <source>J. Pers. Soc. Psychol</source>. <volume>62</volume>:<fpage>378</fpage>. doi: <pub-id pub-id-type="doi">10.1037//0022-3514.62.3.378</pub-id><pub-id pub-id-type="pmid">1560334</pub-id></mixed-citation>
</ref>
<ref id="B46">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Garoufi</surname> <given-names>K.</given-names></name> <name><surname>Koller</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Generation of effective referring expressions in situated context</article-title>. <source>Lang. Cogn. Neurosci</source>. <volume>29</volume>, <fpage>986</fpage>&#x02013;<lpage>1001</lpage>. doi: <pub-id pub-id-type="doi">10.1080/01690965.2013.847190</pub-id></mixed-citation>
</ref>
<ref id="B47">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Garoufi</surname> <given-names>K.</given-names></name> <name><surname>Staudte</surname> <given-names>M.</given-names></name> <name><surname>Koller</surname> <given-names>A.</given-names></name> <name><surname>Crocker</surname> <given-names>M. W.</given-names></name></person-group> (<year>2016</year>). <article-title>Exploiting listener gaze to improve situated communication in dynamic virtual environments</article-title>. <source>Cogn. Sci</source>. <volume>40</volume>, <fpage>1671</fpage>&#x02013;<lpage>1703</lpage>. doi: <pub-id pub-id-type="doi">10.1111/cogs.12298</pub-id><pub-id pub-id-type="pmid">26471391</pub-id></mixed-citation>
</ref>
<ref id="B48">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Gigliobianco</surname> <given-names>S.</given-names></name> <name><surname>Kontogiorgos</surname> <given-names>D.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2024</year>). <article-title>&#x0201C;Learning task-oriented dialogues through various degrees of interactivity,&#x0201D;</article-title> in <source>Proceedings of the 28th Workshop on the Semantics and Pragmatics of Dialogue</source>.</mixed-citation>
</ref>
<ref id="B49">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Gonsior</surname> <given-names>B.</given-names></name> <name><surname>Wollherr</surname> <given-names>D.</given-names></name> <name><surname>Buss</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Towards a dialog strategy for handling miscommunication in human-robot dialog,&#x0201D;</article-title> in <source>19th International Symposium in Robot and Human Interactive Communication</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>264</fpage>&#x02013;<lpage>269</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ROMAN.2010.5598618</pub-id></mixed-citation>
</ref>
<ref id="B50">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Goodman</surname> <given-names>N. D.</given-names></name> <name><surname>Stuhlm&#x000FC;ller</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Knowledge and implicature: Modeling language understanding as social cognition</article-title>. <source>Top. Cogn. Sci</source>. <volume>5</volume>, <fpage>173</fpage>&#x02013;<lpage>184</lpage>. doi: <pub-id pub-id-type="doi">10.1111/tops.12007</pub-id><pub-id pub-id-type="pmid">23335578</pub-id></mixed-citation>
</ref>
<ref id="B51">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Grice</surname> <given-names>H. P.</given-names></name></person-group> (<year>1975</year>). <article-title>&#x0201C;Logic and conversation,&#x0201D;</article-title> in <source>Speech Acts</source> (<publisher-loc>Brill</publisher-loc>), <fpage>41</fpage>&#x02013;<lpage>58</lpage>. doi: <pub-id pub-id-type="doi">10.1163/9789004368811_003</pub-id></mixed-citation>
</ref>
<ref id="B52">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Grice</surname> <given-names>P.</given-names></name></person-group> (<year>1989</year>). <source>Studies in the Way of Words</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Harvard University Press</publisher-name>.</mixed-citation>
</ref>
<ref id="B53">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Q.</given-names></name> <name><surname>Agichtein</surname> <given-names>E.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Exploring mouse movements for inferring query intent,&#x0201D;</article-title> in <source>Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>, <fpage>707</fpage>&#x02013;<lpage>708</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1390334.1390462</pub-id></mixed-citation>
</ref>
<ref id="B54">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Q.</given-names></name> <name><surname>Agichtein</surname> <given-names>E.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Towards predicting web searcher gaze position from mouse movements,&#x0201D;</article-title> in <source>CHI&#x00027;10 Extended Abstracts on Human Factors in Computing Systems</source>, <fpage>3601</fpage>&#x02013;<lpage>3606</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1753846.1754025</pub-id></mixed-citation>
</ref>
<ref id="B55">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Haake</surname> <given-names>K.</given-names></name> <name><surname>Schimke</surname> <given-names>S.</given-names></name> <name><surname>Betz</surname> <given-names>S.</given-names></name> <name><surname>Zarrie&#x000DF;</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Do hesitations facilitate processing of partially defective system utterances? An exploratory eye tracking study,&#x0201D;</article-title> in <source>Proceedings of Interspeech</source>. doi: <pub-id pub-id-type="doi">10.21437/Interspeech.2019-2820</pub-id></mixed-citation>
</ref>
<ref id="B56">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Halliday</surname> <given-names>M. A.</given-names></name></person-group> (<year>1967</year>). <article-title>Notes on transitivity and theme in English: Part 2</article-title>. <source>J. Linguist</source>. <volume>3</volume>, <fpage>199</fpage>&#x02013;<lpage>244</lpage>. doi: <pub-id pub-id-type="doi">10.1017/S0022226700016613</pub-id></mixed-citation>
</ref>
<ref id="B57">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Horvitz</surname> <given-names>E.</given-names></name></person-group> (<year>1999</year>). <article-title>&#x0201C;Principles of mixed-initiative user interfaces,&#x0201D;</article-title> in <source>Proceedings of the SIGCHI conference on Human Factors in Computing Systems</source>, <fpage>159</fpage>&#x02013;<lpage>166</lpage>. doi: <pub-id pub-id-type="doi">10.1145/302979.303030</pub-id></mixed-citation>
</ref>
<ref id="B58">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Horwitz</surname> <given-names>R.</given-names></name> <name><surname>Brockhaus</surname> <given-names>S.</given-names></name> <name><surname>Henninger</surname> <given-names>F.</given-names></name> <name><surname>Kieslich</surname> <given-names>P. J.</given-names></name> <name><surname>Schierholz</surname> <given-names>M.</given-names></name> <name><surname>Keusch</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;Learning from mouse movements: improving questionnaires and respondents&#x00027; user experience through passive data collection,&#x0201D;</article-title> in <source>Advances in Questionnaire Design, Development, Evaluation and Testing</source>, <fpage>403</fpage>&#x02013;<lpage>425</lpage>. doi: <pub-id pub-id-type="doi">10.1002/9781119263685.ch16</pub-id></mixed-citation>
</ref>
<ref id="B59">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>White</surname> <given-names>R.</given-names></name> <name><surname>Buscher</surname> <given-names>G.</given-names></name></person-group> (<year>2012a</year>). <article-title>&#x0201C;User see, user point: gaze and cursor alignment in web search,&#x0201D;</article-title> in <source>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>, <fpage>1341</fpage>&#x02013;<lpage>1350</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2207676.2208591</pub-id></mixed-citation>
</ref>
<ref id="B60">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>White</surname> <given-names>R. W.</given-names></name> <name><surname>Buscher</surname> <given-names>G.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name></person-group> (<year>2012b</year>). <article-title>&#x0201C;Improving searcher models using mouse cursor activity,&#x0201D;</article-title> in <source>Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>, <fpage>195</fpage>&#x02013;<lpage>204</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2348283.2348313</pub-id></mixed-citation>
</ref>
<ref id="B61">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>White</surname> <given-names>R. W.</given-names></name> <name><surname>Dumais</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;No clicks, no problem: using cursor movements to understand and improve search,&#x0201D;</article-title> in <source>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>, <fpage>1225</fpage>&#x02013;<lpage>1234</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1978942.1979125</pub-id></mixed-citation>
</ref>
<ref id="B62">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Isaacs</surname> <given-names>E. A.</given-names></name> <name><surname>Clark</surname> <given-names>H. H.</given-names></name></person-group> (<year>1987</year>). <article-title>References in conversation between experts and novices</article-title>. <source>J. Exper. Psychol</source>. <volume>116</volume>:<fpage>26</fpage>. doi: <pub-id pub-id-type="doi">10.1037//0096-3445.116.1.26</pub-id></mixed-citation>
</ref>
<ref id="B63">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Jensen</surname> <given-names>L. C.</given-names></name> <name><surname>Langedijk</surname> <given-names>R. M.</given-names></name> <name><surname>Fischer</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Understanding the perception of incremental robot response in human-robot interaction,&#x0201D;</article-title> in <source>2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>41</fpage>&#x02013;<lpage>47</lpage>. doi: <pub-id pub-id-type="doi">10.1109/RO-MAN47096.2020.9223615</pub-id></mixed-citation>
</ref>
<ref id="B64">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>A.</given-names></name> <name><surname>Mulder</surname> <given-names>B.</given-names></name> <name><surname>Sijbinga</surname> <given-names>A.</given-names></name> <name><surname>Hulsebos</surname> <given-names>L.</given-names></name></person-group> (<year>2012</year>). <article-title>Action as a window to perception: measuring attention with mouse movements</article-title>. <source>Appl. Cogn. Psychol</source>. <volume>26</volume>, <fpage>802</fpage>&#x02013;<lpage>809</lpage>. doi: <pub-id pub-id-type="doi">10.1002/acp.2862</pub-id></mixed-citation>
</ref>
<ref id="B65">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kelleher</surname> <given-names>J.</given-names></name> <name><surname>Kruijff</surname> <given-names>G.-J. M.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x0201C;Incremental generation of spatial referring expressions in situated dialog,&#x0201D;</article-title> in <source>Proceedings of the 21st international conference on computational linguistics and 44th annual meeting of the association for computational linguistics</source>, <fpage>1041</fpage>&#x02013;<lpage>1048</lpage>. doi: <pub-id pub-id-type="doi">10.3115/1220175.1220306</pub-id></mixed-citation>
</ref>
<ref id="B66">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kennington</surname> <given-names>C.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>A simple generative model of incremental reference resolution for situated dialogue</article-title>. <source>Comput. Speech Lang</source>. <volume>41</volume>, <fpage>43</fpage>&#x02013;<lpage>67</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.csl.2016.04.002</pub-id></mixed-citation>
</ref>
<ref id="B67">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Kieslich</surname> <given-names>P. J.</given-names></name> <name><surname>Henninger</surname> <given-names>F.</given-names></name> <name><surname>Wulff</surname> <given-names>D. U.</given-names></name> <name><surname>Haslbeck</surname> <given-names>J. M.</given-names></name> <name><surname>Schulte-Mecklenbeck</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Mouse-tracking: a practical guide to implementation and analysis 1,&#x0201D;</article-title> in <source>A Handbook of Process Tracing Methods</source> (<publisher-loc>Routledge</publisher-loc>), <fpage>111</fpage>&#x02013;<lpage>130</lpage>. doi: <pub-id pub-id-type="doi">10.4324/9781315160559-9</pub-id></mixed-citation>
</ref>
<ref id="B68">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kirk</surname> <given-names>D. S.</given-names></name> <name><surname>Fraser</surname> <given-names>D. S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;The effects of remote gesturing on distance instruction,&#x0201D;</article-title> in <source>Computer Supported Collaborative Learning 2005: The Next 10 Years</source>! (Routledge), <fpage>301</fpage>&#x02013;<lpage>310</lpage>. doi: <pub-id pub-id-type="doi">10.3115/1149293.1149332</pub-id></mixed-citation>
</ref>
<ref id="B69">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kirsh</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Using mouse movement heatmaps to visualize user attention to words,&#x0201D;</article-title> in <source>Proceedings of the 11th Nordic Conference on Human-Computer Interaction: Shaping Experiences, Shaping Society</source>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3419249.3421250</pub-id></mixed-citation>
</ref>
<ref id="B70">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Koller</surname> <given-names>A.</given-names></name> <name><surname>Garoufi</surname> <given-names>K.</given-names></name> <name><surname>Staudte</surname> <given-names>M.</given-names></name> <name><surname>Crocker</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Enhancing referential success by tracking hearer gaze,&#x0201D;</article-title> in <source>Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue</source>, <fpage>30</fpage>&#x02013;<lpage>39</lpage>.</mixed-citation>
</ref>
<ref id="B71">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kontogiorgos</surname> <given-names>D.</given-names></name></person-group> (<year>2022</year>). <source>Mutual understanding in situated interactions with conversational user interfaces: theory, studies, and computation</source>. PhD thesis, KTH Royal Institute of Technology. doi: <pub-id pub-id-type="doi">10.31237/osf.io/fpts4</pub-id></mixed-citation>
</ref>
<ref id="B72">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kontogiorgos</surname> <given-names>D.</given-names></name> <name><surname>Gustafson</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Measuring collaboration load with pupillary responses-implications for the design of instructions in task-oriented HRI</article-title>. <source>Front. Psychol</source>. <volume>12</volume>:<fpage>623657</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2021.623657</pub-id></mixed-citation>
</ref>
<ref id="B73">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kontogiorgos</surname> <given-names>D.</given-names></name> <name><surname>Pelikan</surname> <given-names>H. R.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Towards adaptive and least-collaborative-effort social robots,&#x0201D;</article-title> in <source>Companion of the 2020 ACM/IEEE International Conference on Human-Robot Interaction</source>, <fpage>311</fpage>&#x02013;<lpage>313</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3371382.3378249</pub-id></mixed-citation>
</ref>
<ref id="B74">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kontogiorgos</surname> <given-names>D.</given-names></name> <name><surname>Pereira</surname> <given-names>A.</given-names></name> <name><surname>Gustafson</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Estimating uncertainty in task-oriented dialogue,&#x0201D;</article-title> in <source>2019 International Conference on Multimodal Interaction</source>, <fpage>414</fpage>&#x02013;<lpage>418</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3340555.3353722</pub-id></mixed-citation>
</ref>
<ref id="B75">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kousidis</surname> <given-names>S.</given-names></name> <name><surname>Kennington</surname> <given-names>C.</given-names></name> <name><surname>Baumann</surname> <given-names>T.</given-names></name> <name><surname>Buschmeier</surname> <given-names>H.</given-names></name> <name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Situationally aware in-car information presentation using incremental speech generation: safer, and more effective,&#x0201D;</article-title> in <source>Proceedings of the EACL 2014 Workshop on Dialogue in Motion</source>, <fpage>68</fpage>&#x02013;<lpage>72</lpage>. doi: <pub-id pub-id-type="doi">10.3115/v1/W14-0212</pub-id></mixed-citation>
</ref>
<ref id="B76">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Krahmer</surname> <given-names>E.</given-names></name> <name><surname>Van Deemter</surname> <given-names>K.</given-names></name></person-group> (<year>2012</year>). <article-title>Computational generation of referring expressions: a survey</article-title>. <source>Comput. Linguist</source>. <volume>38</volume>, <fpage>173</fpage>&#x02013;<lpage>218</lpage>. doi: <pub-id pub-id-type="doi">10.1162/COLI_a_00088</pub-id></mixed-citation>
</ref>
<ref id="B77">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Krassanakis</surname> <given-names>V.</given-names></name> <name><surname>Kesidis</surname> <given-names>A. L.</given-names></name></person-group> (<year>2020</year>). <article-title>Matmouse: a mouse movements tracking and analysis toolbox for visual search experiments</article-title>. <source>Multimodal Technol. Inter</source>. <volume>4</volume>:<fpage>83</fpage>. doi: <pub-id pub-id-type="doi">10.3390/mti4040083</pub-id></mixed-citation>
</ref>
<ref id="B78">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kraut</surname> <given-names>R. E.</given-names></name> <name><surname>Fussell</surname> <given-names>S. R.</given-names></name> <name><surname>Siegel</surname> <given-names>J.</given-names></name></person-group> (<year>2003</year>). <article-title>Visual information as a conversational resource in collaborative physical tasks</article-title>. <source>Hum. Comput. Inter</source>. <volume>18</volume>, <fpage>13</fpage>&#x02013;<lpage>49</lpage>. doi: <pub-id pub-id-type="doi">10.1207/S15327051HCI1812_2</pub-id></mixed-citation>
</ref>
<ref id="B79">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Niu</surname> <given-names>T.</given-names></name> <name><surname>Feng</surname> <given-names>F.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Referring expression generation via visual dialogue,&#x0201D;</article-title> in <source>CCF International Conference on Natural Language Processing and Chinese Computing</source> (<publisher-loc>Springer</publisher-loc>), <fpage>28</fpage>&#x02013;<lpage>40</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-030-60457-8_3</pub-id></mixed-citation>
</ref>
<ref id="B80">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Lindwall</surname> <given-names>O.</given-names></name> <name><surname>Ekstr&#x000F6;m</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Instruction-in-interaction: the teaching and learning of a manual skill</article-title>. <source>Hum. Stud</source>. <volume>35</volume>, <fpage>27</fpage>&#x02013;<lpage>49</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10746-012-9213-5</pub-id></mixed-citation>
</ref>
<ref id="B81">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>M. X.</given-names></name> <name><surname>Sarkar</surname> <given-names>A.</given-names></name> <name><surname>Negreanu</surname> <given-names>C.</given-names></name> <name><surname>Zorn</surname> <given-names>B.</given-names></name> <name><surname>Williams</surname> <given-names>J.</given-names></name> <name><surname>Toronto</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>&#x0201C;what it wants me to say&#x0201D;: Bridging the abstraction gap between end-user programmers and code-generating large language models,&#x0201D;</article-title> in <source>Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems</source>, <fpage>1</fpage>&#x02013;<lpage>31</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3544548.3580817</pub-id></mixed-citation>
</ref>
<ref id="B82">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Magassouba</surname> <given-names>A.</given-names></name> <name><surname>Sugiura</surname> <given-names>K.</given-names></name> <name><surname>Kawai</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>A multimodal classifier generative adversarial network for carry and place tasks from ambiguous language instructions</article-title>. <source>IEEE Robot. Autom. Lett</source>. <volume>3</volume>, <fpage>3113</fpage>&#x02013;<lpage>3120</lpage>. doi: <pub-id pub-id-type="doi">10.1109/LRA.2018.2849607</pub-id></mixed-citation>
</ref>
<ref id="B83">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mitev</surname> <given-names>N.</given-names></name> <name><surname>Renner</surname> <given-names>P.</given-names></name> <name><surname>Pfeiffer</surname> <given-names>T.</given-names></name> <name><surname>Staudte</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Using listener gaze to refer in installments benefits understanding,&#x0201D;</article-title> in <source>CogSci</source>.</mixed-citation>
</ref>
<ref id="B84">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Monaro</surname> <given-names>M.</given-names></name> <name><surname>Gamberini</surname> <given-names>L.</given-names></name> <name><surname>Sartori</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>The detection of faked identity using unexpected questions and mouse dynamics</article-title>. <source>PLoS ONE</source> <volume>12</volume>:<fpage>e0177851</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0177851</pub-id><pub-id pub-id-type="pmid">28542248</pub-id></mixed-citation>
</ref>
<ref id="B85">
<mixed-citation publication-type="web"><person-group person-group-type="author"><name><surname>Morawiec</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <source>sklearn-porter. Transpile trained scikit-learn estimators to C, Java, JavaScript and others</source>. <ext-link ext-link-type="uri" xlink:href="https://github.com/nok/sklearn-porter">https://github.com/nok/sklearn-porter</ext-link></mixed-citation>
</ref>
<ref id="B86">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mueller</surname> <given-names>F.</given-names></name> <name><surname>Lockerd</surname> <given-names>A.</given-names></name></person-group> (<year>2001</year>). <article-title>&#x0201C;Cheese: tracking mouse movement activity on websites, a tool for user modeling,&#x0201D;</article-title> in <source>CHI&#x00027;01 Extended Abstracts on Human Factors in Computing Systems</source>, <fpage>279</fpage>&#x02013;<lpage>280</lpage>. doi: <pub-id pub-id-type="doi">10.1145/634067.634233</pub-id></mixed-citation>
</ref>
<ref id="B87">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>H. J.</given-names></name> <name><surname>Krummenacher</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Visual search and selective attention</article-title>. <source>Vis. Cogn</source>. <volume>14</volume>, <fpage>389</fpage>&#x02013;<lpage>410</lpage>. doi: <pub-id pub-id-type="doi">10.1080/13506280500527676</pub-id></mixed-citation>
</ref>
<ref id="B88">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Pedregosa</surname> <given-names>F.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Gramfort</surname> <given-names>A.</given-names></name> <name><surname>Michel</surname> <given-names>V.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Grisel</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Scikit-learn: machine learning in python</article-title>. <source>J. Mach. Learn. Res</source>. <volume>12</volume>, <fpage>2825</fpage>&#x02013;<lpage>2830</lpage>. doi: <pub-id pub-id-type="doi">10.5555/1953048.2078195</pub-id></mixed-citation>
</ref>
<ref id="B89">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Pelikan</surname> <given-names>H. R.</given-names></name> <name><surname>Broth</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Why that NAO? How humans adapt to a conventional humanoid robot in taking turns-at-talk,&#x0201D;</article-title> in <source>Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems</source>, <fpage>4921</fpage>&#x02013;<lpage>4932</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2858036.2858478</pub-id></mixed-citation>
</ref>
<ref id="B90">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Qvarfordt</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Gaze-informed multimodal interaction,&#x0201D;</article-title> in <source>The Handbook of Multimodal-Multisensor Interfaces: Foundations, User Modeling, and Common Modality Combinations</source>, <fpage>365</fpage>&#x02013;<lpage>402</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3015783.3015794</pub-id></mixed-citation>
</ref>
<ref id="B91">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Reigeluth</surname> <given-names>C. M.</given-names></name> <name><surname>Merrill</surname> <given-names>M. D.</given-names></name> <name><surname>Wilson</surname> <given-names>B. G.</given-names></name> <name><surname>Spiller</surname> <given-names>R. T.</given-names></name></person-group> (<year>1980</year>). <article-title>The elaboration theory of instruction: a model for sequencing and synthesizing instruction</article-title>. <source>Instruct. Sci</source>. <volume>9</volume>, <fpage>195</fpage>&#x02013;<lpage>219</lpage>. doi: <pub-id pub-id-type="doi">10.1007/BF00177327</pub-id></mixed-citation>
</ref>
<ref id="B92">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Reimers</surname> <given-names>N.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name></person-group> (<year>2019</year>). <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>. <source>arXiv preprint arXiv:1908.10084</source>.</mixed-citation>
</ref>
<ref id="B93">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Rheem</surname> <given-names>H.</given-names></name> <name><surname>Verma</surname> <given-names>V.</given-names></name> <name><surname>Becker</surname> <given-names>D. V.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Use of mouse-tracking method to measure cognitive load,&#x0201D;</article-title> in <source>Proceedings of the Human Factors and Ergonomics Society Annual Meeting</source> (<publisher-loc>Los Angeles, CA</publisher-loc>: <publisher-name>SAGE Publications Sage CA</publisher-name>), <fpage>1982</fpage>&#x02013;<lpage>1986</lpage>. doi: <pub-id pub-id-type="doi">10.1177/1541931218621449</pub-id></mixed-citation>
</ref>
<ref id="B94">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Rojowiec</surname> <given-names>R.</given-names></name> <name><surname>G&#x000F6;tze</surname> <given-names>J.</given-names></name> <name><surname>Sadler</surname> <given-names>P.</given-names></name> <name><surname>Voigt</surname> <given-names>H.</given-names></name> <name><surname>Zarrie&#x000DF;</surname> <given-names>S.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;From &#x0201C;before&#x0201D; to &#x0201C;after&#x0201D;: Generating natural language instructions from image pairs in a simple visual domain,&#x0201D;</article-title> in <source>Proceedings of the 13th International Conference on Natural Language Generation</source>, <fpage>316</fpage>&#x02013;<lpage>326</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/2020.inlg-1.38</pub-id></mixed-citation>
</ref>
<ref id="B95">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Rookhuiszen</surname> <given-names>R. B.</given-names></name> <name><surname>Obbink</surname> <given-names>M.</given-names></name> <name><surname>Theune</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Two approaches to give: dynamic level adaptation versus playfulness,&#x0201D;</article-title> in <source>Proceedings of the First NLG Challenge on Generating Instructions in Virtual Environments</source>.</mixed-citation>
</ref>
<ref id="B96">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Sacks</surname> <given-names>H.</given-names></name> <name><surname>Schegloff</surname> <given-names>E. A.</given-names></name> <name><surname>Jefferson</surname> <given-names>G.</given-names></name></person-group> (<year>1978</year>). <article-title>&#x0201C;A simplest systematics for the organization of turn taking for conversation,&#x0201D;</article-title> in <source>Studies in the Organization of Conversational Interaction</source> (<publisher-loc>Elsevier</publisher-loc>), <fpage>7</fpage>&#x02013;<lpage>55</lpage>.</mixed-citation>
</ref>
<ref id="B97">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sadler</surname> <given-names>P.</given-names></name> <name><surname>Hakimov</surname> <given-names>S.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>Yes, this way! Learning to ground referring expressions into actions with intra-episodic feedback from supportive teachers</article-title>. <source>arXiv preprint arXiv:2305.12880</source>.</mixed-citation>
</ref>
<ref id="B98">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sadler</surname> <given-names>P.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Pento-diaref: a diagnostic dataset for learning the incremental algorithm for referring expression generation from examples,&#x0201D;</article-title> in <source>Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics</source>, <fpage>2098</fpage>&#x02013;<lpage>2114</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/2023.eacl-main.154</pub-id></mixed-citation>
</ref>
<ref id="B99">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Saupp</surname> <given-names>A.</given-names></name> <name><surname>Mutlu</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Effective task training strategies for instructional robots,&#x0201D;</article-title> in <source>Proceedings of the 10th Annual Robotics: Science and Systems Conference</source>. doi: <pub-id pub-id-type="doi">10.15607/RSS.2014.X.002</pub-id></mixed-citation>
</ref>
<ref id="B100">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Saupp&#x000E9;</surname> <given-names>A.</given-names></name> <name><surname>Mutlu</surname> <given-names>B.</given-names></name></person-group> (<year>2015</year>). <article-title>Effective task training strategies for human and robot instructors</article-title>. <source>Auton. Robots</source> <volume>39</volume>, <fpage>313</fpage>&#x02013;<lpage>329</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10514-015-9461-0</pub-id></mixed-citation>
</ref>
<ref id="B101">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schlangen</surname> <given-names>D.</given-names></name> <name><surname>Fern&#x000E1;ndez</surname> <given-names>R.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Beyond repair-testing the limits of the conversational repair system,&#x0201D;</article-title> in <source>Proceedings of the 8th SIGdial Workshop on Discourse and Dialogue</source>, <fpage>51</fpage>&#x02013;<lpage>54</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/2007.sigdial-1.10</pub-id></mixed-citation>
</ref>
<ref id="B102">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schlangen</surname> <given-names>D.</given-names></name> <name><surname>Fern&#x000E1;ndez</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <source>The Potsdam Dialogue Corpora: Transcription and Annotation Manual.</source> University of Potsdam</mixed-citation>
</ref>
<ref id="B103">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schober</surname> <given-names>M. F.</given-names></name> <name><surname>Clark</surname> <given-names>H. H.</given-names></name></person-group> (<year>1989</year>). <article-title>Understanding by addressees and overhearers</article-title>. <source>Cogn. Psychol</source>. <volume>21</volume>, <fpage>211</fpage>&#x02013;<lpage>232</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0010-0285(89)90008-X</pub-id></mixed-citation>
</ref>
<ref id="B104">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schoemann</surname> <given-names>M.</given-names></name> <name><surname>O&#x00027;Hora</surname> <given-names>D.</given-names></name> <name><surname>Dale</surname> <given-names>R.</given-names></name> <name><surname>Scherbaum</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Using mouse cursor tracking to investigate online cognition: preserving methodological ingenuity while moving toward reproducible science</article-title>. <source>Psychon. Bull. Rev</source>. <volume>28</volume>, <fpage>766</fpage>&#x02013;<lpage>787</lpage>. doi: <pub-id pub-id-type="doi">10.3758/s13423-020-01851-3</pub-id><pub-id pub-id-type="pmid">33319317</pub-id></mixed-citation>
</ref>
<ref id="B105">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Skantze</surname> <given-names>G.</given-names></name> <name><surname>Hjalmarsson</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Towards incremental speech generation in dialogue systems,&#x0201D;</article-title> in <source>Proceedings of the SIGDIAL 2010 Conference</source>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</mixed-citation>
</ref>
<ref id="B106">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Smucker</surname> <given-names>M. D.</given-names></name> <name><surname>Guo</surname> <given-names>X. S.</given-names></name> <name><surname>Toulis</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Mouse movement during relevance judging: implications for determining user attention,&#x0201D;</article-title> in <source>Proceedings of the 37th International ACM SIGIR Conference on Research &#x00026;Development in Information Retrieval</source>, <fpage>979</fpage>&#x02013;<lpage>982</lpage>. doi: <pub-id pub-id-type="doi">10.1145/2600428.2609489</pub-id></mixed-citation>
</ref>
<ref id="B107">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Staudte</surname> <given-names>M.</given-names></name> <name><surname>Koller</surname> <given-names>A.</given-names></name> <name><surname>Garoufi</surname> <given-names>K.</given-names></name> <name><surname>Crocker</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Using listener gaze to augment speech generation in a virtual 3D environment,&#x0201D;</article-title> in <source>Proceedings of the Annual Meeting of the Cognitive Science Society</source>.</mixed-citation>
</ref>
<ref id="B108">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Stoia</surname> <given-names>L.</given-names></name> <name><surname>Shockley</surname> <given-names>D. M.</given-names></name> <name><surname>Byron</surname> <given-names>D.</given-names></name> <name><surname>Fosler-Lussier</surname> <given-names>E.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x0201C;Noun phrase generation for situated dialogs,&#x0201D;</article-title> in <source>Proceedings of the Fourth International Natural Language Generation Conference</source>, <fpage>81</fpage>&#x02013;<lpage>88</lpage>. doi: <pub-id pub-id-type="doi">10.3115/1706269.1706286</pub-id></mixed-citation>
</ref>
<ref id="B109">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Striegnitz</surname> <given-names>K.</given-names></name> <name><surname>Buschmeier</surname> <given-names>H.</given-names></name> <name><surname>Kopp</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Referring in installments: a corpus study of spoken object references in an interactive virtual environment,&#x0201D;</article-title> in <source>INLG 2012 Proceedings of the Seventh International Natural Language Generation Conference</source>, <fpage>12</fpage>&#x02013;<lpage>16</lpage>.</mixed-citation>
</ref>
<ref id="B110">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sz&#x000E9;kely</surname> <given-names>&#x000C9;.</given-names></name> <name><surname>Henter</surname> <given-names>G. E.</given-names></name> <name><surname>Beskow</surname> <given-names>J.</given-names></name> <name><surname>Gustafson</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;How to train your fillers: uh and um in spontaneous speech synthesis,&#x0201D;</article-title> in <source>The 10th ISCA Speech Synthesis Workshop</source>. doi: <pub-id pub-id-type="doi">10.21437/SSW.2019-44</pub-id></mixed-citation>
</ref>
<ref id="B111">
<mixed-citation publication-type="web"><person-group person-group-type="author"><name><surname>Team</surname> <given-names>R Developement Core.</given-names></name></person-group> (<year>2009</year>). <source>A language and environment for statistical computing</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.R-project.org">http://www.R-project.org</ext-link> (Accessed November 19, 2025).</mixed-citation>
</ref>
<ref id="B112">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tellex</surname> <given-names>S.</given-names></name> <name><surname>Kollar</surname> <given-names>T.</given-names></name> <name><surname>Dickerson</surname> <given-names>S.</given-names></name> <name><surname>Walter</surname> <given-names>M.</given-names></name> <name><surname>Banerjee</surname> <given-names>A.</given-names></name> <name><surname>Teller</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>&#x0201C;Understanding natural language commands for robotic navigation and mobile manipulation,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, <fpage>1507</fpage>&#x02013;<lpage>1514</lpage>. doi: <pub-id pub-id-type="doi">10.1609/aaai.v25i1.7979</pub-id></mixed-citation>
</ref>
<ref id="B113">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tomlinson Jr</surname> <given-names>J. M.</given-names></name> <name><surname>Assimakopoulos</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;The dynamics of pragmatic enrichment during metaphor processing: activation vs. suppression,&#x0201D;</article-title> in <source>Proceedings of the Annual Meeting of the Cognitive Science Society</source>.</mixed-citation>
</ref>
<ref id="B114">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tomlinson Jr</surname> <given-names>J. M.</given-names></name> <name><surname>Bott</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;How intonation contrains pragmatic inference,&#x0201D;</article-title> in <source>Proceedings of the Annual Meeting of the Cognitive Science Society</source>.</mixed-citation>
</ref>
<ref id="B115">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Torrey</surname> <given-names>C.</given-names></name> <name><surname>Fussell</surname> <given-names>S. R.</given-names></name> <name><surname>Kiesler</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;How a robot should give advice,&#x0201D;</article-title> in <source>2013 8th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>275</fpage>&#x02013;<lpage>282</lpage>. doi: <pub-id pub-id-type="doi">10.1109/HRI.2013.6483599</pub-id></mixed-citation>
</ref>
<ref id="B116">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Torrey</surname> <given-names>C.</given-names></name> <name><surname>Powers</surname> <given-names>A.</given-names></name> <name><surname>Fussell</surname> <given-names>S. R.</given-names></name> <name><surname>Kiesler</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Exploring adaptive dialogue based on a robot&#x00027;s awareness of human gaze and task progress,&#x0201D;</article-title> in <source>Proceedings of the ACM/IEEE International Conference on Human-Robot Interaction</source>, <fpage>247</fpage>&#x02013;<lpage>254</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1228716.1228750</pub-id></mixed-citation>
</ref>
<ref id="B117">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Torrey</surname> <given-names>C.</given-names></name> <name><surname>Powers</surname> <given-names>A.</given-names></name> <name><surname>Marge</surname> <given-names>M.</given-names></name> <name><surname>Fussell</surname> <given-names>S. R.</given-names></name> <name><surname>Kiesler</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x0201C;Effects of adaptive robot dialogue on information exchange and social relations,&#x0201D;</article-title> in <source>Proceedings of the 1st ACM SIGCHI/SIGART Conference on Human-Robot Interaction</source>, <fpage>126</fpage>&#x02013;<lpage>133</lpage>. doi: <pub-id pub-id-type="doi">10.1145/1121241.1121264</pub-id></mixed-citation>
</ref>
<ref id="B118">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Traum</surname> <given-names>D. R.</given-names></name> <name><surname>Hinkelman</surname> <given-names>E. A.</given-names></name></person-group> (<year>1992</year>). <article-title>Conversation acts in task-oriented spoken dialogue</article-title>. <source>Comput. Intell</source>. <volume>8</volume>, <fpage>575</fpage>&#x02013;<lpage>599</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-8640.1992.tb00380.x</pub-id></mixed-citation>
</ref>
<ref id="B119">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wachsmuth</surname> <given-names>I.</given-names></name> <name><surname>Lenzen</surname> <given-names>M.</given-names></name> <name><surname>Knoblich</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Embodied communication in humans and machines</article-title>. <source>AI Magaz</source>. <volume>26</volume>, <fpage>85</fpage>&#x02013;<lpage>86</lpage>. doi: <pub-id pub-id-type="doi">10.1093/acprof:oso/9780199231751.001.0001</pub-id></mixed-citation>
</ref>
<ref id="B120">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wagner</surname> <given-names>P.</given-names></name> <name><surname>Trouvain</surname> <given-names>J.</given-names></name> <name><surname>Zimmerer</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>In defense of stylistic diversity in speech research</article-title>. <source>J. Phon</source>. <volume>48</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.wocn.2014.11.001</pub-id></mixed-citation>
</ref>
<ref id="B121">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wallbridge</surname> <given-names>C. D.</given-names></name> <name><surname>Lemaignan</surname> <given-names>S.</given-names></name> <name><surname>Senft</surname> <given-names>E.</given-names></name> <name><surname>Belpaeme</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>Generating spatial referring expressions in a social robot: dynamic vs. non-ambiguous</article-title>. <source>Front. Robot. AI</source> <volume>6</volume>:<fpage>67</fpage>. doi: <pub-id pub-id-type="doi">10.3389/frobt.2019.00067</pub-id><pub-id pub-id-type="pmid">33501082</pub-id></mixed-citation>
</ref>
<ref id="B122">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wallbridge</surname> <given-names>C. D.</given-names></name> <name><surname>Smith</surname> <given-names>A.</given-names></name> <name><surname>Giuliani</surname> <given-names>M.</given-names></name> <name><surname>Melhuish</surname> <given-names>C.</given-names></name> <name><surname>Belpaeme</surname> <given-names>T.</given-names></name> <name><surname>Lemaignan</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>The effectiveness of dynamically processed incremental descriptions in human robot interaction</article-title>. <source>ACM Trans. Hum. Robot Inter</source>. <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>24</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3481628</pub-id></mixed-citation>
</ref>
<ref id="B123">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Weerakoon</surname> <given-names>D.</given-names></name> <name><surname>Subbaraju</surname> <given-names>V.</given-names></name> <name><surname>Karumpulli</surname> <given-names>N.</given-names></name> <name><surname>Tran</surname> <given-names>T.</given-names></name> <name><surname>Xu</surname> <given-names>Q.</given-names></name> <name><surname>Tan</surname> <given-names>U.-X.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;Gesture enhanced comprehension of ambiguous human-to-robot instructions,&#x0201D;</article-title> in <source>Proceedings of the 2020 International Conference on Multimodal Interaction</source>, <fpage>251</fpage>&#x02013;<lpage>259</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3382507.3418863</pub-id></mixed-citation>
</ref>
<ref id="B124">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Whisenand</surname> <given-names>T. G.</given-names></name> <name><surname>Emurian</surname> <given-names>H. H.</given-names></name></person-group> (<year>1999</year>). <article-title>Analysis of cursor movements with a mouse</article-title>. <source>Comput. Human Behav</source>. <volume>15</volume>, <fpage>85</fpage>&#x02013;<lpage>103</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0747-5632(98)00036-3</pub-id></mixed-citation>
</ref>
<ref id="B125">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Williams</surname> <given-names>T.</given-names></name> <name><surname>Scheutz</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Referring expression generation under uncertainty: algorithm and evaluation framework,&#x0201D;</article-title> in <source>Proceedings of the 10th International Conference on Natural Language Generation</source>, <fpage>75</fpage>&#x02013;<lpage>84</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/W17-3511</pub-id></mixed-citation>
</ref>
<ref id="B126">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>K.</given-names></name> <name><surname>Yamauchi</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>Semantic priming revealed by mouse movement trajectories</article-title>. <source>Conscious. Cogn</source>. <volume>27</volume>, <fpage>42</fpage>&#x02013;<lpage>52</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.concog.2014.04.004</pub-id><pub-id pub-id-type="pmid">24797040</pub-id></mixed-citation>
</ref>
<ref id="B127">
<mixed-citation publication-type="book"><person-group person-group-type="author"><name><surname>Yoon</surname> <given-names>E. J.</given-names></name> <name><surname>Tessler</surname> <given-names>M. H.</given-names></name> <name><surname>Goodman</surname> <given-names>N. D.</given-names></name> <name><surname>Frank</surname> <given-names>M. C.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Talking with tact: polite language as a balance between kindness and informativity,&#x0201D;</article-title> in <source>Proceedings of the 38th Annual Conference of the Cognitive Science Society</source> (<publisher-loc>Cognitive Science Society</publisher-loc>), <fpage>2771</fpage>&#x02013;<lpage>2776</lpage>.</mixed-citation>
</ref>
<ref id="B128">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zarrie&#x000DF;</surname> <given-names>S.</given-names></name> <name><surname>Hough</surname> <given-names>J.</given-names></name> <name><surname>Kennington</surname> <given-names>C.</given-names></name> <name><surname>Manuvinakurike</surname> <given-names>R.</given-names></name> <name><surname>DeVault</surname> <given-names>D.</given-names></name> <name><surname>Fern&#x000E1;ndez</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>&#x0201C;Pentoref: a corpus of spoken references in task-oriented dialogues,&#x0201D;</article-title> in <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC&#x00027;16)</source>, <fpage>125</fpage>&#x02013;<lpage>131</lpage>.</mixed-citation>
</ref>
<ref id="B129">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zarrie&#x000DF;</surname> <given-names>S.</given-names></name> <name><surname>Schlangen</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Easy things first: Installments improve referring expression generation for objects in photographs,&#x0201D;</article-title> in <source>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>, <fpage>610</fpage>&#x02013;<lpage>620</lpage>. doi: <pub-id pub-id-type="doi">10.18653/v1/P16-1058</pub-id></mixed-citation>
</ref>
<ref id="B130">
<mixed-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zellner</surname> <given-names>B.</given-names></name></person-group> (<year>1994</year>). <article-title>&#x0201C;Pauses and the temporal structure of speech,&#x0201D;</article-title> in E. Keller (Ed.) <italic>Fundamentals of Speech Synthesis and Speech Recognition</italic> (Chichester: John Wiley), <fpage>41</fpage>&#x02013;<lpage>62</lpage>.</mixed-citation>
</ref>
</ref-list>
<fn-group>
<fn fn-type="custom" custom-type="edited-by" id="fn0001">
<p>Edited by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1577564/overview">Xianmin Wang</ext-link>, Guangzhou University, China</p></fn>
<fn fn-type="custom" custom-type="reviewed-by" id="fn0002">
<p>Reviewed by: <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/3174324/overview">Ayain John</ext-link>, Dayanand Sagar Academy of Technology and Management, India</p>
<p><ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/3188012/overview">Qiaoqiao Ren</ext-link>, Ghent University, Belgium</p></fn>
</fn-group>
<fn-group>
<fn id="fn0003"><label>1</label><p>Referring expressions (REs) are utterances that often involve language identifying entities in the physical space (e.g., &#x0201C;<italic>this one here&#x0201D;</italic>) or abstract entities (e.g., &#x0201C;<italic>Grace Hopper was here&#x0201D;</italic>) (<xref ref-type="bibr" rid="B62">Isaacs and Clark, 1987</xref>).</p></fn>
<fn id="fn0004"><label>2</label><p>For an overview of visual search and attention, see the work of (<xref ref-type="bibr" rid="B87">M&#x000FC;ller and Krummenacher, 2006</xref>) and (<xref ref-type="bibr" rid="B38">Eckstein, 2011</xref>).</p></fn>
<fn id="fn0005"><label>3</label><p>For an overview of the concept of information structure and utterances as informational units, see (<xref ref-type="bibr" rid="B56">Halliday, 1967</xref>).</p></fn>
</fn-group>
</back>
</article>