<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1127626</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2023.1127626</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Does a robot&#x2019;s gaze aversion affect human gaze aversion?</article-title>
<alt-title alt-title-type="left-running-head">Mishra et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frobt.2023.1127626">10.3389/frobt.2023.1127626</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Mishra</surname>
<given-names>Chinmaya</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1673551/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Offrede</surname>
<given-names>Tom</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2146258/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Fuchs</surname>
<given-names>Susanne</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/75815/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mooshammer</surname>
<given-names>Christine</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/829836/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Skantze</surname>
<given-names>Gabriel</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/747537/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Furhat Robotics AB</institution>, <addr-line>Stockholm</addr-line>, <country>Sweden</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Humboldt-Universit&#xe4;t zu Berlin</institution>, <addr-line>Berlin</addr-line>, <country>Germany</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Leibniz-Centre General Linguistics (ZAS)</institution>, <addr-line>Berlin</addr-line>, <country>Germany</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Division of Speech</institution>, <institution>Music and Hearing</institution>, <institution>KTH Royal Institute of Technology</institution>, <addr-line>Stockholm</addr-line>, <country>Sweden</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/166210/overview">Ramana Vinjamuri</ext-link>, University of Maryland, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/143256/overview">Christian Balkenius</ext-link>, Lund University, Sweden</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/267871/overview">Erik A. Billing</ext-link>, University of Sk&#xf6;vde, Sweden</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Gabriel Skantze, <email>skantze@kth.se</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>06</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>10</volume>
<elocation-id>1127626</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>05</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Mishra, Offrede, Fuchs, Mooshammer and Skantze.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Mishra, Offrede, Fuchs, Mooshammer and Skantze</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Gaze cues serve an important role in facilitating human conversations and are generally considered to be one of the most important non-verbal cues. Gaze cues are used to manage turn-taking, coordinate joint attention, regulate intimacy, and signal cognitive effort. In particular, it is well established that gaze aversion is used in conversations to avoid prolonged periods of mutual gaze. Given the numerous functions of gaze cues, there has been extensive work on modelling these cues in social robots. Researchers have also tried to identify the impact of robot gaze on human participants. However, the influence of robot gaze behavior on human gaze behavior has been less explored. We conducted a within-subjects user study (N &#x3d; 33) to verify if a robot&#x2019;s gaze aversion influenced human gaze aversion behavior. Our results show that participants tend to avert their gaze more when the robot keeps staring at them as compared to when the robot exhibits well-timed gaze aversions. We interpret our findings in terms of intimacy regulation: humans try to compensate for the robot&#x2019;s lack of gaze aversion.</p>
</abstract>
<kwd-group>
<kwd>gaze</kwd>
<kwd>gaze aversion</kwd>
<kwd>human-robot interaction</kwd>
<kwd>social robot</kwd>
<kwd>gaze control model</kwd>
<kwd>gaze behavior</kwd>
<kwd>intimacy</kwd>
<kwd>topic intimacy</kwd>
</kwd-group>
<contract-sponsor id="cn001">H2020 Marie Sk&#x142;odowska-Curie Actions<named-content content-type="fundref-id">10.13039/100010665</named-content>
</contract-sponsor>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Human-Robot Interaction</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>It is well established that gaze cues are one of the most important non-verbal cues used in Human-Human Interactions (HHI) (<xref ref-type="bibr" rid="B21">Kendon, 1967</xref>). Several studies have shown the many roles gaze cues play in facilitating human interactions. When interacting with each other, people use gaze to coordinate joint attention, communicating their focus of attention and perceiving their partner&#x2019;s focus to follow (<xref ref-type="bibr" rid="B36">Tomasello, 1995</xref>). <xref ref-type="bibr" rid="B17">Ho et al. (2015)</xref> also showed how people use gaze to manage turn-taking: for instance, gaze directed at or averted from one&#x2019;s interlocutor can indicate whether a speaker is intending to yield or hold the turn (for example, when making a pause), or when the listener is intending to take the turn.</p>
<p>Given the importance of gaze behavior in HHI, researchers in Human-Robot Interaction (HRI) have tried to emulate human-like gaze behaviors in robots. The main motivation behind such Gaze Control Systems (GCS), or models of gaze behavior, has been to exploit the many functionalities of gaze cues in HHI and realize them in HRI. Moreover, thanks to the sophisticated anthropomorphic design of many of today&#x2019;s social robots (e.g., Furhat robot (<xref ref-type="bibr" rid="B28">Moubayed et al., 2013</xref>) or iCub robot (<xref ref-type="bibr" rid="B26">Metta et al., 2010</xref>)), it is possible to model nuanced gaze behaviors with independent eye and head movements. It has been established that robots&#x2019; gaze behaviors are recognized and perceived to be intentional by humans (<xref ref-type="bibr" rid="B4">Andrist et al., 2014</xref>). Robots&#x2019; gaze behaviors have also been found to play an equally important role in HRI as human gaze in HHI (<xref ref-type="bibr" rid="B19">Imai et al., 2002</xref>; <xref ref-type="bibr" rid="B37">Yamazaki et al., 2008</xref>). Thus, researchers have measured the impact of robots&#x2019; gaze behavior on human behavior during HRI. In <xref ref-type="bibr" rid="B33">Schellen et al. (2021)</xref> participants were found to become more honest in subsequent trials if the robot looked at them when they were being deceptive. <xref ref-type="bibr" rid="B43">Skantze (2017)</xref> and <xref ref-type="bibr" rid="B14">Gillet et al. (2021)</xref> observed that robots&#x2019; gaze behavior could lead to more participation during group activities. Most of these works have concentrated on human behavior in general, but not the gaze-to-gaze interaction between robots and humans. This then leads to our research question.<list list-type="simple">
<list-item>
<p>&#x2022;<italic>Does a robot&#x2019;s gaze behavior have any influence on human gaze behavior in a HRI?</italic>
</p>
</list-item>
</list>
</p>
<p>Answering this question is important because it can help in designing better GCS and interactions in HRI. Even though previous works have shown various ways in which humans perceive and respond to robot gaze behavior, whether there are changes in human gaze behavior as a direct influence of robots&#x2019; gaze behavior has remained less explored. Moreover, most of these studies have used head movements instead of eye gaze to model robot gaze behavior, due to physical constraints of the robots used (<xref ref-type="bibr" rid="B4">Andrist et al., 2014</xref>; <xref ref-type="bibr" rid="B25">Mehlmann et al., 2014</xref>; <xref ref-type="bibr" rid="B30">Nakano et al., 2015</xref>). While head orientation is a good approximation of gaze behavior in general, it lacks the rich information ingrained in eye gaze. Additionally, from a motor control perspective, eye gaze is much quicker than head motion and is therefore also more adaptable than moving the head. Thus, we were interested in verifying if subtle gaze cues performed by a robot are perceived by humans and if it had any influence on their own gaze behavior.</p>
<p>In order to verify the impact of robot gaze behavior, we narrowed our focus to gaze aversions for this study. This was mainly motivated by two considerations. First, gaze aversion has been shown to play an important role in human conversations: coordinating turn-taking (<xref ref-type="bibr" rid="B17">Ho et al., 2015</xref>), regulating intimacy (<xref ref-type="bibr" rid="B1">Abele, 1986</xref>) and signalling cognitive load (<xref ref-type="bibr" rid="B13">Doherty-Sneddon and Phelps, 2005</xref>). Secondly, it is an important gaze cue which is relatively easy to perceive and generate during HRI.</p>
<p>In this work, we designed a within-subjects user study to measure if gaze aversion exhibited by a robot has any influence on the gaze aversion behavior of participants. We automated the robot&#x2019;s gaze using the GCS proposed in <xref ref-type="bibr" rid="B27">Mishra and Skantze (2022)</xref> (more details in <xref ref-type="sec" rid="s3">Section 3</xref>) to exhibit time- and context-appropriate gaze aversions. Participants&#x2019; gaze was tracked using eye-tracking glasses throughout the interactions. Subjective responses were also collected from the participants after the experiment, using a questionnaire that asked about their impression of the interaction. Our results show that participants avert their gaze more when the robot doesn&#x2019;t avert its gaze as compared to when it does.</p>
<p>The main contributions of this paper are:<list list-type="simple">
<list-item>
<p>&#x2022; The first study (to the best of our knowledge) that verified the existence of a direct relationship between robot gaze aversion and human gaze aversion.</p>
</list-item>
<list-item>
<p>&#x2022; A study design to measure the influence of a robot&#x2019;s gaze behavior on human gaze behavior.</p>
</list-item>
<list-item>
<p>&#x2022; An exploratory analysis of the eye gaze data, which pointed towards a potential positive correlation between gaze aversion and topic intimacy of the questions.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Background</title>
<p>
<bold>Gaze aversion</bold> is the act of shifting the gaze away from one&#x2019;s interaction partner during a conversation. Speakers tend to look away from the listener more often than the other way around during a conversation. This has been thought to help plan the upcoming utterance and avoid distractions (<xref ref-type="bibr" rid="B5">Argyle and Cook, 1976</xref>). It has been found that holding mutual gaze significantly increases hesitations and false starts (<xref ref-type="bibr" rid="B7">Beattie, 1981</xref>). Speakers process visual information from their interlocutors, produce speech and plan the upcoming speech, all at the same time. Prior studies in HHI have shown that people use gaze aversions to manage cognitive load (<xref ref-type="bibr" rid="B13">Doherty-Sneddon and Phelps, 2005</xref>) because averting gaze reduces the load of processing the visual information. <xref ref-type="bibr" rid="B17">Ho et al. (2015)</xref> found that speakers signal their desire to retain the current turn, i.e., turn-holding, by averting their gaze and that they begin their turns with averted gaze. Additionally, gaze aversion has been found to have a significant contribution in regulating the intimacy level during a conversation (<xref ref-type="bibr" rid="B1">Abele, 1986</xref>). <xref ref-type="bibr" rid="B8">Binetti et al. (2016)</xref> found that the amount of time people can look at each other before starting to feel uncomfortable was 3&#x2013;5 s.</p>
<p>Several studies have modelled gaze aversion behavior in social robots and evaluated their impact. <xref ref-type="bibr" rid="B4">Andrist et al. (2014)</xref> collected gaze data from HHI and used that to model human-like gaze aversions on a NAO robot. They found that well-timed gaze aversions led to better management of the conversational floor and the robot being perceived as more thoughtful. <xref ref-type="bibr" rid="B41">Zhong et al. (2019)</xref> controlled the robot&#x2019;s gaze using a set of heuristics and found that users rated the robot to be more responsive. Subjective evaluation of the gaze system in <xref ref-type="bibr" rid="B22">Lala et al. (2019)</xref> showed that gaze aversions with fillers were preferred when taking turns. On the other hand, there have been a few studies that included gaze aversions as a sub-component of their GCS, but they did not measure any effects of gaze aversion (<xref ref-type="bibr" rid="B29">Mutlu et al., 2012</xref>; <xref ref-type="bibr" rid="B25">Mehlmann et al., 2014</xref>; <xref ref-type="bibr" rid="B40">Zhang et al., 2017</xref>; <xref ref-type="bibr" rid="B32">Pereira et al., 2019</xref>). For example, <xref ref-type="bibr" rid="B25">Mehlmann et al. (2014)</xref> looked at the role of turn-taking gaze behaviors as a whole to evaluate their GCS. However, it is important to note that both <xref ref-type="bibr" rid="B25">Mehlmann et al. (2014)</xref> and <xref ref-type="bibr" rid="B40">Zhang et al. (2017)</xref> used the gaze behavior of participants as feedback to manage the robot&#x2019;s gaze behaviors. <xref ref-type="bibr" rid="B25">Mehlmann et al. (2014)</xref> grounded their architecture on the findings from HHI, whereas <xref ref-type="bibr" rid="B40">Zhang et al. (2017)</xref> relied on findings from human-virtual agent interactions.</p>
<p>Although it has been established that humans perceive robot gaze as similar to human gaze in many cases (<xref ref-type="bibr" rid="B38">Yoshikawa et al., 2006</xref>; <xref ref-type="bibr" rid="B35">Staudte and Crocker, 2009</xref>), it is still important to verify if it holds for different gaze cues and situations in an HRI setting as findings from <xref ref-type="bibr" rid="B3">Admoni et al. (2011)</xref> suggest that robot gaze cues are not reflexively perceived in the same way as human gaze cues. Thus, it is crucial to investigate whether a relationship exists between robot gaze behavior and human gaze behavior, how they are related, and what are the implications of such a relationship. For example, if it is known that lack of gaze aversion by a robot makes people uncomfortable, then we might want to include appropriate gaze aversions when designing a robot for therapeutic intervention. On the other hand, we would probably include fewer gaze aversions when designing an interaction where a robot is training employees to face rude customers. To the best of our knowledge, this is the first work that tries to establish a direct relationship between robot gaze aversion and human gaze aversion behavior.</p>
</sec>
<sec id="s3">
<title>3 Automatic gaze aversion using Gaze control systems</title>
<p>To automate the robot&#x2019;s gaze behavior in this study, we used the GCS proposed in <xref ref-type="bibr" rid="B27">Mishra and Skantze (2022)</xref>. It is a comprehensive GCS that takes into account a wide array of gaze-regulating factors, such as turn-taking, intimacy, and joint attention. The gaze behavior of the robot is planned for a future rolling time window, by giving priorities to different gaze targets (e.g., users, objects, environment), based on various system events related to speaking/listening states and objects being mentioned or moved. At every time step, the GCS makes use of this plan to decide where the robot should be looking and to better coordinate eye&#x2013;head movements.</p>
<p>To model gaze aversion, the model processes the gaze plan at every time step to check if the gaze of the robot is planned to be directed at the user for a duration longer than 3&#x2013;5 s (the preferred mutual gaze duration from HHI (<xref ref-type="bibr" rid="B8">Binetti et al., 2016</xref>)). If that is the case, the model inserts intimacy-regulating gaze aversions into the gaze plan. This results in a quick glance away from the user for about 400 ms using the eye gaze only. Additionally, when the robot&#x2019;s intention is to hold the floor at the beginning of an utterance or at pauses, the GCS also inserts gaze aversions at the appropriate time to model turn-taking and cognitive gaze aversions.</p>
<p>The parameters of the model are either taken from the literature or tuned empirically. This, combined with the novel eye-head coordination, results in a human-like gaze aversion behavior by the robot. In a subjective evaluation of the GCS through a user study, it was found to be preferred over a purely reactive model, and the participants especially found the gaze aversion behavior to be better (<xref ref-type="bibr" rid="B27">Mishra and Skantze, 2022</xref>).</p>
</sec>
<sec id="s4">
<title>4 Hypotheses</title>
<p>
<xref ref-type="bibr" rid="B1">Abele (1986)</xref> found that too much eye gaze directed at an interlocutor would induce discomfort for the speaker and that periodic aversion of gaze would result in a more comfortable interaction. The <italic>Equilibrium Theory</italic> (<xref ref-type="bibr" rid="B6">Argyle and Dean, 1965</xref>) also suggests an inverse relationship between gaze directed at and gaze averted, arguing that increased gaze at an interlocutor would be compensated with more gaze aversions by them. While the theory also discusses other factors such as proxemics, we were interested only in the gaze aspect and in verifying if there is an effect of robot gaze on human gaze behavior. Additionally, it is known that while listening, individuals tend to look more at their speaking interlocutors whereas while speaking, they tend to exhibit more gaze aversions (<xref ref-type="bibr" rid="B5">Argyle and Cook, 1976</xref>; <xref ref-type="bibr" rid="B11">Cook, 1977</xref>; <xref ref-type="bibr" rid="B17">Ho et al., 2015</xref>). Thus, if the robot is not averting its gaze during the interaction, we can expect the participant to produce more gaze aversion while speaking, but not necessarily while listening. Based on this, we formulate the following hypotheses.<list list-type="simple">
<list-item>
<p>&#x2022;<bold>H1</bold>
<italic>Lack of gaze aversions by a robot will lead to an increase in gaze aversions by the participants when they are speaking.</italic>
</p>
<list list-type="simple">
<list-item>
<p>&#x2022;<bold>
<italic>H1a:</italic>
</bold>
<italic>Participants will avert their gaze away from the robot longer in the condition when the robot does not avert its gaze away from the participants. (see</italic>
<xref ref-type="sec" rid="s5">Section 5</xref>
<italic>).</italic>
</p>
</list-item>
<list-item>
<p>&#x2022;<bold>
<italic>H1b:</italic>
</bold>
<italic>Participants will look away from the robot more often when the robot exhibits fixed gaze behavior (does not avert its gaze).</italic>
</p>
</list-item>
</list>
</list-item>
</list>
</p>
</sec>
<sec id="s5">
<title>5 Study design</title>
<p>To investigate the effect of a robot&#x2019;s gaze aversion on human gaze aversion, we designed a within-subjects user study with two conditions. In the control condition, the robot constantly directs its gaze towards the participant, without averting it; we call this the <italic>Fixed Gaze (FG)</italic> condition. In the experimental condition (which we call the <italic>Gaze Aversion (GA)</italic> condition), the robot&#x2019;s gaze is automated using the GCS described in <xref ref-type="sec" rid="s3">Section 3</xref> which is found to be better at exhibiting gaze aversion behavior in a subjective analysis. While the GCS is capable of coordinating individual eye and head movements, the interaction is designed in such a way that it does not require any head movements by the robot when directing its gaze. This is because the interaction involved mainly intimacy-regulating gaze aversions, which necessitate only a quick glance away from the interlocutor (see <xref ref-type="sec" rid="s3">Section 3</xref>). Hence, the robot&#x2019;s head movements are not a factor in the study, which is in line with our aim to verify the effect of robot&#x2019;s eye gaze behavior on human gaze behavior.</p>
<sec id="s5-1">
<title>5.1 Interaction setting</title>
<p>We designed an interview scenario similar to that in <xref ref-type="bibr" rid="B4">Andrist et al. (2014)</xref>, where the robot asked the participant six questions with increasing levels of intimacy (more details in <xref ref-type="sec" rid="s5-2">Subsection 5.2</xref>). While <xref ref-type="bibr" rid="B4">Andrist et al. (2014)</xref> investigated whether appropriate gaze aversions by the robot would elicit more disclosure, we wanted to verify if gaze aversions by a robot would directly elicit lower gaze aversions by humans, signaling more comfort even with highly intimate questions (which are known to induce discomfort). To make the interaction more conversational and less one-sided, the robot also gave an answer to each question after the participant had answered it. Questions with different levels of intimacy were used in order to vary the level to which the participant might feel the need to avert their gaze.</p>
<p>The robot&#x2019;s turns were controlled by the researcher using the Wizard-of-Oz (WoZ) approach. The researcher listened to the participant&#x2019;s responses through a wireless microphone and controlled the robot&#x2019;s response by selecting one of three options, which resulted in varying flows of the conversation script. On selecting &#x201c;Robot answer&#x201d;, the robot would answer the question that was asked to the participant before moving on to ask the next question. The option &#x201c;User declined to answer&#x201d; would prompt the robot to acknowledge the user&#x2019;s choice before moving on to the answer, and then ask the next question. The &#x201c;User asked to repeat question&#x201d; option was used to repeat the question. Having a WoZ paradigm enabled the researcher to control the timing of the robot&#x2019;s turn-taking, resulting in a smooth conversational dynamics. Additionally, it made it possible for the researcher to manage the interaction from a separate room, reducing the influence that the presence of a third-person observer might have on the participants. The robot&#x2019;s responses were handcrafted to be generic enough to account for most of the answers that participants might provide. They always started with an acknowledgement of the participant&#x2019;s answer (e.g., &#x201c;<italic>I appreciate what you say about the weather</italic>&#x201d;). Then a response was chosen at random from previously created pool of handcrafted answers to the question and appended to the acknowledgement. In cases where the participant did not answer the question, the robot always acknowledged that by using phrases like &#x201c;<italic>That&#x2019;s okay</italic>&#x201d; and then appended a random response from the pool of answers.</p>
<p>An example dialog where the participant answered the question is provided below (R denotes the robot, P denotes a participant).<list list-type="simple">
<list-item>
<p>R: <italic>What do you think about the weather today?</italic>
</p>
</list-item>
<list-item>
<p>P: <italic>I think it is perfect. It is neither freezing nor too hot. Just the perfect balance of sunny and cool. I really don&#x2019;t like if it is too hot or too cold.</italic>
</p>
</list-item>
<list-item>
<p>R: <italic>I appreciate what you say about the weather, but honestly, I can&#x2019;t relate. I never get to go outside. Maybe you didn&#x2019;t notice, but I don&#x2019;t have legs. So I never have any idea what the weather is like out in the real world. My dream is to 1 day see the sky. Perhaps my creators will allow me some day.</italic>
</p>
</list-item>
<list-item>
<p>R: <italic>What are your views on pop music?</italic>
</p>
</list-item>
</list>
</p>
<p>We used a Furhat robot for the study, which is a humanoid robot head that projects an animated face onto a translucent mask using back-projection and has a mechanical 3-DoF neck. This makes it possible to generate nuanced gaze behavior using both eye and head movements, as well as facial expressions and accurate lip movements (<xref ref-type="bibr" rid="B28">Moubayed et al., 2013</xref>). For the experimental condition (Gaze Aversion; GA) the robot was named Robert and for the control condition the robot was named Marty. We wanted to give the impression that the participants were interacting with two distinct robots for each condition, but at the same time, we did not want the robots themselves to have an influence on the interaction. This led to the selection of two faces that were similar to each other from the list of characters already available in the robot. Two male voices were selected from the list of available voices based on how natural they sounded when saying the utterances for the tasks. The participants were not informed about the different gaze behaviors of the two robots.</p>
<p>The experiment was conducted in a closed room while restricting any outside distractions. The participants were alone with the robot during the interactions. Participants were asked to sit in a chair that was placed approximately 60&#x2013;90 cm in front of the robot. The robot was carefully positioned such that it was almost at eye level and at a comfortable distance for the participants. A Tobii Pro Glasses 2 eye-tracker was used to record the participants&#x2019; eye gaze during each interaction. We also recorded the speech of the participants using a Zoom H5 multi-track microphone. A pair of Rode Wireless Go microphone systems was also used to stream the audio from the user to the Wizard. <xref ref-type="fig" rid="F1">Figure 1</xref> shows an overview of the experimental setup.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Experimental setup for the interview task.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g001.tif"/>
</fig>
</sec>
<sec id="s5-2">
<title>5.2 Intimacy rating of questions</title>
<p>The questions for the task were selected from <xref ref-type="bibr" rid="B16">Hart et al. (2021)</xref> and <xref ref-type="bibr" rid="B20">Kardas et al. (2021)</xref>, who asked their participants to rate them in terms of sensitivity and intimacy, respectively. In order to account for any influence culture and demography might have on the perceived topic intimacy levels of the questions, an online survey was conducted where residents of Stockholm rated these questions based on their perceived topic intimacy. Participants were recruited using social media forums for Stockholm residents (e.g., Facebook groups, Stockholm SubReddit). Another consideration was to avoid complex questions that would involve a lot of recalling or problem-solving (e.g., &#x201c;What are your views about gun control?&#x201d;). The motivation for this is that people are known to avert their gaze when performing a cognitively challenging task (<xref ref-type="bibr" rid="B13">Doherty-Sneddon and Phelps, 2005</xref>). We wanted to keep the questions as simple as possible so as to restrict the influence on gaze aversions to just the robot&#x2019;s gaze behavior and the question&#x2019;s intimacy level.</p>
<p>A total of 28 questions were selected from the questions in <xref ref-type="bibr" rid="B16">Hart et al. (2021)</xref> and <xref ref-type="bibr" rid="B20">Kardas et al. (2021)</xref>. The participants were asked to rate the questions on how intimate they felt on a 9-point Likert scale ranging from &#x201c;1: Not intimate at all&#x201d; to &#x201c;9: Extremely intimate&#x201d; (question asked: <italic>Please indicate how intimate you find the following questions (1: not intimate at all; 9: extremely intimate). Please don&#x2019;t think too much about each one; just follow your intuition about what you consider personal/intimate</italic>). The responses from 148 participants (68 females, 76 males, one non-binary and two undisclosed), aged between 18 and 50 (mean &#x3d; 29.35, SD &#x3d; 6.89), were then used to order the questions based on their intimacy values. Using linear mixed models, it was verified that gender, age, nationality and L1 did not influence the intimacy ratings. We selected a total of 12 questions out of them and divided them into two sets with similar intimacy distribution which were used evenly across both conditions (<italic>FG</italic> and <italic>GA</italic>). We tried to select simple questions that would not induce a heavy cognitive load. <xref ref-type="table" rid="T1">Table 1</xref> lists the questions and their rated intimacy values from the survey.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Mean intimacy ratings of selected questions used in the study.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Question</th>
<th align="center">Mean</th>
<th align="center">SD</th>
<th align="center">Question set</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">What do you think about the weather today?</td>
<td align="center">1.192</td>
<td align="center">0.558</td>
<td align="center">1</td>
</tr>
<tr>
<td align="left">What are your views on pop music?</td>
<td align="center">1.976</td>
<td align="center">1.372</td>
<td align="center">1</td>
</tr>
<tr>
<td align="left">How did you celebrate last Christmas?</td>
<td align="center">3.023</td>
<td align="center">1.758</td>
<td align="center">1</td>
</tr>
<tr>
<td align="left">Tell me about a conversation you had with another person earlier today</td>
<td align="center">4.330</td>
<td align="center">2.121</td>
<td align="center">1</td>
</tr>
<tr>
<td align="left">For what in your life do you feel most grateful?</td>
<td align="center">5.223</td>
<td align="center">1.917</td>
<td align="center">1</td>
</tr>
<tr>
<td align="left">What is one of the more embarrassing moments in your life?</td>
<td align="center">6.538</td>
<td align="center">2.016</td>
<td align="center">1</td>
</tr>
<tr>
<td align="left">What did you have for breakfast this morning?</td>
<td align="center">1.823</td>
<td align="center">1.308</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">What season do you like the best? Why?</td>
<td align="center">1.838</td>
<td align="center">1.091</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">Do you have anything planned for later today? What will you do?</td>
<td align="center">3.523</td>
<td align="center">1.779</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">What would constitute a perfect day for you?</td>
<td align="center">4.007</td>
<td align="center">1.827</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">Is there something you&#x2019;ve dreamed of doing for a long time? Why haven&#x2019;t you done it?</td>
<td align="center">5.430</td>
<td align="center">2.064</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">Can you describe a time you cried in front of another person?</td>
<td align="center">7.023</td>
<td align="center">1.918</td>
<td align="center">2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5-3">
<title>5.3 Participants</title>
<p>We recorded eye gaze and acoustic data of 33 male participants (sex assigned at birth). The choice for male participants was methodologically and logistically motivated. Firstly, topic intimacy has been found to be perceived differently by people of different genders (<xref ref-type="bibr" rid="B34">Sprague, 1999</xref>). Thus, intimacy during the interaction might be affected by the participants&#x2019; and robot&#x2019;s gender. To reduce the influence of this variable (given that it is not a variable of interest in this study), we controlled it by recruiting participants of only one gender.</p>
<p>The participants were recruited using social media, notice boards and the digital recruitment platform Accindi (<ext-link ext-link-type="uri" xlink:href="https://www.accindi.se/">https://www.accindi.se/</ext-link>). The participants were all residents of Stockholm. The cultural background of participants was not controlled for. Participants&#x2019; ages ranged between 21 and 56 (mean &#x3d; 30.54, SD &#x3d; &#xB1;8.07). They had no hearing or speech impairments, had normal/corrected vision (did not require the use of glasses for face-to-face interactions) and spoke English. They were compensated with a 100SEK gift card on completion of the experiment. The study was approved by the ethics committee of Humboldt-Universit&#xe4;t zu Berlin.</p>
</sec>
<sec id="s5-4">
<title>5.4 Procedure</title>
<p>As described earlier, the study followed a within-subjects design. Each participant interacted with the robot under two conditions; the order of the conditions was randomized. Each set of questions (cf. <xref ref-type="table" rid="T1">Table 1</xref>) was also counterbalanced across the conditions. The participants were asked to give as much information as they could when answering the questions. However, they were not forced to answer any of the questions. In case they did not feel comfortable answering any questions, the robot acknowledged it and moved on to the next question. The interaction always started with the robot introducing itself before moving on to the questions. The entire experiment took approximately 45 min. The experiment&#x2019;s procedure can be broken down into the following steps.<list list-type="simple">
<list-item>
<p>&#x2022;<bold>Step 1:</bold> The participants were informed about the experiment&#x2019;s procedure, compensation, and data protection, both verbally and in writing. They then provided their written consent to participation.</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 2:</bold> The participants were instructed to speak freely about a prompted topic for about 2 min. This recording was used as the baseline speech measure for participants&#x2019; speech before interacting with the robot. The speech data is not discussed in the present work.</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 3:</bold> Next, the participants were asked to put on the eye-tracking glasses, which were then calibrated. After successfully calibrating the glasses, the researcher left the room and initiated the interview task. The robot introduced itself and proceeded with the Q&#x26;A.</p>
</list-item>
<list-item>
<p>The researcher kept track of the participant&#x2019;s responses and timed the robot&#x2019;s turns with the appropriate response using the wizard buttons. Once the interaction came to an end, the researcher returned to the room for the next steps.</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 4:</bold> The participants were then asked to remove the tracking glasses and were provided with a questionnaire to fill in. The questionnaire had 9-point Likert scale questions about the participant&#x2019;s perception of the robot and the flow of conversation (see <xref ref-type="table" rid="T2">Table 2</xref>).</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 5:</bold> Next, they were asked to fill out the Revised NEO Personality Inventory (NEO-PI-R) (<xref ref-type="bibr" rid="B12">Costa Jr and McCrae, 2008</xref>), which measures personality traits. They were also asked to take the LexTALE test (<xref ref-type="bibr" rid="B23">Lemh&#xf6;fer and Broersma, 2012</xref>), which indicates their general level of English proficiency, on an iPad. Both of these tasks served as distractor tasks, providing a break between the two interactions and allowing the participants to focus on the second robot with renewed attention.</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 6:</bold> The participants were then asked to speak freely about another prompted topic for about 2 min. This served as the baseline for the second interaction before the participant interacted with the robot (data not discussed here).</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 7:</bold> After recording the free speech, the participants were asked to put on the eye-tracking glasses and the tracker was calibrated again. The researcher left the room and initiated the next interaction. The robot introduced itself again and proceeded with the Q&#x26;A.</p>
</list-item>
<list-item>
<p>&#x2022;<bold>Step 8:</bold> At the end of the interaction, the researcher returned to the room and provided the participants with the last questionnaire. Apart from the 9-point Likert scale questions about the perception of the robot and the conversation flow, the questionnaire also asked about basic demographic details.</p>
</list-item>
</list>
</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Questionnaire used for subjective evaluation.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Dimension</th>
<th align="left">Question</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="left">Conversation Flow (<bold>D1</bold>)</td>
<td align="left">My conversation with the robot flowed well</td>
</tr>
<tr>
<td align="left">I was able to understand when the robot wanted me to speak</td>
</tr>
<tr>
<td align="left">I was able to understand when robot wanted to keep speaking</td>
</tr>
<tr>
<td align="left">The robot responded to me at the appropriate time</td>
</tr>
<tr>
<td rowspan="4" align="left">Human-Likeness (<bold>D2</bold>)</td>
<td align="left">The robot&#x2019;s face was very human-like</td>
</tr>
<tr>
<td align="left">The robot&#x2019;s voice was very human-like</td>
</tr>
<tr>
<td align="left">The robot&#x2019;s behavior was very human-like</td>
</tr>
<tr>
<td align="left">Throughout the conversation, I was very aware that I was talking to a robot</td>
</tr>
<tr>
<td rowspan="4" align="left">Overall Impression (D3)</td>
<td align="left">I enjoyed talking with the robot</td>
</tr>
<tr>
<td align="left">I felt positively about the robot</td>
</tr>
<tr>
<td align="left">I felt positively about the conversation</td>
</tr>
<tr>
<td align="left">I felt comfortable while talking with the robot</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5-5">
<title>5.5 Measurements</title>
<p>In order to test <bold>H1</bold>, we mainly focused on the behavioral measure of gaze behavior of the participants, which was captured using the eye-tracking glasses. Our experiment had one independent variable, the <italic>gaze aversion</italic> of the robot which was manipulated in a within-subjects design (<italic>GA</italic> &#x26; <italic>FG</italic> condition). The order of the questions remained the same for both the <italic>GA</italic> and <italic>FG</italic> conditions, i.e., increasing intimacy with each subsequent question.</p>
<p>The Tobii Pro Glasses 2 eye-tracker records a video from the point of view of the participant, and provides the 2D gaze points (i.e., where the eyes are directed in the 2D frame of the video). The videos were recorded at 25fps and the eye-tracker sampled the gaze points at a 50 Hz resolution. Both datasets were synchronized to obtain timestamp vs. 2D gaze point ([<italic>ts</italic>, (<italic>x</italic>, <italic>y</italic>)]) data for each recording, i.e., gaze location per timestamp. We used the Haar-cascade algorithm available in the OpenCV library to detect the face of the robot in the videos and obtain the timestamp vs. bounding box of face ([<italic>ts</italic>, (<italic>X</italic>, <italic>Y</italic>, <italic>H</italic>, <italic>W</italic>)], X and Y - lower left corner of the bounding box, H and W - height and width of the bounding box) data.</p>
<p>Gaze Aversion for each time stamp was calculated by verifying if the gaze points (<italic>x</italic>, <italic>y</italic>) were inside the bounding box [<italic>X</italic>, <italic>Y</italic>, <italic>H</italic>, <italic>W</italic>] or not. The parameters for Haar-cascade were manually fine-tuned for each recording to obtain the best fitting bounding boxes for detecting the robot&#x2019;s face. An example of non-gaze aversion detection using the algorithm can be seen in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Example of Gaze Aversion detection using the algorithm. Here the gaze point (<italic>x</italic>, <italic>y</italic>) (the red circle) lies within the face&#x2019;s bounding box [<italic>X</italic>, <italic>Y</italic>, <italic>H</italic>, <italic>W</italic>] (blue rectangle), so it is not a Gaze Aversion.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g002.tif"/>
</fig>
<p>The timing information for the robot&#x2019;s utterances can be obtained from the speech synthesizer. We recorded the robot&#x2019;s responses and their time information for all interactions. This log was used to extract the participant&#x2019;s speaking and listening durations. When the robot is speaking, the participant is the listener and <italic>vice versa</italic>. This information was used to extract the gaze aversion of the participant when they were <italic>Speaking</italic> and <italic>Listening</italic>.</p>
<p>For <bold>
<italic>H1a</italic>
</bold>, we used the % of gaze aversion as the metric of overall gaze aversion. Each timestamp where it was possible to detect whether there was a gaze aversion or not was considered as a <italic>gaze event</italic>. We counted the number of gaze aversions (<italic>gaCount</italic>) and the total number of <italic>gaze events</italic> (<italic>geTotal</italic>) over the duration when the participants were <italic>Speaking</italic> and <italic>Listening</italic>. The % of gaze aversion (<italic>ga%</italic>) is then calculated as:<disp-formula id="e1">
<mml:math id="m1">
<mml:mi>g</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>%</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>T</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>For <bold>
<italic>H1b</italic>
</bold>, we identified individual gaze aversion instances, which are the number of times the participants directed their gaze away from the robot. The duration from when participants looked away from the robot until the time they returned their gaze back at the robot was counted as one gaze aversion instance.</p>
<p>We also collected subjective feedback from the participants for both conditions with a questionnaire. The questionnaire included the 12 questions that were used to measure the responses of the participants under three dimensions on a 9-point Likert scale (see <xref ref-type="table" rid="T2">Table 2</xref>).</p>
<p>The analysis of speech data is beyond the scope of this work and is analyzed in <xref ref-type="bibr" rid="B42">Offrede et al. (2023)</xref>.</p>
</sec>
</sec>
<sec sec-type="results" id="s6">
<title>6 Results</title>
<p>As mentioned in <xref ref-type="sec" rid="s5-1">Section 5.1</xref>, we used a WoZ approach to manage the robot&#x2019;s turns. While the wizard was instructed to behave in the same way for both the conditions, we wanted to make sure that the wizard did not influence the turn taking of the robot, which could in turn influence the gaze aversion behavior of the participants. We calculated the turn gaps (time between when the participant had finished speaking and the robot started to speak) from the audio recordings of the interactions. A Mann-Whitney test indicated that there was no significant difference in turn gaps between condition <italic>FG</italic> (N &#x3d; 249, M &#x3d; 1.89, SD &#x3d; &#xB1;2.67) and condition <italic>GA</italic> (N &#x3d; 243, M &#x3d; 1.77, SD &#x3d; &#xB1;1.63), W &#x3d; 29848, <italic>p</italic> &#x3d; 0.797. This shows that the wizard managed the turns in the same way across conditions.</p>
<sec id="s6-1">
<title>6.1 Effect of Robot&#x2019;s gaze aversion behaviour</title>
<p>Of the 33 participants recorded, we excluded two participants&#x2019; data from the analysis as the gaze data was corrupted due to some technical problems with the eye-tracker. Additionally, gaze data from the eye-trackers were not always available for all timestamps, due to various reasons such as calibration strength and detection efficiency. When averting gaze, participants also moved their head away from their partner&#x2019;s face. This varied a lot from participant to participant and led to instances where the robot&#x2019;s face was out of the eye-tracker&#x2019;s camera frame. Moreover, there were instances where Haar-cascade could not detect the robot&#x2019;s face for some timestamps due to various reasons. These factors resulted in instances where it was not possible to determine if there was a gaze aversion or not. We were able to capture 87.38% of gaze data (data loss &#x3d; 12.62%), which is normal for eye-trackers (<xref ref-type="bibr" rid="B18">Holmqvist, 2017</xref>). Overall, only 1.6% of the data (8170 timestamps out of 504011) was affected by the technical constraints which led to the exclusion of data. Thus, instances of gaze aversion by participants where Furhat was out-of-frame (due to head movement) are not very common. Also, the amount of data lost in this way was the same across conditions so we do not believe that the excluded data had any influence on the results reported.</p>
<p>On average, participants averted their gaze more in the <italic>FG</italic> condition as compared to the <italic>GA</italic> condition when they were <italic>Speaking</italic>. A two-tailed Wilcoxon signed-rank test indicated a significant difference in gaze aversion across conditions when the participants were <italic>Speaking</italic> (<italic>W</italic> &#x3d; 142.0, <italic>p</italic> &#x3d; 0.037), as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. This supported <bold>
<italic>H1a</italic>
</bold>, which predicted that participants would avert their gaze for a longer duration when there is no gaze aversion by the robot (i.e., the <italic>FG</italic> condition). There was no significant difference between conditions when participants were <italic>Listening</italic> (<italic>W</italic> &#x3d; 150.0, <italic>p</italic> &#x3d; 0.194) which is expected (see 4). The mean values of gaze aversion when participants were <italic>Speaking</italic> and <italic>Listening</italic> can be found in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Total % of gaze aversion while participants were <italic>Speaking</italic>.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g003.tif"/>
</fig>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Mean % of gaze aversion per condition.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center"/>
<th colspan="2" align="center">Condition: GA</th>
<th colspan="2" align="center">Condition: FG</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">
<italic>ga%</italic>
</td>
<td align="center">Mean</td>
<td align="center">SD</td>
<td align="center">Mean</td>
<td align="center">SD</td>
</tr>
<tr>
<td align="center">
<italic>Speaking</italic>
</td>
<td align="center">0.399</td>
<td align="center">&#xB1;0.195</td>
<td align="center">0.456</td>
<td align="center">&#xB1;0.199</td>
</tr>
<tr>
<td align="center">
<italic>Listening</italic>
</td>
<td align="center">0.112</td>
<td align="center">&#xB1;0.084</td>
<td align="center">0.137</td>
<td align="center">&#xB1;0.155</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Analyzing the number of gaze aversion instances performed by the participants while <italic>Speaking</italic> showed that participants looked away from the robot more frequently in the <italic>FG</italic> condition (Mean &#x3d; 91.742, SD &#x3d; &#xB1;60.158) as compared to the <italic>GA</italic> condition (Mean &#x3d; 70.774, SD &#x3d; &#xB1;46.141). As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, a two-tailed Wilcoxon signed-rank test indicated a significant difference in the number of gaze aversion instances across conditions (<italic>W</italic> &#x3d; 496.00, <italic>p</italic>
<inline-formula id="inf1">
<mml:math id="m2">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula> 0.001). This supported <bold>
<italic>H1b</italic>
</bold>, which predicted that participants would look away from the robot more often when there is no gaze aversion by the robot. It can be seen that the effect of robot&#x2019;s gaze aversion on participants&#x2019; gaze behavior is stronger and more distinct when analyzing gaze aversion instances. We argue that gaze aversion instance is a better metric to verify the effect.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Number of Gaze Aversion instances while participants were <italic>Speaking</italic>.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g004.tif"/>
</fig>
</sec>
<sec id="s6-2">
<title>6.2 Gaze aversion when participants were Speaking and Listening</title>
<p>It is already known from the HHI literature (<xref ref-type="bibr" rid="B5">Argyle and Cook, 1976</xref>; <xref ref-type="bibr" rid="B11">Cook, 1977</xref>; <xref ref-type="bibr" rid="B17">Ho et al., 2015</xref>) that people exhibit fewer gaze aversions while listening and more while speaking. We were interested to see if there was a similar pattern emerging from the data.</p>
<p>To verify this, we first calculated the % of gaze aversion when participants were listening to and answering each of the robot&#x2019;s questions. Since the durations of both <italic>Speaking</italic> and <italic>Listening</italic> varied from one participant to the other, we normalized the time into 10 intervals for <italic>Speaking</italic> and 10 intervals for <italic>Listening</italic> phase. Next, we found the aggregate % of gaze aversion for all the questions when <italic>Speaking</italic> and <italic>Listening</italic>. The resulting plot can be seen in <xref ref-type="fig" rid="F5">Figure 5</xref>.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>% of Gaze Aversion while participants were <italic>Listening</italic> and <italic>Speaking</italic>, for the two conditions.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g005.tif"/>
</fig>
<p>We can see a clear trend emerging from the plot with the low gaze aversion during the <italic>Listening</italic> phase when the participants listened to the robot. However, just before taking the floor (<italic>Speaking</italic> phase), it can be seen that the gaze aversion starts increasing. This is in line with the findings from <xref ref-type="bibr" rid="B17">Ho et al. (2015)</xref>, who found that speakers usually started their turns with gaze aversion and averted their gaze before taking the turn. We also notice that the gaze aversion peaks at around 20%&#x2013;30% of the speaker&#x2019;s turn, before starting to fall. Towards the end of the turn, we see a sharp decline in gaze aversion. This is consistent with the findings from <xref ref-type="bibr" rid="B17">Ho et al. (2015)</xref>, which show that people end their turns with their gaze directed at the listener. It can also be seen that even though the gaze aversion behavior of participants followed a similar pattern for both conditions, the amount of gaze aversion was lower for the <italic>GA</italic> condition. This further supports hypothesis <bold>H1</bold>.</p>
</sec>
<sec id="s6-3">
<title>6.3 Results from the questionnaire</title>
<p>On analyzing the responses from the questionnaire, all three dimensions were found to have good internal reliability (Cronbach&#x2019;s <italic>&#x3b1;</italic> &#x3d; 0.8, 0.71 &#x26; 0.92 respectively). The participants found the robot under the <italic>FG</italic> condition to be significantly more <italic>Human-Like</italic> (Student&#x2019;s t-test, <italic>p</italic> &#x3d; 0.029). This result was unexpected and is further discussed in <xref ref-type="sec" rid="s7">Section 7</xref>. We did not find any significant differences for the other two dimensions. The mean score for the LexTALE test was 80.515% (SD &#x3d; &#xB1;12.610) which showed that the participants had good English proficiency (<xref ref-type="bibr" rid="B23">Lemh&#xf6;fer and Broersma, 2012</xref>).</p>
</sec>
<sec id="s6-4">
<title>6.4 Exploratory analysis: Topic intimacy</title>
<p>Apart from analyzing the data for <bold>H1</bold>, we were also interested in whether any trends emerged through an exploratory analysis of the topic intimacy of the questions and gaze aversion. By plotting the mean % of gaze aversion values of all participants for each question during the <italic>Speaking</italic> phase, we can see that there is an increase in gaze aversion as the intimacy values increase with the question order (cf. <xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Mean Gaze Aversion per Question while participants were <italic>Speaking</italic>.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g006.tif"/>
</fig>
<p>We fit a GLMM (Generalized Linear Mixed Model) with mean gaze aversion values per question of each participant as the dependent variable (<xref ref-type="bibr" rid="B44">JASP Team, 2023</xref>). The questions&#x2019; order and the conditions were used as the fixed effects variables, and we included random intercepts for participants and random slopes for question order and condition per participant. The model suggested that the <italic>gaze aversions increased as the intimacy values increased</italic> (<italic>&#x3c7;</italic>
<sup>2</sup> &#x3d; 41.32, <italic>df</italic> &#x3d; 5, <italic>p</italic>
<inline-formula id="inf2">
<mml:math id="m3">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula> 0.001). It also suggested that there was more gaze aversion in the <italic>FG</italic> condition as compared to the <italic>GA</italic> condition (<italic>&#x3c7;</italic>
<sup>2</sup> &#x3d; 4.244, <italic>df</italic> &#x3d; 1, <italic>p</italic> &#x3d; 0.039). There were no interaction effects observed. The coefficients of the model can be found in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Fixed effect estimates of the GLMM model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Term</th>
<th align="center">Estimate</th>
<th align="center">SE</th>
<th align="center">t</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Intercept (Question 6)</td>
<td align="center">0.396</td>
<td align="center">0.030</td>
<td align="center">13.138</td>
</tr>
<tr>
<td align="center">Question 1</td>
<td align="center">&#x2212;0.084</td>
<td align="center">0.022</td>
<td align="center">&#x2212;3.846</td>
</tr>
<tr>
<td align="center">Question 2</td>
<td align="center">&#x2212;0.046</td>
<td align="center">0.021</td>
<td align="center">&#x2212;2.224</td>
</tr>
<tr>
<td align="center">Question 3</td>
<td align="center">&#x2212;0.071</td>
<td align="center">0.019</td>
<td align="center">&#x2212;3.731</td>
</tr>
<tr>
<td align="center">Question 4</td>
<td align="center">0.009</td>
<td align="center">0.018</td>
<td align="center">&#x2212;0.515</td>
</tr>
<tr>
<td align="center">Question 5</td>
<td align="center">0.047</td>
<td align="center">0.023</td>
<td align="center">2.085</td>
</tr>
<tr>
<td align="center">Condition:<italic>GA</italic>
</td>
<td align="center">&#x2212;0.033</td>
<td align="center">0.016</td>
<td align="center">&#x2212;2.098</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This points in the direction of a positive correlation between topic intimacy and gaze aversion. One interpretation of this finding is that participants tend to compensate for the discomfort caused by highly intimate questions by averting their gaze. This is in line with previous findings from HHI that suggest that a change in any of the conversational dimensions like proximity, topic intimacy or smiling would be compensated by changing one&#x2019;s behavior in other dimensions (<xref ref-type="bibr" rid="B6">Argyle and Dean, 1965</xref>). <xref ref-type="fig" rid="F7">Figure 7</xref> is a visualization of how gaze aversion varied for each condition under each question. It can be seen that the gaze aversion was higher for <italic>FG</italic> for all the questions (except Q3), and that there is an increase of gaze aversion with the increase in question number (which in turn is the topic intimacy value for the question).</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Distribution of mean gaze aversion per condition per question order.</p>
</caption>
<graphic xlink:href="frobt-10-1127626-g007.tif"/>
</fig>
<p>The finding here is interesting because it could mean that the participants compensated for topic intimacy with gaze aversion even when it is a robot that was asking the questions. However, since we didn&#x2019;t control for the order if the questions, this could also be because of other factors such as a cognitive effort and fatigue. Further studies should narrow down the factors that influenced such behavior.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s7">
<title>7 Discussion</title>
<p>The results suggest that participants averted their gaze significantly more in the <italic>FG</italic> condition. Moreover, they had more gaze aversion instances in the <italic>FG</italic> condition. This was supported by both Wilcoxon signed-rank tests (see <xref ref-type="sec" rid="s6">Section 6</xref>) and an exploratory GLMM (see <xref ref-type="sec" rid="s6-4">Section 6.4</xref>). The results are in line with hypothesis <bold>H1</bold>: people compensate for the lack of robot gaze aversion by producing more gaze aversions themselves (<xref ref-type="bibr" rid="B6">Argyle and Dean, 1965</xref>; <xref ref-type="bibr" rid="B1">Abele, 1986</xref>).</p>
<p>We did not observe a significant difference across conditions in gaze aversion when participants were <italic>Listening</italic>. This could be attributed to the fact that there were too few gaze aversions during this phase to observe a significant difference, which is also suggested by prior studies in HHI (<xref ref-type="bibr" rid="B17">Ho et al., 2015</xref>). The gaze aversions varied between 11%&#x2013;14%, which meant that the participants directed their gaze at the robot for about 86%&#x2013;89% of the time. This is higher than the numbers reported in HHI, where listeners direct their gaze at speakers 30%&#x2013;80% of the time (<xref ref-type="bibr" rid="B21">Kendon, 1967</xref>). Our findings coincide with the findings in <xref ref-type="bibr" rid="B39">Yu et al. (2012)</xref>, where they reported that humans directed their gaze more at a robot than at another human.</p>
<p>Unexpectedly, participants rated the robot in the <italic>FG</italic> condition as more human-like compared to the <italic>GA</italic> condition. A key reason for that could be the way the <italic>GA</italic> interaction started. The GCS used would make the robot keep looking at random places in the environment unless the interaction is started by the researcher. This could have resulted in an unnatural behavior where the robot directs its gaze at random places even though the participant is already sitting in front of it. On the other hand, in the <italic>FG</italic> condition the robot kept on looking straight and only started to track the user when the interaction started. However, since the participant was sitting right in front of the robot, it would be perceived as the robot looking at the participant all the time.</p>
<p>While we did not find any significant differences in the other two dimensions (Conversation Flow &#x26; Overall Impression) assessed in the self-reported questionnaire, we did see a significant difference across conditions from the objective measures (i.e., gaze behavior). This could point to an effect that, even though it might not be explicitly perceived by people, a robot&#x2019;s gaze behavior would implicitly affect human gaze behavior. This could also be an interesting direction for further study.</p>
<p>A further exploratory analysis of the data reveals a positive correlation between gaze aversion and topic intimacy of the questions. Thus, more intimate questions seem to lead to a larger avoidance of eye gaze. In our study, more intimate questions occurred towards the end of the conversation. As the order of the questions was fixed, the order may of course be a confounding factor. However, we are not aware of other work showing that humans would avoid eye gaze more and more over the conversation. We argue that eye gaze is rather related to the topic (intimacy), but further work is needed that controls for this potential confound.</p>
</sec>
<sec id="s8">
<title>8 Limitations and future work</title>
<p>The participants of our study had a rather large age span and we had only male participants. A clear limitation of this study is the lack of a balanced dataset. As the results obtained are only for male participants, these results do not necessarily generalize to other genders. The choice for male participants was methodologically and logistically motivated. Firstly, topic intimacy has been found to be perceived differently by people of different genders (<xref ref-type="bibr" rid="B34">Sprague, 1999</xref>). Thus, intimacy during the interaction might be affected by the participants&#x2019; and robot&#x2019;s gender. To reduce the influence of this variable (given that it is not a variable of interest in this study), we controlled it by recruiting participants of only one gender.</p>
<p>In addition to the participants' gaze behavior, the recorded data were also used to analyze their speech acoustics in relation to that of the robot (<xref ref-type="bibr" rid="B42">Offrede et al., 2023</xref>). Since sex and gender are known to impact acoustic features of speech (<xref ref-type="bibr" rid="B31">P&#xe9;piot, 2014</xref>), all processing and analysis of data need to be carried out separately for males and females. This would reduce the statistical power of the acoustic analysis, leading us to choose participants from only one sex. Given the choice between female or male participants, males were chosen since they are more numerous in the institute where we collected data.</p>
<p>Further studies with a more diverse participant pool and female-presenting robots would be needed to verify this effect in general. However, it is interesting to note that a recent study (<xref ref-type="bibr" rid="B2">Acarturk et al., 2021</xref>) found no difference in gaze aversion behavior due to gender. The authors concluded that GA behavior was independent of gender and suggested &#x201c;that it arises from the social context of the interaction.&#x201d;</p>
<p>It is known that culture also influences our gaze behavior on many levels, such as how we look at faces (<xref ref-type="bibr" rid="B9">Blais et al., 2008</xref>) or interpretation of mutual gaze and gaze aversions (<xref ref-type="bibr" rid="B10">Collett, 1971</xref>; <xref ref-type="bibr" rid="B5">Argyle and Cook, 1976</xref>). <xref ref-type="bibr" rid="B24">McCarthy et al. (2006)</xref> observed that people&#x2019;s mutual gaze and gaze aversion behaviors during thinking differed based on the culture of the individuals. However, recent studies have challenged some of aspects of cultural influences that have been reported previously (<xref ref-type="bibr" rid="B15">Haensel et al., 2022</xref>). Nonetheless, investigating any effect culture of participants may play on their gaze behavior when interacting with a robot could also be an interesting area to look into in the future.</p>
</sec>
<sec sec-type="conclusion" id="s9">
<title>9 Conclusion</title>
<p>In this paper, we investigated whether a robot&#x2019;s gaze behavior can affect human gaze behavior during HRI. We conducted a within-subjects user study and recorded participants&#x2019; gaze data along with participants&#x2019; responses. The analysis of participants&#x2019; eye gaze in both conditions suggests that they tend to avert their gaze more in the absence of gaze aversions by a robot. An exploratory analysis of the data also indicated that more intimate questions may lead to a larger avoidance of mutual gaze. The existence of a direct relationship between robot&#x2019;s gaze behavior and human gaze behavior is an original finding.</p>
<p>The study also shows the importance of modelling gaze aversions in HRI. In the absence of robot gaze aversions, the interaction may become more effortful for the user while trying to avoid frequent mutual gaze with the robot. These findings go hand in hand with the Equilibrium Theory suggesting a trade-off relation between the robot&#x2019;s and user&#x2019;s interactive gaze behavior. Our findings are helpful for designing systems more capable of adapting to the context and situation by taking human gaze behavior into account.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s10">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s11">
<title>Ethics statement</title>
<p>The studies involving human participants were reviewed and approved by the Ethics committee of Humboldt-Universit&#xe4;t zu Berlin. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s12">
<title>Author contributions</title>
<p>All authors designed the eye-tracking experiment. CMi, TO, and GS designed the online study on topic intimacy, CMi and TO wrote the ethical review application and CMi and TO ran the experiment. CMi annotated the gaze data and did the statistical analysis. CMi and TO prepared an initial draft of the paper. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s13">
<title>Funding</title>
<p>This project has received funding from the European Union&#x2019;s Framework Programme for Research and Innovation Horizon 2020 (2014-2020) under the Marie Sk&#x142;odowska-Curie Grant Agreement No. 859588.</p>
</sec>
<sec sec-type="COI-statement" id="s14">
<title>Conflict of interest</title>
<p>GS is a co-founder and CMi is an employee at Furhat Robotics.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s15">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abele</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>Functions of gaze in social interaction: Communication and monitoring</article-title>. <source>J. Nonverbal Behav.</source>
<volume>10</volume>, <fpage>83</fpage>&#x2013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.1007/bf01000006</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Acarturk</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Indurkya</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Nawrocki</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Sniezynski</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Jarosz</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Usal</surname>
<given-names>K. A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Gaze aversion in conversational settings: An investigation based on mock job interview</article-title>. <source>J. Eye Mov. Res.</source>
<volume>14</volume>. <pub-id pub-id-type="doi">10.16910/jemr.14.1.1</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Admoni</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Bank</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Toneva</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Scassellati</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Robot gaze does not reflexively cue human attention</article-title>,&#x201d; in <source>Proceedings of the Annual Meeting of the Cognitive Science Society</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Cognitive Science Society</publisher-name>), <volume>33</volume>, <fpage>1983</fpage>&#x2013;<lpage>1988</lpage>. <comment>Austin, TX</comment>.</citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Andrist</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>X. Z.</given-names>
</name>
<name>
<surname>Gleicher</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mutlu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Conversational gaze aversion for humanlike robots</article-title>,&#x201d; in <source>2014 9th ACM/IEEE International Conference on Human-Robot Interaction (HRI) IEEE</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>25</fpage>&#x2013;<lpage>32</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Argyle</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cook</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cramer</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>1976</year>). <article-title>Gaze and mutual gaze</article-title>. <source>Br. J. Psychiatry</source>
<volume>165</volume>, <fpage>848</fpage>&#x2013;<lpage>850</lpage>. <pub-id pub-id-type="doi">10.1017/s0007125000073980</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Argyle</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Dean</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1965</year>). <article-title>Eye-contact, distance and affiliation</article-title>. <source>Sociometry</source>
<volume>28</volume>, <fpage>289</fpage>&#x2013;<lpage>304</lpage>. <pub-id pub-id-type="doi">10.2307/2786027</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beattie</surname>
<given-names>G. W.</given-names>
</name>
</person-group> (<year>1981</year>). <article-title>A further investigation of the cognitive interference hypothesis of gaze patterns during conversation</article-title>. <source>Br. J. Soc. Psychol.</source>
<volume>20</volume>, <fpage>243</fpage>&#x2013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.1111/j.2044-8309.1981.tb00493.x</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Binetti</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Coutrot</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Johnston</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mareschal</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Pupil dilation as an index of preferred mutual gaze duration</article-title>. <source>R. Soc. Open Sci.</source>
<volume>3</volume>, <fpage>160086</fpage>. <pub-id pub-id-type="doi">10.1098/rsos.160086</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Blais</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Jack</surname>
<given-names>R. E.</given-names>
</name>
<name>
<surname>Scheepers</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Fiset</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Caldara</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Culture shapes how we look at faces</article-title>. <source>PloS one</source>
<volume>3</volume>, <fpage>e3022</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0003022</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Collett</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1971</year>). <article-title>Training englishmen in the non-verbal behaviour of arabs: An experiment on intercultural communication 1</article-title>. <source>Int. J. Psychol.</source>
<volume>6</volume>, <fpage>209</fpage>&#x2013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1080/00207597108246684</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cook</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1977</year>). <article-title>Gaze and mutual gaze in social encounters: &#x201c;How long&#x2014;and when&#x2014;we look others in the eye&#x201d; is one of the main signals in nonverbal communication</article-title>. <source>Am. Sci.</source>
<volume>65</volume>, <fpage>328</fpage>&#x2013;<lpage>333</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Costa</surname>
<given-names>P. T.</given-names>
<suffix>Jr</suffix>
</name>
<name>
<surname>McCrae</surname>
<given-names>R. R.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>The revised neo personality inventory (neo-pi-r)</article-title>. <source>SAGE Handb. Personality Theory Assess.</source>
<volume>2</volume>, <fpage>179</fpage>&#x2013;<lpage>198</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doherty-Sneddon</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Phelps</surname>
<given-names>F. G.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Gaze aversion: A response to cognitive or social difficulty?</article-title>
<source>Mem. Cognition</source>
<volume>33</volume>, <fpage>727</fpage>&#x2013;<lpage>733</lpage>. <pub-id pub-id-type="doi">10.3758/bf03195338</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gillet</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Cumbal</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Pereira</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lopes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Engwall</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Leite</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Robot gaze can mediate participation imbalance in groups with different skill levels</article-title>,&#x201d; in <source>Proceedings of the 2021 ACM/IEEE International Conference on Human-Robot Interaction</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>303</fpage>&#x2013;<lpage>311</lpage>. <pub-id pub-id-type="doi">10.1145/3434073.3444670</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haensel</surname>
<given-names>J. X.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>T. J.</given-names>
</name>
<name>
<surname>Senju</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Cultural differences in mutual gaze during face-to-face interactions: A dual head-mounted eye-tracking study</article-title>. <source>Vis. Cogn.</source>
<volume>30</volume>, <fpage>100</fpage>&#x2013;<lpage>115</lpage>. <pub-id pub-id-type="doi">10.1080/13506285.2021.1928354</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hart</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>VanEpps</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Schweitzer</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>The (better than expected) consequences of asking sensitive questions</article-title>. <source>Organ. Behav. Hum. Decis. Process.</source>
<volume>162</volume>, <fpage>136</fpage>&#x2013;<lpage>154</lpage>. <pub-id pub-id-type="doi">10.1016/j.obhdp.2020.10.014</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ho</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Foulsham</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kingstone</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Speaking and listening with the eyes: Gaze signaling during dyadic interactions</article-title>. <source>PloS one</source>
<volume>10</volume>, <fpage>e0136905</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0136905</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Holmqvist</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Common predictors of accuracy, precision and data loss in 12 eye-trackers</article-title>,&#x201d; in <source>The 7th Scandinavian Workshop on Eye Tracking</source>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Imai</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kanda</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ono</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ishiguro</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mase</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2002</year>). &#x201c;<article-title>Robot mediated round table: Analysis of the effect of robot&#x2019;s gaze</article-title>,&#x201d; in <source>Proceedings. 11th IEEE International Workshop on Robot and Human Interactive Communication</source> (<publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>411</fpage>&#x2013;<lpage>416</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="web">
<collab>JASP Team</collab> (<year>2023</year>). <article-title>JASP (Version 0.17.2) [Computer software]</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://jasp-stats.org/">https://jasp-stats.org/</ext-link>
</comment>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kardas</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Epley</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Overly shallow?: Miscalibrated expectations create a barrier to deeper conversation</article-title>. <source>J. Personality Soc. Psychol.</source>
<volume>122</volume>, <fpage>367</fpage>&#x2013;<lpage>398</lpage>. <pub-id pub-id-type="doi">10.1037/pspa0000281</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kendon</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1967</year>). <article-title>Some functions of gaze-direction in social interaction</article-title>. <source>Acta Psychol.</source>
<volume>26</volume>, <fpage>22</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1016/0001-6918(67)90005-4</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lala</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Inoue</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kawahara</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Smooth turn-taking by a robot using an online continuous model to generate turn-taking cues</article-title>,&#x201d; in <source>Proceedings of the 2019 International Conference on Multimodal Interaction</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>226</fpage>&#x2013;<lpage>234</lpage>. <pub-id pub-id-type="doi">10.1145/3340555.3353727</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lemh&#xf6;fer</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Broersma</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Introducing lextale: A quick and valid lexical test for advanced learners of English</article-title>. <source>Behav. Res. Methods</source>
<volume>44</volume>, <fpage>325</fpage>&#x2013;<lpage>343</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-011-0146-0</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McCarthy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Itakura</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Muir</surname>
<given-names>D. W.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Cultural display rules drive eye gaze during thinking</article-title>. <source>J. Cross-Cultural Psychol.</source>
<volume>37</volume>, <fpage>717</fpage>&#x2013;<lpage>722</lpage>. <pub-id pub-id-type="doi">10.1177/0022022106292079</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mehlmann</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>H&#xe4;ring</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Janowski</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Baur</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gebhard</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Andr&#xe9;</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Exploring a model of gaze for grounding in multimodal hri</article-title>,&#x201d; in <source>Proceedings of the 16th International Conference on Multimodal Interaction</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>247</fpage>&#x2013;<lpage>254</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Metta</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Natale</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nori</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sandini</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Vernon</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Fadiga</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>The icub humanoid robot: An open-systems platform for research in cognitive development</article-title>. <source>Neural Netw.</source>
<volume>23</volume>, <fpage>1125</fpage>&#x2013;<lpage>1134</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2010.08.010</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mishra</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Skantze</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Knowing where to look: A planning-based architecture to automate the gaze behavior of social robots</article-title>,&#x201d; in <source>2022 31st IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)</source> (<publisher-loc>Naples, Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1201</fpage>&#x2013;<lpage>1208</lpage>. <pub-id pub-id-type="doi">10.1109/RO-MAN53752.2022.9900740</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moubayed</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Skantze</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Beskow</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>The furhat back-projected humanoid head&#x2013;lip reading, gaze and multi-party interaction</article-title>. <source>Int. J. Humanoid Robotics</source>
<volume>10</volume>, <fpage>1350005</fpage>. <pub-id pub-id-type="doi">10.1142/s0219843613500059</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mutlu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Kanda</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Forlizzi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hodgins</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ishiguro</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Conversational gaze mechanisms for humanlike robots</article-title>. <source>ACM Trans. Interact. Intelligent Syst. (TiiS)</source>
<volume>1</volume>, <fpage>1</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1145/2070719.2070725</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nakano</surname>
<given-names>Y. I.</given-names>
</name>
<name>
<surname>Yoshino</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yatsushiro</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Takase</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Generating robot gaze on the basis of participation roles and dominance estimation in multiparty interaction</article-title>. <source>ACM Trans. Interact. Intelligent Syst. (TiiS)</source>
<volume>5</volume>, <fpage>1</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1145/2743028</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Offrede</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Mishra</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Skantze</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Fuchs</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mooshammer</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2023</year>). &#x201C;<article-title>Do humans converge phonetically when talking to a robot?</article-title>&#x201d; in <source>Proceedings of the 20th International Congress of Phonetic Sciences (ICPhS)</source>, <conf-loc>Prague, Czech Republic</conf-loc>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>P&#xe9;piot</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Male and female speech: A study of mean f0, f0 range, phonation type and speech rate in parisian French and American English speakers</article-title>,&#x201d; in <source>Speech prosody 7</source> (<publisher-loc>Dublin, Ireland</publisher-loc>: <publisher-name>HAL CCSD</publisher-name>), <fpage>305</fpage>&#x2013;<lpage>309</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pereira</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Oertel</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Fermoselle</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Mendelson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gustafson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Responsive joint attention in human-robot interaction</article-title>,&#x201d; in <source>Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) IEEE</source> (<publisher-loc>Macau, China</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1080</fpage>&#x2013;<lpage>1087</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schellen</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bossi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wykowska</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Robot gaze behavior affects honesty in human-robot interaction</article-title>. <source>Front. Artif. Intell.</source>
<volume>4</volume>, <fpage>663190</fpage>. <pub-id pub-id-type="doi">10.3389/frai.2021.663190</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Skantze</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Predicting and regulating participation equality in human-robot conversations: effects of age and gender</article-title>&#x201d; in <source>2017 ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source>, <conf-loc>New York, USA</conf-loc>, <fpage>196</fpage>&#x2013;<lpage>204</lpage>. <pub-id pub-id-type="doi">10.3389/frai.2021.663190</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sprague</surname>
<given-names>R. J.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>The relationship of gender and topic intimacy to decisions to seek advice from parents</article-title>. <source>Commun. Res. Rep.</source>
<volume>16</volume>, <fpage>276</fpage>&#x2013;<lpage>285</lpage>. <pub-id pub-id-type="doi">10.1080/08824099909388727</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Staudte</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Crocker</surname>
<given-names>M. W.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Visual attention in spoken human-robot interaction</article-title>,&#x201d; in <source>2009 4th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source> (<publisher-loc>La Jolla, CA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>77</fpage>&#x2013;<lpage>84</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tomasello</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1995</year>). &#x201c;<article-title>Joint attention as social cognition</article-title>,&#x201d; in <source>Joint attention: Its origins and role in development</source>, <fpage>103</fpage>&#x2013;<lpage>130</lpage>.</citation>
</ref>
<ref id="B37">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yamazaki</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Yamazaki</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kuno</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Burdelski</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kawashima</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kuzuoka</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Precision timing in human-robot interaction: Coordination of head movement and utterance</article-title>,&#x201d; in <source>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>131</fpage>&#x2013;<lpage>140</lpage>. <pub-id pub-id-type="doi">10.1145/1357054.1357077</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yoshikawa</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Shinozawa</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ishiguro</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hagita</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Miyamoto</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>Responsive robot gaze to interaction partner</article-title>,&#x201d; in <source>Robotics: Science and systems</source> (<publisher-loc>Philadelphia, USA</publisher-loc>: <publisher-name>Robotics: Science and systems</publisher-name>), <fpage>37</fpage>&#x2013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.15607/RSS.2006.II.037</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Schermerhorn</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Scheutz</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Adaptive eye gaze patterns in interactions with human and artificial agents</article-title>. <source>ACM Trans. Interact. Intelligent Syst. (TiiS)</source>
<volume>1</volume>, <fpage>1</fpage>&#x2013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1145/2070719.2070726</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Beskow</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kjellstr&#xf6;m</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Look but don&#x2019;t stare: Mutual gaze interaction in social robots</article-title>,&#x201d; in <source>Proceedings of the 9th International Conference on Social Robotics</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <volume>10652</volume>, <fpage>556</fpage>&#x2013;<lpage>566</lpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname>
<given-names>V. J.</given-names>
</name>
<name>
<surname>Schmiedel</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Dornberger</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Investigating the effects of gaze behavior on the perceived delay of a robot&#x2019;s response</article-title>,&#x201d; in <source>Proceedings of the 11th International Conference on Social Robotics</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>54</fpage>&#x2013;<lpage>63</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>