<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1356477</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2024.1356477</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Perspective</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Building for speech: designing the next-generation of social robots for audio interaction</article-title>
<alt-title alt-title-type="left-running-head">Addlesee and Papaioannou</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frobt.2024.1356477">10.3389/frobt.2024.1356477</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Addlesee</surname>
<given-names>Angus</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2431142/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/Writing - review &#x26; editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Papaioannou</surname>
<given-names>Ioannis</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/Writing - review &#x26; editing/"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Mathematics and Computer Science</institution>, <institution>Heriot-Watt University</institution>, <addr-line>Edinburgh</addr-line>, <country>United Kingdom</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Aveni AI</institution>, <addr-line>Edinburgh</addr-line>, <country>United Kingdom</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/971894/overview">Frank Foerster</ext-link>, University of Hertfordshire, United Kingdom</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/681827/overview">Dimosthenis Kontogiorgos</ext-link>, Royal Institute of Technology, Sweden</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Angus Addlesee&#x2009;, <email>a.addlesee@hw.ac.uk</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>03</day>
<month>01</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>11</volume>
<elocation-id>1356477</elocation-id>
<history>
<date date-type="received">
<day>15</day>
<month>12</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>11</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Addlesee and Papaioannou.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Addlesee and Papaioannou</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>There have been significant advances in robotics, conversational AI, and spoken dialogue systems (SDSs) over the past few years, but we still do not find social robots in public spaces such as train stations, shopping malls, or hospital waiting rooms. In this paper, we argue that early-stage collaboration between robot designers and SDS researchers is crucial for creating social robots that can legitimately be used in real-world environments. We draw from our experiences running experiments with social robots, and the surrounding literature, to highlight recurring issues. Robots need better speakers, a greater number of high-quality microphones, quieter motors, and quieter fans to enable human-robot spoken interaction in the wild. If a robot was designed to meet these requirements, researchers could create SDSs that are more accessible, and able to handle multi-party conversations in populated environments. Robust robot joints are also needed to limit potential harm to older adults and other more vulnerable groups. We suggest practical steps towards future real-world deployments of conversational AI systems for human-robot interaction.</p>
</abstract>
<kwd-group>
<kwd>social robots</kwd>
<kwd>spoken dialogue</kwd>
<kwd>accessibility</kwd>
<kwd>robotics</kwd>
<kwd>human-robot interaction</kwd>
<kwd>conversational AI</kwd>
</kwd-group>
<contract-num rid="cn001">871245</contract-num>
<contract-sponsor id="cn001">European Commission<named-content content-type="fundref-id">10.13039/501100000780</named-content>
</contract-sponsor>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Human-Robot Interaction</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Social robots are not yet found in our public spaces, despite this vision being an imminent reality over 25 years ago (<xref ref-type="bibr" rid="B38">Thrun, 1998</xref>). They do not roam our shopping malls helping lost families find the bathroom, we do not bump into them providing departure times in train stations and airports, and they are not helping patients in hospital waiting rooms with their questions (see <xref ref-type="fig" rid="F1">Figure 1</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>A person asking for directions in a hospital.</p>
</caption>
<graphic xlink:href="frobt-11-1356477-g001.tif"/>
</fig>
<p>Spoken dialogue systems (SDSs) have consistently improved over time (<xref ref-type="bibr" rid="B19">Glass, 1999</xref>; <xref ref-type="bibr" rid="B46">Williams, 2009</xref>; <xref ref-type="bibr" rid="B26">Lemon, 2022</xref>), with many years of peer-reviewed papers containing remarkable results. These models, however, are often evaluated automatically upon collected data, or with users in highly controlled lab settings. Social robots work wonderfully in the lab, but fail when deployed in the real world for experiments or demonstration (<xref ref-type="bibr" rid="B39">Tian and Oviatt, 2021</xref>). Some of these failures stem from the embedded SDS [for example: multi-party interactions (<xref ref-type="bibr" rid="B7">Addlesee et al., 2023a</xref>), socio-affective competence (<xref ref-type="bibr" rid="B39">Tian and Oviatt, 2021</xref>), voice accessibility (<xref ref-type="bibr" rid="B2">Addlesee, 2023</xref>), or trust failures (<xref ref-type="bibr" rid="B40">Tolmeijer et al., 2020</xref>)], but even interactions that the SDS should be able to handle with ease go wrong. These failures are often caused by the design of the social robot itself, classified by <xref ref-type="bibr" rid="B22">Honig and Oron-Gilad (2018)</xref> as technical hardware failures.</p>
<p>The field of robotics has also seen incredible advancements over the past years, today&#x2019;s robots can navigate obstacle courses (<xref ref-type="bibr" rid="B48">Xiao et al., 2022</xref>), manipulate objects in their environment (<xref ref-type="bibr" rid="B15">Chai et al., 2022</xref>), generate human-like gestures (<xref ref-type="bibr" rid="B37">Tatarian et al., 2022</xref>), and follow complex human instructions using large language models (LLMs) (<xref ref-type="bibr" rid="B10">Ahn et al., 2022</xref>). Sadly, while collaboration between these two fields is common, they often begin after the robot has been designed. In our experience from multiple international robot dialogue projects, spoken interaction is not considered during the initial design phase. This lack of early-stage collaboration between robot designers and SDS researchers leaves room for oversight of critical features for spoken interaction, contributing to both performance and social errors (<xref ref-type="bibr" rid="B39">Tian and Oviatt, 2021</xref>). In this paper, we have combed the literature and drawn from our own experiences to highlight underlying and fundamental hardware problems that repeatedly surface when experimenting with social robots and real users. We hope this paper sparks discussion between both communities to create fully functioning social robots that do genuinely work in public spaces in the future, enabling live in-the-wild experiments.</p>
</sec>
<sec id="s2">
<title>2 People struggle to hear robots</title>
<p>The first issue that crops up commonly in the literature is the limited volume of the robot&#x2019;s voice. Robot designers simply attach a speaker to the robot without considering the fact that the world is noisy, and some users, such as older adults, may have hearing loss.</p>
<p>In recent work, researchers deployed a robot to interact with real users in an assisted living facility. The robot had to be fitted with an additional speaker that had a louder maximum volume. The modification was necessary because the users simply could not hear the robot&#x2019;s voice, preventing any basic interaction (<xref ref-type="bibr" rid="B36">Stegner et al., 2023</xref>).</p>
<p>This issue is not constrained to this particular setting, or to one particular robot. For example, researchers had to repeat every sentence the robot said in the lobby of a concert hall, as participants could not hear it <xref ref-type="bibr" rid="B25">Langedijk et al. (2020)</xref>. In various school environments, the robot&#x2019;s volume was not loud enough to enable effective interaction, so external speakers had to be fitted (<xref ref-type="bibr" rid="B29">Nikolopoulos et al., 2011</xref>). When guiding people in an elder care facility, the robot&#x2019;s single speaker faced the wrong direction, so users could not hear it <xref ref-type="bibr" rid="B25">Langedijk et al. (2020)</xref>. Another robot was deployed in the homes of a few older adults, and they noted that its volume was not loud enough. People could not hear a social robot in a gym (<xref ref-type="bibr" rid="B31">Sackl et al., 2022</xref>), and the list goes on.</p>
<p>Robots are expensive, but the speakers that researchers had to retrofit to the robots were inexpensive and readily available. This low-cost change was simple, yet <italic>crucial</italic>, to enable effective communication with a user in a real-world setting. When designing robots for spoken interaction, we recommend fitting multiple speakers (facing various directions) that have a loud maximum volume. This will guarantee that the robot can be heard in public spaces, and ensure its accessibility for people with limited hearing. In the future, parametric array loudspeakers (PALs) could be installed to use ultrasonic transducers (<xref ref-type="bibr" rid="B49">Yang et al., 2005</xref>; <xref ref-type="bibr" rid="B13">Bhorge et al., 2023</xref>). PALs are unidirectional, using the nonlinear interactions between soundwaves to enable directed personal communication to a specific user in a populated environment (<xref ref-type="bibr" rid="B50">Zhu et al., 2023</xref>).</p>
</sec>
<sec id="s3">
<title>3 Robots struggle to hear people</title>
<p>There is another conversation participant that cannot properly hear what their interlocutor is saying&#x2013;the robot. This problem is similar to the one in <xref ref-type="sec" rid="s2">Section 2</xref>, and is also frequently found in the literature. A social robot struggled to hear users in a hotel lobby, for example (<xref ref-type="bibr" rid="B21">Hahkio, 2020</xref>). Many researchers retrofit better microphones to the robot (<xref ref-type="bibr" rid="B42">Villalpando et al., 2018</xref>), or next to the robot (<xref ref-type="bibr" rid="B43">Wagner et al., 2023</xref>), in order to hear the user more clearly.</p>
<p>In an assisted living facility, researchers had to resort to listening to the user through an ajar door to run their experiments. The microphone array could not reliably pick up what users said (<xref ref-type="bibr" rid="B36">Stegner et al., 2023</xref>).</p>
<p>Home voice assistants do successfully hear people in noisy environments, like family homes (<xref ref-type="bibr" rid="B30">Porcheron et al., 2018</xref>), however. They can pick up what the user said when other conversations are happening in the room, and when the TV or radio are on [we are also finding this in ongoing work (<xref ref-type="bibr" rid="B1">Addlesee, 2022</xref>)]. Today&#x2019;s social robots typically have four microphones<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref>, but we argue that this is far too few. Apple&#x2019;s Homepod originally had six microphones (<xref ref-type="bibr" rid="B14">Calore, 2019</xref>), and Amazon&#x2019;s Alexa Echo had seven (<xref ref-type="bibr" rid="B35">Spekking, 2021</xref>). The newest Homepod and Echo have reduced to four microphones for two reasons: (1) These devices are incentivised to keep their device&#x2019;s costs low to encourage adoption by new users (<xref ref-type="bibr" rid="B45">Welch, 2023</xref>); and (2) The device&#x2019;s shape and internal component arrangements have been refined and optimised over many years through experiments with millions of users (<xref ref-type="bibr" rid="B47">Wilson, 2020</xref>). Robots do not share either of these features. Microphones are trivially inexpensive relative to the price of a robot, and instead of helping microphones, the robot&#x2019;s shape actively hinders their performance. The body parts of a social robot often sit between the user and the microphone (for example, when the user is behind the robot, or in a wheelchair). Robots also create a lot of noise themselves, called <italic>ego-noise</italic>. Related research required high-quality audio input from a noisy propellered UAV, so they attached sixteen microphones in various locations around the device (<xref ref-type="bibr" rid="B28">Nakadai et al., 2017</xref>), not just four.</p>
<p>Human-robot spoken communication can also be disrupted by societal or linguistic phenomena, such as overlapping or poorly formed turn-taking conditions (<xref ref-type="bibr" rid="B34">Skantze, 2021</xref>). Such conditions include barging-in (<xref ref-type="bibr" rid="B44">Wagner et al., 2021</xref>) (where the user interrupts the robot mid-sentence, but the robot fails to recognise that the user started speaking), and poor end-of-turn detection (due to long pauses or intermittent speech from the user). In our experience, users sometimes barge-in because of high latency caused by limited computational power onboard the robot, or on-site connectivity issues, in addition to the SDS latency.</p>
<p>Potential approaches to this challenge include incremental dialogue processing (<xref ref-type="bibr" rid="B8">Addlesee et al., 2020</xref>; <xref ref-type="bibr" rid="B12">Aylett et al., 2023</xref>), predictive turn-taking (<xref ref-type="bibr" rid="B24">Inoue et al., 2024</xref>), or explicit turn-taking signals which enable the user to better understand when the robot is actually listening to them. For instance, <xref ref-type="bibr" rid="B18">Foster et al. (2019)</xref> employed a tablet on the robot&#x2019;s torso that was showing <italic>&#x201c;I am listening&#x201d;</italic> and <italic>&#x201c;I am speaking&#x201d;</italic> text to help guide the users in a noisy shopping mall setup.</p>
<p>We therefore recommend fitting multiple high-quality microphones in various locations around the robot&#x2019;s body, as well as using appropriate signal processing techniques, such as beam-forming (<xref ref-type="bibr" rid="B9">Adel et al., 2012</xref>). Latency issues must be addressed within the SDS, and by increasing the robot&#x2019;s computational capabilities. These changes will again ensure that multi-party spoken interactions can realistically take place in public spaces. Robot designers must also consider microphone placement lower down on the robot for shorter users, and for people in wheelchairs, as they are commonly just placed on the top of the robot&#x2019;s head.</p>
<sec id="s3-1">
<title>3.1 Multi-party interaction</title>
<p>All of the above challenges assume that the interaction is dyadic&#x2013;that is, one person conversing with a single system/robot. Conversational AI systems and SDSs are typically designed for this setting, including commercial assistants like Alexa and Siri. However, dyadic interactions can only be guaranteed in specific environments, such as single-occupant homes (and even then, there may be visitors). In the public spaces that social robots are expected to roam in the future (see <xref ref-type="sec" rid="s1">Section 1</xref>), groups of people may approach the robot (<xref ref-type="bibr" rid="B11">Alameda-Pineda et al., 2024</xref>). In multi-party conversations (MPCs), the SDS must track who said an utterance, who the user was addressing, and then generate a suitable response, depending on whether the robot is addressing an individual or the whole group (<xref ref-type="bibr" rid="B41">Traum, 2004</xref>). The robot may also need to decide to remain silent, for example if people are talking to each other, but still monitor the content of their conversation in case it can assist them. Additionally, MPCs introduce unique challenges such as multi-party goal-tracking (<xref ref-type="bibr" rid="B6">Addlesee et al., 2023b</xref>). Groups may have conflicting goals, or share goals (<xref ref-type="bibr" rid="B17">Eshghi and Healey, 2016</xref>). Current social robots are not designed to enable MPCs, since speaker diarization (tracking &#x2018;who said what&#x2019;) is critical (<xref ref-type="bibr" rid="B5">Addlesee et al., 2023c</xref>; <xref ref-type="bibr" rid="B32">Schauer et al., 2023</xref>). The audio from the robot&#x2019;s microphones must not only be clear enough to perform ASR accurately, but clear enough to determine <italic>who</italic> said an utterance (<xref ref-type="bibr" rid="B16">Cooper et al., 2023</xref>). Ideally, the microphones would also provide the angle which the audio originated from. This angle can be combined with the robot&#x2019;s vision to determine which person in view said an utterance. The robot can then look at the user it is addressing when responding.</p>
<p>We recommend that social robots be designed with multi-party interaction in mind&#x2013;this means designing microphone arrays such that speaker diarization is accurate, combining this with person-tracking, and developing NLP systems that can understand and manage multi-party conversations (<xref ref-type="bibr" rid="B26">Lemon, 2022</xref>).</p>
</sec>
</sec>
<sec id="s4">
<title>4 Ego-noise</title>
<p>This issue of ego-noise, introduced in <xref ref-type="sec" rid="s3">Section 3</xref>, is so problematic that an entire field of research has grown to tackle it. Researchers find that ego-noise, noise generated by the robot itself, does not just negatively impact ASR performance, but that ego-noise reduction methods also suppress some of the user&#x2019;s utterance (<xref ref-type="bibr" rid="B23">Ince et al., 2010</xref>; <xref ref-type="bibr" rid="B33">Schmidt et al., 2018</xref>). To clarify, both the ego-noise reduction techniques, and the ego-noise itself negatively impact ASR performance (<xref ref-type="bibr" rid="B11">Alameda-Pineda et al., 2024</xref>).</p>
<p>This issue would be helped by additional speakers and microphones, ideally not placed next to noise sources, allowing both parties to hear each other. An optimal social robot designed for spoken interaction would also have much quieter joint motors and fans. These are more expensive than speakers and microphones, but they would greatly improve the SDSs ability to understand the user. This could be paired with research to repair and understand disrupted sentences (<xref ref-type="bibr" rid="B3">Addlesee and Damonte, 2023a</xref>; <xref ref-type="bibr" rid="B4">Addlesee and Damonte, 2023b</xref>), while quieter motors are developed.</p>
<p>In addition to joint motors and fans, the robot&#x2019;s own voice is another source of ego-noise. Microphones cannot simply be turned off when the robot is talking, as speech can be overlapping, so recognised speech may have to be classified as being produced by itself or another (<xref ref-type="bibr" rid="B27">Lemon and Gruenstein, 2004</xref>).</p>
</sec>
<sec id="s5">
<title>5 Joint robustness</title>
<p>Ego noise obviously does not impact robots that do not have a body. In our view, though, social robots should be able to point to location and objects, guide users, and help users physically. For example, consider a hospital waiting room in a hospital memory clinic (<xref ref-type="bibr" rid="B20">Gunson et al., 2022</xref>). Patients are typically older adults, and may use the robot&#x2019;s arm for stability, like they would with another human (see <xref ref-type="fig" rid="F2">Figure 2</xref>). Current social robots can generate social gestures like waving or holding its hand out for a handshake. If you were to shake the robot&#x2019;s hand, however, it would likely break. Such fragility could potentially harm users if deployed in this setting. People may assume that they can link arms with the robot while being guided, a perfectly natural assumption. When an older adult puts their weight on the robot&#x2019;s joint, though, they might fall. This is clearly a potentially harmful design flaw that must be resolved if we are ever going to find robot assistants in the wild.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>A social robot providing stability to an older adult.</p>
</caption>
<graphic xlink:href="frobt-11-1356477-g002.tif"/>
</fig>
</sec>
<sec sec-type="conclusion" id="s6">
<title>6 Conclusion</title>
<p>Interacting naturally with social robots in public spaces is currently still a sci-fi fantasy. There are challenges that SDS researchers must tackle to reach this goal, but that is not the only bottleneck. Even a perfect SDS would fail if it was embedded within today&#x2019;s social robots. We have highlighted that robots need louder speakers (or parametric array loudspeakers in the future), a greater number of high-quality microphones, quieter fans, and quieter motors to allow both parties to hear each other. These are critical problems that completely block spoken interactions outside a lab setting. We highlighted that robots also need to be more physically robust if they are to be safely applied in the real world, particularly in settings with older adults.</p>
<p>Social robotics research will continue to rely on offline evaluations, wizard-of-oz deployments, or lab-based experiments if these robot hardware issues are not resolved. Our suggestions are not an exhaustive list, but we hope that they spark discussion and encourage collaboration between robot designers and SDS researchers. This collaboration should take place in the initial stages of a robot&#x2019;s design to avoid the retrofitting of hardware and sensors discussed in this paper, and instead enable real in-the-wild experiments.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>AA: Conceptualization, Investigation, Writing&#x2013;original draft, Writing&#x2013;review and editing. IP: Writing&#x2013;original draft, Writing&#x2013;review and editing.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. This work was partially supported by the European Commission under the Horizon 2020 framework programme for Research and Innovation (H2020-ICT-2019-2, GA no. 871245): SPRING project <ext-link ext-link-type="uri" xlink:href="https://spring-h2020.eu/">https://spring-h2020.eu/</ext-link>.</p>
</sec>
<ack>
<p>The images in this paper were generated with Hotpot.io.</p>
</ack>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of interest</title>
<p>Author IP was employed by company Aveni AI. The remaining author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
<p>The handling editor FF and reviewer DK declared a past co-authorship with the author IP.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>We checked the technical specifications of several commonly used social robots, and robots that we have deployed ourselves. We are refraining from naming specific robot creators, as this paper aims to encourage collaboration, and not criticise specific robots.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Securely cpeople&#x2019;s interactions with voice assistants at home: a bespoke tool for ethical data collection </article-title>,&#x201d; in <source>Proceedings of the second workshop on NLP for positive impact (NLP4PI)</source>, <fpage>25</fpage>&#x2013;<lpage>30</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2023</year>). &#x201c;<article-title>Voice assistant accessibility</article-title>,&#x201d; in <source>The international workshop on spoken dialogue systems Technology, IWSDS 2023</source>, <fpage>11</fpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Damonte</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2023a</year>). <source>Understanding disrupted sentences using underspecified abstract meaning representation</source>. <publisher-loc>Dublin, Ireland</publisher-loc>: <publisher-name>Interspeech</publisher-name>, <fpage>5</fpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Damonte</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2023b</year>). <article-title>Understanding and answering incomplete questions</article-title> in <conf-name>Proceedings of the 5th conference on conversational user interfaces</conf-name>, <volume>9</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1145/3571884.3597133</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Denley</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Edmondson</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gunson</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Garcia</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Kha</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2023c</year>). &#x201c;<article-title>Detecting agreement in multi-party dialogue: evaluating speaker diarisation versus a procedural baseline to enhance user engagement</article-title>,&#x201d; in <source>Proceedings of the workshop on advancing GROup UNderstanding and robots aDaptive behaviour (GROUND)</source>, <fpage>7</fpage>.</citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Siei&#x144;ska</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Gunson</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Garcia</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Dondrup</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lemon</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2023b</year>). <article-title>Multi-party goal tracking with LLMs: comparing pre-training, fine-tuning, and prompt engineering</article-title> in <source>Proceedings of the 24th annual meeting of the special interest group on discourse and dialogue</source>, <fpage>13</fpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Siei&#x144;ska</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Gunson</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Hern&#xe1;ndez Garc&#xed;a</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Dondrup</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lemon</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2023a</year>). <article-title>Data collection for multi-party task-based dialogue in social robotics</article-title>. <source>Int. Workshop Spok. Dialogue Syst. Technol. IWSDS</source> <volume>2023</volume>, <fpage>10</fpage>. <pub-id pub-id-type="doi">10.21437/interspeech.2023-307</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Eshghi</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A comprehensive evaluation of incremental speech recognition and diarization for conversational AI</article-title> in <conf-name>Proceedings of the 28th international conference on computational linguistics</conf-name>, <fpage>3492</fpage>&#x2013;<lpage>3503</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adel</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Souad</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Alaqeeli</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hamid</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Beamforming techniques for multichannel audio signal separation</article-title> <volume>6</volume>, <fpage>659</fpage>, <lpage>667</lpage>. <pub-id pub-id-type="doi">10.4156/jdcta.vol6.issue20.72</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ahn</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Brohan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Brown</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Chebotar</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cortes</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>David</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Do as i can, not as i say: grounding language in robotic affordances</article-title>. <comment>arXiv Prepr. arXiv:2204.01691</comment>. <pub-id pub-id-type="doi">10.48550/arXiv.2204.01691</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alameda-Pineda</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Garc&#xed;a</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Reinke</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Arias</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Arrigoni</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2024</year>). <article-title>Socially pertinent robots in gerontological healthcare</article-title>. <comment>arXiv preprint arXiv:2404.07560</comment>.</citation>
</ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Aylett</surname>
<given-names>M. P.</given-names>
</name>
<name>
<surname>Carmantini</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Braude</surname>
<given-names>D. A.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Why is my social robot so slow? How a conversational listener can revolutionize turn-taking</article-title> in <conf-name>Proceedings of the 5th conference on conversational user interfaces</conf-name>, <fpage>4</fpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bhorge</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Patil</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Pawar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Poke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Jambhulkar</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Unidirectional parametric speaker</article-title>. <conf-name>ITM Web Conf. ITM Web of Conferences</conf-name> <publisher-loc>Les Ulis, France</publisher-loc>: <publisher-name>EDP Sciences</publisher-name>, <volume>56</volume>, <fpage>04010</fpage>, <pub-id pub-id-type="doi">10.1051/itmconf/20235604010</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Calore</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Review: apple Homepod</article-title>. <source>Wired</source>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chai</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>A survey of the development of quadruped robots: joint configuration, dynamic locomotion control method and mobile manipulation approach</article-title>. <source>Biomim. Intell. Robotics</source> <volume>2</volume>, <fpage>100029</fpage>. <pub-id pub-id-type="doi">10.1016/j.birob.2021.100029</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cooper</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ros</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Lemaignan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Robotics</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Challenges of deploying assistive robots in real-life scenarios: an industrial perspective</article-title> in <conf-name>The 32nd IEEE international conference on robot and human interactive communication</conf-name>. <publisher-loc>Naples, Italy</publisher-loc>: <publisher-name>RO-MAN</publisher-name>, <fpage>10</fpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eshghi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Healey</surname>
<given-names>P. G.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Collective contexts in conversation: grounding by proxy</article-title>. <source>Cognitive Sci.</source> <volume>40</volume>, <fpage>299</fpage>&#x2013;<lpage>324</lpage>. <pub-id pub-id-type="doi">10.1111/cogs.12225</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Foster</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Craenen</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Deshmukh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lemon</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Bastianelli</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Dondrup</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <source>MuMMER: socially intelligent human-robot interaction in public spaces</source>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Glass</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Challenges for spoken dialogue systems</article-title>. <source>Proc. 1999 IEEE ASRU Workshop (MIT Laboratory Comput. Sci. Camb. MA, USA)</source> <volume>696</volume>, <fpage>10</fpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gunson</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Garcia</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Siei&#x144;ska</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Addlesee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dondrup</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lemon</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). &#x201c;<article-title>A visually-aware conversational robot receptionist</article-title>,&#x201d; in <source>Proceedings of the 23rd annual meeting of the special interest group on discourse and dialogue</source>, <fpage>645</fpage>&#x2013;<lpage>648</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hahkio</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Service robots&#x2019; feasibility in the hotel industry: a case study of Hotel Presidentti</source>. <publisher-loc>Vantaa, Finland</publisher-loc>: <publisher-name>Laurea University of Applied Sciences</publisher-name>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Honig</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Oron-Gilad</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Understanding and resolving failures in human-robot interaction: literature review and model development</article-title>. <source>Front. Psychol.</source> <volume>9</volume>, <fpage>861</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2018.00861</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ince</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Nakadai</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Rodemann</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hasegawa</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tsujino</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>A hybrid framework for ego noise cancellation of a robot</article-title> in <conf-name>2010 IEEE international Conference on Robotics and automation</conf-name>. <publisher-name>IEEE</publisher-name>, <fpage>3623</fpage>&#x2013;<lpage>3628</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Inoue</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Ekstedt</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Kawahara</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Skantze</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Multilingual turn-taking prediction using voice activity projection</article-title> in <conf-name>Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024)</conf-name>, <fpage>11873</fpage>&#x2013;<lpage>11883</lpage>. Available at: <ext-link ext-link-type="uri" xlink:href="https://aclanthology.org/2024.lrec-main.1036/">https://aclanthology.org/2024.lrec-main.1036/</ext-link>
</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Langedijk</surname>
<given-names>R. M.</given-names>
</name>
<name>
<surname>Odabasi</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Graf</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Studying drink-serving service robots in the real world</article-title> in <conf-name>2020 29th IEEE international conference on robot and human interactive communication</conf-name>. <publisher-name>IEEE</publisher-name>, <fpage>788</fpage>&#x2013;<lpage>793</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lemon</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Conversational AI for multi-agent communication in natural language</article-title>. <source>AI Commun.</source> <volume>35</volume>, <fpage>295</fpage>&#x2013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.3233/aic-220147</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lemon</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Gruenstein</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Multithreaded context for robust conversational interfaces: context-sensitive speech recognition and interpretation of corrective fragments</article-title>. <source>ACM Trans. Computer-Human Interact. (TOCHI)</source> <volume>11</volume>, <fpage>241</fpage>&#x2013;<lpage>267</lpage>. <pub-id pub-id-type="doi">10.1145/1017494.1017496</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Nakadai</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kumon</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Okuno</surname>
<given-names>H. G.</given-names>
</name>
<name>
<surname>Hoshiba</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wakabayashi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Washizaki</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Development of microphone-array-embedded UAV for search and rescue task</article-title> in <conf-name>IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</conf-name> <publisher-name>IEEE</publisher-name>, <fpage>5985</fpage>&#x2013;<lpage>5990</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Nikolopoulos</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kuester</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Sheehan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ramteke</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Karmarkar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Thota</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <source>Robotic agents used to help teach social skills to children with autism: the third generation</source>. <publisher-name>IEEE</publisher-name>, <fpage>253</fpage>&#x2013;<lpage>258</lpage>.</citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Porcheron</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>J. E.</given-names>
</name>
<name>
<surname>Reeves</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sharples</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Voice interfaces in everyday life</article-title>,&#x201d; in <conf-name>Proceedings of the 2018 CHI conference on human factors in computing systems</conf-name>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sackl</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pretolesi</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Burger</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ganglbauer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tscheligi</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Social robots as coaches: how human-robot interaction positively impacts motivation in sports training sessions</article-title> in <conf-name>2022 31st IEEE international conference on robot and human interactive communication</conf-name>. <publisher-name>IEEE</publisher-name>, <fpage>141</fpage>&#x2013;<lpage>148</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Schauer</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Sweeny</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lyttle</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Said</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Szeles</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2023</year>). <article-title>Detecting agreement in multi-party conversational AI</article-title> in <source>Proceedings of the workshop on advancing GROup UNderstanding and robots a Daptive behaviour (GROUND)</source>, <fpage>5</fpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Schmidt</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>L&#xf6;llmann</surname>
<given-names>H. W.</given-names>
</name>
<name>
<surname>Kellermann</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A novel ego-noise suppression algorithm for acoustic signal enhancement in autonomous systems</article-title> in <conf-name>2018 IEEE international Conference on acoustics, Speech and signal processing (ICASSP)</conf-name> <publisher-name>IEEE</publisher-name>, <fpage>6583</fpage>&#x2013;<lpage>6587</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Skantze</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Turn-taking in conversational systems and human-robot interaction: a review</article-title>. <source>Comput. Speech and Lang.</source> <volume>67</volume>, <fpage>101178</fpage>. <pub-id pub-id-type="doi">10.1016/j.csl.2020.101178</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Spekking</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <source>Amazon Echo dot (RS03QR) - LED and microphone board</source>. <publisher-loc>Amsterdam, Netherlands</publisher-loc>: <publisher-name>Wikimedia</publisher-name>.</citation>
</ref>
<ref id="B36">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Stegner</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Senft</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mutlu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Situated participatory design: a method for <italic>in situ</italic> design of robotic interaction with older adults</article-title> in <conf-name>Proceedings of the 2023 CHI conference on human factors in computing systems</conf-name>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>.</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tatarian</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Stower</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rudaz</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chamoux</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kappas</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chetouani</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>How does modality matter? investigating the synthesis and effects of multi-modal robot behavior on social intelligence</article-title>. <source>Int. J. Soc. Robotics</source> <volume>14</volume>, <fpage>893</fpage>&#x2013;<lpage>911</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-021-00839-w</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thrun</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>When robots meet people</article-title>. <source>IEEE Intelligent Syst. their Appl.</source> <volume>13</volume>, <fpage>27</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1109/5254.683178</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Oviatt</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A taxonomy of social errors in human-robot interaction</article-title>. <source>ACM Trans. Human-Robot Interact. (THRI)</source> <volume>10</volume>, <fpage>1</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1145/3439720</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tolmeijer</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Weiss</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hanheide</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lindner</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Powers</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Dixon</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Taxonomy of trust-relevant failures and mitigation strategies</article-title>. <source>Proc. 2020 acm/ieee Int. Conf. human-robot Interact.</source>, <fpage>3</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/3319502.3374793</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Traum</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Issues in multiparty dialogues. Advances in agent communication: international workshop on agent communication languages, ACL 2003</article-title> in <source>Revised and invited papers</source>. <publisher-loc>Melbourne, Australia</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>201</fpage>&#x2013;<lpage>211</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Villalpando</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Schillaci</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hafner</surname>
<given-names>V. V.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Predictive models for robot ego-noise learning and imitation</article-title> in <conf-name>2018 joint IEEE 8th international Conference on Development and Learning and epigenetic robotics (ICDL-EpiRob)</conf-name> <publisher-name>IEEE</publisher-name>, <fpage>263</fpage>&#x2013;<lpage>268</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wagner</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kraus</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lindemann</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Minker</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Comparing multi-user interaction strategies in human-robot teamwork</article-title>. <source>Int. Workshop Spok. Dialogue Syst. Technol. IWSDS</source> <volume>2023</volume>, <fpage>12</fpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wagner</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kraus</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rach</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Minker</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2021</year>). <source>How to address humans: system barge-in in multi-user HRI</source>. <publisher-name>Springer Singapore</publisher-name>, <fpage>147</fpage>&#x2013;<lpage>152</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-15-9323-9_13</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Welch</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Apple&#x2019;s new HomePod unsurprisingly sounds close to the original</article-title>. <source>Verge</source>.</citation>
</ref>
<ref id="B46">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Williams</surname>
<given-names>J. D.</given-names>
</name>
</person-group> (<year>2009</year>). <source>Spoken dialogue systems: challenges, and opportunities for research</source>. <publisher-loc>Merano, Italy</publisher-loc>: <publisher-name>ASRU</publisher-name>, <fpage>25</fpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wilson</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Why Amazon radically redesigned the Echo</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Fast Company</publisher-name>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xiao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Warnell</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Stone</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Autonomous ground navigation in highly constrained spaces: lessons learned from the benchmark autonomous robot navigation challenge at ICRA 2022</article-title>. <source>IEEE Robotics and Automation Mag.</source> <volume>29</volume>, <fpage>148</fpage>&#x2013;<lpage>156</lpage>. <pub-id pub-id-type="doi">10.1109/mra.2022.3213466</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gan</surname>
<given-names>W. S.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Er</surname>
<given-names>M. H.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Acoustic beamforming of a parametric speaker comprising ultrasonic transducers</article-title>. <source>Sensors Actuators A Phys.</source> <volume>125</volume>, <fpage>91</fpage>&#x2013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/j.sna.2005.04.037</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Kuang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Optimal audio beam pattern synthesis for an enhanced parametric array loudspeaker</article-title>. <source>J. Acoust. Soc. Am.</source> <volume>154</volume>, <fpage>3210</fpage>&#x2013;<lpage>3222</lpage>. <pub-id pub-id-type="doi">10.1121/10.0022415</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>