<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">733954</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2022.733954</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Perception is Only Real When Shared: A Mathematical Model for Collaborative Shared Perception in Human-Robot Interaction</article-title>
<alt-title alt-title-type="left-running-head">Matarese et al.</alt-title>
<alt-title alt-title-type="right-running-head">Collaborative Shared Perception in HRI</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Matarese</surname>
<given-names>Marco</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1085399/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Rea</surname>
<given-names>Francesco</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/105752/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Sciutti</surname>
<given-names>Alessandra</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/57983/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>DIBRIS Department</institution>, <institution>University of Genoa</institution>, <addr-line>Genoa</addr-line>, <country>Italy</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>RBCS Unit</institution>, <institution>Italian Institute of Technology</institution>, <addr-line>Genoa</addr-line>, <country>Italy</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>CONTACT Unit</institution>, <institution>Italian Institute of Technology</institution>, <addr-line>Genoa</addr-line>, <country>Italy</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/985088/overview">Yukie Nagai</ext-link>, The University of Tokyo, Japan</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1005781/overview">Casey Kennington</ext-link>, Boise State University, United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1535811/overview">Hanzhe Zhang</ext-link>, Michigan State University, United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/374625/overview">Kazunori Terada</ext-link>, Gifu University, Japan</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Marco Matarese, <email>marco.matarese@iit.it</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Human-Robot Interaction, a section of the journal Frontiers in Robotics and AI</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>15</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>9</volume>
<elocation-id>733954</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Matarese, Rea and Sciutti.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Matarese, Rea and Sciutti</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Partners have to build a shared understanding of their environment in everyday collaborative tasks by aligning their perceptions and establishing a common ground. This is one of the aims of shared perception: revealing characteristics of the individual perception to others with whom we share the same environment. In this regard, social cognitive processes, such as joint attention and perspective-taking, form a shared perception. From a Human-Robot Interaction (HRI) perspective, robots would benefit from the ability to establish shared perception with humans and a common understanding of the environment with their partners. In this work, we wanted to assess whether a robot, considering the differences in perception between itself and its partner, could be more effective in its helping role and to what extent this improves task completion and the interaction experience. For this purpose, we designed a mathematical model for a collaborative shared perception that aims to maximise the collaborators&#x2019; knowledge of the environment when there are asymmetries in perception. Moreover, we instantiated and tested our model <italic>via</italic> a real HRI scenario. The experiment consisted of a cooperative game in which participants had to build towers of Lego bricks, while the robot took the role of a suggester. In particular, we conducted experiments using two different robot behaviours. In one condition, based on shared perception, the robot gave suggestions by considering the partners&#x2019; point of view and using its inference about their common ground to select the most informative hint. In the other condition, the robot just indicated the brick that would have yielded a higher score from its individual perspective. The adoption of shared perception in the selection of suggestions led to better performances in all the instances of the game where the visual information was not <italic>a priori</italic> common to both agents. However, the subjective evaluation of the robot&#x2019;s behaviour did not change between conditions.</p>
</abstract>
<kwd-group>
<kwd>shared perception</kwd>
<kwd>human-robot interaction</kwd>
<kwd>theory of mind</kwd>
<kwd>joint attention</kwd>
<kwd>shared autonomy</kwd>
<kwd>nonverbal communication</kwd>
<kwd>gaze cues</kwd>
</kwd-group>
<contract-sponsor id="cn001">H2020 European Research Council<named-content content-type="fundref-id">10.13039/100010663</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>The ability to cooperate and communicate is inherent in human nature. People can easily share information with others to achieve common objectives. However, human interactions require a common ground to succeed in reaching a shared goal (<xref ref-type="bibr" rid="B49">Thomaz et al., 2019</xref>). The lack of this common ground could cause misunderstanding and mistakes even in simple collaborative tasks, e.g., when two agents perceive the same objects differently (<xref ref-type="bibr" rid="B13">Chai et al., 2014</xref>). Indeed, between collaborating agents, even a slight misalignment on the common ground may be due to different perceptions of the shared environment. Each agent can have a peculiar perception of such an environment because of differences in perspective, sensory capabilities (e.g., colour-blindness) or prior knowledge (<xref ref-type="bibr" rid="B36">Mazzola et al., 2020</xref>).</p>
<p>Despite the different perceptions of a shared environment, people can naturally interact with each other. Two collaborators can align on a common ground of beliefs, intentions and perceptions, ideally maximising both performances and the shared knowledge. Establishing shared perception aims to build a common understanding of the environment by bridging the different individual perceptions. For instance, this implies revealing what is hidden in the eyes of a collaborator, annulling perceptual asymmetries. Even when this is not entirely possible, e.g., the hidden item cannot be uncovered, a partner can leverage the shared knowledge to reveal something about covered items to help the collaborator make informed actions. Moreover, suppose the two partners have a good understanding of the other&#x2019;s goals and intentions. In that case, they will also know how to maximise the shared knowledge, selecting when it is crucial to reveal the differences in individual information and when it is more effective to focus only on the shared space.</p>
<p>A crucial aspect of shared perception is perspective-taking (<xref ref-type="bibr" rid="B55">Wolgast et al., 2020</xref>). By taking the point of view of a partner, an agent can understand the differences between their perception and that of their collaborators. Furthermore, the ability to understand partners&#x2019; actions as guided by intentional behaviours and to ascribe to them mental states is called Theory of Mind (ToM) (<xref ref-type="bibr" rid="B22">G&#xf6;r&#xfc;r et al., 2017</xref>). Shared perception is a pivotal part of ToM because, by combining it with perspective-taking, an agent can more easily infer the rationale behind one&#x2019;s actions and - more importantly - understand that a collaborator&#x2019;s unreasonable action could be due to a mismatch between their perceptions.</p>
<p>From the robotic point of view, shared perception can help robots infer their partner&#x2019;s intentions and, taking their point of view, reason over them and consequently select the most effective collaborative action. By providing help in establishing a common understanding of the shared environment between the two partners, shared perception is particularly beneficial when human and robotic perceptions are not identical. This might happen in scenarios where the robot can perceive advantageous characteristics of the environment that the human collaborator can not. In particular, considering a one-to-one Human-Robot Interaction (HRI) scenario, both the human and robot could benefit from building a ToM of the other and establishing a shared perception. Considering ToM, the human needs to translate the robot&#x2019;s actions in terms of objectives, beliefs and intentions (<xref ref-type="bibr" rid="B47">Scassellati, 2002</xref>), while the robot needs to infer its collaborator&#x2019;s mental states (<xref ref-type="bibr" rid="B15">Devin and Alami, 2016</xref>) to anticipate the unfolding of the following actions better. Shared perception is then fundamental to allow each of the two partners to be aware of what the other can perceive and which action should be performed to maximise the potential of achieving such objectives.</p>
<p>Let us imagine a person assembling an Ikea piece of furniture with their domestic robot. This task needs some tools (e.g., screwdrivers, screws, bolts) to assemble the different parts (e.g., shelves). Each tool has different characteristics that make it useful or useless to assemble a particular part of the piece of furniture. The assembling task, <italic>per se</italic>, can make the environment very chaotic, given all the building material. This brings possible obstacles in the person&#x2019;s perceptual space and could result in asymmetries in the perception between the two agents. In this setting, the robot can exploit shared perception by trying to resolve such asymmetries, verbally communicating the characteristics of potentially valuable objects, or handing over the object which is the best according to what it sees and what it infers about what the human partner is seeing.</p>
<p>The field of HRI has dedicated wide attention to the social phenomena that constitute the backbone of shared perception, such as joint attention (<xref ref-type="bibr" rid="B38">Nagai et al., 2003</xref>), perspective-taking (<xref ref-type="bibr" rid="B17">Fischer and Demiris, 2016</xref>), common knowledge (<xref ref-type="bibr" rid="B30">Kiesler, 2005</xref>), communication (<xref ref-type="bibr" rid="B35">Mavridis, 2015</xref>) and ToM (<xref ref-type="bibr" rid="B8">Bianco and Ognibene, 2019</xref>). In this work, we wanted to investigate how these mechanisms work in synergy to lead to shared perception between a human and a robot and what impact shared perception has on HRI.</p>
<p>Hence, we present a mathematical model for shared perception, through which a collaborative robot aims at maximising both performances and the collaborator&#x2019;s knowledge about the environment. To test the model, we asked participants to play a cooperative game with the iCub robot in a real HRI scenario in which they had to build a tower with LEGO bricks (<xref ref-type="fig" rid="F1">Figure 1</xref>). During the task, the robot could either leverage shared perception principles (SP) or just aim at maximising the overall task performance (NSP). We designed our experiments to create specific critical moments&#x2014;the <italic>conflicts</italic>&#x2014;characterised by a mismatch between participants&#x2019; and robot&#x2019;s perceptions.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>A participant and the robot iCub performing the task.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g001.tif"/>
</fig>
<p>The following sections are organised as follows: <xref ref-type="sec" rid="s2">Section 2</xref> presents an overview of related works; <xref ref-type="sec" rid="s3">Section 3</xref> describes the mathematical model, the experiment and the software architecture; <xref ref-type="sec" rid="s4">Section 4</xref> concerns the experimental results. The last two sections are dedicated to discussion and conclusion.</p>
</sec>
<sec id="s2">
<title>2 Background and Related Works</title>
<p>Shared perception is a complex mechanism, that entails a range of social skills. A robot, to establish shared perception, would need the awareness that the collaborator could have a different perception and should also be aware of which are those differences, e.g., what parts of their perceivable environment are hidden to its partner. Furthermore, it would need to know the other&#x2019;s goal and its relation to the objects in the environment. Last, the robot should estimate how the partner understands its own behaviour to provide communications that the collaborator can effectively understand and enact.</p>
<p>One of the fundamental mechanisms of the understanding that others might perceive the world differently is Perspective-Taking (PT). PT is <italic>&#x201c;a multifaceted skill set, involving the disposition, motivation, and contextual attempts to consider and understand other individuals&#x201d;</italic> (<xref ref-type="bibr" rid="B55">Wolgast et al., 2020</xref>). As well as humans, a robot can infer humans&#x2019; perception through mechanisms of PT: it has been proved that PT also improves action recognition performances (<xref ref-type="bibr" rid="B27">Johnson and Demiris, 2005</xref>; <xref ref-type="bibr" rid="B28">Johnson and Demiris, 2007</xref>). Therefore, algorithms for PT in HRI have been proposed to disambiguate whether an object is visible to people facing the robot, using just their head pose (<xref ref-type="bibr" rid="B17">Fischer and Demiris, 2016</xref>). Moreover, several PT-based architectures have been proposed to estimate where a person will execute a future task (<xref ref-type="bibr" rid="B40">Pandey et al., 2013</xref>), to then produce proactive and collaborative behaviours. Other contexts in which PT has been investigated are the military field (<xref ref-type="bibr" rid="B29">Kennedy et al., 2007</xref>), where the robot used those mechanisms to understand if it was visible to an enemy or in human-robot teaching scenarios (<xref ref-type="bibr" rid="B7">Berlin et al., 2006</xref>; <xref ref-type="bibr" rid="B11">Breazeal et al., 2006</xref>). Several works showed that PT is beneficial to disambiguate both things and circumstances (<xref ref-type="bibr" rid="B46">Ros et al., 2010</xref>), such as tools and commands (<xref ref-type="bibr" rid="B50">Trafton et al., 2005a</xref>; <xref ref-type="bibr" rid="B51">Trafton et al., 2005b</xref>). Hence, we can use PT mechanisms to &#x201c;put ourselves in one&#x2019;s shoes&#x201d; so that we can understand their point of view and build a common ground on which to base an efficient collaboration (<xref ref-type="bibr" rid="B12">Brown-Schmidt and Heller, 2018</xref>). In this study, we focus on visual PT, which is the ability to see the world from another person&#x2019;s perspective, taking into account what they see and how they see it (<xref ref-type="bibr" rid="B19">Flavell, 1977</xref>).</p>
<p>To better align the perspectives of two or more collaborating agents, we need that all of them build a reliable Theory of Mind (ToM) of the others (<xref ref-type="bibr" rid="B34">Marchetti et al., 2018</xref>). This means that, in addition to perception, partners have to base their interaction also on shared knowledge: at least, they need to know what the other agents already know so that they can easily anticipate others&#x2019; actions (<xref ref-type="bibr" rid="B54">Winfield, 2018</xref>). Several works in robotics and HRI took inspiration from other fields such as psychology and philosophy to model ToMs for robots. For example, <xref ref-type="bibr" rid="B47">Scassellati, (2002)</xref> discuss the theories presented in (<xref ref-type="bibr" rid="B4">Baron-Cohen, 1997</xref>) and (<xref ref-type="bibr" rid="B32">Leslie et al., 1994</xref>) on developmental ToM in children to build robots with similar capabilities. Rather, more recent works implement ToMs through a Bayesian model (<xref ref-type="bibr" rid="B31">Lee et al., 2019</xref>) to best solve human-robot nonverbal communication issues. In HRI, it has been shown that people appreciate robots that show ToM-like abilities as teammates for their ability to identify the most likely cause of others&#x2019; behaviour (<xref ref-type="bibr" rid="B26">Hiatt et al., 2011</xref>). This is also because people perceive such robots as more capable of aligning themselves to persons by fully recognising their environment (<xref ref-type="bibr" rid="B6">Benninghoff et al., 2013</xref>). Moreover, developmental human-inspired ToMs have been presented to enhance the quality of the HRI itself: e.g., in (<xref ref-type="bibr" rid="B52">Vinanzi et al., 2019</xref>), the authors modelled the trustworthiness of the robot&#x2019;s human collaborator using a probabilistic ToM and a trust model supported by an episodic memory system.</p>
<p>The gaze plays a pivotal role in facilitating the understanding of others&#x2019; goals and enabling intuitive collaboration. Gaze movements have been proved to be very helpful in collaborative scenarios (<xref ref-type="bibr" rid="B16">Fischer et al., 2015</xref>). <xref ref-type="bibr" rid="B41">Pierno et al. (2006)</xref> observed the same neural response in people observing someone cueing an object and in people observing someone reaching an object to grasp it: gaze cues are a powerful indicator of people&#x2019;s intentions. People are sensitive also to robot gazing when this signal is directed at an object in the environment. It has already been proved that people predict which objects to select using referential gaze cues from robots, even if they are not consciously aware of those cues (<xref ref-type="bibr" rid="B37">Mutlu et al., 2009</xref>). Indeed, through gaze cues, a robot could highlight parts of the environment, thus providing information about its perception (<xref ref-type="bibr" rid="B20">Fussell et al., 2003</xref>). Several studies proved that people are very good at identifying the target of partners&#x2019; referential gaze to use this information to predict their future actions (<xref ref-type="bibr" rid="B48">Staudte and Crocker, 2011</xref>; <xref ref-type="bibr" rid="B10">Boucher et al., 2012</xref>). When people refer to objects around them, they look at those objects before manipulating them (<xref ref-type="bibr" rid="B23">Griffin and Bock, 2000</xref>; <xref ref-type="bibr" rid="B24">Hanna and Brennan, 2007</xref>; <xref ref-type="bibr" rid="B56">Yu et al., 2012</xref>) and when partners refer to objects, people use their gaze to predict their following intentions to quickly respond to the partner&#x2019;s reference (<xref ref-type="bibr" rid="B10">Boucher et al., 2012</xref>). In fact, with little information about the partner&#x2019;s gaze, people are slower at responding to their partner&#x2019;s communication (<xref ref-type="bibr" rid="B10">Boucher et al., 2012</xref>). Moreover, objects that are not related to a task are rarely fixated (<xref ref-type="bibr" rid="B25">Hayhoe and Ballard, 2005</xref>). In sum, beyond implicitly revealing an agent&#x2019;s future intentions, gaze movements can be an effective form of nonverbal communication (<xref ref-type="bibr" rid="B44">Rea et al., 2016</xref>; <xref ref-type="bibr" rid="B1">Admoni and Scassellati, 2017</xref>; <xref ref-type="bibr" rid="B53">Wallkotter et al., 2021</xref>).</p>
<p>So far, the used approach in PT studies focused on the disambiguation of tools and commands to help the artificial agents build a common ground with their collaborators. With the current work, we want to move forward in using such mechanisms by considering robots capable of sharing information gathered from their own perspective but communicated by taking into consideration both the perspective of their human partners and the shared knowledge. Through shared perception mechanisms, we aim to go further in this approach by building a model that can enable robots to autonomously resolve situations characterised by asymmetries between their perception and one of their collaborators.</p>
<p>In the current work, contrary to what is already present in the literature about asymmetries in perception in HRI (<xref ref-type="bibr" rid="B13">Chai et al., 2014</xref>), we underline how the issue of creating a common ground also applies to interactions not mediated by language. Even in scenarios where the action possibilities are constrained, and the goal of the task is clear, the mismatch in perception requires a communicative effort to establish a shared understanding. We show that this can also be achieved with non-verbal signals.</p>
<p>For this purpose, in this paper, we provide a mathematical model that supports shared perception-based decisions for a robot helper in a collaborative task. We assess task performance and interaction experience when this model guides the robot hints. We compare them with interactions in which robot behaviour is driven just by the goal of maximising the task score.</p>
</sec>
<sec id="s3">
<title>3 Materials and Methods</title>
<sec id="s3-1">
<title>3.1 The Mathematical Model</title>
<p>The mathematical model for collaborative Shared Perception (SP) adopts a formulation taken from the theory of sets and probability. The elements that characterise the model are presented in a general and abstract way so that they can be instantiated depending on the circumstances. In particular, the model manipulates concepts such as objects, environment, personal/common perception and awareness. Here, we do not provide examples of instantiating those concepts, but later in this section, we discuss how we did it for our experiment.</p>
<p>The model considers only one-to-one interaction; thus, we always have an agent (<italic>a</italic>
<sub>1</sub>) aiming to share their perception with another agent (<italic>a</italic>
<sub>2</sub>). In this work, we consider <italic>a</italic>
<sub>1</sub> as a robot and <italic>a</italic>
<sub>2</sub> as a person. The model&#x2019;s core is a sort of common knowledge between the two agents that we call &#x201c;common awareness&#x201d;. In particular, <italic>a</italic>
<sub>1</sub> exploits elements belonging to this common knowledge to give insights about elements in its individual perception. Hence, the model&#x2019;s objective is to enable robots to exploit SP mechanisms. For this purpose, it aims to maximise partners&#x2019; awareness of objects they can not perceive by using elements belonging to the common ground already established. In particular, the model tries to share its individual perception with the partner, choosing the elements of common awareness that it could most appropriately exploit. In this sense, our model is about decision-making and not just communication.</p>
<p>To present our model, we need first to define its elements. We define the <bold>environment</bold> <italic>X</italic> &#x3d; {<italic>x</italic>: <italic>x</italic>is an object} as a finite set of objects. Thus, we adopt a <italic>closed world</italic> formulation: everything we consider belongs to the environment.</p>
<p>Moreover, we define an <bold>object</bold> <italic>x</italic> &#x2208; <italic>X</italic> as a finite set of characteristics, as follows: <italic>x</italic> &#x3d; {<italic>c</italic>
<sub>1</sub>, &#x2026; , <italic>c</italic>
<sub>
<italic>n</italic>
</sub>: <italic>c</italic>
<sub>
<italic>i</italic>
</sub>is a characteristic<italic>&#x2200;i</italic> &#x3d; 1, &#x2026; , <italic>n</italic>} where, for <italic>characteristics</italic>, we mean features such as colour, shape, etc. An object&#x2019;s characteristic have to be instantiated, <italic>i.e.</italic> if <italic>c</italic>
<sub>1</sub> refers to the object&#x2019;s colour, we could have <italic>x</italic> &#x3d; {<italic>c</italic>
<sub>1</sub> &#x3d; <italic>blue</italic>, &#x2026; }.</p>
<p>Moreover, an agent couples all these characteristics with a probability distribution describing the agent&#x2019;s degree of certainty on each characteristic&#x2019;s instance. Thus, from the agent <italic>a</italic>&#x2019;s point of view, the object <italic>x</italic> is a set of pairs where, to each object&#x2019;s characteristic, it is associated with a probability distribution over the set of all its possible instances: <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. The probability distribution functions associated with the objects&#x2019; characteristics can be derived from the task rules if the collaborative task is constrained. For example, the agent <italic>a</italic> associates to the characteristic &#x201c;colour,&#x201d; let us say <italic>c</italic>
<sub>1</sub>, of the object <italic>x</italic> a probability distribution <inline-formula id="inf2">
<mml:math id="m2">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> over the set of all the possible instances, let us say <italic>blue</italic>, <italic>red</italic> and <italic>black</italic>. Assuming that <italic>x</italic> is <italic>blue</italic>, if <italic>a</italic> knows that <italic>x</italic> is blue, then we would have <inline-formula id="inf3">
<mml:math id="m3">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>; on the other hand, if <italic>a</italic> has no information about the colour of <italic>x</italic>, then we would have <inline-formula id="inf4">
<mml:math id="m4">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.33</mml:mn>
</mml:math>
</inline-formula>, <inline-formula id="inf5">
<mml:math id="m5">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.33</mml:mn>
</mml:math>
</inline-formula>, and <inline-formula id="inf6">
<mml:math id="m6">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.33</mml:mn>
</mml:math>
</inline-formula>. The implicit assumption is that the objects belonging to the environment do not change over time.</p>
<p>With the elements described above, we can define the <bold>Personal Perception</bold> of the agent <italic>a</italic>, <italic>P</italic>
<sub>
<italic>a</italic>
</sub> &#x3d; {<italic>x</italic> &#x2208; <italic>X</italic>: the agent<italic>a</italic>can perceive<italic>x</italic>} as the finite set of the objects belonging to <italic>a</italic>&#x2019;s perception. From the definition of <italic>P</italic>
<sub>
<italic>a</italic>
</sub> follows that <italic>P</italic>
<sub>
<italic>a</italic>
</sub> &#x2286; <italic>X</italic> for each agent <italic>a</italic>. The definition of personal perception depends on the agent&#x2019;s capability: <italic>i.e.</italic>, if the agent <italic>a</italic> is a robot equipped with only a camera, then <italic>P</italic>
<sub>
<italic>a</italic>
</sub> refers to the object the robot can see through its camera.</p>
<p>Similarly, we define the <bold>Awareness Space</bold> of the agent <italic>a</italic>, <italic>W</italic>
<sub>
<italic>a</italic>
</sub> &#x3d; {<italic>x</italic> &#x2208; <italic>X</italic>: <italic>a</italic>is aware that<italic>x</italic> &#x2208; <italic>X</italic>} as the finite set containing the objects of which <italic>a</italic> is aware. We characterise the set <italic>W</italic>
<sub>
<italic>a</italic>
</sub> as follows: <italic>&#x2200;x</italic> &#x2208; <italic>X</italic>, if <italic>x</italic> &#x2208; <italic>P</italic>
<sub>
<italic>a</italic>
</sub>&#x21d2;<italic>x</italic> &#x2208; <italic>W</italic>
<sub>
<italic>a</italic>
</sub> for each agent <italic>a</italic>. Thus, it follows that <italic>P</italic>
<sub>
<italic>a</italic>
</sub> &#x2286; <italic>W</italic>
<sub>
<italic>a</italic>
</sub> &#x2286; <italic>X</italic> for each agent <italic>a</italic>.</p>
<p>We say that the agents <italic>a</italic>
<sub>1</sub> and <italic>a</italic>
<sub>2</sub> both perceive the object <italic>x</italic> if <italic>&#x2203;x</italic> &#x2208; <italic>X</italic>: <inline-formula id="inf7">
<mml:math id="m7">
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf8">
<mml:math id="m8">
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. As well, we say that the agents <italic>a</italic>
<sub>1</sub> and <italic>a</italic>
<sub>2</sub> are both aware of the object <italic>x</italic> if <italic>&#x2203;x</italic> &#x2208; <italic>X</italic>: <inline-formula id="inf9">
<mml:math id="m9">
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf10">
<mml:math id="m10">
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. Thus, we define the <bold>Common Perception</bold> between the agents <italic>a</italic>
<sub>1</sub> and <italic>a</italic>
<sub>2</sub> as follows: <inline-formula id="inf11">
<mml:math id="m11">
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. Similarly, we define the <bold>Common Awareness</bold> between <italic>a</italic>
<sub>1</sub> and <italic>a</italic>
<sub>2</sub> as follows: <inline-formula id="inf12">
<mml:math id="m12">
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. Hence, from the characterisation of <italic>W</italic>, we have that if <inline-formula id="inf13">
<mml:math id="m13">
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x21d2;</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>.</p>
<p>A SP communication from the agent <italic>a</italic>
<sub>1</sub> to the agent <italic>a</italic>
<sub>2</sub> <inline-formula id="inf14">
<mml:math id="m14">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mover>
<mml:mo>&#x2192;</mml:mo>
<mml:mrow>
<mml:mtext>SP</mml:mtext>
</mml:mrow>
</mml:mover>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> aims to maximise the knowledge of <italic>a</italic>
<sub>2</sub> of the objects belonging to the personal perception of <italic>a</italic>
<sub>1</sub>, <italic>P</italic>
<sub>
<italic>a</italic>
</sub>, that do not belong to the awareness space of <italic>a</italic>
<sub>2</sub>, <inline-formula id="inf15">
<mml:math id="m15">
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. To achieve this objective, SP exploits the common characteristics between objects belonging to <italic>a</italic>
<sub>1</sub>&#x2019;s personal perception and those belonging to <italic>a</italic>
<sub>1</sub> and <italic>a</italic>
<sub>2</sub>&#x2019;s common awareness. Thus, we say that the SP is possible if the following condition occurs:</p>
<p>
<inline-formula id="inf16">
<mml:math id="m16">
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mo>&#x2203;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mtext>&#x2009;so&#x2009;that&#x2009;</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mo>&#x2203;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mtext>&#x2009;so&#x2009;that&#x2009;</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>
</p>
<p>(where <inline-formula id="inf17">
<mml:math id="m17">
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="double-struck">N</mml:mi>
</mml:math>
</inline-formula>). In other words, if the object <italic>x</italic>
<sub>1</sub>, belonging to the personal perception of <italic>a</italic>
<sub>1</sub>, shares at least one characteristic with the object <italic>x</italic>
<sub>2</sub>, belonging to the common awareness of the two agents, then it is possible to communicate such common characteristics to give insights about <italic>x</italic>
<sub>1</sub>. If this precondition is respected, we have that <inline-formula id="inf18">
<mml:math id="m18">
<mml:mo>&#x2203;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mo>:</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x21a6;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf19">
<mml:math id="m19">
<mml:mo>&#x2203;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> so that <inline-formula id="inf20">
<mml:math id="m20">
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>: <inline-formula id="inf21">
<mml:math id="m21">
<mml:mi>c</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. It means that <italic>x</italic>
<sub>1</sub> is approximated by <italic>a</italic>
<sub>2</sub> with the object <inline-formula id="inf22">
<mml:math id="m22">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> that contains the characteristic <italic>c</italic> that <italic>a</italic>
<sub>1</sub> shared through its communication. The goals of the task influence the target of the communication. In the current model, we assume that such objectives are common between the two agents because of their collaboration.</p>
<p>The objective of the SP is to maximise <inline-formula id="inf23">
<mml:math id="m23">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mtext>_</mml:mtext>
<mml:mi>c</mml:mi>
<mml:mtext>_</mml:mtext>
<mml:mi>v</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, thus to minimise <italic>a</italic>
<sub>2</sub>&#x2019;s degree of uncertainty about such a characteristic. In the best case, when this uncertainty becomes zero because of <italic>a</italic>
<sub>1</sub>&#x2019;s communication, <italic>a</italic>
<sub>2</sub> can be sure of the value of <italic>c</italic>: it means that <inline-formula id="inf24">
<mml:math id="m24">
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mtext>_</mml:mtext>
<mml:mi>c</mml:mi>
<mml:mtext>_</mml:mtext>
<mml:mi>v</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>. This way <italic>a</italic>
<sub>1</sub> makes <inline-formula id="inf25">
<mml:math id="m25">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> less and less approximated and, once <italic>a</italic>
<sub>1</sub> can communicate all <italic>x</italic>
<sub>1</sub>&#x2019;s characteristics, or once <italic>a</italic>
<sub>2</sub> can infer all of them, we have <inline-formula id="inf26">
<mml:math id="m26">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x21d2;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x21d2;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. However, it is not always possible to make <inline-formula id="inf27">
<mml:math id="m27">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> collapse on <italic>x</italic> because it depends on both the agent&#x2019;s communication capabilities and context. In general, the objective of the shared perception <inline-formula id="inf28">
<mml:math id="m28">
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mover>
<mml:mo>&#x2192;</mml:mo>
<mml:mrow>
<mml:mtext>SP</mml:mtext>
</mml:mrow>
</mml:mover>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> on the object <italic>x</italic> is:</p>
<p>
<italic>&#x2200;c</italic>
<sub>
<italic>i</italic>
</sub> &#x2208; <italic>x</italic>, <inline-formula id="inf29">
<mml:math id="m29">
<mml:mi>max</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>Pr</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mtext>_</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mtext>_</mml:mtext>
<mml:mi>v</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>In most cases, it is not necessary to cancel the uncertainty or communicate all the object&#x2019;s features. Most importantly, we need to communicate the essential characteristics for the goal and improve the probability of having the right information so that the partner can make the right decision. From our formulation, follows that if <inline-formula id="inf30">
<mml:math id="m30">
<mml:mo>&#x2204;</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x21d2;</mml:mo>
</mml:math>
</inline-formula> it is not possible to do shared perception.</p>
<p>We point out that one of the main preconditions for a good functioning of the model is that both the robot and human are capable of using recursive ToM (<xref ref-type="bibr" rid="B42">Woodruff and Premack, 1978</xref>; <xref ref-type="bibr" rid="B3">Arslan et al., 2012</xref>). In short, recursive (or second-order) ToM is the ability to reason over the others&#x2019; estimation of our own mental states. People are good at using recursive ToM with their peers (<italic>e.g.</italic>, in strategic games (<xref ref-type="bibr" rid="B21">Goodie et al., 2012</xref>)) and recently <xref ref-type="bibr" rid="B14">de Weerd et al. (2017)</xref> proved that they spontaneously use recursive ToM also with artificial agents when these latter are capable of second-order ToM as well. Without recursive ToM, the robot could not assume that the human is aware of its own knowledge about the partner&#x2019;s difference in perception. At the same time, the human could not make sense of a robot&#x2019;s suggestion, without the awareness that the robot has a model of what that person perceives. Hence, the proposed model would not be effective without recursive ToM, as the selection of the most informative characteristic, and of how to communicate it, rests on the assumption that both agents entertain such understanding of each other. More in general, without recursive ToM a robot could not exploit at its best the everyday awareness it builds with its human collaborator. Moreover, both the robot and human would be weakly aligned (or, even worse, not aligned at all) on beliefs, desires and intentions.</p>
</sec>
<sec id="s3-2">
<title>3.2 The Experiment</title>
<p>We asked participants to build a tower with a maximum of five LEGO bricks by picking them among the ones available on the table in front of them. The bricks had different colours: we associated a score with each colour (<xref ref-type="table" rid="T1">Table 1</xref>). Participants could put a brick on the tower&#x2019;s top on each round, but only if its value was less or equal to the brick previously on top. The game ended when the tower was complete; i.e., either after five rounds or when all the available bricks had a higher value than the one on the top. The goal of the game was to maximise the score of the tower. The experimenter explained the rules before task initiation, underlining the importance of maximising the score of the tower.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>The list of colours, and their values, that participants could consult during the experiments.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Colour</th>
<th align="center">Value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Pink</td>
<td align="char" char=".">3</td>
</tr>
<tr>
<td align="left">Orange</td>
<td align="char" char=".">4</td>
</tr>
<tr>
<td align="left">Blue</td>
<td align="char" char=".">4</td>
</tr>
<tr>
<td align="left">Yellow</td>
<td align="char" char=".">6</td>
</tr>
<tr>
<td align="left">Black</td>
<td align="char" char=".">6</td>
</tr>
<tr>
<td align="left">Green</td>
<td align="char" char=".">8</td>
</tr>
<tr>
<td align="left">White</td>
<td align="char" char=".">10</td>
</tr>
<tr>
<td align="left">Red</td>
<td align="char" char=".">10</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="fig" rid="F1">Figure 1</xref> shows the experimental setup. Both the participants and iCub sat at a table, facing each other during the experiment. On the table, there was a sheet of paper reporting the values of the colours (<xref ref-type="table" rid="T1">Table 1</xref>), eight coloured bricks and four <italic>obstacles</italic>. The obstacles were little constructions that could hide a brick. Because of these, the iCub could not perceive two bricks, while the participants could not perceive the other two. Then, there were other four bricks that both iCub and the participants could see.</p>
<p>The task was presented as a collaborative game with the robot. The participants were the <italic>builders</italic>: they have to physically take one brick at a time from the table and insert it on the top of the tower. Instead, the robot was the <italic>suggester</italic>: it could suggest which brick to take.</p>
<p>Each round of the game was structured in three distinct phases: i) the inspection, ii) the communication and iii) the action. The inspection phase i) had a fixed duration of 15&#xa0;seconds. During this period, the participants should inspect the table to choose a candidate brick to put on the tower during the action phase. In the communication phase ii), the robot provided its suggestion by looking at the brick it wanted to propose to the partner. We told participants that, during the second phase, they would expect a suggestion from the robot, which was collaborating with them. Three seconds after the robot&#x2019;s gaze motion, an acoustic signal informed the participants of the beginning of the action phase iii). In this latter phase, the participant chose, picked up a brick and positioned it on the tower. At the beginning of this phase, iCub stared at the tower under construction. Once the participants placed the brick on the tower, the robot inspected the table with its gaze. We limited the communication between the participants and the robot to keep the interaction as minimalist as possible: indeed, we allowed them to use only gaze.</p>
<p>We choose the bricks&#x2019; configurations to force critical moments (<italic>conflicts</italic>), where the selection of the best brick was different considering the personal views of the two agents. We ran a series of task simulations to find the configurations that maximised such differences. Once found, we selected two; then, we created the other two configurations by replacing some bricks with ones of a different colour but equal value. Thus, the participants had the feeling of playing with four different configurations; however, they played with only two configurations. This way, each participant could perform the task with both robot&#x2019;s modes (SP and NSP&#x2014;see description below) in both bricks&#x2019; configurations. This simple expedient allowed us to present the same conflicts, for each condition, to all participants. In particular, we performed a within-subject user study so that each participant could face both the experimental conditions. <xref ref-type="fig" rid="F2">Figure 2</xref> shows the configurations we used: the configuration in <xref ref-type="fig" rid="F2">Figure 2A</xref> was equivalent to the one shown in <xref ref-type="fig" rid="F2">Figure 2B</xref>; the same applies to <xref ref-type="fig" rid="F2">Figure 2C</xref> and <xref ref-type="fig" rid="F2">Figure 2D</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The four bricks&#x2019; setups used in the experiments. Setups <bold>(A,B)</bold> and <bold>(C,D)</bold> are equivalent, respectively, meaning that, even if it seems that they are different, they include bricks with the same value in the same position (red and white bricks, and black and yellow ones have the same value).</p>
</caption>
<graphic xlink:href="frobt-09-733954-g002.tif"/>
</fig>
<p>Participants did not know <italic>a priori</italic> which sets of brick colours were present in each session. In addition, <xref ref-type="table" rid="T1">Table 1</xref> also shows pink as a possible brick colour, corresponding to the lowest value of all, but none of the configurations had pink bricks. We added such a colour to avoid settings where participants could know <italic>a priori</italic> that the bricks visible to them corresponded to an overall minimum value.</p>
<p>Before each session, we told the participants that they would interact with a robot powered by a new program, so it would be like interacting with a different robot. In both conditions, the robot had the same knowledge of the environment&#x2014;the position of the bricks it could see (i.e., not occluded to it) and their colour&#x2014;and it used the exact internal representation of the task. Regardless of the robot mode, iCub always suggested one of the most valuable bricks to maximise the tower&#x2019;s value.</p>
<p>The difference between SP and NSP behaviours was in the order of the suggestions in case of multiple best options from the robot&#x2019;s perspective (e.g., during the conflicts). The SP-iCub, following the shared perception model, hinted at the bricks which would maximise the information about the relevant properties of the bricks hidden to participants. On the other hand, the NSP-robot suggested the bricks following its internal representation of the task (see <xref ref-type="sec" rid="s3-2-1-1">Section 3.2.1.1</xref>), without taking into account the perspective of its partner. In the current experiment, such internal representation actually led the robot to behave in the conflicts in opposite ways in the SP and NSP settings. This way, we ensured the maximum difference between the two robotic behaviours.</p>
<p>All the participants had to perform four trials, one for each configuration of the bricks: during one session of two consecutive trials, they had to interact with the SP-robot and, during the other two consecutive trials, with the NSP-robot. We counterbalanced the order of the presented setups and the robot mode.</p>
<p>Before and after each experiment, we asked participants how many bricks they thought the robot could see, how many they could see, and how many bricks they thought were on the table. All participants answered that there were eight bricks on the table (during the instruction, we told them that there was a brick behind each obstacle), that they could see six of them, and that also iCub could see just six bricks. Thus, at both the beginning and the end of the experiments, participants understood that iCub could not see certain bricks. In our setup, we assumed that participants would accept the robot&#x2019;s suggestions as the best options from its point of view. As we ensured after the experiments, all participants perceived the robot&#x2019;s gazing as suggestions to take those bricks; they also reported to us that they assumed it behaved like that based on the colour of the bricks.</p>
<p>Familiar situations that reflect the mechanisms of our task are competitive team card games. In a card game, each player has both private individual perception (the cards in their hand) and a common perception with the others (the cards on the board). In team-based games, people in the same team aim to maximise both the chances of victory and partners&#x2019; knowledge of the cards in their hands. We want to take as an example the game of <italic>&#x201c;Briscola&#x201d;</italic>
<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref>: a famous Italian competitive turn-based card game involving two teams, each consisting of two players. Here, players have to discard one card <italic>per</italic> turn, and both value and seed of the cards determine which team scores points. During the game, the partners try to inform each other of their private cards by strategically selecting which cards to discard. The game rules forbid players to inform the partner about the cards in their hand directly; thus, they have to play aiming at maximising both the probability of winning and the possibility for the partner to infer their cards. In our experiment, the SP-robot acted with this double aim, while the NSP-robot played only with the first objective.</p>
<sec id="s3-2-1">
<title>3.2.1 The Experimental Software Architecture</title>
<p>To perform our experiment, we developed a software architecture composed of three main modules: the knowledge module, the communication module, and the reasoner.</p>
<sec id="s3-2-1-1">
<title>3.2.1.1 The Knowledge Module</title>
<p>The knowledge module aimed to manage the knowledge base; also, it provided information to the reasoner. The knowledge base was defined once at the beginning of the experiment and updated online after each move. It maintained a graph to represent the task and a stack to represent the available bricks. The stack stored the bricks in decreasing value order to consider them in such decreasing order: it maintained only the bricks that the robot could perceive. In the stack, the bricks with the same value respected their positioning order on the table. Moreover, through the graph, the robot could track the task&#x2019;s progress and the next possible moves. The graph&#x2019;s vertices represented the bricks on the table, while the arrows represented the possible moves: the vertex <italic>i</italic> had an arrow towards the vertex <italic>j</italic> if and only if it could stack the brick <italic>j</italic> after the brick <italic>i</italic>. According to participants&#x2019; choices, both the graph and the stack were updated during the task.</p>
</sec>
<sec id="s3-2-1-2">
<title>3.2.1.2 The Communication Module</title>
<p>The communication module mainly aimed at sending commands to a control module that accounted for both robot&#x2019;s neck and eyes kinematics: the <italic>iKinGazeCtrl</italic> module (<xref ref-type="bibr" rid="B45">Roncone et al., 2016</xref>). It combines those independent controls to ensure the convergence of the robot&#x2019;s fixation point on its target. The <italic>iKinGazeCtrl</italic> module allowed the robot to have biologically inspired movements: this makes the robot&#x2019;s movements more natural. We used a combined approach because eye-based estimation of the observed location has been proven to be much more informative than head-based one, at least for human observations (<xref ref-type="bibr" rid="B39">Palinko et al., 2016</xref>).</p>
</sec>
<sec id="s3-2-1-3">
<title>3.2.1.3 The Reasoner</title>
<p>The Reasoner aimed to guide robot behaviour by collecting information from the Knowledge Module, reasoning over them and then deciding the robot&#x2019;s actions. In particular, it gathered from the Knowledge Module information about the possible next moves, the available bricks on the table, and what bricks were visible from its own perspective. We provided <italic>a priori</italic> this information to the robot. By exploiting this information, the reasoner knew what brick the robot should indicate to participants at each moment of the task.</p>
<p>In particular, in NSP mode, the reasoner used the following heuristic: it indicated the brick currently on the top of the knowledge base&#x2019;s stack (according to the task&#x2019;s rules). Such a brick was always one of the bricks with the higher value among those the robot could perceive (compared to the one currently on the top of the tower, since we designed the reasoner so that it would follow the game rules). Thus, in NSP mode, the robot considered only its perspective and task rules.</p>
<p>On the other hand, in SP mode, the reasoner had a more complex approach, also considering the collaborator&#x2019;s perspective. Instead of considering only the brick on the top of the knowledge base&#x2019;s stack&#x2014;the <italic>candidate</italic> brick&#x2014;it also reasoned on the bricks perceivable by both the robot and participants. If the second brick on the stack had the same value as the one on the top, then a conflict may have occurred: thus, the reasoner asked for information from the knowledge base to understand that. If it was not the case, the reasoner decided to behave as in NSP mode. In the former case, the reasoner picked the brick to indicate based on the SP model described in <xref ref-type="sec" rid="s3-1">Section 3.1</xref>, aiming at minimising participants&#x2019; uncertainty about the relevant properties of the hidden brick.</p>
</sec>
</sec>
<sec id="s3-2-2">
<title>3.2.2 Conflicts</title>
<p>We designed our experiments to elicit two critical moments for each task: the <italic>conflicts</italic>. During the conflicts, there was a mismatch between participants&#x2019; perception and the robot&#x2019;s since the best brick was hidden from the robot&#x2019;s view or the participants&#x2019;. This mismatch could lead to different brick choices because some important information was unavailable to one of the two agents involved. We designed three types of conflict: the <italic>Main Conflict</italic>, the <italic>First Brick Conflict</italic>, and the <italic>Mid-Game Conflict</italic>.</p>
<sec id="s3-2-2-1">
<title>3.2.2.1 The Main Conflict</title>
<p>The main conflict occurred in both the setups toward the end of the task: participants faced this type of conflict during each session, for a total amount of four main conflicts <italic>per</italic> participant. <xref ref-type="fig" rid="F3">Figure 3</xref> (right) shows an example of the bricks configuration during this conflict. Since this conflict was practically identical in both of the setups, here we present only the configuration of the first one.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Average and standard error of the Main Conflict&#x2019;s covered brick taking percentage. In this case, the covered bricks were the E and G ones.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g003.tif"/>
</fig>
<p>The main conflict presented a situation where the robot could not see the best choice (the yellow bricks, G and E), which were instead visible to participants. At the same time, only the robot could see what lay behind the occlusion (C, the blue brick). This represented the second-best choice for the participant but one of the best choices in the robot&#x2019;s view (together with the other blue brick D, visible to the participant as well).</p>
<p>This conflict yielded two different hints for the SP-based robot and the NSP-based one. At that step of the game, participants had a high degree of uncertainty about the colour of the covered blue brick (C). Based on the participants&#x2019; point of view, it could have been green, black, yellow, blue or orange because, at this point of the task, the last brick taken was green. The robot did not suggest that covered brick yet; thus, it could be the same colour as the last brick taken (or one of the colours with a lower value than green). Thus, to the brick C, which is the <inline-formula id="inf31">
<mml:math id="m31">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> of this conflict, the participants associated a uniform probability distribution over the colours listed above.</p>
<p>The shared perception model aims at minimising the participants&#x2019; uncertainty about the covered object&#x2019;s characteristics. Following the model, the SP-robot indicated the blue brick visible to both agents (D). Indeed, the robot revealed that it could not perceive anything better than a blue brick: the covered brick C had necessarily a lower or equal value than the brick D: it could only be blue or orange. Pursuing the same uncertainty minimisation goal, if participants took D, iCub indicated the other blue brick (C) in the next move. Otherwise, if they took one of the yellow bricks (E-G), it indicated the visible blue (D) again. On the other hand, during NSP sessions, the robot decided only on its own knowledge and data structures, indicating at first the blue brick (C), which was hidden from the participants&#x2019; view. If participants took it, iCub then indicated the other blue brick (D); otherwise, if they took one of the yellow bricks (E-G) or the visible blue one (D), it then indicated the hidden brick (C) again.</p>
</sec>
<sec id="s3-2-2-2">
<title>3.2.2.2 The First Brick Conflict</title>
<p>This conflict occurred at the beginning of sessions based on the first setup; this means that each participant faced twice this type of conflict. The bricks&#x2019; configuration related to this conflict is shown in <xref ref-type="fig" rid="F4">Figure 4</xref> (right). As we can see from the figure, the core of the conflict was the red brick covered to the participants but visible to the robot (F). The red colour corresponds to the highest value on the scale, representing one of the best choices to start a tower. At the beginning of the task, the brick F could be of any colour with the same probability. Thus, if we assume that the brick F was the <inline-formula id="inf32">
<mml:math id="m32">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> of this conflict, the participants associated a uniform probability distribution over all the colours reported on the list.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Participants&#x2019; score in both SP and NSP conditions during the conflicts. <bold>(A)</bold> is referred to the Main Conflict, <bold>(B)</bold> to the Mid-Game Conflict, and <bold>(C)</bold> to the First Brick Conflict. We applied a Gaussian random noise (<italic>&#x3bc;</italic> &#x3d; 0, <italic>&#x3c3;</italic> &#x3d; 0.05 on both the <italic>x</italic> and <italic>y</italic> axis) to make all of them visible. It is also shown the average score with standard error.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g004.tif"/>
</fig>
<p>In the SP condition, the robot followed the model and selected the hint aimed at minimising the participants&#x2019; uncertainty about the covered object&#x2019;s characteristics by leveraging on the characteristics of the objects visible to both the participant and itself. Hence iCub indicated it (F) at first to reveal that the covered brick F could not be worse than the visible red brick (A), which means that it could be only either red or white (corresponding to the same value). If participants took that brick in the first round, iCub then indicated the visible red brick (A); otherwise, if participants took at first the visible red brick (A), it then indicated the hidden brick (F) again. Instead, during the NSP sessions, the robot indicated the red brick visible to both (A) firstly and afterwards it indicated the hidden red brick (F).</p>
</sec>
<sec id="s3-2-2-3">
<title>3.2.2.3 The Mid-game Conflict</title>
<p>This conflict occurred in the middle of game sessions based on the second setup; this means that each participant faced twice this type of conflict. An example of the bricks configuration during this conflict is shown in <xref ref-type="fig" rid="F5">Figure 5</xref> (right side). The core of the conflict was the green brick (M), hidden to the participant but visible to the robot. The colour green was associated with the highest value available on the table in that portion of the game. At this point of the game, participants had a high uncertainty about the colour of the covered green brick M: from the participants&#x2019; perspective, it could be red, white, green, black, yellow, blue or orange.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Average and standard error of the First Brick Conflict&#x2019;s covered brick taking percentage. In this case, the covered brick was the F one.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g005.tif"/>
</fig>
<p>The robot, guided by the SP model, aimed to inform participants that there was something interesting that they could not see. Hence, iCub indicated at first the green brick (M) hidden to the partner. This way, the robot revealed that the covered brick M was equal to or better than the visible ones, which means that it could be either red, white or green. The robot minimised the participants&#x2019; uncertainty about the covered object&#x2019;s characteristics through its communication signal. If participants took it, iCub then indicated the other green brick (N) visible to both; otherwise, if they took the visible green block (N) in the first round, iCub then indicated the hidden one (M) again. On the other hand, during NSP sessions, the robot indicated what looked best from its viewpoint, irrespective of the human point of view. In particular, it first indicated the green brick visible to both (N) and then the hidden green brick (M).</p>
</sec>
<sec id="s3-2-2-4">
<title>3.2.2.4 Step-by-step Task Simulation</title>
<p>For clarity, we provide here a simulation step by step of a session with the robot in both conditions. For simplicity, in these simulations, we consider that the participants always follow the robot&#x2019;s suggestions. We start from the SP-robot and the bricks setup 1 (<xref ref-type="fig" rid="F2">Figure 2</xref>). First, we encounter the configuration of the <italic>First Brick Conflict</italic>. Thus, the robot would suggest the brick F. Due to the presence of the brick A in the robot and the participant&#x2019;s shared perception; such a suggestion shows that the brick F (with a value of 10) is better (or no worse) than the brick A (with value 10) because otherwise, the robot would have suggested this latter. Then, the robot would suggest brick A, which is the best choice from its perspective. Then, it would suggest the bricks B and H (both with value 8) for the same reason. Finally, the configuration of the bricks becomes that of the <italic>Main Conflict</italic>. In this case, the robot would suggest the brick D: such a suggestion shows that the brick D (with a value of 4) is better (or no worse) than the brick C (with a value of 4), which is covered to the participant. This concludes the session. Now, we move to the NSP behaviour using the same setup. First, the robot would suggest brick A (with a value of 10), which is one of its best choice and the first brick according to its internal representation. Then, it would suggest the bricks F (with a value of 10 but covered to the participant), B, and H (both with a value of 8) in this order for the same reason. Lastly, the robot would suggest the brick C (with value 4), which is one of the best choices from its perspective and the first one according to its internal representation of the task but covered to the participant.</p>
</sec>
</sec>
<sec id="s3-2-3">
<title>3.2.3 Pilot Experiments</title>
<p>Before running the experiments with the robot, we performed a pilot study with eight colleagues in a human-human configuration. One participant took the role of suggester, while the other took the builder&#x2019;s role. The pilot aimed to study the nature of the signals used by people in exploiting SP mechanisms. Before starting the task, we asked participants to use only nonverbal communication.</p>
<p>Through videos, we noted that all participants used gaze cues to indicate the objects: sometimes, a movement with the eyebrows followed the gazing. The suggesters tried to attract the other participant&#x2019;s attention by establishing eye contact; then, they gazed at the candidate brick. At the end of this gazing exchange, the builders followed the suggesters&#x2019; hints. Sometimes, especially in the first trials, the builders asked for a confirmation by pointing or gazing at the object they wanted to take. The suggesters attempted to make the builders aware of their hidden bricks in every trial.</p>
</sec>
</sec>
<sec id="s3-3">
<title>3.3 Participants</title>
<p>We had 22 participants (9 males and 13 females) with an average age of 26.5 years (SD: 7.8). Two participants failed to understand the experiment instructions (<italic>i.e.</italic>, did not choose the highest valued brick as the first element of a tower) and were therefore discarded from the analysis. All participants gave written informed consent before participating and received a fixed refund of <italic>&#xa3;</italic>15. The experimental protocol was approved by Regione Liguria&#x2019;s regional ethic committee.</p>
</sec>
<sec id="s3-4">
<title>3.4 Measures</title>
<p>During the experiment, we collected some behavioural measures such as the participants&#x2019; score, the number of times participants followed the robot&#x2019;s hints, the time needed to take a brick (calculated as the time between the beginning of the action phase and the grip of the brick), and what bricks the robot indicated. In particular, we focused on the conflicts where the mechanism of shared perception could have had an impact.</p>
<sec id="s3-4-1">
<title>3.4.1 Questionnaires</title>
<p>We submitted questionnaires to participants before the beginning of the experiment and after each interactive session with the robot. Before the experiment, we asked participants to reply to the Seventeen-Item Scale for Robotic Needs (SISRN) questionnaire (<xref ref-type="bibr" rid="B33">Manzi et al., 2021</xref>). We chose the SISRN questionnaire to know what participants thought about generic robots&#x2019; capabilities. Furthermore, both before the experiment and after each experimental session, we submitted to participants the Godspeed (<xref ref-type="bibr" rid="B5">Bartneck et al., 2009</xref>) and the Inclusion of Other in the Self (IOS) questionnaires (<xref ref-type="bibr" rid="B2">Aron et al., 1992</xref>). The questionnaire web page contained a video of the iCub robot<xref ref-type="fn" rid="fn2">
<sup>2</sup>
</xref>: we showed it to participants to provide them with enough information about the robot before a real interaction with it. The Godspeed questionnaire was chosen to collect participants&#x2019; impressions about robot&#x2019;s <italic>anthropomorphism</italic>, <italic>animacy</italic>, <italic>likeability</italic>, and <italic>perceived intelligence</italic> before and after the interactive sessions. The IOS questionnaire was used to understand if participants felt closer to the robot in some of the two experimental conditions.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>4 Results</title>
<p>The collaborative game with the iCub robot was characterised by perceptual asymmetries between the human-builder and the robot-suggester. This asymmetry was particularly critical in certain choices during the game (the <italic>conflicts</italic>, see <xref ref-type="sec" rid="s3">Section 3</xref>), where the best brick to take could differ between the two agents&#x2019; perspectives. We focus our analysis on these specific moments in the game: the <italic>Main Conflict</italic> (<xref ref-type="sec" rid="s3-2-2-1">Section 3.2.2.1</xref>), the <italic>First Brick Conflict</italic> (<xref ref-type="sec" rid="s3-2-2-2">Section 3.2.2.2</xref>) and the <italic>Mid-Game Conflict</italic> (<xref ref-type="sec" rid="s3-2-2-3">Section 3.2.2.3</xref>), to assess the impact of the robot&#x2019;s suggestions to the partner, when they are based on a shared perception mechanism (SP) or not (NSP). The former considers the brick visibility to the human in selecting which brick to suggest, whereas the NSP-robot just relies on its internal representation of the task.</p>
<p>First, we checked whether the configuration of the bricks affected participants&#x2019; performances. For this purpose, we split the total scores obtained in each bricks&#x2019; configuration for each robot mode. We conducted paired t-tests and we found no significant differences between the scores obtained with the two setups, with the robot in SP mode (<italic>paired t-test</italic> <italic>t</italic> (19) &#x3d; &#x2212;0.567, <italic>p</italic> &#x3d; 0.578); and in NSP mode (<italic>paired t-test</italic> <italic>t</italic> (19) &#x3d; 0.837, <italic>p</italic> &#x3d; 0.413)).</p>
<sec id="s4-1">
<title>4.1 The Main Conflict</title>
<p>
<xref ref-type="fig" rid="F3">Figure 3</xref> (left) shows the average percentage of time participants picked one of the yellow bricks (either E or G)&#x2014;covered to the robot&#x2019;s sight. These corresponded to the best choice for the participant, as these bricks had the highest value. During the SP sessions, more than 80% of the participants took a yellow brick; conversely, only 40% of the participants took one of the current best bricks during NSP sessions (<italic>&#x3bc;</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 84.21, <italic>SE</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 5.99; <italic>&#x3bc;</italic>
<sub>
<italic>nsp</italic>
</sub> &#x3d; 38.09, <italic>SE</italic>
<sub>
<italic>nsp</italic>
</sub> &#x3d; 7.58). Instead, they took the brick iCub was indicating them: the blue brick C. The difference between the two conditions is significant (<italic>two-tailed z-test,</italic> <italic>z</italic> &#x3d; 2.21, <italic>p</italic> &#x3d; 0.02).</p>
<p>There was no difference in the time employed to pick the brick in the two conditions. The timing was computed only for participants who picked the covered block, hence on the percentages of participants reported in <xref ref-type="fig" rid="F3">Figure 3</xref>. <xref ref-type="fig" rid="F4">Figure 4A</xref> shows the scores of participants collected during the Main Conflict in both the experimental conditions. As we can see from the plot, around 40% of the participant scored more during SP sessions than during the NSP ones.</p>
</sec>
<sec id="s4-2">
<title>4.2 The First Brick Conflict</title>
<p>
<xref ref-type="fig" rid="F5">Figure 5</xref> (left side) shows the percentage of times in which participants picked the brick covered to them (F). As we can see from the figure, during the SP sessions, participants resolved the conflict properly almost all the time: more than 30% of the time took the brick F as their first move (<italic>&#x3bc;</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 31.57, <italic>SE</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 5.4), and around 65% of the time took it as their second move (<italic>&#x3bc;</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 63.15, <italic>SE</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 5.6). The remaining 5% of the time, participants did not take the brick F; instead, they preferred to take a green brick. All the participants who did not take the brick F firstly during the SP sessions took the brick A as their first move. On the other hand, during the NSP sessions, nearly 50% of the time, participants could take the brick F (<italic>&#x3bc;</italic>
<sub>
<italic>nsp</italic>
</sub> &#x3d; 47.61, <italic>SE</italic>
<sub>
<italic>nsp</italic>
</sub> &#x3d; 5.7) While the remaining have opted to take a green one. All participants took the brick A as their first move during the NSP sessions. The difference between the two conditions is significant (<italic>two-tailed z-test</italic>, <italic>z</italic> &#x3d; 3.81, <italic>p</italic> &#x3c; 0.001), while the difference between the percentage referred to the covered brick taken at move two was not.</p>
<p>Also, there were no significant differences in the time employed to pick the hidden brick between conditions for this conflict. <xref ref-type="fig" rid="F4">Figure 4C</xref> shows the scores participants collected during the First Brick Conflict in both the experimental conditions. As we can see from the plot, around 50% of the participant scored more during SP sessions than during the NSP ones.</p>
</sec>
<sec id="s4-3">
<title>4.3 The Mid-game Conflict</title>
<p>
<xref ref-type="fig" rid="F6">Figure 6</xref> shows the percentage of times in which participants took the covered green brick M. Participants behaved quite the same as during the previous conflict: during the SP sessions, 60% of the time, participants took the brick M as their first move (<italic>&#x3bc;</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 60, <italic>SE</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 5.4), and the remaining 40% of the time they took it in the next move (<italic>&#x3bc;</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 40, <italic>SE</italic>
<sub>
<italic>sp</italic>
</sub> &#x3d; 5.6). Thus, all participants could resolve the <italic>Mid-Game Conflict</italic> properly during the SP sessions. On the other hand, during the NSP sessions, only 60% of the time, participants could take the brick M (<italic>&#x3bc;</italic>
<sub>
<italic>nsp</italic>
</sub> &#x3d; 61.9, <italic>SE</italic>
<sub>
<italic>nsp</italic>
</sub> &#x3d; 5.7), while the remaining 40% opted to take a yellow one. As happened in the previous conflict, all participants took the brick N as their first move during the NSP sessions. The difference between the two conditions is significant (<italic>two-tailed z-test</italic>, <italic>z</italic> &#x3d; 4.03, <italic>p</italic> &#x3c; 0.001), while the difference between the percentage referred to the covered brick taken at move two was not significant.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Average and standard error of the Mid-Game Conflict&#x2019;s covered brick taking percentage. In this case, the covered brick was the M one.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g006.tif"/>
</fig>
<p>Also in this case, the time employed to pick the hidden brick did not differ between conditions. <xref ref-type="fig" rid="F4">Figure 4B</xref> shows the scores participants collected during the Mid-Game Conflict in both the experimental conditions. As we can see from the plot, around 40% of the participant scored more during SP sessions than during the NSP ones.</p>
</sec>
<sec id="s4-4">
<title>4.4 Questionnaires</title>
<p>
<xref ref-type="fig" rid="F7">Figure 7</xref> shows the average and the standard error of their answers to the IOS image. As we can see, there is a difference between the answers given before the experiment and the ones given after the interactive sessions, regardless of the robot mode. These difference turned out to be statistically significant for both PRE-SP and PRE-NSP groups (<italic>repeated measures ANOVA</italic> <italic>F</italic> (21) &#x3d; 6.82, <italic>p</italic> &#x3d; 0.01 and <italic>F</italic> (21) &#x3d; 6.6, <italic>p</italic> &#x3d; 0.01, respectively). Nonetheless, we found no significant differences between the answers given after the SP sessions and those given after the NSP sessions.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Average and standard error of participants&#x2019; answers to the IOS image.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g007.tif"/>
</fig>
<p>Similarly, there were no significant differences in the Godspeed questionnaire (<xref ref-type="fig" rid="F8">Figure 8</xref>) scores among the different phases (PRE, POST-SP, POST-NSP). We found no correlations between the answers given to the SISRN and the other questionnaires or the participants&#x2019; performance during the conflicts.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Average and standard error of participants&#x2019; answers to the Godspeed questionnaire.</p>
</caption>
<graphic xlink:href="frobt-09-733954-g008.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<title>5 Discussion</title>
<p>In this work, we assessed whether a robot attempting to establish a shared perception with its human partners is better evaluated and ensures a more effective collaboration. The results suggested that shared perception leads to higher performances in the task. A robot considerate of the partners&#x2019; viewpoints and goals can facilitate selecting the best action.</p>
<p>The proposed mathematical model for shared perception, which guided iCub selection of the hints to be provided to the partner in the SP condition, seemed to be effective. The robot indicated with its gaze the brick that minimised participants&#x2019; uncertainty about the properties of the objects hidden from their view. This did not imply that the players picked the suggested object right away. Instead, it successfully ensured that participants took into account also the covered bricks in their reasoning, thanks to the implicit knowledge shared by the robot. As a result, in the SP conditions, the vast majority of the time, participants did not miss any of the highest value elements while building their Lego tower.</p>
<p>We hypothesised that participants would use a recursive (or second-order) Theory of Mind (ToM) when reasoning about the robot&#x2019;s suggestions: e.g., participants knew that the robot knew that they could not perceive the bricks covered to them (for example, the brick F in <xref ref-type="fig" rid="F2">Figure 2A</xref>). Indeed, several models have been presented that make use of recursive ToM (<xref ref-type="bibr" rid="B43">Pynadath and Marsella, 2005</xref>; <xref ref-type="bibr" rid="B9">Bosse et al., 2011</xref>; <xref ref-type="bibr" rid="B14">de Weerd et al., 2017</xref>). In particular, <xref ref-type="bibr" rid="B9">Bosse et al. (2011)</xref> presented a model for multilevel ToM based on BDI concepts that they tested in three different case studies: social manipulation, predators&#x2019; behaviour, and emergent soap stories. Instead, <xref ref-type="bibr" rid="B43">Pynadath and Marsella, (2005)</xref> proposed PsychSim, a multi-agent simulation tool for modelling interaction and influence that makes use of a recursive model of other agents. Moreover, <xref ref-type="bibr" rid="B14">de Weerd et al. (2017)</xref> proved that agents using second-order ToM lead to higher effectiveness than agents capable of only first-order ToM. Moreover, more importantly for us, they discovered that people spontaneously use recursive ToM when their partner is capable of second-order ToM as well.</p>
<p>Instead, without shared perception, the robot hints were less informative. It followed the rationale of indicating the highest valued brick from its perspective. In particular, the robot relied only on its internal representation of the task, which led it to behave during conflicts in the opposite way to the SP setting. Thus, we ensured the maximum difference between the SP and NSP behaviours. This led to errors, in particular when the best bricks were not visible to the robot (<italic>Main Conflict</italic>). However, also in situations in which the asymmetry in perception was not so critical (<italic>First-Brick</italic> and <italic>Mid-Game</italic> conflicts), as the best blocks were hidden to participants but not to the robot, its hints were less effective. A significantly lower percentage of players gathered all the best bricks in the NSP condition than in the SP.</p>
<p>We speculate that during SP sessions, participants built a more precise representation of the robot&#x2019;s perspective than during NSP sessions. In the former sessions, they better understood that when the robot was indicating a brick visible to all, it was because it had nothing better to suggest. Consequently, they could resolve the conflict correctly and pick the best option, even when they could not see it directly. In NSP sessions, the participants did not have enough information to understand the reasons guiding the robot suggestion and gauge their validity. As a result, they blindly followed the robot&#x2019;s hints in some cases. In particular, in 50% of the cases in the <italic>First Brick conflict</italic> and 60% of those in the <italic>Mid-Game conflict</italic>, participants&#x2019; second move followed iCub&#x2019;s indication toward the covered object even if there was very limited information about what brick the robot was indicating to them. In those conflicts, this choice was still valid, as the suggested brick had a value as high as the visible ones. In the <italic>Main conflict</italic> instead, the excessive trust led in about 60% of the cases to a sub-optimal choice. In other cases, the lack of understanding of the robot&#x2019;s motives led participants to disregard the robots&#x2019; suggestions, missing out on valuable blocks hidden from participants&#x2019; view. Indeed, about 40%&#x2013;50% of the time, participants in one of the conflicts went on picking visible blocks, whereas the robot was pointing at the highest one behind an occlusion. This means that the absence of a reliable common ground makes people unable to fully exploit the collaboration.</p>
<p>An interesting reflection about participants&#x2019; trust toward the robot can be afforded by the <italic>First Brick</italic> task configuration. In this conflict, participants&#x2019; choice could not be driven by the outcome of previous moves, as it regarded the first move of a game. Furthermore, the games previously played with the robot should not have had any influence since, at the beginning of each session, the experimenter instructed the participants that a new program controlled the robot. In this case, for the game&#x2019;s first move, two bricks of the highest possible value are present on the scene, one visible and one hidden from the participants&#x2019; view. Since there were no bricks with higher value in the game, the most rational first choice would have been picking the visible highest value brick. Despite this, when the robot indicated the hidden item (i.e., in the SP condition) in around 30% of cases, participants opted to pick that instead. Overall, this result suggests that for a good portion of participants, the robot indication was sufficient to make them abandon a sure optimal choice, to pick something unknown. We ascribe the choice to participants&#x2019; proneness to trust (or better over-trust) the robot, a phenomenon often observed in interactions with robots. An alternative explanation is that this choice was driven by an attempt to behave kindly toward the humanoid. Recent evidence points out that also in HRI, mechanisms like reciprocity play a role&#x2014;with humans overtly following the robot&#x2019;s advice despite disagreeing with it, to ensure its future benevolence, as it happens between humans (<xref ref-type="bibr" rid="B58">Zonca et al., 2021b</xref>).</p>
<p>It is also relevant to notice that in the experiment, there was also an analogous configuration in which the iCub indicated first the visible highest value brick and then the hidden highest value brick (i.e., in the NSP condition). In this case, the proportion of participants who followed the second indication and managed to pick also the hidden high-value item reached about 60% of the cases. Although higher, this implies that in about 40% of the situations, seeing that the robot indications were meaningful in the first move was not enough to induce participants to trust its indications in the next move. In other words, the fact that iCub indicated the best brick among those the participants could perceive did not convince participants to select the item it indicated when it was not visible. Such a lack of trust led those participants to miss a relevant opportunity. Apparently, the less trusting participants did not receive enough information about what drove the robot&#x2019;s suggestions to follow them. These findings underline the importance of the robot selecting its hints properly by revealing as much as possible to the partner its own understanding of the environment. By doing so, the robot can avoid, on the one hand, over-trust and the other excessive lack of trust.</p>
<p>Despite the difference in overall performances between SP and NSP sessions, we registered no differences between the answers to the post-session questionnaires. This contradiction shows us SP mechanisms&#x2019; essential and implicit nature: people exploited SP mechanisms, but they were not fully aware of them. In fact, participants reported no particular differences between the robot&#x2019;s behaviours according to the different experimental behaviours.</p>
<p>In our experiments, we defined two robot behaviours that resulted in being at the antipodes, with the NSP robot actually being a <italic>anti-</italic>SP mode. A fairer NSP behaviour would provide randomised choices from the robot. However, such a less controlled design (e.g., with a robot selecting randomly in case of conflict in the NSP condition) would have required a much larger sample to enable reliable testing of all the possible conditions. To explore this option, we ran a series of simulations.</p>
<p>More precisely, we conducted 1,000 simulations in which we tested a fair NSP robot mode using the results of users&#x2019; behaviour that we obtained in our experiments. In such simulations, in case of multiple best options, we let the NSP robot choose randomly between them. On the other hand, we defined the simulated users&#x2019; behaviour based on the experimental results. This means that, in NSP-mode, when the hidden brick is suggested as a second move, users resolve the conflict 50% of the time during the First-Brick conflict, 60% during the Mid-Game conflict, and 40% during the Main conflict. The same applies for the SP-mode: 100% during the Mid-Game conflict, 95% during the First-Brick conflict and 83% during the Main conflict. Then, we analysed how many times the simulated users resolved the conflicts. <xref ref-type="table" rid="T2">Table 2</xref> shows the statistics about the resolution of such simulated conflicts. As we can see, we obtained results comparable with those derived from the real experiments.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Results of the simulation of the experiment with the fair NSP robot behaviour. The statistics are applied on the simulated data.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left"/>
<th colspan="2" align="center">% of Resolution</th>
<th colspan="2" align="center">
<italic>Two-tailed z-test</italic>
</th>
</tr>
<tr>
<th align="center">SP</th>
<th align="center">NSP</th>
<th align="center">
<italic>z</italic>
</th>
<th align="center">
<italic>p</italic>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Main Conflict</td>
<td align="char" char=".">80.2</td>
<td align="char" char=".">64.3</td>
<td align="char" char=".">8.97</td>
<td align="char" char=".">&#x3c;0.001</td>
</tr>
<tr>
<td align="left">First-Brick Conflict</td>
<td align="char" char=".">90.8</td>
<td align="char" char=".">71.4</td>
<td align="char" char=".">11.08</td>
<td align="char" char=".">&#x3c;0.001</td>
</tr>
<tr>
<td align="left">Mid-Game Conflict</td>
<td align="char" char=".">100</td>
<td align="char" char=".">59</td>
<td align="char" char=".">22.7</td>
<td align="char" char=".">&#x3c;0.001</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In this work, we considered a simplified scenario capturing the main elements of intriguing HRI collaborative situations: asymmetries in perception and objects having properties that make them more or less suitable to be used next. To focus on these central aspects, we opted for simplifying the settings: objects have a single relevant property (their colour); and the communication is minimised: we allowed only gazing to indicate the proposed object. However, in principle, the model we propose could be generalised to more complex settings, as far as the robot can 1) estimate the impact of its suggestions on the uncertainty of the partner&#x2019;s representation, 2) have a measure of relative (to the task) importance to assign to each object&#x2019;s characteristic, and 3) have richer communication capabilities. We refer the reader to <xref ref-type="sec" rid="s5-1">Section 5.1</xref>, where we better discuss what a more generalised approach could concern. In fact, we can easily map our experimental task into the more complex assembly task that we give as an example in <xref ref-type="sec" rid="s1">Section 1</xref>. As long as the person and the robot are aligned on the assembly step, the robot can use our model to choose what to say or which object to pass in order to maximise the flow of information. Let us imagine a scenario in which two sets of screws are appropriate for the assembly, one well visible to both the robot and the participant and one partially occluded to the latter. Through our model, the robot can decide it is worth suggesting or handing over the screws from the semi-occluded set to maximise both performance and the person&#x2019;s knowledge about the available tools in the chaotic environment. This way, the robot can lower the person&#x2019;s uncertainty about objects not included&#x2014;or partially included&#x2014;in their perception.</p>
<p>The scenario in which humans and robots have a misaligned perception of the shared environment has been addressed by <xref ref-type="bibr" rid="B13">Chai et al., 2014</xref>. Their work focused on allowing a robot to acquire knowledge about common ground via collaborative dialogue with its human partner: the more the communication proceeds, the more the robot can improve its internal representation of the shared environment. The main aim of their approach was to help the robot in lowering its uncertainty about semi-occluded objects. Our work addresses the task proposed by <xref ref-type="bibr" rid="B13">Chai et al., 2014</xref> from the opposite perspective by proposing the robot as a suggester, helping the human in resolving asymmetries in the shared environment. Moreover, we explored interactions that do not involve the use of speech. We addressed the shared perception problem using more primitive communicative ways, thus without considering the language.</p>
<p>To conclude, it is essential to note that a direct prediction of our model is that if two agents share nothing in their common awareness spaces, then it is impossible to obtain shared perception. Hence, we can claim that establishing common ground is key to pursuing a collaborative HRI task. Thus, it becomes crucial to building shared knowledge about both the environment&#x2014;which can present asymmetries in perception, as in our experiments&#x2014;and the collaborating agents&#x2019; mental states in terms of objectives, beliefs and intentions.</p>
<sec id="s5-1">
<title>5.1 Limitations</title>
<p>The first limitation of our work regards the simplifications we made in our experimental setting. In particular, the interaction was constrained during the experiments, and only nonverbal communication was allowed. We allowed only a turn-based speechless communication to maintain careful control of the information exchange with the participants and to ensure that all participants faced the &#x201c;conflict&#x201d; instances with the same amount of information. This is obviously a simplification: communication between partners is usually more complicated than this. However, we believe that there are forms of real interactions which are not too far from the settings we proposed, such as some turn-based card games (e.g., &#x201c;Briscola&#x201d; we mentioned above) and assembling tasks, where the context constrains the interaction.</p>
<p>Other limitations regard some assumptions of the model. In particular, the model assumes that the objects&#x2019; features do not change over time. This limitation can be overcome by introducing memory-based and/or probabilistic measures of uncertainty regarding the objects and their characteristics. In particular, for what regards the object the robot is aware of but that not perceive anymore, such a measure of uncertainty could model how stronger the robot believes the object is still where it remembers (e.g., a measure that could worsen over time). A similar argument could be applied to mutable characteristics of particular objects. A fluid measure of uncertainty can manage how much the robot is sure about an object&#x2019;s characteristics (e.g., the shape of a partially-occluded object could be challenging to understand but easily guessable).</p>
<p>Furthermore, the robot already knew which bricks the participants could perceive and which ones they could not. To make the architecture more autonomous, we could use perspective-taking to allow the robot to automatically infer the objects belonging to the partner&#x2019;s personal perception (<xref ref-type="bibr" rid="B18">Fischer and Demiris, 2020</xref>). Also occlusions (the obstacles in our experiment) could be detected through perspective-taking algorithms.</p>
<p>Finally, we assumed that participants would accept the suggestions the robot gave as the bests to achieve the goal of the task. Before starting the experiment, we presented the robot as a collaborator but, in general, we should take into account the level of trust people have towards robots (<xref ref-type="bibr" rid="B57">Zonca et al., 2021a</xref>).</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s6">
<title>6 Conclusion</title>
<p>We investigated the role of Shared Perception (SP) in Human-Robot Interaction (HRI). In particular, with the present work, we aimed to 1) understand whether and how humans would exploit SP mechanisms with a robot during a cooperative game characterised by an asymmetry in the perception of the environment and 2) propose a computational model for SP. Indeed, we designed a mathematical model for cooperative SP. We tested it via a user study in which the robot and participants had to collaborate to build a tower with LEGO bricks. Some of those were visible by both agents, others were covered to the participants, and the remaining were covered to the robot. We designed our experiment to elicit critical moments that we called <italic>conflicts</italic>, and we investigated the differences between a robot with SP (SP-iCub) and a robot unable to use SP mechanisms (NSP-iCub) when the perceptions of the interacting agents differed.</p>
<p>Our results show that humans can potentially exploit SP mechanisms with robots as they do with other humans. For all conflicts, SP-iCub resulted to be more informative than NSP-iCub. Indeed, with the former, people could correctly resolve conflicts most of the time. Conversely, only a minority of the participants could make the best move in such critical situations with the latter. However, despite the clear difference between the experimental conditions and the resulting strategies that we registered, our participants did not report perceiving the robot&#x2019;s behaviours differently. This effect highlights the implicit nature of SP: people exploit SP mechanisms but are unaware of their decision process.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusion of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s8">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by Comitato Etico Regione Liguria. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>All the authors contributed to conception and design of the study. MM designed the model, that has been refined with FR and AS. MM performed the experiments and the statistical analysis. MM wrote the first draft of the manuscript. FR and AS revised the manuscript. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec id="s10">
<title>Funding</title>
<p>This work has been supported by a Starting Grant from the European Research Council (ERC) under the European Union&#x2019;s Horizon 2020 research and innovation programme. G.A. No 804388, wHiSPER.</p>
</sec>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
<p>The handling editor is currently co-organizing a Research Topic with one of the authors AS, and confirms the absence of any other collaboration.</p>
</sec>
<sec sec-type="disclaimer" id="s12">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://en.wikipedia.org/wiki/Briscola">https://en.wikipedia.org/wiki/Briscola</ext-link>.</p>
</fn>
<fn id="fn2">
<label>2</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://www.youtube.com/watch?v=3N1oCMwtz8w">https://www.youtube.com/watch?v&#x3d;3N1oCMwtz8w</ext-link>.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Admoni</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Scassellati</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Social Eye Gaze in Human-Robot Interaction: a Review</article-title>. <source>J. Human-Robot Interact.</source> <volume>6</volume>, <fpage>25</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.5898/JHRI.6.1.Admoni</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aron</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Aron</surname>
<given-names>E. N.</given-names>
</name>
<name>
<surname>Smollan</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>1992</year>). <article-title>Inclusion of Other in the Self Scale and the Structure of Interpersonal Closeness</article-title>. <source>J. personality Soc. Psychol.</source> <volume>63</volume>, <fpage>596</fpage>&#x2013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.63.4.596</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Arslan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Hohenberger</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Verbrugge</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>The Development of Second-Order Social Cognition and its Relation with Complex Language Understanding and Memory</article-title>,&#x201d; in <conf-name>Proceedings of the Annual Meeting of the Cognitive Science Society</conf-name>. </citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Baron-Cohen</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1997</year>). <source>Mindblindness: An Essay on Autism and Theory of Mind</source>. <publisher-loc>Cambridge, Massachusetts</publisher-loc>: <publisher-name>MIT press</publisher-name>. </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bartneck</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kuli&#x107;</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Croft</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Zoghbi</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Measurement Instruments for the Anthropomorphism, Animacy, Likeability, Perceived Intelligence, and Perceived Safety of Robots</article-title>. <source>Int J Soc Robotics</source> <volume>1</volume>, <fpage>71</fpage>&#x2013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-008-0001-3</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Benninghoff</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Kulms</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hoffmann</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kr&#xe4;mer</surname>
<given-names>N. C.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Theory of Mind in Human-Robot-Communication: Appreciated or Not?</article-title> <source>Kognitive Syst.</source> <volume>2013</volume> (<issue>1</issue>). <pub-id pub-id-type="doi">10.17185/duepublico/31357</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Berlin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gray</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Thomaz</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Breazeal</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Perspective Taking: An Organizing Principle for Learning in Human-Robot Interaction</article-title>. <source>AAAI</source> <volume>2</volume>, <fpage>1444</fpage>&#x2013;<lpage>1450</lpage>. <pub-id pub-id-type="doi">10.5555/1597348.1597418</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bianco</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ognibene</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Functional Advantages of an Adaptive Theory of Mind for Robotics: a Review of Current Architectures</article-title>,&#x201d; in <conf-name>2019 11th Computer Science and Electronic Engineering (CEEC)</conf-name>, <fpage>139</fpage>&#x2013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.1109/CEEC47804.2019.8974334</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bosse</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Memon</surname>
<given-names>Z. A.</given-names>
</name>
<name>
<surname>Treur</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A Recursive Bdi Agent Model for Theory of Mind and its Applications</article-title>. <source>Appl. Artif. Intell.</source> <volume>25</volume>, <fpage>1</fpage>&#x2013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1080/08839514.2010.529259</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boucher</surname>
<given-names>J.-D.</given-names>
</name>
<name>
<surname>Pattacini</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Lelong</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bailly</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Elisei</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Fagel</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>I Reach Faster When I See You Look: Gaze Effects in Human-Human and Human-Robot Face-To-Face Cooperation</article-title>. <source>Front. Neurorobot.</source> <volume>6</volume>, <fpage>3</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2012.00003</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Breazeal</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Berlin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Brooks</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gray</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Thomaz</surname>
<given-names>A. L.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Using Perspective Taking to Learn from Ambiguous Demonstrations</article-title>. <source>Robotics Aut. Syst.</source> <volume>54</volume>, <fpage>385</fpage>&#x2013;<lpage>393</lpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2006.02.004</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brown-Schmidt</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Heller</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Perspective-taking during Conversation</article-title>. <source>Oxf. Handb. Psycholinguist.</source> <volume>551</volume>, <fpage>548</fpage>&#x2013;<lpage>572</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordhb/9780198786825.013.23</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chai</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>She</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ottarson</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Littley</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). &#x201c;<article-title>Collaborative Effort towards Common Ground in Situated Human-Robot Dialogue</article-title>,&#x201d; in <conf-name>2014 9th ACM/IEEE International Conference on Human-Robot Interaction (HRI) (IEEE)</conf-name>, <fpage>33</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1145/2559636.2559677</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Weerd</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Verbrugge</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Verheij</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Negotiating with Other Minds: the Role of Recursive Theory of Mind in Negotiation with Incomplete Information</article-title>. <source>Auton. Agent Multi-Agent Syst.</source> <volume>31</volume>, <fpage>250</fpage>&#x2013;<lpage>287</lpage>. <pub-id pub-id-type="doi">10.1007/s10458-015-9317-1</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Devin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Alami</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>An Implemented Theory of Mind to Improve Human-Robot Shared Plans Execution</article-title>,&#x201d; in <conf-name>2016 11th ACM/IEEE International Conference on Human-Robot Interaction (HRI) (IEEE)</conf-name>, <fpage>319</fpage>&#x2013;<lpage>326</lpage>. <pub-id pub-id-type="doi">10.1109/HRI.2016.7451768</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fischer</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Jensen</surname>
<given-names>L. C.</given-names>
</name>
<name>
<surname>Kirstein</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Stabinger</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Erkent</surname>
<given-names>&#xd6;.</given-names>
</name>
<name>
<surname>Shukla</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). &#x201c;<article-title>The Effects of Social Gaze in Human-Robot Collaborative Assembly</article-title>,&#x201d; in <conf-name>International Conference on Social Robotics</conf-name> (<publisher-name>Springer</publisher-name>), <fpage>204</fpage>&#x2013;<lpage>213</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-25554-5_21</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fischer</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Demiris</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Markerless Perspective Taking for Humanoid Robots in Unconstrained Environments</article-title>,&#x201d; in <conf-name>2016 IEEE International Conference on Robotics and Automation (ICRA) (IEEE)</conf-name>, <fpage>3309</fpage>&#x2013;<lpage>3316</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA.2016.7487504</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fischer</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Demiris</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Computational Modeling of Embodied Visual Perspective Taking</article-title>. <source>IEEE Trans. Cogn. Dev. Syst.</source> <volume>12</volume>, <fpage>723</fpage>&#x2013;<lpage>732</lpage>. <pub-id pub-id-type="doi">10.1109/TCDS.2019.2949861</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Flavell</surname>
<given-names>J. H.</given-names>
</name>
</person-group> (<year>1977</year>). &#x201c;<article-title>The Development of Knowledge about Visual Perception</article-title>,&#x201d; in <source>Nebraska Symposium on Motivation</source> (<publisher-name>University of Nebraska Press</publisher-name>). </citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fussell</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Setlock</surname>
<given-names>L. D.</given-names>
</name>
<name>
<surname>Parker</surname>
<given-names>E. M.</given-names>
</name>
</person-group> (<year>2003</year>). &#x201c;<article-title>Where Do Helpers Look?</article-title>,&#x201d; in <conf-name>CHI&#x2019;03 Extended Abstracts on Human Factors in Computing Systems</conf-name>, <fpage>768</fpage>&#x2013;<lpage>769</lpage>. <pub-id pub-id-type="doi">10.1145/765891.765980</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goodie</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Doshi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>D. L.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Levels of Theory-Of-Mind Reasoning in Competitive Games</article-title>. <source>J. Behav. Decis. Mak.</source> <volume>25</volume>, <fpage>95</fpage>&#x2013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1002/bdm.717</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>G&#xf6;r&#xfc;r</surname>
<given-names>O. C.</given-names>
</name>
<name>
<surname>Rosman</surname>
<given-names>B. S.</given-names>
</name>
<name>
<surname>Hoffman</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Albayrak</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Toward Integrating Theory of Mind into Adaptive Decision-Making of Social Robots to Understand Human Intention</article-title>,&#x201d; in <conf-name>Workshop on the Role of Intentions in Human-Robot Interaction at the International Conference on Human-Robot Interaction</conf-name>, <conf-loc>Vienna, Austria</conf-loc>, <conf-date>6 March 2017</conf-date>. </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Griffin</surname>
<given-names>Z. M.</given-names>
</name>
<name>
<surname>Bock</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>What the Eyes Say about Speaking</article-title>. <source>Psychol. Sci.</source> <volume>11</volume>, <fpage>274</fpage>&#x2013;<lpage>279</lpage>. <pub-id pub-id-type="doi">10.1111/1467-9280.00255</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hanna</surname>
<given-names>J. E.</given-names>
</name>
<name>
<surname>Brennan</surname>
<given-names>S. E.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Speakers&#x27; Eye Gaze Disambiguates Referring Expressions Early during Face-To-Face Conversation</article-title>. <source>J. Mem. Lang.</source> <volume>57</volume>, <fpage>596</fpage>&#x2013;<lpage>615</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2007.01.008</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hayhoe</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ballard</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Eye Movements in Natural Behavior</article-title>. <source>Trends cognitive Sci.</source> <volume>9</volume>, <fpage>188</fpage>&#x2013;<lpage>194</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2005.02.009</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hiatt</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Trafton</surname>
<given-names>J. G.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Accommodating Human Variability in Human-Robot Teams through Theory of Mind</article-title>,&#x201d; in <conf-name>Twenty-Second International Joint Conference on Artificial Intelligence</conf-name>. <pub-id pub-id-type="doi">10.5591/978-1-57735-516-8/IJCAI11-345</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Demiris</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Perceptual Perspective Taking and Action Recognition</article-title>. <source>Int. J. Adv. Robotic Syst.</source> <volume>2</volume>, <fpage>32</fpage>. <pub-id pub-id-type="doi">10.5772/5775</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Demiris</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Visuo-Cognitive Perspective Taking for Action Recognition (AISB)</article-title>. <source>Int. J. Adv. Robot. Syst.</source> <volume>2</volume> (<issue>4</issue>), <fpage>32</fpage>. <pub-id pub-id-type="doi">10.5772/5775</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kennedy</surname>
<given-names>W. G.</given-names>
</name>
<name>
<surname>Bugajska</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Marge</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Adams</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Fransen</surname>
<given-names>B. R.</given-names>
</name>
<name>
<surname>Perzanowski</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>Spatial Representation and Reasoning for Human-Robot Collaboration</article-title>. <source>AAAI</source> <volume>7</volume>, <fpage>1554</fpage>&#x2013;<lpage>1559</lpage>. <pub-id pub-id-type="doi">10.5555/1619797.1619894</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kiesler</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Fostering Common Ground in Human-Robot Interaction</article-title>,&#x201d; in <conf-name>ROMAN 2005. IEEE International Workshop on Robot and Human Interactive Communication</conf-name>, <fpage>729</fpage>&#x2013;<lpage>734</lpage>. <pub-id pub-id-type="doi">10.1109/ROMAN.2005.1513866</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Sha</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Breazeal</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>A Bayesian Theory of Mind Approach to Nonverbal Communication</article-title>,&#x201d; in <conf-name>2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</conf-name>, <fpage>487</fpage>&#x2013;<lpage>496</lpage>. <pub-id pub-id-type="doi">10.1109/HRI.2019.8673023</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Leslie</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Hirschfeld</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Gelman</surname>
<given-names>S. A.</given-names>
</name>
</person-group> (<year>1994</year>). <source>Mapping the Mind: Domain Specificity in Cognition and Culture</source>. <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. <pub-id pub-id-type="doi">10.1017/CBO9780511752902.009</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manzi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sorgente</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Massaro</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Villani</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Di Lernia</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Malighetti</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Emerging Adults&#x2019; Expectations about the Next Generation of Robots: Exploring Robotic Needs through a Latent Profile Analysis</article-title>. <source>Cyberpsychology, Behav. Soc. Netw.</source> <volume>24</volume> (<issue>5</issue>) <pub-id pub-id-type="doi">10.1089/cyber.2020.0161</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marchetti</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Manzi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Itakura</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Massaro</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Theory of Mind and Humanoid Robots from a Lifespan Perspective</article-title>. <source>Z. f&#xfc;r Psychol.</source> <volume>226</volume>, <fpage>98</fpage>&#x2013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1027/2151-2604/a000326</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mavridis</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>A Review of Verbal and Non-verbal Human-Robot Interactive Communication</article-title>. <source>Robotics Aut. Syst.</source> <volume>63</volume>, <fpage>22</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2014.09.031</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Mazzola</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Aroyo</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Rea</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sciutti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Interacting with a Social Robot Affects Visual Perception of Space</article-title>,&#x201d; in <conf-name>Proceedings of the 2020 ACM/IEEE International Conference on Human-Robot Interaction</conf-name>, <fpage>549</fpage>&#x2013;<lpage>557</lpage>. <pub-id pub-id-type="doi">10.1145/3319502.3374819</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Mutlu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Shiwa</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kanda</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ishiguro</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hagita</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Footing in Human-Robot Conversations</article-title>,&#x201d; in <conf-name>Proceedings of the 4th ACM/IEEE international conference on Human robot interaction</conf-name>, <fpage>61</fpage>&#x2013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1145/1514095.1514109</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nagai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hosoda</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Morita</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Asada</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>A Constructive Model for the Development of Joint Attention</article-title>. <source>Connect. Sci.</source> <volume>15</volume>, <fpage>211</fpage>&#x2013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1080/09540090310001655101</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Palinko</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Rea</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sandini</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sciutti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Robot Reading Human Gaze: Why Eye Tracking Is Better than Head Tracking for Human-Robot Collaboration</article-title>,&#x201d; in <conf-name>2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE)</conf-name>, <fpage>5048</fpage>&#x2013;<lpage>5054</lpage>. <pub-id pub-id-type="doi">10.1109/iros.2016.7759741</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pandey</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Ali</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Alami</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Towards a Task-Aware Proactive Sociable Robot Based on Multi-State Perspective-Taking</article-title>. <source>Int J Soc Robotics</source> <volume>5</volume>, <fpage>215</fpage>&#x2013;<lpage>236</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-013-0181-3</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pierno</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Becchio</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wall</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>A. T.</given-names>
</name>
<name>
<surname>Turella</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Castiello</surname>
<given-names>U.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>When Gaze Turns into Grasp</article-title>. <source>J. Cognitive Neurosci.</source> <volume>18</volume>, <fpage>2130</fpage>&#x2013;<lpage>2137</lpage>. <pub-id pub-id-type="doi">10.1162/jocn.2006.18.12.2130</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Premack</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Woodruff</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>1978</year>). <article-title>Does the Chimpanzee Have a Theory of Mind?</article-title> <source>Behav. Brain Sci.</source> <volume>1</volume>, <fpage>515</fpage>&#x2013;<lpage>526</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X00076512</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Pynadath</surname>
<given-names>D. V.</given-names>
</name>
<name>
<surname>Marsella</surname>
<given-names>S. C.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Psychsim: Modeling Theory of Mind with Decision-Theoretic Agents</article-title>,&#x201d; in <conf-name>Proceedings of the 19th International Joint Conference on Artificial Intelligence</conf-name> (<publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>Morgan Kaufmann Publishers Inc.</publisher-name>), <fpage>1181</fpage>&#x2013;<lpage>1186</lpage>. <pub-id pub-id-type="doi">10.5555/1642293.1642482</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Rea</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Muratore</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Sciutti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>13-year-olds Approach Human-Robot Interaction like Adults</article-title>,&#x201d; in <conf-name>2016 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob)</conf-name>, <fpage>138</fpage>&#x2013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.1109/DEVLRN.2016.7846805</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Roncone</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pattacini</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Metta</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Natale</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Cartesian 6-dof Gaze Controller for Humanoid Robots</article-title>. <source>Robotics Sci. Syst.</source> <volume>2016</volume>. <pub-id pub-id-type="doi">10.15607/RSS.2016.XII.022</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ros</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sisbot</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Alami</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Steinwender</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hamann</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Warneken</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Solving Ambiguities with Perspective Taking</article-title>,&#x201d; in <conf-name>2010 5th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>181</fpage>&#x2013;<lpage>182</lpage>. <pub-id pub-id-type="doi">10.1109/HRI.2010.5453204</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Scassellati</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Theory of Mind for a Humanoid Robot</article-title>. <source>Aut. Robots</source> <volume>12</volume>, <fpage>13</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1023/A:1013298507114</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Staudte</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Crocker</surname>
<given-names>M. W.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Investigating Joint Attention Mechanisms through Spoken Human-Robot Interaction</article-title>. <source>Cognition</source> <volume>120</volume>, <fpage>268</fpage>&#x2013;<lpage>291</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2011.05.005</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Thomaz</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Lieven</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Cakmak</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chai</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Garrod</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gray</surname>
<given-names>W. D.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>Interaction for Task Instruction and Learning</article-title>,&#x201d; in <conf-name>Interactive task learning: Humans, robots, and agents acquiring new tasks through natural interactions</conf-name> (<publisher-name>MIT Press</publisher-name>), <fpage>91</fpage>&#x2013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.7551/mitpress/11956.003.0011</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Trafton</surname>
<given-names>J. G.</given-names>
</name>
<name>
<surname>Cassimatis</surname>
<given-names>N. L.</given-names>
</name>
<name>
<surname>Bugajska</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Brock</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Mintz</surname>
<given-names>F. E.</given-names>
</name>
<name>
<surname>Schultz</surname>
<given-names>A. C.</given-names>
</name>
</person-group> (<year>2005a</year>). <article-title>Enabling Effective Human-Robot Interaction Using Perspective-Taking in Robots</article-title>. <source>IEEE Trans. Syst. Man. Cybern. A</source> <volume>35</volume>, <fpage>460</fpage>&#x2013;<lpage>470</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCA.2005.850592</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Trafton</surname>
<given-names>J. G.</given-names>
</name>
<name>
<surname>Schultz</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Bugajska</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mintz</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2005b</year>). &#x201c;<article-title>Perspective-taking with Robots: Experiments and Models</article-title>,&#x201d; in <conf-name>ROMAN 2005. IEEE International Workshop on Robot and Human Interactive Communication, 2005</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>580</fpage>&#x2013;<lpage>584</lpage>. <pub-id pub-id-type="doi">10.1109/ROMAN.2005.1513842</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vinanzi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Patacchiola</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chella</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cangelosi</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Would a Robot Trust You? Developmental Robotics Model of Trust and Theory of Mind</article-title>. <source>Phil. Trans. R. Soc. B</source> <volume>374</volume>, <fpage>20180032</fpage>. <pub-id pub-id-type="doi">10.1098/rstb.2018.0032</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wallkotter</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tulli</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Castellano</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Paiva</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chetouani</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Explainable Agents through Social Cues: A Review</article-title>. <source>J. Hum.-Robot Interact.</source> <volume>10</volume> (<issue>3</issue>), <fpage>24</fpage>. <pub-id pub-id-type="doi">10.1145/3457188</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Winfield</surname>
<given-names>A. F. T.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Experiments in Artificial Theory of Mind: From Safety to Story-Telling</article-title>. <source>Front. Robot. AI</source> <volume>5</volume>, <fpage>75</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2018.00075</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolgast</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tandler</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Umlauft</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Adults&#x27; Dispositional and Situational Perspective-Taking: a Systematic Review</article-title>. <source>Educ. Psychol. Rev.</source> <volume>32</volume>, <fpage>353</fpage>&#x2013;<lpage>389</lpage>. <pub-id pub-id-type="doi">10.1007/s10648-019-09507-y</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Schermerhorn</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Scheutz</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Adaptive Eye Gaze Patterns in Interactions with Human and Artificial Agents</article-title>. <source>ACM Trans. Interact. Intell. Syst.</source> <volume>1</volume>, <fpage>1</fpage>&#x2013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1145/2070719.2070726</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zonca</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fols&#xf8;</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sciutti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021a</year>). <article-title>Dynamic Modulation of Social Influence by Indirect Reciprocity</article-title>. <source>Sci. Rep.</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-021-90656-y</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zonca</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Folso</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sciutti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021b</year>). <article-title>If You Trust Me, I Will Trust You: the Role of Reciprocity in Human-Robot Trust</article-title>. <comment>
<italic>arXiv preprint arXiv:2106.14832</italic>
</comment> </citation>
</ref>
</ref-list>
</back>
</article>