<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frobt.2017.00016</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Personality Perception of Robot Avatar Teleoperators in Solo and Dyadic Tasks</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Bremner</surname> <given-names>Paul Adam</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/256799"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Celiktutan</surname> <given-names>Oya</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/380982"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Gunes</surname> <given-names>Hatice</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/181977"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Bristol Robotics Laboratory, University of West England</institution>, <addr-line>Bristol</addr-line>, <country>UK</country></aff>
<aff id="aff2"><sup>2</sup><institution>Computer Laboratory, University of Cambridge</institution>, <addr-line>Cambridge</addr-line>, <country>UK</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Giuseppe Carbone, University of Cassino, Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Fulvio Mastrogiovanni, University of Genoa, Italy; Paolo Boscariol, University of Udine, Italy</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Paul Adam Bremner, <email>paul.bremner&#x00040;brl.ac.uk</email></corresp>
<fn fn-type="other" id="fn001"><p><sup>&#x02020;</sup>These authors have contributed equally to this work.</p></fn>
<fn fn-type="other" id="fn002"><p>Specialty section: This article was submitted to Humanoid Robotics, a section of the journal Frontiers in Robotics and AI</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>05</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>4</volume>
<elocation-id>16</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>12</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>04</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Bremner, Celiktutan and Gunes.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Bremner, Celiktutan and Gunes</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Humanoid robot avatars are a potential new telecommunication tool, whereby a user is remotely represented by a robot that replicates their arm, head, and possible face movements. They have been shown to have a number of benefits over more traditional media such as phones or video calls. However, using a teleoperated humanoid as a communication medium inherently changes the appearance of the operator, and appearance-based stereotypes are used in interpersonal judgments (whether consciously or unconsciously). One such judgment that plays a key role in how people interact is personality. Hence, we have been motivated to investigate if and how using a robot avatar alters the perceived personality of teleoperators. To do so, we carried out two studies where participants performed 3 communication tasks, solo in study one and dyadic in study two, and were recorded on video both with and without robot mediation. Judges recruited using online crowdsourcing services then made personality judgments of the participants in the video clips. We observed that judges were able to make internally consistent trait judgments in both communication conditions. However, judge agreement was affected by robot mediation, although which traits were affected was highly task dependent. Our most important finding was that in dyadic tasks personality trait perception was shifted to incorporate cues relating to the robot&#x02019;s appearance when it was used to communicate. Our findings have important implications for telepresence robot design and personality expression in autonomous robots.</p>
</abstract>
<kwd-group>
<kwd>telepresence</kwd>
<kwd>Big Five personality traits</kwd>
<kwd>personality perception</kwd>
</kwd-group>
<contract-num rid="cn01">EP/L00416X/1</contract-num>
<contract-sponsor id="cn01">Engineering and Physical Sciences Research Council<named-content content-type="fundref-id">10.13039/501100000266</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="5"/>
<equation-count count="0"/>
<ref-count count="52"/>
<page-count count="16"/>
<word-count count="13171"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<label>1</label> <title>Introduction</title>
<p>Telecommunication is omnipresent in today&#x02019;s society, with people desiring to be able to communicate with one another, regardless of distance, for a variety of social and practical reasons. While video-enabled communication offers a number of benefits over voice-only communication, it is still lacking compared to face-to-face interactions (Daly-Jones et al., <xref ref-type="bibr" rid="B20">1998</xref>). For example, remotely located team members are less included in cooperative activities than colocated team members (Daly-Jones et al., <xref ref-type="bibr" rid="B20">1998</xref>) and have fewer conversational turns and speaking time in group conversations (O&#x02019;Conaill et al., <xref ref-type="bibr" rid="B39">1993</xref>). Suggested reasons for these disparities are a lack of social presence of these remote group members, reduced engagement, and reduced awareness of actions (Tang et al., <xref ref-type="bibr" rid="B47">2004</xref>). A suggested underlying cause for the disparities found in traditional telecommunication is a lack of physical presence. An alternative is the use of teleoperated robots as communication media. A common approach to such embodied telecommunication is the use of mobile remote presence (MRP) devices: a screen displaying the operators face mounted on a stalk attached to a wheeled base (Kristoffersson et al., <xref ref-type="bibr" rid="B31">2013</xref>). Though studies examining the utility of MRPs have found that there are some improvements in social presence, different social norms are observed when people use them to interact, and there are impacts on trust and rapport (Lee and Takayama, <xref ref-type="bibr" rid="B33">2011</xref>; Rae et al., <xref ref-type="bibr" rid="B41">2013</xref>). Further, such systems are not able to effectively transmit non-verbal communication cues, a key element of human communication not only for information conveyance but also in maintaining engagement and building rapport (Salam et al., <xref ref-type="bibr" rid="B44">2016</xref>).</p>
<p>A proposed method for further improving social presence and effectively transmitting body language is to use a humanoid robot as a communication medium. In such a system, the operator&#x02019;s body language is duplicated on a humanoid robot such that it is comprehensible and highly salient (Bremner and Leonards, <xref ref-type="bibr" rid="B15">2016</xref>; Bremner et al., <xref ref-type="bibr" rid="B13">2016b</xref>). Using a humanoid robot as a communications avatar has benefits with regard to engagement of conversational partners (Hossen Mamode et al., <xref ref-type="bibr" rid="B29">2013</xref>), social presence (Adalgeirsson and Breazeal, <xref ref-type="bibr" rid="B1">2010</xref>), group interaction (Hossen Mamode et al., <xref ref-type="bibr" rid="B29">2013</xref>), and trust (Bevan and Stanton Fraser, <xref ref-type="bibr" rid="B8">2015</xref>).</p>
<p>However, when using a robot as a remote proxy for communication, the operator is represented with a different physical appearance, much as computer generated avatars do in virtual environments. Appearance has been observed to be utilized in making interpersonal judgments (Naumann et al., <xref ref-type="bibr" rid="B38">2009</xref>), and this can extend to virtual avatars (Wang et al., <xref ref-type="bibr" rid="B50">2013</xref>; Fong and Mar, <xref ref-type="bibr" rid="B24">2015</xref>). It was observed that judges made relatively consistent inferences based on avatar appearance alone (Wang et al., <xref ref-type="bibr" rid="B50">2013</xref>; Fong and Mar, <xref ref-type="bibr" rid="B24">2015</xref>), and more attractive avatars were rated more highly in an interview scenario (Behrend et al., <xref ref-type="bibr" rid="B7">2012</xref>). How this might manifest with robot avatars, in particular in the interaction between a robot appearance and human voice communication, remains unclear and is yet to be explored.</p>
<p>Here, the particular judgment we are concerned with is that of personality perception, an important facet of communication. Researchers in psychology have shown that personality plays a key role in forming interpersonal relationships, and predicting future behaviors (Borkenau et al., <xref ref-type="bibr" rid="B11">2004</xref>). These findings have motivated a significant body of work for how people judge others&#x02019; personalities based on their observable behaviors. A key component of these social cues for personality are non-verbal behaviors. We aim to investigate if such non-verbal personality cues transmitted by a teleoperated humanoid robot continue to be utilized in personality judgments, and how they interact with verbal cues. Non-verbal cues can be transmitted as our robot teleoperation system utilizes a motion capture-based approach so that arm and head movements the operator performs while talking are recreated with minimal delay on a NAO humanoid robot (Bremner and Leonards, <xref ref-type="bibr" rid="B15">2016</xref>). The control system is intuitive and immersive, and we observe people behaving similarly to how they do face-to-face (Bremner et al., <xref ref-type="bibr" rid="B13">2016b</xref>).</p>
<p>We designed two experiments which follow an experimental methodology common in the personality analysis literature, i.e., videos of participants performing different communication tasks are shown to external observers (judges) for personality assessment (e.g., Borkenau et al. (<xref ref-type="bibr" rid="B11">2004</xref>)). Personality judgments are made on the so-called big five traits, <italic>extroversion, conscientiousness, agreeableness, neuroticism</italic>, and <italic>openness</italic> (multiple questions relate to each trait). We varied communication media between judges, either video only or robot mediated (also recorded on video). Two main measures are used to see whether there was an effect of communication condition on personality judgments: (1) judge consistency in how they evaluate a given trait, both within and between judge (low consistency indicates lack of cues or conflicting cues); and (2) personality shifts between high and low classification for each trait between the video and robot conditions.</p>
<p>Hence we address the following research questions:
<list list-type="bullet">
<list-item><p><bold>RQ1</bold>: Are there differences in judges&#x02019; consistency in assessing personality traits (within-judge consistency)?</p></list-item>
<list-item><p><bold>RQ2</bold>: Are there differences in how much judges agree with one another on personality judgments (between-judge consistency)?</p></list-item>
<list-item><p><bold>RQ3</bold>: Are personality judgments less accurate compared to self-ratings (self-other agreement)?</p></list-item>
<list-item><p><bold>RQ4</bold>: Are perceived personalities systematically shifted to incorporate characteristics associated with the robot&#x02019;s appearance (personality shifts)?</p></list-item>
</list></p>
<p>This paper is an extended version of our work published by Bremner et al. (<xref ref-type="bibr" rid="B12">2016a</xref>). We extended our previous work by adding a second experiment that refined our experimental procedure and used dyadic rather than solo tasks. Our discussions and conclusions are extended to include both experiments, evaluating all our results to give a clearer picture.</p>
<p>In the first experiment, three tasks are performed direct to camera, i.e., solo tasks. In the second experiment, participants performed three tasks that involved interaction with a confederate, i.e., dyadic. The first experiment provided some limited evidence for shifts in personality perception. Further, by adding an audio-only communication condition, we were able to show that the robot was not simply ignored, and gesture cues performed on the robot were utilized. An important finding from the first experiment was that effects were very task dependent, as the literature suggested. Borkenau et al. (<xref ref-type="bibr" rid="B11">2004</xref>) found that <italic>openness</italic> is better inferred in more ability-demanding tasks such as pantomime task. Hence, the second experiment used additional tasks, which by being dyadic will engender personality cues differently; it is also a refinement of our experimental procedure, improving the reliability of our results. It produced compelling evidence that cues related to the robot&#x02019;s appearance were incorporated in personality judgments, causing consistent shifts in perceived personality.</p>
</sec>
<sec id="S2">
<label>2</label> <title>Related Work</title>
<p>A common approach to investigating personality judgments is first impression or thin slice personality analysis. It is a body of research that studies the accuracy with which people are able to make personality judgments of others based only on short behavioral episodes (termed thin slices). This approach is taken as it is believed that these judgments provide insight into the assessments people make in everyday interactions (Funder and Sneed, <xref ref-type="bibr" rid="B27">1993</xref>; Borkenau et al., <xref ref-type="bibr" rid="B11">2004</xref>). In such studies, targets are typically asked to perform a range of communication tasks, either solo performances to camera or dyadic with confederates, and are filmed while doing so. <italic>Judges</italic> then observe the video clips and complete personality assessment questionnaires. Ratings of judges are compared with target self-ratings, acquaintance ratings, and for inter-judge agreement. For many traits, there is sufficient inter-judge agreement for the method to be useful in assessing the impressions a person creates on those they interact with (Borkenau et al., <xref ref-type="bibr" rid="B11">2004</xref>); however, the accuracy of judge ratings to self/acquaintance ratings is typically a lot lower, as self/acquaintance ratings are error prone, and use different sources to make their judgments (Vinciarelli and Mohammadi, <xref ref-type="bibr" rid="B49">2014</xref>).</p>
<p>Often analyzed in thin slice personality studies are the cues that appear to be utilized in people making their judgments. Appearance, speaking style, gaze, head movements, and hand gestures have been frequently reported to be significant predictors of personality (Riggio and Friedman, <xref ref-type="bibr" rid="B43">1986</xref>; Borkenau and Liebler, <xref ref-type="bibr" rid="B10">1992</xref>; Borkenau et al., <xref ref-type="bibr" rid="B11">2004</xref>). Indeed, this sort of analysis forms the basis for automated personality analysis systems. Aran and Gatica-Perez (<xref ref-type="bibr" rid="B4">2013</xref>) focused on personality perception in a small group meeting scenario. They extracted a set of multimodal features including speaking turn, pitch, energy, head and body activity, and social attention features. Thin slice analysis yielded the highest accuracy for <italic>extroversion</italic>, while <italic>openness</italic> was better modeled by longer time scales. With regard to the related work in personality computing, the closest approach was presented in the study by Batrinca et al. (<xref ref-type="bibr" rid="B6">2016</xref>). In order to analyze the Big Five personality traits, Batrinca et al. conducted a study where a set of participants were asked to interact with a computer, which was controlled by an experimenter, and then a different set of participants were asked to interact with the experimenter face-to-face to collaborate on completing a map task. In order to elicit the participants&#x02019; personality traits, the experimenter exhibited four different levels of collaborative behaviors from fully collaborative to fully non-collaborative. Self-reported personality traits were used to study the manifestation of traits from audiovisual cues. In the human-machine interaction setting, their results showed that (1) extroversion and neuroticism can be predicted with a high level of accuracy, regardless of the collaboration modality; (2) prediction of the agreeableness and conscientiousness traits depends on the collaboration modality; and (3) openness was the only trait that cannot be modeled. In contrast to their findings in the human&#x02013;machine interaction setting, they showed that openness was the trait that can be predicted with highest accuracy in the human&#x02013;human interaction setting.</p>
<p>Applying such personality perception analysis to robot teleoperators has so far been limited. Perception of teleoperator&#x02019;s personality is important not only in social interactions but is also crucial where teleoperated robots are used in a service capacity such as for elderly care (Yamazaki et al., <xref ref-type="bibr" rid="B51">2012</xref>), and search and rescue (Martins and Ventura, <xref ref-type="bibr" rid="B35">2009</xref>). In these settings, perception of the operator will effect system utility for carrying out the desired service and achieving the desired outcome. In the study by Celiktutan et al. (<xref ref-type="bibr" rid="B17">2016</xref>), we showed that many of the aforementioned personality cues can be transmitted by a telepresence robot. We trained support vector machine classifiers with a set of features extracted from participants&#x02019; voice and body movements. We found that the use of a robot avatar helps to discriminate between different personality types (e.g., extroverted vs.introverted) better than audio-only mediated communication for extroversion (65%) and conscientiousness (60%).</p>
<p>Studies with Mobile Remote Presence devices (MRPs) have briefly mentioned perceiving the operator&#x02019;s personality (Lee and Takayama, <xref ref-type="bibr" rid="B33">2011</xref>), but it has not been deliberately studied as we do here. There are two studies that look directly at personality perception of teleoperators. Kuwamura et al. (<xref ref-type="bibr" rid="B32">2012</xref>) examined an effect that they term <italic>personality distortion</italic>, demonstrated by reduction in internal consistency of the personality questionnaire they used, for two different robot platforms and communication using video. They use 3 tasks: (1) an experimenter talks freely with the participant, (2) a different experimenter introduces and talks about themselves, and (3) a third experimenter interviews the participant. They only observed <italic>personality distortion</italic> for one of the robot platforms, for <italic>extroversion</italic> in the interview task, and for <italic>agreeableness</italic> in the introduction task. Using a single fixed person for each task, particularly members of the experimental team who are aware of the goals of the study, greatly reduces the ecological validity of their results. In contrast, here we use a large number of na&#x000EF;ve targets performing naturalistic communication, and conduct far more in-depth data analysis.</p>
<p>In a study with a teleoperated, highly humanlike robot, Straub et al. (<xref ref-type="bibr" rid="B46">2010</xref>) examined both how participant teleoperators incorporate the fact that they are operating a robot into their presented identity, and how interlocutors at the robot&#x02019;s location blend operator and robot identities. They used language analysis to make their assessments. They observed that many operators pretended they themselves were a robot, and interlocutors often referred to the operator as a robot. These behaviors are different from what we typically observe with our teleoperation system, where most operators appeared to act naturally as themselves (Bremner et al., <xref ref-type="bibr" rid="B13">2016b</xref>).</p>
</sec>
<sec id="S3" sec-type="materials|methods">
<label>3</label> <title>Materials and Methods</title>
<p>We designed a two-stage experimental method for assessing changes in perceived personality that we used in two studies. First, a set of participants (targets) were recorded performing three communication tasks in two conditions, directly visible on video camera (audiovisual condition) and communicating using the teleoperated robot (teleoperated robot condition, also recorded on camera). This ensures that we have a large set of natural communication behaviors, and hence personality cues, for a range of personality types, that can be viewed directly or when mediated by a robot.</p>
<p>In the second stage of the study, the recorded data were used to create a set of video clips for each target in each communication condition. The video clips were pseudorandomly assigned to a set of surveys in such a way as to have one of each task and communication condition combinations present, with a given target only appearing once in a given survey (i.e., communication condition was varied between surveys). Each survey was viewed by a set of 10 judges, who after watching each clip assessed the personality of that target. We used an online crowdsourcing service to have the clips assessed. Employing judges <italic>via</italic> online crowdsourcing services has recently gained popularity due to its efficiency and practicality as it enables collecting responses from a large group of people within a short period of time (Biel and Gatica-Perez, <xref ref-type="bibr" rid="B9">2013</xref>; Salam et al., <xref ref-type="bibr" rid="B44">2016</xref>).</p>
<p>Personality was assessed by a questionnaire that aims to gather an assessment along the widely known Big Five personality traits (Vinciarelli and Mohammadi, <xref ref-type="bibr" rid="B49">2014</xref>). These five personality traits are <italic>extroversion</italic> (EX&#x02014;assertive, outgoing, energetic, friendly, socially active), <italic>agreeableness</italic> (AG&#x02014;cooperative, compliant, trustworthy), <italic>conscientiousness</italic> (CO&#x02014;self-disciplined, organized, reliable, consistent), <italic>neuroticism</italic> (NE&#x02014;having tendency to negative emotions such as anxiety, depression, or anger), and <italic>openness</italic> (OP&#x02014;having tendency to changing experience, adventure, new ideas). Each trait is measured using a set of items (the BFI-10 (Rammstedt and John, <xref ref-type="bibr" rid="B42">2007</xref>) with 2 per trait in the Solo Tasks Study, and the IPIP-BFM-20 (Topolewska et al., <xref ref-type="bibr" rid="B48">2014</xref>) with 4 per trait in the Dyadic Tasks Study) scored on 10-point Likert scales. As well as being assessed by external observers, each target completed the personality questionnaire for self-assessment.</p>
<sec id="S3-1">
<label>3.1</label> <title>Teleoperation System</title>
<p>In order to reproduce the gestures of targets on the NAO humanoid robot platform from Softbank Robotics (Gouaillier et al., <xref ref-type="bibr" rid="B28">2009</xref>), we used a motion capture-based teleoperation system. Previously we have demonstrated the system to be capable of producing comprehensible gestures (Bremner and Leonards, <xref ref-type="bibr" rid="B14">2015</xref>, <xref ref-type="bibr" rid="B15">2016</xref>). The arm motion of the targets is recorded using a Microsoft Kinect and Polhemus Patriot,<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> and used to produce equivalent motion on the robot. Arm link end points at the wrist, elbow, and shoulder are tracked and were used to calculate joint angles for the robot so that its upper and lower arm links reproduce human arm link positions and motion. This method ensures that joint coordination, and hand trajectories are as similar as possible between the human and the robot within the constraints of the NAO robot platform. Figure <xref ref-type="fig" rid="F1">1</xref> shows a gesture produced by one of the targets, and the equivalent gesture on the NAO.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>Snapshots from the Solo Tasks Study</bold>. Left hand side: a target perceived to be <italic>extroverted</italic> by judges. Right hand side: a target perceived to be <italic>introverted</italic> by judges.</p></caption>
<graphic xlink:href="frobt-04-00016-g001.tif"/>
</fig>
</sec>
<sec id="S3-2">
<label>3.2</label> <title>Solo Tasks Study</title>
<sec id="S3-2-1">
<label>3.2.1</label> <title>Tasks</title>
<p>In the first study, the three tasks performed by participants involved them performing directly to the camera, i.e., solo, and were based upon a subset of tasks used by Borkenau et al. (<xref ref-type="bibr" rid="B11">2004</xref>). Each of the tasks was framed as an interaction with the experimenter who stood beside the video camera used in the recordings, and provided non-verbal feedback and prompt questions to ensure as natural communicative behaviors as possible. Targets were instructed to speak for as long as they felt able, with a maximum time of 2&#x02009;min for each task. The majority of the targets talked for 30&#x02013;60&#x02009;s on each task, with occasional prompts for missing information. Prior to performing tasks, we asked the targets to introduce themselves and give some information about themselves, e.g., where they work, what they do, their family, etc. This stage was purely to help naturalize the target to the experimental setting. It was not used to produce clips for judge rating.</p>
<sec id="S3-2-1-1">
<label>3.2.1.1</label> <title>Task 1 (Hobby)</title>
<p>This task asked targets to describe one of their hobbies, providing as much detail as possible. Suggested detail included what their hobby involves, why they like it, how long have they been doing it for, etc. Example personality cues we anticipated from this task include what targets have as their hobby, and what detail and the depth of detail they provide while describing their hobby.</p>
</sec>
<sec id="S3-2-1-2">
<label>3.2.1.2</label> <title>Task 2 (Story)</title>
<p>This task is based on Murray&#x02019;s thematic apperception test (TAT), where the target is shown a picture and is asked to tell a dramatic story based on a picture (Murray, <xref ref-type="bibr" rid="B37">1943</xref>). They are asked what is happening in the picture,<xref ref-type="fn" rid="fn2"><sup>2</sup></xref> what are the characters thinking and feeling, what happens before the events in the picture and what happens after. The picture is purposely designed to be ambiguous so that the target has the scope to interpret the picture as they see fit, and has to be creative in their story telling. It is a projective test, where the details given by the target, and how they relate the actions of the characters, provide cues about their personality.</p>
</sec>
<sec id="S3-2-1-3">
<label>3.2.1.3</label> <title>Task 3 (Mime)</title>
<p>This task required the targets to mime preparing and cooking a meal of their choice. This was different from the mime task used by Borkenau et al. (<xref ref-type="bibr" rid="B11">2004</xref>), where targets had to mime alternative uses for a brick. Our pretests showed little variability between targets for that task. Instead, the chosen task gave the desired variability, and the gestures were better suited to performance on the NAO robot. Which meal was selected, and the complexity of the mime, are example personality cues we anticipated from this task.</p>
</sec>
</sec>
<sec id="S3-2-2">
<label>3.2.2</label> <title>Participants</title>
<p>Twenty-six participants were recorded as targets (16 female, mean age&#x02009;&#x0003D;&#x02009;30.85, SD&#x02009;&#x0003D;&#x02009;7.58) and gave written informed consent for their participation, they were reimbursed with a &#x000A3;5 gift voucher for their time. Recordings for 20 of the targets were used to create the clips used for judgments (6 targets were omitted due to recording problems). The study was approved by the ethics committee of the Faculty of Environment and Technology of The University of the West of England.</p>
<p>Clip ratings were undertaken by 143 judges recruited through the CrowdFlower online crowdsourcing platform.<xref ref-type="fn" rid="fn3"><sup>3</sup></xref> Judges were compensated 50 cents for annotating a total of four clips.</p>
</sec>
<sec id="S3-2-3">
<label>3.2.3</label> <title>Recordings</title>
<p>All tasks were recorded by one RGB video camera and the motion capture system used for teleoperation. The recorded motion capture data were then used to produce robot-mediated versions of the targets&#x02019; performances on the NAO robot using the aforementioned teleoperation system, which were also recorded on video.</p>
<p>In addition to the audiovisual and teleoperated robot conditions, an audio-only condition was created using the audio from hobby and story tasks. Hence, each target had a total of 8 clips split over 3 communication conditions: 3 clips for the audiovisual condition, 2 clips for the audio-only condition, and 3 clips for the teleoperated robot condition. This resulted in a total of 158 clips (two clips became corrupted).</p>
<p>To avoid confusion, prompt questions were edited out of the clips. Further, for the few tasks where performance exceeded 60&#x02009;s, clips were edited to be close to this length as pretests showed a decrease in the reliability of judgments with overly long clips. Mean clip duration was 50&#x02009;s (SD&#x02009;&#x0003D;&#x02009;20&#x02009;s).</p>
<p>The clips were split up into surveys each containing four clips: one of each task and one of the audio-only clips, each of a unique target. Communication condition was pseudo-randomized across the three tasks in each survey, but always contained at least one of each communication condition.</p>
</sec>
</sec>
<sec id="S3-3">
<label>3.3</label> <title>Dyadic Tasks Study</title>
<sec id="S3-3-1">
<label>3.3.1</label> <title>The Extended Teleoperation System</title>
<p>The teleoperation system was extended to enable interactive multimodal communication. The first addition made was a stereo camera helmet on the NAO robot, the images from which are displayed in an Oculus Rift head-mounted display (HMD). Coupled with using the Rift&#x02019;s inertial measurement unit to drive the robot&#x02019;s head, meant the operator could see from the robots point of view, and their gaze direction and head motion could be observed on the robot. Secondly we used a voice over IP communication system to allow full duplex audio communication. Finally, due to feedback from participants in the Solo Tasks Study, we did not use the Polhemus Patriot in the Dyadic Tasks Study to make behaviors more natural; importantly, wrist rotation was only really needed for the mime task in the Solo Tasks Study, and is less important for normal gesturing. Figure <xref ref-type="fig" rid="F2">2</xref> shows the teleoperation system and the setup during performance of dyadic tasks in the teleoperation (TO) condition.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Snapshots from the Dyadic Tasks Study</bold>. Upper row: illustration of teleoperation (TO) room and interaction room. Lower row: snapshots from the dyadic interaction sequences.</p></caption>
<graphic xlink:href="frobt-04-00016-g002.tif"/>
</fig>
</sec>
<sec id="S3-3-2">
<label>3.3.2</label> <title>Tasks</title>
<p>In the second study, the three tasks performed by participants involved interacting with a confederate, i.e., dyadic. A confederate was used to ensure that each participant had the same interactive partner, giving us a measure of control over the interactions, while still seeming natural to the participants. The three selected tasks were based on the suggestions by Funder et al. (<xref ref-type="bibr" rid="B26">2000</xref>) of having an informative task, a competitive task, and a cooperative task. The intention of these task types is that they each engender personality cues in different ways.</p>
<p>The three tasks were briefly explained to the participant and the confederate together, and more detailed written instructions were provided to be used during the experimental session. This was done to ensure that the experimenters could leave the room for the participant and confederate to converse alone. The two communication conditions (audiovisual and teleoperated robot) were performed sequentially, in a pseudorandomized order, in the same room. The audiovisual condition was recorded face-to-face, i.e., with both participant and confederate seated across a table from one another. In the teleoperated robot condition, the participant moved to an adjoining room where the teleoperation controls were located, while the confederate sat at a table across from the robot.</p>
<sec id="S3-3-2-1">
<label>3.3.2.1</label> <title>Task 1 (Informative)</title>
<p>Participants watched a clip from a Sylvester and Tweety cartoon, which they then had to describe to the confederate. This is a task commonly used to examine gesturing (Alibali, <xref ref-type="bibr" rid="B2">2001</xref>), as describing the action filled cartoon often engenders gestures, which may be useful personality cues that can be produced by the robot. Another key reason for this task choice was that all participants have the same things to talk about: in the previously used hobby task several participants struggled to find much to say without significant prompting. Two different Sylvester and Tweety cartoons were used, one for each communication condition; cartoon assignment was randomized between conditions. We expected there to be an abundance of gestural cues, as well as cues related to the participants&#x02019; verbal behavior (such as how detailed the description was).</p>
</sec>
<sec id="S3-3-2-2">
<label>3.3.2.2</label> <title>Task 2 (Competitive)</title>
<p>The participants and the confederate played a memory-based word game adapted from the traditional <italic>Grandmothers Trunk</italic> game. The first player says &#x0201C;My Grandmother went on holiday and she&#x02026;&#x0201D; and adds something she did, accompanied by a gesture, the other player then repeats what the first said and their gesture, and adds something else she did. Play continues alternating between players who repeat the whole list of things and perform the gestures, adding a new thing each time, until one player forgets something and that player loses. How they approach the competitive nature of the task, and the actions they select are personality cues we expected from this task.</p>
</sec>
<sec id="S3-3-2-3">
<label>3.3.2.3</label> <title>Task 3 (Cooperative)</title>
<p>The participants and the confederate cooperated to put a set of 5 items into utility order for surviving in a given scenario. There were two scenarios each with its own set of items, surviving a ship wreck, and surviving a crash landing on the moon. One scenario was presented per communication condition and was randomly assigned. How agreement is reached, and how the task is approached are the main cues we expect from this task.</p>
</sec>
</sec>
<sec id="S3-3-3">
<label>3.3.3</label> <title>Participants</title>
<p>Thirty participants were recorded as targets (13 female, mean age&#x02009;&#x0003D;&#x02009;25.01, SD&#x02009;&#x0003D;&#x02009;4.2), and gave written informed consent for their participation, they were reimbursed with a &#x000A3;5 gift voucher for their time. Recordings for 25 of the targets were used to create the clips used for judgments (5 targets were omitted due to recording problems). The study was approved by the ethics committee of the University of Cambridge.</p>
<p>Clip ratings were undertaken by 250 judges recruited through the Prolific Academic online crowdsourcing platform.<xref ref-type="fn" rid="fn4"><sup>4</sup></xref> Each judge rated 6 clips and was compensated &#x000A3;2 for their time.</p>
</sec>
<sec id="S3-3-4">
<label>3.3.4</label> <title>Recordings</title>
<p>In all tasks, both the confederate and the participant were recorded by separate RGB video cameras. The confederate was only recorded to obscure the fact that she was a confederate. In the teleoperated robot condition, a video camera recorded the robot instead of the participant. In order to produce videos of identical length for all targets and tasks, the video clips were further edited to select a 60&#x02009;s segment from the beginning of the Informative task and from the end of Competitive and Cooperative tasks. This is in line with suggestions by Carney et al. (2007b) for using clips of this length of a task to maximize consistent judgment conditions for each target. Thus, each target had a set of three 60&#x02009;s clips for each of the two communication conditions. One survey consisted of a pseudo-randomized set of 6 clips, 1 example of each task in each communication condition, with unique targets in each clip. Additionally a practice clip of the confederate was added to the start of all surveys to use as a measure of judge reliability, it also served to demonstrate her voice such that it could be ignored when she spoke during the target clips.</p>
<p>In Table <xref ref-type="table" rid="T1">1</xref>, we summarized both studies in terms of number of participants, tasks, communication conditions, and communicated cues.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Summary of the conducted studies</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Study</th>
<th valign="top" align="center">Number of participants</th>
<th valign="top" align="left">Tasks</th>
<th valign="top" align="left">Communication conditions</th>
<th valign="top" align="left">Communicated cues</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Solo</td>
<td align="center" valign="top">26</td>
<td align="left" valign="top">Hobby, story, mime</td>
<td align="left" valign="top">AO, AV, TO</td>
<td align="left" valign="top">Wrist, elbow, shoulder motion, wrist orientation</td>
</tr>
<tr>
<td align="left" valign="top">Dyadic</td>
<td align="center" valign="top">30</td>
<td align="left" valign="top">Informative, competitive, cooperative</td>
<td align="left" valign="top">AV, TO</td>
<td align="left" valign="top">Wrist, elbow, shoulder motion; head motion; gaze direction</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>AO, audio-only; AV, audiovisual; TO, teleoperation</italic>.</p></table-wrap-foot></table-wrap>
</sec>
</sec>
</sec>
<sec id="S4">
<label>4</label> <title>Results and Analysis</title>
<p>To address the research questions introduced in Section <xref ref-type="sec" rid="S1">1</xref>, we analyzed the level of agreement and the extent of shifts with respect to different communication conditions (e.g., audiovisual/AV, audio-only/AO, teleoperation/TO) and different tasks for each personality trait. We evaluated personality judgments to measure intra-/inter-agreement, self-other agreement, and personality shifts as below.</p>
<list list-type="bullet">
<list-item><p><italic>Intra-judge Agreement</italic>: Intra-judge agreement (also known as internal consistency) evaluates the quality of personality judgments based on correlations between different questionnaire items that contribute to measuring the same personality trait by each judge. We measured intra-judge agreement in terms of standardized Cronbach&#x02019;s <italic>&#x003B1;</italic>: <inline-formula><mml:math id="M1"><mml:mrow><mml:mn>&#x003B1;</mml:mn><mml:mo>=</mml:mo><mml:mstyle scriptlevel='+1'><mml:mfrac><mml:mrow><mml:mi>K</mml:mi><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>K</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula> where <italic>K</italic> is the number of the items (<italic>K</italic>&#x02009;&#x0003D;&#x02009;2 in the Solo Tasks Study, and <italic>K</italic>&#x02009;&#x0003D;&#x02009;4 in the Dyadic Tasks Study) and <inline-formula><mml:math id="M2"><mml:mover accent='true'><mml:mi>r</mml:mi><mml:mo>&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> is the mean of pairwise correlations between values assigned. The resulting <italic>&#x003B1;</italic> coefficient ranges from 0 to 1; higher values are associated with higher internal consistency and values less than 0.5 are usually unacceptable (McKeown et al., <xref ref-type="bibr" rid="B36">2012</xref>).</p></list-item>
<list-item><p><italic>Inter-judge Agreement</italic>: Inter-judge agreement refers to the level of consensus among judges. We computed the inter-judge agreement in terms of intraclass correlation (ICC) (Shrout and Fleiss, <xref ref-type="bibr" rid="B45">1979</xref>). ICC assesses the reliability of the judges by comparing the variability of different ratings of the same target to the total variation across all ratings and all targets. We used ICC(1,k) as in our experiments each target subject was rated by a different set of k judges, randomly sampled from a larger population of judges. ICC(1,k) measures the degree of agreement for ratings that are averages of <italic>k</italic> independent ratings on the target subjects.</p></list-item>
<list-item><p><italic>Self-other Agreement</italic>: Self-other agreement measures the similarity between the personality judgments made by self and others. We computed self-other agreement in terms of Pearson correlation and tested the significance of correlations using Student&#x02019;s <italic>t</italic> distribution. Pearson correlation was computed between the target&#x02019;s self-reported responses and the mean of the others&#x02019; scores per trait.</p></list-item>
<list-item><p><italic>Personality Shifts</italic>: Personality shift refers to the extent to which people shifted from one personality class to another, in judges&#x02019; perception, between AV and TO conditions. In order to measure shifts, we first classified each target into low or high (e.g., <italic>introverted</italic> or <italic>extroverted</italic>) for each trait according to if their average judge rating for each task was above or below the mean for all targets in AV. For each trait, each target was grouped according to their classification in both conditions, creating 4 groups (i.e., AV: high and TO: high, AV: high and TO: low, etc.). We presented these results in terms of contingency tables and tested the significance using McNemar&#x02019;s test with Edwards&#x02019;s correction (Edwards, <xref ref-type="bibr" rid="B22">1948</xref>).</p></list-item>
</list>
<p>In the following subsections, we present these results for each study (solo and dyadic) separately.</p>
<sec id="S4-1">
<label>4.1</label> <title>Solo Tasks Study</title>
<sec id="S4-1-1">
<label>4.1.1</label> <title>Elimination of Low-Quality Judges</title>
<p>Although crowdsourcing techniques have many advantages, identifying annotators who assign labels without looking at the content (low-quality judges or spammers) is necessary to get informative results. As a first measure, we eliminated judges who incorrectly answered a test question about the content of the clips. After this elimination mean-judges-per-clip was 7.9 (SD&#x02009;&#x0003D;&#x02009;1.5), with minimum judges-per-clip being 5.</p>
<p>To assess whether there remained further low-quality judges we calculated within-judge consistency for the AV clips using Cronbach&#x02019;s <italic>&#x003B1;</italic>, which measures whether the values assigned to the items that contribute to the same trait are correlated. The average value across all tasks was lower than we expected (less than 0.5), indicating some judges answer randomly. With no low-quality judges, we would expect values for the AV clips greater than 0.5, i.e., in line with values reported in the literature for the BFI-10 with video clips assessed by online judges (Cred&#x000E9; et al., <xref ref-type="bibr" rid="B19">2012</xref>). We therefore used a judge selection method to remove these additional low-quality judges. We used a ranking-based method based on pairwise correlations instead of standard methods for outlier detection. For each clip, we calculated an average correlation score for each judge from pairwise correlations (using all 10 questions in the BFI-10) with the remaining judges. Judges with low correlation scores are deemed to be spammers. The judges were then ranked in order of correlation score and the <italic>k</italic> highest ranked selected.</p>
<p>To evaluate the efficacy of this ranking procedure we calculated within-judge consistency results for the AV clips for different judge numbers ranging from <italic>k</italic>&#x02009;&#x0003D;&#x02009;10 (without elimination) to <italic>k</italic>&#x02009;&#x0003D;&#x02009;3. These values averaged over all tasks are presented in Figure <xref ref-type="fig" rid="F3">3</xref>A. We further validated this by computing ICC with varying number of judges, Figure <xref ref-type="fig" rid="F3">3</xref>C. Selecting 5 judges per clip (based on pairwise comparisons) was found to be sufficient to increase reliability to acceptable levels for the AV clips (greater than 0.5) for all traits except for <italic>openness</italic>. We use 5 judges as it allows us to exclude all judges who failed the test question while having the same number of judges for all clips [5 judges is common in this type of study, e.g., Borkenau and Liebler (<xref ref-type="bibr" rid="B10">1992</xref>)].</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Changes in Cronbach&#x02019;s <italic>&#x003B1;</italic> values (A,B) and ICC values (C,D) as a function of number selected judges (k) for different traits in the AV communication condition for Solo Tasks Study (A&#x02013;C) and Dyadic Tasks Study (B&#x02013;D)</bold>.</p></caption>
<graphic xlink:href="frobt-04-00016-g003.tif"/>
</fig>
</sec>
<sec id="S4-1-2">
<label>4.1.2</label> <title>Within-Judge Consistency</title>
<p>Within-judge consistency was measured in terms of Cronbach&#x02019;s <italic>&#x003B1;</italic>. For the selected 5 judges per clip, the detailed results with respect to different communication conditions and tasks are presented in Table <xref ref-type="table" rid="T2">2</xref>(a), where <italic>&#x003B1;</italic> values that indicate sufficient reliability for the BFI-10 (greater than 0.5, in line with values reported in the literature (Cred&#x000E9; et al., <xref ref-type="bibr" rid="B19">2012</xref>)) are highlighted in bold. To compare <italic>&#x003B1;</italic> values between communication conditions we follow the method suggested by Feldt et al. (<xref ref-type="bibr" rid="B23">1987</xref>): 95% confidence intervals are calculated for each <italic>&#x003B1;</italic> value, and if the value from one condition falls outside the confidence intervals from a condition it is being compared to, this suggests it is significantly less consistent. Comparing AO with AV for the hobby task, values for all traits, except for <italic>agreeableness</italic>, fall outside the 95% confidence intervals of the AV values. Comparing TO with AV for the mime task, values for all traits, except for <italic>conscientiousness</italic>, fall outside the 95% confidence intervals of the AV values. This indicates AV is found to be more consistent as compared to AO for the hobby task (except for <italic>agreeableness</italic>) and TO for the mime task (except for <italic>conscientiousness</italic>). No other comparisons indicate significant differences.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><bold>Analysis of personality judgments across 3 communication conditions and 3 tasks</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center" colspan="4">Audiovisual (AV)<hr/></th>
<th valign="top" align="center" colspan="3">Audio-only (AO)<hr/></th>
<th valign="top" align="center" colspan="4">Teleoperation (TO)<hr/></th>
</tr><tr>
<th valign="top" align="center"/>
<th valign="top" align="center">Hobby</th>
<th valign="top" align="center">Story</th>
<th valign="top" align="center">Mime</th>
<th valign="top" align="center">All</th>
<th valign="top" align="center">Hobby</th>
<th valign="top" align="center">Story</th>
<th valign="top" align="center">All</th>
<th valign="top" align="center">Hobby</th>
<th valign="top" align="center">Story</th>
<th valign="top" align="center">Mime</th>
<th valign="top" align="center">All</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" colspan="12"><bold>(a) Within-judge</bold></td>
</tr>
<tr>
<td align="left" valign="top">EX</td>
<td align="center" valign="top"><bold>0.64</bold></td>
<td align="center" valign="top"><bold>0.56</bold></td>
<td align="center" valign="top"><bold>0.63</bold></td>
<td align="center" valign="top"><bold>0.62</bold></td>
<td align="center" valign="top"><bold>0.57</bold></td>
<td align="center" valign="top">&#x02212;0.15</td>
<td align="center" valign="top">0.34</td>
<td align="center" valign="top"><bold>0.61</bold></td>
<td align="center" valign="top">0.39</td>
<td align="center" valign="top">0.19</td>
<td align="center" valign="top">0.47</td>
</tr>
<tr>
<td align="left" valign="top">AG</td>
<td align="center" valign="top"><bold>0.54</bold></td>
<td align="center" valign="top">0.41</td>
<td align="center" valign="top"><bold>0.60</bold></td>
<td align="center" valign="top"><bold>0.52</bold></td>
<td align="center" valign="top"><bold>0.61</bold></td>
<td align="center" valign="top">0.33</td>
<td align="center" valign="top"><bold>0.52</bold></td>
<td align="center" valign="top">0.40</td>
<td align="center" valign="top"><bold>0.56</bold></td>
<td align="center" valign="top">0.37</td>
<td align="center" valign="top">0.44</td>
</tr>
<tr>
<td align="left" valign="top">CO</td>
<td align="center" valign="top">0.47</td>
<td align="center" valign="top"><bold>0.60</bold></td>
<td align="center" valign="top"><bold>0.54</bold></td>
<td align="center" valign="top"><bold>0.55</bold></td>
<td align="center" valign="top"><bold>0.50</bold></td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">0.39</td>
<td align="center" valign="top"><bold>0.54</bold></td>
<td align="center" valign="top"><bold>0.56</bold></td>
<td align="center" valign="top"><bold>0.57</bold></td>
<td align="center" valign="top"><bold>0.55</bold></td>
</tr>
<tr>
<td align="left" valign="top">NE</td>
<td align="center" valign="top"><bold>0.76</bold></td>
<td align="center" valign="top"><bold>0.76</bold></td>
<td align="center" valign="top"><bold>0.78</bold></td>
<td align="center" valign="top"><bold>0.78</bold></td>
<td align="center" valign="top"><bold>0.75</bold></td>
<td align="center" valign="top">0.42</td>
<td align="center" valign="top"><bold>0.63</bold></td>
<td align="center" valign="top"><bold>0.66</bold></td>
<td align="center" valign="top"><bold>0.54</bold></td>
<td align="center" valign="top">0.30</td>
<td align="center" valign="top"><bold>0.50</bold></td>
</tr>
<tr>
<td align="left" valign="top">OP</td>
<td align="center" valign="top">&#x02212;0.6</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">0.22</td>
<td align="center" valign="top">&#x02212;0.04</td>
<td align="center" valign="top">&#x02212;0.14</td>
<td align="center" valign="top">0.12</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">0.17</td>
<td align="center" valign="top">&#x02212;0.24</td>
<td align="center" valign="top">&#x02212;0.14</td>
<td align="center" valign="top">&#x02212;0.07</td>
</tr>
<tr>
<td align="left" valign="top" colspan="12"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="12"><bold>(b) Between-judge</bold></td>
</tr>
<tr>
<td align="left" valign="top">EX</td>
<td align="center" valign="top">0.84&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.81&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.74&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.81&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.72&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.51&#x0002A;</td>
<td align="center" valign="top">0.70&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.72&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.63&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.12</td>
<td align="center" valign="top">0.66&#x0002A;&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">AG</td>
<td align="center" valign="top">0.46&#x0002A;</td>
<td align="center" valign="top">0.61&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.40</td>
<td align="center" valign="top">0.55&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.25</td>
<td align="center" valign="top">&#x02212;0.15</td>
<td align="center" valign="top">0.32</td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">0.54&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.95</td>
<td align="center" valign="top">0.39&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">CO</td>
<td align="center" valign="top">0.78&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.67&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.71&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.72&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.37</td>
<td align="center" valign="top">&#x02212;0.10</td>
<td align="center" valign="top">0.22</td>
<td align="center" valign="top">0.32</td>
<td align="center" valign="top">0.65&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.35</td>
<td align="center" valign="top">0.36&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">NE</td>
<td align="center" valign="top">0.80&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.71&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.55&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.75&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.57&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.12</td>
<td align="center" valign="top">0.55&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.70&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.36</td>
<td align="center" valign="top">&#x02212;0.56</td>
<td align="center" valign="top">0.44&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">OP</td>
<td align="center" valign="top">0.12</td>
<td align="center" valign="top">0.67&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.40</td>
<td align="center" valign="top">0.52&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.49</td>
<td align="center" valign="top">0.40</td>
<td align="center" valign="top">0.55&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.34</td>
<td align="center" valign="top">0.17</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.36&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top" colspan="12"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="12"><bold>(c) Self-other</bold></td>
</tr>
<tr>
<td align="left" valign="top">EX</td>
<td align="center" valign="top">0.34&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.32&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.26&#x0002A;</td>
<td align="center" valign="top">0.30&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.44&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.24&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.12</td>
<td align="center" valign="top">&#x02212;0.02</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.05</td>
</tr>
<tr>
<td align="left" valign="top">AG</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.13</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.28&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.05</td>
<td align="center" valign="top">0.12</td>
<td align="center" valign="top">0.08</td>
<td align="center" valign="top">&#x02212;0.01</td>
<td align="center" valign="top">0.10</td>
<td align="center" valign="top">0.06</td>
</tr>
<tr>
<td align="left" valign="top">CO</td>
<td align="center" valign="top">&#x02212;0.17</td>
<td align="center" valign="top">0.09</td>
<td align="center" valign="top">0.16</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">0.13</td>
<td align="center" valign="top">&#x02212;0.13</td>
<td align="center" valign="top">0.01</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">0.16</td>
<td align="center" valign="top">&#x02212;0.16</td>
<td align="center" valign="top">0.01</td>
</tr>
<tr>
<td align="left" valign="top">NE</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">&#x02212;0.07</td>
<td align="center" valign="top">0.05</td>
<td align="center" valign="top">&#x02212;0.01</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.09</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.02</td>
<td align="center" valign="top">&#x02212;0.08</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.00</td>
</tr>
<tr>
<td align="left" valign="top">OP</td>
<td align="center" valign="top">0.06</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">0.00</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">0.10</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.16</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.03</td>
<td align="center" valign="top">0.09</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>(a) Within-judge consistency in terms of Cronbach&#x02019;s <italic>&#x003B1;</italic> (good reliability&#x02009;&#x0003E;&#x02009;0.80 is highlighted in bold); (b) Between-judge consistency in terms of ICC(1,k) (at a significance level of &#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05, &#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01, &#x0002A;&#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001); (c) Self-other agreement in terms of Pearson correlation (at a significance level of &#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05, &#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01, and &#x0002A;&#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001)</italic>.</p></table-wrap-foot></table-wrap>
</sec>
<sec id="S4-1-3">
<label>4.1.3</label> <title>Between-Judge Consistency</title>
<p>We computed between-judge consistency in terms of intraclass correlation, ICC(1,k) proposed by Shrout and Fleiss (<xref ref-type="bibr" rid="B45">1979</xref>), where <italic>k</italic>&#x02009;&#x0003D;&#x02009;5. Our judge selection method uses the <italic>k</italic> most correlated judges so might bias the ICC results (see Section <xref ref-type="sec" rid="S4-1-1">4.1.1</xref>). To evaluate this, we calculated ICC for <italic>k</italic>&#x02009;&#x0003D;&#x02009;(10, &#x02026;3) for the AV condition. Figure <xref ref-type="fig" rid="F3">3</xref>B shows that, for <italic>extroversion, conscientiousness</italic>, and <italic>neuroticism</italic>, ICC does not change meaningfully as the number of judges varies, while selecting the 5 most correlated judges slightly biases the results for <italic>agreeableness</italic> and <italic>openness</italic>.</p>
<p>The detailed results for the selected 5 judges per clip are presented in Table <xref ref-type="table" rid="T2">2</xref>(b). We obtained significant correlations for most traits in the AV condition, with values in the same range (0.40&#x02009;&#x0003C;&#x02009;<italic>ICC</italic>(1, <italic>k</italic>)&#x02009;&#x0003C;&#x02009;0.81) as reported in the literature for online judges using a 10-item test (0.42&#x02009;&#x0003C;&#x02009;<italic>ICC</italic>(1, <italic>k</italic>)&#x02009;&#x0003C;&#x02009;0.76) (Biel and Gatica-Perez, <xref ref-type="bibr" rid="B9">2013</xref>). Fewer significant correlations were observed in the other communication conditions, particularly in the story task for AO and the mime task for TO. <italic>Extroversion</italic> was the only trait that consistently maintained correlation across conditions.</p>
</sec>
<sec id="S4-1-4">
<label>4.1.4</label> <title>Self-Other Agreement</title>
<p>We examined the extent to which judges agree with the target&#x02019;s self-assessment. Pearson correlations between the self-ratings and the judge&#x02019;s ratings of conditions and tasks are reported in Table <xref ref-type="table" rid="T2">2</xref>(c) for the selected 5 judges per clip. We observed that the judge&#x02019;s ratings bear a significant relation to the target&#x02019;s self-ratings for <italic>extroversion</italic> only (<italic>r</italic>&#x02009;&#x0003D;&#x02009;0.24&#x02009;&#x02212;&#x02009;0.44 and <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05). However, we did not obtain any significant correlations in the TO condition (all <italic>r</italic>&#x02009;&#x0003C;&#x02009;0.2 and <italic>p</italic>&#x02009;&#x0003E;&#x02009;0.05).</p>
</sec>
<sec id="S4-1-5">
<label>4.1.5</label> <title>Personality Shifts</title>
<p>We examined the extent to which people shifted from one personality class to another, in judges&#x02019; perception, between AV and TO conditions, in the hobby and story tasks for the selected 5 judges per clip. We did not examine shifts involving AO or Mime task as the ICC scores indicated that personality ratings in this condition would be too unreliable. These results are presented in Table <xref ref-type="table" rid="T3">3</xref> as 2&#x02009;&#x000D7;&#x02009;2 contingency tables. To aid analysis we have also illustrated each shift as a proportional change (%) both from high to low (HIGH2LOW) and from low to high (LOW2HIGH) in Figure <xref ref-type="fig" rid="F4">4</xref> (see the figure on the left hand side).</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p><bold>Contingency tables for each trait (at a significance level of &#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td valign="top" align="left"><bold>EX</bold></td>
<td valign="top" align="center"><bold>TO: high</bold></td>
<td valign="top" align="center"><bold>TO: low</bold></td>
<td valign="top" align="left"><bold>AG</bold></td>
<td valign="top" align="center"><bold>TO: high</bold></td>
<td valign="top" align="center"><bold>TO: low</bold></td>
<td valign="top" align="left"><bold>CO</bold></td>
<td valign="top" align="center"><bold>TO: high</bold></td>
<td valign="top" align="center"><bold>TO: low</bold></td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">16</td>
<td align="center" valign="top"><bold>6</bold></td>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">16</td>
<td align="center" valign="top"><bold>11</bold></td>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">13</td>
<td align="center" valign="top"><bold>9</bold></td>
</tr>
<tr>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>10</bold></td>
<td align="center" valign="top">8</td>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>5</bold></td>
<td align="center" valign="top">8</td>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>12</bold></td>
<td align="center" valign="top">6</td>
</tr>
<tr>
<td align="left" valign="top" colspan="9"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><bold>NE</bold></td>
<td align="center" valign="top"><bold>TO: high</bold></td>
<td align="center" valign="top"><bold>TO: low</bold></td>
<td align="left" valign="top"><bold>OP</bold></td>
<td align="center" valign="top"><bold>TO: high</bold></td>
<td align="center" valign="top"><bold>TO: low</bold></td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"/>
</tr>
<tr>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">6</td>
<td align="center" valign="top"><bold>14</bold>&#x0002A;</td>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">13</td>
<td align="center" valign="top"><bold>6</bold></td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>1</bold>&#x0002A;</td>
<td align="center" valign="top">19</td>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>12</bold></td>
<td align="center" valign="top">9</td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>Shift between two classes (from high to low or vice versa) are highlighted in bold</italic>.</p></table-wrap-foot></table-wrap>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Amount of shifts (%) from high to low (HIGH2LOW) and from low to high (LOW2HIGH) (&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05, &#x0002A;&#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001) between AV and TO: solo tasks (left hand side) versus dyadic tasks (right hand side)</bold>.</p></caption>
<graphic xlink:href="frobt-04-00016-g004.tif"/>
</fig>
<p>We found a significant shift from high to low for <italic>neuroticism</italic> (70%). Note that the corrected McNemar&#x02019;s test is very conservative in estimating significance, particularly for small sample sizes. Although not statistically significant, we observed large shifts from low to high for <italic>extroversion</italic> (56%), <italic>conscientiousness</italic> (67%), and <italic>openness</italic> (57%).</p>
</sec>
</sec>
<sec id="S4-2">
<label>4.2</label> <title>Dyadic Tasks Study</title>
<p>As in the Solo Tasks Study, we assessed whether there existed low-quality judges (spammers) in the judge pool used for the Dyadic Tasks Study. To do so, we repeated the same method that we used for the Solo Tasks Study, where we evaluated ICC values, and used judge rating techniques to selectively remove judges. These results are presented in Figures <xref ref-type="fig" rid="F3">3</xref>B,D. As we observed ICC values for the AV condition in line with expectation with all judges included, and cannot observe large changes in the Cronbach&#x02019;s <italic>&#x003B1;</italic> values and the ICC values, by excluding judges, we concluded that the judges were reliable. Hence, we present the results for the Dyadic Tasks Study without eliminating any judges.</p>
<sec id="S4-2-1">
<label>4.2.1</label> <title>Within-Judge Consistency</title>
<p>Within-judge consistency was measured in terms of Cronbach&#x02019;s <italic>&#x003B1;</italic>. The detailed results with respect to different communication conditions and tasks are presented in Table <xref ref-type="table" rid="T4">4</xref>(a), where <italic>&#x003B1;</italic> values that indicate sufficient reliability for the IPIP-BFM-20 (greater than 0.75, in line with values reported in the literature (Cred&#x000E9; et al., <xref ref-type="bibr" rid="B19">2012</xref>)) are highlighted in bold. Values are above or close to good reliability (&#x0003E;0.7) for all traits except for <italic>neuroticism</italic>. Comparing values across communication conditions, we observe little difference, hence judges were able to make consistent trait evaluations when the robot is used for communication.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p><bold>Analysis of personality judgments across 2 communication conditions and 3 tasks</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center" colspan="4">Audiovisual (AV)<hr/></th>
<th valign="top" align="center" colspan="4">Teleoperation (TO)<hr/></th>
</tr><tr>
<th valign="top" align="center"/>
<th valign="top" align="center">Informative</th>
<th valign="top" align="center">Competitive</th>
<th valign="top" align="center">Cooperative</th>
<th valign="top" align="center">All</th>
<th valign="top" align="center">Informative</th>
<th valign="top" align="center">Competitive</th>
<th valign="top" align="center">Cooperative</th>
<th valign="top" align="center">All</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" colspan="9"><bold>(a) Within-judge</bold></td>
</tr>
<tr>
<td align="left" valign="top">EX</td>
<td align="center" valign="top"><bold>0.85</bold></td>
<td align="center" valign="top"><bold>0.87</bold></td>
<td align="center" valign="top"><bold>0.85</bold></td>
<td align="center" valign="top"><bold>0.87</bold></td>
<td align="center" valign="top"><bold>0.84</bold></td>
<td align="center" valign="top"><bold>0.85</bold></td>
<td align="center" valign="top"><bold>0.84</bold></td>
<td align="center" valign="top"><bold>0.86</bold></td>
</tr>
<tr>
<td align="left" valign="top">AG</td>
<td align="center" valign="top"><bold>0.77</bold></td>
<td align="center" valign="top"><bold>0.80</bold></td>
<td align="center" valign="top"><bold>0.84</bold></td>
<td align="center" valign="top"><bold>0.83</bold></td>
<td align="center" valign="top"><bold>0.86</bold></td>
<td align="center" valign="top"><bold>0.84</bold></td>
<td align="center" valign="top"><bold>0.81</bold></td>
<td align="center" valign="top"><bold>0.84</bold></td>
</tr>
<tr>
<td align="left" valign="top">CO</td>
<td align="center" valign="top">0.71</td>
<td align="center" valign="top"><bold>0.75</bold></td>
<td align="center" valign="top">0<bold>.77</bold></td>
<td align="center" valign="top">0.74</td>
<td align="center" valign="top"><bold>0.76</bold></td>
<td align="center" valign="top">0.70</td>
<td align="center" valign="top">0.72</td>
<td align="center" valign="top">0.73</td>
</tr>
<tr>
<td align="left" valign="top">NE</td>
<td align="center" valign="top">0.57</td>
<td align="center" valign="top">0.60</td>
<td align="center" valign="top">0.54</td>
<td align="center" valign="top">0.57</td>
<td align="center" valign="top">0.54</td>
<td align="center" valign="top">0.64</td>
<td align="center" valign="top">0.60</td>
<td align="center" valign="top">0.59</td>
</tr>
<tr>
<td align="left" valign="top">OP</td>
<td align="center" valign="top"><bold>0.78</bold></td>
<td align="center" valign="top"><bold>0.82</bold></td>
<td align="center" valign="top"><bold>0.87</bold></td>
<td align="center" valign="top"><bold>0.85</bold></td>
<td align="center" valign="top"><bold>0.75</bold></td>
<td align="center" valign="top"><bold>0.79</bold></td>
<td align="center" valign="top"><bold>0.85</bold></td>
<td align="center" valign="top"><bold>0.81</bold></td>
</tr>
<tr>
<td align="left" valign="top" colspan="9"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="9"><bold>(b) Between-judge</bold></td>
</tr>
<tr>
<td align="left" valign="top">EX</td>
<td align="center" valign="top">0.83&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.84&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.70&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.85&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.61&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.78&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.78&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.82&#x0002A;&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">AG</td>
<td align="center" valign="top">0.18</td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">0.58&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.51&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.08</td>
<td align="center" valign="top">0.35</td>
<td align="center" valign="top">0.37&#x0002A;</td>
<td align="center" valign="top">0.41&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">CO</td>
<td align="center" valign="top">0.27</td>
<td align="center" valign="top">0.28</td>
<td align="center" valign="top">0.48&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.61&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.24</td>
<td align="center" valign="top">&#x02212;0.11</td>
<td align="center" valign="top">0.24</td>
<td align="center" valign="top">&#x02212;0.26</td>
</tr>
<tr>
<td align="left" valign="top">NE</td>
<td align="center" valign="top">0.52&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.53&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.22</td>
<td align="center" valign="top">0.66&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.38&#x0002A;</td>
<td align="center" valign="top">0.13</td>
<td align="center" valign="top">&#x02212;0.35</td>
<td align="center" valign="top">0.46&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">OP</td>
<td align="center" valign="top">0.21</td>
<td align="center" valign="top">0.67&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.57&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.51&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.55&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.47&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.29</td>
<td align="center" valign="top">0.52&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top" colspan="9"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="9"><bold>(c) Self-other</bold></td>
</tr>
<tr>
<td align="left" valign="top">EX</td>
<td align="center" valign="top">0.29&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.12</td>
<td align="center" valign="top">&#x02212;0.29&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.06</td>
<td align="center" valign="top">0.32&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.21&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.15</td>
<td align="center" valign="top">0.18</td>
</tr>
<tr>
<td align="left" valign="top">AG</td>
<td align="center" valign="top">0.74&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.73&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.44&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.75&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.57&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.65&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.27&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.63&#x0002A;&#x0002A;&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">CO</td>
<td align="center" valign="top">0.22&#x0002A;</td>
<td align="center" valign="top">0.28&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.31&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.31&#x0002A;&#x0002A;</td>
<td align="center" valign="top">&#x02212;0.01</td>
<td align="center" valign="top">0.27&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.14</td>
<td align="center" valign="top">0.17</td>
</tr>
<tr>
<td align="left" valign="top">NE</td>
<td align="center" valign="top">0.16</td>
<td align="center" valign="top">0.18</td>
<td align="center" valign="top">0.28&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.24&#x0002A;</td>
<td align="center" valign="top">0.24&#x0002A;</td>
<td align="center" valign="top">0.19</td>
<td align="center" valign="top">0.07</td>
<td align="center" valign="top">0.23&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">OP</td>
<td align="center" valign="top">0.68&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.61&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.17</td>
<td align="center" valign="top">0.71&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.51&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.37&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">0.04</td>
<td align="center" valign="top">0.46&#x0002A;&#x0002A;&#x0002A;</td>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>(a) Intra-judge consistency in terms of Cronbach&#x02019;s <italic>&#x003B1;</italic> (good reliability&#x02009;&#x0003E;&#x02009;0.80 is highlighted in bold); (b) Inter-judge consistency in terms of ICC(1,k) (at a significance level of &#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05, &#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01, &#x0002A;&#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001); (c) Self-other agreement in terms of Pearson correlation (at a significance level of &#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05, &#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01, and &#x0002A;&#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001)</italic>.</p></table-wrap-foot></table-wrap>
</sec>
<sec id="S4-2-2">
<label>4.2.2</label> <title>Between-Judge Consistency</title>
<p>We computed between-judge consistency in terms of intraclass correlation, ICC(1,k), where <italic>k</italic>&#x02009;&#x0003D;&#x02009;10 (Shrout and Fleiss, <xref ref-type="bibr" rid="B45">1979</xref>). The detailed results for the 10 judges per clip are presented in Table <xref ref-type="table" rid="T4">4</xref>(b). <italic>Extroversion</italic> and <italic>openness</italic> are the only traits with significant agreement across most tasks and both conditions (0.47&#x02009;&#x02264;&#x02009;<italic>ICC</italic>(1, <italic>k</italic>)&#x02009;&#x02264;&#x02009;0.85 at a significance level of <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01). Other traits vary between tasks and conditions as to where significant agreement is achieved. A clearer picture can be gained from the all task results, where it can be seen that agreement on <italic>conscientiousness</italic> deteriorates in the TO condition relative to AV (a drastic drop from 0.61 to &#x02212;0.26 over all tasks).</p>
</sec>
<sec id="S4-2-3">
<label>4.2.3</label> <title>Self-Other Agreement</title>
<p>We examined the extent to which judges agree with the target&#x02019;s self-assessment. Pearson correlations between the self-ratings and the judge&#x02019;s ratings of conditions and tasks are reported in Table <xref ref-type="table" rid="T4">4</xref>(c). Significant agreement was found for <italic>agreeableness</italic> and <italic>openness</italic> across most tasks and both conditions (<italic>r<sub>ag</sub></italic>&#x02009;&#x0003D;&#x02009;0.75 and <italic>r<sub>op</sub></italic>&#x02009;&#x0003D;&#x02009;0.71 over all tasks), although agreement is much lower in the TO condition (<italic>r<sub>ag</sub></italic>&#x02009;&#x0003D;&#x02009;0.63 and <italic>r<sub>op</sub></italic>&#x02009;&#x0003D;&#x02009;0.46 over all tasks). For <italic>extroversion</italic> and <italic>neuroticism</italic>, agreement is much lower than for other traits, and this is fairly consistent across conditions. Again we observe the larger difference across conditions for <italic>conscientiousness</italic> (<italic>r<sub>co</sub></italic>&#x02009;&#x0003D;&#x02009;0.17), with almost no significant agreement in the TO condition compared to significant agreement across all tasks in the AV condition (<italic>r<sub>co</sub></italic>&#x02009;&#x0003D;&#x02009;0.31).</p>
</sec>
<sec id="S4-2-4">
<label>4.2.4</label> <title>Personality Shifts</title>
<p>We examined the extent to which people shifted from one personality trait classification to another, in judges&#x02019; perception, between AV and TO conditions for each task. These results are presented in Tables <xref ref-type="table" rid="T3">3</xref> and <xref ref-type="table" rid="T5">5</xref> as 2&#x02009;&#x000D7;&#x02009;2 contingency tables. To aid analysis, we have also illustrated each shift as a proportional change (%) both from high to low (HIGH2LOW) and from low to high (LOW2HIGH) in Figure <xref ref-type="fig" rid="F4">4</xref> (see the figure on the right hand side). We found a significant shift from high to low for <italic>agreeableness</italic> (65%), <italic>conscientiousness</italic> (67%) and <italic>openness</italic> (56%). Although not statistically significant, we observed a large shift from high to low for <italic>neuroticism</italic> (57%).</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p><bold>Contingency tables for each trait (at a significance level of &#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05 and &#x0002A;&#x0002A;&#x0002A;<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td valign="top" align="left"><bold>EX</bold></td>
<td valign="top" align="center"><bold>TO: high</bold></td>
<td valign="top" align="center"><bold>TO: low</bold></td>
<td valign="top" align="left"><bold>AG</bold></td>
<td valign="top" align="center"><bold>TO: high</bold></td>
<td valign="top" align="center"><bold>TO: low</bold></td>
<td valign="top" align="left"><bold>CO</bold></td>
<td valign="top" align="center"><bold>TO: high</bold></td>
<td valign="top" align="center"><bold>TO: low</bold></td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">31</td>
<td align="center" valign="top"><bold>5</bold></td>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">14</td>
<td align="center" valign="top"><bold>26</bold>&#x0002A;&#x0002A;&#x0002A;</td>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">12</td>
<td align="center" valign="top"><bold>24</bold>&#x0002A;</td>
</tr>
<tr>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>13</bold></td>
<td align="center" valign="top">26</td>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>5</bold>&#x0002A;&#x0002A;&#x0002A;</td>
<td align="center" valign="top">30</td>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>10</bold>&#x0002A;</td>
<td align="center" valign="top">29</td>
</tr>
<tr>
<td align="left" valign="top" colspan="9"><hr/></td>
</tr>
<tr>
<td align="left" valign="top"><bold>NE</bold></td>
<td align="center" valign="top"><bold>TO: high</bold></td>
<td align="center" valign="top"><bold>TO: low</bold></td>
<td align="left" valign="top"><bold>OP</bold></td>
<td align="center" valign="top"><bold>TO: high</bold></td>
<td align="center" valign="top"><bold>TO: low</bold></td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"><hr/></td>
<td align="left" valign="top" colspan="3"/>
</tr>
<tr>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">16</td>
<td align="center" valign="top"><bold>21</bold></td>
<td align="left" valign="top">AV: high</td>
<td align="center" valign="top">18</td>
<td align="center" valign="top"><bold>23</bold>&#x0002A;</td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
</tr>
<tr>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>10</bold></td>
<td align="center" valign="top">28</td>
<td align="left" valign="top">AV: low</td>
<td align="center" valign="top"><bold>10</bold>&#x0002A;</td>
<td align="center" valign="top">24</td>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
<td align="center" valign="top"/>
</tr>
</tbody>
</table>
<table-wrap-foot><p><italic>Shift between two classes (from high to low or vice versa) are highlighted in bold</italic>.</p></table-wrap-foot></table-wrap>
</sec>
</sec>
</sec>
<sec id="S5" sec-type="discussion">
<label>5</label> <title>Discussion</title>
<p>In this section, we discuss our results, including comparisons with related work introduced in Section <xref ref-type="sec" rid="S2">2</xref>. We present in-depth discussion of meta-data (i.e., judge ratings, self-ratings) in terms of intra/inter-judge agreement, accuracy of judgments and personality shifts, with regard to different communication conditions (i.e., AO: audio-only, AV: audiovisual, and TO: teleoperation) and different tasks (i.e., solo and dyadic tasks). Note that in the majority of related works results were not directly comparable as personality recognition accuracy is typically the reported metric, as opposed to agreement as used here; accuracy as measured by comparing human responses with machine learning systems (e.g., Aran and Gatica-Perez (<xref ref-type="bibr" rid="B4">2013</xref>), Batrinca et al. (<xref ref-type="bibr" rid="B6">2016</xref>)), or between self-ratings and judge ratings (e.g., Funder (<xref ref-type="bibr" rid="B25">1995</xref>), Borkenau et al. (<xref ref-type="bibr" rid="B11">2004</xref>)). Nevertheless, for which traits this reported accuracy is high or low helps provide some explanation for our findings.</p>
<sec id="S5-1">
<label>5.1</label> <title>Intra-Judge Agreement</title>
<p>Consistency within judges for how each trait is judged (Table <xref ref-type="table" rid="T2">2</xref>(a) and Table <xref ref-type="table" rid="T4">4</xref>(a)) is used to address RQ1. In both studies, judges were sufficiently consistent in their trait ratings in the audiovisual condition (AV), with the exception of <italic>openness</italic> in the Solo Tasks Study, and to a lesser extent <italic>neuroticism</italic> in the Dyadic Tasks Study for us to conclude that the tasks and judges&#x02019; behaviors were reliable. Batrinca et al. (<xref ref-type="bibr" rid="B6">2016</xref>) also reported a similar finding that openness was not modeled successfully in the human-machine interaction, whereas, in the human&#x02013;human interaction setting, it was the only trait that could be predicted with a high accuracy over all collaboration tasks. In our case, the difference between the two studies with regard to consistent judgment of the <italic>openness</italic> trait indicates that cues for this trait may be more evident in dyadic tasks. Some researchers have suggested that one aspect of <italic>openness</italic> is intellect, where intellect incorporates the facets of intelligence, intellectual engagement, and creativity (DeYoung, <xref ref-type="bibr" rid="B21">2011</xref>), and the tasks in the Dyadic Tasks Study are more conducive to displaying these facets.</p>
<p>In the Solo Tasks Study, there were some notable differences between the audio-only (AO) and the teleoperated robot (TO) conditions. For the hobby task, judges remained consistent in both the AO and TO conditions, indicating they were able to use audio cues to make judgments for this task, and robot appearance had no effect on consistency. However, for the story task, judges were much less consistent in the AO than in the AV condition, for all traits except for <italic>agreeableness</italic>. This is in contrast to the teleoperated robot condition (TO), where they remained as consistent as in the AV condition. The only additional cues available with the robot compared to audio only are gestures and appearance. The results indicate that such cues are used to aid judgments in the same way that they do in the AV condition, though their utility appears to be task dependent (only of apparent benefit in the story task). Importantly, the fact that they are utilized provides good evidence that the robot is not simply ignored when making judgments. Hence, the findings of high levels of agreement across both conditions in all tasks in the Dyadic Tasks Study indicate that in dyadic tasks the robot transmits sufficient cues to make judgments as consistently as observing the target directly.</p>
<p>The use of gesture to aid personality judgments appears to be dependent on it accompanying speech, as in the Solo Tasks Study ratings in the TO condition are far less consistent than in the AV condition for the mime task. That is to say, gestures alone do not provide sufficient information for judging personality. This was in contrast to what was reported by Aran and Gatica-Perez (<xref ref-type="bibr" rid="B4">2013</xref>), where the best results were achieved when they used visual cues only for predicting personality traits, and using audio cues or combining them with visual cues resulted in lower accuracy. This showed that either other behavior cues not transmitted by the robot are needed, or appearance cues are used which conflict with gesture cues in the TO condition.</p>
<p>Taking the results from both studies together, it is apparent that judges are able to remain consistent in their judgments of a given trait whether they are observing someone directly or their communication relayed through a teleoperated robot. Indeed, where there are slight shifts in consistency between AV and TO conditions, they are not large; the one exception being for the mime task in the Solo Tasks Study. Hence, each judge appears to formulate a relatively consistent evaluation of a given targets&#x02019; personality traits based on speech, gesture, and appearance, combining them to assess each trait facet. This finding is in contrast to the study by Kuwamura et al. (<xref ref-type="bibr" rid="B32">2012</xref>) where they suggested small shifts in intra-judge consistency provided evidence of robot appearance effects on personality perception. While in subsequent sections we do observe evidence for effects of robot mediation on perception, we do not find such small shifts in intra-judge consistency convincing in this regard.</p>
</sec>
<sec id="S5-2">
<label>5.2</label> <title>Inter-Judge Agreement</title>
<p>Looking at inter-judge agreement results to address RQ2 (Table <xref ref-type="table" rid="T2">2</xref>(b) and Table <xref ref-type="table" rid="T4">4</xref>(b)), <italic>extroversion</italic> was the only trait on which judges reached consensus in both studies, regardless of the communication condition, and task (the mime task in the Solo Tasks Study being the one exception). This result is in line with the widely accepted idea that <italic>extroversion</italic> is the easiest trait to infer upon others (Barrick et al., <xref ref-type="bibr" rid="B5">2000</xref>). Hence, the strength of the available cues was sufficient to overcome any conflict between appearance, vocal, and gesture based cues. Indeed it indicates that judges had a common set of interpretations for the available cues.</p>
<p>On the other hand, where agreement was reached on <italic>agreeableness, conscientiousness</italic>, and <italic>neuroticism</italic> for some tasks in the AV condition in each study, it had mostly deteriorated in the TO condition, and the AO condition in the Solo Tasks Study. The clearest example of this is for <italic>conscientiousness</italic> taking all three tasks together in the Dyadic Tasks Study (and to some extent in the Solo Tasks Study as well), where agreement drastically deteriorated in the TO condition as compared to the AV condition. As explained in the study by Macrae et al. (<xref ref-type="bibr" rid="B34">1996</xref>), physical appearance based impressions (facial and vocal features) are often used in the judgment of <italic>conscientiousness</italic>. In particular, low <italic>conscientiousness</italic> is conveyed by a childlike face (Macrae et al., <xref ref-type="bibr" rid="B34">1996</xref>), which the face of the NAO robot can be considered to have, and this may conflict with the vocal cues of the operator. <italic>Neuroticism</italic> is mainly related to emotions, and <italic>agreeableness</italic> is related to trust, cooperation and sympathy (Zillig et al., <xref ref-type="bibr" rid="B52">2002</xref>), both of which it seems reasonable to suggest judges might perceive as being low for a robot (particularly NAO with its lack of facial expressions), again creating conflicts. It would appear that judges do not have a consistent manner with which to resolve such conflicts.</p>
<p>Task-based analyzes in the Solo Tasks Study show that for <italic>agreeableness</italic> and <italic>conscientiousness</italic> the story task provides sufficient cues for agreement to be maintained in the TO condition, whereas the hobby task does so for <italic>neuroticism</italic>. As agreement being maintained in the TO condition indicates sufficient cues to overcome appearance/behavior conflicts, it is instructive to consider how those tasks might relate to the traits. In telling the story, targets might demonstrate their morality, and relation to others, components of <italic>agreeableness</italic> (Zillig et al., <xref ref-type="bibr" rid="B52">2002</xref>). How well structured and clear the story is could relate to facets of the <italic>conscientiousness</italic> trait. The hobby task on the other hand might demonstrate how self-conscious a person is about their hobby, a facet of <italic>neuroticism</italic> (Zillig et al., <xref ref-type="bibr" rid="B52">2002</xref>). While these two tasks might provide some cues for facets of the traits for which consistency was not maintained, they appear to do so in a way that conflicts with cues related to the robot.</p>
<p>We also compared differences in agreement between the TO and AO conditions in the Solo Tasks Study. Where there is agreement in TO for <italic>agreeableness, conscientiousness</italic>, and <italic>neuroticism</italic>, we found it was greatly reduced for <italic>agreeableness</italic> and <italic>conscientiousness</italic>, and to a lesser extent for <italic>neuroticism</italic>. This provides further evidence that physical cues, be they behavioral or appearance based, are utilized in the TO condition. Again, this appears to be dependent on the presence of speech: in the mime task for the Solo Tasks Study, judges were unable to provide a consistent rating for any trait in the TO condition, in contrast to the consistent ratings for <italic>extroversion, conscientiousness</italic>, and <italic>neuroticism</italic> in the AV condition. A likely reason for this observation is that without vocal cues there is an increased reliance on appearance based cues, often based on stereotypes (Kenny et al., <xref ref-type="bibr" rid="B30">1994</xref>), and judges do not have consistent stereotypes relating to robot appearance.</p>
<p>Batrinca et al. (<xref ref-type="bibr" rid="B6">2016</xref>) showed that the prediction of agreeableness and conscientiousness in the human-machine interaction setting and the prediction of conscientiousness and neuroticism were highly dependent on the collaboration task, where the extroversion trait was the only trait yielding consistent results over all tasks in both settings. Similarly, our task-based analyses in the Dyadic Tasks Study show that in the AV condition, while the cooperative task provided a higher level of agreement for <italic>agreeableness</italic> and <italic>conscientiousness</italic>, the competitive task yielded better results for <italic>neuroticism</italic> and <italic>openness</italic>. Indeed, the results are somewhat expected given the nature of the tasks: the cooperative task was to agree upon how to order five items in a survival scenario, in which participants were expected to exhibit the <italic>agreeableness</italic> facet of personality; the competitive task was more related to creativity and intelligence, that are strongly associated with <italic>openness</italic> (Zillig et al., <xref ref-type="bibr" rid="B52">2002</xref>). Though agreement is lower, it is still maintained for <italic>agreeableness</italic> in the cooperative task and <italic>openness</italic> in the competitive task in the TO condition. This indicates that in these cases, for at least some of the judges, either the vocal cues override the visual cues, or movement cues are utilized (with the vocal cues).</p>
<p>Taken together, the findings from both studies indicate that the ability of judges to make judgments based on a common interpretation of cues is affected not only by communication condition but is also dependent on the task. While in some cases it is apparent that a particular task is conducive to providing more verbal cues than another for a particular trait (as indicated by higher agreement, and inferred from the literature), whether these override the physical cues in the TO condition is hard to predict. Indeed, whether clear cues in the AV condition translate into agreement in the TO condition vary a great deal between all tasks. Hence, it seems reasonable to suggest that whether inter-judge consistency is observed also depends on how much appearance cues are utilized for a given task and trait, and thus how all the cues interact. This complex interaction effect provides strong evidence that personality perception is likely to be altered when communicating <italic>via</italic> a robot, and this depends on what cues are produced.</p>
</sec>
<sec id="S5-3">
<label>5.3</label> <title>Accuracy of Judgments</title>
<p>In order to assess RQ3, we analyzed the extent to which judge ratings correlated with self-ratings provided by target participants (Table <xref ref-type="table" rid="T2">2</xref>(c) and Table <xref ref-type="table" rid="T4">4</xref>(c)). In general in the Solo Tasks Study, there was very little correlation between self and other ratings. This is in contrast to previous findings where they found low, but significant, self-other correlation (0.11&#x02009;&#x02212;&#x02009;0.42) (Carney et al., <xref ref-type="bibr" rid="B16">2007a</xref>). The one exception to this was self-other correlation for <italic>extroversion</italic> in the AV condition. This suggests that participant targets did not present cues relating to their self-perception in the tasks we used, other than for <italic>extroversion</italic> which is commonly reported as the trait with the most available cues. Audio cues were sufficient for this correlation to be maintained in the hobby task in the AO condition, but not in the story task, or in either task in the TO condition.</p>
<p>In contrast to the tasks used in the Solo Tasks Study, the tasks of the Dyadic Tasks Study resulted in self-other agreement for <italic>extroversion, agreeableness, conscientiousness</italic>, and <italic>openness</italic> in the majority of tasks for the AV condition. This indicates that the tasks we used in the Dyadic Tasks Study were better at engendering more naturalistic behavior, and hence personality cues than the tasks in the Solo Tasks Study. Indeed, an important factor in thin slice personality analysis is how easy a person is to judge (Funder, <xref ref-type="bibr" rid="B25">1995</xref>), and people behaving more naturally produce better cues. However, despite these apparently better cues, there was a large reduction in agreement for <italic>conscientiousness, neuroticism</italic>, and <italic>openness</italic> (and to a lesser extent <italic>agreeableness</italic>) in the TO condition relative to the AV condition. This finding combined with those of the Solo Tasks Study suggests that there is a shift in the way personality cues are interpreted caused by their interaction with the appearance of the robot, and the way non-verbal communication cues are reproduced on it.</p>
</sec>
<sec id="S5-4">
<label>5.4</label> <title>Personality Shifts</title>
<p>In order to address RQ4, we analyzed the difference in perceived personality in terms of the occurrences of personality shifts. We principally consider the results from the Dyadic Tasks Study as it provides the more compelling evidence. The main reason for this assertion is that more naturalistic cues appeared to be produced in the Dyadic Tasks Study (see previous section), and we consider such cues and their interaction with the TO condition more ecologically valid. In addition, by being able to consider three tasks rather than the two considered in the Solo Tasks Study we have increased statistical power. The shifts we observed (Figure <xref ref-type="fig" rid="F4">4</xref>) provide evidence that cues related to the robots appearance are incorporated into, or even override personality judgments based on speech. Indeed, this is somewhat to be expected given that (Behrend et al., <xref ref-type="bibr" rid="B7">2012</xref>) observed that, in judgments of suitability, attractiveness of a graphical avatar superseded qualities perceived in an interviewees words.</p>
<p>There are two likely causal factors in the perceived personalities being shifted, first human-based physical appearance stereotypes (inferred from humanlike characteristics of the robot) might be applied, second characteristics related to robots might be applied. Here, we will discuss possible underlying causes for the shifts observed in the Dyadic Tasks Study. In the case of <italic>conscientiousness</italic> and <italic>neuroticism</italic> a childlike face, as the NAO might be considered to have, conveys low ratings for both traits (Borkenau and Liebler, <xref ref-type="bibr" rid="B10">1992</xref>; Macrae et al., <xref ref-type="bibr" rid="B34">1996</xref>). Further, <italic>conscientiousness</italic> and <italic>neuroticism</italic> were also observed to be influenced by face shape in graphical avatars (Fong and Mar, <xref ref-type="bibr" rid="B24">2015</xref>), and as the NAO has a face shape that differs from a human, hence this could lead to distortions in perceptions of these traits. Additionally, <italic>neuroticism</italic> is mainly related to emotions (Zillig et al., <xref ref-type="bibr" rid="B52">2002</xref>), something which robots are rarely considered to have. Also linked to emotions is <italic>openness</italic>, which combined with its other facets of imagination and creativity, might also be reasonably expected to be low for a robot, which could also be considered to have <italic>hard facial linaments</italic>, also linked to low <italic>openness</italic> (Borkenau and Liebler, <xref ref-type="bibr" rid="B10">1992</xref>). The NAO robot could also be considered male in appearance, and male avatars have been found to cue for lower <italic>conscientiousness</italic> and <italic>openness</italic> (Fong and Mar, <xref ref-type="bibr" rid="B24">2015</xref>). Low <italic>agreeableness</italic> is more difficult to rationalize, but one facet is trustworthiness (Zillig et al., <xref ref-type="bibr" rid="B52">2002</xref>), and judges may have perceived using a robot to communicate as less trustworthy. The vocal cues for <italic>extroversion</italic> appeared to be very strong, and this might explain why little influence on this trait was observed.</p>
<p>An important thing to note from these findings is that people appear to be attributing personality stereotypes to NAO for characteristics other than the <italic>extroversion</italic> trait, which has been previously examined (Park et al., <xref ref-type="bibr" rid="B40">2012</xref>; Aly and Tapus, <xref ref-type="bibr" rid="B3">2013</xref>; Celiktutan and Gunes, <xref ref-type="bibr" rid="B18">2015</xref>). Hence, in future work in which a desired personality is to be expressed by an autonomous robot, its appearance based cues must be considered alongside any behavioral cues expressed. We suggest that strong behavioral cues may be required to overcome such stereotypes.</p>
</sec>
<sec id="S5-5">
<label>5.5</label> <title>Conclusion</title>
<p>In this paper, we have shown that judges are able to make personality trait judgments that are as consistent with a robot avatar as when the same people are viewed on video in contrast to past work (Kuwamura et al., <xref ref-type="bibr" rid="B32">2012</xref>). One possible reason for this difference in findings is that our teleoperation system allows reproduction of some non-verbal communication cues on the robot which might improve the ease with which judges can assess personality. Hence, we suggest that it is important for telepresence systems to be able to transmit non-verbal communication cues, whether this be actuation of physical systems, or large enough screens on remote presence devices.</p>
<p>We have shown that the appearance of a teleoperated robot avatar influences how the personality of its controller is perceived, i.e., robot appearance based personality cues are utilized along with cues in the speech of the operators. Hence, the perceived personality of a teleoperator is shifted toward that related to the robot&#x02019;s appearance. In light of these findings, we suggest that robot avatar appearance and behavior be carefully considered relative to the person who will be controlling it, and this needs to be done on an individual basis. Training of operators to produce clear cues, or having some cues appropriate to the operator&#x02019;s personality autonomously generated, might allow some control of appearance effects.</p>
<p>Having the correct robot personality has been found to have a positive effect on interactions with people (Park et al., <xref ref-type="bibr" rid="B40">2012</xref>; Aly and Tapus, <xref ref-type="bibr" rid="B3">2013</xref>; Celiktutan and Gunes, <xref ref-type="bibr" rid="B18">2015</xref>), and our findings also have implications for such autonomous robot personality expression. It is important to consider what appearance cues for personality a robot has, as we have observed humanlike personality inferences, and whether the planned behavioral cues might conflict with them. Cues that work on one platform may not be transferable to another. Additionally, we suggest that future experiments on robots expressing personality need to carefully consider tasks undertaken, as we observed that intra-judge agreement on personality perception was highly task dependent.</p>
</sec>
<sec id="S5-6">
<label>5.6</label> <title>Limitations and Future Work</title>
<p>While this paper provides evidence for how personality perception is affected for people teleoperating a humanoid robot avatar, it has a number of limitations we hope to address in future work.</p>
<p>One area of limitation in our work relates to the movement capabilities of the NAO robot, and the inherent differences with human movement capabilities. Although our previous work showed reproduced gestures are comprehensible (Bremner and Leonards, <xref ref-type="bibr" rid="B14">2015</xref>, <xref ref-type="bibr" rid="B15">2016</xref>), there are clearly appreciable differences in the way some movements are reproduced. Indeed, while these differences have limited affect on perceived meaning, they likely contribute to the observed distortions in personality. The main limitations in this regard are in elbow flexion, movement speed, and wrist and hand motion: the NAO elbow can only bend to &#x0007E;90&#x000B0;, the main effect of which being a reduction in vertical travel of the hand for some gestures; humans are capable of extremely rapid motions that the robot cannot match, consequently it will catch up as best it can, but the usual response will be to not express some motions due to the method of motion processing; wrist flexion and hand shape are clearly of utility in many gestures, and their absence (as well as wrist rotation in study 2) restricts the expression of components of some gestures. These movement restrictions are added to by limitations in the Kinect sensor and software processing: movements that result in hand occlusions can lead to imprecision, as well as noise in the sensor data can lead to some added jitter on the robot (though this is filtered as much as possible).</p>
<p>It is also important to note that robot operators had little to no awareness of the limitations of the robot as none of them had prior experience with NAO, and when in control of it they could not observe its motion. The only instruction given pertaining to system capabilities was to not to rest with the arms flat against the body or behind the back as tracking would be lost. While this resulted in some initial poses that were a bit unnatural (video of which was not used in the studies), participants soon reverted to &#x0201C;normal&#x0201D; behavior. Indeed, qualitative comparison of participants in the dyadic study in each condition (video of participants recorded while they were operating the robot allowed this) reveals little difference in gesturing behavior for the majority of participants. Exceptions were the two participants with prior experience working with robots who moved more than they did face-to-face. In further work, we aim to more closely examine the data for any differences (which may be subtle), and if present test how they contribute to the observed personality distortion effects.</p>
<p>In the study by Celiktutan et al. (<xref ref-type="bibr" rid="B17">2016</xref>), our AV condition results showed that face gestures and head activity play an important role in the recognition of the extroversion, agreeableness and conscientiousness traits. This implies another limitation of the robotic platform used in this study. To convey the teleoperators personality traits more accurately, the robot should portray head pose or facial activity together with audio and arm gestures.</p>
<p>A further limitation is that there are some differences between our two studies, the Dyadic Tasks Study has a slightly different design due to correcting issues we encountered in the Solo Tasks Study, making the study comparison slightly less fair. In particular, we addressed the issue with low-quality judges, by utilizing a different recruiting platform which allowed us to recruit better quality judges, and thus did not require a judge removal process. In the Solo Tasks Study, the issues with low-quality judges meant we used a judge selection method based on the gathered responses. The procedure we used had a slight biasing effect on the between-judge consistency (ICC) result for <italic>agreeableness</italic> and <italic>openness</italic>. This bias means that where ICC values are not significant it is strong evidence that there is either a lack of cues or conflicting cues, as even amongst the most agreeing judges consensus of opinion was not possible. Where there is significant agreement, it indicates there are cues for that trait in the particular task and condition and some judges are able to pick up on these cues. Indeed, Funder points out that there exists good and bad judges of personality (Funder, <xref ref-type="bibr" rid="B25">1995</xref>), and we suggest our selection method allowed us to bias toward good judges. This limits the generalizability of our results to judges more adept at picking up on personality cues. By changing crowdsourcing platforms we were able to remove the need for this selection process in the Dyadic Tasks Study.</p>
<p>In addition to recruiting better quality judges, we also utilized a larger personality questionnaire, making our results more accurate, especially with regard to measuring intra-judge and inter-judge consistency.</p>
<p>In the work reported here, it is not clear how different cues are utilized in the aforementioned personality perception. Given that there was such high variability in affects of robot appearance dependent on the task, it seems likely this is due to differences in use of audio and visual cues. Hence, we intend to analyze in-depth the behaviors of targets relative to their judged personality for different tasks. To facilitate this, we aim to extend our work on automatic personality classification, which can extract and identify useful cues automatically (Celiktutan et al., <xref ref-type="bibr" rid="B17">2016</xref>), and apply it to the recordings from the Dyadic Tasks Study. A comparative cue analysis could not only allow us to gain a better understanding of the causes of personality shifts, but could also be useful in synthesizing robot personality behavioral cues.</p>
</sec>
</sec>
<sec id="S6">
<title>Ethics Statement</title>
<p>This study was carried out in accordance with the recommendations of the ethics committees of the University of the West of England and the University of Cambridge with written informed consent from all subjects. All subjects gave written informed consent in accordance with the Declaration of Helsinki. The protocol was approved by the ethics committees of the University of the West of England and the University of Cambridge.</p>
</sec>
<sec id="S7" sec-type="author-contributor">
<title>Author Contributions</title>
<p>PB: substantial contributions to the conception and design of the work, the acquisition, analysis, and interpretation of data; drafting the work; final approval of the version to be published; and agreement to be accountable. OC: substantial contributions to the conception and design of the work, the acquisition, analysis, and interpretation of data; drafting the work; final approval of the version to be published; and agreement to be accountable. HG: substantial contributions to the design of the work, analysis, and interpretation of data; revising the work critically for important intellectual content; final approval of the version to be published; and agreement to be accountable.</p>
</sec>
<sec id="S8">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was funded by the EPSRC under its IDEAS Factory Sandpits call on Digital Personhood (Grant Ref: EP/L00416X/1).</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Adalgeirsson</surname> <given-names>S. O.</given-names></name> <name><surname>Breazeal</surname> <given-names>C.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;MeBot: a robotic platform for socially embodied telepresence,&#x0201D;</article-title> in <source>Proc. of Int. Conf. Human Robot Interaction</source> (<publisher-loc>Osaka</publisher-loc>: <publisher-name>ACM/IEEE</publisher-name>), <fpage>15</fpage>&#x02013;<lpage>22</lpage>.</citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alibali</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>Effects of visibility between speaker and listener on gesture production: some gestures are meant to be seen</article-title>. <source>J. Mem. Lang.</source> <volume>44</volume>, <fpage>169</fpage>&#x02013;<lpage>188</lpage>.<pub-id pub-id-type="doi">10.1006/jmla.2000.2752</pub-id></citation></ref>
<ref id="B3"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Aly</surname> <given-names>A.</given-names></name> <name><surname>Tapus</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;A model for synthesizing a combined verbal and nonverbal behavior based on personality traits in human-robot interaction,&#x0201D;</article-title> in <source>Proc. of ACM/IEEE Int. Conf. on Human-Robot Interaction</source>, <publisher-loc>Tokyo</publisher-loc>.</citation></ref>
<ref id="B4"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Aran</surname> <given-names>O.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;One of a kind: inferring personality impressions in meetings,&#x0201D;</article-title> in <source>Proc. of ACM Int. Conf. on Multimodal Interaction</source>, <publisher-loc>Sydney</publisher-loc>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barrick</surname> <given-names>M. R.</given-names></name> <name><surname>Patton</surname> <given-names>G. K.</given-names></name> <name><surname>Haugland</surname> <given-names>S. N.</given-names></name></person-group> (<year>2000</year>). <article-title>Accuracy of interviewer judgments of job applicant personality traits</article-title>. <source>Personnel Psychol.</source> <volume>53</volume>, <fpage>925</fpage>&#x02013;<lpage>951</lpage>.<pub-id pub-id-type="doi">10.1111/j.1744-6570.2000.tb02424.x</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Batrinca</surname> <given-names>L.</given-names></name> <name><surname>Mana</surname> <given-names>N.</given-names></name> <name><surname>Lepri</surname> <given-names>B.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name> <name><surname>Pianesi</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>Multimodal personality recognition in collaborative goal-oriented tasks</article-title>. <source>IEEE Trans. Multimedia</source> <volume>18</volume>, <fpage>659</fpage>&#x02013;<lpage>673</lpage>.<pub-id pub-id-type="doi">10.1109/TMM.2016.2522763</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Behrend</surname> <given-names>T.</given-names></name> <name><surname>Toaddy</surname> <given-names>S.</given-names></name> <name><surname>Thompson</surname> <given-names>L. F.</given-names></name> <name><surname>Sharek</surname> <given-names>D. J.</given-names></name></person-group> (<year>2012</year>). <article-title>The effects of avatar appearance on interviewer ratings in virtual employment interviews</article-title>. <source>Comput. Human Behav.</source> <volume>28</volume>, <fpage>2128</fpage>&#x02013;<lpage>2133</lpage>.<pub-id pub-id-type="doi">10.1016/j.chb.2012.06.017</pub-id></citation></ref>
<ref id="B8"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bevan</surname> <given-names>C.</given-names></name> <name><surname>Stanton Fraser</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Shaking hands and cooperation in tele-present human-robot negotiation,&#x0201D;</article-title> in <source>Proc. of Int. Conf. Human Robot Interaction</source> (<publisher-loc>Portland</publisher-loc>: <publisher-name>ACM/IEEE</publisher-name>), <fpage>247</fpage>&#x02013;<lpage>254</lpage>.</citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Biel</surname> <given-names>J.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>The YouTube lens: crowdsourced personality impressions and audiovisual analysis of Vlogs</article-title>. <source>IEEE Trans. Multimedia</source> <volume>15</volume>, <fpage>41</fpage>&#x02013;<lpage>55</lpage>.<pub-id pub-id-type="doi">10.1109/TMM.2012.2225032</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borkenau</surname> <given-names>P.</given-names></name> <name><surname>Liebler</surname> <given-names>A.</given-names></name></person-group> (<year>1992</year>). <article-title>Trait inferences: sources of validity at zero acquaintance</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>62</volume>, <fpage>645</fpage>&#x02013;<lpage>657</lpage>.<pub-id pub-id-type="doi">10.1037/0022-3514.62.4.645</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borkenau</surname> <given-names>P.</given-names></name> <name><surname>Mauer</surname> <given-names>N.</given-names></name> <name><surname>Riemann</surname> <given-names>R.</given-names></name> <name><surname>Spinath</surname> <given-names>F. M.</given-names></name> <name><surname>Angleitner</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Thin slices of behavior as cues of personality and intelligence</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>86</volume>, <fpage>599</fpage>&#x02013;<lpage>614</lpage>.<pub-id pub-id-type="doi">10.1037/0022-3514.86.4.599</pub-id><pub-id pub-id-type="pmid">15053708</pub-id></citation></ref>
<ref id="B12"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Bremner</surname> <given-names>P.</given-names></name> <name><surname>Celiktutan</surname> <given-names>O.</given-names></name> <name><surname>Gunes</surname> <given-names>H.</given-names></name></person-group> (<year>2016a</year>). <article-title>&#x0201C;Personality perception of robot avatar tele-operators,&#x0201D;</article-title> in <conf-name>The Eleventh ACM/IEEE International Conference on Human Robot Interaction, HRI &#x02019;16</conf-name> (<conf-loc>Christchurch</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>), <fpage>141</fpage>&#x02013;<lpage>148</lpage>.</citation></ref>
<ref id="B13"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bremner</surname> <given-names>P.</given-names></name> <name><surname>Koschate</surname> <given-names>M.</given-names></name> <name><surname>Levine</surname> <given-names>M.</given-names></name></person-group> (<year>2016b</year>). <article-title>&#x0201C;Humanoid robot avatars: an &#x02018;in the wild&#x02019; usability study,&#x0201D;</article-title> in <source>RO-MAN</source> (<publisher-loc>New Zealand</publisher-loc>: <publisher-name>IEEE</publisher-name>).</citation></ref>
<ref id="B14"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bremner</surname> <given-names>P.</given-names></name> <name><surname>Leonards</surname> <given-names>U.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Efficiency of speech and iconic gesture integration for robotic and human communicators &#x02013; a direct comparison,&#x0201D;</article-title> in <source>Proc. of IEEE Int. Conf. on Robotics and Automation</source> (<publisher-loc>Seattle</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1999</fpage>&#x02013;<lpage>2006</lpage>.</citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bremner</surname> <given-names>P.</given-names></name> <name><surname>Leonards</surname> <given-names>U.</given-names></name></person-group> (<year>2016</year>). <article-title>Iconic gestures for robot avatars, recognition and integration with speech</article-title>. <source>Front. Psychol.</source> <volume>7</volume>:<fpage>183</fpage>.<pub-id pub-id-type="doi">10.3389/fpsyg.2016.00183</pub-id><pub-id pub-id-type="pmid">26925010</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carney</surname> <given-names>D. R.</given-names></name> <name><surname>Colvin</surname> <given-names>C. R.</given-names></name> <name><surname>Hall</surname> <given-names>J. A.</given-names></name></person-group> (<year>2007a</year>). <article-title>A thin slice perspective on the accuracy of first impressions</article-title>. <source>J. Res. Pers.</source> <volume>41</volume>, <fpage>1054</fpage>&#x02013;<lpage>1072</lpage>.<pub-id pub-id-type="doi">10.1016/j.jrp.2007.01.004</pub-id></citation></ref>
<ref id="B17"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Celiktutan</surname> <given-names>O.</given-names></name> <name><surname>Bremner</surname> <given-names>P.</given-names></name> <name><surname>Gunes</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Personality classification from robot-mediated communication cues,&#x0201D;</article-title> in <source>25th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)</source>, <publisher-loc>New York</publisher-loc>.</citation></ref>
<ref id="B18"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Celiktutan</surname> <given-names>O.</given-names></name> <name><surname>Gunes</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Computational analysis of human-robot interactions through first-person vision: personality and interaction experience,&#x0201D;</article-title> in <source>24th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)</source> (<publisher-loc>Kobe</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>815</fpage>&#x02013;<lpage>820</lpage>.</citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cred&#x000E9;</surname> <given-names>M.</given-names></name> <name><surname>Harms</surname> <given-names>P.</given-names></name> <name><surname>Niehorster</surname> <given-names>S.</given-names></name> <name><surname>Gaye-Valentine</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>An evaluation of the consequences of using short measures of the big five personality traits</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>102</volume>, <fpage>874</fpage>&#x02013;<lpage>888</lpage>.<pub-id pub-id-type="doi">10.1037/a0027403</pub-id><pub-id pub-id-type="pmid">22352328</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Daly-Jones</surname> <given-names>O.</given-names></name> <name><surname>Monk</surname> <given-names>A.</given-names></name> <name><surname>Watts</surname> <given-names>L.</given-names></name></person-group> (<year>1998</year>). <article-title>Some advantages of video conferencing over high-quality audio conferencing: fluency and awareness of attentional focus</article-title>. <source>Int. J. Human Comput. Stud.</source> <volume>49</volume>, <fpage>21</fpage>&#x02013;<lpage>58</lpage>.<pub-id pub-id-type="doi">10.1006/ijhc.1998.0195</pub-id></citation></ref>
<ref id="B21"><citation citation-type="book"><person-group person-group-type="author"><name><surname>DeYoung</surname> <given-names>C. D.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Intelligence and personality,&#x0201D;</article-title> in <source>The Cambridge Handbook of Intelligence</source>, eds <person-group person-group-type="editor"><name><surname>Sternberg</surname> <given-names>R. J.</given-names></name> <name><surname>Kaufman</surname> <given-names>S. B.</given-names></name></person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>), <fpage>711</fpage>&#x02013;<lpage>737</lpage>.</citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Edwards</surname> <given-names>A. L.</given-names></name></person-group> (<year>1948</year>). <article-title>Note on the correction for continuity in testing the significance of the difference between correlated proportions</article-title>. <source>Psychometrika</source> <volume>13</volume>, <fpage>185</fpage>&#x02013;<lpage>187</lpage>.<pub-id pub-id-type="doi">10.1007/BF02289261</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feldt</surname> <given-names>L. S.</given-names></name> <name><surname>Woodruff</surname> <given-names>D. J.</given-names></name> <name><surname>Salih</surname> <given-names>F. A.</given-names></name></person-group> (<year>1987</year>). <article-title>Statistical inference for coefficient alpha</article-title>. <source>Appl. Psychol. Measure.</source> <volume>11</volume>, <fpage>93</fpage>&#x02013;<lpage>103</lpage>.<pub-id pub-id-type="doi">10.1177/014662168701100107</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fong</surname> <given-names>K.</given-names></name> <name><surname>Mar</surname> <given-names>R. A.</given-names></name></person-group> (<year>2015</year>). <article-title>What does my avatar say about me? Inferring personality from avatars</article-title>. <source>Pers. Soc. Psychol. Bull.</source> <volume>41</volume>, <fpage>237</fpage>&#x02013;<lpage>249</lpage>.<pub-id pub-id-type="doi">10.1177/0146167214562761</pub-id><pub-id pub-id-type="pmid">25576173</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funder</surname> <given-names>D. C.</given-names></name></person-group> (<year>1995</year>). <article-title>On the accuracy of personality judgment: a realistic approach</article-title>. <source>Psychol. Rev.</source> <volume>102</volume>, <fpage>652</fpage>&#x02013;<lpage>670</lpage>.<pub-id pub-id-type="doi">10.1037/0033-295X.102.4.652</pub-id><pub-id pub-id-type="pmid">7480467</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funder</surname> <given-names>D. C.</given-names></name> <name><surname>Furr</surname> <given-names>R. M.</given-names></name> <name><surname>Colvin</surname> <given-names>C. R.</given-names></name></person-group> (<year>2000</year>). <article-title>The riverside behavioral q-sort: a tool for the description of social behavior</article-title>. <source>J. Pers.</source> <volume>68</volume>, <fpage>451</fpage>&#x02013;<lpage>489</lpage>.<pub-id pub-id-type="doi">10.1111/1467-6494.00103</pub-id><pub-id pub-id-type="pmid">10831309</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funder</surname> <given-names>D. C.</given-names></name> <name><surname>Sneed</surname> <given-names>C. D.</given-names></name></person-group> (<year>1993</year>). <article-title>Behavioral manifestations of personality: an ecological approach to judgmental accuracy</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>64</volume>, <fpage>479</fpage>&#x02013;<lpage>490</lpage>.<pub-id pub-id-type="doi">10.1037/0022-3514.64.3.479</pub-id><pub-id pub-id-type="pmid">8468673</pub-id></citation></ref>
<ref id="B28"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Gouaillier</surname> <given-names>D.</given-names></name> <name><surname>Hugel</surname> <given-names>V.</given-names></name> <name><surname>Blazevic</surname> <given-names>P.</given-names></name> <name><surname>Kilner</surname> <given-names>C.</given-names></name> <name><surname>Monceaux</surname> <given-names>J.</given-names></name> <name><surname>Lafourcade</surname> <given-names>P.</given-names></name> <etal/></person-group> (<year>2009</year>). <article-title>&#x0201C;Mechatronic design of NAO humanoid,&#x0201D;</article-title> in <source>Proc of IEEE Int. Conf. on Robotics and Automation</source> (<publisher-loc>Kobe</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>769</fpage>&#x02013;<lpage>774</lpage>.</citation></ref>
<ref id="B29"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Hossen Mamode</surname> <given-names>H. Z.</given-names></name> <name><surname>Bremner</surname> <given-names>P.</given-names></name> <name><surname>Pipe</surname> <given-names>A. G.</given-names></name> <name><surname>Carse</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Cooperative tabletop working for humans and humanoid robots: group interaction with an avatar,&#x0201D;</article-title> in <source>IEEE Int. Conf. on Robotics and Automation</source> (<publisher-loc>Karlsruhe</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>184</fpage>&#x02013;<lpage>190</lpage>.</citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kenny</surname> <given-names>D. A.</given-names></name> <name><surname>Albright</surname> <given-names>L.</given-names></name> <name><surname>Malloy</surname> <given-names>T. E.</given-names></name> <name><surname>Kashy</surname> <given-names>D. A.</given-names></name></person-group> (<year>1994</year>). <article-title>Consensus in interpersonal perception: acquaintance and the big five</article-title>. <source>Psychol. Bull.</source> <volume>116</volume>, <fpage>245</fpage>&#x02013;<lpage>258</lpage>.<pub-id pub-id-type="doi">10.1037/0033-2909.116.2.245</pub-id><pub-id pub-id-type="pmid">7972592</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kristoffersson</surname> <given-names>A.</given-names></name> <name><surname>Coradeschi</surname> <given-names>S.</given-names></name> <name><surname>Loutfi</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>A review of mobile robotic telepresence</article-title>. <source>Adv. Human-Comput. Interact.</source> <volume>2013</volume>, <fpage>17</fpage>.<pub-id pub-id-type="doi">10.1155/2013/902316</pub-id></citation></ref>
<ref id="B32"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Kuwamura</surname> <given-names>K.</given-names></name> <name><surname>Minato</surname> <given-names>T.</given-names></name> <name><surname>Nishio</surname> <given-names>S.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Personality distortion in communication through teleoperated robots,&#x0201D;</article-title> in <source>Proc of IEEE Int. Symp. on Robot and Human Interactive Communication</source> (<publisher-loc>Paris</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>49</fpage>&#x02013;<lpage>54</lpage>.</citation></ref>
<ref id="B33"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>M. K.</given-names></name> <name><surname>Takayama</surname> <given-names>L.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Now, I have a body,&#x0201D;</article-title> in <source>Proc. of the Conf. on Human Factors in Computing Systems</source> (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>ACM Press</publisher-name>), <fpage>33</fpage>.</citation></ref>
<ref id="B34"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Macrae</surname> <given-names>C. N.</given-names></name> <name><surname>Stangor</surname> <given-names>C.</given-names></name> <name><surname>Hewstone</surname> <given-names>M.</given-names></name></person-group> (<year>1996</year>). <source>Stereotypes and Stereotyping</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>The Guilford Press</publisher-name>).</citation></ref>
<ref id="B35"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Martins</surname> <given-names>H.</given-names></name> <name><surname>Ventura</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Immersive 3-d teleoperation of a search and rescue robot using a head-mounted display,&#x0201D;</article-title> in <source>IEEE Conf. on Emerging Technologies Factory Automation (ETFA)</source> (<publisher-loc>Mallorca</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McKeown</surname> <given-names>G.</given-names></name> <name><surname>Valstar</surname> <given-names>M.</given-names></name> <name><surname>Cowie</surname> <given-names>R.</given-names></name> <name><surname>Pantic</surname> <given-names>M.</given-names></name> <name><surname>Schroder</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>The semaine database: annotated multimodal records of emotionally colored conversations between a person and a limited agent</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>3</volume>, <fpage>5</fpage>&#x02013;<lpage>17</lpage>.<pub-id pub-id-type="doi">10.1109/T-AFFC.2011.20</pub-id></citation></ref>
<ref id="B37"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Murray</surname> <given-names>H. A.</given-names></name></person-group> (<year>1943</year>). <source>Thematic Apperception Test</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Harvard University Press</publisher-name>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naumann</surname> <given-names>L. P.</given-names></name> <name><surname>Vazire</surname> <given-names>S.</given-names></name> <name><surname>Rentfrow</surname> <given-names>P. J.</given-names></name> <name><surname>Gosling</surname> <given-names>S. D.</given-names></name></person-group> (<year>2009</year>). <article-title>Personality judgments based on physical appearance</article-title>. <source>Pers. Soc. Psychol. Bull.</source> <volume>35</volume>, <fpage>1661</fpage>&#x02013;<lpage>1671</lpage>.<pub-id pub-id-type="doi">10.1177/0146167209346309</pub-id><pub-id pub-id-type="pmid">19762717</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x02019;Conaill</surname> <given-names>B.</given-names></name> <name><surname>Whittaker</surname> <given-names>S.</given-names></name> <name><surname>Wilbur</surname> <given-names>S.</given-names></name></person-group> (<year>1993</year>). <article-title>Conversations over video conferences: an evaluation of the spoken aspects of video-mediated communication</article-title>. <source>Human Comput. Interact.</source> <volume>8</volume>, <fpage>389</fpage>&#x02013;<lpage>428</lpage>.<pub-id pub-id-type="doi">10.1207/s15327051hci0804_4</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>E.</given-names></name> <name><surname>Jin</surname> <given-names>D.</given-names></name> <name><surname>del Pobil</surname> <given-names>A. P.</given-names></name></person-group> (<year>2012</year>). <article-title>The law of attraction in human-robot interaction</article-title>. <source>Int. J. Adv. Rob. Syst.</source> <volume>9</volume>, <fpage>1</fpage>&#x02013;<lpage>7</lpage>.<pub-id pub-id-type="doi">10.5772/50228</pub-id></citation></ref>
<ref id="B41"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Rae</surname> <given-names>I.</given-names></name> <name><surname>Takayama</surname> <given-names>L.</given-names></name> <name><surname>Mutlu</surname> <given-names>B.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;In-body experiences,&#x0201D;</article-title> in <conf-name>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems &#x02013; CHI &#x02019;13</conf-name> (<conf-loc>New York, NY</conf-loc>: <conf-sponsor>ACM Press</conf-sponsor>), <fpage>1921</fpage>&#x02013;<lpage>1930</lpage>.</citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rammstedt</surname> <given-names>B.</given-names></name> <name><surname>John</surname> <given-names>O. P.</given-names></name></person-group> (<year>2007</year>). <article-title>Measuring personality in one minute or less: a 10-item short version of the big five inventory in English and German</article-title>. <source>J. Res. Pers.</source> <volume>41</volume>, <fpage>203</fpage>&#x02013;<lpage>212</lpage>.<pub-id pub-id-type="doi">10.1016/j.jrp.2006.02.001</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riggio</surname> <given-names>R. E.</given-names></name> <name><surname>Friedman</surname> <given-names>H. S.</given-names></name></person-group> (<year>1986</year>). <article-title>Impression formation: the role of expressive behavior</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>50</volume>, <fpage>421</fpage>&#x02013;<lpage>427</lpage>.<pub-id pub-id-type="doi">10.1037/0022-3514.50.2.421</pub-id><pub-id pub-id-type="pmid">3517289</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salam</surname> <given-names>H.</given-names></name> <name><surname>Celiktutan</surname> <given-names>O.</given-names></name> <name><surname>Hupont</surname> <given-names>I.</given-names></name> <name><surname>Gunes</surname> <given-names>H.</given-names></name> <name><surname>Chetouani</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Fully automatic analysis of engagement and its relationship to personality in human-robot interactions</article-title>. <source>IEEE Access</source> <volume>5</volume>, <fpage>705</fpage>&#x02013;<lpage>721</lpage>.<pub-id pub-id-type="doi">10.1109/ACCESS.2016.2614525</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shrout</surname> <given-names>P.</given-names></name> <name><surname>Fleiss</surname> <given-names>J.</given-names></name></person-group> (<year>1979</year>). <article-title>Intraclass correlations: uses in assessing rater reliability</article-title>. <source>Psychol. Bull.</source> <volume>86</volume>, <fpage>420</fpage>&#x02013;<lpage>428</lpage>.<pub-id pub-id-type="doi">10.1037/0033-2909.86.2.420</pub-id><pub-id pub-id-type="pmid">18839484</pub-id></citation></ref>
<ref id="B46"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Straub</surname> <given-names>I.</given-names></name> <name><surname>Nishio</surname> <given-names>S.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Incorporated identity in interaction with a teleoperated android robot: a case study,&#x0201D;</article-title> in <source>Proc of Int. Symp. in Robot and Human Interactive Communication</source> (<publisher-loc>Viareggio</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>119</fpage>&#x02013;<lpage>124</lpage>.</citation></ref>
<ref id="B47"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>A.</given-names></name> <name><surname>Boyle</surname> <given-names>M.</given-names></name> <name><surname>Greenberg</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <article-title>&#x0201C;Display and presence disparity in mixed presence groupware,&#x0201D;</article-title> in <source>Proc. of Australasian User Interface Conf</source> (<publisher-loc>Dunedin</publisher-loc>: <publisher-name>Australian Computer Society, Inc.</publisher-name>), <fpage>73</fpage>&#x02013;<lpage>82</lpage>.</citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Topolewska</surname> <given-names>E.</given-names></name> <name><surname>Skiminia</surname> <given-names>E.</given-names></name> <name><surname>Strus</surname> <given-names>W.</given-names></name> <name><surname>Cieciuch</surname> <given-names>J.</given-names></name> <name><surname>Rowinski</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>The short ipip-bfm-20 questionnaire for measuring the big five</article-title>. <source>Ann. Psychol.</source> <volume>2</volume>, <fpage>385</fpage>&#x02013;<lpage>402</lpage>.</citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vinciarelli</surname> <given-names>A.</given-names></name> <name><surname>Mohammadi</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>A survey of personality computing</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>5</volume>, <fpage>273</fpage>&#x02013;<lpage>291</lpage>.<pub-id pub-id-type="doi">10.1109/TAFFC.2014.2330816</pub-id></citation></ref>
<ref id="B50"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Geigel</surname> <given-names>J.</given-names></name> <name><surname>Herbert</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Reading personality: avatar vs. human faces,&#x0201D;</article-title> in <source>Proc. of HAC Conf. on Affective Computing and Intelligent Interaction</source> (<publisher-loc>Geneva</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>479</fpage>&#x02013;<lpage>484</lpage>.</citation></ref>
<ref id="B51"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Yamazaki</surname> <given-names>R.</given-names></name> <name><surname>Nishio</surname> <given-names>S.</given-names></name> <name><surname>Ogawa</surname> <given-names>K.</given-names></name> <name><surname>Ishigur</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Teleoperated android as an embodied communication medium: a case study with demented elderlies in a care facility,&#x0201D;</article-title> in <source>RO-MAN</source> (<publisher-loc>Paris</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1066</fpage>&#x02013;<lpage>1071</lpage>.</citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zillig</surname> <given-names>L. M. P.</given-names></name> <name><surname>Hemenover</surname> <given-names>S. H.</given-names></name> <name><surname>Dienstbier</surname> <given-names>R. A.</given-names></name></person-group> (<year>2002</year>). <article-title>What do we assess when we assess a big 5 trait? A content analysis of the affective, behavioral, and cognitive processes represented in big 5 personality inventories</article-title>. <source>Pers. Soc. Psychol. Bull.</source> <volume>28</volume>, <fpage>847</fpage>&#x02013;<lpage>858</lpage>.<pub-id pub-id-type="doi">10.1177/0146167202289013</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup>Product of <uri xlink:href="http://polhemus.com/">http://polhemus.com/</uri>.</p></fn>
<fn id="fn2"><p><sup>2</sup>Image used was <uri xlink:href="https://www.flickr.com/photos/bassclarinetist/">https://www.flickr.com/photos/bassclarinetist/</uri>, used under creative commons licence.</p></fn>
<fn id="fn3"><p><sup>3</sup>CrowdFlower, a data enrichment, data mining and crowdsourcing company, <uri xlink:href="http://www.crowdflower.com/">http://www.crowdflower.com/</uri>.</p></fn>
<fn id="fn4"><p><sup>4</sup>Prolific Academic online crowd sourcing platform, <uri xlink:href="https://www.prolific.ac/">https://www.prolific.ac/</uri>.</p></fn>
</fn-group>
</back>
</article>
