<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frobt.2017.00049</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Novel Speech Motion Generation by Modeling Dynamics of Human Speech Production</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Sakai</surname> <given-names>Kurima</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x0002A;</xref>
<uri xlink:href="http://frontiersin.org/people/u/360663"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Minato</surname> <given-names>Takashi</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/123778"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ishi</surname> <given-names>Carlos T.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/485091"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ishiguro</surname> <given-names>Hiroshi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://frontiersin.org/people/u/148426"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Graduate School of Engineering Science, Osaka University</institution>, <addr-line>Toyonaka</addr-line>, <country>Japan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Advanced Telecommunications Research Institute International</institution>, <addr-line>Keihanna Science City</addr-line>, <country>Japan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Francesco Becchi, Telerobot Labs s.r.l., Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Felix Reinhart, Bielefeld University, Germany; Manfred Hild, Beuth University of Applied Sciences, Germany</p></fn>
<corresp content-type="corresp" id="cor1">&#x0002A;Correspondence: Kurima Sakai, <email>kurima.sakai&#x00040;atr.jp</email></corresp>
<fn fn-type="other" id="fn001"><p>Specialty section: This article was submitted to Humanoid Robotics, a section of the journal Frontiers in Robotics and AI</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>10</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date><volume>4</volume>
<elocation-id>49</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>07</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>09</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Sakai, Minato, Ishi and Ishiguro.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Sakai, Minato, Ishi and Ishiguro</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>We developed a method to automatically generate humanlike trunk motions based on speech (i.e., the neck and waist motions involved in speech) for a conversational android from its speech in real time. To generate humanlike movements, the android&#x02019;s mechanical limitation (i.e., limited number of joints) needs to be compensated for. By enforcing the synchronization of speech and motion in the android, the method enables us to compensate for its mechanical limitations. Moreover, motion can be modulated to express emotions by tuning the parameters in the dynamical model. This method is based on a spring-damper dynamical model driven by voice features to simulate the human trunk movements involved in speech. In contrast to the existing methods based on machine learning, our system can easily modulate the motions generated due to speech patterns because the model&#x02019;s parameters correspond to muscle stiffness. The experimental results show that the android motions generated by our model can be perceived as more natural and thus motivate users to talk longer with it compared to a system that simply copies human motions. In addition, our model generates emotional speech motions by tuning its parameters.</p>
</abstract>
<kwd-group>
<kwd>humanlike motion</kwd>
<kwd>speech-driven system</kwd>
<kwd>head motion</kwd>
<kwd>android</kwd>
<kwd>emotional motion</kwd>
</kwd-group>
<contract-num rid="cn01">JPMJER1401</contract-num>
<contract-sponsor id="cn01">Japan Science and Technology Agency<named-content content-type="fundref-id">10.13039/501100002241</named-content></contract-sponsor>
<counts>
<fig-count count="15"/>
<table-count count="1"/>
<equation-count count="10"/>
<ref-count count="48"/>
<page-count count="14"/>
<word-count count="10134"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<label>1</label> <title>Introduction</title>
<p>Humanoid robots, especially android robots, are expected to join daily human activities since they can interact with people in a humanlike manner. Androids that resemble humans are suitable for social roles that require rapport and reliability (Prakash and Rogers, <xref ref-type="bibr" rid="B35">2014</xref>), and recent studies have used them as a guide at an event site (Kondo et al., <xref ref-type="bibr" rid="B22">2013</xref>), a salesperson at a department store (Watanabe et al., <xref ref-type="bibr" rid="B43">2015</xref>), a bystander in a medical diagnosis (Yoshikawa et al., <xref ref-type="bibr" rid="B46">2011</xref>), and a receptionist (Hashimoto and Kobayashi, <xref ref-type="bibr" rid="B15">2009</xref>). One critical issue in developing androids is the design of behaviors that people can accept. People tend to expect an agent&#x02019;s behaviors to be based on its appearance (Komatsu and Yamada, <xref ref-type="bibr" rid="B21">2011</xref>); accordingly, humanlike or natural behaviors are expected from androids. Here, natural means humanlike since their appearance is very humanlike and humanlike motion matches with the androids. In other words, people tend to have negative impressions of androids when they do not show the expected humanlike motions. However, an android actually suffers from a clear technical limitation originating in its mechanics; it is not possible to produce motions that are exactly the same as those of humans. The number of controllable joints, the range of joint motions, and achievable joint velocity are limited, and thus an android&#x02019;s motion is much less complex than a human&#x02019;s motion.</p>
<p>This issue can be serious for an android. To effectively exploit its humanlike characteristics, androids must be designed with motions that can foster rapport and reliability. This requires the capability to express subtle changes of facial expressions and body motions. Concerning nonverbal behavior, a human conveys information to a partner not only by gestures but also by varying such motions slightly according to emotions or attitudes. Therefore, in the design of android motions, we must solve the following issues:
<list list-type="bullet">
<list-item><p>The android needs to produce humanlike motions despite its mechanical limitations (e.g., restricted degrees of freedom).</p></list-item>
<list-item><p>The motions must be modulated (i.e., changeable motion properties) based on the android&#x02019;s internal state (i.e., expressing emotions or attitudes).</p></list-item>
</list></p>
<p>For the first issue, androids must provide humanlike impressions even though their motions are not exactly identical as those of humans. To develop a model that can generate such movements, this study focuses on the synchronization between such multimodal expressions as speech and body movement. By enforcing multimodal synchronization in an android, we hypothesize that people will feel a motion&#x02019;s human-likeness even though it is mechanically restricted. This idea comes from existing knowledge that the multiplicative integration of multimodal signals is an effective cognitive strategy for humans to recognize an object&#x02019;s material or texture (e.g., Ernst and Banks, <xref ref-type="bibr" rid="B8">2002</xref>; Fujisaki et al., <xref ref-type="bibr" rid="B12">2014</xref>). In the development of a motion generation model, we also need to consider how to evaluate the human-likeness of the generated motion. If we can reproduce human motions in androids, we can develop a model by minimizing the measurable differences between the generated motion and the original human motion. Unfortunately, this approach does not achieve humanlike motions due to the mechanical limitations. Therefore, this study develops a model based on subjective human evaluations; we selected model parameters whose generated motions were subjectively judged for their human-likeness. The second issue can be solved by parameterizing the motion generation. Accordingly, this study develops an analytical model with parameters that influences the characteristics of the generated motions.</p>
<p>This study&#x02019;s target is a conversational android that mainly talks with people. It needs to evoke and maintain people&#x02019;s motivation to talk with it by creating a sense of reliability. Moreover, it needs to give the impression that it is producing its own utterances. It should also produce movements involved in speaking (speech motion). In addition, such speech motions must be changeable based on the android&#x02019;s current emotion. This study proposes a method of generating speech motions by modeling the dynamics of the human&#x02019;s trunk (neck and waist) motions involved in speech production. The synchronization between multimodal expressions (speech and body movements) improves the naturalness of the android&#x02019;s motion and reduces the negative impressions of non-complex motions caused by its mechanical limitations. People might accept that the android itself is speaking if the movements of its lips, neck, chest, and abdomen, which are normally involved in human vocalization, are produced. This approach also allows motion variation by tuning the parameters in the dynamical model. The motion can be easily modulated to express the desired emotion since the parameters can be associated with motion changes owing to various mental states. Another important issue is real-time processing. In some cases, the utterance data cannot be prepared in advance for an autonomous robot: that is, the motions cannot be generated in advance. We design motion generation to occur simultaneously when the android speaks.</p>
<p>To find which speech and head motion features are synchronized during speaking, we first investigates human speech motions in Section <xref ref-type="sec" rid="S3">3</xref>. In Section <xref ref-type="sec" rid="S4">4</xref>, we develop a motion generation model that enforces the synchronization found in Section <xref ref-type="sec" rid="S3">3</xref>. The naturalness of generated motion needs to be subjectively evaluated. In Section <xref ref-type="sec" rid="S5">5</xref>, a psychological experiment verifies that the proposed model can generate android speech motions that can motivate people to talk with it. Section <xref ref-type="sec" rid="S6">6</xref> shows that speech motion can be changed to express emotions by tuning the model parameters.</p>
</sec>
<sec id="S2">
<label>2</label> <title>Related Works</title>
<p>Many studies have tackled the issue of human motion transfer in humanoid robots while overcoming the limitation of robot kinematics. For example, Pollard et al. (<xref ref-type="bibr" rid="B34">2002</xref>) proposed a method to scale the joint angles and velocities of measured human motions to the capabilities of a humanoid robot. Their robot successfully mimicked the dancing motions of performers while preserving their movement styles. Even though those studies basically aimed to reproduce a human performer&#x02019;s motion, they could not modulate the motions and/or mix them with other motions according to a given situation. Furthermore, most studies have focused on transferring a motion&#x02019;s gestural meaning and not its human-likeness or changes of motion properties.</p>
<p>Another approach to the human-likeness of android motion is multimodal expression, that is, displaying motions synchronized with other expressions, such as speech. Salem et al. (<xref ref-type="bibr" rid="B39">2012</xref>) proposed that generated a pointing gesture by a communication robot that matches its utterance (for example, the robot points to a vase and says &#x0201C;pick up that vase&#x0201D;). This study showed the importance of gesture-level synchronization. Sakai et al. (<xref ref-type="bibr" rid="B38">2015</xref>) made a robot&#x02019;s head motion more natural by enforcing the synchronization of speech and motion. Their method automatically added head gestures (e.g., nodding and tilting) to the robot motions transferred from human motions, where the head gestures were synchronized with a human&#x02019;s speech acts. In their experiments with their method, people had more natural impressions of the robot motion than with the transferred motion. This result suggests that enforcing the synchronization of speech and motion improves the naturalness and human-likeness of an android&#x02019;s motion.</p>
<p>Some studies of virtual agents have proposed methods to automatically generate the neck movements of agents matched with their speech. Le et al. (<xref ref-type="bibr" rid="B25">2012</xref>) modeled the relationship between human neck movements (roll, pitch, and yaw) and the voice information (power and pitch) using a Gaussian mixture model (GMM) for real-time neck-motion generation. Other models using a Hidden Markov Model (HMM) have also been proposed: Sargin et al. (<xref ref-type="bibr" rid="B40">2008</xref>), Busso et al. (<xref ref-type="bibr" rid="B4">2005</xref>), and Foster and Oberlander (<xref ref-type="bibr" rid="B9">2007</xref>). Watanabe et al. (<xref ref-type="bibr" rid="B44">2004</xref>) proposed a different kind of model that estimates the speaker&#x02019;s nodding timing based on historical on&#x02013;off patterns in the voice. These methods can successfully generate humanlike motions, but they are restricted to certain situations in which a model&#x02019;s learning data can be collected. Generally, human speech motions depend on the relationship between the speaker and listener and the speech context (for example, in happy situations, the magnitude and speed of human speech motions might be larger than usual). To alter the agent motion to fit the given situation, data must be collected for any situation, which is impractical, unfortunately.</p>
<p>The situation-dependency issue of can be solved by modulating the android&#x02019;s motions. Jia et al. (<xref ref-type="bibr" rid="B19">2014</xref>) synthesized the head and facial gestures of a talking agent with emotional expressions. In their method, the nodding motion involved in the utterance of stressed syllables is modulated according to the emotion represented in a PAD model (an extension of Russell&#x02019;s emotion model, Russell, <xref ref-type="bibr" rid="B36">1980</xref>). However, their study focused only on positive emotions. Moreover, only the amplitude of motion was modulated, although the other motion features (e.g., velocity) should be varied according to the emotion. Masuda and Kato (<xref ref-type="bibr" rid="B28">2010</xref>) developed a method that changed the gestural motions of a robot by associating Laban theory (Laban, <xref ref-type="bibr" rid="B23">1988</xref>) with Russell&#x02019;s emotion model. Unfortunately, since their method produced exaggerated gestures that cannot be expressed by humans, it is unsuitable for very humanlike androids. Other researchers have studied the relationship between motion characteristics and emotion in walking (Gross et al., <xref ref-type="bibr" rid="B14">2012</xref>) and kicking (Amaya et al., <xref ref-type="bibr" rid="B1">1996</xref>) motions. Such studies commonly conclude that the magnitude and speech of motion vary depending on the emotion.</p>
<p>Physiological studies are also helpful to understand motions of situation dependency, which is, the relationship between internal states and motion characteristics. Many studies reported that psychological pressure influences the human&#x02019;s musculoskeletal system and produces jerky motions. For example, anxiety increases muscle stiffness (Fridlund et al., <xref ref-type="bibr" rid="B11">1986</xref>). Nakano and Hoshino (<xref ref-type="bibr" rid="B31">2007</xref>) showed that a human&#x02019;s waist moves slowly in a relaxed state but jerkily in a nervous state. This also suggests that the musculoskeletal system is influenced by mental pressure. Consequently, speech motion that is dependent on internal states might be generated by simulating the dynamics of the musculoskeletal system.</p>
<p>In this study, we develop an analytical model based on the physical constraints (kinematic relations) of a human&#x02019;s speech production. We focus on the motions of the body&#x02019;s trunk (neck and waist) on the sagittal plane, since lateral motions usually depend on the social situation, e.g., neck yawing due to gaze aversion that depends on the conversation context (Andrist et al., <xref ref-type="bibr" rid="B2">2014</xref>) as well as on the speaker&#x02019;s personality (Larsen and Shackelford, <xref ref-type="bibr" rid="B24">1996</xref>). The motion generation is formulated by a spring-damper model that simulates musculoskeletal dynamics. The generated motion can be easily modulated by tuning the parameters of muscle stiffness to express internal states. Motions are triggered by utterances in this model. By observing a person&#x02019;s utterances, this study found the physical constraints between speech production and the movements of the human head and mouth (opening/closing) to implement this trigger.</p>
</sec>
<sec id="S3">
<label>3</label> <title>Relations Between Prosodic Features and Head Motions</title>
<sec id="S3-1">
<label>3.1</label> <title>Basic Idea of Generation Model</title>
<p>We generate a speech motion synchronized with an utterance to enforce multimodal synchronization to avoid unnatural motions caused by hardware limitations. Such motions should also be generated by an analytic model to easily modulate them based on the android&#x02019;s emotion. This section first finds the relation between prosodic features and head motions, which is not influenced by social situation. Human head movements are synchronized with prosodic features when a person speaks; in particular, voice power and pitch features have high correlation with head movements (Bolinger, <xref ref-type="bibr" rid="B3">1985</xref>). However, for our current Japanese targets, the correlation between prosodic features and head movements is not very high (Yehia et al., <xref ref-type="bibr" rid="B45">2002</xref>). Anatomical research has also reported that the human head moves when the mouth is opened (Eriksson et al., <xref ref-type="bibr" rid="B7">1998</xref>). Since mouth openness strongly depends on vowels in Japanese, we expect to find a relationship between head movements and prosodic information that involves Japanese vowels. Here, we reveal what features of human movements are related to the three voice features (power, pitch, and vowels) when a person is speaking without interacting with anyone. Those relations underlie the model to generate context-free head motions in speaking (Section <xref ref-type="sec" rid="S4">4</xref>). These motions can be modulated to fit a social situation by turning the model parameters (Section <xref ref-type="sec" rid="S6">6</xref>).</p>
</sec>
<sec id="S3-2">
<label>3.2</label> <title>Experimental Setup</title>
<p>To examine the relation between head motion and voice features, we recorded the speech of human subjects while controlling their voice features. Before the experiment, a preliminary test checked whether people could pronounce the required vowels (a, i, u, e, o) with the required pitch and power. We found that people did not move their heads at all when the power was low. Based on this result, we identified a relationship where the more loudly people speak, the larger the head movement becomes. We asked the subjects to produce loud voices in the experiments. The preliminary test also suggested that the vowel changes affect head motions when people loudly pronounced the vowels, but they had some difficulty doing so while keeping the required pitch. We prepared two tests: pitch and vowel. The former measured the head motion when the participants pronounced a set of any syllables (here we chose five vowels) with the required pitch (3&#x02009;s for each syllable) to investigate the relation between the head motion and pitch. The latter measured the head motion when the participants loudly pronounced the required vowel for 3&#x02009;s with any pitch to investigate the relation between the head motion and vowels.</p>
<p>The head movements were measured by an inertia measurement unit (IMU) (InterSense InertiaCube4) attached to the top of the head at a measurement sampling rate of 100&#x02009;Hz. The participants were given the following instructions: &#x0201C;Clearly pronounce a high-voice-pitch &#x02018;a&#x0201D;&#x02019; or &#x0201C;clearly pronounce a low-voice-pitch &#x02018;i&#x02019;.&#x0201D; The frequency of the pitch was not fixed since they had difficulty producing the required frequency. They were also required to reset their posture forward before each pronunciation. In the pitch test, they loudly pronounced five vowels (each for 3&#x02009;s) with high, middle, and low pitch. In the vowel test, they loudly pronounced each vowel for 3&#x02009;s with whatever pitch they could easily pronounce.</p>
<p>The participants conducted each test twice. The first trial acclimated them to speaking in the experimental room and checked that the recording system.</p>
</sec>
<sec id="S3-3">
<label>3.3</label> <title>Results</title>
<p>Eleven Japanese speakers (six males, five females, average age: 22.0, SD: 0.54) participated in the experiment. They were recruited from a job-offering site for university students with different backgrounds. We removed one male from our analysis since he pronounced all the words with the same tone although we instructed him to pronounce them in high, middle, and low tones.</p>
<p>In the pitch test, the participants tended to maintain almost constant head posture while speaking. We calculated the average head-elevation angle during their utterances (Figure <xref ref-type="fig" rid="F1">1</xref>). They tended to look up when pronouncing with a high pitch and look forward or down when pronouncing with a low pitch. A one-way analysis of variance (ANOVA) with one within-subject factor revealed a significant main effect (<italic>F</italic>(2, 18)&#x02009;&#x0003D;&#x02009;12.84, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01). Hereafter, <italic>F</italic> and <italic>p</italic> denote <italic>F</italic> statistics and significance level. Following this result, we conducted multiple comparisons by the Bonferroni method and found significant differences in the head angle as the angles increased with a greater pitch level. This means that they respectively tended to move their heads up and down when they spoke with high and low pitch.</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>Head postures related to voice pitch. The angle is zero when the participant is facing forward.</p></caption>
<graphic xlink:href="frobt-04-00049-g001.tif"/>
</fig>
<p>Figure <xref ref-type="fig" rid="F2">2</xref> shows the results of the vowel test results. The average head-elevation angle during speech was measured as well as in the pitch test. The vertical axis shows the head angle displacements and how much the head angle changed while pronouncing the vowels compared with the before pronouncing them. We categorized the vowels into two groups: wide group (a, e, o) and narrow group (i, u). We found a significant difference between the two groups (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05).<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> The participants greatly moved their heads when they opened their mouths. The observed motions seem to be made purely for pronouncing utterances since there was no social context; they were pronouncing meaningless syllables without any interaction. Therefore, the revealed relations might not depend on the social situations. However, the speech motions did change depending on the social situation. The generated motions of the android based on the relations must match the social situation.</p>
<fig position="float" id="F2">
<label>Figure 2</label>
<caption><p>Head displacements related to vowels.</p></caption>
<graphic xlink:href="frobt-04-00049-g002.tif"/>
</fig>
<p>Most participants in the vowel test pronounced with a middle pitch. The angle displacements in Figure <xref ref-type="fig" rid="F2">2</xref> are less than 5&#x000B0; but the head angles of middle pitch in Figure <xref ref-type="fig" rid="F1">1</xref> exceed 5&#x000B0;, because they slightly raised their head postures and moved their heads when they opened their mouths.</p>
</sec>
</sec>
<sec id="S4">
<label>4</label> <title>Speech-Driven Trunk Motion Generation</title>
<sec id="S4-4">
<label>4.1</label> <title>Generating Smooth Motion from Prosodic Features</title>
<p>In this section, we develop a model to generate a context-free trunk motion. First, we made a head motion-generation model based on the above results and extended it to both trunk and neck motions. We found a strong relation between the head angle and the prosodic features in Section <xref ref-type="sec" rid="S3">3</xref>, but a simple mapping is not appropriate for generating a smooth motion, which is essential for human-likeness (Shimada and Ishiguro, <xref ref-type="bibr" rid="B41">2008</xref>; Piwek et al., <xref ref-type="bibr" rid="B33">2014</xref>). This is because prosodic features are sometimes intermittent and change rapidly. The model needs to generate smooth motions based on the rapidly changing discrete sound input. Concerning humanlike motions, people perceive naturalness in the second-order dynamic motions of a virtual agent (Nakazawa et al., <xref ref-type="bibr" rid="B32">2009</xref>). Therefore, we use a spring-damper dynamical model (Figure <xref ref-type="fig" rid="F3">3</xref>; equation (<xref ref-type="disp-formula" rid="E1">1</xref>)) to generate a smooth motion from the non-smooth prosodic features
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mi>j</mml:mi><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo class="MathClass-op">&#x002D9;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo class="MathClass-op">&#x002D9;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mi>k</mml:mi><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mi mathvariant="italic">dir</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<fig position="float" id="F3">
<label>Figure 3</label>
<caption><p>Spring-damper model for speech motion generation.</p></caption>
<graphic xlink:href="frobt-04-00049-g003.tif"/>
</fig>
<p>Head angle <italic>&#x003B8;<sub>base</sub></italic> is driven by the external force <italic>&#x003C4;</italic>(<italic>t</italic>)<italic>dir</italic>(<italic>t</italic>), where <italic>t</italic> denotes discrete time. The absolute values of force <italic>&#x003C4;</italic>(<italic>t</italic>) and force direction <italic>dir</italic>(<italic>t</italic>) are separately described below for convenience of explanation. The prosodic information is associated with the force <italic>&#x003C4;</italic>(<italic>t</italic>)<italic>dir</italic>(<italic>t</italic>) based on the results in Section <xref ref-type="sec" rid="S3">3</xref>. For example, if the model is given, a large upward force is made by widely opening the mouth, and a large upward neck movement is generated, following second-order dynamics. Parameters <italic>j</italic>, <italic>k</italic>, and <italic>d</italic> in the equation are equivalent to head weight, muscle stiffness, and muscle viscosity, respectively, based on using the spring-damper model of human muscle dynamics (Linder, <xref ref-type="bibr" rid="B27">2000</xref>; Liang and Chiang, <xref ref-type="bibr" rid="B26">2006</xref>). The meaning of each parameter is easy to grasp, and we can intuitively modulate the motion by changing the parameters. Furthermore, muscle stiffness is related to a person&#x02019;s internal states such as emotion and tension. We assume that we can modulate the motion to express the speaker&#x02019;s internal states.</p>
</sec>
<sec id="S4-5">
<label>4.2</label> <title>Head-Motion Generation Based on Spring-Damper Model</title>
<p>This section defines the external force in equation (<xref ref-type="disp-formula" rid="E1">1</xref>) using prosodic information. The results in Section <xref ref-type="sec" rid="S3-3">3.3</xref> show that pronouncing vowels with a wide opened mouth produces a large head movement. Hereafter, to express the vowel information with a continuous value, we use the value of mouth-openness value. As the mouth is opening, the force <italic>m</italic>(<italic>t</italic>), which is proportional to the degree of mouth openness <italic>&#x003B5;</italic>(<italic>t</italic>), is given to the model. <italic>&#x003B5;</italic>(<italic>t</italic>) can be estimated from the voice, as described in Section <xref ref-type="sec" rid="S4-6">4.3</xref>. On the contrary, no force is given to the model while the mouth is closing, and thus the head smoothly returns to its original position by using the restoring force of the spring, as shown in equation (<xref ref-type="disp-formula" rid="E2">2</xref>). In this study, the sampling time is 10&#x02009;ms. As described in Section <xref ref-type="sec" rid="S4-6">4.3</xref>, speaking with large power produces a large head movement. While the voice power is increasing or keeping the same level, force <italic>p</italic>(<italic>t</italic>), which is proportional to it, is given to the model. As with the vowel, no force is given to the model with the power shown in equation (<xref ref-type="disp-formula" rid="E3">3</xref>). In total, the force is expressed as equation (<xref ref-type="disp-formula" rid="E4">4</xref>), where <italic>v</italic> and <italic>l</italic> are constant values (no physical meaning) to balance different scales of values (voice power and degree of mouth openness). Figure <xref ref-type="fig" rid="F4">4</xref> shows an example of the input force generated from the voice:
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mi>m</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003E;</mml:mo><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>0</mml:mn><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">otherwise</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">,</mml:mo></mml:math></disp-formula>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mi>p</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003E;</mml:mo><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>0</mml:mn><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">otherwise</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">,</mml:mo></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi mathvariant="italic">vp</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mi mathvariant="italic">lm</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<fig position="float" id="F4">
<label>Figure 4</label>
<caption><p>Input force generated from the voice data.</p></caption>
<graphic xlink:href="frobt-04-00049-g004.tif"/>
</fig>
<p>From the voice pitch results in Section <xref ref-type="sec" rid="S3-3">3.3</xref>, we associate the pitch with the force&#x02019;s direction, as shown in equation (<xref ref-type="disp-formula" rid="E5">5</xref>), where <italic>p<sub>t</sub></italic> denotes the state of pitch (<italic>p<sub>t</sub></italic>&#x02009;&#x02208;&#x02009;{<italic>High</italic>, <italic>Middle</italic>, <italic>Low</italic>}). Since the participants in the pitch test tended to change their postures more than usual, we focused on the direction of movement. Figure <xref ref-type="fig" rid="F1">1</xref> shows that the average head angle is a low pitch with a positive value, but we defined <italic>dir</italic>(<italic>t</italic>), which became minus when the pitch was low, to enforce the relationship between the pitch and head direction.</p>
<p>Equation <xref ref-type="disp-formula" rid="E5">5</xref> means that an upward force is given to the head when the pitch is high and a head-lowering movement when the pitch is low. When the pitch is in the mid-range, no force is given to the model and only a restoring movement is generated. The next section describes how we classify the pitch into three groups
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mi mathvariant="italic">dir</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>1</mml:mn><mml:mspace width="1em" class="quad"/><mml:mspace width="2.56804pt" class="tmspace"/><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">Head</mml:mi><mml:mtext>&#x02009;</mml:mtext><mml:mi mathvariant="italic">up</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi mathvariant="italic">High</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mspace width="0.5em" class="quad"/><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">Head</mml:mi><mml:mtext>&#x02009;</mml:mtext><mml:mi mathvariant="italic">down</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi mathvariant="italic">Low</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>0</mml:mn><mml:mspace width="1em" class="quad"/><mml:mspace width="2.56804pt" class="tmspace"/><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">Restoring</mml:mi><mml:mtext>&#x02009;</mml:mtext><mml:mi mathvariant="italic">movement</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi mathvariant="italic">Middle</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="S4-6">
<label>4.3</label> <title>Prosodic Information Extraction</title>
<p>Fundamental frequency F0 was extracted in a 32-ms frame size every 10&#x02009;ms. We used a conventional method that extracts the autocorrelation peaks of residual signals calculated by a linear predictive coding (LPC) inverse filter. Then value <inline-formula><mml:math id="M"><mml:mrow><mml:mover accent='true'><mml:mrow><mml:mi>F</mml:mi><mml:mn>0</mml:mn></mml:mrow><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is calculated by averaging <italic>F</italic>0 over the last 100&#x02009;ms. The pitch is classified as follows:
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi mathvariant="italic">High</mml:mi><mml:mspace width="2.56804pt" class="tmspace"/><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>F</mml:mi><mml:mn>0</mml:mn></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover><mml:mo class="MathClass-rel">&#x0003E;</mml:mo><mml:mi>F</mml:mi><mml:msub><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="italic">high</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi mathvariant="italic">Low</mml:mi><mml:mspace width="2.56804pt" class="tmspace"/><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>F</mml:mi><mml:mn>0</mml:mn></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover><mml:mo class="MathClass-rel">&#x0003C;</mml:mo><mml:mi>F</mml:mi><mml:msub><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="italic">low</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi mathvariant="italic">Middle</mml:mi><mml:mspace width="2.56804pt" class="tmspace"/><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<p>The frequency range was empirically determined since the voice&#x02019;s fundamental frequency depends on speaker&#x02019;s gender and age.</p>
<p>We estimated the mouth openness using a lip-motion-generating system (Ishi et al., <xref ref-type="bibr" rid="B18">2012</xref>) based the voice&#x02019;s formant information. These methods can extract the prosodic features in real time and simultaneously create motion generation when the android produces a voice.</p>
</sec>
<sec id="S4-7">
<label>4.4</label> <title>Trunk Motion Generation</title>
<p>Zafar et al. (<xref ref-type="bibr" rid="B48">2002</xref>) revealed that an up-and-down movement of the head also produces a front-and-back movement of the body. This suggests that a cooperative movement between the neck and waist might produce a more humanlike impression. Zafar et al. (<xref ref-type="bibr" rid="B47">2000</xref>) also reported that lip movement is followed by a neck movement, although with a slight delay. From these findings, we defined trunk motions (including head motion) as phase-shifted motions of <italic>&#x003B8;<sub>base</sub></italic>, as shown in equation (<xref ref-type="disp-formula" rid="E7">7</xref>). <italic>act<sub>i</sub></italic> indicates the <italic>i</italic>-th actuator of the robot (Figure <xref ref-type="fig" rid="F5">5</xref>), and <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">ac</mml:mi><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">ac</mml:mi><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> are parameters that determine the coordination between the joints:
<disp-formula id="E7"><label>(7)</label><mml:math id="M9"><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">ac</mml:mi><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">ac</mml:mi><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">ac</mml:mi><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<fig position="float" id="F5">
<label>Figure 5</label>
<caption><p>Multi-joint control.</p></caption>
<graphic xlink:href="frobt-04-00049-g005.tif"/>
</fig>
</sec>
<sec id="S4-8">
<label>4.5</label> <title>Improvement of the Model</title>
<p>To verify the model, we recorded the voice and motion of the speakers and compared the motions generated by our model and the originals. We tested two female speakers: Speakers M and H. We tested female voices because we used a female type of android is used in the latter experiment. The speakers read a self-introduction whose content was identical (approximately 50&#x02009;s), and their trunk movements (neck and waist angles) were measured by the IMU attached to their head and torso.</p>
<p>To test the proposed model, we controlled the android&#x02019;s neck joint by the model. Parameters <italic>j</italic> and <italic>d</italic> in the model were set to 0.0676 and 0.52. <italic>k</italic> was set to 0.195 (<italic>&#x003B8;<sub>base</sub></italic>&#x02009;&#x02265;&#x02009;0) and 0.065 (<italic>&#x003B8;<sub>base</sub></italic>&#x02009;&#x0003C;&#x02009;0). We assumed that the head can move easily when it looks downward since the head posture is not defying gravity. To implement these characteristics we set a smaller <italic>k</italic> value for <italic>&#x003B8;<sub>base</sub></italic>&#x02009;&#x02265;&#x02009;0. Balance parameters <italic>v</italic> and <italic>l</italic> in equation (<xref ref-type="disp-formula" rid="E4">4</xref>) were set to 0.001 and 0.0005, where the voice power ranged approximately from 10 to 20&#x02009;dB and the mouth openness ranged approximately from 0 to 200 (non-dimensional value). Thresholds <italic>F</italic>0<italic><sub>high</sub></italic> and <italic>F</italic>0<italic><sub>low</sub></italic> in equation (<xref ref-type="disp-formula" rid="E6">6</xref>) were, respectively, set to 256 and 215&#x02009;Hz. These parameters described in this session were set by trial and error. The experimenter selected the parameters in a preliminary trial for a natural neck motion that matched the speech. Mouth openness (<italic>&#x003B5;</italic>(<italic>t</italic>)) was calculated from the voice by Ishi et al.&#x02019;s method (Ishi, <xref ref-type="bibr" rid="B16">2005</xref>). Generated angle <italic>&#x003B8;<sub>base</sub></italic> was directly used to control the neck joint (i.e., <italic>&#x003B8;<sub>neck</sub></italic>(<italic>t</italic>)&#x02009;&#x0003D;&#x02009;<italic>&#x003B8;<sub>base</sub></italic>(<italic>t</italic>)), but no other trunk joints were controlled in this experiment.</p>
<p>Figure <xref ref-type="fig" rid="F6">6</xref> shows the head angles generated from two speakers&#x02019; voices. Here, original indicates the original motion and non-restricted indicates the motion (<italic>&#x003B8;<sub>neck</sub></italic>(<italic>t</italic>)) generated by the proposed model. In Figure <xref ref-type="fig" rid="F6">6</xref>B, there is a clear difference between the trajectories of the original and proposed method, where the original angle was positive in most cases (Speaker M looked up during the speech), but the angle by the proposed method was negative in most cases. This is because Speaker M reset her posture again just before speaking and consequently looked a little bit up during the speech, although she was required to face forward before speaking and to keep her gaze during the speech. In this research, since we focus on the correlation between the changes of the prosodic information and the neck pitch angle, our system cannot reproduce the direction in which the original speaker looked. This study did not consider the average neck angle in its later analysis. The original motions tended to consist of large movements followed by small vibrating motions. On the other hand, the proposed model seems to generate a simple cyclic pattern, which produces a robot-like impression. We assumed that the speakers would prominently make a large head movement when they make a large voice and widely open their mouths. In other words, the amount of motion is not simply proportional to the magnitude of the prosodic features; large changes in the prosodic features produce larger movements. To clearly express this relationship, we introduce thresholds to restrict the motions shown in equations (<xref ref-type="disp-formula" rid="E8">8</xref>) and (<xref ref-type="disp-formula" rid="E9">9</xref>) by which a large driving force is only given to the model when the voice power and/or mouth openness are largely increasing. In this model, larger movements are generated when the prosodic features exceed the thresholds (<italic>p</italic>_<italic>th</italic> and <italic>m</italic>_<italic>th</italic>), and otherwise the restoring force generates small vibrating movements. The plot of restricted in Figure <xref ref-type="fig" rid="F6">6</xref> shows the movements by this model, where the thresholds <italic>p</italic>_<italic>th</italic> and <italic>m</italic>_<italic>th</italic> were empirically set to 1 and 10 so that large head movements prominently appear synchronized with a loud voice. With this improvement, motions similar to the original ones can be generated:
<disp-formula id="E8"><label>(8)</label><mml:math id="M10"><mml:mi>p</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mi>&#x003C1;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003E;</mml:mo><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>p</mml:mi><mml:mtext>_</mml:mtext><mml:mi mathvariant="italic">th</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>0</mml:mn><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">otherwise</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">,</mml:mo></mml:math></disp-formula>
<disp-formula id="E9"><label>(9)</label><mml:math id="M11"><mml:mi>m</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfenced separators="" open="{" close=""><mml:mrow><mml:mtable equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mi>&#x003F5;</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo class="MathClass-bin">&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003E;</mml:mo><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>m</mml:mi><mml:mtext>_</mml:mtext><mml:mi mathvariant="italic">th</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd class="array" columnalign="left"><mml:mn>0</mml:mn><mml:mspace width="1em" class="quad"/></mml:mtd><mml:mtd class="array" columnalign="left"><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi mathvariant="italic">otherwise</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mfenced><mml:mo class="MathClass-punc">.</mml:mo></mml:math></disp-formula></p>
<fig position="float" id="F6">
<label>Figure 6</label>
<caption><p>Human head movements during speaking (top plot shows speech waveform and bottom plot shows generated head motion). <bold>(A)</bold> Speaker H. <bold>(B)</bold> Speaker M.</p></caption>
<graphic xlink:href="frobt-04-00049-g006.tif"/>
</fig>
</sec>
</sec>
<sec id="S5">
<label>5</label> <title>Evaluation of Proposed System</title>
<sec id="S5-9">
<label>5.1</label> <title>Experimental Setup</title>
<p>First, we evaluated the impressions of the android&#x02019;s movements generated by the proposed method to show that it could generate natural and humanlike motions. The next section shows that the generated movements can be modulated for expressing internal states. The existing systems described in Section <xref ref-type="sec" rid="S2">2</xref> basically reproduce the original human movements. Such methods do not enforce the synchronization between speech and motion. To show how our method improved the impressions of the androids, we compared it with a method in which the original speaker motions were reproduced in the android. This method copied the speaker&#x02019;s head and waist angles measured by IMU to the corresponding joint angles of the android (copy condition). Furthermore, to verify whether an unnatural impression is produced when the motion is not synchronized with the voice, we prepared another android motion by copying the original with a 1-s delay (non-synchronized condition) and two conditions for our model: with improvement of equations (<xref ref-type="disp-formula" rid="E8">8</xref>) and (<xref ref-type="disp-formula" rid="E9">9</xref>) and another without it. Then, we compared four conditions (proposed method with improvement, proposed method without improvement, copy, and non-synchronized, hereafter referred to as Proposed-I, Proposed-NI, Copy, and NoSync). We used a female type of android, named ERICA (Figure <xref ref-type="fig" rid="F7">7</xref>).</p>
<fig position="float" id="F7">
<label>Figure 7</label>
<caption><p>Android ERICA.</p></caption>
<graphic xlink:href="frobt-04-00049-g007.tif"/>
</fig>
<p>The participants evaluated the android&#x02019;s movements without any interactions (videotaped movements were used). We used the recorded human motion data and voice data mentioned in Section <xref ref-type="sec" rid="S5">5</xref>. To verify that the method generated motions from any speaker&#x02019;s voice, this experiment used the data from both speakers. A comparison of Figures <xref ref-type="fig" rid="F6">6</xref>A,B revealed that the movement of Speaker M was larger and more frequent than that of Speaker H. Consequently, we examined the movement of different types of speakers.</p>
<p>The parameters in the model were the same as those used in Section <xref ref-type="sec" rid="S5">5</xref>. The head and waist motions in the Proposed-I and Proposed-NI conditions were generated from the speakers&#x02019; voice data and used to control ERICA&#x02019;s corresponding joints (i.e., we used two joints). The neck joint was controlled as <italic>&#x003B8;<sub>neck</sub></italic>(<italic>t</italic>)&#x02009;&#x0003D;&#x02009;<italic>&#x003B8;<sub>base</sub></italic>(<italic>t</italic>). No phase-shift between the neck and waist motions was assigned, that is, <italic>&#x003B8;<sub>waist</sub></italic>(<italic>t</italic>)&#x02009;&#x0003D;&#x02009;<italic>&#x003B1;<sub>waist</sub>&#x003B8;<sub>base</sub></italic>(<italic>t</italic>), where <italic>&#x003B1;<sub>waist</sub></italic>&#x02009;&#x0003D;&#x02009;0.1. The android&#x02019;s lip movements were automatically generated from the voice by Ishi et al.&#x02019;s method (Ishi, <xref ref-type="bibr" rid="B16">2005</xref>) in all the conditions. Furthermore, we added blinking at random intervals normally distributed with a mean of 4&#x02009;s and a SD of 0.5&#x02009;s. Figure <xref ref-type="fig" rid="F8">8</xref> shows the kinematic structure of the android. We used joint 1 for blinking, joints 15 and 18 for trunk movements, and joint 13 for lip movements. The waist joint&#x02019;s control input was processed by a low-pass filter that moved over an average of the past 200&#x02009;ms, since the time constant in the waist control is larger than in the neck control. The android&#x02019;s voice was provided by playing the original voice. Its movements were videotaped by covering the front waist-up image of its body (Figure <xref ref-type="fig" rid="F9">9</xref>), and the videos were used as the experimental stimuli.</p>
<fig position="float" id="F8">
<label>Figure 8</label>
<caption><p>Kinematic structure of proposed android.</p></caption>
<graphic xlink:href="frobt-04-00049-g008.tif"/>
</fig>
<fig position="float" id="F9">
<label>Figure 9</label>
<caption><p>Video stimulus.</p></caption>
<graphic xlink:href="frobt-04-00049-g009.tif"/>
</fig>
<p>Each participant evaluated all eight conditions (four types of motion&#x02009;&#x000D7;&#x02009;two types of speaker). The order of speakers was fixed (Speaker H was first) because we assumed that there would be no order effect related to the speakers. We counterbalanced the order of the four motion conditions. The participants could watch the video stimuli as many times as they liked.</p>
</sec>
<sec id="S5-10">
<label>5.2</label> <title>Evaluation Measurements</title>
<p>The participants rated the naturalness and impression of each motion in the video on a 7-point Likert scale. For naturalness, the questionnaire asked whether &#x0201C;the neck and waist movements are natural&#x0201D; (naturalness). Piwek et al. (<xref ref-type="bibr" rid="B33">2014</xref>) revealed that natural movements improve the sense of intimacy toward agents. Then, the questionnaire asked participants to rate the statement, &#x0201C;I want to interact with the android,&#x0201D; (will) as the willingness to interact with it.</p>
</sec>
<sec id="S5-11">
<label>5.3</label> <title>Results</title>
<p>Fifteen Japanese speakers (12 males, 3 females, average age: 21.5, SD: 1.6) participated in the experiment. They were recruited in the same manner as described earlier. We conducted a two-way ANOVA with two within-subject factors (motion and speaker) and found no significant effect for the scores of will, but there was a significant trend of the motion factor for naturalness scores (<italic>F</italic>(3, 42)&#x02009;&#x0003D;&#x02009;2.33, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.1). In further examining the scores for Proposed-I and Proposed-NI scores, we found that some gave higher scores for Proposed-I, while others assigned opposite scores. This means that the effect of motion restriction in equations (<xref ref-type="disp-formula" rid="E8">8</xref>) and (<xref ref-type="disp-formula" rid="E9">9</xref>) was subjective. To comprehensively verify the effect of the proposed method, we combined the scores of Proposed-I and Proposed-NI scores into the scores of the &#x0201C;Proposed&#x0201D; condition by extracting the higher score between Proposed-I and Proposed-NI within-participants for each evaluation measurement.</p>
<p>A two-way ANOVA with two within-subject factors was conducted. Regarding the will score, there was a significant effect of motion factor (<italic>F</italic>(2, 28)&#x02009;&#x0003D;&#x02009;4.90, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05). Figure <xref ref-type="fig" rid="F10">10</xref> compares the scores between the motion factors. Multiple comparisons by the Holm method revealed significant differences between conditions as Proposed&#x02009;&#x0003E;&#x02009;Copy (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05) (meaning the Proposed condition score is larger than that in the Copy condition, and the same hereafter) and Proposed&#x02009;&#x0003E;&#x02009;NoSync (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05). Regarding Naturalness, there was a significant effect of motion factor (<italic>F</italic>(2, 28)&#x02009;&#x0003D;&#x02009;6.51, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01) as well as significant differences between conditions as Proposed&#x02009;&#x0003E;&#x02009;Copy (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05) and Proposed&#x02009;&#x0003E;&#x02009;NoSync (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01) (Figure <xref ref-type="fig" rid="F11">11</xref>).</p>
<fig position="float" id="F10">
<label>Figure 10</label>
<caption><p>Willingness to talk with android.</p></caption>
<graphic xlink:href="frobt-04-00049-g010.tif"/>
</fig>
<fig position="float" id="F11">
<label>Figure 11</label>
<caption><p>Naturalness of neck and waist motions.</p></caption>
<graphic xlink:href="frobt-04-00049-g011.tif"/>
</fig>
<p>We also found an interaction effect between motion and speaker factors for naturalness (<italic>F</italic>(2, 28)&#x02009;&#x0003D;&#x02009;5.89, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01). Figure <xref ref-type="fig" rid="F12">12</xref> shows the average naturalness scores of six conditions. In the NoSync condition, the Speaker H score was significantly higher than that for Speaker M (<italic>F</italic>(1, 14)&#x02009;&#x0003D;&#x02009;12.64, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01). In the Speaker M condition, there was a significant simple main effect for the motion factor (<italic>F</italic>(2, 28)&#x02009;&#x0003D;&#x02009;14.40, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01). We conducted multiple comparisons by the Holm method and found significant differences among the conditions: Proposed&#x02009;&#x0003E;&#x02009;Copy (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05), Proposed&#x02009;&#x0003E;&#x02009;NoSync (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001), and Copy&#x02009;&#x0003E;&#x02009;NoSync (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05). The experimental results showed that the android motions generated by our model appeared more natural and motivated participants to talk more compared with copying of the human motions. This also means that our method outperformed the existing method by ideally reproducing the original human movements. We tested two speakers and found this effect for both of them suggesting our model generated natural speech motions for various speech patterns.</p>
<fig position="float" id="F12">
<label>Figure 12</label>
<caption><p>Naturalness of neck and waist motions for each speaker.</p></caption>
<graphic xlink:href="frobt-04-00049-g012.tif"/>
</fig>
</sec>
<sec id="S5-12">
<label>5.4</label> <title>Discussion</title>
<p>Why was the speech motion of our model evaluated as more natural than the copying of the human motions? We usually do not feel unnaturalness when we see a person who barely moves her head while speaking. We usually rationalize such a situation, for example, her body is tense due to some strain. However, when we see an android whose head barely moves, we might blame the system. We assumed that people would differently attribute the reason for the movement by a human or an android. If people clearly sense the correlation between the motion and prosodic features, which are typical in human speech, we can expect to avoid the attribution of a reason for non-humanlike motions by the android. This might explain why the proposed method&#x02019;s motions were better than the copied motions. Humans have many joints and muscles in their bodies and always make complex and subtle body motions, which cannot be perfectly reproduced in an android due to the limitations of its degrees of freedom and the characteristics of its actuators. That is, incomplete copying of human motions produces a more negative impression than uncopied motions. The proposed method, on the other hand, might be able to compensate for such incompleteness by expressively showing the motions that are strongly related to certain voice features. A similar result was obtained in a study of speech-driven lip motion generation by an android (Ishi et al., <xref ref-type="bibr" rid="B17">2011</xref>). Perhaps the synchronization between voice and motion captures the essence of human-likeness in speech motions. In the future, with our model, we will investigate which kinds of motions essentially contribute to human-likeness by modulating the model parameters. In this sense, our proposed analytical model is helpful for studying the mechanism of human&#x02013;robot interaction.</p>
<p>We found no significant differences between the Copy and NoSync conditions, although we expected the latter to be worse (e.g., since asynchrony between voice and lip movement on decreases the human-likeness of human and virtual characters (Tinwell et al., <xref ref-type="bibr" rid="B42">2015</xref>)). However, this depended on the speaker. As shown in Figure <xref ref-type="fig" rid="F12">12</xref>, NoSync-Speaker M had a significantly lower score in naturalness score. We inferred a strong correlation between speech and motion for Speaker M but not for Speaker H. Since the participants attributed some meaning to the delayed motions in the NoSync-Speaker H condition, and thus they did not give bad scores. This result does not negate the necessity for a temporal synchronization between speech and motion, but we need further study to determine for which kinds of speech patterns this synchronization is necessary. Kirchhof (<xref ref-type="bibr" rid="B20">2014</xref>) suggested that the necessity of temporal synchrony between speech and gestures becomes smaller by loosening their semantic synchrony. A semantical correspondence between speech and motion might govern the participant impressions.</p>
<p>Comparing Figure <xref ref-type="fig" rid="F6">6</xref>B with Figure <xref ref-type="fig" rid="F6">6</xref>A, Speaker M&#x02019;s motions tend to have a larger magnitude. This means that she has a higher speech-motion correlation and moves her body more. Nevertheless, Speaker M&#x02019;s naturalness score is lower than that of Speaker H in the copy condition although the difference is not significant. Perhaps motion copying failed to take into account the coordination throughout the entire body. Speaker M might greatly move not only her neck and waist but also other body parts to balance her entire body during speech; however, the system does not generate those motions. The participants might feel some unnaturalness in the android&#x02019;s entire body movements with the loss of coordination. Since the proposed method reproduces well-coordinated motions with only the neck and waist, the participants felt that they were natural, even though they were different from the original speaker&#x02019;s motion. In their study of humanoid motion generation, Gielniak et al. (<xref ref-type="bibr" rid="B13">2013</xref>) showed that the coordinated motions on multiple joints produced humanlike motions. Their method emulates the coordinated effects of human joints that are connected by muscles, and their experimental result showed that this coordination provides the impression of human-likeness in the robot motions. These results suggest that motion coordination under a physical body&#x02019;s constraints is critical for giving a natural impression of motions.</p>
<p>The preference for Proposed-I or Proposed-NI motion depended on the participants. Some believed that Proposed-NI is better since it moved more, but others felt the opposite. People tend to prefer a person who mimics them: the chameleon effect (Chartrand and Bargh, <xref ref-type="bibr" rid="B5">1999</xref>). Accordingly, we infer that the participants preferred motions that resembled their own. We could choose either Proposed-I or Proposed-NI for the android motion based on the personality of the conversation partner by prior personality questionnaires or estimations based on a multimodal sensor system. Further study is needed to investigate how to generate an android&#x02019;s speech motion to suit the personality of the conversation partner.</p>
</sec>
</sec>
<sec id="S6">
<label>6</label> <title>Evaluation of Capability of Emotional Expression</title>
<sec id="S6-13">
<label>6.1</label> <title>Purpose of Experiment</title>
<p>This section shows how the generated motions can be modulated to express the internal state of an android. The parameters in the proposed model correspond to muscle stiffness, which depends on the speaker&#x02019;s mental state, e.g., stress and emotion (Sainsbury and Gibson, <xref ref-type="bibr" rid="B37">1954</xref>; Fridlund et al., <xref ref-type="bibr" rid="B11">1986</xref>). In other words, perhaps we can modulate the speech motions based on the mental states by tuning the parameters of the spring-damper model shown in equation (<xref ref-type="disp-formula" rid="E1">1</xref>).</p>
<p>We experimentally showed that our participants properly modulated the android&#x02019;s motions based on its desired internal states.</p>
</sec>
<sec id="S6-14">
<label>6.2</label> <title>Experimental Setup</title>
<p>Some researchers revealed that motion properties, e.g., magnitude and speed, vary by emotional states (Amaya et al., <xref ref-type="bibr" rid="B1">1996</xref>; Michalak et al., <xref ref-type="bibr" rid="B29">2009</xref>; Masuda and Kato, <xref ref-type="bibr" rid="B28">2010</xref>; Gross et al., <xref ref-type="bibr" rid="B14">2012</xref>). Consequently, we transformed independent variables <italic>j</italic>, <italic>d</italic>, and <italic>k</italic> in equation (<xref ref-type="disp-formula" rid="E1">1</xref>) to <italic>&#x003C9;</italic><sub>0</sub>, <italic>&#x003BE;</italic>, and <italic>&#x003D5;</italic> in equation (<xref ref-type="disp-formula" rid="E10">10</xref>) so that the participants can intuitively change the speed and magnitude of the android&#x02019;s motions. <italic>&#x003C9;</italic><sub>0</sub>, which is the natural angular frequency, is equivalent to the time that a motion needs to be converged. <italic>&#x003BE;</italic> is the damping ratio. The damping property depends on <italic>&#x003BE;</italic> (<italic>&#x003BE;</italic>&#x02009;&#x0003E;&#x02009;1: over damping, <italic>&#x003BE;</italic>&#x02009;&#x0003D;&#x02009;1: critical damping, <italic>&#x003BE;</italic>&#x02009;&#x0003C;&#x02009;1: damped oscillation). <italic>&#x003D5;</italic> is the reciprocal of inertia, and the magnitude of motion is proportional to <italic>&#x003D5;</italic>:
<disp-formula id="E10"><label>(10)</label><mml:math id="M12"><mml:mtable columnalign="left" class="align"><mml:mtr><mml:mtd columnalign="left" class="align-odd"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo class="MathClass-op">&#x002D9;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x003BE;</mml:mi><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo class="MathClass-op">&#x002D9;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-bin">&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="italic">base</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mi mathvariant="italic">Dir</mml:mi><mml:mrow><mml:mo class="MathClass-open">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo class="MathClass-close">)</mml:mo></mml:mrow><mml:mspace width="0.5em"/><mml:mi>j</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow></mml:mfrac><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>k</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow></mml:mfrac><mml:mo class="MathClass-punc">,</mml:mo><mml:mi>d</mml:mi><mml:mo class="MathClass-rel">&#x0003D;</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003BE;</mml:mi><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow></mml:mfrac><mml:mo class="MathClass-punc">.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In this experiment, we set the damping ratio (<italic>&#x003BE;</italic>&#x02009;&#x0003D;&#x02009;1). The human head rhythmically moves in synchronization to speech. Since this movement is not damped, the damping ratio should be <italic>&#x003BE;</italic>&#x02009;&#x02248;&#x02009;1. The participants modulated <italic>&#x003C9;</italic><sub>0</sub> and <italic>&#x003D5;</italic> through the graphical user interface (GUI) in Figure <xref ref-type="fig" rid="F13">13</xref>A. The range of <italic>&#x003C9;</italic><sub>0</sub> was set at 1&#x02013;10 in 0.5 steps, and the range of <italic>&#x003D5;</italic> was set to 10<sup>0</sup>&#x02013;10<sup>2</sup> in 10<sup>0.05</sup> steps. The android speaks by playing a prerecorded voice while its neck and mouth move in synchrony with its voice. The voice was played from speakers on its head. We delayed the voice by 333&#x02009;ms for synchronization with the motion because of the calculation delay of the prosodic features and the communication time with the android. The setting&#x02019;s block diagram is shown in Figure <xref ref-type="fig" rid="F13">13</xref>B. The motion properties were changed immediately in response to the changes in <italic>&#x003C9;</italic><sub>0</sub> and <italic>&#x003D5;</italic>. The participants explored the GUI to find proper <italic>&#x003C9;</italic><sub>0</sub> and <italic>&#x003D5;</italic> values for expressing the desired emotion while looking at the android&#x02019;s motions.</p>
<fig position="float" id="F13">
<label>Figure 13</label>
<caption><p>Experimental setup. <bold>(A)</bold> Interface for adjusting motion parameters. <bold>(B)</bold> Block diagram.</p></caption>
<graphic xlink:href="frobt-04-00049-g013.tif"/>
</fig>
</sec>
<sec id="S6-15">
<label>6.3</label> <title>Experimental Conditions and Measurements</title>
<p>We chose four emotional states (happy, bored, relaxed, and tense) based on Russell&#x02019;s emotion model (Russell, <xref ref-type="bibr" rid="B36">1980</xref>), and the participants searched for the values of <italic>&#x003C9;</italic><sub>0</sub> and <italic>&#x003D5;</italic> values to feel the android&#x02019;s motion expressed by each emotion.</p>
<p>The difficulty of tuning the motions depends on the speech voice (the tuning was easy for some voices but not for the others). Therefore, the speech voice was fixed in all the emotions and for all participants. This voice sample was approximately a 1-min recording of a female experimental assistant reading a news story aloud while maintaining a neutral mental state.</p>
<p>The android changed its eye direction and facial expression to express its emotional state. This is because people felt difficulty exploring the parameters without facial expression in the preliminary trial. The eye movements and facial expression follow Ekman&#x02019;s report (Ekman and Friesen, <xref ref-type="bibr" rid="B6">1981</xref>). In happy and relaxed emotions, the android smiled by lifting the angle of her mouth up and cyclically moving her eyes horizontally. In bored and tense emotions, it grimaced with her eyelids down and rolled her eyes downward.</p>
<p>The participants sat in front of the android and adopted the parameters <italic>&#x003C9;</italic><sub>0</sub> and <italic>&#x003D5;</italic> using the GUI shown in Figure <xref ref-type="fig" rid="F13">13</xref>A. For each emotion preset, they searched for the parameter values that produced the android&#x02019;s desired emotional expression. The order of the four emotion conditions was counterbalanced among the participants.</p>
<p>To evaluate the degree of easiness of this modulation, we measured how long it took to modulate the motions. To the evaluate the degree of satisfaction with the modulation, the participants also answered a 7-point Likert scale questionnaire (1: unsatisfactory, &#x0007E;4: not sure, &#x0007E;7: satisfactory). Based on satisfactory scores, we wanted to evaluate whether android motions were produced as the participants expected. High scores denote successfully produced motions.</p>
</sec>
<sec id="S6-16">
<label>6.4</label> <title>Results</title>
<p>Twelve subjects (six males, six females, average age: 20.4, SD: 1.0) participated in the experiment. They were recruited in the same manner for the above two experiments.</p>
<p>Table <xref ref-type="table" rid="T1">1</xref> shows the elapsed needed time to modulate the motions. The voice used in this experiment was about 1&#x02009;min, and the results revealed that the participants fixed the parameters within 3&#x02013;5 repetitions. They did not take that much time for motion tuning, even though the results were not compared with those of other methods. Figure <xref ref-type="fig" rid="F14">14</xref> shows the average degree of satisfaction with the modulated motions. Because the average scores of all the emotion exceeded four, the participants felt the modulated motions expressed the desired emotion. Even tough some experienced difficulty tuning the motion due to incongruity between the manner of speaking (speech in neutral emotion) and the emotion, they successfully produced the desired emotional motion. A statistical test revealed that the average scores of happy, relaxed, and tense were significantly higher than four (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01), but not bored (<italic>p</italic>&#x02009;&#x0003C;&#x02009;0.1).<xref ref-type="fn" rid="fn2"><sup>2</sup></xref> Expressing the bored motion was difficult with the trunk motion.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Elapsed time to modulate the motions [s].</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="center"/>
<th align="center"><bold>Happy</bold></th>
<th align="center"><bold>Bored</bold></th>
<th align="center"><bold>Relaxed</bold></th>
<th align="center"><bold>Tense</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Average</td>
<td align="center">305</td>
<td align="center">294</td>
<td align="center">188</td>
<td align="center">278</td>
</tr>
<tr>
<td align="left">SD</td>
<td align="center">205</td>
<td align="center">266</td>
<td align="center">132</td>
<td align="center">213</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig position="float" id="F14">
<label>Figure 14</label>
<caption><p>Degree of satisfaction with the modulated motions.</p></caption>
<graphic xlink:href="frobt-04-00049-g014.tif"/>
</fig>
<p>Figure <xref ref-type="fig" rid="F15">15</xref> shows examples of modulated motion (time series of head angle) and corresponding voice data. There is a tendency for the magnitude and speed of the motions to increase in the order of bored, relaxed, happy, and tense. This result shows that the modulated motion properties are different based on the emotional states.</p>
<fig position="float" id="F15">
<label>Figure 15</label>
<caption><p>Examples of emotional motion modulated by the participants. <bold>(A)</bold> Bored motion, <bold>(B)</bold> relaxed motion, <bold>(C)</bold> happy motion, and <bold>(D)</bold> tense motion.</p></caption>
<graphic xlink:href="frobt-04-00049-g015.tif"/>
</fig>
</sec>
<sec id="S6-17">
<label>6.5</label> <title>Discussion</title>
<p>One the other hand, the existing learning-based systems need to additionally collect human emotional motion data when we generate the android&#x02019;s emotional expressions that are not included in the learning data. The proposed model has an advantage with facility in the motion development.</p>
<p>Figure <xref ref-type="fig" rid="F14">14</xref> shows the participants were not sufficiently satisfied by the generated motions of bored. Such motions were expressed by small trunk movements, but they might have thought that just making small motions is insufficient for the bored state. They wanted to add a facial expression or a voice to distinctively express boredom but they could not; therefore, their degree of satisfaction fell. To precisely express the emotional states in the android, the facial expression, loudness of voice and speech rate must be involved in the motion modulation. In the future, our system needs to implement this idea.</p>
<p>The speech voice was fixed in the experiment, but the speech pattern is usually changed owing to the speaker&#x02019;s emotional state. For example, the voice usually becomes louder and faster in the tense and happy states and lower and slower in the bored and relaxed states. This is similar to the relation between the magnitude and speed of the motions and the emotional state shown in Figure <xref ref-type="fig" rid="F15">15</xref>. A tense state brings larger and faster motions and a relaxed state brings smaller and slower motions. This suggests that the motion and speech patterns are similarly changed by the emotional state, such as the magnitude, and the speed of voice and motion become larger and faster in the tense and happy states and smaller and slower in the bored and relaxed states. If this relation generally holds, the proposed model can generate emotional speech motions based on emotional voices without tuning the parameters. Furthermore, developing a system of automatic emotional voice and motion generation is possible by integrating our model with the emotional voice synthesization (Murray and Arnott, <xref ref-type="bibr" rid="B30">2008</xref>). Future work should reveal the relation for this system. Here, the relation might depend on the age and gender of the speaker since a gender difference exists in non-verbal behaviors (Frances, <xref ref-type="bibr" rid="B10">1979</xref>). For example, male speakers might express more emotional motions. To use our model in different types of androids, such dependencies must be investigated in future work.</p>
</sec>
</sec>
<sec id="S7">
<label>7</label> <title>Future Works</title>
<p>People tend to anticipate some properties of a speaker&#x02019;s motion when they hear her voice, for example, a vigorous movement is expected from a cheerful voice. We found fewer differences between the proposed and copy methods in the Speaker H comparison (Figure <xref ref-type="fig" rid="F12">12</xref>) because the participants might have expected much movement from her voice. The proposed model generates the motion based on pitch, power, and vowel information, but we also expect other voice features to be related. Identifying the relation between voice features and expected motion is critical. Moreover, human movement has some randomness that should be implemented in the method. This study tested the proposed method for Japanese speech, but the idea of synchronization between motions and voice features could more effectively work for English speech, since English speakers show a higher correlation between speech and head movement than Japanese speakers (Yehia et al., <xref ref-type="bibr" rid="B45">2002</xref>). Furthermore, the results here were only obtained from university students. Although elderly people prefer a human-like to a robot-like appearance, younger adults prefer the opposite (Prakash and Rogers, <xref ref-type="bibr" rid="B35">2014</xref>). Therefore, the influence of the android&#x02019;s appearance and the participants&#x02019; age must also be investigated in future work.</p>
<p>We evaluated our proposed method for subjective impressions but not for behavioral aspects. The participant behavior or attitude toward the android might change if they had a feeling of human-likeness about it. For instance, gaze behavior (Shimada and Ishiguro, <xref ref-type="bibr" rid="B41">2008</xref>) and posture are influenced by the relationship between conversation partners. In future work, a behavioral evaluation is necessary to scrutinize the method&#x02019;s effects.</p>
</sec>
<sec id="S8">
<label>8</label> <title>Conclusion</title>
<p>We developed a method that automatically generates humanlike trunk motions of a conversational android from its speech in real time. By enforcing the speech motions that are strongly related to prosodic features, the method compensates for the negative impressions caused by incomplete copying of human motions from which conventional methods suffer. By simulating human trunk movement based on a spring-damper dynamical model, the motion can be modulated based on the android&#x02019;s internal states. Our experimental results show that the android motions generated by the model appear natural and motivate people to talk more with the android, even though the generated motions are different from the original human motions. The results also suggest that the model can generate humanlike motions for any speech pattern. The additional capability of enforcing multimodal synchronization must be applicable to other types of humanoid robots and other modalities than speech and motion. Such possibilities must be investigated in future work.</p>
</sec>
<sec id="S9">
<title>Ethics Statement</title>
<p>The study was approved by the Ethics Committee of the Advanced Telecommunications Research Institute International (Kyoto, Japan). In the beginning of the experiment, we explained about this study to the subjects who were university students and received informed consent from them. We used a job-offering site for university students, and the subjects were recruited.</p>
</sec>
<sec id="S10">
<title>Author Contributions</title>
<p>KS proposed the idea of this speech-driven trunk motion generating system, build the experiment system, conducted the experiment, analyzed the result, and wrote this article. TM designed the experiment, analyzed and evaluated the result. CI worked in the system development, especially sound signal analysis. HI proposed the basic idea to generate humanlike movements under a mechanical limitation of the android (i.e., limited number of joint).</p>
</sec>
<sec id="S11">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This research was supported by the Japan Science and Technology Agency, ERATO, ISHIGURO symbiotic Human-Robot Interaction Project, Grant Number JPMJER1401.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amaya</surname> <given-names>K.</given-names></name> <name><surname>Bruderlin</surname> <given-names>A.</given-names></name> <name><surname>Calvert</surname> <given-names>T.</given-names></name></person-group> (<year>1996</year>). &#x0201C;<article-title>Emotion from motion</article-title>,&#x0201D; in <source>Graphics Interface</source>, Vol. <volume>96</volume>, <publisher-loc>Toronto</publisher-loc>, <fpage>222</fpage>&#x02013;<lpage>229</lpage>.</citation></ref>
<ref id="B2"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Andrist</surname> <given-names>S.</given-names></name> <name><surname>Tan</surname> <given-names>X. Z.</given-names></name> <name><surname>Gleicher</surname> <given-names>M.</given-names></name> <name><surname>Mutlu</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). &#x0201C;<article-title>Conversational gaze aversion for humanlike robots</article-title>,&#x0201D; in <conf-name>Proceedings of the 2014 ACM/IEEE International Conference on Human-Robot Interaction</conf-name>, <conf-loc>New York, NY</conf-loc>, <fpage>25</fpage>&#x02013;<lpage>32</lpage>.</citation></ref>
<ref id="B3"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Bolinger</surname> <given-names>D.</given-names></name></person-group> (<year>1985</year>). <source>Intonation and Its Parts: Melody in Spoken English</source>. <publisher-loc>Stanford</publisher-loc>: <publisher-name>Stanford University Press</publisher-name>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Busso</surname> <given-names>C.</given-names></name> <name><surname>Deng</surname> <given-names>Z.</given-names></name> <name><surname>Neumann</surname> <given-names>U.</given-names></name> <name><surname>Narayanan</surname> <given-names>S.</given-names></name></person-group> (<year>2005</year>). <article-title>Natural head motion synthesis driven by acoustic prosodic features</article-title>. <source>Comput. Animat. Virtual Worlds</source> <volume>16</volume>, <fpage>283</fpage>&#x02013;<lpage>290</lpage>.<pub-id pub-id-type="doi">10.1002/cav.80</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chartrand</surname> <given-names>T. L.</given-names></name> <name><surname>Bargh</surname> <given-names>J. A.</given-names></name></person-group> (<year>1999</year>). <article-title>The chameleon effect</article-title>. <source>J. Pers. Soc. Psychol.</source> <volume>76</volume>, <fpage>893</fpage>&#x02013;<lpage>910</lpage>.<pub-id pub-id-type="doi">10.1037/0022-3514.76.6.893</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ekman</surname> <given-names>P.</given-names></name> <name><surname>Friesen</surname> <given-names>W. V.</given-names></name></person-group> (<year>1981</year>). <article-title>The repertoire of nonverbal behavior: categories, origins, usage, and coding</article-title>. <source>Nonverbal Commun. Interact. Gesture</source> <fpage>57</fpage>&#x02013;<lpage>106</lpage>.</citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eriksson</surname> <given-names>P.-O.</given-names></name> <name><surname>Zafar</surname> <given-names>H.</given-names></name> <name><surname>Nordh</surname> <given-names>E.</given-names></name></person-group> (<year>1998</year>). <article-title>Concomitant mandibular and head-neck movements during jaw opening-closing in man</article-title>. <source>J. Oral Rehabil.</source> <volume>25</volume>, <fpage>859</fpage>&#x02013;<lpage>870</lpage>.<pub-id pub-id-type="doi">10.1046/j.1365-2842.1998.00333.x</pub-id><pub-id pub-id-type="pmid">9846906</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ernst</surname> <given-names>M. O.</given-names></name> <name><surname>Banks</surname> <given-names>M. S.</given-names></name></person-group> (<year>2002</year>). <article-title>Humans integrate visual and haptic information in a statistically optimal fashion</article-title>. <source>Nature</source> <volume>415</volume>, <fpage>429</fpage>&#x02013;<lpage>433</lpage>.<pub-id pub-id-type="doi">10.1038/415429a</pub-id><pub-id pub-id-type="pmid">11807554</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Foster</surname> <given-names>M. E.</given-names></name> <name><surname>Oberlander</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Corpus-based generation of head and eyebrow motion for an embodied conversational agent</article-title>. <source>Lang. Resour. Eval.</source> <volume>41</volume>, <fpage>305</fpage>&#x02013;<lpage>323</lpage>.<pub-id pub-id-type="doi">10.1007/s10579-007-9055-3</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frances</surname> <given-names>S. J.</given-names></name></person-group> (<year>1979</year>). <article-title>Sex differences in nonverbal behavior</article-title>. <source>Sex Roles</source> <volume>5</volume>, <fpage>519</fpage>&#x02013;<lpage>535</lpage>.<pub-id pub-id-type="doi">10.1007/BF00287326</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fridlund</surname> <given-names>A. J.</given-names></name> <name><surname>Hatfield</surname> <given-names>M. E.</given-names></name> <name><surname>Cottam</surname> <given-names>G. L.</given-names></name> <name><surname>Fowler</surname> <given-names>S. C.</given-names></name></person-group> (<year>1986</year>). <article-title>Anxiety and striate-muscle activation: evidence from electromyographic pattern analysis</article-title>. <source>J. Abnorm. Psychol.</source> <volume>95</volume>, <fpage>228</fpage>.<pub-id pub-id-type="doi">10.1037/0021-843X.95.3.228</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fujisaki</surname> <given-names>W.</given-names></name> <name><surname>Goda</surname> <given-names>N.</given-names></name> <name><surname>Motoyoshi</surname> <given-names>I.</given-names></name> <name><surname>Komatsu</surname> <given-names>H.</given-names></name> <name><surname>Nishida</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>Audiovisual integration in the human perception of materials</article-title>. <source>J. Vis.</source> <volume>14</volume>, <fpage>1</fpage>&#x02013;<lpage>20</lpage>.<pub-id pub-id-type="doi">10.1167/14.4.12</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gielniak</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>K.</given-names></name> <name><surname>Thomaz</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Generating human-like motion for robots</article-title>. <source>Int. J. Robot. Res.</source> <volume>32</volume>, <fpage>1275</fpage>&#x02013;<lpage>1301</lpage>.<pub-id pub-id-type="doi">10.1177/0278364913490533</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gross</surname> <given-names>M. M.</given-names></name> <name><surname>Crane</surname> <given-names>E. A.</given-names></name> <name><surname>Fredrickson</surname> <given-names>B. L.</given-names></name></person-group> (<year>2012</year>). <article-title>Effort-Shape and kinematic assessment of bodily expression of emotion during gait</article-title>. <source>Hum. Mov. Sci.</source> <volume>31</volume>, <fpage>202</fpage>&#x02013;<lpage>221</lpage>.<pub-id pub-id-type="doi">10.1016/j.humov.2011.05.001</pub-id><pub-id pub-id-type="pmid">21835480</pub-id></citation></ref>
<ref id="B15"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Hashimoto</surname> <given-names>T.</given-names></name> <name><surname>Kobayashi</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). &#x0201C;<article-title>Study on natural head motion in waiting state with receptionist robot SAYA that has human-like appearance</article-title>,&#x0201D; in <conf-name>Proceedings of 2009 IEEE Workshop on Robotic Intelligence in Informationally Structured Space</conf-name>, <conf-loc>Nashville, TN</conf-loc>, <fpage>93</fpage>&#x02013;<lpage>98</lpage>.</citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ishi</surname> <given-names>C. T.</given-names></name></person-group> (<year>2005</year>). <article-title>Perceptually-related F0 parameters for automatic classification of phrase final tones</article-title>. <source>IEICE Trans. Inform. Syst.</source> <volume>88</volume>, <fpage>481</fpage>&#x02013;<lpage>488</lpage>.<pub-id pub-id-type="doi">10.1093/ietisy/e88-d.3.481</pub-id></citation></ref>
<ref id="B17"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ishi</surname> <given-names>C. T.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name> <name><surname>Hagita</surname> <given-names>N.</given-names></name></person-group> (<year>2011</year>). &#x0201C;<article-title>Speech-driven lip motion generation for tele-operated humanoid robots</article-title>,&#x0201D; in <source>Auditory-Visual Speech Processing</source>, <publisher-loc>Volterra</publisher-loc>, <fpage>131</fpage>&#x02013;<lpage>135</lpage>.</citation></ref>
<ref id="B18"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ishi</surname> <given-names>C. T.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name> <name><surname>Hagita</surname> <given-names>N.</given-names></name> <name><surname>Robotics</surname> <given-names>I.</given-names></name> <name><surname>Labs</surname> <given-names>C.</given-names></name></person-group> (<year>2012</year>). <source>Evaluation of Formant-Based Lip Motion Generation in Tele-Operated Humanoid Robots</source>. <publisher-loc>Vilamoura</publisher-loc>: <publisher-name>IEEE</publisher-name>, <fpage>2377</fpage>&#x02013;<lpage>2382</lpage>.</citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Meng</surname> <given-names>H. M.</given-names></name> <name><surname>Cai</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Head and facial gestures synthesis using PAD model for an expressive talking avatar</article-title>. <source>Multimed. Tools Appl.</source> <volume>73</volume>, <fpage>439</fpage>&#x02013;<lpage>461</lpage>.<pub-id pub-id-type="doi">10.1007/s11042-013-1604-8</pub-id></citation></ref>
<ref id="B20"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Kirchhof</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). &#x0201C;<article-title>Desynchronized speech-gesture signals still get the message across</article-title>,&#x0201D; in <conf-name>International Conference on Multimodality</conf-name>, <conf-loc>Hongkong</conf-loc>.</citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Komatsu</surname> <given-names>T.</given-names></name> <name><surname>Yamada</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Adaptation gap hypothesis: how differences between users&#x02019; expected and perceived agent functions affect their subjective impression</article-title>. <source>J. Syst. Cybern. Inform.</source> <volume>9</volume>, <fpage>67</fpage>&#x02013;<lpage>74</lpage>.</citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kondo</surname> <given-names>Y.</given-names></name> <name><surname>Takemura</surname> <given-names>K.</given-names></name> <name><surname>Takamatsu</surname> <given-names>J.</given-names></name> <name><surname>Ogasawara</surname> <given-names>T.</given-names></name></person-group> (<year>2013</year>). <article-title>A gesture-centric android system for multi-party human-robot interaction</article-title>. <source>J. Hum.Robot Interact.</source> <volume>2</volume>, <fpage>133</fpage>&#x02013;<lpage>151</lpage>.<pub-id pub-id-type="doi">10.5898/JHRI.2.1.Kondo</pub-id></citation></ref>
<ref id="B23"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Laban</surname> <given-names>R. V.</given-names></name></person-group> (<year>1988</year>). <source>The Mastery of Movement</source>. <publisher-name>Princeton Book Co. Pub</publisher-name>.</citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Larsen</surname> <given-names>R. J.</given-names></name> <name><surname>Shackelford</surname> <given-names>T. K.</given-names></name></person-group> (<year>1996</year>). <article-title>Gaze avoidance: personality and social judgments of people who avoid direct face-to-face contact</article-title>. <source>Pers. Individ. Dif.</source> <volume>21</volume>, <fpage>907</fpage>&#x02013;<lpage>917</lpage>.<pub-id pub-id-type="doi">10.1016/S0191-8869(96)00148-1</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>B. H.</given-names></name> <name><surname>Ma</surname> <given-names>X.</given-names></name> <name><surname>Deng</surname> <given-names>Z.</given-names></name></person-group> (<year>2012</year>). <article-title>Live speech driven head-and-eye motion generators</article-title>. <source>Vis. Comput. Graph.</source> <volume>18</volume>, <fpage>1902</fpage>&#x02013;<lpage>1914</lpage>.<pub-id pub-id-type="doi">10.1109/TVCG.2012.74</pub-id><pub-id pub-id-type="pmid">22392712</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>C.-C.</given-names></name> <name><surname>Chiang</surname> <given-names>C.-F.</given-names></name></person-group> (<year>2006</year>). <article-title>A study on biodynamic models of seated human subjects exposed to vertical vibration</article-title>. <source>Int. J. Ind. Ergon.</source> <volume>36</volume>, <fpage>869</fpage>&#x02013;<lpage>890</lpage>.<pub-id pub-id-type="doi">10.1016/j.ergon.2006.06.008</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Linder</surname> <given-names>A.</given-names></name></person-group> (<year>2000</year>). <article-title>A new mathematical neck model for a low-velocity rear-end impact dummy: evaluation of components influencing head kinematics</article-title>. <source>Accid. Anal. Prev.</source> <volume>32</volume>, <fpage>261</fpage>&#x02013;<lpage>269</lpage>.<pub-id pub-id-type="doi">10.1016/S0001-4575(99)00085-8</pub-id><pub-id pub-id-type="pmid">10688482</pub-id></citation></ref>
<ref id="B28"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Masuda</surname> <given-names>M.</given-names></name> <name><surname>Kato</surname> <given-names>S.</given-names></name></person-group> (<year>2010</year>). &#x0201C;<article-title>Motion rendering system for emotion expression of human form robots based on Laban movement analysis</article-title>,&#x0201D; in <conf-name>Proceedings of the 19th International Symposium in Robot and Human Interactive Communication</conf-name>, <conf-loc>Viareggio</conf-loc>, <fpage>324</fpage>&#x02013;<lpage>329</lpage>.</citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Michalak</surname> <given-names>J.</given-names></name> <name><surname>Troje</surname> <given-names>N. F.</given-names></name> <name><surname>Fischer</surname> <given-names>J.</given-names></name> <name><surname>Vollmar</surname> <given-names>P.</given-names></name> <name><surname>Heidenreich</surname> <given-names>T.</given-names></name> <name><surname>Schulte</surname> <given-names>D.</given-names></name></person-group> (<year>2009</year>). <article-title>Embodiment of sadness and depression&#x02013;gait patterns associated with dysphoric mood</article-title>. <source>Psychosom. Med.</source> <volume>71</volume>, <fpage>580</fpage>&#x02013;<lpage>587</lpage>.<pub-id pub-id-type="doi">10.1097/PSY.0b013e3181a2515c</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murray</surname> <given-names>I. R.</given-names></name> <name><surname>Arnott</surname> <given-names>J. L.</given-names></name></person-group> (<year>2008</year>). <article-title>Applying an analysis of acted vocal emotions to improve the simulation of synthetic speech</article-title>. <source>Comput. Speech Lang.</source> <volume>22</volume>, <fpage>107</fpage>&#x02013;<lpage>129</lpage>.<pub-id pub-id-type="doi">10.1016/j.csl.2007.06.001</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakano</surname> <given-names>A.</given-names></name> <name><surname>Hoshino</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Composite conversation gesture synthesis using layered planning</article-title>. <source>Syst. Comput. Japan</source> <volume>38</volume>, <fpage>58</fpage>&#x02013;<lpage>68</lpage>.<pub-id pub-id-type="doi">10.1002/scj.20532</pub-id></citation></ref>
<ref id="B32"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Nakazawa</surname> <given-names>M.</given-names></name> <name><surname>Nishimoto</surname> <given-names>T.</given-names></name> <name><surname>Sagayama</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). &#x0201C;<article-title>Behavior generation for spoken dialogue agent by dynamical model</article-title>,&#x0201D; in <conf-name>Proceedings of Human-Agent Interaction Symposium (in Japanese)</conf-name>, <conf-loc>Tokyo</conf-loc>, <fpage>2C</fpage>&#x02013;<lpage>1</lpage>.</citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Piwek</surname> <given-names>L.</given-names></name> <name><surname>McKay</surname> <given-names>L. S.</given-names></name> <name><surname>Pollick</surname> <given-names>F. E.</given-names></name></person-group> (<year>2014</year>). <article-title>Empirical evaluation of the uncanny valley hypothesis fails to confirm the predicted effect of motion</article-title>. <source>Cognition</source> <volume>130</volume>, <fpage>271</fpage>&#x02013;<lpage>277</lpage>.<pub-id pub-id-type="doi">10.1016/j.cognition.2013.11.001</pub-id><pub-id pub-id-type="pmid">24374019</pub-id></citation></ref>
<ref id="B34"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Pollard</surname> <given-names>N. S.</given-names></name> <name><surname>Hodgins</surname> <given-names>J. K.</given-names></name> <name><surname>Riley</surname> <given-names>M. J.</given-names></name> <name><surname>Atkeson</surname> <given-names>C. G.</given-names></name></person-group> (<year>2002</year>). &#x0201C;<article-title>Adapting human motion for the control of a humanoid robot</article-title>,&#x0201D; in <conf-name>Proceedings of the 2002 IEEE International Conference on Robotics and Automation</conf-name>, Vol. <volume>2</volume>, <conf-loc>Washington, DC</conf-loc>, <fpage>1390</fpage>&#x02013;<lpage>1397</lpage>.</citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prakash</surname> <given-names>A.</given-names></name> <name><surname>Rogers</surname> <given-names>W. A.</given-names></name></person-group> (<year>2014</year>). <article-title>Why some humanoid faces are perceived more positively than others: effects of human-likeness and task</article-title>. <source>Int. J. Soc. Robot.</source> <volume>7</volume>, <fpage>309</fpage>&#x02013;<lpage>331</lpage>.<pub-id pub-id-type="doi">10.1007/s12369-014-0269-4</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russell</surname> <given-names>J. A.</given-names></name></person-group> (<year>1980</year>). <article-title>A circumplex model of affect</article-title>. <source>Personal. Soc. Psychol.</source> <volume>39</volume>, <fpage>1161</fpage>&#x02013;<lpage>1178</lpage>.<pub-id pub-id-type="doi">10.1037/h0077714</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sainsbury</surname> <given-names>P.</given-names></name> <name><surname>Gibson</surname> <given-names>J. G.</given-names></name></person-group> (<year>1954</year>). <article-title>Symptoms of anxiety and tension and the accompanying physiological changes in the muscular system</article-title>. <source>J. Neurol. Neurosurg. Psychiatr.</source> <volume>17</volume>, <fpage>216</fpage>&#x02013;<lpage>224</lpage>.<pub-id pub-id-type="doi">10.1136/jnnp.17.3.216</pub-id></citation></ref>
<ref id="B38"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Sakai</surname> <given-names>K.</given-names></name> <name><surname>Ishi</surname> <given-names>C. T.</given-names></name> <name><surname>Minato</surname> <given-names>T.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). &#x0201C;<article-title>Online speech-driven head motion generating system and evaluation on a tele-operated robot</article-title>,&#x0201D; in <conf-name>Proceedings of the 24th International Symposium in Robot and Human Interactive Communication</conf-name>, <conf-loc>Kobe</conf-loc>, <fpage>529</fpage>&#x02013;<lpage>534</lpage>.</citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salem</surname> <given-names>M.</given-names></name> <name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>I.</given-names></name> <name><surname>Rohlfing</surname> <given-names>K.</given-names></name> <name><surname>Joublin</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>). <article-title>Generation and evaluation of communicative robot gesture</article-title>. <source>Int. J. Soc. Robot.</source> <volume>4</volume>, <fpage>201</fpage>&#x02013;<lpage>217</lpage>.<pub-id pub-id-type="doi">10.1007/s12369-011-0124-9</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sargin</surname> <given-names>M. E.</given-names></name> <name><surname>Yemez</surname> <given-names>Y.</given-names></name> <name><surname>Erzin</surname> <given-names>E.</given-names></name> <name><surname>Tekalp</surname> <given-names>A. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Analysis of head gesture and prosody patterns for prosody-driven head-gesture animation</article-title>. <source>IEEE. Trans. Pattern. Anal. Mach. Intell.</source> <volume>30</volume>, <fpage>1330</fpage>&#x02013;<lpage>1345</lpage>.<pub-id pub-id-type="doi">10.1109/TPAMI.2007.70797</pub-id></citation></ref>
<ref id="B41"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Shimada</surname> <given-names>M.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name></person-group> (<year>2008</year>). &#x0201C;<article-title>Motion behavior and its influence on human-likeness in an android robot</article-title>,&#x0201D; in <conf-name>Proceedings of the 30th Annual Meeting of the Cognitive Science Society</conf-name>, <conf-loc>Washington, DC</conf-loc>, <fpage>2468</fpage>&#x02013;<lpage>2473</lpage>.</citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tinwell</surname> <given-names>A.</given-names></name> <name><surname>Grimshaw</surname> <given-names>M.</given-names></name> <name><surname>Nabi</surname> <given-names>D. A.</given-names></name></person-group> (<year>2015</year>). <article-title>The effect of onset asynchrony in audio visual speech and the uncanny valley in virtual characters</article-title>. <source>Int. J. Mech. Robot. Syst.</source> <volume>2</volume>, <fpage>97</fpage>&#x02013;<lpage>110</lpage>.<pub-id pub-id-type="doi">10.1504/IJMRS.2015.068991</pub-id></citation></ref>
<ref id="B43"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Watanabe</surname> <given-names>M.</given-names></name> <name><surname>Ogawa</surname> <given-names>K.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). &#x0201C;<article-title>Can androids be salespeople in the real world?</article-title>&#x0201D; in <conf-name>Proceedings of the ACM Conference Extended Abstracts on Human Factors in Computing Systems</conf-name>, <conf-loc>Seoul</conf-loc>, <fpage>781</fpage>&#x02013;<lpage>788</lpage>.</citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Watanabe</surname> <given-names>T.</given-names></name> <name><surname>Okubo</surname> <given-names>M.</given-names></name> <name><surname>Nakashige</surname> <given-names>M.</given-names></name> <name><surname>Danbara</surname> <given-names>R.</given-names></name></person-group> (<year>2004</year>). <article-title>InterActor: speech-driven embodied interactive actor</article-title>. <source>Int. J. Hum. Comput. Interact.</source> <volume>17</volume>, <fpage>43</fpage>&#x02013;<lpage>60</lpage>.<pub-id pub-id-type="doi">10.1207/s15327590ijhc1701_4</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yehia</surname> <given-names>H. C.</given-names></name> <name><surname>Kuratate</surname> <given-names>T.</given-names></name> <name><surname>Vatikiotis-Bateson</surname> <given-names>E.</given-names></name></person-group> (<year>2002</year>). <article-title>Linking facial animation, head motion and speech acoustics</article-title>. <source>J. Phon.</source> <volume>30</volume>, <fpage>555</fpage>&#x02013;<lpage>568</lpage>.<pub-id pub-id-type="doi">10.1006/jpho.2002.0165</pub-id></citation></ref>
<ref id="B46"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Yoshikawa</surname> <given-names>M.</given-names></name> <name><surname>Matsumoto</surname> <given-names>Y.</given-names></name> <name><surname>Sumitani</surname> <given-names>M.</given-names></name> <name><surname>Ishiguro</surname> <given-names>H.</given-names></name></person-group> (<year>2011</year>). &#x0201C;<article-title>Development of an android robot for psychological support in medical and welfare fields</article-title>,&#x0201D; in <conf-name>Proceedings of the 2011 IEEE International Conference on Robotics and Biomimetics</conf-name>, <conf-loc>Phuket</conf-loc>, <fpage>2378</fpage>&#x02013;<lpage>2383</lpage>.</citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zafar</surname> <given-names>H.</given-names></name> <name><surname>Nordh</surname> <given-names>E.</given-names></name> <name><surname>Eriksson</surname> <given-names>P.-O.</given-names></name></person-group> (<year>2000</year>). <article-title>Temporal coordination between mandibular and head-neck movements during jaw opening-closing tasks in man</article-title>. <source>Arch. Oral Biol.</source> <volume>45</volume>, <fpage>675</fpage>&#x02013;<lpage>682</lpage>.<pub-id pub-id-type="doi">10.1016/S0003-9969(00)00032-7</pub-id><pub-id pub-id-type="pmid">10869479</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zafar</surname> <given-names>H.</given-names></name> <name><surname>Nordh</surname> <given-names>E.</given-names></name> <name><surname>Eriksson</surname> <given-names>P.-O.</given-names></name></person-group> (<year>2002</year>). <article-title>Spatiotemporal consistency of human mandibular and head-neck movement trajectories during jaw opening-closing tasks</article-title>. <source>Exp. Brain Res.</source> <volume>146</volume>, <fpage>70</fpage>&#x02013;<lpage>76</lpage>.<pub-id pub-id-type="doi">10.1007/s00221-002-1157-y</pub-id><pub-id pub-id-type="pmid">12192580</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup>We use the Wilcoxon rank sum test because the normality was not assumed by a Shapiro&#x02013;Wilk test.</p></fn>
<fn id="fn2"><p><sup>2</sup>We used <italic>t</italic>-test if the normality was assumed by the Shapiro&#x02013;Wilk test; otherwise the Wilcoxon rank sum test was used.</p></fn>
</fn-group>
</back>
</article>