<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="review-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2023.1084000</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Recent advancements in multimodal human&#x02013;robot interaction</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Su</surname> <given-names>Hang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/948852/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Qi</surname> <given-names>Wen</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1105924/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chen</surname> <given-names>Jiahao</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/777598/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Yang</surname> <given-names>Chenguang</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/571027/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Sandoval</surname> <given-names>Juan</given-names></name>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1382972/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Laribi</surname> <given-names>Med Amine</given-names></name>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/558466/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Electronics, Information and Bioengineering, Politecnico di Milano</institution>, <addr-line>Milan</addr-line>, <country>Italy</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Future Technology, South China University of Technology</institution>, <addr-line>Guangzhou</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>State Key Laboratory of Management and Control for Complex Systems, Institute of Automation, Chinese Academy of Sciences</institution>, <addr-line>Beijing</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>Bristol Robotics Laboratory, University of the West of England</institution>, <addr-line>Bristol</addr-line>, <country>United Kingdom</country></aff>
<aff id="aff5"><sup>5</sup><institution>Department of GMSC, Pprime Institute, CNRS, ENSMA, University of Poitiers</institution>, <addr-line>Poitiers</addr-line>, <country>France</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Zhenshan Bing, Technical University of Munich, Germany</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Dinesh Bhatia, North Eastern Hill University, India; Nancie A. Gunson, Heriot-Watt University, United Kingdom</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Wen Qi <email>wenqi&#x00040;scut.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>05</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1084000</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>04</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Su, Qi, Chen, Yang, Sandoval and Laribi.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Su, Qi, Chen, Yang, Sandoval and Laribi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Robotics have advanced significantly over the years, and human&#x02013;robot interaction (HRI) is now playing an important role in delivering the best user experience, cutting down on laborious tasks, and raising public acceptance of robots. New HRI approaches are necessary to promote the evolution of robots, with a more natural and flexible interaction manner clearly the most crucial. As a newly emerging approach to HRI, multimodal HRI is a method for individuals to communicate with a robot using various modalities, including voice, image, text, eye movement, and touch, as well as bio-signals like EEG and ECG. It is a broad field closely related to cognitive science, ergonomics, multimedia technology, and virtual reality, with numerous applications springing up each year. However, little research has been done to summarize the current development and future trend of HRI. To this end, this paper systematically reviews the state of the art of multimodal HRI on its applications by summing up the latest research articles relevant to this field. Moreover, the research development in terms of the input signal and the output signal is also covered in this manuscript.</p>
</abstract>
<kwd-group>
<kwd>multi-modal signal processing</kwd>
<kwd>multi-modal feedback</kwd>
<kwd>multi-modal human&#x02013;robot interaction</kwd>
<kwd>physical human&#x02013;robot interaction</kwd>
<kwd>human&#x02013;robot interaction</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="160"/>
<page-count count="21"/>
<word-count count="16688"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Recent years have witnessed a huge leap in the advancement of robotics, yet it is quite challenging to build a robot that can communicate with individuals naturally and synthesize understandable multimodal motions in a variety of interaction scenarios. To deliver appropriate feedback, the robot requires a high level of multimodal recognition in order to comprehend the person&#x00027;s inner moods, goals, and character. Devices for human&#x02013;robot interaction (HRI) have become a common part of everyday life thanks to the growth of the Internet of Things. The input and output of a single sense modality, such as sight, touch, sound, scent, or flavor, is no longer the only option for HRI.</p>
<p>The goal of multimodal HRI is to communicate with a robot utilizing various multimodal signals (<xref ref-type="fig" rid="F1">Figure 1</xref>), including voice, image, text, eye movement, and touch. Multimodal HRI is a broad field that is closely associated with cognitive science, ergonomics, communication technologies, and virtual reality. It includes both multimodal input signals from humans to robots and multimodal output signals from robots to humans. As the carrier of the Internet of Things in the era of big data, multimodal HRI is closely connected to the advancement of visual effects, AI, sentimental data processing, psychological and physiological appraisal, distance education, as well as medical rehabilitative services.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The various signals of multimodal HRI.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1084000-g0001.tif"/>
</fig>
<p>The earliest studies on multimodal HRI date back to the 1990s, and several publications offer an interactive approach that combines voice and gesture. Furthermore, the rise of immersive visualization opens up a new multimodal interactive interface for HRI: an immersive world that blends visual, aural, tactile, and other sense modalities. Immersive visualization, which integrates multimodal channels and multimodalities, has become an inseparable part of high-dimensional big data visualization.</p>
<p>Great strides have been made in robotics over the past years, with human-computer interaction technology playing a critical role in improving the user experience, reducing tiresome processes, and promoting the acceptance of robots. Novel human-computer interaction strategies are necessary to further robotics progress, and a more natural and adaptable interaction style is particularly important (Fang et al., <xref ref-type="bibr" rid="B38">2019</xref>). In many application areas, robots must process output signals in the same way as human beings. Visual and auditory signals are the most straightforward methods for individuals to interact with home robots. With the advancement of statistical modeling, speech recognition has been increasingly employed in robotics and smart gadgets to enable natural language-based HRI. Furthermore, significant progress in picture recognition has been made (Xie et al., <xref ref-type="bibr" rid="B149">2020</xref>), with some robots able to comprehend instructions given to them in human language and perform necessary activities by combining visual and aural input signals.</p>
<p>This paper systematically follows the state of the art of multimodal HRI and thoroughly reviews the research progress in terms of the signal input, the signal output, and the applications of multimodal HRI (<xref ref-type="fig" rid="F2">Figure 2</xref>). Specifically, this article elaborates on the research progress of signal input of multimodal HRI from three perspectives: gesture input and recognition, speech input and recognition, as well as emotion input and recognition. In terms of information output, gesture generation and emotional expression generation are covered. The latest applications of multimodal HRI, including assistive mobile robots, robotic exoskeletons, as well as robotic prostheses, will also be introduced.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Major application areas of multimodal HRI (Manna and Bhaumik, <xref ref-type="bibr" rid="B88">2013</xref>; Klauer et al., <xref ref-type="bibr" rid="B68">2014</xref>; Gopinathan et al., <xref ref-type="bibr" rid="B46">2017</xref>; K&#x00027;&#x00301;ut&#x00027;&#x00301;uk et al., <xref ref-type="bibr" rid="B74">2019</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1084000-g0002.tif"/>
</fig>
</sec>
<sec sec-type="methods" id="s2">
<title>2. Methodology</title>
<p>By using Preferred Reporting Items for Systematic Reviews and Meta-analysis (PRISMA) guidelines, a systematic review of recently published literature was conducted on recent advancements in multimodal human&#x02013;robot Interaction (Page et al., <xref ref-type="bibr" rid="B101">2021</xref>). The inclusion criteria were: (i) publications indexed in the Web of Science, Scopus, and ProQuest databases; (ii) publication dates between 2008 and 2022; (iii) written in English; (iv) being a review paper or an innovative empirical study; and (v) certain search terms covered. The exclusion criteria were: (i) editorial materials, (ii) conference proceedings, and (iii) books were removed from the research. The Systematic Review Data Repository (SRDR), a software program for the collection, processing, and inspection of data for our systematic review, was employed. The quality of the specified scholarly sources was evaluated by using the Mixed Method Appraisal Tool. After extracting and analyzing publicly accessible papers as evidence, no institutional ethics approval was required before starting our research (<xref ref-type="fig" rid="F3">Figure 3</xref>).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>PRISMA flow diagram describing the search results and screening (Source: Processed by authors).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1084000-g0003.tif"/>
</fig>
<p>Throughout April 2008 and October 2022 (mostly in 2022), a systematic literature review of the Web of Science, ProQuest, and Scopus databases was performed, with search terms including &#x0201C;multimodal human&#x02013;robot interaction,&#x0201D; &#x0201C;multimodal HRI techniques,&#x0201D; &#x0201C;multimodalities used in HRI,&#x0201D; &#x0201C;speech recognition in HR,&#x0201D; &#x0201C;application of multimodal HRI,&#x0201D; &#x0201C;gesture recognition in HRI,&#x0201D; and &#x0201C;multimodal feedback in HRI.&#x0201D; The search keywords were determined as the most frequently used words or phrases in the researched literature. Because the examined research was published between 2008 and 2022, only 359 publications met the qualifying requirements. We chose 227 primarily empirical sources by excluding ambiguous or controversial findings (insufficient/irrelevant data), outcomes unsubstantiated by replication, excessively broad material, or having nearly identical titles (<xref ref-type="fig" rid="F4">Figure 4</xref>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Topics and types of paper identified and selected.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1084000-g0004.tif"/>
</fig>
</sec>
<sec id="s3">
<title>3. Modalities used in human&#x02013;robot interaction</title>
<p>There are several modalities that are currently used in human&#x02013;robot interaction, including audio, visual, haptic, kinesthetic, and proprioceptive modality (Navarro et al., <xref ref-type="bibr" rid="B98">2015</xref>; Li and Zhang, <xref ref-type="bibr" rid="B80">2017</xref>; Ferlinc et al., <xref ref-type="bibr" rid="B40">2019</xref>; Deuerlein et al., <xref ref-type="bibr" rid="B37">2021</xref>; Groechel et al., <xref ref-type="bibr" rid="B48">2021</xref>). These modalities can be used alone or in combination to enable different forms of human&#x02013;robot interaction, such as voice commands, visual gestures, and physical touch. Additionally, some researchers work on improving the quality of interaction and the perceived &#x0201C;intelligence&#x0201D; of the robot by incorporating tools like natural language processing, cognitive architectures, and social signal processing.</p>
<sec>
<title>3.1. Audio modality</title>
<p>The audio modality is an important aspect of human&#x02013;robot interaction as it allows for verbal communication between humans and robots. In order for robots to effectively understand and respond to human speech, they must be equipped with speech recognition and natural language processing (NLP) capabilities.</p>
<p>Robots that use audio modalities can recognize and generate human speech through the use of speech recognition and synthesis technologies (Lackey et al., <xref ref-type="bibr" rid="B75">2011</xref>; Luo et al., <xref ref-type="bibr" rid="B85">2011</xref>; Zhao et al., <xref ref-type="bibr" rid="B159">2012</xref>; Tsiami et al., <xref ref-type="bibr" rid="B139">2018</xref>; Deuerlein et al., <xref ref-type="bibr" rid="B37">2021</xref>). Speech recognition allows the robot to understand spoken commands or questions from a human, while speech synthesis allows the robot to generate spoken responses or instructions. This modality is used in several application such as voice assistants, voice-controlled robots, and even some language tutor robots (House et al., <xref ref-type="bibr" rid="B60">2009</xref>; Belpaeme et al., <xref ref-type="bibr" rid="B16">2018</xref>; Humphry and Chesher, <xref ref-type="bibr" rid="B62">2021</xref>).</p>
</sec>
<sec>
<title>3.2. Visual modality</title>
<p>Visual modality allows the robot to perceive and interpret visual cues such as facial expressions, gestures, body language, and gaze direction. Robots that employ visual modalities can perceive their environment using cameras and process visual information using computer vision algorithms (Hasanuzzaman et al., <xref ref-type="bibr" rid="B54">2007</xref>; Li and Zhang, <xref ref-type="bibr" rid="B80">2017</xref>). These algorithms can be used to recognize objects, faces, and gestures, as well as to track the motion of humans and other objects. This modality is used in applications such as robot navigation, surveillance, and human&#x02013;robot interaction.</p>
<p>Recent advances in computer vision and deep learning have led to significant improvements in the ability of robots to recognize and interpret visual cues, making them more effective in human&#x02013;robot interactions (Celiktutan et al., <xref ref-type="bibr" rid="B25">2018</xref>). One area of research in visual modality is the use of facial expression recognition, which enables a robot to understand a person&#x00027;s emotions and respond accordingly. This can make the interaction more natural and intuitive for the human. Another area of research is the use of gesture recognition, which allows a robot to understand and respond to human gestures, such as pointing or nodding. This can be useful in tasks such as navigation or object manipulation. In addition, visual saliency detection, which allows the robot to focus on the most important aspects of the visual scene, and object recognition, which enables the robot to identify and locate objects in the environment, are also important areas of research in visual modality.</p>
</sec>
<sec>
<title>3.3. Haptic modality</title>
<p>Haptic modality enables touch-based communication between humans and robots, including the robot&#x00027;s ability to sense and respond to touch and to apply force or vibrations to the human. Recent advances in haptic technology have led to the development of more advanced haptic interfaces, such as force feedback devices and tactile sensors (Navarro et al., <xref ref-type="bibr" rid="B98">2015</xref>; Pyo et al., <xref ref-type="bibr" rid="B105">2021</xref>). These devices allow robots to provide a wider range of haptic cues, which can be used in applications such as robotic surgery, prosthetics, and tactile communication. One area of research in haptic modality is the use of force feedback, which allows a robot to apply forces to a person, making the interaction more natural and intuitive. Another area of research is the use of tactile sensing, which allows a robot to sense the texture, shape, and temperature of objects, and to respond accordingly.</p>
</sec>
<sec>
<title>3.4. Kinesthetic modality</title>
<p>The kinesthetic modality is an aspect of human&#x02013;robot interaction that relates to the ability of the robot to sense and respond to motion and movement. This includes the ability of the robot to sense and respond to the motion of the human body, such as posture, gait, and joint angles. Robots that use kinesthetic modalities can sense and control their own movement. This can be done by using sensors to measure the position and movement of the robot&#x00027;s joints, and actuators to control those joints (Groechel et al., <xref ref-type="bibr" rid="B48">2021</xref>). This modality is used in applications such as industrial robots, bipedal robots, and robots for search and rescue.</p>
</sec>
<sec>
<title>3.5. Proprioceptive modality</title>
<p>Proprioception refers to the ability of an organism to sense the position, orientation, and movement of its own body parts (Ferlinc et al., <xref ref-type="bibr" rid="B40">2019</xref>). In human&#x02013;robot interaction, proprioception can be used to allow robots to sense and respond to the position and movement of their own body parts in relation to the environment and the human. Robots that use proprioceptive modalities can sense their internal state (Hoffman and Breazeal, <xref ref-type="bibr" rid="B57">2008</xref>). This can include, for example, the position of their joints and the forces acting on their body. This information can be used to control the robot&#x00027;s movements, to detect and diagnose failures, and to plan its actions. For example, Malinovsk&#x000E1; et al. (<xref ref-type="bibr" rid="B86">2022</xref>) have developed a neural network model that can learn proprioceptive-tactile representations on a simulated humanoid robot, demonstrating the ability to accurately predict touch and its location from proprioceptive information. However, further work is needed to address the model&#x00027;s limitations.</p>
<p>All these modalities can be combined in different ways to provide robots with a wide range of capabilities and enhance their ability to interact with humans in natural ways. Additionally, for better human robot interaction, using modalities that are congruent with human communication, like visual and auditory modality, are preferred as it makes the interaction more intuitive and easy for human participants.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Techniques for multimodal human&#x02013;robot interaction</title>
<p>In multimodal HRI, social robots frequently use multimodal interaction methods comparable to those utilized by individuals: speech generation (through speakers), voice recognition (through microphones), gesture creation (through physical embodiment), and gesture recognition (via cameras or motion trackers) (Mead and Matari&#x00107;, <xref ref-type="bibr" rid="B92">2017</xref>). This section will provide a brief review on the signal input, signal output, as well as the practical application of multimodal human&#x02013;robot interaction (<xref ref-type="fig" rid="F5">Figure 5</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Multimodal HRI (Castillo et al., <xref ref-type="bibr" rid="B22">2013</xref>; Yu et al., <xref ref-type="bibr" rid="B155">2016</xref>; Lannoy et al., <xref ref-type="bibr" rid="B76">2017</xref>; Legrand et al., <xref ref-type="bibr" rid="B78">2018</xref>; Stephens-Fripp et al., <xref ref-type="bibr" rid="B129">2018</xref>; Mohebbi, <xref ref-type="bibr" rid="B95">2020</xref>; Khalifa et al., <xref ref-type="bibr" rid="B65">2022</xref>; Ot&#x000E1;lora et al., <xref ref-type="bibr" rid="B100">2022</xref>; Strazdas et al., <xref ref-type="bibr" rid="B132">2022</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1084000-g0005.tif"/>
</fig>
<sec>
<title>4.1. Multimodal signal processing for human&#x02013;robot interaction</title>
<p>The last decade has seen great advancement in human&#x02013;robot interaction. Nowadays, with more sophisticated and intelligent sensors, speech, gestures, images, videos, as well as physiological signals like electroencephalography (EEG) and electrocardiogram (ECG), can be input into robots and recognized by them. A brief introduction to the most common input signals employed in HRI will be given in this section.</p>
<sec>
<title>4.1.1. Computer vision</title>
<p>Robots equipped with cameras can recognize and track human faces and movements, allowing them to respond to visual cues and gestures (Andhare and Rawat, <xref ref-type="bibr" rid="B8">2016</xref>; Maroto-G&#x000F3;mez et al., <xref ref-type="bibr" rid="B89">2023</xref>). It is an important modality in multimodal human&#x02013;robot interaction (HRI) as it allows robots to perceive and understand their environment and the actions of humans.</p>
<p>Object detection and tracking: Robots can use computer vision to detect and track objects and people in their environment (Redmon et al., <xref ref-type="bibr" rid="B110">2016</xref>). This can be used for tasks such as following a person, avoiding obstacles, or manipulating objects.</p>
<p>Facial recognition: Robots can use computer vision to recognize and identify specific individuals by analyzing their facial features (Schroff et al., <xref ref-type="bibr" rid="B122">2015</xref>). This can be used for tasks such as personalization, security, or tracking attendance.</p>
<p>Gesture recognition: Robots can use computer vision to recognize and interpret human hand and body gestures (Mitra and Acharya, <xref ref-type="bibr" rid="B93">2007</xref>). This can be used as an additional modality for controlling the robot or issuing commands, rather than using speech or buttons, which is particularly useful in noisy environments.</p>
<p>Face and body language: Robots can use computer vision to detect and interpret facial expressions and body language, which can be used to infer the emotions or intent of a human, and generate appropriate responses, which is known as affective computing (Pantic and Rothkrantz, <xref ref-type="bibr" rid="B102">2000</xref>).</p>
<p>Gaze tracking: Robots can use computer vision to track the gaze of a human in order to understand where their attention is focused (Smith et al., <xref ref-type="bibr" rid="B128">2013</xref>). This can be used to infer human attention and interest or anticipate the next action, for example, a robot assistant can know that the human is going to pick an object by following their gaze.</p>
<p>Multi-Camera system: Using multiple cameras can enable a robot to track and understand the 3D space and provide more robust performance, such as enabling robots to walk without colliding with obstacles (Heikkila, <xref ref-type="bibr" rid="B56">2000</xref>).</p>
<p>These are just a few examples of how computer vision can be used in HRI, and many other applications are being developed and explored in the field. Computer vision systems can be integrated with other modalities, such as speech recognition or haptic feedback, to create a more comprehensive HRI experience.</p>
<p>It&#x00027;s important to note, however, that computer vision can be a challenging technology to implement, especially when it comes to dealing with variations in lighting, occlusion, or viewpoint. It also could be affected by the environment, such as reflections, glare, or shadows, that can make it difficult for the robot to accurately interpret visual data (Tian et al., <xref ref-type="bibr" rid="B137">2020</xref>).</p>
</sec>
<sec>
<title>4.1.2. Natural language processing</title>
<p>Natural language processing (NLP) is an important modality in multimodal human&#x02013;robot interaction (HRI) as it allows robots to understand and respond to human speech in a way that is more natural and intuitive (Scalise et al., <xref ref-type="bibr" rid="B120">2018</xref>). Here are a few ways that NLP can be used in HRI.</p>
<p>Speech recognition: NLP can be used to convert speech into text, which can then be used to interpret human commands, queries, or requests. This allows the robot to understand and respond to spoken commands, such as &#x0201C;turn on the lights&#x0201D; or &#x0201C;navigate to the kitchen.&#x0201D;</p>
<p>Natural Language Understanding (NLU): After the speech is converted into text, the NLP system uses NLU to extract the intent and entities from the text (K&#x000FC;bler et al., <xref ref-type="bibr" rid="B71">2011</xref>; Bastianelli et al., <xref ref-type="bibr" rid="B15">2014</xref>). This allows the robot to understand the intent of the command and the objects or actions referred to by the entities, such as &#x0201C;set the temperature to 20 degrees&#x0201D; intent is &#x0201C;set&#x0201D; and &#x0201C;temperature&#x0201D; is the entities.</p>
<p>Natural Language Generation (NLG): Natural Language Generation (NLG) is a subfield of Natural language processing (NLP) that enables computers to produce natural language responses to humans. NLG has become an important aspect of human&#x02013;robot interaction (HRI) due to its ability to allow robots to communicate with humans in a more human-like manner. However, the process of generating natural language responses involves retrieving and synthesizing relevant information from various data sources, such as open data repositories, domain-specific databases, and knowledge graphs. In addition, the NLG process involves the use of complex algorithms and statistical models to generate natural language responses that are contextually appropriate and grammatically correct.</p>
<p>Question answering: NLU and NLG together allows the robot to understand and generate answers to questions, for example &#x0201C;What&#x00027;s the weather today?&#x0201D;</p>
<p>Dialogue management: NLP can be used to manage the dialogue between the human and the robot, for example to track the state of the conversation, and allow for more seamless interactions, for example by remembering the context of the previous turns in the conversation and using it to generate appropriate responses.</p>
<p>Language translation: NLP can be used to translate text from one language to another, which can enable robots to interact with people who speak different languages.</p>
<p>NLP can be a challenging technology to implement, especially when it comes to handling variations in accent, dialect, or speech patterns, also, NLP models rely heavily on the training data, and thus the performance may not be accurate when it comes to handling new or unseen words, entities or concepts (Khurana et al., <xref ref-type="bibr" rid="B66">2022</xref>).</p>
</sec>
<sec>
<title>4.1.3. Gesture recognition</title>
<p>Gesture identification is a critical step in gesture recognition after unprocessed signals from sensors are obtained. Gesture identification is the discovery of gestural signals in raw data and the separation of the relevant gestural inputs. Popular solutions for solving the issue of gesture recognition are grounded in visual features, ML algorithms, and skeletal models (Mitra and Acharya, <xref ref-type="bibr" rid="B93">2007</xref>; Rautaray and Agrawal, <xref ref-type="bibr" rid="B109">2015</xref>). When it comes to detecting body gestures, the comprehensive representation of the body is ineffective from time to time. In contrast to the preceding methodologies, the skeleton model methodology uses a human skeleton to discover human body positions. The skeletal model technique is also advantageous for categorizing gestures. With the benefits listed above, the skeletal model method has emerged as an appealing solution for sensing devices (Mitra and Acharya, <xref ref-type="bibr" rid="B93">2007</xref>; Cheng et al., <xref ref-type="bibr" rid="B29">2015</xref>).</p>
<p>Among alternative communication modalities for human&#x02013;robot and inter-robot interaction, hand gesture recognition is mostly employed. Hand gesture recognition can be divided into two categories: static hand gestures and dynamic hand gestures. Static hand gestures refer to specific hand postures or shapes that convey meaning without the need for movement. These gestures can be simple, such as a thumbs-up or a peace sign, or more complex, such as those used in sign languages for the deaf community. Static hand gestures have their advantages in specific contexts, such as low computational complexity and less dependency on temporal information.</p>
<p>In comparison to static hand gestures, the dynamic hand gestures of robots are more humanoid. Dynamic hand gestures are particularly versatile since the robotic hand may move in any direction and bend at practically any angle in all available coordinates; static hand gestures, on the other hand, are constrained to much fewer movements (Rautaray and Agrawal, <xref ref-type="bibr" rid="B109">2015</xref>). A wide range of applications, involving smart homes, video surveillance, sign language recognition, human&#x02013;robot interaction, and health care, have recently embraced dynamic hand gestures. All of these applications require high levels of accuracy against a busy background, optimum recognition, and temporal precision (Huenerfauth and Lu, <xref ref-type="bibr" rid="B61">2014</xref>; Ur Rehman et al., <xref ref-type="bibr" rid="B142">2022</xref>).</p>
</sec>
<sec>
<title>4.1.4. Emotion recognition</title>
<p>Emotions are inherent human characteristics that impact choices along with behaviors, and they are crucial in interaction and emotional intelligence (Salovey and Mayer, <xref ref-type="bibr" rid="B117">2004</xref>), i.e., the capacity to comprehend, utilize, and command feelings, is substantial for effective relationships. Affective computation seeks to provide robots with emotional intelligence, aiming to improve natural human&#x02013;robot interaction. Humanoid competencies of observation, comprehension, and feeling output are sought in the context of human&#x02013;robot interaction. Emotions in HRI can be examined from three distinct perspectives, as follows.</p>
<p>Formalization of the robot&#x00027;s internal psychological state: Adding sentimental characteristics to individuals and bots can increase their efficacy, adaptability, and plausibility. In recent years, robots have been produced to mimic feelings by determining neurocomputational frameworks, formalizing them in pre-existing cognitive architectures, modifying well-known mental representations, or developing specific affective designs (Saunderson and Nejat, <xref ref-type="bibr" rid="B119">2019</xref>).</p>
<p>The emotional response of robots: The capacity of robots to display recognizable emotional responses has a significant influence on human&#x02013;robot interaction in complicated communication scenarios (Rossi et al., <xref ref-type="bibr" rid="B114">2020</xref>). Numerous research examined how individuals perceive and identify sentimental reactions through modalities (postures, expressions, motions, and voices) might transmit emotional signals from robots to humans.</p>
<p>Robotic applications that can detect and comprehend human feelings are competent in social interactions. Recent research aims to develop algorithms for categorizing psychological states from many input signals, including speech, body language, expression, and physiological signals (Cavallo et al., <xref ref-type="bibr" rid="B23">2018</xref>).</p>
<p>Furthermore, sentiment identification is a multidisciplinary area that necessitates expertise from a variety of disciplines, including psychological science, neurology, data processing, electronics, and AI. It may be handled using multimodal signals, including physiological signals like EEG, GSR, or heart rate fluctuations measured by BVP or EKG. As with BVP and GSR, these are inner signals that represent the equilibrium of the parasympathetic and sympathetic nervous systems, whereas EEG shows variations in the cortex parts of the brain (Das et al., <xref ref-type="bibr" rid="B34">2016</xref>). Externally visible indications, on the other hand, include facial expressions, bodily motions, and voice. While internal signals are thought to be more impartial due to the inherent qualities of several operational parts of the central nervous system, external signals remain subjective measures of expressed feelings (Yao, <xref ref-type="bibr" rid="B152">2016</xref>).</p>
</sec>
</sec>
<sec>
<title>4.2. Multimodal feedback for human&#x02013;robot interaction</title>
<p>Developing a socially competent robot capable of interacting naturally with individuals and synthesizing adequately intelligible multimodal actions in a wide range of interaction scenarios is a difficult task. This necessitates a high degree of multimodal perception of robots, since they must comprehend the human&#x00027;s mental moods, goals, and character aspects in order to provide proper feedback.</p>
<sec>
<title>4.2.1. Speech synthesis</title>
<p>Speech synthesis, also known as text-to-speech (TTS), is an important approach in multimodal human&#x02013;robot interaction (HRI) as it allows robots to provide verbal feedback or instructions to the human user in a way that is similar to how a human would (Luo et al., <xref ref-type="bibr" rid="B85">2011</xref>; Ashok et al., <xref ref-type="bibr" rid="B11">2022</xref>). Robots can use text-to-speech (TTS) technology to generate spoken responses to humans. Here is how speech synthesis works in more detail:</p>
<list list-type="order">
<list-item><p>Text is generated by the robot&#x00027;s onboard computer in response to a user request, or based on data the robot needs to communicate. However, the generation of text to be uttered (Natural Language Generation) is a research field. On a superficial level, responses can either be template-based (i.e., scripted by humans), retrieved from knowledge sources (typically, the Internet) or generated using large-language models.</p></list-item>
<list-item><p>Once the text has been generated, it is passed to a Text-to-Speech (TTS) engine, which uses a set of rules, or a machine learning model, to convert the text into speech. This process involves transforming the written text into a phonetic representation that can be pronounced by the robot.</p></list-item>
<list-item><p>The TTS engine can be tuned to mimic different human voices, genders, and even create a virtual robot voice. For example, many systems make use of widely available TTS engines (e.g., Acapela, Cereproc, Google TTS), which offer a range of voices in different languages and accents.</p></list-item>
<list-item><p>The generated speech is then output through the robot&#x00027;s speakers, allowing the user to hear the response. The quality of the output speech is dependent on the TTS engine, the quality of the audio hardware, and the environmental conditions in which the robot is operating.</p></list-item>
<list-item><p>The output can be in different languages, depending on the specific application and the need. For instance, some robots may be designed to operate in multilingual environments and require the ability to speak multiple languages to communicate effectively with users.</p></list-item>
</list>
<p>Speech synthesis can be integrated with other modalities, such as computer vision, or natural language processing (NLP), to create a more comprehensive HRI experience. For example, a robot that uses speech recognition and NLP to understand spoken commands can use speech synthesis to provide verbal feedback, such as &#x0201C;I&#x00027;m sorry, I didn&#x00027;t understand that command.&#x0201D;</p>
<p>Speech synthesis can also be used to provide instructions, such as &#x0201C;Please put the object on the tray,&#x0201D; or to answer questions, such as &#x0201C;The current temperature is 20 degrees.&#x0201D;</p>
<p>Speech synthesis can enhance the user experience by making the interaction with the robot more natural, intuitive, and engaging. It can also be used to provide information or instructions in a variety of languages, making the robot accessible to a wider range of users.</p>
</sec>
<sec>
<title>4.2.2. Visual feedback</title>
<p>Visual feedback is an important output modality in multimodal human&#x02013;robot interaction (HRI) as it allows robots to provide feedback to the human user through visual cues (Gams and Ude, <xref ref-type="bibr" rid="B42">2016</xref>; Yoon et al., <xref ref-type="bibr" rid="B154">2017</xref>). Here are a few ways that visual feedback can be used in HRI.</p>
<p>Status indication: Robots can use lights, displays, or other visual cues to indicate the status of the robot, such as when the robot is ready to receive commands, when it is performing a task, or when it has completed a task (Admoni and Scassellati, <xref ref-type="bibr" rid="B1">2016</xref>).</p>
<p>Error indication: Robots can use visual cues such as flashing lights or error messages to indicate an error or problem with the robot, for example when the robot can&#x00027;t complete a task due to an obstacle or error (Kim et al., <xref ref-type="bibr" rid="B67">2016</xref>).</p>
<p>Wayfinding: Robots can use visual cues such as arrows or maps, to indicate a path or a location, this can help the user to navigate and orient themselves in the environment (Giudice and Legge, <xref ref-type="bibr" rid="B45">2008</xref>).</p>
<p>Object recognition and tracking: Robots can use visual cues such as highlighted boxes, to indicate the objects or areas of interest the robot is tracking or recognizing (Cazzato et al., <xref ref-type="bibr" rid="B24">2020</xref>).</p>
<p>Expressions and emotions: Robots can use visual cues such as facial expressions or body language, to indicate the robot&#x00027;s emotions or intent, similar to the way humans communicate non-verbally (Al-Nafjan et al., <xref ref-type="bibr" rid="B4">2017</xref>).</p>
<p>Multi-Camera Systems: Robots can use multiple cameras to provide visual feedback, by showing multiple views of the environment, or provide 3D information, which can help the user to understand the robot&#x00027;s perception of the environment (Feng et al., <xref ref-type="bibr" rid="B39">2017</xref>).</p>
<p>Displaying conversation: In addition to the above, visual feedback can also be used to display the content of the conversation in HRI, such as what the robot &#x0201C;hears&#x0201D; through automatic speech recognition (ASR) and what it is saying. This can be particularly useful for individuals with hearing impairments or in noisy environments where auditory feedback may not be sufficient (Rasouli et al., <xref ref-type="bibr" rid="B108">2018</xref>).</p>
</sec>
<sec>
<title>4.2.3. Visual feedback</title>
<p>Visual feedback is an important output modality in multimodal human&#x02013;robot interaction (HRI) as it allows robots to provide feedback to the human user through visual cues (Gams and Ude, <xref ref-type="bibr" rid="B42">2016</xref>; Yoon et al., <xref ref-type="bibr" rid="B154">2017</xref>). Here are a few ways that visual feedback can be used in HRI.</p>
<p>Status indication: Robots can use lights, displays, or other visual cues to indicate the status of the robot, such as when the robot is ready to receive commands, when it is performing a task, or when it has completed a task.</p>
<p>Error indication: Robots can use visual cues such as flashing lights or error messages to indicate an error or problem with the robot, for example when the robot can&#x00027;t complete a task due to an obstacle or error.</p>
<p>Wayfinding: Robots can use visual cues such as arrows or maps, to indicate a path or a location, this can help the user to navigate and orient themselves in the environment.</p>
<p>Object recognition and tracking: Robots can use visual cues such as highlighted boxes, to indicate the objects or areas of interest the robot is tracking or recognizing.</p>
<p>Expressions and emotions: Robots can use visual cues such as facial expressions or body language, to indicate the robot&#x00027;s emotions or intent, similar to the way humans communicate non-verbally.</p>
<p>Multi-Camera Systems: Robots can use multiple cameras to provide visual feedback, by showing multiple views of the environment, or provide 3D information, which can help the user to understand the robot&#x00027;s perception of the environment.</p>
<p>Displaying conversation: In addition to the above, visual feedback can also be used to display the content of the conversation in HRI, such as what the robot &#x0201C;hears&#x0201D; through ASR (automatic speech recognition) and what it is saying. This can be particularly useful for individuals with hearing impairments or in noisy environments where auditory feedback may not be sufficient.</p>
</sec>
<sec>
<title>4.2.4. Gesture generation</title>
<p>In general, gesture generation is an area that remains largely underdeveloped in robotics research, with most of the focus being on gesture recognition. In conventional robotics, recognition always predominates over gesture synthesis. The term &#x0201C;gesture&#x0201D; has been commonly utilized to refer to item manipulation tasks instead of non-verbal expressive behaviors among the few extant systems that are really devoted to gesture synthesizing. Computational techniques to synthesize multimodal action may be divided into three steps: identifying what to express, deciding how to transmit it, and lastly, acting on it (Covington, <xref ref-type="bibr" rid="B33">2001</xref>). Although the Articulated Communicator Engine acts at the behavioral realization layer, the entire system employed by the digital assistant Max consists of a combined content and behavioral planning architecture (Kopp et al., <xref ref-type="bibr" rid="B69">2008</xref>).</p>
<p>Utilizing multimodal Utterance Representation Markup Language, gesture expressions inside the Articulated Communicator Engine (ACE) framework may be defined in two distinct ways (Salem et al., <xref ref-type="bibr" rid="B116">2012</xref>). A gesture&#x00027;s exterior representation, such as the posture of the gesture stroke, can be clearly articulated in verbal words and co-verbal gestures, which are classified as feature-based explanations (Gozzi et al., <xref ref-type="bibr" rid="B47">2022</xref>). By correlating temporal markers, gesture association to certain language pieces is discovered. Secondly, gestures may be defined as keyframe animation, where each keyframe defines a &#x0201C;key posture,&#x0201D; a component of the general gesture motion that describes the condition of each joint at that particular moment. Assigned time IDs are used to gather speed data for the interpolation between every two key postures and the related association to portions of speech. In ACE, keyframe animations may be created manually or via motion-capture data from a human presenter, enabling real-time animation of virtual agents. Each pitch phrase and co-expressive gesture expression in a multimodal utterance reflects a single thought unit, often known as a chunk of speech-gesture production (Kopp and Wachsmuth, <xref ref-type="bibr" rid="B70">2004</xref>).</p>
<p>The ACE engine uses the following timing for gestures online: The basic way to establish synchronicity within a chunk is to modify the gesture to match the pace and structure of speech. To this end, the ACE scheduler gets millisecond-level scheduling details about the synthesized voice and uses those details to determine the beginning and end of the gesture stroke. Each individual gesture component receives an automated propagation of these timing limitations (Kopp and Wachsmuth, <xref ref-type="bibr" rid="B70">2004</xref>). Chae et al. (<xref ref-type="bibr" rid="B26">2022</xref>) developed a methodology that enables robots to generate co-speech gestures automatically, based on a morphemic analysis of the sentence of utterance. After determining the expression unit and the corresponding gesture type, a database of motion primitives is used to retrieve an appropriate gesture that conveys the robot&#x00027;s thoughts and feelings. The method showed promising results, with 83% accuracy in determining expression units and gesture types, and positive feedback from a user study with a humanoid robot.</p>
</sec>
<sec>
<title>4.2.5. Emotional expression generation</title>
<p>In-home robot and service robot has received much attention recently, and the demand for service robots is expected to expand rapidly in the coming years. Human-centerd operations are among the most intriguing aspects of smart service robots. Smart interaction is an important characteristic of service robots in care services, companionship, and entertainment. In real-world settings, emotional intelligence will be critical for a robot to participate in an amicable conversation. Furthermore, there has been a surge in interest in researching robotic mood-generating methods that aim to offer a robot more human-like behavioral patterns.</p>
<p>Previous research in this field demonstrates a number of effective techniques for creating emotional robots. It has been found that a smooth transition between mood states is crucial for the development of robotic emotions (Stock-Homburg, <xref ref-type="bibr" rid="B131">2022</xref>). The engagement activity of the robots and the user&#x00027;s perception of the robot are both directly influenced by the robot&#x00027;s emotional shift from one mood to another. The empathy of a robot must still be shown through responsive interaction actions. A fixed one-to-one link between the emotional state of a robot and its response is inappropriate. The shift between mood states would be more intriguing and realistic if the robot&#x00027;s expression remained constant. In order to create truly sociable robots, Rincon et al. (<xref ref-type="bibr" rid="B112">2019</xref>) developed a social robot that aims to assist older people in their daily activities while also being able to perceive and display emotions in a human-like way. The robot is currently being tested in a daycare center in the northern region of Portugal. Shao et al. (<xref ref-type="bibr" rid="B124">2020</xref>) proposed a novel affect elicitation and detection method for social robots in HRIs, which used non-verbal emotional behaviors of the robot to elicit user affect and directly measure it through EEG signals. The study conducted experiments with younger and older adults to evaluate the affect elicitation technique and compare two affect detection models utilizing multilayer perceptron neural networks (NNs) and support vector machines (SVMs).</p>
<p>Rather than being established randomly, the correlations between the emotive response of a robot and its emotional state may be modeled from emotional analysis and used to develop patterns of interaction in the creation of communicative behaviors (Han M. J. et al., <xref ref-type="bibr" rid="B52">2012</xref>).</p>
</sec>
<sec>
<title>4.2.6. Multi-modal feedback</title>
<p>Robots can use a combination of multiple modalities to provide feedback, for example, using speech synthesis and visual feedback to indicate status. Multi-modal feedback is a key aspect of multimodal human&#x02013;robot interaction (HRI), as it allows robots to convey information or commands to the human user through multiple modalities simultaneously (Andronas et al., <xref ref-type="bibr" rid="B9">2021</xref>). This can provide a more comprehensive and engaging user experience. Here are a few ways that multi-modal feedback can be used in HRI.</p>
<p>Multi-modal status indication: Robots can use a combination of multiple modalities such as audio cues and visual cues, to indicate the status of the robot, such as a beep sound and a flashing light when the robot is ready to receive commands, and a different sound and light when it has completed a task.</p>
<p>Multi-modal error indication: Robots can use a combination of multiple modalities, such as a warning tone and a flashing light, to indicate an error or problem with the robot.</p>
<p>Multi-modal cues and prompts: Robots can use a combination of modalities such as speech synthesis, visual cues and sound to prompt the user to perform a specific action, this can make the instruction clear and easy to follow. For example, a robot assistant in a factory might use a combination of a flashing light and speech synthesis to prompt the user to perform a specific task (Cherubini et al., <xref ref-type="bibr" rid="B30">2019</xref>).</p>
<p>Multi-modal social presence: Robots can use a combination of modalities such as speech synthesis, facial expressions and sound effects to create a sense of social presence and make the robot more relatable and human-like.</p>
<p>Multi-modal information: Robots can use a combination of modalities such as speech synthesis, visual cues, and haptic feedback to convey information, this can make the information more intuitive and easy to understand. For example, a robot designed to provide directions might use a combination of speech synthesis and visual cues to display a map and provide turn-by-turn directions.</p>
<p>Multi-modal dialogue management: Robots can use a combination of modalities such as speech recognition, computer vision, and haptic feedback to manage the dialogue between the human and the robot, this can allow for more seamless and natural interactions. For example, a robot assistant in a hospital might use speech recognition to understand the user&#x00027;s request, computer vision to locate the necessary supplies, and haptic feedback to alert the user when the supplies have been retrieved (Ahn et al., <xref ref-type="bibr" rid="B2">2019</xref>).</p>
</sec>
</sec>
<sec>
<title>4.3. Application of multimodal HRI</title>
<p>With the fast advancement in sensors and HRI innovations, numerous applications of multimodal HRI have sprung up in recent years. In this section, four major applications will be briefly introduced, that is, industrial robots, assistive mobile robots, robotic exoskeletons, as well as robotic prothesis, as shown in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Major applications of multimodal human&#x02013;robot interaction (Hahne et al., <xref ref-type="bibr" rid="B50">2020</xref>; Kavalieros et al., <xref ref-type="bibr" rid="B64">2022</xref>; Mocan et al., <xref ref-type="bibr" rid="B94">2022</xref>; Pawu&#x0015B; and Paszkiel, <xref ref-type="bibr" rid="B103">2022</xref>; Sasaki et al., <xref ref-type="bibr" rid="B118">2022</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1084000-g0006.tif"/>
</fig>
<sec>
<title>4.3.1. Industrial robots</title>
<p>In the last few years, there has been a huge leap in the productivity and marketability of industrial robots, and the use of industrial cobots has significantly aided in the growth of the industry. Industrial cobots, or collaborative robots, are designed to work alongside human operators in various tasks and environments. As a result of the development and popularity of Industry 4.0, industrial cobots are now expected to be increasingly independent and smart to complete more complicated and flexible jobs. Industrial robot growth is dependent on the development of several technologies, of which sensing technologies are a crucial component. Sensors can be employed to gather a wealth of data to assist industrial cobots in carrying out their duties, that is to say, industrial cobots need sensors to carry out their functions.</p>
<p>There are four kinds of sensors used on industrial robots: visual sensors, tactile sensors, laser sensors, and encoders (Li and Liu, <xref ref-type="bibr" rid="B79">2019</xref>). Apart from the four types of sensors, other sensors used in industrial cobots to perform various activities include proximity sensors, ultrasonic sensors, torque sensors, inertial sensors, acoustic sensors, magnetic sensors, and so on.</p>
<p>Sensors are widely employed to aid users in controlling industrial cobots to perform assigned activities that include human&#x02013;robot collaboration (HRC), adaptive cruise control, manipulator control, and so on. The notion of HRC was recently introduced to actualize the joint operation of employees and robots. This form of production can increase the flexibility and agility of manufacturing systems by combining human cognition and strain capacity with the precision and tirelessness of robots. The fundamental issue that HRC must address is safety. Robotic machines should be capable of detecting and recognizing things in order to prevent conflicts or to stop movement instantly in the event of a collision. Vision sensors, proximity sensors, laser sensors, torque sensors, and tactile sensors are popular sensors used to execute this function. For example, in this paper (O&#x00027;Neill et al., <xref ref-type="bibr" rid="B99">2015</xref>; Fritzsche et al., <xref ref-type="bibr" rid="B41">2016</xref>), tactile sensors are used to detect physical touch and pinpoint the location of accidents in order to protect personnel who are collaborating with the robot. In Popov et al. (<xref ref-type="bibr" rid="B104">2017</xref>), inner joint torque sensors are used to identify and categorize collisions by calculating external forces.</p>
<p>Interaction between robots and humans can be crucial in human&#x02013;robot collaboration. Workers can successfully control computer programming via HRI. In Kurian (<xref ref-type="bibr" rid="B73">2014</xref>), for example, voice recognition supported by acoustical sensors is utilized to assist people in interacting with robots. However, if the surrounding environment is noisy, such a method may not work well. To address this issue, hand gesture identification using vision sensors has been presented in Tang et al. (<xref ref-type="bibr" rid="B134">2015</xref>). The integration of multiple sensor types and technologies allows cobots to adapt to various working conditions and enhances the efficiency and safety of human&#x02013;robot collaboration in industrial settings.</p>
</sec>
<sec>
<title>4.3.2. Assistive mobile robots</title>
<p>For more than two decades, researchers have been adapting mobile robotic principles to assistance devices, which corresponds to two key applications: intelligent wheelchairs and assistive walkers. The smart wheelchair is among the most commonly used assistive equipment, with an estimated user base of 65 million globally. Wheelchairs can be either manual or power-driven. Many wheelchair users find it difficult to utilize their wheelchairs autonomously due to the lack of skills, muscles, or vision. An intelligent wheelchair is simply a motorized wheelchair outfitted with sensors plus digital control systems.</p>
<p>The advancement of navigation algorithms for obstacle detection, automated user transportation, and aided steering of the wheel through cutting-edge human&#x02013;robot interfaces are all outcomes of studies on intelligent wheelchairs. The idea that the regularly used joysticks are not always helpful, especially for users with a low degree of neuro-muscular competence, is the primary driving force behind these studies. Smart wheelchairs include a variety of sensors, including cameras, infrared, lasers, and ultrasonic (Desai et al., <xref ref-type="bibr" rid="B36">2017</xref>). Modern technologies have increasingly been incorporated into user interfaces in an attempt to enhance user independence or entirely automate the product&#x00027;s navigation.</p>
<p>Touch screens, voice recognition systems, and aided joysticks are employed to transmit the locations or routes to the robotic machine (Schwesinger et al., <xref ref-type="bibr" rid="B123">2017</xref>). Some research projects focused on creating frameworks that let users or the controller receive force input from the surroundings using haptic interface for individuals with vision problems (Chuy et al., <xref ref-type="bibr" rid="B31">2019</xref>). Other more recent methods used speech commands and audible feedback to communicate choices to the operator, such as how to navigate around obstacles, safely approach items, and reach items from a certain angle (Sharifuddin et al., <xref ref-type="bibr" rid="B125">2019</xref>). Many advancements have been made to convert users&#x00027; eye, facial, and body motions into orders for the wheelchair using the visual output for sufferers who are unable to handle a normal joystick (Rabhi et al., <xref ref-type="bibr" rid="B106">2018a</xref>,<xref ref-type="bibr" rid="B107">b</xref>). The same method of classifying and recognizing gestures is used to collect surface EMG (Kumar et al., <xref ref-type="bibr" rid="B72">2019</xref>) and EEG signals (Zgallai et al., <xref ref-type="bibr" rid="B157">2019</xref>). Some research projects that employ many input sources, including biosignals and feedback sensors, in tandem to conduct aided navigation also examine multimodal sensory integration (Reis et al., <xref ref-type="bibr" rid="B111">2015</xref>). For aided navigation and steering of walkers, ML algorithms are combined with sensors and bio-signal gathering frameworks, comparable to the HRI techniques for intelligent wheelchairs (Alves et al., <xref ref-type="bibr" rid="B5">2016</xref>; Caetano et al., <xref ref-type="bibr" rid="B21">2016</xref>; Wachaja et al., <xref ref-type="bibr" rid="B143">2017</xref>). These intelligent walkers may be used by individuals with visual impairments for secure outdoor and indoor activity.</p>
</sec>
<sec>
<title>4.3.3. Robotic exoskeletons</title>
<p>A robotic exoskeleton serves as an active orthosis device that should be transportable in everyday life situations to assist patients with movement and control limitations (der Loos et al., <xref ref-type="bibr" rid="B35">2016</xref>). Furthermore, robotic exoskeletons might be a feasible option for industrial cargo bearing, as industrial personnel do repetitive physical duties exposing them to musculoskeletal problems (Treussart et al., <xref ref-type="bibr" rid="B138">2020</xref>). An essential component in the control of robotic exoskeletons is the acquisition and identification of human intent, which is carried out by different means of human&#x02013;robot interaction and acts as an input to the control system.</p>
<p>Cognitive human&#x02013;robot interaction (cHRI) utilizes EEG signals from the central nervous system to the musculoskeletal system, or surface EMG signals, to recognize the client&#x00027;s needs prior to any real body movements and then estimate the appropriate torque or positional inputs. When compared with the lower-limb exoskeletons, research activities on upper-limb exoskeletons concentrate on developing interface and decoding methodologies to enable accuracy and agility in a larger range of motions. ML techniques are very effective for recognizing the user&#x00027;s mobility intentions grounded in categorized biological signals and may be used to operate such equipment in live time (Nagahanumaiah, <xref ref-type="bibr" rid="B97">2022</xref>).</p>
<p>Physical human&#x02013;robot interaction (pHRI) employs force measures or alterations in joint locations caused by musculoskeletal system movement as control inputs to the robotic exoskeleton. In such circumstances, the robot&#x00027;s controller seeks to minimize the effort required to complete the tasks, resulting in compliant action. To be more precise, in HRI, minimal contact pressures are preferred and task-tracking mistakes should be avoided. To achieve this, interacting forces are usually controlled using resistance or admission controllers that employ a virtual impedance term to simulate HRI, as described by Hogan (<xref ref-type="bibr" rid="B58">1984</xref>).</p>
</sec>
<sec>
<title>4.3.4. Robotic prothesis</title>
<p>A robotic prosthetic limb is a robot that is linked to a sufferer&#x00027;s body and replicates its capabilities in everyday routines (Lawson et al., <xref ref-type="bibr" rid="B77">2014</xref>). Robotic prostheses come into direct contact with the body since their functions are often controlled and directed in real-time by clients by muscular or cerebral impulses. Physical specifications, an anthropomorphic appearance, deciphering user&#x00027;s intention, and replicating movements, force efforts, or grip shapes of the actual body are all key characteristics to consider while constructing a robotic prosthesis. Many recent studies have concentrated on producing prosthetics that more nearly resemble the abilities of a lost organic limb (Masteller et al., <xref ref-type="bibr" rid="B90">2021</xref>). The key to effective growth is to acquire a precise technique of recognizing the client&#x00027;s needs, ultimately, perceiving the surroundings and transforming that need into action. The mobility and dynamic control systems of the robot may capture a mixture of bio-signals to construct an identification scheme and actualize the user intention. These data are biometric records of residual limb muscular electrical activity, brain function, or contact stress in sockets. The prosthesis&#x00027;s control system gets complicated input patterns from the client and makes real-time motor control decisions based on the learnt forecast of the user&#x00027;s purpose. Pattern recognition technologies applied to myoelectric or other biosignals are used to identify user-intentioned behaviors. Typically, a classifier is taught to distinguish various robotic prosthesis joint actuators using patterns from multi-channel EMG data.</p>
<p>In order to strengthen the resilience of the activities performed by the equipment, the study on this topic is primarily focused on enhancing the understanding of myoelectric patterns and the concurrent pattern identification and management of numerous functions. Essentially, assignments from everyday routines have many degrees of freedom to move simultaneously. Hence, integrated joint movements must be categorized differently. Deep learning (DL) techniques have recently been used as a novel tool to conduct classification and regression tasks straight from high-dimensional raw EMG signals without locating and recognizing any signal characteristics (Ameri et al., <xref ref-type="bibr" rid="B7">2018</xref>).</p>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5. Recent advancements of application for multi-modal human&#x02013;robot interaction</title>
<p>Multimodal HRI has advanced greatly in the past decade, with numerous research progress and applications coming into existence each year. In this section, we systematically review the state of the art of multimodal HRI, and thoroughly comb the research progress in terms of the signal input, the signal output, as well as the applications of multimodal HRI by listing and summarizing relevant articles.</p>
<sec>
<title>5.1. Multimodal input for human&#x02013;robot interaction</title>
<p>This section presents a systematic literature review summarizing the latest research progress in signal input for multimodal human&#x02013;robot interaction (HRI).</p>
<sec>
<title>5.1.1. Intuitive user interfaces and multimodal interfaces</title>
<p>In recent years, numerous studies have been conducted to improve signal input in HRI. Salem et al. (<xref ref-type="bibr" rid="B115">2010</xref>) describe a method for enabling the bipedal robot ASIMO to generate voice and co-verbal gestures freely at runtime without being constrained to a predefined repertoire of motor motions. Berg and Lu (<xref ref-type="bibr" rid="B17">2020</xref>) summarize research methodologies on HRI in service and industrial robotics, emphasizing that advancements in human&#x02013;robot interfaces have brought us closer to intuitive user interfaces, particularly when adopting multimodal interfaces that include voice and gesture detection.</p>
</sec>
<sec>
<title>5.1.2. Audiovisual UI and multimodal feeling identification</title>
<p>Ince et al. (<xref ref-type="bibr" rid="B63">2021</xref>) present research on an audiovisual UI-based drumming platform for multimodal HRI, creating an audiovisual communicative interface by combining communicative multimodal drumming with humanoid robots. Chen et al. (<xref ref-type="bibr" rid="B28">2021</xref>) explore multimodal emotion identification and intent understanding, presenting various modalities of mental feature extraction and emotion recognition methods, and applying them in practice to achieve HRI.</p>
</sec>
<sec>
<title>5.1.3. Natural interaction framework and seamless communication</title>
<p>Andronas et al. aim to develop and implement a natural interaction framework for human-system and system-human communication, allowing seamless communication between controllers and &#x0201C;robot companions.&#x0201D; An automobile sector scenario evaluates the framework&#x00027;s performance, showing how an intuitive interface framework can enhance the effectiveness of both humans and robots (Andronas et al., <xref ref-type="bibr" rid="B9">2021</xref>).</p>
</sec>
<sec>
<title>5.1.4. Nonverbal communication, locomotion training, emotional messaging, and multimodal robotic UI</title>
<p>Various techniques and systems have been explored to improve human&#x02013;robot interaction through nonverbal communication, locomotion training, emotional messaging, and multimodal robotic UI. Han J. et al. (<xref ref-type="bibr" rid="B51">2012</xref>) introduce a novel method to investigate the application of nonverbal signals in HRI using the Nao system, which includes an array of sensors, controllers, and interfaces. The findings suggest that individuals are more inclined to interact with a robot that can understand and communicate through nonverbal channels.</p>
</sec>
<sec>
<title>5.1.5. Active engagement and multimodal HRI solutions</title>
<p>Gui et al. develops a locomotion trainer with multiple walking patterns that can be regulated by participants&#x00027; active movement intent. A multimodal HRI solution, including cHRI and pHRI, is designed to enhance subjects&#x00027; active engagement during therapy (Gui et al., <xref ref-type="bibr" rid="B49">2017</xref>). Additionally, a MEC-HRI system featuring various emotional messaging channels, such as voice, gesture, and expression, is presented. The robots in the MEC-HRI platform can understand human emotions and respond accordingly (Liu et al., <xref ref-type="bibr" rid="B83">2016</xref>).</p>
</sec>
<sec>
<title>5.1.6. Spatial language and multimodal robotic UI</title>
<p>Research on robot spatial relations employs a multimodal robotic UI. They demonstrate how to extract other geographical information, such as linguistic geographical descriptions, from the evidence grid map. Examples of spatial language are provided for both human-to-robot input and robot-to-human output (Skubic et al., <xref ref-type="bibr" rid="B127">2004</xref>). It can be said without doubt that signal input is an indispensable part of human&#x02013;robot interaction.</p>
</sec>
<sec>
<title>5.1.7. Prosody cues, tactile communication, and proxemics computational method</title>
<p>Recently, significant progress has been made in the field of multimodal HRI, with signal recognition becoming a hot topic. Aly and Tapus investigate the relationship between nonverbal and para-verbal interaction by connecting prosody cues to arm motions. Their method for synthesizing arm gestures employs coupled hidden Markov models, which can be thought of as a cluster of HMMs representing the streams of divided prosodic qualities and segmented rotational features of the two arms&#x00027; expressions (Aly and Tapus, <xref ref-type="bibr" rid="B6">2012</xref>). Tactile communication might be used in multimodal communication networks for HRI. Two studies were carried out to evaluate the viability of employing a vocabulary of standard tactons within a phrase for robot-to-human interaction in tactile speech (Barber et al., <xref ref-type="bibr" rid="B13">2015</xref>).</p>
</sec>
<sec>
<title>5.1.8. Collaborative data extraction and computational framework of proxemics</title>
<p>Whitney et al. (<xref ref-type="bibr" rid="B148">2017</xref>) offer a model that relies on agent collaboration to achieve richer data extraction from observations. This paper proposes a mathematical formulation for an item-fetching area that enables a robot to improve the speed and precision with which it interprets a person&#x00027;s demands by speculating about its own ambiguity and processing implicit messages. Mead and Mataric (<xref ref-type="bibr" rid="B91">2012</xref>) present a computational framework of proxemics based on data-driven probabilistic models of how social signals (speech and gestures) are produced by a human and perceived by a robot. The framework and models were implemented as autonomous proxemic behavior systems for sociable robots.</p>
</sec>
<sec>
<title>5.1.9. Latest advancements in input signals for multimodal HRI</title>
<p>The following studies highlight the latest advancements in input signals used to improve multimodal human&#x02013;robot interaction. Alghowinem et al. (<xref ref-type="bibr" rid="B3">2021</xref>) provide a proxemics computational method based on info-driven probabilistic models of how humans make and robots receive social signals, including gestures and speech. The structure and models were applied as sociable robots&#x00027; autonomous proxemic behavior systems. Tuli et al. (<xref ref-type="bibr" rid="B140">2021</xref>) provide a notion for semantic visualization of human activities and intent forecast in a domain knowledge-based semantic info hub utilizing a flexible task ontology interface. Bolotnikova (<xref ref-type="bibr" rid="B19">2021</xref>) study the subject of whole-body anthropomorphic robot postural planning in the setting of assistive physical HRI. They extend the non-linear optimization-based stance-generating system with the elements required to design a robot stance in communication with a human point cloud.</p>
</sec>
<sec>
<title>5.1.10. Multimodal features and routine development</title>
<p>Barricelli et al. (<xref ref-type="bibr" rid="B14">2022</xref>) provide a fresh method for routine development that takes advantage of Amazon Alexa on Echo Show devices&#x00027; multimodal characteristics (sight, voice, and touch). It then shows how the suggested technique makes it easier for end users to construct routines than the traditional engagement with the Alexa app.</p>
<p>The articles listed above have summarized the latest application of input signals used in multimodal HRI.</p>
</sec>
</sec>
<sec>
<title>5.2. Multimodal output for human&#x02013;robot interaction</title>
<p>This section follows the recent research development in multimodal HRI signal input by presenting a comprehensive literature review and identifying related studies in this field.</p>
<p>As robots become increasingly sophisticated, they are capable of generating multimodal signals, drawing the attention of researchers in the field. Gao et al. (<xref ref-type="bibr" rid="B43">2021</xref>) present a strategy based on multimodal information fusion and multiscale parallel CNN to increase the precision and validity of hand gesture identification. Yongda et al. (<xref ref-type="bibr" rid="B153">2018</xref>) describe a multimodal HRI method that combines voice and gesture, creating a robot control system that converts human speech and gestures into instructions for the robot to perform. Li et al. (<xref ref-type="bibr" rid="B81">2022</xref>) present a unique Multimodal Perception Tracker for monitoring speakers using both auditory and visual modalities, leveraging a lens model to map sound signals to a localization space congruent with visual information.</p>
<p>The latest developments in multimodal output for multimodal human&#x02013;robot interaction include unique emotion identification systems, multimodal conversation handling, and voice and gesture recognition systems for natural interaction with humans. Cid et al. (<xref ref-type="bibr" rid="B32">2015</xref>) offer a unique multimodal emotion identification system that relies on visual and aural input processing to assess five different affective states. Stiefelhagen et al. (<xref ref-type="bibr" rid="B130">2007</xref>) present systems for recognizing utterances, multimodal conversation handling, and visual processing of a user, including localization, tracking, and recognition of the user, identification of pointing gestures, and recognition of a person&#x00027;s head orientation. Zlatintsi et al. (<xref ref-type="bibr" rid="B160">2018</xref>) investigate new aspects of smart HRI by automatically recognizing and validating voice and gestures in a natural interface, providing a thorough structure and resources for a real-world scenario with elderly individuals assisted by an assistive bath robot.</p>
<p>Rodomagoulakis et al. (<xref ref-type="bibr" rid="B113">2016</xref>) develop a smart interface featuring multimodal sensory processing abilities for human action detection within the context of assistive robots, exploring cutting-edge techniques for automated-localization cognition and visual activity recognition to multimodally identify commands and activities. Loth et al. (<xref ref-type="bibr" rid="B84">2015</xref>) measure the recognizer modes that are important at various levels of human&#x02013;robot interaction, providing insight into social behavior in humans to create socially adept robots.</p>
<p>In recent years, many contributions have been made to investigate how multimodal signal output influences HRI. Bird (<xref ref-type="bibr" rid="B18">2021</xref>) explore ways to give a robot social understanding through emotional perception for both verbal and non-verbal interaction, demonstrating how the framework&#x00027;s technology, organizational structure, and interactional examples address several outstanding concerns in the field. Yadav et al. (<xref ref-type="bibr" rid="B150">2021</xref>) provide a thorough analysis of various multimodal techniques for motion identification, using different sensors and analytical strategies with methodological fusion methods. Liu et al. (<xref ref-type="bibr" rid="B82">2022</xref>) investigate multimodal information-driven robot control for cooperative assembly between humans and robots, creating a human&#x02013;robot interface free of programming using function blocks to combine multimodal human instructions that precisely activate specified robot control modes.</p>
<p>Khalifa et al. (<xref ref-type="bibr" rid="B65">2022</xref>) present a robust framework for face tracking and identification in unrestricted environments, designing their framework based on lightweight CNNs to increase accuracy while preserving real-time capabilities essential for HRI systems. Shenoy et al. (<xref ref-type="bibr" rid="B126">2021</xref>) improve the interaction capabilities of Nao humanoid robots by combining detection models for facial expression and speech quality, using the microphone and camera to assess pain and mood in children receiving procedural therapy. Tziafas and Kasaei (<xref ref-type="bibr" rid="B141">2021</xref>) introduce a software architecture that isolates a target object from a congested scene based on vocal cues from a user, employing a multimodal deep neural net as the system&#x00027;s core for visual grounding. The research proposes the CFBRL-KCCA multimodal material recognition framework for object recognition challenges, demonstrating that the suggested fusion algorithm provides a useful method for material discovery (Wang et al., <xref ref-type="bibr" rid="B146">2021</xref>).</p>
<p>The application of multimodal signal output in HRI has become a fast-growing field; however, there are still many challenges to be addressed. As human&#x02013;robot interaction continues to be a hot research topic, researchers will undoubtedly explore new methods and solutions to enhance multimodal signal output and improve the overall HRI experience.</p>
</sec>
<sec>
<title>5.3. Major application areas of multimodal HRI</title>
<p>This section provides a comprehensive review and summarizes recent research advances in the application of multimodal HRI by examining numerous recent publications.</p>
<p>The emergence of sensing technologies and the increasing popularity of robotics have enabled researchers to study multimodal HRI, with numerous papers published each year. Gast et al. (<xref ref-type="bibr" rid="B44">2009</xref>) present a novel outline for real-time multimodal information processing, designed for scenarios involving human-human or human&#x02013;robot interaction and including modules for various output and input signals. Chen et al. (<xref ref-type="bibr" rid="B27">2022</xref>) develop a real-time, multi-model HRC scheme using voice and gestures, creating a collection of 16 dynamic gestures for human-to-industrial robot interaction and making a data collection of dynamic gestures publicly available.</p>
<p>Haninger et al. (<xref ref-type="bibr" rid="B53">2022</xref>) introduce a unique approach for multimodal pHRI, creating a Gaussian process model for human power in each state of a joint effort, and applying these frameworks to model predictive command and Bayesian inference of the style to forecast robot responses. Thomas et al. (<xref ref-type="bibr" rid="B136">2022</xref>) present a multimodal HRI platform that combines speech and hand sign input to control a UGV, translating vocal instructions into the ROS environment to drive the Argo Atlas J8 UGV using Mycroft, an accessible digital assistant. &#x00160;vec et al. (<xref ref-type="bibr" rid="B133">2022</xref>) introduces a multimodal cloud-based system for HRI, with the key contribution being the construction of the architecture based on industry-recognized frameworks, protocols, and JSON messages that have been verified.</p>
<p>A self-tuning multimodal fusion method is proposed to address the issue of helping robots achieve better intention comprehension. This method is not constrained by the manifestations of interacting individuals and surroundings, making it applicable to diverse platforms (Hou et al., <xref ref-type="bibr" rid="B59">2022</xref>). Weerakoon et al. (<xref ref-type="bibr" rid="B147">2022</xref>) present the COSM2IC system, which uses a compact Task Complexity Predictor and multiple sensor data input to evaluate the instructional richness to reduce loss in precision. This structure dynamically switches between a collection of models with different computational intensities so that computationally less demanding models are instantiated whenever viable.</p>
<p>Jooyeun Ham et al. introduce a versatile and elastic multimodal sensor system coupled with a soft bionic arm. They employ a manufacturing strategy that uses both UV laser metallic ablation and plastic cutting concurrently to construct sensor electrode designs and elastic conducting wires in a Kirigami pattern, implementing the layout of wired sensors on an adjustable metalized film (Bao et al., <xref ref-type="bibr" rid="B12">2022</xref>). Bucker et al. (<xref ref-type="bibr" rid="B20">2022</xref>) provide a versatile language-based user interface for HRC, taking advantage of recent developments in big language models to encapsulate the operator command, and employing multimodal focus transformers to integrate these characteristics with trajectory data. The mobility signal of the robot and the client&#x00027;s cardiac signal are gathered and combined to provide multimodal data as the input node vector of the DL framework, which is utilized for the control system&#x00027;s model of HRI (Wang W. et al., <xref ref-type="bibr" rid="B145">2022</xref>). Maniscalco et al. (<xref ref-type="bibr" rid="B87">2022</xref>) evaluate and suitably filter all the robotic sensory data required to fulfill their interaction model, paying careful attention to backchannel interaction, making it bilateral and visible through audio and visual cues. Wang R. et al. (<xref ref-type="bibr" rid="B144">2022</xref>) offer Husformer, a multimodal transformer architecture for multimodal human condition identification, suggesting the use of cross-modal transformers, which motivate one signal to strengthen itself by directly responding to latent relevancy disclosed in other signals. The focus on multimodal HRI has brought many concepts into practice.</p>
<p>Multimodal HRI has been developing rapidly, with numerous new methods for HRI using different modalities being proposed. Strazdas et al. (<xref ref-type="bibr" rid="B132">2022</xref>) create and test a novel multimodal scheme for non-contact human-machine interaction based on voice, face, and gesture detection, assessing the user experience and communication efficiency of their current scheme in a large study with many participants. Zeng and Luo suggest a solution for enhancing the precision of multimodal haptic signal detection by improving the SVM multi-classifier using a binary tree. The modified particle swarm clustered technique is utilized to optimize the binary tree structure, minimize the error piling of the binary leaf node SVM multi-classifier, and increase multimodal haptic signal identification accuracy (Zeng and Luo, <xref ref-type="bibr" rid="B156">2022</xref>). Nagahanumaiah (<xref ref-type="bibr" rid="B97">2022</xref>) develops a tiredness detection algorithm based on real-time information collected from wearable sensors, with the goal of understanding more about how humans feel fatigued in a supervisory human-machine setting, examining machine learning techniques for tiredness identification, and employing robots to modify their interactions.</p>
<p>Schreiter et al. (<xref ref-type="bibr" rid="B121">2022</xref>) aim to deliver high-quality tracking data from activity capture, eye-gaze trackers, and robotic sensors in a semantically rich context, using loosely scripted tasks to produce natural behavior in the videotaped participants, which leads the attendees to move through the changing lab setting in a natural and deliberate manner. In an HRC scenario, Armleder et al. (<xref ref-type="bibr" rid="B10">2022</xref>) develop and implement a control scheme that can enable the implementation of large-scale robotic skin, demonstrating how entire tactile feedback may enhance robot abilities during dynamic interplay by delivering information about various contacts throughout the robot&#x00027;s exterior.</p>
<p>The application of multimodal HRI is extensive, including using multiple sensors and inputs to evaluate social interactions, incorporating time delay and context data to improve recognition and emotional depiction, and developing unique models that combine different modalities. Tatarian et al. (<xref ref-type="bibr" rid="B135">2022</xref>) provide a multimodal interaction that focuses on proxemics of interpersonal navigating, gaze mechanics, kinesics, and social conversation, examining the impact of multimodal actions on relative social IQ using both subjective and objective assessments in a seven-minute encounter with 105 participants. Moroto et al. (<xref ref-type="bibr" rid="B96">2022</xref>) develop a recognition approach that considers the time delay to get genuinely near the reality of the occurring mechanism of feelings, with experimental findings demonstrating the usefulness of taking into account the time lag between gazing and brain function data.</p>
<p>He et al. (<xref ref-type="bibr" rid="B55">2022</xref>) present a unique multimodal M2NN model using the merging of EEG and fNIRS inputs to increase the recognition speed and generalization capacity of MI, combining spatial-temporal extraction of features, multimodal feature synthesis, and MTL. Zhang et al. (<xref ref-type="bibr" rid="B158">2022</xref>) retrieve effective active parts from sEMG data acquired by the MYO wristband using active element detection, then extracting five time-domain parameters from the main section signal: the root average square value, wave duration, number of zero-crossing spots, mean absolute value, and maximum-minimum value. Yang et al. incorporate context data into the current speech by embedding prior statements between interlocutors, which improves the emotional depiction of the present utterance. The suggested cross-modal converter module then focuses on the interconnections between text and auditory modalities, adaptively fostering modality fusion (Yang et al., <xref ref-type="bibr" rid="B151">2022</xref>). Based on the proposed papers listed above, it is clear that multimodality currently plays a significant role in HRI research.</p>
<p>In conclusion, multimodal HRI has seen rapid development and a wide range of applications in recent years. Researchers are exploring various methods and techniques to improve human&#x02013;robot interaction by using multiple modalities, such as voice, gestures, and facial expressions. As more advancements are made in this field, it is expected that multimodal HRI will continue to play a crucial role in shaping the future of human&#x02013;robot interaction. The application of multimodal HRI has expanded across various fields, including robotics, healthcare, COVID-19 diagnosis, secure planning/control, and co-adaptation. Researchers have explored the use of multiple modalities in emotion recognition, gesture recognition, EEG and fNIRS data merging, sensor data processing, speech recognition, and human mobility assessment. Additionally, multimodal HRI has shown potential in medical diagnosis and prognosis, such as epilepsy, creating robots with advanced multimodal mobility, AI-aided fashion design, and the integration of robotics and neuroscience.</p>
<p>The advantages of multimodal HRI include natural and intuitive interaction between humans and robots, increased accuracy and robustness in sensing and control, and the ability to handle complex tasks and situations. However, challenges remain, such as data fusion, algorithm development, and system integration.</p>
<p>Multimodal HRI is a growing field with many areas yet to be explored. As research continues, it is expected that multimodal HRI will play a crucial role in shaping the future of human&#x02013;robot interaction, leading to more efficient, user-friendly, and versatile robotic systems.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s6">
<title>6. Discussion</title>
<p>Multimodal human&#x02013;robot interaction is a field of research that aims to improve the way humans and robots communicate with each other. It is based on the idea that humans use multiple modalities, such as speech, gesture, and facial expression, to convey meaning and that robots should be able to understand and respond to these modalities in a natural and intuitive way.</p>
<sec>
<title>6.1. Natural language processing and computer vision</title>
<p>Natural language processing is widely used in multimodal HRI for speech recognition and understanding, which has the advantage of being able to handle a wide range of spoken languages. However, NLP&#x00027;s accuracy and performance are heavily dependent on the quality and quantity of the training data, which can be a challenge for rare or dialectal languages. Moreover, the recognition of ambiguous phrases or slang can lead to incorrect interpretations.</p>
<p>Computer vision techniques, such as gesture and facial expression recognition, have shown great potential in enhancing the naturalness and expressiveness of robot interactions. These techniques can detect subtle and nuanced movements that may be difficult for humans to perceive. However, limitations of computer vision include its sensitivity to lighting conditions, occlusions, and variations in appearance across individuals. Furthermore, these techniques require high computational power, making them unsuitable for resource-constrained robots.</p>
</sec>
<sec>
<title>6.2. Machine learning and haptic feedback</title>
<p>Machine learning techniques are essential for integrating and interpreting different modalities, including speech, vision, and haptic feedback. ML algorithms enable the robot to recognize and understand complex patterns in multimodal data, making it possible to provide natural and adaptive interactions. However, models may be biased or fail to generalize to unseen data, leading to reduced performance in real-world scenarios.</p>
<p>Haptic feedback and motion planning techniques are particularly useful for physical interaction between humans and robots. Haptic feedback provides a sense of touch, allowing robots to respond to human gestures and movements in a natural way. Motion planning algorithms enable the robot to navigate in a human environment safely and efficiently. However, haptic feedback and motion planning require high precision and accuracy, which can be challenging to achieve in complex and dynamic environments.</p>
</sec>
<sec>
<title>6.3. Deep learning and touch-based interaction</title>
<p>Although deep learning techniques have shown great potential in recognizing and interpreting human gestures and expressions, there are still some challenges that need to be addressed. One challenge is the need for a large amount of labeled data to train deep learning models, which can be time-consuming and expensive to obtain. Another challenge is the need for robustness to variations in lighting, background, and appearance of human gestures and expressions. Despite these challenges, deep learning techniques have the potential to significantly improve the accuracy and robustness of gesture and expression recognition in human&#x02013;robot interaction.</p>
<p>The use of haptic feedback for touch-based interaction has great potential for improving the naturalness and intuitiveness of human&#x02013;robot communication. However, there are still challenges that need to be addressed, such as the need for high-quality and responsive haptic feedback that can mimic human touch, and the need for effective motion planning algorithms that can ensure safe and efficient interactions between humans and robots. Nevertheless, with the ongoing advancements in haptic technology and motion planning algorithms, it is expected that touch-based interaction will become an increasingly important aspect of multimodal human&#x02013;robot interaction in the future.</p>
</sec>
<sec>
<title>6.4. Future directions</title>
<p>In the future, there will be more emphasis on creating more natural and intuitive interaction, as well as improving the robots&#x00027; ability to understand and respond to human emotions. This will be achieved through the integration of emotion recognition and generation algorithms, making robots more human-like. Another trend will be the use of multi-robot systems, in which multiple robots work together to accomplish a task. This will allow for more complex and efficient interactions between humans and robots.</p>
<p>In addition to the integration of emotion recognition and generation algorithms, there will also be a focus on creating robots that can adapt to individual differences in communication style and preferences. This could be achieved through personalized learning and adaptation techniques. Finally, ethical considerations in human&#x02013;robot interaction will become increasingly important, and there will be a need for ethical guidelines and regulations to ensure the safe and responsible use of robots in various applications.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s7">
<title>7. Conclusion</title>
<p>The current state and emerging directions of multimodal human&#x02013;robot interaction is thoroughly discussed in this paper. Also, we have thoroughly combed the research progress in terms of the information input for multimodal HRI, the information output for multimodal HRI, as well as the concrete applications of multimodal HRI. Specifically, this review elaborates on the research progress of information input for multimodal HRI from three perspectives: gesture recognition, speech recognition, as well as emotion recognition. In terms of information output, gesture generation and emotional expression generation are covered. Research in this area has focused on developing various modalities, such as speech, gesture, and facial expression, to enable robots to understand better and respond to human intentions and emotions. The integration of multiple modalities is also crucial for achieving robust and flexible human&#x02013;robot interaction. The major limitation of the study lies in the limited number of real-world deployments of multimodal human&#x02013;robot interaction systems, so the impact of the technology on users may not be well understood. Also, there are technical challenges, such as high computational requirements and system complexity that limit the scalability of multimodal human&#x02013;robot interaction systems. Hopefully, this paper will reflect the current research trend in human&#x02013;robot interaction and provide guidance for future research.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>HS and WQ contributed to the draft writing and the other authors contributed to correcting, supervising, and proof reading, etc. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Admoni</surname> <given-names>H.</given-names></name> <name><surname>Scassellati</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>Social eye gaze in human-robot interaction: a review</article-title>. <source>J. Hum. Robot Interact</source>. <volume>6</volume>, <fpage>25</fpage>&#x02013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.5898/JHRI.6.1.Admoni</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ahn</surname> <given-names>H. S.</given-names></name> <name><surname>Yep</surname> <given-names>W.</given-names></name> <name><surname>Lim</surname> <given-names>J.</given-names></name> <name><surname>Ahn</surname> <given-names>B. K.</given-names></name> <name><surname>Johanson</surname> <given-names>D. L.</given-names></name> <name><surname>Hwang</surname> <given-names>E. J.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Hospital receptionist robot v2: design for enhancing verbal interaction with social skills,&#x0201D;</article-title> in <source>2019 28th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)</source> (<publisher-loc>New Delhi</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/RO-MAN46459.2019.8956300</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alghowinem</surname> <given-names>S.</given-names></name> <name><surname>Jeong</surname> <given-names>S.</given-names></name> <name><surname>Arias</surname> <given-names>K.</given-names></name> <name><surname>Picard</surname> <given-names>R.</given-names></name> <name><surname>Breazeal</surname> <given-names>C.</given-names></name> <name><surname>Park</surname> <given-names>H. W.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Beyond the words: analysis and detection of self-disclosure behavior during robot positive psychology interaction,&#x0201D;</article-title> in <source>2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021)</source> (<publisher-loc>Jodhpur</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>01</fpage>&#x02013;<lpage>08</lpage>. <pub-id pub-id-type="doi">10.1109/FG52635.2021.9666969</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Al-Nafjan</surname> <given-names>A.</given-names></name> <name><surname>Hosny</surname> <given-names>M.</given-names></name> <name><surname>Al-Ohali</surname> <given-names>Y.</given-names></name> <name><surname>Al-Wabil</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Review and classification of emotion recognition based on EEG brain-computer interface system research: a systematic review</article-title>. <source>Appl. Sci.</source> <volume>7</volume>, <fpage>1239</fpage>. <pub-id pub-id-type="doi">10.3390/app7121239</pub-id><pub-id pub-id-type="pmid">32955675</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alves</surname> <given-names>J.</given-names></name> <name><surname>Seabra</surname> <given-names>E.</given-names></name> <name><surname>Caetano</surname> <given-names>I.</given-names></name> <name><surname>Gon&#x000E7;alves</surname> <given-names>J.</given-names></name> <name><surname>Serra</surname> <given-names>J.</given-names></name> <name><surname>Martins</surname> <given-names>M.</given-names></name> <name><surname>Santos</surname> <given-names>C. P.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Considerations and mechanical modifications on a smart walker,&#x0201D;</article-title> in <source>2016 International Conference on Autonomous Robot Systems and Competitions (ICARSC)</source> (<publisher-loc>Bragan&#x000E7;a</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>247</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1109/ICARSC.2016.30</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Aly</surname> <given-names>A.</given-names></name> <name><surname>Tapus</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Prosody-driven robot arm gestures generation in human-robot interaction,&#x0201D;</article-title> in <source>Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>257</fpage>&#x02013;<lpage>258</lpage>. <pub-id pub-id-type="doi">10.1145/2157689.2157783</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ameri</surname> <given-names>A.</given-names></name> <name><surname>Akhaee</surname> <given-names>M. A.</given-names></name> <name><surname>Scheme</surname> <given-names>E.</given-names></name> <name><surname>Englehart</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Real-time, simultaneous myoelectric control using a convolutional neural network</article-title>. <source>PLoS ONE</source>, <volume>13</volume>, <fpage>e0203835</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0203835</pub-id><pub-id pub-id-type="pmid">30212573</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Andhare</surname> <given-names>P.</given-names></name> <name><surname>Rawat</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Pick and place industrial robot controller with computer vision,&#x0201D;</article-title> in <source>2016 International Conference on Computing Communication Control and automation (ICCUBEA)</source> (<publisher-loc>Pune</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1109/ICCUBEA.2016.7860048</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andronas</surname> <given-names>D.</given-names></name> <name><surname>Apostolopoulos</surname> <given-names>G.</given-names></name> <name><surname>Fourtakas</surname> <given-names>N.</given-names></name> <name><surname>Makris</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Multi-modal interfaces for natural human-robot interaction</article-title>. <source>Procedia Manuf</source>. <volume>54</volume>, <fpage>197</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1016/j.promfg.2021.07.030</pub-id><pub-id pub-id-type="pmid">30356005</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Armleder</surname> <given-names>S.</given-names></name> <name><surname>Dean-Leon</surname> <given-names>E.</given-names></name> <name><surname>Bergner</surname> <given-names>F.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name></person-group> (<year>2022</year>). <article-title>Interactive force control based on multimodal robot skin for physical human- robot collaboration</article-title>. <source>Adv. Intell. Syst</source>. <volume>4</volume>, <fpage>2100047</fpage>. <pub-id pub-id-type="doi">10.1002/aisy.202100047</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashok</surname> <given-names>K.</given-names></name> <name><surname>Ashraf</surname> <given-names>M.</given-names></name> <name><surname>Thimmia Raja</surname> <given-names>J.</given-names></name> <name><surname>Hussain</surname> <given-names>M. Z.</given-names></name> <name><surname>Singh</surname> <given-names>D. K.</given-names></name> <name><surname>Haldorai</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Collaborative analysis of audio-visual speech synthesis with sensor measurements for regulating human-robot interaction</article-title>. <source>Int. J. Syst. Assur. Eng. Manag</source>. <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1007/s13198-022-01709-y</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>Z.</given-names></name> <name><surname>Ham</surname> <given-names>J.</given-names></name> <name><surname>Cutkosky</surname> <given-names>M.</given-names></name> <name><surname>Han</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Flexible and stretchable multi-modal sensor network for soft robot interaction</article-title>. <source>Res. Squ. [Preprint]</source>. <pub-id pub-id-type="doi">10.21203/rs.3.rs-1654721/v1</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barber</surname> <given-names>D. J.</given-names></name> <name><surname>Reinerman-Jones</surname> <given-names>L. E.</given-names></name> <name><surname>Matthews</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Toward a tactile language for human-robot interaction: two studies of tacton learning and performance</article-title>. <source>Hum. Factors</source> <volume>57</volume>, <fpage>471</fpage>&#x02013;<lpage>490</lpage>. <pub-id pub-id-type="doi">10.1177/0018720814548063</pub-id><pub-id pub-id-type="pmid">25875436</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barricelli</surname> <given-names>B. R.</given-names></name> <name><surname>Fogli</surname> <given-names>D.</given-names></name> <name><surname>Iemmolo</surname> <given-names>L.</given-names></name> <name><surname>Locoro</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;A multi-modal approach to creating routines for smart speakers,&#x0201D;</article-title> in <source>Proceedings of the 2022 International Conference on Advanced Visual Interfaces</source> (<publisher-loc>Rome</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1145/3531073.3531168</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bastianelli</surname> <given-names>E.</given-names></name> <name><surname>Castellucci</surname> <given-names>G.</given-names></name> <name><surname>Croce</surname> <given-names>D.</given-names></name> <name><surname>Basili</surname> <given-names>R.</given-names></name> <name><surname>Nardi</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <source>Effective and Robust Natural Language Understanding for Human-Robot Interaction</source> (<publisher-loc>Prague</publisher-loc>: <publisher-name>ECAI</publisher-name>), <fpage>57</fpage>&#x02013;<lpage>62</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Belpaeme</surname> <given-names>T.</given-names></name> <name><surname>Vogt</surname> <given-names>P.</given-names></name> <name><surname>Van den Berghe</surname> <given-names>R.</given-names></name> <name><surname>Bergmann</surname> <given-names>K.</given-names></name> <name><surname>G&#x000F6;ksun</surname> <given-names>T.</given-names></name> <name><surname>De Haas</surname> <given-names>M.</given-names></name> <name><surname>Kanero</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Guidelines for designing social robots as second language tutors</article-title>. <source>Int. J. Soc. Robot</source>. <volume>10</volume>, <fpage>325</fpage>&#x02013;<lpage>341</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-018-0467-6</pub-id><pub-id pub-id-type="pmid">30996752</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berg</surname> <given-names>J.</given-names></name> <name><surname>Lu</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Review of interfaces for industrial human-robot interaction</article-title>. <source>Curr. Robot. Rep</source>. <volume>1</volume>, <fpage>27</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1007/s43154-020-00005-6</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bird</surname> <given-names>J. J.</given-names></name></person-group> (<year>2021</year>). <source>A Socially Interactive Multimodal Human-Robot Interaction Framework through Studies on Machine and Deep Learning</source> [PhD thesis]. <publisher-loc>Birmingham</publisher-loc>: <publisher-name>Aston University</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bolotnikova</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <source>Frail Human Assistance by a Humanoid Robot Using Multi-contact Planning and Physical Interaction</source> [PhD thesis]. <publisher-loc>Montpellier</publisher-loc>: <publisher-name>Universit&#x000E9; de Montpellier</publisher-name>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bucker</surname> <given-names>A.</given-names></name> <name><surname>Figueredo</surname> <given-names>L.</given-names></name> <name><surname>Haddadinl</surname> <given-names>S.</given-names></name> <name><surname>Kapoor</surname> <given-names>A.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <name><surname>Bonatti</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Reshaping robot trajectories using natural language commands: A study of multi-modal data alignment using transformers,&#x0201D;</article-title> in <source>2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IEEE)</source>, <fpage>978</fpage>&#x02013;<lpage>984</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Caetano</surname> <given-names>I.</given-names></name> <name><surname>Alves</surname> <given-names>J. Gon&#x000E7;alves, J.</given-names></name> <name><surname>Martins</surname> <given-names>M.</given-names></name> <name><surname>Santos</surname> <given-names>C. P.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Development of a biofeedback approach using body tracking with active depth sensor in asbgo smart walker,&#x0201D;</article-title> in <source>2016 International Conference on Autonomous Robot Systems and Competitions (ICARSC)</source> (<publisher-loc>Bragan&#x000E7;a</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>241</fpage>&#x02013;<lpage>246</lpage>. <pub-id pub-id-type="doi">10.1109/ICARSC.2016.34</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Castillo</surname> <given-names>E.</given-names></name> <name><surname>Morales</surname> <given-names>D. P. Garc&#x000ED;a, A.</given-names></name> <name><surname>Mart&#x000ED;nez-Mart&#x000ED;</surname> <given-names>F.</given-names></name> <name><surname>Parrilla</surname> <given-names>L.</given-names></name> <name><surname>Palma</surname> <given-names>A. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Noise suppression in ECG signals through efficient one-step wavelet processing techniques</article-title>. <source>J. Appl. Math</source>. <volume>2013</volume>, <fpage>763903</fpage>. <pub-id pub-id-type="doi">10.1155/2013/763903</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cavallo</surname> <given-names>F.</given-names></name> <name><surname>Semeraro</surname> <given-names>F.</given-names></name> <name><surname>Fiorini</surname> <given-names>L.</given-names></name> <name><surname>Magyar</surname> <given-names>G. Sin&#x0010D;&#x000E1;k, P.</given-names></name> <name><surname>Dario</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Emotion modelling for social robotics applications: a review</article-title>. <source>J. Bionic Eng</source>. <volume>15</volume>, <fpage>185</fpage>&#x02013;<lpage>203</lpage>. <pub-id pub-id-type="doi">10.1007/s42235-018-0015-y</pub-id><pub-id pub-id-type="pmid">34264899</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cazzato</surname> <given-names>D.</given-names></name> <name><surname>Cimarelli</surname> <given-names>C.</given-names></name> <name><surname>Sanchez-Lopez</surname> <given-names>J. L.</given-names></name> <name><surname>Voos</surname> <given-names>H.</given-names></name> <name><surname>Leo</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey of computer vision methods for 2d object detection from unmanned aerial vehicles</article-title>. <source>J. Imag.</source> <volume>6</volume>, <fpage>78</fpage>. <pub-id pub-id-type="doi">10.3390/jimaging6080078</pub-id><pub-id pub-id-type="pmid">34460693</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Celiktutan</surname> <given-names>O.</given-names></name> <name><surname>Sariyanidi</surname> <given-names>E.</given-names></name> <name><surname>Gunes</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Computational analysis of affect, personality, and engagement in human-robot interactions,&#x0201D;</article-title> in <source>Computer Vision for Assistive Healthcare</source>, eds <person-group person-group-type="editor"><name><surname>Marco</surname> <given-names>L.</given-names></name> <name><surname>Farinella</surname> <given-names>G. M.</given-names></name></person-group> (<publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Elsevier</publisher-name>), <fpage>283</fpage>&#x02013;<lpage>318</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-813445-0.00010-1</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chae</surname> <given-names>Y.-J.</given-names></name> <name><surname>Nam</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>D.</given-names></name> <name><surname>Sin</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>C.</given-names></name> <name><surname>Park</surname> <given-names>S.-K.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Generation of co-speech gestures of robot based on morphemic analysis</article-title>. <source>Rob. Auton. Syst</source>. <volume>155</volume>, <fpage>104154</fpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2022.104154</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Leu</surname> <given-names>M. C.</given-names></name> <name><surname>Yin</surname> <given-names>Z.</given-names></name></person-group> (<year>2022</year>). <article-title>Real-time multi-modal human-robot collaboration using gestures and speech</article-title>. <source>J. Manuf. Sci. Eng</source>. <volume>144</volume>, <fpage>1</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1115/1.4054297</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>M.</given-names></name> <name><surname>Hirota</surname> <given-names>K.</given-names></name> <name><surname>Pedrycz</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Multimodal emotion recognition and intention understanding in human-robot interaction,&#x0201D;</article-title> in <source>Developments in Advanced Control and Intelligent Automation for Complex Systems</source>, eds <person-group person-group-type="editor"><name><surname>Wu</surname> <given-names>M.</given-names></name> <name><surname>Pedrycz</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>255</fpage>&#x02013;<lpage>288</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-62147-6_10</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name></person-group> (<year>2015</year>). <article-title>Survey on 3D hand gesture recognition</article-title>. <source>IEEE Trans. Circuits Syst. Video Technol.</source> <volume>26</volume>, <fpage>1659</fpage>&#x02013;<lpage>1673</lpage>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cherubini</surname> <given-names>A.</given-names></name> <name><surname>Passama</surname> <given-names>R.</given-names></name> <name><surname>Navarro</surname> <given-names>B.</given-names></name> <name><surname>Sorour</surname> <given-names>M.</given-names></name> <name><surname>Khelloufi</surname> <given-names>A.</given-names></name> <name><surname>Mazhar</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>A collaborative robot for the factory of the future: bazar</article-title>. <source>Int. J. Adv. Manuf. Technol</source>. <volume>105</volume>, <fpage>3643</fpage>&#x02013;<lpage>3659</lpage>. <pub-id pub-id-type="doi">10.1007/s00170-019-03806-y</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chuy</surname> <given-names>O. Y.</given-names> <suffix>Jr.</suffix></name> <name><surname>Herrero</surname> <given-names>J.</given-names></name> <name><surname>Al-Selwadi</surname> <given-names>A.</given-names></name> <name><surname>Mooers</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Control and evaluation of a motorized attendant wheelchair with haptic interface</article-title>. <source>J. Med. Device</source> <volume>13</volume>, <fpage>011002</fpage>. <pub-id pub-id-type="doi">10.1115/1.4041336</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cid</surname> <given-names>F.</given-names></name> <name><surname>Manso</surname> <given-names>L. J.</given-names></name> <name><surname>N&#x000FA;nez</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;A novel multimodal emotion recognition approach for affective human robot interaction,&#x0201D;</article-title> in <source>Proceedings of Fine</source> (<publisher-loc>Hamburg</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>9</lpage>.<pub-id pub-id-type="pmid">27534393</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Covington</surname> <given-names>M. A.</given-names></name></person-group> (<year>2001</year>). <article-title>Building natural language generation systems</article-title>. <source>Language</source> <volume>77</volume>, <fpage>611</fpage>&#x02013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1353/lan.2001.0146</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Das</surname> <given-names>P.</given-names></name> <name><surname>Khasnobish</surname> <given-names>A.</given-names></name> <name><surname>Tibarewala</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Emotion recognition employing ECG and GSR signals as markers of ans,&#x0201D;</article-title> in <source>2016 Conference on Advances in Signal Processing (CASP)</source> (<publisher-loc>Pune</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>37</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1109/CASP.2016.7746134</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>der Loos</surname> <given-names>V.</given-names></name> <name><surname>Machiel</surname> <given-names>H.</given-names></name> <name><surname>Reinkensmeyer</surname> <given-names>D. J.</given-names></name> <name><surname>Guglielmelli</surname> <given-names>E.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Rehabilitation and health care robotics,&#x0201D;</article-title> in <source>Springer Handbook of Robotics</source>, eds <person-group person-group-type="editor"><name><surname>Siciliano</surname> <given-names>B.</given-names></name> <name><surname>Khatib</surname> <given-names>O.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>1685</fpage>&#x02013;<lpage>1728</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-32552-1_64</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Desai</surname> <given-names>S.</given-names></name> <name><surname>Mantha</surname> <given-names>S.</given-names></name> <name><surname>Phalle</surname> <given-names>V.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Advances in smart wheelchair technology,&#x0201D;</article-title> in <source>2017 International Conference on Nascent Technologies in Engineering (ICNTE)</source> (<publisher-loc>Vashi</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1109/ICNTE.2017.7947914</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deuerlein</surname> <given-names>C.</given-names></name> <name><surname>Langer</surname> <given-names>M. Se&#x000DF;ner, J.</given-names></name> <name><surname>He&#x000DF;</surname> <given-names>P.</given-names></name> <name><surname>Franke</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Human-robot-interaction using cloud-based speech recognition systems</article-title>. <source>Procedia CIRP</source> <volume>97</volume>, <fpage>130</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1016/j.procir.2020.05.214</pub-id><pub-id pub-id-type="pmid">33816568</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>B.</given-names></name> <name><surname>Wei</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>F.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Skill learning for human-robot interaction using wearable device</article-title>. <source>Tsinghua Sci. Technol</source>. <volume>24</volume>, <fpage>654</fpage>&#x02013;<lpage>662</lpage>. <pub-id pub-id-type="doi">10.26599/TST.2018.9010096</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feng</surname> <given-names>M.</given-names></name> <name><surname>Huang</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>B.</given-names></name> <name><surname>Zheng</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Accurate calibration of a multi-camera system based on flat refractive geometry</article-title>. <source>Appl. Opt.</source> <volume>56</volume>, <fpage>9724</fpage>&#x02013;<lpage>9734</lpage>.<pub-id pub-id-type="pmid">29240118</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferlinc</surname> <given-names>A.</given-names></name> <name><surname>Fabiani</surname> <given-names>E.</given-names></name> <name><surname>Velnar</surname> <given-names>T.</given-names></name> <name><surname>Gradisnik</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>The importance and role of proprioception in the elderly: a short review</article-title>. <source>Mater. Sociomed</source>. <volume>31</volume>, <fpage>219</fpage>. <pub-id pub-id-type="doi">10.5455/msm.2019.31.219-221</pub-id><pub-id pub-id-type="pmid">31762707</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fritzsche</surname> <given-names>M.</given-names></name> <name><surname>Saenz</surname> <given-names>J.</given-names></name> <name><surname>Penzlin</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;A large scale tactile sensor for safe mobile robot manipulation,&#x0201D;</article-title> in <source>2016 11th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source> (<publisher-loc>Christchurch</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>427</fpage>&#x02013;<lpage>428</lpage>. <pub-id pub-id-type="doi">10.1109/HRI.2016.7451789</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gams</surname> <given-names>A.</given-names></name> <name><surname>Ude</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;On-line coaching of robots through visual and physical interaction: analysis of effectiveness of human-robot interaction strategies,&#x0201D;</article-title> in <source>2016 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Stockholm</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3028</fpage>&#x02013;<lpage>3034</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA.2016.7487467</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>Q.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Ju</surname> <given-names>Z.</given-names></name></person-group> (<year>2021</year>). <article-title>Hand gesture recognition using multimodal data fusion and multiscale parallel convolutional neural network for human-robot interaction</article-title>. <source>Expert Syst</source>. <volume>38</volume>, <fpage>e12490</fpage>. <pub-id pub-id-type="doi">10.1111/exsy.12490</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gast</surname> <given-names>J.</given-names></name> <name><surname>Bannat</surname> <given-names>A.</given-names></name> <name><surname>Rehrl</surname> <given-names>T.</given-names></name> <name><surname>Wallhoff</surname> <given-names>F.</given-names></name> <name><surname>Rigoll</surname> <given-names>G.</given-names></name> <name><surname>Wendt</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>&#x0201C;Real-time framework for multimodal human-robot interaction,&#x0201D;</article-title> in <source>2009 2nd Conference on Human System Interactions</source> (<publisher-loc>Catania</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>276</fpage>&#x02013;<lpage>283</lpage>. <pub-id pub-id-type="doi">10.1109/HSI.2009.5090992</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Giudice</surname> <given-names>N. A.</given-names></name> <name><surname>Legge</surname> <given-names>G. E.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Blind navigation and the role of technology,&#x0201D;</article-title> in <source>The Engineering Handbook of Smart Technology for Aging, Disability, and Independence</source>. p. <fpage>479</fpage>&#x02013;<lpage>500</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gopinathan</surname> <given-names>S.</given-names></name> <name><surname>&#x000D6;tting</surname> <given-names>S. K.</given-names></name> <name><surname>Steil</surname> <given-names>J. J.</given-names></name></person-group> (<year>2017</year>). <article-title>A user study on personalized stiffness control and task specificity in physical human-robot interaction</article-title>. <source>Front. Robot. AI</source> <volume>4</volume>, <fpage>58</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2017.00058</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gozzi</surname> <given-names>N.</given-names></name> <name><surname>Malandri</surname> <given-names>L.</given-names></name> <name><surname>Mercorio</surname> <given-names>F.</given-names></name> <name><surname>Pedrocchi</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Xai for myo-controlled prosthesis: explaining emg data for hand gesture classification</article-title>. <source>Knowl. Based Syst</source>. <volume>240</volume>, <fpage>108053</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2021.108053</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Groechel</surname> <given-names>T.</given-names></name> <name><surname>Pakkar</surname> <given-names>R.</given-names></name> <name><surname>Dasgupta</surname> <given-names>R.</given-names></name> <name><surname>Kuo</surname> <given-names>C.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Cordero</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Kinesthetic curiosity: towards personalized embodied learning with a robot tutor teaching programming in mixed reality,&#x0201D;</article-title> in <source>International Symposium on Experimental Robotics</source> (<publisher-loc>La Valletta</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>245</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-71151-1_22</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gui</surname> <given-names>K.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Toward multimodal human-robot interaction to enhance active participation of users in gait rehabilitation</article-title>. <source>IEEE Trans. Neural Syst. Rehabil. Eng</source>. <volume>25</volume>, <fpage>2054</fpage>&#x02013;<lpage>2066</lpage>. <pub-id pub-id-type="doi">10.1109/TNSRE.2017.2703586</pub-id><pub-id pub-id-type="pmid">28504943</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hahne</surname> <given-names>J. M.</given-names></name> <name><surname>Wilke</surname> <given-names>M. A.</given-names></name> <name><surname>Koppe</surname> <given-names>M.</given-names></name> <name><surname>Farina</surname> <given-names>D.</given-names></name> <name><surname>Schilling</surname> <given-names>A. F.</given-names></name></person-group> (<year>2020</year>). <article-title>Longitudinal case study of regression-based hand prosthesis control in daily life</article-title>. <source>Front. Neurosci</source>. <volume>14</volume>, <fpage>600</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2020.00600</pub-id><pub-id pub-id-type="pmid">32636734</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Campbell</surname> <given-names>N.</given-names></name> <name><surname>Jokinen</surname> <given-names>K.</given-names></name> <name><surname>Wilcock</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Investigating the use of non-verbal cues in human-robot interaction with a nao robot,&#x0201D;</article-title> in <source>2012 IEEE 3rd International Conference on Cognitive Infocommunications (CogInfoCom)</source> (<publisher-loc>Kosice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>679</fpage>&#x02013;<lpage>683</lpage>. <pub-id pub-id-type="doi">10.1109/CogInfoCom.2012.6421937</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>M.-J.</given-names></name> <name><surname>Lin</surname> <given-names>C.-H.</given-names></name> <name><surname>Song</surname> <given-names>K.-T.</given-names></name></person-group> (<year>2012</year>). <article-title>Robotic emotional expression generation based on mood transition and personality model</article-title>. <source>IEEE Trans. Cybern</source>. <volume>43</volume>, <fpage>1290</fpage>&#x02013;<lpage>1303</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCB.2012.2228851</pub-id><pub-id pub-id-type="pmid">26502437</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haninger</surname> <given-names>K.</given-names></name> <name><surname>Hegeler</surname> <given-names>C.</given-names></name> <name><surname>Peternel</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Model predictive control with gaussian processes for flexible multi-modal physical human robot interaction,&#x0201D;</article-title> in <source>2022 International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Philadelphia, PA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6948</fpage>&#x02013;<lpage>6955</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA46639.2022.9811590</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hasanuzzaman</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>T.</given-names></name> <name><surname>Ampornaramveth</surname> <given-names>V.</given-names></name> <name><surname>Gotoda</surname> <given-names>H.</given-names></name> <name><surname>Shirai</surname> <given-names>Y.</given-names></name> <name><surname>Ueno</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>Adaptive visual gesture recognition for human-robot interaction using a knowledge-based software platform</article-title>. <source>Rob. Auton. Syst</source>. <volume>55</volume>, <fpage>643</fpage>&#x02013;<lpage>657</lpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2007.03.002</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>Q.</given-names></name> <name><surname>Feng</surname> <given-names>L.</given-names></name> <name><surname>Jiang</surname> <given-names>G.</given-names></name> <name><surname>Xie</surname> <given-names>P.</given-names></name></person-group> (<year>2022</year>). <article-title>Multimodal multitask neural network for motor imagery classification with EEG and fNIRS signals</article-title>. <source>IEEE Sensors J.</source> <volume>22</volume>, <fpage>20695</fpage>&#x02013;<lpage>20706</lpage>. <pub-id pub-id-type="doi">10.1109/JSEN.2022.3205956</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heikkila</surname> <given-names>J.</given-names></name></person-group> (<year>2000</year>). <article-title>Geometric camera calibration using circular control points</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>22</volume>, <fpage>1066</fpage>&#x02013;<lpage>1077</lpage>. <pub-id pub-id-type="doi">10.1109/34.879788</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hoffman</surname> <given-names>G.</given-names></name> <name><surname>Breazeal</surname> <given-names>C.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Achieving fluency through perceptual-symbol practice in human-robot collaboration,&#x0201D;</article-title> in <source>2008 3rd ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1145/1349822.1349824</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hogan</surname> <given-names>N.</given-names></name></person-group> (<year>1984</year>). <article-title>&#x0201C;Impedance control: an approach to manipulation,&#x0201D;</article-title> in <source>1984 American Control Conference</source> (<publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>304</fpage>&#x02013;<lpage>313</lpage>. <pub-id pub-id-type="doi">10.23919/ACC.1984.4788393</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>Y.</given-names></name> <name><surname>Feng</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>T.</given-names></name> <name><surname>Qiu</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Stmmi: a self-tuning multi-modal fusion algorithm applied in assist robot interaction</article-title>. <source>Sci. Program</source>. <volume>2022</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1155/2022/3952758</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>House</surname> <given-names>B.</given-names></name> <name><surname>Malkin</surname> <given-names>J.</given-names></name> <name><surname>Bilmes</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;The voicebot: a voice controlled robot arm,&#x0201D;</article-title> in <source>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source> (<publisher-loc>Boston, MA</publisher-loc>), <fpage>183</fpage>&#x02013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1145/1518701.1518731</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huenerfauth</surname> <given-names>M.</given-names></name> <name><surname>Lu</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>Evaluation of a psycholinguistically motivated timing model for animations of american sign language</article-title>. <source>ACM Trans. Access. Comput</source>. <volume>5</volume>, <fpage>1</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1145/1414471.1414496</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Humphry</surname> <given-names>J.</given-names></name> <name><surname>Chesher</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>Preparing for smart voice assistants: cultural histories and media innovations</article-title>. <source>New Media Soc</source>. <volume>23</volume>, <fpage>1971</fpage>&#x02013;<lpage>1988</lpage>. <pub-id pub-id-type="doi">10.1177/1461444820923679</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ince</surname> <given-names>G.</given-names></name> <name><surname>Yorganci</surname> <given-names>R.</given-names></name> <name><surname>Ozkul</surname> <given-names>A.</given-names></name> <name><surname>Duman</surname> <given-names>T. B.</given-names></name> <name><surname>K&#x000F6;se</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>An audiovisual interface-based drumming system for multimodal human-robot interaction</article-title>. <source>J. Multimodal User Interfaces</source> <volume>15</volume>, <fpage>413</fpage>&#x02013;<lpage>428</lpage>. <pub-id pub-id-type="doi">10.1007/s12193-020-00352-w</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kavalieros</surname> <given-names>D.</given-names></name> <name><surname>Kapothanasis</surname> <given-names>E.</given-names></name> <name><surname>Kakarountas</surname> <given-names>A.</given-names></name> <name><surname>Loukopoulos</surname> <given-names>T.</given-names></name></person-group> (<year>2022</year>). <article-title>Methodology for selecting the appropriate electric motor for robotic modular systems for lower extremities</article-title>. <source>Healthcare</source> <volume>10</volume>, <fpage>2054</fpage>. <pub-id pub-id-type="doi">10.3390/healthcare10102054</pub-id><pub-id pub-id-type="pmid">36292506</pub-id></citation></ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khalifa</surname> <given-names>A.</given-names></name> <name><surname>Abdelrahman</surname> <given-names>A. A.</given-names></name> <name><surname>Strazdas</surname> <given-names>D.</given-names></name> <name><surname>Hintz</surname> <given-names>J.</given-names></name> <name><surname>Hempel</surname> <given-names>T.</given-names></name> <name><surname>Al-Hamadi</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Face recognition and tracking framework for human-robot interaction</article-title>. <source>Appl. Sci</source>. <volume>12</volume>, <fpage>5568</fpage>. <pub-id pub-id-type="doi">10.3390/app12115568</pub-id><pub-id pub-id-type="pmid">35528364</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khurana</surname> <given-names>D.</given-names></name> <name><surname>Koli</surname> <given-names>A.</given-names></name> <name><surname>Khatter</surname> <given-names>K.</given-names></name> <name><surname>Singh</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>Natural language processing: state of the art, current trends and challenges</article-title>. <source>Multimed. Tools Appl</source>. <volume>82</volume>, <fpage>3713</fpage>&#x02013;<lpage>3744</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-022-13428-4</pub-id><pub-id pub-id-type="pmid">35855771</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Szafir</surname> <given-names>D.</given-names></name> <name><surname>Mutlu</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;The impact of robot&#x00027;s expressive behavior on user&#x00027;s task performance,&#x0201D;</article-title> in <source>Proceedings of the 2016 ACM/IEEE International Conference on Human-Robot Interaction</source> (<publisher-loc>Christchurch</publisher-loc>), <fpage>168</fpage>&#x02013;<lpage>175</lpage>.</citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klauer</surname> <given-names>C.</given-names></name> <name><surname>Schauer</surname> <given-names>T.</given-names></name> <name><surname>Reichenfelser</surname> <given-names>W.</given-names></name> <name><surname>Karner</surname> <given-names>J.</given-names></name> <name><surname>Zwicker</surname> <given-names>S.</given-names></name> <name><surname>Gandolla</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Feedback control of arm movements using neuro-muscular electrical stimulation (NMES) combined with a lockable, passive exoskeleton for gravity compensation</article-title>. <source>Front. Neurosci</source>. <volume>8</volume>, <fpage>262</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2014.00262</pub-id><pub-id pub-id-type="pmid">25228853</pub-id></citation></ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Bergmann</surname> <given-names>K.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>I.</given-names></name></person-group> (<year>2008</year>). <article-title>Multimodal communication from multimodal thinking&#x02013;towards an integrated model of speech and gesture production</article-title>. <source>Int. J. Semant. Comput</source>. <volume>2</volume>, <fpage>115</fpage>&#x02013;<lpage>136</lpage>. <pub-id pub-id-type="doi">10.1142/S1793351X08000361</pub-id></citation>
</ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>I.</given-names></name></person-group> (<year>2004</year>). <article-title>Synthesizing multimodal utterances for conversational agents</article-title>. <source>Comput. Animat. Virtual Worlds</source> <volume>15</volume>, <fpage>39</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1002/cav.6</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>K&#x000FC;bler</surname> <given-names>S.</given-names></name> <name><surname>Cantrell</surname> <given-names>R.</given-names></name> <name><surname>Scheutz</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Actions speak louder than words: Evaluating parsers in the context of natural language understanding systems for human-robot interaction,&#x0201D;</article-title> in <source>Proceedings of the International Conference Recent Advances in Natural Language Processing 2011</source> (<publisher-loc>Hissar</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>56</fpage>&#x02013;<lpage>62</lpage>.</citation>
</ref>
<ref id="B72">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>B.</given-names></name> <name><surname>Paul</surname> <given-names>Y.</given-names></name> <name><surname>Jaswal</surname> <given-names>R. A.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Development of emg controlled electric wheelchair using svm and knn classifier for sci patients,&#x0201D;</article-title> in <source>International Conference on Advanced Informatics for Computing Research</source> (<publisher-loc>Shimla</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>75</fpage>&#x02013;<lpage>83</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-15-0111-1_8</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kurian</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>A review on technological development of automatic speech recognition</article-title>. <source>Int. J. Soft Comput. Eng.</source> <volume>4</volume>, <fpage>80</fpage>&#x02013;<lpage>86</lpage>.</citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000FC;t&#x000FC;k</surname> <given-names>M. E.</given-names></name> <name><surname>D&#x000FC;lger</surname> <given-names>L. C.</given-names></name> <name><surname>Da&#x0015F;</surname> <given-names>M. T.</given-names></name></person-group> (<year>2019</year>). <article-title>Design of a robot-assisted exoskeleton for passive wrist and forearm rehabilitation</article-title>. <source>Mech. Sci</source>. <volume>10</volume>, <fpage>107</fpage>&#x02013;<lpage>118</lpage>. <pub-id pub-id-type="doi">10.5194/ms-10-107-2019</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lackey</surname> <given-names>S.</given-names></name> <name><surname>Barber</surname> <given-names>D.</given-names></name> <name><surname>Reinerman</surname> <given-names>L.</given-names></name> <name><surname>Badler</surname> <given-names>N. I.</given-names></name> <name><surname>Hudson</surname> <given-names>I.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Defining next-generation multi-modal communication in human robot interaction,&#x0201D;</article-title> in <source>Proceedings of the Human Factors and Ergonomics Society Annual Meeting</source>, <volume>Vol. 55</volume> (<publisher-loc>Los Angeles, CA</publisher-loc>: <publisher-name>SAGE Publications</publisher-name>), <fpage>461</fpage>&#x02013;<lpage>464</lpage>. <pub-id pub-id-type="doi">10.1177/1071181311551095</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lannoy</surname> <given-names>S.</given-names></name> <name><surname>Dormal</surname> <given-names>V.</given-names></name> <name><surname>Brion</surname> <given-names>M.</given-names></name> <name><surname>Billieux</surname> <given-names>J.</given-names></name> <name><surname>Maurage</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Preserved crossmodal integration of emotional signals in binge drinking</article-title>. <source>Front. Psychol</source>. <volume>8</volume>, <fpage>984</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2017.00984</pub-id><pub-id pub-id-type="pmid">28663732</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lawson</surname> <given-names>B. E.</given-names></name> <name><surname>Mitchell</surname> <given-names>J.</given-names></name> <name><surname>Truex</surname> <given-names>D.</given-names></name> <name><surname>Shultz</surname> <given-names>A.</given-names></name> <name><surname>Ledoux</surname> <given-names>E.</given-names></name> <name><surname>Goldfarb</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>A robotic leg prosthesis: design, control, and implementation</article-title>. <source>IEEE Robot. Autom. Mag</source>. <volume>21</volume>, <fpage>70</fpage>&#x02013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1109/MRA.2014.2360303</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Legrand</surname> <given-names>M.</given-names></name> <name><surname>Merad</surname> <given-names>M.</given-names></name> <name><surname>De Montalivet</surname> <given-names>E.</given-names></name> <name><surname>Roby-Brami</surname> <given-names>A.</given-names></name> <name><surname>Jarrass&#x000E9;</surname> <given-names>N.</given-names></name></person-group> (<year>2018</year>). <article-title>Movement-based control for upper-limb prosthetics: is the regression technique the key to a robust and accurate control?</article-title> <source>Front. Neurorobot</source>. <volume>12</volume>, <fpage>41</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2018.00041</pub-id><pub-id pub-id-type="pmid">30093857</pub-id></citation></ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>P.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Common sensors in industrial robots: a review</article-title>. <source>J. Phys. Conf. Ser</source>. <volume>1267</volume>, <fpage>012036</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/1267/1/012036</pub-id></citation>
</ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>Implicit intention communication in human-robot interaction through visual behavior studies</article-title>. <source>IEEE Trans. Hum. Mach. Syst</source>. <volume>47</volume>, <fpage>437</fpage>&#x02013;<lpage>448</lpage>. <pub-id pub-id-type="doi">10.1109/THMS.2017.2647882</pub-id></citation>
</ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Tang</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>Multi-modal perception attention network with self-supervised learning for audio-visual speaker tracking</article-title>. <source>Proc. AAAI Conf. Artif. Intell</source>. <volume>36</volume>, <fpage>1456</fpage>&#x02013;<lpage>1463</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v36i2.20035</pub-id></citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Vincent Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2022</year>). <article-title>Multimodal data-driven robot control for human-robot collaborative assembly</article-title>. <source>J. Manuf. Sci. Eng</source>. <volume>144</volume>, <fpage>051012</fpage>. <pub-id pub-id-type="doi">10.1115/1.4053806</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.-T.</given-names></name> <name><surname>Pan</surname> <given-names>F.-F.</given-names></name> <name><surname>Wu</surname> <given-names>M.</given-names></name> <name><surname>Cao</surname> <given-names>W.-H.</given-names></name> <name><surname>Chen</surname> <given-names>L.-F.</given-names></name> <name><surname>Xu</surname> <given-names>J.-P.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>&#x0201C;A multimodal emotional communication based humans-robots interaction system,&#x0201D;</article-title> in <source>2016 35th Chinese Control Conference (CCC)</source> (<publisher-loc>Chengdu</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6363</fpage>&#x02013;<lpage>6368</lpage>. <pub-id pub-id-type="doi">10.1109/ChiCC.2016.7554357</pub-id></citation>
</ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Loth</surname> <given-names>S.</given-names></name> <name><surname>Jettka</surname> <given-names>K.</given-names></name> <name><surname>Giuliani</surname> <given-names>M.</given-names></name> <name><surname>De Ruiter</surname> <given-names>J. P.</given-names></name></person-group> (<year>2015</year>). <article-title>Ghost-in-the-machine reveals human social signals for human-robot interaction</article-title>. <source>Front. Psychol</source>. <volume>6</volume>, <fpage>1641</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2015.01641</pub-id><pub-id pub-id-type="pmid">26582998</pub-id></citation></ref>
<ref id="B85">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>R. C.</given-names></name> <name><surname>Chang</surname> <given-names>S.-R.</given-names></name> <name><surname>Huang</surname> <given-names>C.-C.</given-names></name> <name><surname>Yang</surname> <given-names>Y.-P.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Human robot interactions using speech synthesis and recognition with lip synchronization,&#x0201D;</article-title> in <source>IECON 2011-37th Annual Conference of the IEEE Industrial Electronics Society</source> (<publisher-loc>Melbourne, VIC</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>171</fpage>&#x02013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1109/IECON.2011.6119307</pub-id></citation>
</ref>
<ref id="B86">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Malinovsk&#x000E1;</surname> <given-names>K.</given-names></name> <name><surname>Farka&#x00161;</surname> <given-names>I.</given-names></name> <name><surname>Harvanov&#x000E1;</surname> <given-names>J.</given-names></name> <name><surname>Hoffmann</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;A connectionist model of associating proprioceptive and tactile modalities in a humanoid robot,&#x0201D;</article-title> in <source>2022 IEEE International Conference on Development and Learning (ICDL)</source> (<publisher-loc>London</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>336</fpage>&#x02013;<lpage>342</lpage>. <pub-id pub-id-type="doi">10.1109/ICDL53763.2022.9962195</pub-id></citation>
</ref>
<ref id="B87">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maniscalco</surname> <given-names>U.</given-names></name> <name><surname>Storniolo</surname> <given-names>P.</given-names></name> <name><surname>Messina</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Bidirectional multi-modal signs of checking human-robot engagement and interaction</article-title>. <source>Int. J. Soc. Robot</source>. <volume>14</volume>, <fpage>1295</fpage>&#x02013;<lpage>1309</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-021-00855-w</pub-id></citation>
</ref>
<ref id="B88">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Manna</surname> <given-names>S. K.</given-names></name> <name><surname>Bhaumik</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>A bioinspired 10 dof wearable powered arm exoskeleton for rehabilitation</article-title>. <source>J. Robot</source>. <volume>2013</volume>, <fpage>741359</fpage>. <pub-id pub-id-type="doi">10.1155/2013/741359</pub-id></citation>
</ref>
<ref id="B89">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maroto-G&#x000F3;mez</surname> <given-names>M.</given-names></name> <name><surname>Marqu&#x000E9;s-Villaroya</surname> <given-names>S.</given-names></name> <name><surname>Castillo</surname> <given-names>J. C.</given-names></name> <name><surname>Castro-Gonz&#x000E1;lez</surname> <given-names>&#x000C1;.</given-names></name> <name><surname>Malfaz</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>Active learning based on computer vision and human-robot interaction for the user profiling and behavior personalization of an autonomous social robot</article-title>. <source>Eng. Appl. Artif. Intell</source>. <volume>117</volume>, <fpage>105631</fpage>. <pub-id pub-id-type="doi">10.1016/j.engappai.2022.105631</pub-id></citation>
</ref>
<ref id="B90">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Masteller</surname> <given-names>A.</given-names></name> <name><surname>Sankar</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>H. B.</given-names></name> <name><surname>Ding</surname> <given-names>K.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>All</surname> <given-names>A. H.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Recent developments in prosthesis sensors, texture recognition, and sensory stimulation for upper limb prostheses</article-title>. <source>Ann. Biomed. Eng</source>. <volume>49</volume>, <fpage>57</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1007/s10439-020-02678-8</pub-id><pub-id pub-id-type="pmid">33140242</pub-id></citation></ref>
<ref id="B91">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mead</surname> <given-names>R.</given-names></name> <name><surname>Mataric</surname> <given-names>M. J.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;A probabilistic framework for autonomous proxemic control in situated and mobile human-robot interaction,&#x0201D;</article-title> in <source>Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>193</fpage>&#x02013;<lpage>194</lpage>. <pub-id pub-id-type="doi">10.1145/2157689.2157751</pub-id></citation>
</ref>
<ref id="B92">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mead</surname> <given-names>R.</given-names></name> <name><surname>Matari&#x00107;</surname> <given-names>M. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Autonomous human-robot proxemics: socially aware navigation based on interaction potential</article-title>. <source>Auton. Robots</source> <volume>41</volume>, <fpage>1189</fpage>&#x02013;<lpage>1201</lpage>. <pub-id pub-id-type="doi">10.1007/s10514-016-9572-2</pub-id></citation>
</ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mitra</surname> <given-names>S.</given-names></name> <name><surname>Acharya</surname> <given-names>T.</given-names></name></person-group> (<year>2007</year>). <article-title>Gesture recognition: a survey</article-title>. <source>IEEE Trans. Syst. Man Cybern. Part C Appl. Rev</source>. <volume>37</volume>, <fpage>311</fpage>&#x02013;<lpage>324</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCC.2007.893280</pub-id></citation>
</ref>
<ref id="B94">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mocan</surname> <given-names>B.</given-names></name> <name><surname>Mocan</surname> <given-names>M.</given-names></name> <name><surname>Fulea</surname> <given-names>M.</given-names></name> <name><surname>Murar</surname> <given-names>M.</given-names></name> <name><surname>Feier</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>Home-based robotic upper limbs cardiac telerehabilitation system</article-title>. <source>Int. J. Environ. Res. Public Health</source> <volume>19</volume>, <fpage>11628</fpage>. <pub-id pub-id-type="doi">10.3390/ijerph191811628</pub-id><pub-id pub-id-type="pmid">36141899</pub-id></citation></ref>
<ref id="B95">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohebbi</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Human-robot interaction in rehabilitation and assistance: a review</article-title>. <source>Curr. Robot. Rep</source>. <volume>1</volume>, <fpage>131</fpage>&#x02013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1007/s43154-020-00015-4</pub-id><pub-id pub-id-type="pmid">31287340</pub-id></citation></ref>
<ref id="B96">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Moroto</surname> <given-names>Y.</given-names></name> <name><surname>Maeda</surname> <given-names>K.</given-names></name> <name><surname>Ogawa</surname> <given-names>T.</given-names></name> <name><surname>Haseyama</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Human emotion recognition using multi-modal biological signals based on time lag-considered correlation maximization,&#x0201D;</article-title> in <source>ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4683</fpage>&#x02013;<lpage>4687</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP43922.2022.9746128</pub-id></citation>
</ref>
<ref id="B97">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nagahanumaiah</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <source>Multi-modal Human Fatigue Classification using Wearable Sensors for Human-Robot Teams</source> [PhD thesis]. <publisher-loc>Rochester, NY</publisher-loc>: <publisher-name>Rochester Institute of Technology</publisher-name>. <pub-id pub-id-type="doi">10.1109/SOSE55472.2022.9812694</pub-id></citation>
</ref>
<ref id="B98">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Navarro</surname> <given-names>S. E.</given-names></name> <name><surname>Hein</surname> <given-names>B.</given-names></name> <name><surname>W&#x000F6;rn</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Capacitive tactile proximity sensing: from signal processing to applications in manipulation and safe human-robot interaction,&#x0201D;</article-title> in <source>Soft Robotics</source>, eds <person-group person-group-type="editor"><name><surname>Verl</surname> <given-names>A.</given-names></name> <name><surname>Albu-Sch&#x000E4;ffer</surname> <given-names>A</given-names></name> <name><surname>Brock</surname> <given-names>O.</given-names></name> <name><surname>Raatz</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>54</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-662-44506-8_6</pub-id></citation>
</ref>
<ref id="B99">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>O&#x00027;Neill</surname> <given-names>J.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Dockter</surname> <given-names>R.</given-names></name> <name><surname>Kowalewski</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Practical, stretchable smart skin sensors for contact-aware robots in safe and collaborative interactions,&#x0201D;</article-title> in <source>2015 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>624</fpage>&#x02013;<lpage>629</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA.2015.7139244</pub-id></citation>
</ref>
<ref id="B100">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ot&#x000E1;lora</surname> <given-names>S.</given-names></name> <name><surname>Ballen-Moreno</surname> <given-names>F.</given-names></name> <name><surname>Arciniegas-Mayag</surname> <given-names>L.</given-names></name> <name><surname>Cifuentes</surname> <given-names>C. A.</given-names></name> <name><surname>M&#x000FA;nera</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Biomechanical effects of adding an ankle soft actuation in a unilateral exoskeleton</article-title>. <source>Biosensors</source> <volume>12</volume>, <fpage>873</fpage>. <pub-id pub-id-type="doi">10.3390/bios12100873</pub-id><pub-id pub-id-type="pmid">36291010</pub-id></citation></ref>
<ref id="B101">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Page</surname> <given-names>M. J.</given-names></name> <name><surname>McKenzie</surname> <given-names>J. E.</given-names></name> <name><surname>Bossuyt</surname> <given-names>P. M.</given-names></name> <name><surname>Boutron</surname> <given-names>I.</given-names></name> <name><surname>Hoffmann</surname> <given-names>T. C.</given-names></name> <name><surname>Mulrow</surname> <given-names>C. D.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>The prisma 2020 statement: an updated guideline for reporting systematic reviews</article-title>. <source>Syst. Rev</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1186/s13643-021-01626-4</pub-id><pub-id pub-id-type="pmid">34446261</pub-id></citation></ref>
<ref id="B102">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pantic</surname> <given-names>M.</given-names></name> <name><surname>Rothkrantz</surname> <given-names>L. J.</given-names></name></person-group> (<year>2000</year>). <article-title>Expert system for automatic analysis of facial expressions</article-title>. <source>Image Vis. Comput</source>. <volume>18</volume>, <fpage>881</fpage>&#x02013;<lpage>905</lpage>. <pub-id pub-id-type="doi">10.1016/S0262-8856(00)00034-2</pub-id><pub-id pub-id-type="pmid">34756219</pub-id></citation></ref>
<ref id="B103">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pawu&#x0015B;</surname> <given-names>D.</given-names></name> <name><surname>Paszkiel</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>BCI wheelchair control using expert system classifying EEG signals based on power spectrum estimation and nervous tics detection</article-title>. <source>Appl. Sci</source>. <volume>12</volume>, <fpage>10385</fpage>. <pub-id pub-id-type="doi">10.3390/app122010385</pub-id></citation>
</ref>
<ref id="B104">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Popov</surname> <given-names>D.</given-names></name> <name><surname>Klimchik</surname> <given-names>A.</given-names></name> <name><surname>Mavridis</surname> <given-names>N.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Collision detection, localization &#x00026;classification for industrial robots with joint torque sensors,&#x0201D;</article-title> in <source>2017 26th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)</source> (<publisher-loc>Lisbon</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>838</fpage>&#x02013;<lpage>843</lpage>. <pub-id pub-id-type="doi">10.1109/ROMAN.2017.8172400</pub-id></citation>
</ref>
<ref id="B105">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pyo</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Bae</surname> <given-names>K.</given-names></name> <name><surname>Sim</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Recent progress in flexible tactile sensors for human-interactive systems: from sensors to advanced applications</article-title>. <source>Adv. Mater</source>. <volume>33</volume>, <fpage>2005902</fpage>. <pub-id pub-id-type="doi">10.1002/adma.202005902</pub-id><pub-id pub-id-type="pmid">33887803</pub-id></citation></ref>
<ref id="B106">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rabhi</surname> <given-names>Y.</given-names></name> <name><surname>Mrabet</surname> <given-names>M.</given-names></name> <name><surname>Fnaiech</surname> <given-names>F.</given-names></name></person-group> (<year>2018a</year>). <article-title>A facial expression controlled wheelchair for people with disabilities</article-title>. <source>Comput. Methods Programs Biomed</source>. <volume>165</volume>, <fpage>89</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1016/j.cmpb.2018.08.013</pub-id><pub-id pub-id-type="pmid">30337084</pub-id></citation></ref>
<ref id="B107">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rabhi</surname> <given-names>Y.</given-names></name> <name><surname>Mrabet</surname> <given-names>M.</given-names></name> <name><surname>Fnaiech</surname> <given-names>F.</given-names></name></person-group> (<year>2018b</year>). <article-title>Intelligent control wheelchair using a new visual joystick</article-title>. <source>J. Healthc. Eng</source>. <volume>2018</volume>, <fpage>6083565</fpage>. <pub-id pub-id-type="doi">10.1155/2018/6083565</pub-id><pub-id pub-id-type="pmid">29599953</pub-id></citation></ref>
<ref id="B108">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rasouli</surname> <given-names>N.</given-names></name> <name><surname>Trott</surname> <given-names>S.</given-names></name> <name><surname>Soto</surname> <given-names>V.</given-names></name> <name><surname>Alonso</surname> <given-names>R. E.</given-names></name></person-group> (<year>2018</year>). <article-title>Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems</article-title>. <source>ACL</source> <volume>2018</volume>, <fpage>189</fpage>&#x02013;<lpage>198</lpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1804.06512</pub-id></citation>
</ref>
<ref id="B109">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rautaray</surname> <given-names>S. S.</given-names></name> <name><surname>Agrawal</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Vision based hand gesture recognition for human computer interaction: a survey</article-title>. <source>Artif. Intell. Rev</source>. <volume>43</volume>, <fpage>1</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-012-9356-9</pub-id></citation>
</ref>
<ref id="B110">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Redmon</surname> <given-names>J.</given-names></name> <name><surname>Divvala</surname> <given-names>S.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Farhadi</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;You only look once: unified, real-time object detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>779</fpage>&#x02013;<lpage>788</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.91</pub-id></citation>
</ref>
<ref id="B111">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Reis</surname> <given-names>L. P.</given-names></name> <name><surname>Faria</surname> <given-names>B. M.</given-names></name> <name><surname>Vasconcelos</surname> <given-names>S.</given-names></name> <name><surname>Lau</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Multimodal interface for an intelligent wheelchair,&#x0201D;</article-title> in <source>Informatics in Control, Automation and Robotics</source>, eds <person-group person-group-type="editor"><name><surname>Ferrier</surname> <given-names>J. L.</given-names></name> <name><surname>Gusikhin</surname> <given-names>O.</given-names></name> <name><surname>Madani</surname> <given-names>K.</given-names></name> <name><surname>Sasiadek</surname> <given-names>J.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-10891-9_1</pub-id><pub-id pub-id-type="pmid">23923691</pub-id></citation></ref>
<ref id="B112">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rincon</surname> <given-names>J. A.</given-names></name> <name><surname>Costa</surname> <given-names>A.</given-names></name> <name><surname>Novais</surname> <given-names>P.</given-names></name> <name><surname>Julian</surname> <given-names>V.</given-names></name> <name><surname>Carrascosa</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>A new emotional robot assistant that facilitates human interaction and persuasion</article-title>. <source>Knowl. Inf. Syst</source>. <volume>60</volume>, <fpage>363</fpage>&#x02013;<lpage>383</lpage>. <pub-id pub-id-type="doi">10.1007/s10115-018-1231-9</pub-id></citation>
</ref>
<ref id="B113">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rodomagoulakis</surname> <given-names>I.</given-names></name> <name><surname>Kardaris</surname> <given-names>N.</given-names></name> <name><surname>Pitsikalis</surname> <given-names>V.</given-names></name> <name><surname>Mavroudi</surname> <given-names>E.</given-names></name> <name><surname>Katsamanis</surname> <given-names>A.</given-names></name> <name><surname>Tsiami</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>&#x0201C;Multimodal human action recognition in assistive human-robot interaction,&#x0201D;</article-title> in <source>2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source> (<publisher-loc>Shanghai</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2702</fpage>&#x02013;<lpage>2706</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.2016.7472168</pub-id><pub-id pub-id-type="pmid">34887740</pub-id></citation></ref>
<ref id="B114">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rossi</surname> <given-names>S.</given-names></name> <name><surname>Larafa</surname> <given-names>M.</given-names></name> <name><surname>Ruocco</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Emotional and behavioural distraction by a social robot for children anxiety reduction during vaccination</article-title>. <source>Int. J. Soc. Robot</source>. <volume>12</volume>, <fpage>765</fpage>&#x02013;<lpage>777</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-019-00616-w</pub-id></citation>
</ref>
<ref id="B115">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Salem</surname> <given-names>M.</given-names></name> <name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>I.</given-names></name> <name><surname>Joublin</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Towards an integrated model of speech and gesture production for multi-modal robot behavior,&#x0201D;</article-title> in <source>19th International Symposium in Robot and Human Interactive Communication</source> (<publisher-loc>Viareggio</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>614</fpage>&#x02013;<lpage>619</lpage>. <pub-id pub-id-type="doi">10.1109/ROMAN.2010.5598665</pub-id></citation>
</ref>
<ref id="B116">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salem</surname> <given-names>M.</given-names></name> <name><surname>Kopp</surname> <given-names>S.</given-names></name> <name><surname>Wachsmuth</surname> <given-names>I.</given-names></name> <name><surname>Rohlfing</surname> <given-names>K.</given-names></name> <name><surname>Joublin</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>). <article-title>Generation and evaluation of communicative robot gesture</article-title>. <source>Int. J. Soc. Robot</source>. <volume>4</volume>, <fpage>201</fpage>&#x02013;<lpage>217</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-011-0124-9</pub-id></citation>
</ref>
<ref id="B117">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Salovey</surname> <given-names>P.</given-names></name> <name><surname>Mayer</surname> <given-names>J. D.</given-names></name></person-group> (<year>2004</year>). <source>Emotional Intelligence</source>. <publisher-loc>Nashville, TN</publisher-loc>: <publisher-name>Dude publishing</publisher-name>. <pub-id pub-id-type="doi">10.1017/CBO9780511806582.019</pub-id></citation>
</ref>
<ref id="B118">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sasaki</surname> <given-names>K.</given-names></name> <name><surname>Guerra</surname> <given-names>G.</given-names></name> <name><surname>Lei Phyu</surname> <given-names>W.</given-names></name> <name><surname>Chaisumritchoke</surname> <given-names>S.</given-names></name> <name><surname>Sutdet</surname> <given-names>P.</given-names></name> <name><surname>Kaewtip</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Assessment of socket pressure during walking in rapid fit prosthetic sockets</article-title>. <source>Sensors</source> <volume>22</volume>, <fpage>5224</fpage>. <pub-id pub-id-type="doi">10.3390/s22145224</pub-id><pub-id pub-id-type="pmid">35890905</pub-id></citation></ref>
<ref id="B119">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saunderson</surname> <given-names>S.</given-names></name> <name><surname>Nejat</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>How robots influence humans: a survey of nonverbal communication in social human-robot interaction</article-title>. <source>Int. J. Soc. Robot</source>. <volume>11</volume>, <fpage>575</fpage>&#x02013;<lpage>608</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-019-00523-0</pub-id></citation>
</ref>
<ref id="B120">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scalise</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Admoni</surname> <given-names>H.</given-names></name> <name><surname>Rosenthal</surname> <given-names>S.</given-names></name> <name><surname>Srinivasa</surname> <given-names>S. S.</given-names></name></person-group> (<year>2018</year>). <article-title>Natural language instructions for human-robot collaborative manipulation</article-title>. <source>Int. J. Rob. Res</source>. <volume>37</volume>, <fpage>558</fpage>&#x02013;<lpage>565</lpage>. <pub-id pub-id-type="doi">10.1177/0278364918760992</pub-id></citation>
</ref>
<ref id="B121">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schreiter</surname> <given-names>T.</given-names></name> <name><surname>de Almeida</surname> <given-names>T. R.</given-names></name> <name><surname>Zhu</surname> <given-names>Y</given-names></name> <name><surname>Maestro</surname> <given-names>E. G.</given-names></name> <name><surname>Morillo-Mendez</surname> <given-names>L.</given-names></name> <name><surname>Rudenko</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>The magni human motion dataset: accurate, complex, multi-modal, natural, semantically-rich and contextualized</article-title>. <source>arXiv</source>. [preprint]. <pub-id pub-id-type="doi">10.48550/arXiv.2208.14925</pub-id></citation>
</ref>
<ref id="B122">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schroff</surname> <given-names>F.</given-names></name> <name><surname>Kalenichenko</surname> <given-names>D.</given-names></name> <name><surname>Philbin</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Facenet: a unified embedding for face recognition and clustering,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer VISION and Pattern Recognition</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>815</fpage>&#x02013;<lpage>823</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298682</pub-id></citation>
</ref>
<ref id="B123">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwesinger</surname> <given-names>D.</given-names></name> <name><surname>Shariati</surname> <given-names>A.</given-names></name> <name><surname>Montella</surname> <given-names>C.</given-names></name> <name><surname>Spletzer</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>A smart wheelchair ecosystem for autonomous navigation in urban environments</article-title>. <source>Auton. Robots</source> <volume>41</volume>, <fpage>519</fpage>&#x02013;<lpage>538</lpage>. <pub-id pub-id-type="doi">10.1007/s10514-016-9549-1</pub-id></citation>
</ref>
<ref id="B124">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shao</surname> <given-names>M.</given-names></name> <name><surname>Snyder</surname> <given-names>M.</given-names></name> <name><surname>Nejat</surname> <given-names>G.</given-names></name> <name><surname>Benhabib</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>User affect elicitation with a socially emotional robot</article-title>. <source>Robotics</source> <volume>9</volume>, <fpage>44</fpage>. <pub-id pub-id-type="doi">10.3390/robotics9020044</pub-id></citation>
</ref>
<ref id="B125">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sharifuddin</surname> <given-names>M. S. I.</given-names></name> <name><surname>Nordin</surname> <given-names>S.</given-names></name> <name><surname>Ali</surname> <given-names>A. M.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Voice control intelligent wheelchair movement using CNNS,&#x0201D;</article-title> in <source>2019 1st International Conference on Artificial Intelligence and Data Sciences (AiDAS)</source> (<publisher-loc>Ipoh</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>40</fpage>&#x02013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1109/AiDAS47888.2019.8970865</pub-id></citation>
</ref>
<ref id="B126">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shenoy</surname> <given-names>S.</given-names></name> <name><surname>Hou</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Nikseresht</surname> <given-names>F.</given-names></name> <name><surname>Doryab</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Adaptive humanoid robots for pain management in children,&#x0201D;</article-title> in <source>Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>510</fpage>&#x02013;<lpage>514</lpage>. <pub-id pub-id-type="doi">10.1145/3434074.3447224</pub-id></citation>
</ref>
<ref id="B127">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Skubic</surname> <given-names>M.</given-names></name> <name><surname>Perzanowski</surname> <given-names>D.</given-names></name> <name><surname>Blisard</surname> <given-names>S.</given-names></name> <name><surname>Schultz</surname> <given-names>A.</given-names></name> <name><surname>Adams</surname> <given-names>W.</given-names></name> <name><surname>Bugajska</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>Spatial language for human-robot dialogs</article-title>. <source>IEEE Trans. Syst. Man Cybernetics Part C Appl. Rev</source>. <volume>34</volume>, <fpage>154</fpage>&#x02013;<lpage>167</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCC.2004.826273</pub-id></citation>
</ref>
<ref id="B128">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>C.</given-names></name> <name><surname>Croon</surname> <given-names>G. S.</given-names></name> <name><surname>Hjalmarsson</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Gaze-based human-robot communication,&#x0201D;</article-title> in <source>Proceedings of the SIGDIAL 2013 Conference</source> (<publisher-loc>Metz</publisher-loc>), <fpage>104</fpage>&#x02013;<lpage>112</lpage>.</citation>
</ref>
<ref id="B129">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stephens-Fripp</surname> <given-names>B.</given-names></name> <name><surname>Sencadas</surname> <given-names>V.</given-names></name> <name><surname>Mutlu</surname> <given-names>R.</given-names></name> <name><surname>Alici</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>Reusable flexible concentric electrodes coated with a conductive graphene ink for electrotactile stimulation</article-title>. <source>Front. Bioeng. Biotechnol</source>. <volume>6</volume>, <fpage>179</fpage>. <pub-id pub-id-type="doi">10.3389/fbioe.2018.00179</pub-id><pub-id pub-id-type="pmid">30560123</pub-id></citation></ref>
<ref id="B130">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stiefelhagen</surname> <given-names>R.</given-names></name> <name><surname>Ekenel</surname> <given-names>H. K.</given-names></name> <name><surname>Fugen</surname> <given-names>C.</given-names></name> <name><surname>Gieselmann</surname> <given-names>P.</given-names></name> <name><surname>Holzapfel</surname> <given-names>H.</given-names></name> <name><surname>Kraft</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>Enabling multimodal human-robot interaction for the karlsruhe humanoid robot</article-title>. <source>IEEE Trans. Robot</source>. <volume>23</volume>, <fpage>840</fpage>&#x02013;<lpage>851</lpage>. <pub-id pub-id-type="doi">10.1109/TRO.2007.907484</pub-id></citation>
</ref>
<ref id="B131">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stock-Homburg</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>Survey of emotions in human-robot interactions: perspectives from robotic psychology on 20 years of research</article-title>. <source>Int. J. Soc. Robot</source>. <volume>14</volume>, <fpage>389</fpage>&#x02013;<lpage>411</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-021-00778-6</pub-id></citation>
</ref>
<ref id="B132">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Strazdas</surname> <given-names>D.</given-names></name> <name><surname>Hintz</surname> <given-names>J.</given-names></name> <name><surname>Khalifa</surname> <given-names>A.</given-names></name> <name><surname>Abdelrahman</surname> <given-names>A. A.</given-names></name> <name><surname>Hempel</surname> <given-names>T.</given-names></name> <name><surname>Al-Hamadi</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Robot system assistant (ROSA): towards intuitive multi-modal and multi-device human-robot interaction</article-title>. <source>Sensors</source> <volume>22</volume>, <fpage>923</fpage>. <pub-id pub-id-type="doi">10.3390/s22030923</pub-id><pub-id pub-id-type="pmid">35161671</pub-id></citation></ref>
<ref id="B133">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x00160;vec</surname> <given-names>J.</given-names></name> <name><surname>Neduchal</surname> <given-names>P.</given-names></name> <name><surname>Hr&#x000FA;z</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>Multi-modal communication system for mobile robot</article-title>. <source>IFAC-Pap</source>. <volume>55</volume>, <fpage>133</fpage>&#x02013;<lpage>138</lpage>. <pub-id pub-id-type="doi">10.1016/j.ifacol.2022.06.022</pub-id></citation>
</ref>
<ref id="B134">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>G.</given-names></name> <name><surname>Asif</surname> <given-names>S.</given-names></name> <name><surname>Webb</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>The integration of contactless static pose recognition and dynamic hand motion tracking control system for industrial human and robot collaboration</article-title>. <source>Ind. Robot</source>. <volume>42</volume>, <fpage>416</fpage>&#x02013;<lpage>428</lpage>. <pub-id pub-id-type="doi">10.1108/IR-03-2015-0059</pub-id></citation>
</ref>
<ref id="B135">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tatarian</surname> <given-names>K.</given-names></name> <name><surname>Stower</surname> <given-names>R.</given-names></name> <name><surname>Rudaz</surname> <given-names>D.</given-names></name> <name><surname>Chamoux</surname> <given-names>M.</given-names></name> <name><surname>Kappas</surname> <given-names>A.</given-names></name> <name><surname>Chetouani</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>How does modality matter? investigating the synthesis and effects of multi-modal robot behavior on social intelligence</article-title>. <source>Int. J. Soc. Robot</source>. <volume>14</volume>, <fpage>893</fpage>&#x02013;<lpage>911</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-021-00839-w</pub-id></citation>
</ref>
<ref id="B136">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thomas</surname> <given-names>M.</given-names></name> <name><surname>Collier</surname> <given-names>J.</given-names></name> <name><surname>Monckton</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <source>Multi-modal Human-robot Interaction</source>.</citation>
</ref>
<ref id="B137">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Qiao</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Computer vision technology in agricultural automation&#x02013;a review</article-title>. <source>Inf. Proces. Agric</source>. <volume>7</volume>, <fpage>1</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1016/j.inpa.2019.09.006</pub-id></citation>
</ref>
<ref id="B138">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Treussart</surname> <given-names>B.</given-names></name> <name><surname>Geffard</surname> <given-names>F.</given-names></name> <name><surname>Vignais</surname> <given-names>N.</given-names></name> <name><surname>Marin</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Controlling an upper-limb exoskeleton by emg signal while carrying unknown load,&#x0201D;</article-title> in <source>2020 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Paris</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>9107</fpage>&#x02013;<lpage>9113</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA40945.2020.9197087</pub-id><pub-id pub-id-type="pmid">29060416</pub-id></citation></ref>
<ref id="B139">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tsiami</surname> <given-names>A.</given-names></name> <name><surname>Filntisis</surname> <given-names>P. P.</given-names></name> <name><surname>Efthymiou</surname> <given-names>N.</given-names></name> <name><surname>Koutras</surname> <given-names>P.</given-names></name> <name><surname>Potamianos</surname> <given-names>G.</given-names></name> <name><surname>Maragos</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Far-field audio-visual scene perception of multi-party human-robot interaction for children and adults,&#x0201D;</article-title> in <source>2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source> (<publisher-loc>Calgary, AB</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6568</fpage>&#x02013;<lpage>6572</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.2018.8462425</pub-id></citation>
</ref>
<ref id="B140">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tuli</surname> <given-names>T. B.</given-names></name> <name><surname>Kohl</surname> <given-names>L.</given-names></name> <name><surname>Chala</surname> <given-names>S. A.</given-names></name> <name><surname>Manns</surname> <given-names>M.</given-names></name> <name><surname>Ansari</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Knowledge-based digital twin for predicting interactions in human-robot collaboration,&#x0201D;</article-title> in <source>2021 26th IEEE International Conference on Emerging Technologies and Factory Automation (ETFA)</source> (<publisher-loc>Vasteras</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/ETFA45728.2021.9613342</pub-id></citation>
</ref>
<ref id="B141">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tziafas</surname> <given-names>G.</given-names></name> <name><surname>Kasaei</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Few-shot visual grounding for natural human-robot interaction,&#x0201D;</article-title> in <source>2021 IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC</source>) (<publisher-loc>Santa Maria da Feira</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>50</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1109/ICARSC52212.2021.9429801</pub-id></citation>
</ref>
<ref id="B142">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ur Rehman</surname> <given-names>M.</given-names></name> <name><surname>Ahmed</surname> <given-names>F.</given-names></name> <name><surname>Attique Khan</surname> <given-names>M.</given-names></name> <name><surname>Tariq</surname> <given-names>U.</given-names></name> <name><surname>Abdulaziz Alfouzan</surname> <given-names>F.</given-names></name> <name><surname>Alzahrani</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Dynamic hand gesture recognition using 3d-cnn and lstm networks</article-title>. <source>Comput. Mater. Contin.</source> <volume>70</volume>, <fpage>4675</fpage>&#x02013;<lpage>4690</lpage>. <pub-id pub-id-type="doi">10.32604/cmc.2022.019586</pub-id></citation>
</ref>
<ref id="B143">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wachaja</surname> <given-names>A.</given-names></name> <name><surname>Agarwal</surname> <given-names>P.</given-names></name> <name><surname>Zink</surname> <given-names>M.</given-names></name> <name><surname>Adame</surname> <given-names>M. R.</given-names></name> <name><surname>M&#x000F6;ller</surname> <given-names>K.</given-names></name> <name><surname>Burgard</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>Navigating blind people with walking impairments using a smart walker</article-title>. <source>Auton. Robots</source> <volume>41</volume>, <fpage>555</fpage>&#x02013;<lpage>573</lpage>. <pub-id pub-id-type="doi">10.1007/s10514-016-9595-8</pub-id></citation>
</ref>
<ref id="B144">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>R.</given-names></name> <name><surname>Jo</surname> <given-names>W.</given-names></name> <name><surname>Zhao</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Yang</surname> <given-names>B.</given-names></name> <name><surname>Chen</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Husformer: a multi-modal transformer for multi-modal human state recognition</article-title>. <source>arXiv</source>. [preprint]. <pub-id pub-id-type="doi">10.48550/arXiv.2209.15182</pub-id></citation>
</ref>
<ref id="B145">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Yuan</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name></person-group> (<year>2022</year>). <article-title>Motion intensity modeling and trajectory control of upper limb rehabilitation exoskeleton robot based on multi-modal information</article-title>. <source>Complex Intell. Syst</source>. <volume>8</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1007/s40747-021-00632-2</pub-id></citation>
</ref>
<ref id="B146">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>Multi-modal broad learning for material recognition</article-title>. <source>Cogn. Comput. Syst</source>. <volume>3</volume>, <fpage>123</fpage>&#x02013;<lpage>130</lpage>. <pub-id pub-id-type="doi">10.1049/ccs2.12004</pub-id></citation>
</ref>
<ref id="B147">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weerakoon</surname> <given-names>D.</given-names></name> <name><surname>Subbaraju</surname> <given-names>V.</given-names></name> <name><surname>Tran</surname> <given-names>T.</given-names></name> <name><surname>Misra</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Cosm2ic: optimizing real-time multi-modal instruction comprehension</article-title>. <source>IEEE Robot. Autom. Lett</source>. <volume>7</volume>, <fpage>10697</fpage>&#x02013;<lpage>10704</lpage>. <pub-id pub-id-type="doi">10.1109/LRA.2022.3194683</pub-id></citation>
</ref>
<ref id="B148">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Whitney</surname> <given-names>D.</given-names></name> <name><surname>Rosen</surname> <given-names>E.</given-names></name> <name><surname>MacGlashan</surname> <given-names>J.</given-names></name> <name><surname>Wong</surname> <given-names>L. L.</given-names></name> <name><surname>Tellex</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Reducing errors in object-fetching interactions through social feedback,&#x0201D;</article-title> in <source>2017 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1006</fpage>&#x02013;<lpage>1013</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA.2017.7989121</pub-id></citation>
</ref>
<ref id="B149">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>E.</given-names></name> <name><surname>Sun</surname> <given-names>P.</given-names></name> <name><surname>Song</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Liang</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;Polarmask: single shot instance segmentation with polar representation,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>12193</fpage>&#x02013;<lpage>12202</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01221</pub-id><pub-id pub-id-type="pmid">33989151</pub-id></citation></ref>
<ref id="B150">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yadav</surname> <given-names>S. K.</given-names></name> <name><surname>Tiwari</surname> <given-names>K.</given-names></name> <name><surname>Pandey</surname> <given-names>H. M.</given-names></name> <name><surname>Akbar</surname> <given-names>S. A.</given-names></name></person-group> (<year>2021</year>). <article-title>A review of multimodal human activity recognition with special emphasis on classification, applications, challenges and future directions</article-title>. <source>Knowl. Based Syst</source>. <volume>223</volume>, <fpage>106970</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2021.106970</pub-id></citation>
</ref>
<ref id="B151">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>D.</given-names></name> <name><surname>Huang</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Contextual and cross-modal interaction for multi-modal speech emotion recognition</article-title>. <source>IEEE Signal Proces. Lett</source>. <volume>29</volume>, <fpage>2093</fpage>&#x02013;<lpage>2097</lpage>. <pub-id pub-id-type="doi">10.1109/LSP.2022.3210836</pub-id></citation>
</ref>
<ref id="B152">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>Q.</given-names></name></person-group> (<year>2016</year>). <source>Multi-sensory Emotion Recognition with Speech and Facial Expression</source>.</citation>
</ref>
<ref id="B153">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yongda</surname> <given-names>D.</given-names></name> <name><surname>Fang</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>Research on multimodal human-robot interaction based on speech and gesture</article-title>. <source>Comput. Electr. Eng</source>. <volume>72</volume>, <fpage>443</fpage>&#x02013;<lpage>454</lpage>. <pub-id pub-id-type="doi">10.1016/j.compeleceng.2018.09.014</pub-id></citation>
</ref>
<ref id="B154">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoon</surname> <given-names>H. U.</given-names></name> <name><surname>Wang</surname> <given-names>R. F.</given-names></name> <name><surname>Hutchinson</surname> <given-names>S. A.</given-names></name> <name><surname>Hur</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Customizing haptic and visual feedback for assistive human-robot interface and the effects on performance improvement</article-title>. <source>Rob. Auton. Syst</source>. <volume>91</volume>, <fpage>258</fpage>&#x02013;<lpage>269</lpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2017.01.015</pub-id></citation>
</ref>
<ref id="B155">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Q.</given-names></name> <name><surname>Wu</surname> <given-names>L.</given-names></name> <name><surname>Bridwell</surname> <given-names>D. A.</given-names></name> <name><surname>Erhardt</surname> <given-names>E. B.</given-names></name> <name><surname>Du</surname> <given-names>Y.</given-names></name> <name><surname>He</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Building an EEG-fMRI multi-modal brain graph: a concurrent EEG-fMRI study</article-title>. <source>Front. Hum. Neurosci</source>. <volume>10</volume>, <fpage>476</fpage>. <pub-id pub-id-type="doi">10.3389/fnhum.2016.00476</pub-id><pub-id pub-id-type="pmid">27733821</pub-id></citation></ref>
<ref id="B156">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>H.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Construction of multi-modal perception model of communicative robot in non-structural cyber physical system environment based on optimized BT-SVM model</article-title>. <source>Comput. Commun</source>. <volume>181</volume>, <fpage>182</fpage>&#x02013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1016/j.comcom.2021.10.019</pub-id></citation>
</ref>
<ref id="B157">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zgallai</surname> <given-names>W.</given-names></name> <name><surname>Brown</surname> <given-names>J. T.</given-names></name> <name><surname>Ibrahim</surname> <given-names>A.</given-names></name> <name><surname>Mahmood</surname> <given-names>F.</given-names></name> <name><surname>Mohammad</surname> <given-names>K.</given-names></name> <name><surname>Khalfan</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Deep learning ai application to an EEG driven bci smart wheelchair,&#x0201D;</article-title> in <source>2019 Advances in Science and Engineering Technology International Conferences (ASET)</source> (<publisher-loc>Dubai</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/ICASET.2019.8714373</pub-id></citation>
</ref>
<ref id="B158">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Meng</surname> <given-names>G.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Intelligent perception recognition of multi-modal emg signals based on machine learning,&#x0201D;</article-title> in <source>2022 2nd International Conference on Bioinformatics and Intelligent Computing</source> (<publisher-loc>Harbin</publisher-loc>), <fpage>389</fpage>&#x02013;<lpage>396</lpage>. <pub-id pub-id-type="doi">10.1145/3523286.3524576</pub-id></citation>
</ref>
<ref id="B159">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Ji</surname> <given-names>Q.</given-names></name></person-group> (<year>2012</year>). <article-title>Audio-visual tibetan speech recognition based on a deep dynamic bayesian network for natural human robot interaction</article-title>. <source>Int. J. Adv. Robot. Syst</source>. <volume>9</volume>, <fpage>258</fpage>. <pub-id pub-id-type="doi">10.5772/54000</pub-id></citation>
</ref>
<ref id="B160">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zlatintsi</surname> <given-names>A.</given-names></name> <name><surname>Rodomagoulakis</surname> <given-names>I.</given-names></name> <name><surname>Koutras</surname> <given-names>P.</given-names></name> <name><surname>Dometios</surname> <given-names>A.</given-names></name> <name><surname>Pitsikalis</surname> <given-names>V.</given-names></name> <name><surname>Tzafestas</surname> <given-names>C. S.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Multimodal signal processing and learning aspects of human-robot interaction for an assistive bathing robot,&#x0201D;</article-title> in <source>2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source> (<publisher-loc>Calgary, AB</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3171</fpage>&#x02013;<lpage>3175</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.2018.8461568</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>
