<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Robot. AI</journal-id>
<journal-title>Frontiers in Robotics and AI</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Robot. AI</abbrev-journal-title>
<issn pub-type="epub">2296-9144</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1123374</article-id>
<article-id pub-id-type="doi">10.3389/frobt.2023.1123374</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Robotics and AI</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Multi-dimensional task recognition for human-robot teaming: literature review</article-title>
<alt-title alt-title-type="left-running-head">Baskaran and Adams</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/frobt.2023.1123374">10.3389/frobt.2023.1123374</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Baskaran</surname>
<given-names>Prakash</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1149154/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Adams</surname>
<given-names>Julie A.</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/833175/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Collaborative Robotics and Intelligent Systems Institute</institution>, <institution>Oregon State University</institution>, <addr-line>Corvallis</addr-line>, <addr-line>OR</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2117420/overview">Fouad Yacef</ext-link>, Center for Development of Advanced Technologies (CDTA), Algeria</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1898275/overview">Yassine Himeur</ext-link>, University of Dubai, United Arab Emirates</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/158904/overview">Tim Oates</ext-link>, University of Maryland, Baltimore County, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Prakash Baskaran, <email>baskarap@oregonstate.edu</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>10</volume>
<elocation-id>1123374</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>17</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Baskaran and Adams.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Baskaran and Adams</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Human-robot teams collaborating to achieve tasks under various conditions, especially in unstructured, dynamic environments will require robots to adapt autonomously to a human teammate&#x2019;s state. An important element of such adaptation is the robot&#x2019;s ability to infer the human teammate&#x2019;s tasks. Environmentally embedded sensors (e.g., motion capture and cameras) are infeasible in such environments for task recognition, but wearable sensors are a viable task recognition alternative. Human-robot teams will perform a wide variety of composite and atomic tasks, involving multiple activity components (i.e., gross motor, fine-grained motor, tactile, visual, cognitive, speech and auditory) that may occur concurrently. A robot&#x2019;s ability to recognize the human&#x2019;s composite, concurrent tasks is a key requirement for realizing successful teaming. Over a hundred task recognition algorithms across multiple activity components are evaluated based on six criteria: sensitivity, suitability, generalizability, composite factor, concurrency and anomaly awareness. The majority of the reviewed task recognition algorithms are not viable for human-robot teams in unstructured, dynamic environments, as they only detect tasks from a subset of activity components, incorporate non-wearable sensors, and rarely detect composite, concurrent tasks across multiple activity components.</p>
</abstract>
<kwd-group>
<kwd>human-robot teaming</kwd>
<kwd>task recognition</kwd>
<kwd>activity recognition</kwd>
<kwd>wearable sensors</kwd>
<kwd>machine learning</kwd>
</kwd-group>
<contract-sponsor id="cn001">Office of Naval Research<named-content content-type="fundref-id">10.13039/100000006</named-content>
</contract-sponsor>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Human-Robot Interaction</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Human-robot teams (HRTs) are often required to operate in dynamic, unstructured environments, which poses a unique set of task recognition challenges that must be overcome to enable a successful collaboration. Consider a post-tornado disaster response that requires locating, triaging and transporting victims, securing infrastructure, clearing debris, and locating and securing potentially dangerous goods from looting (e.g., weapons at a firearms shop, drugs at a pharmacy). Achieving effective human-robot collaboration in such scenarios requires natural teaming between the human first responders and their robot teammates (e.g., quadrotors or ground vehicles). An important element of such teaming is the robot teammates&#x2019; ability to infer the tasks performed by the human teammates in order to adapt their interactions autonomously based on their human teammates&#x2019; state.</p>
<p>Tasks performed by human teammates can be classified into two categories: 1) Atomic, and 2) Composite. <italic>Atomic tasks</italic> are simple activities that are either short in duration, or involve repetitive actions that cannot be decomposed further. A <italic>Composite task</italic> aggregates multiple atomic actions or activities into a more complex task (<xref ref-type="bibr" rid="B29">Chen K. et al., 2021</xref>). For example, <italic>Clearing a dangerous item</italic> is a composite task comprised of several atomic tasks: a) Scanning the environment for suspicious items, b) Evaluating the threat level, c) Taking a picture, d) Discussing with the incident commander via walkie-talkie, and e) moving on to the next area.</p>
<p>Humans conduct tasks using a breadth of their capabilities, as depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>. Depending on the complexity, the tasks can involve multiple activity components: gross motor, fine-grained motor, tactile, cognitive, visual, speech, and auditory (<xref ref-type="bibr" rid="B67">Heard et al., 2018a</xref>). For example, the <italic>Clearing a dangerous item</italic> task aggregates a <italic>visual</italic> component of scanning the environment to locate the item, a <italic>cognitive</italic> component of evaluating the item&#x2019;s threat level, taking a picture task involves <italic>gross motor</italic>, <italic>fine-grained motor</italic> and <italic>tactile</italic> components, while discussing the next steps with the incident commander via walkie-talkie involves <italic>auditory</italic> and <italic>speech</italic> components for listening to queries and providing information.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>An example of a multi-dimensional task recognition framework enabling a robot to detect a human teammate&#x2019;s tasks across all activity components when the HRT is clearing a dangerous item.</p>
</caption>
<graphic xlink:href="frobt-10-1123374-g001.tif"/>
</fig>
<p>HRTs often perform a wide variety of tasks, such that the set of tasks performed by human teammates may involve differing combinations of multi-dimensional activity components. The focus on disaster response environments, which are dynamic, uncertain, and unstructured, does not permit the use of environmentally embedded sensors (e.g., motion capture systems and cameras). Wearable sensors are a viable alternative that facilitate gathering the objective data necessary to identify a human&#x2019;s task. Robots will require a multi-dimensional task recognition algorithm capable of detecting composite tasks composed of differing combinations of the activity components. Current state-of-the-art HRT task recognition using wearable sensors generally focuses on gross motor and some fine-grained atomic tasks. Prior research identified visual, cognitive, and some auditory tasks using wearable sensors; however, none of those methods recognize composite tasks across all activity components.</p>
<p>Humans often complete multiple tasks concurrently (<xref ref-type="bibr" rid="B29">Chen K. et al., 2021</xref>). For example, while <italic>clearing dangerous items</italic>, a human teammate may receive a <italic>communication request</italic> from the incident command center seeking important information and the teammate will be required to do both tasks simultaneously. Detecting this task concurrency will allow robots to better adapt to their teammate&#x2019;s interactions, priorities or appropriations, which will improve the team&#x2019;s overall collaboration and performance. A limitation of most existing algorithms assume that the human performs a single composite task at a time.</p>
<p>Individual differences are expected, even if the humans receive identical training, resulting in differing task completion steps, such as different atomic tasks being completed in different orders or with differing completion times. These differences can result in one task being mapped to multiple different sensor readings. For example, while <italic>applying tourniquet</italic>, a first responder may skip securing the excess band, as it is not a necessary step to stop the bleeding, while another responder may pack the wound with gauze, an additional step. Algorithmically identifying such individual differences is challenging. The existing approaches to addressing individual differences only considered trivial gross-motor and fine-grained motor atomic tasks (<xref ref-type="bibr" rid="B131">Mannini and Intille, 2018</xref>; <xref ref-type="bibr" rid="B5">Akbari and Jafari, 2020</xref>; <xref ref-type="bibr" rid="B50">Ferrari et al., 2020</xref>), not composite tasks involving multiple activity components.</p>
<p>Human responders train to respond to disasters, but each disaster differs and often requires the completion of unique tasks or tasks in a different manner. Performing tasks for which the robot has not been previously trained are called <italic>out-of-class tasks</italic> (<xref ref-type="bibr" rid="B109">Laput and Harrison, 2019</xref>). Misclassification of an out-of-class task can result in a robot adapting its behavior incorrectly, causing more harm than good.</p>
<p>This literature review evaluates task recognition algorithms within the context of HRTs, operating in uncertain, dynamic environments, by developing a set of relevant criteria. <xref ref-type="sec" rid="s2">Sections 2</xref>, <xref ref-type="sec" rid="s3">3</xref> provide the necessary background on task recognition and discuss prior literature reviews, respectively. <xref ref-type="sec" rid="s4">Section 4</xref> provides an overview of the relevant task recognition metrics, while <xref ref-type="sec" rid="s5">Section 5</xref> reviews relevant algorithms. <xref ref-type="sec" rid="s6">Section 6</xref> discusses the limitations of the current approaches, while <xref ref-type="sec" rid="s7">Sections 7</xref>, <xref ref-type="sec" rid="s8">8</xref> provide insights for future directions and concluding remarks, respectively.</p>
</sec>
<sec id="s2">
<title>2 Background</title>
<p>Task recognition involves classifying a human&#x2019;s task action based on a set of domain relevant activities, or tasks (<xref ref-type="bibr" rid="B206">Wang et al., 2019</xref>). A task, <italic>t</italic>
<sub>
<italic>k</italic>
</sub>, belongs to a task set <italic>T</italic>, for which a sequence of sensor readings <italic>S</italic>
<sub>
<italic>k</italic>
</sub> corresponds to the task. Task recognition is intended to identify a function <italic>f</italic> that predicts the task performed based on the sensor readings <italic>S</italic>
<sub>
<italic>k</italic>
</sub>, such that the discrepancy between the predicted task <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> and the ground truth <italic>t</italic>
<sub>
<italic>k</italic>
</sub> is minimized. <italic>f</italic> does not usually take <italic>S</italic>
<sub>
<italic>k</italic>
</sub> as direct input, as the sensor readings often require a processing function <italic>&#x3a6;</italic> that converts the sensor readings <italic>S</italic>
<sub>
<italic>k</italic>
</sub> into a <italic>d</italic>-dimensional feature vector <inline-formula id="inf2">
<mml:math id="m2">
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2250;</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> by extracting meaningful features (<xref ref-type="bibr" rid="B206">Wang et al., 2019</xref>). The function <italic>f</italic> inputs the feature vector <bold>x</bold> to predict the task <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>. Due to human&#x2019;s individual differences multiple different feature vectors can be mapped to a single task. Therefore, machine learning algorithms are adopted for learning the function <italic>f</italic> (<xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>).</p>
<p>Generally, human&#x2019;s tasks encompass multiple activity components: Physical movements, cognitive, visual, speech and auditory. Physical movements can be categorized based on motion granularity: 1) Gross motor, 2) Fine-grained motor, and 3) Tactile. Gross motor tasks&#x2019; physical movements displace the entire body, such as walking, running, and climbing stairs, or major portions, such as swinging an arm (e.g., <xref ref-type="bibr" rid="B36">Cleland et al., 2013</xref>; <xref ref-type="bibr" rid="B34">Chen and Xue, 2015</xref>; <xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>). Fine-grained motor tasks involve body extremities&#x2019; motion (i.e., wrists and fingers), such as grasping and object manipulation (e.g., <xref ref-type="bibr" rid="B216">Zhang et al., 2008</xref>; <xref ref-type="bibr" rid="B49">Fathi et al., 2011</xref>; <xref ref-type="bibr" rid="B109">Laput and Harrison, 2019</xref>). Tactile tasks&#x2019; physical movements result in sense of touch, such as mouse clicks, keyboard strokes, and carrying a backpack (e.g., <xref ref-type="bibr" rid="B72">Hsiao et al., 2015</xref>; <xref ref-type="bibr" rid="B114">Lee and Nicholls, 1999</xref>). Cognitive tasks use the brain to process new information, as well as recall or retrieve information from memory (e.g., <xref ref-type="bibr" rid="B99">Kunze et al., 2013a</xref>; <xref ref-type="bibr" rid="B214">Yuan et al., 2014</xref>; <xref ref-type="bibr" rid="B180">Salehzadeh et al., 2020</xref>). Visual task examples include identifying different objects, and reading (e.g., <xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B76">Ishimaru et al., 2017</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>). Speech-reliant tasks are voice articulation dependent, such as communicating over the radio (e.g., <xref ref-type="bibr" rid="B39">Dash et al., 2019</xref>; <xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>), while auditory tasks are acoustic events in the environment, such as an important announcement or emergency sounds (e.g., <xref ref-type="bibr" rid="B193">Stork et al., 2012</xref>; <xref ref-type="bibr" rid="B70">Hershey et al., 2017</xref>; <xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>). Most existing task recognition approaches focus primarily on detecting physical gross motor and fine-grained motor tasks. However, some tasks involve little to no physical movement. Robots need a holistic understanding of tasks&#x2019; various activity components to detect them accurately.</p>
</sec>
<sec id="s3">
<title>3 Related work</title>
<p>Human task recognition has been an active field of research for more than a decade; therefore, a number of review papers with a wide range of scope and objectives exist in the literature (<xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>; <xref ref-type="bibr" rid="B37">Cornacchia et al., 2017</xref>; <xref ref-type="bibr" rid="B149">Nweke et al., 2018</xref>; <xref ref-type="bibr" rid="B206">Wang et al., 2019</xref>; <xref ref-type="bibr" rid="B29">Chen K. et al., 2021</xref>; <xref ref-type="bibr" rid="B169">Ramanujam et al., 2021</xref>; <xref ref-type="bibr" rid="B16">Bian et al., 2022</xref>; <xref ref-type="bibr" rid="B217">Zhang S. et al., 2022</xref>). This manuscript focuses on wearable sensor-based human task recognition; thus, this section is only focused on this domain. A brief overview of the latest task recognition surveys from the last decade is presented.</p>
<p>One of the earliest surveys provided an overall task recognition framework, along with its primary components incorporating wearable sensors (<xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>). The survey categorized the manuscripts based on their learning approach (supervised or semi-supervised) and response time (offline or online), and qualitatively evaluated them in terms of recognition performance, energy consumption, obtrusiveness, and flexibility. The highlighted open problems included the need for composite and concurrent task recognition. Other comprehensive surveys of task recognition using wearable sensors discussed how different task types can be detected by i) breadth of sensing modalities, ii) choosing appropriate on-body sensor locations, iii) and learning approaches (<xref ref-type="bibr" rid="B37">Cornacchia et al., 2017</xref>; <xref ref-type="bibr" rid="B45">Elbasiony and Gomaa, 2020</xref>).</p>
<p>The prior reviews primarily surveyed classical machine learning based approaches (e.g., Support Vector Machines and Decision Trees) for task recognition. More recent surveys reviewed manuscripts that leveraged deep learning (<xref ref-type="bibr" rid="B149">Nweke et al., 2018</xref>; <xref ref-type="bibr" rid="B206">Wang et al., 2019</xref>; <xref ref-type="bibr" rid="B29">Chen K. et al., 2021</xref>; <xref ref-type="bibr" rid="B169">Ramanujam et al., 2021</xref>; <xref ref-type="bibr" rid="B217">Zhang S. et al., 2022</xref>). <xref ref-type="bibr" rid="B149">Nweke et al. (2018)</xref> and <xref ref-type="bibr" rid="B206">Wang et al. (2019)</xref> presented a taxonomy of generative, discriminative, and hybrid deep learning algorithms for task recognition, while <xref ref-type="bibr" rid="B169">Ramanujam et al. (2021)</xref> categorized the deep learning algorithms by convolutional neural network, long short-term memory, and hybrid methods to conduct an in-depth analysis on the benchmark datasets. A comprehensive review of the deep learning challenges and opportunities was presented (<xref ref-type="bibr" rid="B29">Chen K. et al., 2021</xref>). <xref ref-type="bibr" rid="B217">Zhang S. et al. (2022)</xref> focused on the most recent cutting-edge deep learning methods, such as generative adversarial networks and deep reinforcement learning, along with a thorough analysis in terms of model comparison, selection, and deployment.</p>
<p>Previous task recognition surveys primarily focused on discussing algorithms using wearable sensors. An in-depth understanding of the state-of-art sensing modalities is as important as the algorithmic solutions. A recent survey categorized task recognition-related sensing modalities into five classes: mechanical kinematic sensing, field-based sensing, wave-based sensing, physiological sensing, and hybrid (<xref ref-type="bibr" rid="B16">Bian et al., 2022</xref>). Specific sensing modalities were presented by category, along with the strengths and weaknesses of each modality across the categorization.</p>
<p>Other surveys focused solely on video-based task recognition algorithms, typically using surveillance-based datasets (<xref ref-type="bibr" rid="B89">Ke et al., 2013</xref>; <xref ref-type="bibr" rid="B210">Wu et al., 2017</xref>; <xref ref-type="bibr" rid="B215">Zhang H.-B. et al., 2019</xref>; <xref ref-type="bibr" rid="B155">Pareek and Thakkar, 2021</xref>), while <xref ref-type="bibr" rid="B126">Liu et al. (2019)</xref> reviewed algorithms that leveraged the change in wireless signals, such as the received signal strength indicator and Doppler shift to recognize tasks. The <italic>in situ</italic>, potentially deconstructed (e.g., 2023 Turkey&#x2013;Syria earthquake) first response HRT cannot rely on such sensors for task recognition.</p>
<p>Most existing literature reviews focus primarily on algorithms that detect tasks involving physical movements; thus, the reviewed algorithms are biased to only include gross and fine-grained motor task recognition methods. Compared to the prior literature reviews, this manuscript acknowledges that tasks have multiple dimensions and proposes a task taxonomy based on motion granularity (i.e., gross motor, fine-grained motor, and tactile), as well as other task component channels (i.e., visual, cognitive, auditory and speech). The manuscript&#x2019;s primary contributions are summarized below.<list list-type="simple">
<list-item>
<p>&#x2022; A comprehensive list of wearable sensor-based metrics are evaluated to assess the metrics&#x2019; ability to detect tasks within the context of the first-response HRT domain.</p>
</list-item>
<list-item>
<p>&#x2022; A systematic review of relevant manuscripts over the years that are categorized by the seven activity components and grouped based on the machine learning methods.</p>
</list-item>
<list-item>
<p>&#x2022; A set of criteria to evaluate the reviewed manuscripts&#x2019; ability to recognize tasks performed by first-response human teammates operating in HRTs.</p>
</list-item>
</list>The criteria developed to evaluate the metrics and algorithms may appear restrictive, as they are grounded in the first-response domain; however, the classifications provided are widely applicable for domains that involve unstructured, dynamic environments that require wearable sensors and demand a holistic understanding of a human&#x2019;s task state.</p>
</sec>
<sec id="s4">
<title>4 Task recognition metrics</title>
<p>Task recognition algorithms require metrics (e.g., inertial measurements, pupil dilation, and heart-rate) to detect the tasks performed by humans. The task recognition metrics incorporated by an algorithm inform how accurately a given set of tasks and the associated activity components can be detected for a particular task domain; therefore, selecting the right set of metrics takes precedence over algorithm development. Thirty task recognition metrics were identified across the existing literature. Three criteria were developed to evaluate the metrics&#x2019; ability to detect tasks performed by first-response HRTs operating in uncertain, unstructured environments.</p>
<p>
<italic>Sensitivity</italic> refers to a metric&#x2019;s ability to detect tasks reliably. A metric&#x2019;s sensitivity is classified as <italic>High</italic> if at least three citations indicate that the metric detects tasks with <inline-formula id="inf4">
<mml:math id="m4">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy, while a metric is classified as <italic>Medium</italic> if the task detection accuracy is <inline-formula id="inf5">
<mml:math id="m5">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>70</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula>, but <inline-formula id="inf6">
<mml:math id="m6">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula>. <italic>Low</italic> metric sensitivity occurs if the metric detects tasks with <inline-formula id="inf7">
<mml:math id="m7">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>70</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy. Metrics without sufficient citations to determine their sensitivity are classified as <italic>Indeterminate</italic>, and additional evidence is required to substantiate the metric&#x2019;s sensitivity.</p>
<p>
<italic>Versatility</italic> refers to a metric&#x2019;s ability to detect tasks across different task domains. A metric&#x2019;s versatility is <italic>High</italic> if the metric is cited for discriminating tasks in at least two or more task domains. Similarly, if the metric was used for classifying tasks belonging to only one task domain, the versatility is <italic>Low</italic>.</p>
<p>
<italic>Suitability</italic> evaluates a metric&#x2019;s feasibility to detect tasks in various physical environments (i.e., structured vs unstructured), which depends on the sensor technology for gathering the metric. Some metrics (e.g., eye gaze) can be acquired using many different sensors and technologies, but the review&#x2019;s focus on disaster response encourages the use of wearable sensors over environmentally embedded sensors. A metric&#x2019;s <italic>suitability</italic> is <italic>conforming</italic> if it is cited to be gathered by a <italic>wearable</italic> sensor that is unaffected by disturbances (e.g., sensor displacement noise, excessive perspiration, and change in lighting conditions), while <italic>non-conforming</italic> otherwise.</p>
<p>The task recognition metrics and the corresponding sensitivity, versatility, and suitability classifications are provided in <xref ref-type="table" rid="T1">Table 1</xref>. The <italic>Activity Component</italic> column in <xref ref-type="table" rid="T1">Table 1</xref> indicates which component(s) (i.e., gross motor, fine-grained motor, tactile, visual, cognitive, auditory, and speech) are associated with the metric. The metrics are categorized based on their sensing properties. Each metric is evaluated based on the three evaluation criteria in order to identify the most reliable, minimal set of metrics required to recognize HRT tasks.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Metrics evaluation overview by Sensitivity (<bold>Sens.</bold>), Versatility (<bold>Verst.</bold>), and Suitability (<bold>Suit.</bold>), where &#x22c1;, <bold>
<italic>&#x220f;</italic>
</bold>, and &#x22c0;, represent Low, Medium, and High, respectively <bold>(.)</bold> indicates Indeterminate, while <bold>&#x2a;</bold> indicate hypothesis predicted for the particular metric. Suitability is classified as conforming (<bold>C</bold>) or non-conforming (<bold>NC</bold>).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Category</th>
<th align="left">Metrics</th>
<th align="center">Sens</th>
<th align="center">Verst</th>
<th align="center">Suit</th>
<th align="left">Activity component</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="5" align="left">Inertial</td>
<td rowspan="3" align="left">Acceleration</td>
<td rowspan="3" align="center">&#x22c0;</td>
<td rowspan="3" align="center">&#x22c0;</td>
<td rowspan="3" align="center">C</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td align="left">Tactile</td>
</tr>
<tr>
<td rowspan="2" align="left">Orientation</td>
<td rowspan="2" align="center">(&#x22c0;&#x2a;)</td>
<td rowspan="2" align="center">&#x22c0;</td>
<td rowspan="2" align="center">C</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td rowspan="7" align="left">Eye Gaze</td>
<td align="left">Fixation</td>
<td align="center">&#x22c0;</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="left">Visual</td>
</tr>
<tr>
<td align="left">Saccades</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="left">Visual</td>
</tr>
<tr>
<td align="left">Scanpath</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="left">Visual</td>
</tr>
<tr>
<td rowspan="2" align="left">Blink rate</td>
<td align="center">(&#x22c1;&#x2a;)</td>
<td align="center">&#x22c0;&#x2a;</td>
<td align="center">C</td>
<td align="left">Visual</td>
</tr>
<tr>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">&#x22c0;&#x2a;</td>
<td align="center">C</td>
<td align="left">Cognitive&#x2a;</td>
</tr>
<tr>
<td rowspan="2" align="left">Pupil dilation</td>
<td rowspan="2" align="center">(&#x22c0;&#x2a;)</td>
<td rowspan="2" align="center">(&#x22c0;&#x2a;)</td>
<td rowspan="2" align="center">NC</td>
<td align="left">Visual&#x2a;</td>
</tr>
<tr>
<td align="left">Cognitive&#x2a;</td>
</tr>
<tr>
<td rowspan="9" align="left">Electro-physiological</td>
<td rowspan="2" align="left">EOG</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Visual</td>
</tr>
<tr>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Cognitive</td>
</tr>
<tr>
<td rowspan="3" align="left">sEMG</td>
<td rowspan="2" align="center">&#x22c0;</td>
<td rowspan="2" align="center">&#x22c0;</td>
<td rowspan="2" align="center">NC</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Tactile</td>
</tr>
<tr>
<td align="left">EEG</td>
<td align="center">&#x22c0;</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="left">Cognitive</td>
</tr>
<tr>
<td align="left">ECG</td>
<td align="center">&#x22c1;</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left">Heart-rate</td>
<td align="center">&#x22c1;</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left">Heart-rate variability</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Cognitive&#x2a;</td>
</tr>
<tr>
<td rowspan="6" align="left">Vision</td>
<td rowspan="2" align="left">Optical flow</td>
<td align="center">&#x22c0;</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td rowspan="2" align="left">Human-body pose</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="center">&#x22c0;</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td rowspan="2" align="left">Object detection</td>
<td align="center">&#x22c1;</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td rowspan="3" align="left">Acoustic</td>
<td align="left">Spectrogram</td>
<td align="center">&#x22c0;</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="left">Auditory</td>
</tr>
<tr>
<td align="left">MFCCs</td>
<td align="center">&#x22c0;</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="left">Auditory</td>
</tr>
<tr>
<td align="left">Noise level</td>
<td align="center">(&#x22c1;&#x2a;)</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Auditory</td>
</tr>
<tr>
<td rowspan="5" align="left">Speech</td>
<td align="left">Transcript</td>
<td align="center">(<italic>&#x220f;</italic>&#x2a;)</td>
<td align="center">(&#x22c1;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Speech</td>
</tr>
<tr>
<td align="left">Keywords</td>
<td align="center">(<italic>&#x220f;</italic>&#x2a;)</td>
<td align="center">(&#x22c1;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Speech</td>
</tr>
<tr>
<td align="left">Speech rate</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Speech</td>
</tr>
<tr>
<td align="left">Voice intensity</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Speech</td>
</tr>
<tr>
<td align="left">Voice pitch</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">(&#x22c0;&#x2a;)</td>
<td align="center">C</td>
<td align="left">Speech</td>
</tr>
<tr>
<td rowspan="3" align="left">Localization</td>
<td align="left">Outdoor localization</td>
<td align="center">&#x22c1;</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="left">Gross</td>
</tr>
<tr>
<td rowspan="2" align="left">Indoor localization</td>
<td rowspan="2" align="center">&#x22c0;</td>
<td rowspan="2" align="center">&#x22c0;</td>
<td rowspan="2" align="center">NC</td>
<td align="left">Gross</td>
</tr>
<tr>
<td align="left">Fine-grained</td>
</tr>
<tr>
<td align="left">Miscellaneous</td>
<td align="left">Physiological</td>
<td align="center">&#x22c1;</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="left">Gross</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s4-1">
<title>4.1 Inertial metrics</title>
<p>
<italic>Inertial</italic> metrics consists of: i) linear acceleration, which measures a body region&#x2019;s three-dimensional movement via an accelerometer; and ii) the body part&#x2019;s three-dimensional orientation (i.e., rotation and rotational rate) using gyroscope and magnetometer. Inertial metrics are primarily used for detecting physical tasks, which includes gross motor, fine-grained motor and tactile tasks (<xref ref-type="bibr" rid="B36">Cleland et al., 2013</xref>; <xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>; <xref ref-type="bibr" rid="B202">Vepakomma et al., 2015</xref>; <xref ref-type="bibr" rid="B37">Cornacchia et al., 2017</xref>).</p>
<p>Linear acceleration is the most widely employed, and can be used as a standalone task recognition metric. Orientation is often used in combination with linear acceleration. The type and number of tasks detected by the inertial metrics can be linked to the number and placement of the sensors on the body (<xref ref-type="bibr" rid="B12">Atallah et al., 2010</xref>; <xref ref-type="bibr" rid="B36">Cleland et al., 2013</xref>). For example, inertial metric sensors to detect gross motor tasks are placed at central or lower body locations (i.e., chest, waist and thighs) (<xref ref-type="bibr" rid="B170">Ravi et al., 2005</xref>; <xref ref-type="bibr" rid="B102">Kwapisz et al., 2011</xref>; <xref ref-type="bibr" rid="B132">Mannini and Sabatini, 2011</xref>; <xref ref-type="bibr" rid="B58">Gjoreski et al., 2014</xref>), while fine-grained motor tasks require the sensor on the forearms and wrists (<xref ref-type="bibr" rid="B216">Zhang et al., 2008</xref>; <xref ref-type="bibr" rid="B95">Koskimaki et al., 2009</xref>; <xref ref-type="bibr" rid="B143">Min and Cho, 2011</xref>; <xref ref-type="bibr" rid="B66">Heard et al., 2019a</xref>), and tactile tasks place the sensors at the hand&#x2019;s dorsal side and fingers (<xref ref-type="bibr" rid="B85">Jing et al., 2011</xref>; <xref ref-type="bibr" rid="B26">Cha et al., 2018</xref>; <xref ref-type="bibr" rid="B128">Liu et al., 2018</xref>). Linear acceleration has high sensitivity, while standalone orientation is indeterminate. The inertial metrics have high versatility and conform with suitability criteria.</p>
</sec>
<sec id="s4-2">
<title>4.2 Eye gaze metrics</title>
<p>
<italic>Eye gaze</italic> metrics record the coordinates (<italic>g</italic>
<sub>
<italic>x</italic>
</sub>, <italic>g</italic>
<sub>
<italic>y</italic>
</sub>) of the gaze point over time. Raw eye gaze data is often processed to yield eye movement metrics representative of a human&#x2019;s visual behavior and can be leveraged for task recognition. Fixations, saccades, scanpath, blink rate, and pupil dilation represent some of the important eye gaze-based metrics (<xref ref-type="bibr" rid="B100">Kunze et al., 2013b</xref>; <xref ref-type="bibr" rid="B192">Steil and Bulling, 2015</xref>; <xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>; <xref ref-type="bibr" rid="B71">Hevesi et al., 2018</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>; <xref ref-type="bibr" rid="B90">Kelton et al., 2019</xref>; <xref ref-type="bibr" rid="B107">Landsmann et al., 2019</xref>). These metrics are commonly used for recognizing visual tasks, while some prior research have also detected cognitive tasks (<xref ref-type="bibr" rid="B101">Kunze et al., 2013c</xref>; <xref ref-type="bibr" rid="B79">Islam et al., 2021</xref>).</p>
<p>
<italic>Fixations</italic> are stationary eye states during which gaze is held upon a particular location (<xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>), while <italic>saccades</italic> are the simultaneous movement of both eyes between two fixations (<xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>). Fixation has high sensitivity (e.g., <xref ref-type="bibr" rid="B192">Steil and Bulling, 2015</xref>; <xref ref-type="bibr" rid="B90">Kelton et al., 2019</xref>; <xref ref-type="bibr" rid="B107">Landsmann et al., 2019</xref>), while a saccade has medium sensitivity (e.g., <xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B76">Ishimaru et al., 2017</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>). Both metrics have high versatility and conform with suitability.</p>
<p>A <italic>scanpath</italic> is a fixation-saccade-fixation sequence (<xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>). Scanpaths have medium sensitivity (<xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>; <xref ref-type="bibr" rid="B79">Islam et al., 2021</xref>), conform with suitability, and have low versatility.</p>
<p>
<italic>Blink rate</italic> represents the number of blinks (i.e., opening and closing eyelids) per unit time (<xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>). Blink rate is often used in conjunction with fixations and saccades to provide further context. The metric&#x2019;s sensitivity is indeterminate, but it is hypothesized to have low sensitivity and is used predominantly to detect desktop or office-based visual tasks (e.g., <xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B77">Ishimaru et al., 2014a</xref>). Prior research reviews indicate that blink rate highly correlates with the cognitive workload (<xref ref-type="bibr" rid="B133">Marquart et al., 2015</xref>; <xref ref-type="bibr" rid="B67">Heard et al., 2018a</xref>); therefore, blink rate is hypothesized to have high sensitivity toward cognitive task recognition. This metric is also hypothesized to have high versatility given its potential to be used in multiple task domains, and it conforms with suitability.</p>
<p>
<italic>Pupil dilation</italic>, or pupillometry, is the change in pupil diameter. The metric&#x2019;s sensitivity is indeterminate; however, hypothesized to have high sensitivity and versatility toward cognitive and visual task recognition based on its ability to reliably detect cognitive workload (<xref ref-type="bibr" rid="B4">Ahlstrom and Friedman-Berg, 2006</xref>; <xref ref-type="bibr" rid="B133">Marquart et al., 2015</xref>; <xref ref-type="bibr" rid="B67">Heard et al., 2018a</xref>), and its high correlation in various visual search tasks (<xref ref-type="bibr" rid="B164">Porter et al., 2007</xref>; <xref ref-type="bibr" rid="B167">Privitera et al., 2010</xref>; <xref ref-type="bibr" rid="B203">Wahn et al., 2016</xref>). Environmental lighting changes can significantly impact the metric&#x2019;s acquisition, so it does not conform with suitability.</p>
</sec>
<sec id="s4-3">
<title>4.3 Electrophysiological metrics</title>
<p>The <italic>electrophysiological</italic> metrics refer to the electrical signals associated with various body parts (e.g., muscles, brain and eyes). These signals can be leveraged for task recognition, as they are highly correlated with tasks humans conduct. The most common electrophysiological metrics are electromyography, electrooculography, electroencephalography, and electrocardiography.</p>
<p>
<italic>Electromyography</italic> measures the potential difference caused by contracting and relaxing muscle tissues. Surface-electromyography (sEMG) is a non-invasive technique, wherein electrodes placed on the skin measure the electromyography signals. A forearm positioned sEMG commonly detects fine-grained motor tasks (e.g., <xref ref-type="bibr" rid="B96">Koskim&#xe4;ki et al., 2017</xref>; <xref ref-type="bibr" rid="B69">Heard et al., 2019b</xref>; <xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>), while upper limb positioned sEMG can detect gross motor tasks (e.g., <xref ref-type="bibr" rid="B184">Scheme and Englehart, 2011</xref>; <xref ref-type="bibr" rid="B198">Trigili et al., 2019</xref>). sEMG has high sensitivity and versatility for detecting gross and fine-grained motor tasks. The metric has been employed for detecting various finger and intricate hand motions (e.g., <xref ref-type="bibr" rid="B32">Chen et al. (2007a</xref>; <xref ref-type="bibr" rid="B33">b)</xref>; <xref ref-type="bibr" rid="B219">Zhang et al. (2009)</xref>); thus, it is hypothesized to have high sensitivity for detecting tactile tasks. Finally, the metric does not conform with suitability, as sweat accumulation underneath the electrodes may compromise the sEMG sensor&#x2019;s adherence to the skin, as well as the associated signal fidelity (<xref ref-type="bibr" rid="B1">Abdoli-Eramaki et al., 2012</xref>).</p>
<p>The <italic>Electrooculography (EOG)</italic> metric measures the potential difference between the cornea and the retina caused by eye movements. The metric has medium sensitivity for classifying visual tasks (e.g., typing, web browsing, reading and watching videos) (<xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B78">Ishimaru et al., 2014b</xref>; <xref ref-type="bibr" rid="B79">Islam et al., 2021</xref>). The metric is capable of detecting cognitive tasks (e.g., <xref ref-type="bibr" rid="B40">Datta et al., 2014</xref>; <xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>), but its cognitive sensitivity is indeterminate. The metric has high versatility (<xref ref-type="bibr" rid="B40">Datta et al., 2014</xref>; <xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>). EOG does not conform with suitability, because it is susceptible to noise introduced by facial muscle movements (<xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>).</p>
<p>
<italic>Electroencephalography (EEG)</italic> collects electrical neurophysiological signals from different parts of the brain. EEG measures two different metrics: i) the event-related potential measures the voltage signal produced by the brain in response to a stimulus (e.g., <xref ref-type="bibr" rid="B220">Zhang X. et al., 2019</xref>; <xref ref-type="bibr" rid="B180">Salehzadeh et al., 2020</xref>); and ii) the power spectral density measures the power present in the signal spectrum (e.g., <xref ref-type="bibr" rid="B99">Kunze et al., 2013a</xref>; <xref ref-type="bibr" rid="B182">Sarkar et al., 2016</xref>). Both metrics have high sensitivity. EEG signals may be inaccurate when a human is physically active, so the metrics are best suited for detecting cognitive tasks in a sedentary environment. Therefore, the EEG metrics have low versatility. EEG signals suffer from low signal-to-noise ratios (<xref ref-type="bibr" rid="B186">Schirrmeister et al., 2017</xref>; <xref ref-type="bibr" rid="B220">Zhang X. et al., 2019</xref>), and incorrect sensor placement can create inaccuracies; therefore, EEG metrics do not conform with suitability.</p>
<p>
<italic>Electrocardiography</italic> (ECG) measures the heart&#x2019;s electrical activity. Standalone ECG signals are not sensitive enough to detect tasks; therefore, ECG is often used in conjunction with inertial metrics to detect gross motor tasks (<xref ref-type="bibr" rid="B84">Jia and Liu, 2013</xref>; <xref ref-type="bibr" rid="B91">Kher et al., 2013</xref>). ECG signal has low versatility, as it has been used to detect only ambulatory tasks (e.g., <xref ref-type="bibr" rid="B84">Jia and Liu, 2013</xref>; <xref ref-type="bibr" rid="B91">Kher et al., 2013</xref>), but it does conform with suitability.</p>
<p>The ECG signals can be used to measure two other metrics: i) <italic>heart-rate</italic> measures the number of heart beats per minute, while ii) <italic>heart-rate variability</italic> measures the variation in the heart-rate&#x2019;s beat-to-beat interval. Heart-rate has low sensitivity for detecting gross motor tasks (e.g., <xref ref-type="bibr" rid="B195">Tapia et al., 2007</xref>; <xref ref-type="bibr" rid="B156">Park et al., 2017</xref>; <xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>) and is often used for distinguishing the humans&#x2019; intensity when performing physical tasks (<xref ref-type="bibr" rid="B195">Tapia et al., 2007</xref>; <xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>). Heart-rate has low versatility, as it can only detect ambulatory tasks. The heart-rate metric conforms with suitability if a human&#x2019;s stress and fatigue levels remain constant. The heart-rate variability metric has seldom been used for task recognition (<xref ref-type="bibr" rid="B156">Park et al., 2017</xref>), but it is sensitive to large variations in cognitive workload (<xref ref-type="bibr" rid="B67">Heard et al., 2018a</xref>). Therefore, the metric is hypothesized to have high sensitivity for cognitive task recognition. The metric conforms with suitability, and is hypothesized to have high versatility.</p>
</sec>
<sec id="s4-4">
<title>4.4 Vision-based metrics</title>
<p>
<italic>Vision-based</italic> metrics (e.g., optical flow, human-body pose and object detection) use videos and images containing human motions in order to infer the tasks being performed (<xref ref-type="bibr" rid="B49">Fathi et al., 2011</xref>; <xref ref-type="bibr" rid="B55">Garcia-Hernando et al., 2018</xref>; <xref ref-type="bibr" rid="B200">Ullah et al., 2018</xref>; <xref ref-type="bibr" rid="B147">Neili Boualia and Essoukri Ben Amara, 2021</xref>). These metrics are acquired via environmentally embedded cameras installed at fixed locations, or using wearable cameras mounted on a human&#x2019;s shoulders, head, or chest (e.g., <xref ref-type="bibr" rid="B137">Mayol and Murray, 2005</xref>; <xref ref-type="bibr" rid="B49">Fathi et al., 2011</xref>; <xref ref-type="bibr" rid="B136">Matsuo et al., 2014</xref>).</p>
<p>Vision-based metrics detect gross and fine-grained motor tasks by enabling various computer vision algorithms [e.g., object detection, localization, and motion tracking <xref ref-type="bibr" rid="B65">He et al. (2016)</xref>; <xref ref-type="bibr" rid="B41">Deng et al. (2009)</xref>; <xref ref-type="bibr" rid="B172">Redmon et al. (2016)</xref>], which are relevant for task recognition. The metrics&#x2019; use is discouraged for the intended HRT domain, because i) environmentally embedded cameras are not readily available in unstructured domains; ii) high susceptibility to background noise from lighting, vibrations, and occlusion; iii) raise privacy concerns, and iv) computationally expensive to process (<xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>). The metrics have medium to high sensitivity, high versatility, and non-conforming with suitability.</p>
</sec>
<sec id="s4-5">
<title>4.5 Acoustic metrics</title>
<p>Acoustic metrics leverage the characteristic sounds in order to detect auditory events in the surrounding environment. Auditory event recognition algorithms commonly use two types of frequency-domain metrics. A <italic>spectrogram</italic> is a three-dimensional acoustic metric representing a sound signal&#x2019;s amplitude over time at various frequencies (<xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>). The spectrogram has high sensitivity and versatility (e.g., <xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>; <xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>; <xref ref-type="bibr" rid="B122">Liang and Thomaz, 2019</xref>). Cepstrum represents the short-term power spectrum of a sound, and is obtained by applying an inverse Fourier transformation on a sound wave&#x2019;s spectrum. The <italic>Mel Frequency Cepstral Coefficients (MFCCs)</italic> represent the amplitudes of the resulting cepstrum on a Mel scale. The MFCC metric has high sensitivity and versatility (e.g., <xref ref-type="bibr" rid="B141">Min et al., 2008</xref>; <xref ref-type="bibr" rid="B193">Stork et al., 2012</xref>). Both metrics conform with suitability, as long as the audio is captured via a wearable microphone.</p>
<p>
<italic>Noise level</italic> measures a task environment&#x2019;s loudness in decibels. Noise level correlates to an increase in auditory workload (<xref ref-type="bibr" rid="B68">Heard et al., 2018b</xref>), but has not been used for task recognition; however, the metric is hypothesized to detect auditory events when the events are fewer <inline-formula id="inf8">
<mml:math id="m8">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. The metric&#x2019;s sensitivity, versatility, and suitability are hypothesized to be low, high, and conforming, respectively.</p>
</sec>
<sec id="s4-6">
<title>4.6 Speech metrics</title>
<p>Communication exchanges between human teammates can be translated into text, or a <italic>Verbal transcript</italic>, such that the message is captured as it was spoken. Transcripts can be generated manually (<xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>), or using an automatic speech recognition tool [e.g., SPHINX (<xref ref-type="bibr" rid="B113">Lee et al., 1990</xref>), Kaldi (<xref ref-type="bibr" rid="B165">Povey et al., 2011</xref>), Wav2Letter&#x2b;&#x2b; (<xref ref-type="bibr" rid="B166">Pratap et al., 2019</xref>)]. The transcribed words are encoded into <italic>n</italic> &#x2212; dimensional vectors [e.g., <italic>GloVe</italic> vector embeddings (<xref ref-type="bibr" rid="B161">Pennington et al., 2014</xref>)] to be used as inputs for detecting speech-reliant tasks (<xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>). Representative <italic>keywords</italic> that are spoken more frequently can be used for detecting tasks (<xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>). Keywords can be detected for every utterance automatically using word-spotting (<xref ref-type="bibr" rid="B199">Tsai and Hao, 2019</xref>; <xref ref-type="bibr" rid="B54">Gao et al., 2020</xref>). Identifying keywords for each task is non-trivial and requires considerable human effort. Both transcript and keywords metrics&#x2019; sensitivity is indeterminate, but is hypothesized to be medium (<xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>; <xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>). The metrics conform with suitability, provided the speech audio is obtained using a wearable microphone that minimize extraneous ambient noise. The metrics are exceptionally domain specific; therefore, their versatility is hypothesized to be low.</p>
<p>Several speech-related metrics (e.g., speech rate, pitch, and voice intensity) that do not rely on natural language processing have proven effective for estimating speech workload (<xref ref-type="bibr" rid="B67">Heard et al., 2018a</xref>; <xref ref-type="bibr" rid="B66">Heard et al., 2019a</xref>; <xref ref-type="bibr" rid="B52">Fortune et al., 2020</xref>). <italic>Speech rate</italic> captures verbal communications&#x2019; articulation rate by measuring the number of syllables uttered per unit time (<xref ref-type="bibr" rid="B52">Fortune et al., 2020</xref>). <italic>Voice intensity</italic> is the speech signal&#x2019;s root-mean-square value, while <italic>Pitch</italic> is the signal&#x2019;s dominant frequency over a time period (<xref ref-type="bibr" rid="B66">Heard et al., 2019a</xref>). These metrics have not been used for task recognition; therefore, additional evidence is required to substantiate their evaluation criteria. The metrics&#x2019; sensitivity and versatility are hypothesized to be high. The suitability criterion is conforming, assuming that the speech audio is obtained via wearable microphones.</p>
</sec>
<sec id="s4-7">
<title>4.7 Localization-based metrics</title>
<p>Localization-based metrics infer tasks by analyzing either the absolute or relative position of items of interest, including humans. The outdoor localization metric measures a human&#x2019;s absolute location (i.e., latitude and longitude coordinates) using satellite navigation systems. This metric has low sensitivity, (<xref ref-type="bibr" rid="B124">Liao et al., 2007</xref>), but can support task recognition by providing context (<xref ref-type="bibr" rid="B171">Reddy et al., 2010</xref>; <xref ref-type="bibr" rid="B174">Riboni and Bettini, 2011</xref>). The metric is highly versatile and non-conforming with suitability (<xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>).</p>
<p>Indoor localization determines the relative position of items, including humans, relative to a known reference point in indoor environments by acquiring the change in radio signals. Radio Frequency Identification (RFID) tags and wireless modems installed at stationary locations are the standard options. The indoor localization metric infers tasks by determining humans&#x2019; location or identifying objects lying in close proximity (<xref ref-type="bibr" rid="B37">Cornacchia et al., 2017</xref>; <xref ref-type="bibr" rid="B48">Fan et al., 2019</xref>). The metric has high sensitivity and versatility. The metric requires environmentally embedded sensors; therefore, it is non-conforming for suitability.</p>
</sec>
<sec id="s4-8">
<title>4.8 Physiological metrics</title>
<p>Physiological metrics provide precise information about a human&#x2019;s vital state. Several physiological metrics exist: i) <italic>Galvanic skin response</italic>, which measures the skin&#x2019;s conductivity, ii) <italic>Respiration rate</italic>, which represents the number of breaths taken per minute, iii) <italic>Posture Magnitude</italic>, which measures a human&#x2019;s trunk flexion (leaning forward) and extension (leaning backward) angle in degrees, and iv) <italic>Skin temperature</italic>, are the most commonly used physiological metrics for task recognition (<xref ref-type="bibr" rid="B143">Min and Cho, 2011</xref>; <xref ref-type="bibr" rid="B112">Lara et al., 2012</xref>). The physiological metrics have low sensitivity, as they react to activity changes with a time delay. The metrics have high versatility and conform with suitability.</p>
</sec>
</sec>
<sec id="s5">
<title>5 Task recognition algorithms</title>
<p>Over one hundred task recognition algorithms across different activity components and task domains were identified and reviewed. The algorithms are evaluated using the following criteria: sensitivity, suitability, generalizability, composite factor, concurrency, and anomaly awareness. The evaluation criteria and the corresponding requirements were chosen in order to assess an algorithm&#x2019;s viability for detecting tasks in a human-robot teaming domain.</p>
<sec id="s5-1">
<title>5.1 Evaluation criteria</title>
<p>
<italic>Sensitivity</italic> refers to an algorithm&#x2019;s ability to detect tasks reliably. An algorithm&#x2019;s sensitivity is classified as <italic>High</italic> if the algorithm detects tasks with <inline-formula id="inf9">
<mml:math id="m9">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy, while <italic>Medium</italic> if the algorithm&#x2019;s accuracy is <inline-formula id="inf10">
<mml:math id="m10">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>70</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula>, but <inline-formula id="inf11">
<mml:math id="m11">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula>, and <italic>Low</italic> if the accuracy is <inline-formula id="inf12">
<mml:math id="m12">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>70</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula>. The accuracy thresholds were chosen by fitting a skewed Gaussian curve on the reviewed task recognition algorithms&#x2019; accuracies.</p>
<p>An algorithm&#x2019;s <italic>suitability</italic> evaluates its feasibility for detecting tasks in various physical environments. An algorithm conforms if it can detect tasks independent of the environment by incorporating wearable, reliable metrics, and is non-conforming otherwise.</p>
<p>
<italic>Generalizability</italic> represents an algorithm&#x2019;s ability to identify tasks across humans. The generalizability criterion depends on the achieved accuracy, given the algorithm&#x2019;s validation method. An algorithm conforms if it achieves <inline-formula id="inf13">
<mml:math id="m13">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy with <italic>leave-one-subject-out</italic> cross-validation or <italic>in-the-wild</italic> validation.</p>
<p>The <italic>Composite factor</italic> criterion determines whether an algorithm can detect tasks composed of multiple atomic activities. If a detected task incorporates two or more atomic activities, then the algorithm conforms with the composite task criterion. Typically, long duration tasks that incorporate multiple action sequences per task are composite in nature.</p>
<p>
<italic>Concurrency</italic> determines if the algorithm can detect tasks executed simultaneously. Concurrency has multiple forms: i) a task may be initiated prior to completing a task, such that a portion of the task overlaps with the prior task (i.e., <italic>interleaved tasks</italic>), and ii) multiple tasks performed at the same time (i.e., <italic>simultaneous tasks</italic>) (<xref ref-type="bibr" rid="B8">Allen and Ferguson, 1994</xref>; <xref ref-type="bibr" rid="B127">Liu et al., 2017</xref>). An algorithm conforms if it can detect at least one form of concurrency, and is non-conforming otherwise.</p>
<p>
<italic>Anomaly Awareness</italic> determines an algorithm&#x2019;s ability to detect an out-of-class task instance, which arises when an algorithm encounters sensor data that does not correspond to any of the algorithm&#x2019;s learned tasks. An algorithm conforms with anomaly awareness if it can detect out-of-class instances.</p>
<p>Most task recognition algorithms can only detect a predefined set of atomic tasks and are unable to detect concurrent tasks or out-of-class instances (<xref ref-type="bibr" rid="B111">Lara and Labrador, 2013</xref>; <xref ref-type="bibr" rid="B37">Cornacchia et al., 2017</xref>). Thus, unless identified otherwise, the reviewed algorithms do not conform with composite factor, concurrency and anomaly awareness.</p>
</sec>
<sec id="s5-2">
<title>5.2 Overview of task recognition algorithm categories</title>
<p>Task recognition algorithms typically incorporate supervised machine learning to identify the tasks from the sensor data (see <xref ref-type="sec" rid="s2">Section 2</xref>). These algorithms can be grouped into several categories based on feature extraction, ability to handle uncertainty, and heuristics. Three common data-driven task recognition algorithm categories exist in the literature, which are.<list list-type="simple">
<list-item>
<p>&#x2022; <italic>Classical machine learning</italic> rely on features extracted from raw sensor data to learn a prediction model. Classical approaches are suitable when there is sufficient domain knowledge to extract meaningful features, and the training dataset is small.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>Deep learning</italic> avoids designing handcrafted features, learns the features automatically (<xref ref-type="bibr" rid="B59">Goodfellow et al., 2016</xref>), and is generally suitable when a large amount of data is available for training the model. Deep learning approaches leverage data to extract high-level features, while simultaneously training a model to predict the tasks.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>Probabilistic graphical models</italic> utilize probabilistic network structures (e.g., Bayesian Networks (<xref ref-type="bibr" rid="B44">Du et al., 2006</xref>), Hidden Markov Models (<xref ref-type="bibr" rid="B35">Chung and Liu, 2008</xref>), Conditional Random Fields (<xref ref-type="bibr" rid="B201">Vail et al., 2007</xref>)) to model uncertainties and the tasks&#x2019; temporal relationships, while also identifying composite, concurrent tasks.</p>
</list-item>
</list>The data-driven models&#x2019; primary limitations are that they i) cannot be interpreted easily, and ii) may require large amount of training data to be robust enough to handle individual differences across humans and generalize across multiple domains.</p>
<p>
<italic>Knowledge-driven</italic> task recognition models exploit heuristics and domain knowledge to recognize the tasks using reasoning-based approaches [e.g., ontology and first-order logic (<xref ref-type="bibr" rid="B197">Triboan et al., 2017</xref>; <xref ref-type="bibr" rid="B177">Safyan et al., 2019</xref>; <xref ref-type="bibr" rid="B194">Tang et al., 2019</xref>)]. Knowledge-driven models are logically elegant and easier to interpret, but do not have enough expressive power to model uncertainties. Additionally, creating logical rules to model temporal relations becomes impractical when there are a large number of tasks with intricate relationships (<xref ref-type="bibr" rid="B31">Chen and Nugent, 2019</xref>; <xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>).</p>
</sec>
<sec id="s5-3">
<title>5.3 Gross motor tasks</title>
<p>Gross motor tasks occur across multiple task categories, such as Activities of Daily Living (ADL), <italic>fitness</italic>, and, <italic>industrial</italic>. A high-level overview of the reviewed algorithms with regard to the evaluation criteria is presented by algorithm category in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Gross motor task recognition algorithms evaluation overview by Sensitivity (Sens.), Suitability (Suit.), Generalizability (Genr.), Composite Factor (Comp.), Concurrency (Conc.), and Anomaly Awareness (Anom.). Sensitivity is classified as Low (&#x22c1;), Medium (<italic>&#x220f;</italic>), or High (&#x22c0;), while other criteria are classified as conforming (C), non-conforming (NC), or requiring additional evidence (RE).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Classical Machine Learning</td>
</tr>
<tr>
<td align="left">Artificial neural network</td>
<td align="left">
<xref ref-type="bibr" rid="B91">Kher et al. (2013)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">Decision trees</td>
<td align="left">
<xref ref-type="bibr" rid="B157">Parkka et al. (2006)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B195">Tapia et al. (2007)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Ensemble</td>
<td align="left">
<xref ref-type="bibr" rid="B146">Nandy et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">k-Nearest Neighbors</td>
<td align="left">
<xref ref-type="bibr" rid="B97">Kubota et al. (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B104">Ladjailia et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Logistic regression</td>
<td align="left">
<xref ref-type="bibr" rid="B112">Lara et al. (2012)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Plurality voting</td>
<td align="left">
<xref ref-type="bibr" rid="B170">Ravi et al. (2005)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Random forest</td>
<td align="left">
<xref ref-type="bibr" rid="B7">Allahbakhshi et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Recurrent neural network</td>
<td align="left">
<xref ref-type="bibr" rid="B14">Batzianoulis et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Relevance vector machines</td>
<td align="left">
<xref ref-type="bibr" rid="B84">Jia and Liu (2013)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="3" align="left">SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B156">Park et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B98">Kumar and John (2016)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B187">Schuldt et al. (2004)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td rowspan="4" align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B34">Chen and Xue (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B9">Alsheikh et al. (2016)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B74">Ignatov (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B115">Lee et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B75">Inoue et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="3" align="left">CNN &#x2b; LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B48">Fan et al. (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B160">Peng et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B30">Chen et al. (2021b)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">CNN &#x2b; GRU</td>
<td align="left">
<xref ref-type="bibr" rid="B212">Xu et al. (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Transformer</td>
<td align="left">
<xref ref-type="bibr" rid="B43">Dirgov&#xe1; Lupt&#xe1;kov&#xe1; et al. (2022)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Probabilistic Graphical Model</td>
</tr>
<tr>
<td align="left">Bayesian network</td>
<td align="left">
<xref ref-type="bibr" rid="B223">Zhang et al. (2013)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Conditional random field</td>
<td align="left">
<xref ref-type="bibr" rid="B73">Hu et al. (2016)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Gaussian mixture model</td>
<td align="left">
<xref ref-type="bibr" rid="B198">Trigili et al. (2019)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">Hidden Markov model</td>
<td align="left">
<xref ref-type="bibr" rid="B83">Jalal et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B93">Kolekar and Dash (2016)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Knowledge-driven</td>
</tr>
<tr>
<td align="left">Dynamic time warping</td>
<td align="left">
<xref ref-type="bibr" rid="B42">Ding et al. (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Principal component analysis</td>
<td align="left">
<xref ref-type="bibr" rid="B158">Pawar et al. (2007)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Trigger-based</td>
<td align="left">
<xref ref-type="bibr" rid="B151">Orr et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-3-1">
<title>5.3.1 Classical machine learning</title>
<p>Most gross motor task recognition algorithms incorporate classical machine learning using inertial metrics, often measured at central and lower body locations (<xref ref-type="bibr" rid="B37">Cornacchia et al., 2017</xref>), such as the chest (e.g., <xref ref-type="bibr" rid="B118">Li et al., 2010</xref>; <xref ref-type="bibr" rid="B112">Lara et al., 2012</xref>; <xref ref-type="bibr" rid="B19">Braojos et al., 2014</xref>), waist (e.g., <xref ref-type="bibr" rid="B11">Arif et al., 2014</xref>; <xref ref-type="bibr" rid="B209">Weng et al., 2014</xref>; <xref ref-type="bibr" rid="B24">Capela et al., 2015</xref>; <xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>), and thighs (e.g., <xref ref-type="bibr" rid="B19">Braojos et al., 2014</xref>; <xref ref-type="bibr" rid="B58">Gjoreski et al., 2014</xref>; <xref ref-type="bibr" rid="B209">Weng et al., 2014</xref>). Generally, inertial metrics measured at upper peripheral locations (e.g., forearms and wrists) are not suited for detecting gross motor tasks (<xref ref-type="bibr" rid="B97">Kubota et al., 2019</xref>).</p>
<p>Algorithms may also combine inertial data with physiological metrics, such as ECG, heart-rate, respiration rate, or skin temperature (e.g., <xref ref-type="bibr" rid="B157">Parkka et al., 2006</xref>; <xref ref-type="bibr" rid="B84">Jia and Liu, 2013</xref>; <xref ref-type="bibr" rid="B91">Kher et al., 2013</xref>; <xref ref-type="bibr" rid="B156">Park et al., 2017</xref>; <xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>). These algorithms extract time- and frequency-domain features and use conventional classifiers [e.g., Support Vector Machine (SVM) (<xref ref-type="bibr" rid="B84">Jia and Liu, 2013</xref>; <xref ref-type="bibr" rid="B156">Park et al., 2017</xref>), Decision Trees (<xref ref-type="bibr" rid="B157">Parkka et al., 2006</xref>; <xref ref-type="bibr" rid="B195">Tapia et al., 2007</xref>), Random Forest (<xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>; <xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>), or Logistic Regression (<xref ref-type="bibr" rid="B112">Lara et al., 2012</xref>)]. Physiological data can increase recognition accuracy by providing additional context, such as distinguishing between intensity levels [e.g., <italic>running</italic> and <italic>running with weights</italic> (<xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>)]. However, the metrics may also disrupt real-time task recognition, as they are not sensitive to sudden changes in physical activity. For example, incorporating heart-rate reduced performance when heart-rate remained high after performing physically demanding tasks, even when the human was lying or sitting (<xref ref-type="bibr" rid="B195">Tapia et al., 2007</xref>).</p>
<p>Classical machine learning algorithms involving vision-based metrics leverage optical flow extracted from stationary cameras for gross motor task recognition (e.g., <xref ref-type="bibr" rid="B98">Kumar and John, 2016</xref>; <xref ref-type="bibr" rid="B104">Ladjailia et al., 2020</xref>). Task specific motion descriptors derived from optical flow are used as features to train a machine learning classifier (e.g., SVM (<xref ref-type="bibr" rid="B98">Kumar and John, 2016</xref>) or k-Nearest Neighbors (<xref ref-type="bibr" rid="B104">Ladjailia et al., 2020</xref>)).</p>
<p>Generally, classical machine learning based gross motor task detection algorithms typically have high <italic>sensitivity</italic>, primarily due to the atomic and repetitive nature of gross motor tasks. These algorithms conform with <italic>suitability</italic> when the metrics incorporated are wearable and reliable (<xref ref-type="bibr" rid="B170">Ravi et al., 2005</xref>; <xref ref-type="bibr" rid="B157">Parkka et al., 2006</xref>; <xref ref-type="bibr" rid="B195">Tapia et al., 2007</xref>; <xref ref-type="bibr" rid="B112">Lara et al., 2012</xref>; <xref ref-type="bibr" rid="B84">Jia and Liu, 2013</xref>; <xref ref-type="bibr" rid="B91">Kher et al., 2013</xref>; <xref ref-type="bibr" rid="B156">Park et al., 2017</xref>; <xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>; <xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>), and are non-conforming otherwise (<xref ref-type="bibr" rid="B187">Schuldt et al., 2004</xref>; <xref ref-type="bibr" rid="B98">Kumar and John, 2016</xref>; <xref ref-type="bibr" rid="B97">Kubota et al., 2019</xref>; <xref ref-type="bibr" rid="B104">Ladjailia et al., 2020</xref>). Overall, the algorithms conform with <italic>generalizability</italic>, as they typically achieved high accuracy using a leave-one-subject-out cross-validation (<xref ref-type="bibr" rid="B157">Parkka et al., 2006</xref>; <xref ref-type="bibr" rid="B98">Kumar and John, 2016</xref>; <xref ref-type="bibr" rid="B156">Park et al., 2017</xref>; <xref ref-type="bibr" rid="B97">Kubota et al., 2019</xref>; <xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>; <xref ref-type="bibr" rid="B146">Nandy et al., 2020</xref>). All evaluated algorithms are non-conforming for the <italic>concurrency</italic>, <italic>composite factor</italic>, and <italic>anomaly awareness</italic> criteria.</p>
</sec>
<sec id="s5-3-2">
<title>5.3.2 Deep learning methods</title>
<p>Classical machine learning algorithms require handcrafted features that are highly problem-specific, and generalize poorly across task categories (<xref ref-type="bibr" rid="B176">Saez et al., 2017</xref>). Additionally, those algorithms cannot represent the composite relationships among atomic tasks, and require significant human effort to select features and sensor data thresholding (<xref ref-type="bibr" rid="B176">Saez et al., 2017</xref>). Comparative studies indicate deep learning algorithms outperform classical machine learning when large amount of training data is available (<xref ref-type="bibr" rid="B57">Gjoreski et al., 2016</xref>; <xref ref-type="bibr" rid="B176">Saez et al., 2017</xref>; <xref ref-type="bibr" rid="B189">Shakya et al., 2018</xref>).</p>
<p>Deep learning algorithms involving inertial metrics typically require little to no sensor data preprocessing. A Convolutional Neural Network (CNN) detected eight gross motor tasks (e.g., falling, running, jumping, walking, ascending and descending a staircase) using raw acceleration data (<xref ref-type="bibr" rid="B34">Chen and Xue, 2015</xref>). Although inertial data preprocessing is not required, it may be advantageous in some situations. For example, a CNN algorithm transformed the <italic>x</italic>, <italic>y</italic>, and <italic>z</italic> acceleration into vector magnitude data in order to minimize the acceleration&#x2019;s rotational interference (<xref ref-type="bibr" rid="B115">Lee et al., 2017</xref>). The acceleration signal&#x2019;s <italic>spectrogram</italic>, which is a three dimensional representation of changes in the acceleration signal&#x2019;s energy as a function of frequency and time, was used to train a CNN model (<xref ref-type="bibr" rid="B9">Alsheikh et al., 2016</xref>). Employing the spectrogram improved the classification accuracy and reduced the computational complexity significantly (<xref ref-type="bibr" rid="B9">Alsheikh et al., 2016</xref>).</p>
<p>Most recent algorithms leverage publicly available huge benchmark datasets (e.g., <xref ref-type="bibr" rid="B175">Roggen et al., 2010</xref>; <xref ref-type="bibr" rid="B173">Reiss and Stricker, 2012</xref>) to build deeper and more complex task recognition models. Deep learning algorithms combine CNNs with sequential modeling networks (e.g., Long Short-term Memory (LSTMs) (<xref ref-type="bibr" rid="B160">Peng et al., 2018</xref>; <xref ref-type="bibr" rid="B30">Chen L. et al., 2021</xref>), Gated Recurrent Units (GRUs) (<xref ref-type="bibr" rid="B212">Xu et al., 2019</xref>)) to detect composite gross motor tasks from inertial data. The <italic>DEBONAIR</italic> algorithm (<xref ref-type="bibr" rid="B30">Chen L. et al., 2021</xref>) incorporated multiple convolutional sub-networks to extract features based on the input metrics&#x2019; dynamicity and passed the sub-networks&#x2019; feature maps to LSTM networks to detect composite gross motor tasks (e.g., vacuuming, nordic walking, and rope jumping). The <italic>AROMA</italic> algorithm (<xref ref-type="bibr" rid="B160">Peng et al., 2018</xref>) recognized atomic and composite tasks jointly by adopting a CNN &#x2b; LSTM architecture, while <italic>InnoHAR</italic> algorithm (<xref ref-type="bibr" rid="B212">Xu et al., 2019</xref>) combined the <italic>Inception</italic> CNN module with GRUs to detect composite gross motor tasks. Several other algorithms draw inspiration from natural language processing to detect gross motor task transitions (<xref ref-type="bibr" rid="B196">Thu and Han, 2021</xref>) and concurrency (<xref ref-type="bibr" rid="B43">Dirgov&#xe1; Lupt&#xe1;kov&#xe1; et al., 2022</xref>) by utilizing bi-directional LSTMs and Transformers, respectively. Bi-directional LSTMs concatenate information from positive as well as negative time directions in order to predict tasks, whereas Transformers incorporate self-attention mechanisms to draw long-term dependencies by focusing on the most relevant parts of the input sequence.</p>
<p>RFID indoor localization is common for task recognition (e.g., <xref ref-type="bibr" rid="B42">Ding et al., 2015</xref>; <xref ref-type="bibr" rid="B73">Hu et al., 2016</xref>; <xref ref-type="bibr" rid="B48">Fan et al., 2019</xref>). The RFID&#x2019;s <italic>received signal strength indicator</italic> and <italic>phase angle</italic> metrics are used to determine the relative distance and orientation of the tags with respect to the associated embedded environment readers (<xref ref-type="bibr" rid="B185">Scherh&#xe4;ufl et al., 2014</xref>). The two common task identification methods are: i) tag-attached, and ii) tag-free (<xref ref-type="bibr" rid="B48">Fan et al., 2019</xref>). <italic>DeepTag</italic> (<xref ref-type="bibr" rid="B48">Fan et al., 2019</xref>) introduced an advanced RFID-based task recognition algorithm that identified tasks in both tag-attached and tag-free scenarios. The deep learning-based algorithm used a preprocessed <italic>received signal strength indicator</italic> and <italic>phase angle</italic> information that combined a CNN with LSTMs in order to predict seven ADL tasks. Generally, the gross motor task recognition algorithms involving indoor localization have high <italic>sensitivity</italic>, but do not conform with <italic>suitability</italic> and <italic>composite factor</italic>.</p>
<p>Deep learning algorithms&#x2019; increased network complexity and abstraction alleviates most of the classical machine learning algorithms&#x2019; limitation, resulting in high sensitivity, especially when the data is abundant (<xref ref-type="bibr" rid="B57">Gjoreski et al., 2016</xref>; <xref ref-type="bibr" rid="B176">Saez et al., 2017</xref>); however, caution must be exercised to not overfit the algorithms. Deep learning algorithms can achieve high classification accuracy on multi-modal sensor data without requiring special feature engineering for each modality. For example, a hybrid deep learning algorithm trained using an 8-channel sEMG and inertial data detected thirty gym exercises (e.g., dips, bench press, rowing) (<xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>). Deep learning algorithms rarely validate their results via leave-one-subject-out cross-validation, as in most cases the algorithms are validated by splitting all the available data randomly into training and validation datasets; therefore, the algorithms&#x2019; <italic>generalizability</italic> criteria either requires additional evidence, or is non-conforming.</p>
</sec>
<sec id="s5-3-3">
<title>5.3.3 Probabilistic graphical models</title>
<p>Algorithms&#x2019; task predictions are not always accurate, as there is always some uncertainty associated with the predictions, especially when tasks overlap with one another, or share similar motion patterns (e.g., running vs running with weights). Additionally, humans may perform two or more tasks simultaneously, which complicates task identification when using classical and deep learning methods that are typically trained to predict only one task occurring at a time. Probabilistic graphical task recognition algorithms are adept at managing these uncertainties, and have the ability to model simultaneous tasks.</p>
<p>Probabilistic graphical models can detect gross motor tasks across various metrics [e.g., indoor localization (<xref ref-type="bibr" rid="B73">Hu et al., 2016</xref>), sEMG (<xref ref-type="bibr" rid="B198">Trigili et al., 2019</xref>), inertial (<xref ref-type="bibr" rid="B116">Lee and Cho, 2011</xref>; <xref ref-type="bibr" rid="B92">Kim et al., 2015</xref>), human-body pose (<xref ref-type="bibr" rid="B83">Jalal et al., 2017</xref>), optical flow (<xref ref-type="bibr" rid="B93">Kolekar and Dash, 2016</xref>), and object detection (<xref ref-type="bibr" rid="B223">Zhang et al., 2013</xref>)]. Hidden Markov Models are the most widely utilized probabilistic graphical algorithm for gross motor task recognition (e.g., <xref ref-type="bibr" rid="B116">Lee and Cho, 2011</xref>; <xref ref-type="bibr" rid="B92">Kim et al., 2015</xref>; <xref ref-type="bibr" rid="B93">Kolekar and Dash, 2016</xref>; <xref ref-type="bibr" rid="B83">Jalal et al., 2017</xref>), because Hidden Markov Model&#x2019;s sequence modeling properties can be exploited for continuous task recognition (<xref ref-type="bibr" rid="B92">Kim et al., 2015</xref>). Hidden Markov Models also allow for modeling the tasks hierarchically (<xref ref-type="bibr" rid="B116">Lee and Cho, 2011</xref>), and can distinguish tasks with intra-class variances and inter-class similarities (<xref ref-type="bibr" rid="B92">Kim et al., 2015</xref>). Other probabilistic models [e.g., Gaussian Mixture Models (<xref ref-type="bibr" rid="B198">Trigili et al., 2019</xref>)] can also detect gross motor tasks. A probabilistic graphical model, the <italic>Interval-temporal Bayesian Network</italic>, unified Bayesian network&#x2019;s probabilistic representation with interval algebra&#x2019;s (<xref ref-type="bibr" rid="B8">Allen and Ferguson, 1994</xref>) ability to represent temporal relationships between atomic events (<xref ref-type="bibr" rid="B223">Zhang et al., 2013</xref>) to detect composite and concurrent gross motor tasks. The algorithm&#x2019;s <italic>sensitivity</italic> and <italic>generalizability</italic> are low and non-conforming, respectively. The algorithm&#x2019;s <italic>suitability</italic> is non-conforming, as it employed vision-based metrics. Finally, the algorithm&#x2019;s <italic>composite factor</italic> and <italic>concurrency</italic> conform.</p>
</sec>
<sec id="s5-3-4">
<title>5.3.4 Knowledge-driven algorithms</title>
<p>Gross motor rule-based task recognition algorithms incorporate template matching or thresholding to recognize tasks. A <italic>Dynamic Time Warping</italic> (<xref ref-type="bibr" rid="B181">Salvador and Chan, 2007</xref>) based algorithm detected free-weight exercises by computing the similarity between Doppler shift profiles of the reflected RFID signals (<xref ref-type="bibr" rid="B42">Ding et al., 2015</xref>). A principal component analysis thresholding algorithm detected ambulatory task transitions by analyzing the motion artifacts in ECG data induced by body movements (<xref ref-type="bibr" rid="B158">Pawar et al., 2007</xref>; <xref ref-type="bibr" rid="B159">2006</xref>).</p>
<p>Rule-based algorithms can detect concurrent tasks, if the rules are relatively simple to derive using the sensor data. A multiagent algorithm (<xref ref-type="bibr" rid="B151">Orr et al., 2018</xref>) detected up to seven gross motor atomic tasks (e.g., dressing, cleaning, and food preparation). The algorithm detected up to two concurrent tasks using environmentally-embedded proximity sensors.</p>
<p>Rule-based systems are ideal for gross motor task detection when the sensor data is limited and can be comprehended in a relatively straightforward manner. For example, the prior rule-based multiagent algorithm detected concurrent tasks, as it was easy to form the rules using the proximity sensor data. Rule-based algorithms are unsuitable when the sensor data cannot be interpreted easily (i.e., instances of high dimensionality), or when there are a large number of tasks that have intricate relationships.</p>
</sec>
<sec id="s5-3-5">
<title>5.3.5 Discussion</title>
<p>Most machine learning based algorithms can detect gross motor tasks reliably with acceptable suitability and generalizability when the tasks are atomic and non-concurrent with repetitive motions (e.g., <xref ref-type="bibr" rid="B157">Parkka et al., 2006</xref>; <xref ref-type="bibr" rid="B34">Chen and Xue, 2015</xref>; <xref ref-type="bibr" rid="B156">Park et al., 2017</xref>; <xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>). The human-robot teaming domain often involves composite tasks that may occur concurrently. None of the existing gross motor task detection algorithms satisfy all the required criteria for the intended domain.</p>
<p>The interval-temporal algorithm (<xref ref-type="bibr" rid="B223">Zhang et al., 2013</xref>) is the preferred approach for gross motor task detection. The algorithm can detect concurrent and composite tasks, but had low sensitivity and is non-conforming for suitability and generalizability, which can be attributed to the vision-based metrics and low-level Bayesian network&#x2019;s poor classification accuracy. However, the algorithm is independent of the metrics (<xref ref-type="bibr" rid="B223">Zhang et al., 2013</xref>), as it operates hierarchically, utilizing the low-level atomic event predictions. Therefore, a modified version more suited to the intended domain may incorporate a classical machine learning algorithm [e.g., Random Forest (<xref ref-type="bibr" rid="B7">Allahbakhshi et al., 2020</xref>)] or a deep network [e.g., CNN (<xref ref-type="bibr" rid="B34">Chen and Xue, 2015</xref>)], depending on the amount of data available, to detect the low-level atomic tasks using inertial metrics. The interval-temporal algorithm can be used to detect the composite and concurrent gross motor tasks.</p>
</sec>
</sec>
<sec id="s5-4">
<title>5.4 Fine-grained motor tasks</title>
<p>Fine-grained motor tasks often involve highly articulated and dexterous motions that can be performed in multiple ways. The execution and the time taken to complete the tasks differ from one human to the other. These aspects of fine-grained motor tasks can create ambiguity in the sensor data, making it difficult for the algorithms to detect such tasks; therefore, a wide range of methods adopting various sensing modalities exist for detecting fine-grained tasks accurately. The evaluation criteria for each reviewed fine-grained task recognition algorithm by algorithm category is provided in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Fine-grained motor task recognition algorithms&#x2019; evaluation overview.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Classical Machine Learning</td>
</tr>
<tr>
<td align="left">Ensemble</td>
<td align="left">
<xref ref-type="bibr" rid="B143">Min and Cho (2011)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="3" align="left">k-Nearest Neighbors</td>
<td align="left">
<xref ref-type="bibr" rid="B95">Koskimaki et al. (2009)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B97">Kubota et al. (2019)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B104">Ladjailia et al. (2020)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="3" align="left">Random forest</td>
<td align="left">
<xref ref-type="bibr" rid="B204">Wang et al. (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B120">Li et al. (2016a)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B69">Heard et al. (2019b)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="4" align="left">SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B110">Laput et al. (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B163">Pirsiavash and Ramanan (2012)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B136">Matsuo et al. (2014)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B221">Zhang and Harrison (2015)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td rowspan="5" align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B109">Laput and Harrison (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B121">Li et al. (2016b)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B117">Li et al. (2020a)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B130">Ma et al. (2016)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B25">Castro et al. (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">CNN &#x2b; LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B200">Ullah et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B53">Frank et al. (2019)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B55">Garcia-Hernando et al. (2018)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">LSTM bi-directional</td>
<td align="left">
<xref ref-type="bibr" rid="B225">Zhao et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">Residual &#x2b; Attention</td>
<td align="left">
<xref ref-type="bibr" rid="B138">Mekruksavanich et al. (2022)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B6">Al-qaness et al. (2022)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Transformer</td>
<td align="left">
<xref ref-type="bibr" rid="B222">Zhang et al. (2022b)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Probabilistic Graphical Model</td>
</tr>
<tr>
<td rowspan="4" align="left">Bayesian network</td>
<td align="left">
<xref ref-type="bibr" rid="B127">Liu et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">RE</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B211">Wu et al. (2007)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B51">Fortin-Simard et al. (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B71">Hevesi et al. (2018)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Conditional random field</td>
<td align="left">
<xref ref-type="bibr" rid="B49">Fathi et al. (2011)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">Gaussian mixture model</td>
<td align="left">
<xref ref-type="bibr" rid="B142">Min et al. (2007)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B141">Min et al. (2008)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
</tr>
<tr>
<td align="left">Hierarchical latent SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B125">Lillo et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Temporal memory</td>
<td align="left">
<xref ref-type="bibr" rid="B216">Zhang et al. (2008)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Markov chains</td>
<td align="left">
<xref ref-type="bibr" rid="B178">Saguna et al. (2013)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Probabilistic NN</td>
<td align="left">
<xref ref-type="bibr" rid="B207">Wang et al. (2012)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Temporal graph</td>
<td align="left">
<xref ref-type="bibr" rid="B123">Liao et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">RE</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-4-1">
<title>5.4.1 Classical machine learning</title>
<p>Classical machine learning algorithms are suitable for detecting fine-grained motor tasks only when the tasks are short in duration, atomic, or repetitive (<xref ref-type="bibr" rid="B66">Heard et al., 2019a</xref>). Among the classical machine learning algorithms, k-Nearest Neighbors (e.g., <xref ref-type="bibr" rid="B95">Koskimaki et al., 2009</xref>; <xref ref-type="bibr" rid="B97">Kubota et al., 2019</xref>; <xref ref-type="bibr" rid="B104">Ladjailia et al., 2020</xref>), Random Forest (e.g., <xref ref-type="bibr" rid="B204">Wang et al., 2015</xref>; <xref ref-type="bibr" rid="B120">Li et al., 2016a</xref>; <xref ref-type="bibr" rid="B69">Heard et al., 2019b</xref>), and SVM (e.g., <xref ref-type="bibr" rid="B143">Min and Cho, 2011</xref>; <xref ref-type="bibr" rid="B163">Pirsiavash and Ramanan, 2012</xref>; <xref ref-type="bibr" rid="B136">Matsuo et al., 2014</xref>) are the most popular choices for fine-grained motor task detection.</p>
<p>Several classical machine learning algorithms use egocentric wearable camera videos for detecting ADL tasks (e.g., <xref ref-type="bibr" rid="B163">Pirsiavash and Ramanan, 2012</xref>; <xref ref-type="bibr" rid="B136">Matsuo et al., 2014</xref>; <xref ref-type="bibr" rid="B55">Garcia-Hernando et al., 2018</xref>). Image processing techniques (e.g., histogram of orientation or spatial pyramids) are used to detect objects and conventional machine learning algorithms recognize the tasks from the detected objects. These algorithms may also incorporate saliency detectors (<xref ref-type="bibr" rid="B136">Matsuo et al., 2014</xref>), or depth information (<xref ref-type="bibr" rid="B55">Garcia-Hernando et al., 2018</xref>) to identify the objects being manipulated. Temporal motion descriptive features from optical flow can also be used for recognizing fine-grained tasks. A k-Nearest Neighbors algorithm classified the fine-grained motor tasks (<xref ref-type="bibr" rid="B104">Ladjailia et al., 2020</xref>) based on a histogram constructed using the motion descriptors from optical flow.</p>
<p>Forearm sEMG signals can detect tasks that are difficult for a vision-based algorithm to differentiate when using the same conventional classifiers. A comparison between an sEMG [i.e., Myo armband (<xref ref-type="bibr" rid="B183">Sathiyanarayanan and Rajan, 2016</xref>)] and a motion capture sensor revealed that the former had higher efficacy in recognizing fine-grained motions (e.g., grasps and assembly part manipulation tasks) (<xref ref-type="bibr" rid="B97">Kubota et al., 2019</xref>). The classifiers with the sEMG data detected the minute variation in the muscle associated with each grasp; resulting in significantly higher recognition accuracy than using the motion capture data.</p>
<p>Some classical machine learning algorithms that use a single Inertial Measurement Unit (IMU) can classify fine-grained ADL tasks [e.g., eating and drinking (<xref ref-type="bibr" rid="B216">Zhang et al., 2008</xref>)], and assembly line activities [e.g., hammering and tightening screws <xref ref-type="bibr" rid="B95">Koskimaki et al. (2009)</xref>]. This approach is suitable for tasks involving a single hand (i.e., the dominant), when the number of recognized tasks is small (e.g., <inline-formula id="inf14">
<mml:math id="m14">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>5</mml:mn>
</mml:math>
</inline-formula>). For instance, five assembly line tasks were recognized using a wrist worn IMU&#x2019;s acceleration and angular velocity data (<xref ref-type="bibr" rid="B95">Koskimaki et al., 2009</xref>). The associated time- and frequency-domain features were used to train a k-Nearest Neighbors algorithm to classify the tasks. A two-stage classification approach using acceleration metrics obtained by a wrist-worn accelerometer recognized eating and drinking (<xref ref-type="bibr" rid="B216">Zhang et al., 2008</xref>). However, when the tasks are composite or larger in number, the algorithms augment the IMU with different sensing modalities. Algorithms typically combine IMU with sEMG metrics measured at upper peripheral locations, such as the forearms and wrists, in order to capture highly articulated motions (<xref ref-type="bibr" rid="B97">Kubota et al., 2019</xref>). Increasing sensing modalities provides more task context, enabling an algorithm to discriminate a broader set of tasks.</p>
<p>A recent system attempted to recognize twenty-three composite clinical procedures by using metrics from two Myo armbands and statically embedded cameras (<xref ref-type="bibr" rid="B66">Heard et al., 2019a</xref>). The Myo&#x2019;s sEMG and inertial metrics were combined with the camera&#x2019;s human body pose metric to train a Random Forest classifier with majority voting. Many clinical procedures require multiple articulated fine-grained motions that range from <inline-formula id="inf15">
<mml:math id="m15">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>10</mml:mn>
<mml:mi>s</mml:mi>
</mml:math>
</inline-formula> to <inline-formula id="inf16">
<mml:math id="m16">
<mml:mo>&#x3e;</mml:mo>
<mml:mn>60</mml:mn>
<mml:mi>s</mml:mi>
</mml:math>
</inline-formula> to complete. Long-duration fine-grained procedures are difficult to detect due to intra-class variability, inter-class similarity, and individual differences among participants. The video provided contextual information that improved the procedure recognition accuracy by alleviating intra-class variance and inter-class similarity.</p>
<p>A multi-modal framework, incorporating five inertial sensors, data gloves and a bio-signal sensor, detected eleven atomic tasks (e.g., writing, brushing, typing) and eight composite tasks (e.g., exercising, working, meeting) (<xref ref-type="bibr" rid="B143">Min and Cho, 2011</xref>). A hybrid ensemble approach combined classifier selection and output fusion. The sensors&#x2019; inputs were initially recognized by a Naive Bayes selection module. The selection module&#x2019;s task probabilities chose a set of task-specific SVM classifiers that fused their predictions into a matrix in order to identify the tasks.</p>
<p>Classical machine learning algorithms&#x2019; classification accuracies range between 45% and 65%; thus, they generally have low <italic>sensitivity</italic>. The algorithms&#x2019; <italic>suitability</italic> criterion depend on the metrics employed. The <italic>composite factor</italic> and <italic>generalizability</italic> criteria also vary across algorithms, as they depend on the tasks detected and the validation methodology. Overall, most algorithms are non-conforming for the <italic>concurrency</italic> and <italic>composite</italic> factors, making them unsuitable for detecting fine-grained motor tasks for the intended HRT domain.</p>
</sec>
<sec id="s5-4-2">
<title>5.4.2 Deep learning</title>
<p>The ambiguous, convoluted sensor data from fine-grained motor tasks causes the feature engineering and extraction to be laborious. Deep learning algorithms overcome this limitation by automating the feature extraction process. There are three different types of deep learning algorithms for fine-grained motor task recognition: i) Convolutional, ii) Recurrent, and iii) Hybrid. Convolutional algorithms typically incorporate only CNNs to learn the spatial features from sensor data for each task and distinguish them by comparing the spatial patterns (e.g., <xref ref-type="bibr" rid="B25">Castro et al., 2015</xref>; <xref ref-type="bibr" rid="B121">Li et al., 2016b</xref>; <xref ref-type="bibr" rid="B130">Ma et al., 2016</xref>; <xref ref-type="bibr" rid="B109">Laput and Harrison, 2019</xref>; <xref ref-type="bibr" rid="B117">Li L. et al., 2020</xref>). Recurrent algorithms detect the tasks by capturing the sequential information present in the sensor data, typically using memory cells (e.g., <xref ref-type="bibr" rid="B55">Garcia-Hernando et al., 2018</xref>; <xref ref-type="bibr" rid="B225">Zhao et al., 2018</xref>). Hybrid algorithms extract spatial features and learn the temporal relationships simultaneously by combining convolutional and recurrent networks (<xref ref-type="bibr" rid="B200">Ullah et al., 2018</xref>; <xref ref-type="bibr" rid="B48">Fan et al., 2019</xref>; <xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>).</p>
<p>Deep learning algorithms using egocentric videos from wearable cameras combine object detection with task recognition. A CNN with a late fusion ensemble predicted the tasks from a chest-mounted wearable camera (<xref ref-type="bibr" rid="B25">Castro et al., 2015</xref>) by incorporating relevant contextual information (e.g., time and day of the week) to boost the classification accuracy. Two separate CNNs were combined together to recognize objects of interest and hand motions (<xref ref-type="bibr" rid="B130">Ma et al., 2016</xref>). The networks were fine tuned jointly using a triplet loss function to recognize fine-grained ADL tasks with medium to high <italic>sensitivity</italic>.</p>
<p>Analyzing changes in body poses spatially and temporally can provide important cues for fine-grained motor task recognition (<xref ref-type="bibr" rid="B125">Lillo et al., 2017</xref>). An end-to-end CNN network exploited camera images for estimating fifteen upper body joint positions (<xref ref-type="bibr" rid="B147">Neili Boualia and Essoukri Ben Amara, 2021</xref>). The estimated joint positions permitted discriminating features to recognize tasks. The CNN architecture had two levels: i) fully-convolutional layers that extracted the salient feature, or heat maps, and ii) fusion layers that learned the spatial dependencies between the joints by concatenating the convolutional layers. The CNN-estimated joint positions served as input to train a multi-class SVM that predicted twelve ADL tasks with high <italic>sensitivity</italic>.</p>
<p>Hybrid deep learning algorithms are becoming increasingly popular for task recognition across metrics (<xref ref-type="bibr" rid="B200">Ullah et al., 2018</xref>; <xref ref-type="bibr" rid="B48">Fan et al., 2019</xref>; <xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>). An optical flow-based algorithm (<xref ref-type="bibr" rid="B200">Ullah et al., 2018</xref>) leveraged deep learning to extract temporal optical flow features from the salient frames, and incorporated a multilayer LSTM to predict the tasks using the temporal optical flow features. Another hybrid deep learning algorithm (<xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>) trained on the sEMG and inertial metrics detected assembly tasks. The algorithm&#x2019;s CNN layers extracted spatial features from the merits at each timestep, while the LSTM layers learned how the spatial features evolved temporally. Hybrid algorithms can provide excellent expressive and predictive capabilities; however, these algorithms&#x2019; performance relies heavily on the size of training dataset (<xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>).</p>
<p>A CNN-based algorithm incorporated inertial metrics from an off-the-shelf smartwatch to detect twenty-five atomic tasks (e.g., operating a drill, cutting paper, and writing) (<xref ref-type="bibr" rid="B109">Laput and Harrison, 2019</xref>). A Fourier transform was applied to the acceleration data to obtain the corresponding spectrograms. The CNN identified the spatial-temporal relationships encoded in the spectrograms by generating distinctive activation patterns for each task. The algorithm also rejected (i.e., detected) unknown instances.</p>
<p>Deep learning algorithms can recognize concurrent and composite fine-grained motor tasks directly from raw sensor data using complex network architectures, provided sufficient data is available (<xref ref-type="bibr" rid="B150">Ord&#xf3;&#xf1;ez and Roggen, 2016</xref>; <xref ref-type="bibr" rid="B225">Zhao et al., 2018</xref>). Human task trajectories are continuous in that the current task depends on both past and future information. A deep residual bidirectional LSTM algorithm (<xref ref-type="bibr" rid="B225">Zhao et al., 2018</xref>) detected the <italic>Opportunity</italic> dataset&#x2019;s composite tasks by incorporating information from positive as well as negative time directions. The dataset contains five composite ADLs (e.g., relaxation, preparing coffee, preparing breakfast, grooming, cleaning), involving a total number of 211 atomic events (e.g., walk, sit, lying, open doors, reach for an object). Several metrics, including acceleration and orientation of various body parts, and three-dimensional indoor position were gathered.</p>
<p>Recent deep learning algorithms leverage attention mechanisms to model long-term dependencies from inertial data (<xref ref-type="bibr" rid="B6">Al-qaness et al., 2022</xref>; <xref ref-type="bibr" rid="B222">Zhang Y. et al., 2022</xref>; <xref ref-type="bibr" rid="B138">Mekruksavanich et al., 2022</xref>). The <italic>ResNet-SE</italic> algorithm (<xref ref-type="bibr" rid="B138">Mekruksavanich et al., 2022</xref>) classified composite fine-grained motor tasks on three publicly available datasets. The algorithm incorporated residual networks to address loss degradation, followed by a squeeze-and-excite attention function to modulate the relevance of each residual feature map. The <italic>Multi-ResAtt</italic> algorithm (<xref ref-type="bibr" rid="B6">Al-qaness et al., 2022</xref>) incorporated residual networks to process inertial metrics from IMUs distributed over different body locations, followed by bidirectional GRUs with attention mechanism to learn time-series features.</p>
<p>Generally, deep learning algorithms are highly effective at detecting atomic fine-grained motor tasks, but their ability to detect composite and concurrent tasks reliably is indeterminate. The latter may be due to insufficient ecologically-valid composite, concurrent task recognition datasets available publicly. Utilizing generative adversarial networks (<xref ref-type="bibr" rid="B205">Wang et al., 2018</xref>; <xref ref-type="bibr" rid="B119">Li X. et al., 2020</xref>) to expand datasets by producing synthetic sensor data may alleviate the issue.</p>
</sec>
<sec id="s5-4-3">
<title>5.4.3 Probabilistic graphical models</title>
<p>Bayesian networks are the most common probabilistic graphical models for fine-grained motor task detection (e.g., <xref ref-type="bibr" rid="B211">Wu et al., 2007</xref>; <xref ref-type="bibr" rid="B51">Fortin-Simard et al., 2015</xref>; <xref ref-type="bibr" rid="B127">Liu et al., 2017</xref>; <xref ref-type="bibr" rid="B71">Hevesi et al., 2018</xref>), followed by Gaussian Mixture Models (e.g., <xref ref-type="bibr" rid="B142">Min et al., 2007</xref>; <xref ref-type="bibr" rid="B141">Min et al., 2008</xref>). Many such algorithms augment the inertial data with a different sensing modality (e.g., <xref ref-type="bibr" rid="B142">Min et al., 2007</xref>; <xref ref-type="bibr" rid="B211">Wu et al., 2007</xref>; <xref ref-type="bibr" rid="B141">Min et al., 2008</xref>; <xref ref-type="bibr" rid="B71">Hevesi et al., 2018</xref>) in order to provide more task context, which enables discriminating a broader set of tasks. Recognition of up to three day-to-day early morning tasks (<xref ref-type="bibr" rid="B142">Min et al., 2007</xref>) augmented with a microphone, resulted in the recognition of six tasks (<xref ref-type="bibr" rid="B141">Min et al., 2008</xref>). The intended HRT domain requires multiple sensors to detect tasks belonging to different activity components, although adding new modalities arbitrarily may deteriorate the classifier performance (<xref ref-type="bibr" rid="B53">Frank et al., 2019</xref>).</p>
<p>Hierarchical graphical models detect composite tasks by decomposing them into a set of smaller classification problems. <xref ref-type="bibr" rid="B49">Fathi et al.&#x2019;s (2011)</xref> meal preparation task detection algorithm decomposed hand manipulations into numerous atomic actions, and learned tasks from a hierarchical action sequence using conditional random fields. Another hierarchical model that operated at three levels of abstraction detected composite, concurrent tasks using body poses (<xref ref-type="bibr" rid="B125">Lillo et al., 2017</xref>).</p>
<p>Identifying the causality (i.e., action and reaction pair) between two events allows for easier human interpretation, and for modeling far more intricate temporal relationships (<xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>). A graphical algorithm incorporated the Granger-causality (<xref ref-type="bibr" rid="B61">Granger, 1969</xref>; <xref ref-type="bibr" rid="B62">1980</xref>) test for uncovering cause-effect relationships among atomic events (<xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>). The algorithm employed a generic Bayesian network to detect the atomic events. A temporal causal graph was generated via the Granger-causality test between atomic events. Each graph represented a particular task instance. The graph nodes represented the atomic events and directed links with weights represented the cause-effect relationships between the atomic events. An artificial neural network is trained using these graphs as inputs to predict the composite, concurrent tasks. The algorithm was evaluated on the <italic>Opportunity</italic> (<xref ref-type="bibr" rid="B175">Roggen et al., 2010</xref>) and <italic>OSUPEL</italic> (<xref ref-type="bibr" rid="B20">Brendel et al., 2011</xref>) datasets, indicating that the algorithm is independent of the metrics.</p>
<p>Overall, probabilistic graphical models typically have high <italic>sensitivity</italic> for detecting fine-grained motor tasks. The algorithms, especially hierarchical (e.g., <xref ref-type="bibr" rid="B125">Lillo et al., 2017</xref>) and the Granger-causality based temporal graph (<xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>), are independent of the metrics due to data abstraction; therefore, their <italic>suitability</italic> is classified as conforming. Most task recognition algorithms are susceptible to individual differences (see <xref ref-type="table" rid="T3">Table 3</xref>). Even those that conform with generalizability may experience a significant decrease in accuracy when classifying an unknown human&#x2019;s data (<xref ref-type="bibr" rid="B109">Laput and Harrison, 2019</xref>); thus, the <italic>generalizability</italic> criterion requires additional evidence. Algorithms can only identify tasks reliably for humans on which they were trained, suggesting that online and self-learning mechanisms are needed to accommodate new humans (<xref ref-type="bibr" rid="B207">Wang et al., 2012</xref>). The <italic>composite factor</italic> and <italic>concurrency</italic> vary across algorithms, but are non-conforming overall. The <italic>anomaly awareness</italic> criterion is classified as non-conforming, as most probabilistic graphical models do not detect out-of-class tasks.</p>
</sec>
<sec id="s5-4-4">
<title>5.4.4 Discussion</title>
<p>Classical machine learning algorithms are unreliable for detecting fine-grained motor tasks due to poor sensitivity and generalizability. Deep learning algorithms can detect the atomic fine-grained motor tasks reliably, but not composite, concurrent tasks. Moreover, deep learning typically requires a large number of parameters, very large datasets and can be difficult to train (<xref ref-type="bibr" rid="B125">Lillo et al., 2017</xref>). Deep learning&#x2019;s automatic feature learning capability prohibits exploiting explicit relationships among tasks and semantic knowledge, making it difficult to detect composite, concurrent fine-grained motor tasks. Probabilistic graphical models offer some suitable alternatives; however, none of the existing algorithms satisfy all the required criteria for the intended domain.</p>
<p>The Granger-causality based temporal graph algorithm (<xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>) and the three-level hierarchical algorithm (<xref ref-type="bibr" rid="B125">Lillo et al., 2017</xref>) are the most suitable for fine-grained motor task detection given all the other algorithms. Both algorithms have high sensitivity and can detect concurrent and composite tasks. The Granger-causality algorithm conforms with suitability, but requires additional evidence to substantiate its generalizability. The hierarchical algorithm conforms with generalizability, but is non-conforming with suitability, as it employed a vision-based system for estimating human-body pose metric. However, the metric can be estimated using a series of inertial motion trackers (<xref ref-type="bibr" rid="B47">Faisal et al., 2019</xref>); therefore, a human-robot teaming domain friendly version of both algorithms can be developed theoretically.</p>
</sec>
</sec>
<sec id="s5-5">
<title>5.5 Tactile tasks</title>
<p>Tactile interaction occurs when humans interact with objects around them (e.g., keyboard typing, mouse-clicking and finger gestures). Individual classifications for each tactile task algorithm by its category are provided in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Tactile task recognition algorithms&#x2019; evaluation overview.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Classical Machine Learning</td>
</tr>
<tr>
<td align="left">Decision trees</td>
<td align="left">
<xref ref-type="bibr" rid="B85">Jing et al. (2011)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Ensemble</td>
<td align="left">
<xref ref-type="bibr" rid="B26">Cha et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B72">Hsiao et al. (2015)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Voting</td>
<td align="left">
<xref ref-type="bibr" rid="B188">Sha et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B38">C&#xf4;t&#xe9;-Allard et al. (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">CNN &#x2b; LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B168">Rahimian et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Probabilistic Graphical Model</td>
</tr>
<tr>
<td align="left">Gaussian mixture model</td>
<td align="left">
<xref ref-type="bibr" rid="B86">Ju and Liu (2014)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Hidden Markov model</td>
<td align="left">
<xref ref-type="bibr" rid="B218">Zhang et al. (2011)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">Naive bayes classifier</td>
<td align="left">
<xref ref-type="bibr" rid="B33">Chen et al. (2007b)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B32">Chen et al. (2007a)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-5-1">
<title>5.5.1 Classical machine learning</title>
<p>Most classical machine learning algorithms incorporate inertial metrics measured at the fingers or dorsal side of the hand. These approaches typically detect finger gestures and keystrokes depending on the measurement site (e.g., <xref ref-type="bibr" rid="B85">Jing et al., 2011</xref>; <xref ref-type="bibr" rid="B26">Cha et al., 2018</xref>; <xref ref-type="bibr" rid="B128">Liu et al., 2018</xref>; <xref ref-type="bibr" rid="B188">Sha et al., 2020</xref>). Several of these approaches use multiple ring-like accelerometer device worn on the fingers (e.g., <xref ref-type="bibr" rid="B85">Jing et al., 2011</xref>; <xref ref-type="bibr" rid="B226">Zhou et al., 2015</xref>; <xref ref-type="bibr" rid="B188">Sha et al., 2020</xref>). The time- and frequency-domain features (e.g., minimum, maximum, standard deviation, energy, and entropy) extracted from the acceleration signals were used to train classical machine learning algorithms (e.g., decision tree classifier and majority voting) to detect finger gestures (e.g., finger rotation and bending) and keystrokes. Although these approaches incorporated inertial metrics, none conform with <italic>suitability</italic> due to lack of reproducibility (i.e., the ring-like sensor is not commercially available) and wearing a ring-like device may hinder humans&#x2019; dexterity, impacting task performance negatively.</p>
<p>Inertial metrics from the dorsal side of the hand detected seven office tasks (e.g., keyboard typing, mouse-clicking, writing) (<xref ref-type="bibr" rid="B26">Cha et al., 2018</xref>). Time- and frequency-domain features extracted from the acceleration signals were used to train an ensemble classifier. The algorithm achieved high accuracy <inline-formula id="inf17">
<mml:math id="m17">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>90</mml:mn>
<mml:mi>%</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> in an in-the-wild evaluation. Most misclassifications occurred during transitions between tasks, implying that inertial-based tactile task recognition may be susceptible to task transitions due to signal variations. The high error rates during transitions can lead to lower classification accuracy, especially when tasks switch frequently.</p>
<p>Classical machine learning algorithms&#x2019; generally have high <italic>sensitivity</italic>. The algorithms&#x2019; are typically non-conforming for the <italic>suitability</italic> criterion, as many supporting research efforts focus on developing and validating new sensor technology for sensing tactility, rather than detecting tactile tasks (e.g., <xref ref-type="bibr" rid="B80">Iwamoto and Shinoda, 2007</xref>; <xref ref-type="bibr" rid="B152">Ozioko et al., 2017</xref>; <xref ref-type="bibr" rid="B88">Kawazoe et al., 2019</xref>). The <italic>generalizability</italic> criteria also vary across algorithms, as they depend on the validation methodology. Finally, the algorithms are non-conforming for the <italic>concurrency</italic> and <italic>composite</italic> factors, making them unsuitable for detecting tactile tasks for the intended HRT domain.</p>
</sec>
<sec id="s5-5-2">
<title>5.5.2 Deep learning</title>
<p>Several publicly available sEMG-based hand gesture datasets (e.g., <xref ref-type="bibr" rid="B13">Atzori et al., 2014</xref>; <xref ref-type="bibr" rid="B10">Amma et al., 2015</xref>; <xref ref-type="bibr" rid="B87">Kaczmarek et al., 2019</xref>) support deep learning algorithms to detect tactile hand gestures (e.g., <xref ref-type="bibr" rid="B38">C&#xf4;t&#xe9;-Allard et al., 2019</xref>; <xref ref-type="bibr" rid="B168">Rahimian et al., 2020</xref>). A hybrid deep learning model consisting of two parallel paths (i.e., one LSTM path and one CNN path) was developed (<xref ref-type="bibr" rid="B168">Rahimian et al., 2020</xref>). A fully connected multilayer fusion network combined the outputs of the two paths to classify the hand gestures.</p>
<p>Recognizing tactile tasks is an under-developed area of research, as the tasks are nuanced and often overshadowed by fine-grained motor tasks. Generally, deep learning algorithms have high <italic>sensitivity</italic>; however, the incorporated sEMG metrics with a random dataset split for validation cause them to not conform with the <italic>suitability</italic> and <italic>generalizability</italic> criteria.</p>
</sec>
<sec id="s5-5-3">
<title>5.5.3 Probabilistic graphical model</title>
<p>Probabilistic graphical models for tactile task recognition typically involve simple algorithms (e.g., Hidden Markov Models (<xref ref-type="bibr" rid="B218">Zhang et al., 2011</xref>), Gaussian Mixture Models (<xref ref-type="bibr" rid="B86">Ju and Liu, 2014</xref>), and Bayesian Networks (<xref ref-type="bibr" rid="B32">Chen et al., 2007a</xref>; <xref ref-type="bibr" rid="B33">b</xref>)) when compared to the prior gross motor and fine-grained motor sections (see Sections 5.3.3 and 5.4.3), as the tasks detected are inherently atomic (e.g., hand and finger gestures). sEMG signals are one of the most frequently used metrics for detecting hand and finger gestures (e.g., <xref ref-type="bibr" rid="B33">Chen et al., 2007b</xref>; <xref ref-type="bibr" rid="B218">Zhang et al., 2011</xref>; <xref ref-type="bibr" rid="B38">C&#xf4;t&#xe9;-Allard et al., 2019</xref>; <xref ref-type="bibr" rid="B168">Rahimian et al., 2020</xref>). <xref ref-type="bibr" rid="B33">Chen et al.&#x2019;s (2007b)</xref> gesture recognition algorithm pioneered the use of sEMG signals. Twenty-five hand gestures (i.e., six wrist actions and seventeen finger gestures) were detected using a 2-channel sEMG placed on the forearm. A Bayesian classifier was trained using the mean absolute value and autoregressive model coefficients extracted from the sEMG. The algorithm was extended to include two accelerometers, one placed on the wrist and the other placed on the dorsal side of the hand (<xref ref-type="bibr" rid="B32">Chen et al., 2007a</xref>).</p>
<p>Overall, probabilistic graphical models also tend to have high <italic>sensitivity</italic> for detecting tactile tasks. Most algorithms are susceptible to individual differences and incorporate sEMG metrics; thus, the algorithms are non-conforming for the <italic>suitability</italic> and <italic>generalizability</italic> criteria. Additionally, the algorithms are non-conforming for the <italic>concurrency</italic> and <italic>composite</italic> factors, as the evaluated tactile tasks are inherently atomic.</p>
</sec>
<sec id="s5-5-4">
<title>5.5.4 Discussion</title>
<p>All data-driven algorithms can detect tactile tasks with <inline-formula id="inf18">
<mml:math id="m18">
<mml:mo>&#x3e;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy (<xref ref-type="bibr" rid="B32">Chen et al., 2007a</xref>; <xref ref-type="bibr" rid="B85">Jing et al., 2011</xref>; <xref ref-type="bibr" rid="B218">Zhang et al., 2011</xref>; <xref ref-type="bibr" rid="B86">Ju and Liu, 2014</xref>; <xref ref-type="bibr" rid="B72">Hsiao et al., 2015</xref>; <xref ref-type="bibr" rid="B26">Cha et al., 2018</xref>; <xref ref-type="bibr" rid="B168">Rahimian et al., 2020</xref>; <xref ref-type="bibr" rid="B188">Sha et al., 2020</xref>) primarily because the detected tasks (i.e., finger and hand gestures) were atomic; therefore, the algorithms have high <italic>sensitivity</italic>. Except for the office-based tactile task classifier (<xref ref-type="bibr" rid="B26">Cha et al., 2018</xref>), none of the existing algorithms conform with <italic>suitability</italic>, because either the sensors incorporated were commercially unavailable for reproducibility, or the metrics employed were unreliable. All the algorithms are non-conforming with the <italic>concurrency</italic> and <italic>composite factor</italic> criteria, because tactile tasks are rarely composite or concurrent. Finally, none of the algorithms detect out-of-class instances; therefore, they do not conform with <italic>anomaly awareness</italic>. A recommended tactile task detection algorithm to support the intended domain is the interval-temporal algorithm (<xref ref-type="bibr" rid="B223">Zhang et al., 2013</xref>), or the Granger-causality based temporal graph (<xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>) with the inclusion of inertial metrics measured at the dorsal side of the hand to capture the tactile component, along with the fine-grained motor component.</p>
</sec>
</sec>
<sec id="s5-6">
<title>5.6 Visual tasks</title>
<p>Eye movement is closely associated with humans&#x2019; goals, tasks, and intentions, as almost all tasks performed by humans involve visual observation. This association makes oculography a rich source of information for task recognition. Fixation, saccades, blink rate, and scanpath are the most commonly used metrics for detecting visual tasks (<xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>), followed by EOG potentials (<xref ref-type="bibr" rid="B78">Ishimaru et al., 2014b</xref>; <xref ref-type="bibr" rid="B76">Ishimaru et al., 2017</xref>; <xref ref-type="bibr" rid="B129">Lu et al., 2018</xref>). Visual tasks typically occur in <italic>office or desktop-based</italic> environments, where the participants are sedentary. The classifications of the reviewed visual task recognition algorithms are presented by algorithm category in <xref ref-type="table" rid="T5">Table 5</xref>.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Visual task recognition algorithms&#x2019; evaluation overview.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Classical Machine Learning</td>
</tr>
<tr>
<td align="left">Auto-context model</td>
<td align="left">
<xref ref-type="bibr" rid="B134">Martinez et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">Decision trees</td>
<td align="left">
<xref ref-type="bibr" rid="B107">Landsmann et al. (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B77">Ishimaru et al. (2014a)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">k- Nearest Neighbors</td>
<td align="left">
<xref ref-type="bibr" rid="B78">Ishimaru et al. (2014b)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Random Forest</td>
<td align="left">
<xref ref-type="bibr" rid="B191">Srivastava et al. (2018)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="3" align="left">SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B90">Kelton et al. (2019)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B21">Bulling et al. (2010)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B129">Lu et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B79">Islam et al. (2021)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">CNN &#x2b; LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B76">Ishimaru et al. (2017)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Graph CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B106">Lan et al. (2020)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Encoder-Decoder</td>
<td align="left">
<xref ref-type="bibr" rid="B140">Meyer et al. (2022)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-6-1">
<title>5.6.1 Classical machine learning</title>
<p>Classical machine learning using eye gaze metrics (e.g., saccades, fixation, and blink rate) for visual task recognition was pioneered by <xref ref-type="bibr" rid="B21">Bulling et al. (2010)</xref>. Statistical features (e.g., mean, max, variance) extracted from the gaze metrics, as well as the character-based representation to encode eye movement patterns, were used to train a SVM classifier to detect five office-based tasks. The algorithm&#x2019;s primary limitation is that the classification is provided at each time instance <italic>t</italic> independently and does not integrate long-range contextual information continuously (<xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>). A temporal contextual learning algorithm, the <italic>Auto-context</italic> model, overcame this limitation by including the past and future decision values from the discriminative classifiers (e.g., SVM and k-Nearest Neighbors) recursively until convergence (<xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>).</p>
<p>Low-level eye movement metrics (e.g., saccades and fixations) are versatile and easy to compute, but are vulnerable to overfitting, whereas high-level metrics (e.g., Area-of-Focus) may offer better abstraction, but requires domain and environment knowledge (<xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>). These limitations can be mitigated by exploiting low-level metrics to yield <italic>mid-level</italic> metrics that provide additional context. The mid-level metrics were built on intuitions about expected task relevant eye movements. Two different mid-level metrics were identified: <italic>shape-based pattern</italic> and <italic>distance-based pattern</italic> (<xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>). The shape-based pattern metrics were based on encoding different combinations of saccade and scanpath, while distance-based pattern metrics were generated using consecutive fixations. The low- and mid-level metrics were combined to train a Random Forest classifier to detect eight office-based tasks, including five desktop-based tasks and three software engineering tasks.</p>
<p>Various algorithms were developed focused solely on detecting reading tasks using classical machine learning (e.g., <xref ref-type="bibr" rid="B90">Kelton et al., 2019</xref>; <xref ref-type="bibr" rid="B107">Landsmann et al., 2019</xref>). The complexity of the reading task varied across algorithms. Reading detection can be as rudimentary as classifying active reading or not (<xref ref-type="bibr" rid="B107">Landsmann et al., 2019</xref>), or as complex as distinguishing between reading thoroughly vs skimming text (<xref ref-type="bibr" rid="B90">Kelton et al., 2019</xref>).</p>
<p>Based on feature mining, existing reading detection algorithms can be categorized into two methods: i) Global methods that mine eye movement metrics over an extended period <inline-formula id="inf19">
<mml:math id="m19">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>30</mml:mn>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> to build a reading detector (e.g., <xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B76">Ishimaru et al., 2017</xref>; <xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>), and ii) Local methods that extract the metrics within a narrow temporal window <inline-formula id="inf20">
<mml:math id="m20">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>3</mml:mn>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> (e.g., <xref ref-type="bibr" rid="B94">Kollmorgen and Holmqvist, 2007</xref>; <xref ref-type="bibr" rid="B17">Biedert et al., 2012</xref>). Global methods result in better accuracy, but do not detect reading in real-time, due to longer window sizes, while local methods allow for (near) real-time reading detection, but have low accuracy (<xref ref-type="bibr" rid="B90">Kelton et al., 2019</xref>).</p>
<p>Classical machine learning algorithms (e.g., <xref ref-type="bibr" rid="B21">Bulling et al., 2010</xref>; <xref ref-type="bibr" rid="B78">Ishimaru et al., 2014b</xref>; <xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>; <xref ref-type="bibr" rid="B191">Srivastava et al., 2018</xref>; <xref ref-type="bibr" rid="B107">Landsmann et al., 2019</xref>) have medium to high sensitivity. Most algorithms conform with the <italic>suitability</italic> criterion, while rarely conforming with the <italic>generalizability</italic> criterion. All algorithms are non-conforming for the <italic>concurrency</italic> and <italic>composite</italic> factors, making them unsuitable for the intended HRT domain.</p>
</sec>
<sec id="s5-6-2">
<title>5.6.2 Deep learning</title>
<p>Recent deep learning algorithms leverage CNNs to detect visual tasks directly using raw 2D gaze data obtained via wearable eye trackers. <italic>GazeGraph</italic> (<xref ref-type="bibr" rid="B106">Lan et al., 2020</xref>) algorithm converted 2D eye gaze sequence into a spatial-temporal graph representation that preserved important eye movement details, but rejected large irrelevant variations. A three-layered CNN trained on this representation detected various desktop and document reading tasks. An encoder-decoder based convolutional network detected seven mixed physical and visual tasks by combining 2D gaze data with head inertial metrics (<xref ref-type="bibr" rid="B140">Meyer et al., 2022</xref>).</p>
<p>Several other algorithms apply deep learning techniques using EOG potentials to detect reading task (<xref ref-type="bibr" rid="B76">Ishimaru et al., 2017</xref>; <xref ref-type="bibr" rid="B79">Islam et al., 2021</xref>). Two deep networks, a CNN and a LSTM, were developed to recognize reading in a natural setting (i.e., outside of the laboratory). Three metrics (i.e., blink rate, 2-channel EOG signals, and acceleration) from wearable EOG glasses were used to train the deep learning models.</p>
<p>Obtaining datasets at a large scale is difficult, due to high annotation costs and human effort, while lack of labeled data inhibits deep learning methods&#x2019; effectiveness. A sample efficient, <italic>self-supervised CNN</italic> detected reading task (<xref ref-type="bibr" rid="B79">Islam et al., 2021</xref>) using less labeled data. The self-supervised CNN employed a &#x201c;pretext&#x201d; task to bootstrap the network before training it for the actual target task. Three reading tasks (i.e., reading English documents, reading Japanese documents, both horizontally and vertically), as well as a no reading class, were detected by the self-supervised network. The pretext task recognized the transformation (i.e., rotational, translational, noise addition) applied to the input signal. The pretext pre-training phase initialized the network with good weights, which were fine-tuned by training the network on the target task (i.e., reading detection task) dataset.</p>
<p>The deep learning algorithms (e.g., <xref ref-type="bibr" rid="B76">Ishimaru et al., 2017</xref>; <xref ref-type="bibr" rid="B106">Lan et al., 2020</xref>; <xref ref-type="bibr" rid="B79">Islam et al., 2021</xref>) typically tend to have low sensitivity. Further, the EOG deep learning algorithms do not conform with <italic>suitability</italic>, as the employed metrics are unreliable, as it is susceptible to noise introduced by facial muscle movements (<xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>). These limitations discourage the use of deep learning for visual task recognition.</p>
</sec>
<sec id="s5-6-3">
<title>5.6.3 Discussion</title>
<p>None of the existing algorithms detected visual tasks within the targeted HRT context. The two classical machine learning algorithms: i) <italic>Auto-context</italic> model (<xref ref-type="bibr" rid="B134">Martinez et al., 2017</xref>) and ii) <xref ref-type="bibr" rid="B191">Srivastava et al.&#x2019;s (2018)</xref> algorithm appear to be more appropriate for detecting visual tasks. Both algorithms had <inline-formula id="inf21">
<mml:math id="m21">
<mml:mo>&#x3e;</mml:mo>
<mml:mn>70</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracies across a range of visual tasks and employed eye gaze metrics; thus, conforming with <italic>suitability</italic> and partially with <italic>sensitivity</italic>. The algorithms&#x2019; <italic>generalizability</italic> criterion requires additional evidence, as the former&#x2019;s validation scheme is unclear, while the latter does not have sufficient accuracy. None of the reviewed algorithms conform with the <italic>concurrency</italic> and <italic>composite factor</italic> criteria. The encoder-decoder algorithm (<xref ref-type="bibr" rid="B140">Meyer et al., 2022</xref>) is also a viable alternative, as it achieved <inline-formula id="inf22">
<mml:math id="m22">
<mml:mo>&#x3e;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> with leave-one-subject-out cross validation using eye gaze metrics.</p>
</sec>
</sec>
<sec id="s5-7">
<title>5.7 Cognitive tasks</title>
<p>Cognition describes mental processes, including reasoning, awareness, perception, knowledge, intuition, and judgment (<xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>), as such, most tasks require some cognitive capability. For instance, although tasks, such as reading, writing, watching videos predominantly involve visual, fine-grained motor, or tactile components, they also entail a cognitive component. Therefore, it is impractical to disregard the cognitive task elements, but classifying all such tasks as cognitive is also infeasible. Thus, only those algorithms that explicitly mention identifying the tasks&#x2019; cognitive aspect are reviewed. The evaluation of the reviewed cognitive task recognition algorithms is presented in <xref ref-type="table" rid="T6">Table 6</xref>.</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Cognitive task recognition algorithms&#x2019; evaluation overview.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Classical Machine Learning</td>
</tr>
<tr>
<td align="left">Decision trees</td>
<td align="left">
<xref ref-type="bibr" rid="B77">Ishimaru et al. (2014a)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">Ensemble</td>
<td align="left">
<xref ref-type="bibr" rid="B60">Grana et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">k-Nearest Neighbors</td>
<td align="left">
<xref ref-type="bibr" rid="B99">Kunze et al. (2013a)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td rowspan="2" align="left">SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B105">Lagodzinski et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B40">Datta et al. (2014)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td rowspan="2" align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B220">Zhang et al. (2019b)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B182">Sarkar et al. (2016)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">CNN &#x2b; LSTM</td>
<td align="left">
<xref ref-type="bibr" rid="B180">Salehzadeh et al. (2020)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-7-1">
<title>5.7.1 Classical machine learning</title>
<p>EEG potentials, obtained by placing non-invasive electrodes on humans&#x2019; scalp, are the primary electrophysiological metrics used to detect cognitive tasks. Features (e.g., amplitude and power spectral density) extracted from the EEG frequency bands (i.e., alpha (8&#x2013;12 Hz), beta (13&#x2013;30 Hz), theta (4&#x2013;8 Hz), and delta (<inline-formula id="inf23">
<mml:math id="m23">
<mml:mo>&#x3c;</mml:mo>
</mml:math>
</inline-formula> 4 Hz)) can be used to train classical machine learning algorithms (<xref ref-type="bibr" rid="B144">Mostow et al., 2011</xref>; <xref ref-type="bibr" rid="B100">Kunze et al., 2013b</xref>; <xref ref-type="bibr" rid="B60">Grana et al., 2020</xref>). Reading is the most widely detected cognitive task using EEG sensing. A k-Nearest Neighbors classifier distinguished between reading and non-reading tasks (e.g., drawing, watching a video and listening to music), as well as distinguishing reading different kinds of document, using a wearable EEG sensor (<xref ref-type="bibr" rid="B101">Kunze et al., 2013c</xref>).</p>
<p>EOG is the other electrophysiological metric employed for detecting cognitive tasks (<xref ref-type="bibr" rid="B40">Datta et al., 2014</xref>; <xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>). The efficacy of different EOG features in detecting cognitive tasks was investigated (<xref ref-type="bibr" rid="B40">Datta et al., 2014</xref>). Three features (i.e., adaptive autoregressive parameters, wavelet coefficients and Hjorth parameters) were extracted from a laboratory-developed two-channel EOG signal acquisition device. These features were used independently and in combinations to train a SVM classifier to detect eight cognitive tasks (e.g., reading, writing, copying a text, web browsing, watching a video, playing an online game, and word search).</p>
<p>Several other algorithms combine data from multiple sensing modalities to improve cognitive task recognition accuracy (<xref ref-type="bibr" rid="B77">Ishimaru et al., 2014a</xref>; <xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>; <xref ref-type="bibr" rid="B60">Grana et al., 2020</xref>). For example, combining blink rate and head motion by fusing the eye gaze data with acceleration data improved a decision tree classifier&#x2019;s accuracy in detecting four cognitive tasks (e.g., reading, solving a math problem, watching a video and talking) (<xref ref-type="bibr" rid="B77">Ishimaru et al., 2014a</xref>). The <italic>Codebook</italic> algorithm recognized six cognitive tasks (e.g., reading a printed page, watching a video, engaging conversation, writing handwritten notes and sorting numbers) by clustering the subsequences sampled from a data sequence based on similarity (<xref ref-type="bibr" rid="B105">Lagodzinski et al., 2018</xref>). The resulting cluster centers act as the set of codewords (i.e., codebook). A SVM classifier was trained to classify the histogram reflecting the codeword frequency to predict the tasks.</p>
<p>Classical machine learning algorithms typically have high <italic>sensitivity</italic>; however, the incorporated metrics (i.e., EOG and EEG) are unreliable and cannot accommodate individual differences. Therefore, the algorithms&#x2019; <italic>suitability</italic> and <italic>generalizability</italic> are non-conforming. The <italic>composite factor</italic> and <italic>concurrency</italic> are non-conforming as well.</p>
</sec>
<sec id="s5-7-2">
<title>5.7.2 Deep learning</title>
<p>Recent deep learning advances facilitate detecting cognitive tasks using EEG potentials acquired from off-the-shelf, wireless, wearable EEG devices. Most EEG wearable devices (e.g., <xref ref-type="bibr" rid="B46">EMOTIV, 2021</xref>; <xref ref-type="bibr" rid="B145">MUSE, 2021</xref>; <xref ref-type="bibr" rid="B148">Neurosky, 2021</xref>) record prefrontal EEG signals, which are correlated to a human&#x2019;s intellectual, emotional and cognitive states (<xref ref-type="bibr" rid="B180">Salehzadeh et al., 2020</xref>). A deep EEG network detected three cognitive tasks (e.g., reading, speaking, and watching a video) using data collected from a wearable EEG sensor&#x2019;s (<xref ref-type="bibr" rid="B145">MUSE, 2021</xref>) two prefrontal EEG channels (<xref ref-type="bibr" rid="B180">Salehzadeh et al., 2020</xref>). The hybrid deep learning algorithm incorporated a CNN to populate the feature maps from raw EEG potentials, followed by a LSTM network for modeling the temporal state of the EEG feature maps. Most existing EEG-based algorithms focus on application-specific classification algorithms, which may not translate to other domains. A transferable EEG-based cognitive task recognition algorithm that can adaptively support varying EEG channels as input and operate on a wide range of cognitive applications was developed (<xref ref-type="bibr" rid="B220">Zhang X. et al., 2019</xref>). The algorithm combined deep reinforcement learning with an attention mechanism to extract robust and distinct deep features.</p>
<p>Detecting cognitive tasks with fewer EEG sensors in an unconstrained, natural environment is a challenging task, due to low signal-to-noise ratio, lack of baseline availability, change of baseline due to domain environment and individual differences, as well as uncontrolled mixing of various tasks (<xref ref-type="bibr" rid="B182">Sarkar et al., 2016</xref>). A deep learning algorithm (<xref ref-type="bibr" rid="B182">Sarkar et al., 2016</xref>) revealed that the backward sensor selection (<xref ref-type="bibr" rid="B18">Bishop and Nasrabadi, 2006</xref>) technique can reduce the sensor suite significantly (i.e., from nine probes to three) without compromising accuracy. Two deep neural networks, a deep belief network and a CNN, were trained using the EEG power spectral density to distinguish between listening and watching tasks.</p>
<p>Similar to the classical machine learning approaches, the deep learning algorithms also tend to have high <italic>sensitivity</italic>, but incorporate EEG metrics that suffer from low-signal-to-noise and individual differences; therefore, the algorithms&#x2019; <italic>suitability</italic> and <italic>generalizability</italic> are non-conforming. Additionally, none of reviewed algorithms conform with the <italic>concurrency</italic> and <italic>composite factor</italic>.</p>
</sec>
<sec id="s5-7-3">
<title>5.7.3 Discussion</title>
<p>Generally, cognitive tasks can be classified with <inline-formula id="inf24">
<mml:math id="m24">
<mml:mo>&#x3e;</mml:mo>
<mml:mn>80</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy; therefore, the algorithms&#x2019; typically have high <italic>sensitivity</italic>. Excluding <xref ref-type="bibr" rid="B77">Ishimaru et al.&#x2019;s (2014a)</xref> decision tree classifier, none of the other algorithms conform with suitability, as the metrics employed were EEG or EOG. Thus, none of the discussed algorithms are appropriate for detecting cognitive tasks for the intended HRT domain. Given (a) that cognitive and visual tasks are closely associated, and (b) the efficacy of multimodality sensing (<xref ref-type="bibr" rid="B77">Ishimaru et al., 2014a</xref>), it is hypothesized that a classical machine learning algorithm that incorporates metrics, such as pupil dilation, blink latency and blink rate, as well as heart-rate variability will be viable for detecting cognitive tasks.</p>
</sec>
</sec>
<sec id="s5-8">
<title>5.8 Auditory tasks</title>
<p>Auditory event recognition involves identifying characteristic ambient sounds in order to detect tasks in an environment (<xref ref-type="bibr" rid="B139">Mesaros et al., 2017</xref>; <xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>). Auditory event recognition algorithms typically use microphone sensor(s) to detect the sound events. Individual classifications for each auditory event detection algorithm by its category are provided in <xref ref-type="table" rid="T7">Table 7</xref>.</p>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>Auditory task recognition algorithms&#x2019; evaluation overview.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Classical Machine Learning</td>
</tr>
<tr>
<td align="left">Random Forest</td>
<td align="left">
<xref ref-type="bibr" rid="B193">Stork et al. (2012)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">
<xref ref-type="bibr" rid="B213">Yatani and Truong (2012)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td rowspan="5" align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B108">Laput et al. (2018)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">C</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B70">Hershey et al. (2017)</xref>
</td>
<td align="center">&#x22c0;</td>
<td align="center">RE</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B122">Liang and Thomaz (2019)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B64">Haubrick and Ye (2019)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">C</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">C</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B179">Salamon and Bello (2017)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">RE</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-8-1">
<title>5.8.1 Classical machine learning</title>
<p>Auditory event detection algorithms that incorporate MFCCs typically use classical machine learning (e.g., <xref ref-type="bibr" rid="B193">Stork et al. (2012)</xref> (<xref ref-type="bibr" rid="B213">Yatani and Truong, 2012</xref>)), while deep learning algorithms incorporate spectrograms (e.g., <xref ref-type="bibr" rid="B70">Hershey et al., 2017</xref>; <xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>; <xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>; <xref ref-type="bibr" rid="B122">Liang and Thomaz, 2019</xref>). A Random Forest based voting algorithm, <italic>Non-Markovian Ensemble Voting</italic>, used MFCCs to recognize characteristic sounds produced by twenty-two ADL tasks (<xref ref-type="bibr" rid="B193">Stork et al., 2012</xref>). The predictions were refined over time by collecting consensus via voting from the past and future predictions.</p>
<p>A wearable acoustic sensor, <italic>BodyScope</italic>, worn around the neck classified several ADL tasks (<xref ref-type="bibr" rid="B213">Yatani and Truong, 2012</xref>). The <italic>BodyScope</italic> sensor contained a microphone surrounded by a stethoscope chest piece for sound amplifications in order to exploit the sounds that occurred at a human&#x2019;s mouth and throat regions to recognize the tasks. For instance, when a person speaks to someone, they generate vocal sounds, while eating and drinking produce chewing, sipping, and swallowing sounds. Several time- and frequency-domain features (e.g., zero-crossing rate and MFCCs) were used to train an SVM classifier.</p>
<p>Both algorithms achieved <inline-formula id="inf25">
<mml:math id="m25">
<mml:mo>&#x3e;</mml:mo>
<mml:mn>70</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> accuracy during an in-the-wild study; therefore, the algorithms&#x2019; <italic>sensitivity</italic> is medium to high. Although the metrics used were reliable, the ensemble voting algorithm incorporated an environmentally embedded microphone, while the <italic>BodyScope</italic> sensor suffers from non-reproducibility; therefore, the algorithms&#x2019; <italic>suitability</italic> is classified as non-conforming. Finally, the algorithms&#x2019; <italic>composite factor</italic> and <italic>concurrency</italic> are classified as non-conforming.</p>
</sec>
<sec id="s5-8-2">
<title>5.8.2 Deep learning</title>
<p>Deep learning algorithms can leverage the time, frequency and amplitude information in an audio signal&#x2019;s spectrogram to extract the spatio-temporal features. Most auditory event detection deep learning algorithms (e.g., <xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>; <xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>; <xref ref-type="bibr" rid="B122">Liang and Thomaz, 2019</xref>) leverage transfer learning. These algorithms fine-tune the existing <italic>VGGish model</italic> (<xref ref-type="bibr" rid="B70">Hershey et al., 2017</xref>) [i.e., pre-trained on the <italic>YouTube Audio Set</italic> (<xref ref-type="bibr" rid="B56">Gemmeke et al., 2017</xref>)] with additional layers to detect the target auditory events. The <italic>VGGish model</italic> used a log Mel spectrogram as input to output a 128-dimensional feature embedding for every second of an audio sample. A transfer learning framework detected fifteen ADL tasks (e.g., talking, watching television, brushing, shaving, and listening to music) from the audio recorded using an off-the-shelf smartphone (<xref ref-type="bibr" rid="B122">Liang and Thomaz, 2019</xref>). A five-layer CNN was added to the <italic>VGGish</italic>&#x2019;s feature embedding to predict the ADL tasks.</p>
<p>Polyphonic event detection algorithms recognize multiple auditory events occurring simultaneously (<xref ref-type="bibr" rid="B28">Chan and Chin, 2020</xref>), typically via deep neural networks. A <italic>VGGish</italic>-based algorithm stacked multiple binary classifiers to the feature vector (<xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>), while others have developed various recurrent and hybrid deep neural networks to detect polyphonic sound events (<xref ref-type="bibr" rid="B22">Cakir et al., 2015</xref>; <xref ref-type="bibr" rid="B154">Parascandolo et al., 2016</xref>; <xref ref-type="bibr" rid="B23">Cak&#x131;r et al., 2017</xref>).</p>
<p>Audio augmentation can be exploited by deep learning to improve the recognition rate. Augmenting the original audio with a set of deformations (e.g., time stretching, pitch shifting, dynamic range compression, and background noise mixing) improved a CNN&#x2019;s classification accuracy significantly on a range of environmental sound classification tasks (<xref ref-type="bibr" rid="B179">Salamon and Bello, 2017</xref>). <italic>Ubicoustics</italic>, a real-time, auditory event recognition algorithm was trained by incorporating various augmentation techniques to simulate the sounds that resemble real-world audio samples (<xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>).</p>
<p>Deep learning algorithms tend to outperform classical machine learning approaches in terms of accuracy. Thus, deep learning algorithms, especially the <italic>VGGish</italic>-based models (<xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>; <xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>; <xref ref-type="bibr" rid="B122">Liang and Thomaz, 2019</xref>), are the most suited for detecting auditory events, primarily due to their feature extraction capability and the availability of abundant audio datasets (<xref ref-type="bibr" rid="B56">Gemmeke et al., 2017</xref>; <xref ref-type="bibr" rid="B139">Mesaros et al., 2017</xref>). Most algorithms are validated by splitting all the available data randomly into training and validation datasets; therefore, the algorithms&#x2019; <italic>generalizability</italic> criteria either require additional evidence, or are non-conforming. All evaluated algorithms are non-conforming for the <italic>composite factor</italic> and <italic>anomaly awareness</italic> criteria.</p>
</sec>
<sec id="s5-8-3">
<title>5.8.3 Discussion</title>
<p>Most auditory event detection algorithms typically have medium to high <italic>sensitivity</italic> (e.g., <xref ref-type="bibr" rid="B193">Stork et al., 2012</xref>; <xref ref-type="bibr" rid="B213">Yatani and Truong, 2012</xref>; <xref ref-type="bibr" rid="B70">Hershey et al., 2017</xref>; <xref ref-type="bibr" rid="B108">Laput et al., 2018</xref>; <xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>). The algorithms&#x2019; <italic>suitability</italic> criterion depends on whether the microphone is worn or embedded in the environment. The algorithms&#x2019; <italic>generalizability</italic> criteria either require additional evidence or non-conforming. The polyphonic detection algorithms conform with concurrency; therefore, a deep polyphonic detection algorithm (e.g., <xref ref-type="bibr" rid="B64">Haubrick and Ye, 2019</xref>) is recommended for the intended HRT domain, as it is more likely to contain multiple, simultaneous sound sources (<xref ref-type="bibr" rid="B28">Chan and Chin, 2020</xref>). None of the reviewed algorithms conform with <italic>composite factor</italic> and <italic>anomaly awareness</italic>.</p>
</sec>
</sec>
<sec id="s5-9">
<title>5.9 Speech tasks</title>
<p>Verbal communication plays a key role in task performance, such as assigning tasks, sharing or confirming important information, and reporting task completion, especially in a dynamic environment (<xref ref-type="bibr" rid="B224">Zhang and Sarcevic, 2015</xref>). Speech-reliant task recognition in a highly dynamic environment [e.g., trauma resuscitation (<xref ref-type="bibr" rid="B15">Bergs et al., 2005</xref>)] encounters several challenges, such as inconsistent verbal reports between tasks, the potentially succinct and non-grammatical nature of verbal communication, overlapping multi-person speech, and interleaved verbal exchanges due to multi-tasking (<xref ref-type="bibr" rid="B81">Jagannath et al., 2019</xref>). Only two speech task recognition algorithms were identified, their classifications are cited in <xref ref-type="table" rid="T8">Table 8</xref>.</p>
<table-wrap id="T8" position="float">
<label>TABLE 8</label>
<caption>
<p>Speech task recognition algorithms&#x2019; evaluation overview.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Category</th>
<th rowspan="2" align="center">Paper</th>
<th rowspan="2" align="center">Sens</th>
<th rowspan="2" align="center">Suit</th>
<th rowspan="2" align="center">Genr</th>
<th rowspan="2" align="center">Comp</th>
<th rowspan="2" align="center">Conc</th>
<th rowspan="2" align="center">Anom</th>
</tr>
<tr>
<th align="center">Algorithm</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td colspan="8" align="left">Deep Learning</td>
</tr>
<tr>
<td align="left">Attention</td>
<td align="left">
<xref ref-type="bibr" rid="B63">Gu et al. (2019)</xref>
</td>
<td align="center">
<italic>&#x220f;</italic>
</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
<tr>
<td align="left">CNN</td>
<td align="left">
<xref ref-type="bibr" rid="B2">Abdulbaqi et al. (2020)</xref>
</td>
<td align="center">&#x22c1;</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
<td align="center">NC</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5-9-1">
<title>5.9.1 Deep learning</title>
<p>Little research exists on speech-reliant task recognition (<xref ref-type="bibr" rid="B82">Jagannath et al., 2018</xref>), with the existing algorithms being based on deep learning (e.g., <xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>; <xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>). A text-based task recognition algorithm employed a verbal transcript derived from a medical team&#x2019;s communication as input to predict ten trauma resuscitation tasks (<xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>). The algorithm used both speech and ambient sounds for task prediction. A multimodal attention network was applied to process the transcribed spoken language and the ambient sounds in order to predict the tasks. The main limitation was reliance on manually-generated transcripts, which is infeasible for a contemporaneous task recognition system (<xref ref-type="bibr" rid="B3">Abdulbaqi, 2020</xref>). Automatic transcript generation requires a computationally expensive speech recognition tool. Additionally, poor audio quality caused by distant talking, ambient noise, succinct and non-grammatical speaking can increase an automatic speech recognition tool&#x2019;s error rate; thus, the algorithm&#x2019;s performance in real-world scenarios is expected to be lower than the cited result (<xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>).</p>
<p>An alternative speech-reliant task recognition algorithm depended only on one keyword to detect trauma tasks (<xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>). The speech-reliant algorithm used one representative keyword per utterance as input to a deep neural network, in addition to the ambient sounds. This keyword was determined by calculating the most frequent words list for each task based on the premise that frequently occurring words for particular tasks can serve as features for the neural network&#x2019;s task prediction. Word-spotting tools (e.g., <xref ref-type="bibr" rid="B199">Tsai and Hao, 2019</xref>; <xref ref-type="bibr" rid="B54">Gao et al., 2020</xref>) can extract keywords efficiently, reducing the reliance on traditional speech recognition. The deep learning architecture consisted of an audio network, a keyword network, and a fusion network. The audio network adapted a modified <italic>VGGish</italic> deep network (<xref ref-type="bibr" rid="B190">Simonyan and Zisserman, 2014</xref>) to extract features from the audio spectrogram, while the keyword network extracted important verbal features from the keyword list. The fusion network concatenated the output of both networks to predict the speech-reliant tasks.</p>
<p>The use of speech to recognize tasks is an under-developed area. The reviewed algorithms had low to medium <italic>sensitivity</italic> (e.g., <xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>; <xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>). The algorithms&#x2019; <italic>suitability</italic> criterion was classified as non-conforming, because the incorporated metrics are domain and task specific. Both algorithms&#x2019; are non-conforming with <italic>generalizability</italic>, <italic>composite factor</italic>, <italic>concurrency</italic> and <italic>anomaly awareness</italic>.</p>
</sec>
<sec id="s5-9-2">
<title>5.9.2 Discussion</title>
<p>Verbal communication in a highly dynamic setting [e.g., trauma medical team (<xref ref-type="bibr" rid="B224">Zhang and Sarcevic, 2015</xref>)] occurs at a high level (e.g., discussing task plans and intentions) and at a low level (e.g., coordinating, executing and reporting task completion) (<xref ref-type="bibr" rid="B82">Jagannath et al., 2018</xref>). Identifying speech patterns and keywords during verbal communication can detect the tasks, as well as track their progress (i.e., preparing-performing-reporting). Incorporating verbal exchanges [e.g., transcripts (<xref ref-type="bibr" rid="B63">Gu et al., 2019</xref>) or keywords (<xref ref-type="bibr" rid="B2">Abdulbaqi et al., 2020</xref>)] in tandem with the audio stream increased the accuracy by at least 15% in both algorithms, indicating that speech patterns and keywords can serve as differentiators for speech-reliant tasks. The algorithms&#x2019; main limitations are that the metrics are highly domain-specific and require natural language processing, along with substantial manual effort to identify task specific sensitive keywords (see <xref ref-type="sec" rid="s4">Section 4</xref>). Therefore, the algorithms cannot be readily transferred across domains. A suitable alternative may involve modifying the algorithm to use the audio stream in conjunction with speech workload metrics (e.g., speech rate, voice intensity, and voice pitch) instead of verbal exchanges.</p>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s6">
<title>6 Discussion</title>
<p>HRTs often perform a wide range of tasks, such that the set of all tasks performed by human teammates may involve combinations of all activity components. Thus, an overall task recognition model capable of detecting tasks with multiple different activity components is hypothesized to improve the task recognition in more complex domains. None of the reviewed algorithms meet all the criteria necessary to achieve this goal, due to identifying tasks with a limited set of activity components, and the algorithms&#x2019; limitations with regard to the evaluation criteria: sensitivity, suitability, generalizability, composite factor, concurrency, and anomaly awareness.</p>
<p>Two algorithms, <xref ref-type="bibr" rid="B60">Grana et al.&#x2019;s (2020)</xref> and <xref ref-type="bibr" rid="B77">Ishimaru et al.&#x2019;s (2014a)</xref>, come the closest to addressing the issues, as they identify four activity components. The former identified a gross motor task, while the latter identified a fine-grained motor task in addition to detecting visual, cognitive, and auditory tasks. However, both algorithms failed to satisfy all the required evaluation criteria. Most other algorithms detected tasks involving at most two activity components: gross and fine-grained motor, fine-grained motor and tactile, or visual and cognitive tasks. Finally, none of the reviewed algorithms detected tasks across all the identified activity components.</p>
<p>Several algorithms conform with sensitivity, suitability, and generalizability. Other than the cognitive and speech task recognition algorithms, there exists at least one algorithm that can detect tasks reliably while conforming with suitability and generalizability for each individual activity component. Existing algorithms are highly limited in satisfying the other three criteria: composite, concurrent, and anomaly awareness.</p>
<p>HRT tasks are typically composite in nature in that they are composed of differing combinations of multiple activity components; therefore, understanding the tasks&#x2019; various components is required in order to detect the tasks accurately, as well as track their progress. Thirteen algorithms, all of which were either gross or fine-grained motor task recognition algorithms, attempted to detect composite tasks. Eight algorithms detected composite tasks with high accuracy, while only three (<xref ref-type="bibr" rid="B143">Min and Cho, 2011</xref>; <xref ref-type="bibr" rid="B127">Liu et al., 2017</xref>; <xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>) managed to detect tasks with high accuracy using reliable metrics from wearable sensors. A similar trend was observed for concurrency detection. Eight algorithms attempted to detect concurrent tasks, seven of which belonged to gross or fine-grained motor categories. Five of those algorithms detected tasks with high sensitivity, while only two (<xref ref-type="bibr" rid="B127">Liu et al., 2017</xref>; <xref ref-type="bibr" rid="B123">Liao et al., 2020</xref>) satisfied the sensitivity and suitability criteria.</p>
<p>Only two algorithms (<xref ref-type="bibr" rid="B141">Min et al., 2008</xref>; <xref ref-type="bibr" rid="B109">Laput and Harrison, 2019</xref>) conform with anomaly awareness. The remaining algorithms&#x2019; primary limitation for anomaly awareness was the assumption that humans will not perform tasks outside of a predefined set, which is not the case for the intended domain. Therefore, there remains a need for a task recognition algorithm that can detect out-of-class instances reliably. Such an algorithm can draw from existing novelty and anomaly detection research (e.g., <xref ref-type="bibr" rid="B135">Masana et al., 2018</xref>; <xref ref-type="bibr" rid="B27">Chalapathy and Chawla, 2019</xref>; <xref ref-type="bibr" rid="B103">Kwon et al., 2019</xref>; <xref ref-type="bibr" rid="B162">Perera and Patel, 2019</xref>; <xref ref-type="bibr" rid="B153">Pang et al., 2021</xref>).</p>
</sec>
<sec id="s7">
<title>7 Challenges and future directions</title>
<p>Despite the rapid growth in task recognition research, there are still a number of largely unexplored challenges. Some research challenges and future directions that can help realize the full potential of task recognition for uncertain, dynamic environments, especially within the context of first-response HRTs are presented.<list list-type="simple">
<list-item>
<p>&#x2022; <bold>Composite Task Recognition:</bold> Existing task recognition methods can detect short-duration or repetitive atomic tasks (e.g., walking, running, reading) with high accuracy; however, detecting long-duration composite tasks (e.g., triaging a victim, digging a wildland fire trench) that involve multiple activity components remain difficult. A task recognition algorithm that incorporates several sensors as input to recognize the composite task&#x2019;s multiple activity components is desired. For example, the algorithm can incorporate inertial and sEMG metrics to detect gross motor, fine-grained motor, and tactile tasks, as well as integrate pupil dilation, fixation, and saccades from a wearable eye tracker to detect visual tasks.</p>
</list-item>
<list-item>
<p>&#x2022; <bold>Leverage interactions between activity components:</bold> Most composite tasks can be decomposed into atomic tasks across activity components. These atomic tasks typically follow a temporal ordering within the components. For example, <italic>taking a picture of a suspicious object</italic> requires <italic>visually</italic> scanning and <italic>cognitively</italic> evaluating the object prior to the <italic>fine-grained motor</italic> and <italic>tactile</italic> picture capturing task. A composite task recognition algorithm can leverage the interactions between the activity components in order to simultaneously (a) optimize each component&#x2019;s atomic task detections, and (b) construct the composite tasks from the low-level atomic task detections.</p>
</list-item>
<list-item>
<p>&#x2022; <bold>Adaptively segmenting metrics:</bold> Most existing task recognition algorithms segment the sensor data into temporal chunks using a fixed window size as input to the machine learning algorithms. A short-duration task may require a smaller window size, so that the task is not overshadowed (e.g., confused) by unrelated data, while a long-duration task may require a larger window size to have sufficient context. Therefore, it may be necessary for a task recognition algorithm to use an adaptive sliding window approach. This approach will permit expanding and contracting the window size based on the task, which may lead to more accurate detection. An ensemble learning algorithm may also be leveraged, where the algorithm makes predictions over multiple fixed window sizes and fuses the predictions across the window sizes intelligently to detect the tasks.</p>
</list-item>
<list-item>
<p>&#x2022; <bold>Lack of an ecologically-valid HRT dataset:</bold> Machine learning algorithms only perform as well as the quality of the training data. Training and evaluating algorithms, especially deep learning models require large datasets. Most task recognition studies depend on existing benchmark datasets (e.g., <italic>Opportunity</italic> (<xref ref-type="bibr" rid="B175">Roggen et al., 2010</xref>), <italic>WISDM</italic> (<xref ref-type="bibr" rid="B208">Weiss et al., 2019</xref>), <italic>PAMAP2</italic> (<xref ref-type="bibr" rid="B173">Reiss and Stricker, 2012</xref>) that only involve day-to-day physical tasks and seldom incorporate composite tasks with multiple activity components. Additionally, most of these datasets only entail inertial and heart-rate metrics gathered in a highly controlled environment, which limits the algorithms&#x2019; capability to detect tasks across activity components. Datasets involving components other than the gross and fine-grained motor components are not available publicly. A composite, concurrent HRT task recognition evaluation that incorporates tasks involving multiple activity components performed in an ecologically valid setting is desired. A multimodal dataset that incorporates metrics acquired from multiple wearable sensors (e.g., whole-body inertial measurements, sEMG, eye tracking, and microphone audio) is required.</p>
</list-item>
</list>
</p>
</sec>
<sec sec-type="conclusion" id="s8">
<title>8 Conclusion</title>
<p>HRTs collaborating to achieve tasks under various conditions, especially in unstructured, dynamic environments, will require robots to adapt autonomously to a human teammate&#x2019;s state. An important element of such adaptation is the robot&#x2019;s ability to infer the human teammate&#x2019;s current tasks, as understanding human actions and their interactions with the world provides the robot with more context as to what type of assistance the human may need. A multi-dimensional task recognition algorithm is needed to identify composite and atomic tasks performed by HRTs working in unstructured, dynamic environments in order to trigger such autonomous robot adaptations that can result in improved teaming outcomes.</p>
<p>This literature review classified task recognition algorithms based on six criteria: sensitivity, suitability, generalizability, composite factor, concurrency and anomaly awareness. The algorithms&#x2019; limitations include recognizing tasks with a limited set of activity component, not detecting composite and concurrent tasks across components, not being viable for HRT domain, and not accounting for individual differences.</p>
</sec>
</body>
<back>
<sec id="s9">
<title>Author contributions</title>
<p>PB read the surveyed papers, collected the required data to organize and structure the literature review, and wrote the manuscript. JAA read the manuscript, and provided continuous feedback and guidance to PB in order to organize and write the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s10">
<title>Funding</title>
<p>This research was partially supported by an Office of Naval Research award N0004-21-1-2190.</p>
</sec>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s13">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s12">
<title>Author disclaimer</title>
<p>The contents are those of the authors and do not represent the official views of, nor an endorsement, by the Office of Naval Research.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abdoli-Eramaki</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Damecour</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Christenson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Stevenson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>The effect of perspiration on the sEMG amplitude and power spectrum</article-title>. <source>J. Electromyogr. Kinesiol.</source> <volume>22</volume>, <fpage>908</fpage>&#x2013;<lpage>913</lpage>. <pub-id pub-id-type="doi">10.1016/j.jelekin.2012.04.009</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abdulbaqi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Marsic</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Burd</surname>
<given-names>R. S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Speech-based activity recognition for trauma resuscitation</article-title>. <source>IEEE Int. Conf. Healthc. Inf.</source> <volume>2020</volume>, <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/ichi48887.2020.9374372</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Abdulbaqi</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Speech-based activity recognition for medical teamwork</source>. <comment>Ph.D. thesis</comment>. <publisher-name>Rutgers University-School of Graduate Studies</publisher-name>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ahlstrom</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Friedman-Berg</surname>
<given-names>F. J.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Using eye movement activity as a correlate of cognitive workload</article-title>. <source>Int. J. Industrial Ergonomics</source> <volume>36</volume>, <fpage>623</fpage>&#x2013;<lpage>636</lpage>. <pub-id pub-id-type="doi">10.1016/j.ergon.2006.04.002</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akbari</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Jafari</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Personalizing activity recognition models through quantifying different types of uncertainty using wearable sensors</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>67</volume>, <fpage>2530</fpage>&#x2013;<lpage>2541</lpage>. <pub-id pub-id-type="doi">10.1109/tbme.2019.2963816</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Al-qaness</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Dahou</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Abd Elaziz</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Helmi</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Multi-ResAtt: Multilevel residual network with attention for human activity recognition using wearable sensors</article-title>. <source>IEEE Trans. Industrial Inf.</source> <volume>19</volume>, <fpage>144</fpage>&#x2013;<lpage>152</lpage>. <pub-id pub-id-type="doi">10.1109/tii.2022.3165875</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Allahbakhshi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Conrow</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Naimi</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Weibel</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Using accelerometer and GPS data for real-life physical activity type detection</article-title>. <source>Sensors</source> <volume>20</volume>, <fpage>588</fpage>. <pub-id pub-id-type="doi">10.3390/s20030588</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Allen</surname>
<given-names>J. F.</given-names>
</name>
<name>
<surname>Ferguson</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>Actions and events in interval temporal logic</article-title>. <source>J. Log. Comput.</source> <volume>4</volume>, <fpage>531</fpage>&#x2013;<lpage>579</lpage>. <pub-id pub-id-type="doi">10.1093/logcom/4.5.531</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Alsheikh</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Selim</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Niyato</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Doyle</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>H.-P.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deep activity recognition models with triaxial accelerometers</article-title>,&#x201d; in <source>Workshops at the AAAI conference on artificial intelligence applied to assistive technologies and smart environments. WS-16-01</source>.</citation>
</ref>
<ref id="B10">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Amma</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Krings</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>B&#xf6;er</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Schultz</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Advancing muscle-computer interfaces with high-density electromyography</article-title>,&#x201d; in <source>ACM conference on human factors in computing systems</source>, <fpage>929</fpage>&#x2013;<lpage>938</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arif</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bilal</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kattan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ahamed</surname>
<given-names>S. I.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Better physical activity classification using smartphone acceleration sensor</article-title>. <source>J. Med. Syst.</source> <volume>38</volume>, <fpage>95</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1007/s10916-014-0095-0</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Atallah</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>King</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>G.-Z.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Sensor placement for activity detection using wearable accelerometers</article-title>,&#x201d; in <source>IEEE international conference on body sensor networks</source>, <fpage>24</fpage>&#x2013;<lpage>29</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Atzori</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gijsberts</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Castellini</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Caputo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Hager</surname>
<given-names>A.-G. M.</given-names>
</name>
<name>
<surname>Elsig</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Electromyography data for non-invasive naturally-controlled robotic hand prostheses</article-title>. <source>Sci. Data</source> <volume>1</volume>, <fpage>140053</fpage>&#x2013;<lpage>140113</lpage>. <pub-id pub-id-type="doi">10.1038/sdata.2014.53</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Batzianoulis</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>El-Khoury</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Pirondini</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Coscia</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Micera</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Billard</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>EMG-based decoding of grasp gestures in reaching-to-grasping motions</article-title>. <source>Robotics Aut. Syst.</source> <volume>91</volume>, <fpage>59</fpage>&#x2013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1016/j.robot.2016.12.014</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bergs</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Rutten</surname>
<given-names>F. L.</given-names>
</name>
<name>
<surname>Tadros</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Krijnen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Schipper</surname>
<given-names>I. B.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Communication during trauma resuscitation: Do we know what is happening?</article-title> <source>Injury</source> <volume>36</volume>, <fpage>905</fpage>&#x2013;<lpage>911</lpage>. <pub-id pub-id-type="doi">10.1016/j.injury.2004.12.047</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bian</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lukowicz</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>The state-of-the-art sensing techniques in human activity recognition: A survey</article-title>. <source>Sensors</source> <volume>22</volume>, <fpage>4596</fpage>. <pub-id pub-id-type="doi">10.3390/s22124596</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Biedert</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Hees</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dengel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Buscher</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>A robust realtime reading-skimming classifier</article-title>,&#x201d; in <source>Eye tracking research and applications symposium</source>, <fpage>123</fpage>&#x2013;<lpage>130</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bishop</surname>
<given-names>C. M.</given-names>
</name>
<name>
<surname>Nasrabadi</surname>
<given-names>N. M.</given-names>
</name>
</person-group> (<year>2006</year>). <source>Pattern recognition and machine learning, vol. 4</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Braojos</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Beretta</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Constantin</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Burg</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Atienza</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>A wireless body sensor network for activity monitoring with low transmission overhead</article-title>,&#x201d; in <source>IEEE international conference on embedded and ubiquitous computing</source>, <fpage>265</fpage>&#x2013;<lpage>272</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brendel</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Fern</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Todorovic</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Probabilistic event logic for interval-based event recognition</article-title>,&#x201d; in <source>IEEE international conference on computer vision and pattern recognition</source>, <fpage>3329</fpage>&#x2013;<lpage>3336</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bulling</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ward</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Gellersen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Tr&#xf6;ster</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Eye movement analysis for activity recognition using electrooculography</article-title>. <source>IEEE Trans. Pattern Analysis Mach. Intell.</source> <volume>33</volume>, <fpage>741</fpage>&#x2013;<lpage>753</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2010.86</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Cakir</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Heittola</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Huttunen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Virtanen</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Polyphonic sound event detection using multi label deep neural networks</article-title>,&#x201d; in <source>IEEE international joint conference on neural networks</source>, <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cak&#x131;r</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Parascandolo</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Heittola</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Huttunen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Virtanen</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Convolutional recurrent neural networks for polyphonic sound event detection</article-title>. <source>IEEE/ACM Trans. Audio, Speech, Lang. Process.</source> <volume>25</volume>, <fpage>1291</fpage>&#x2013;<lpage>1303</lpage>. <pub-id pub-id-type="doi">10.1109/taslp.2017.2690575</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Capela</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Lemaire</surname>
<given-names>E. D.</given-names>
</name>
<name>
<surname>Baddour</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Feature selection for wearable smartphone-based human activity recognition with able bodied, elderly, and stroke patients</article-title>. <source>PLoS One</source> <volume>10</volume>, <fpage>e0124414</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0124414</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Castro</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hickson</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bettadapura</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Thomaz</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Abowd</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Christensen</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Predicting daily activities from egocentric images using deep learning</article-title>. <source>ACM Int. Symposium Wearable Comput.</source> <volume>2015</volume>, <fpage>75</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1145/2802083.2808398</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cha</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Seo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Baek</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Koo</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Towards a well-planned, activity-based work environment: Automated recognition of office activities using accelerometers</article-title>. <source>Build. Environ.</source> <volume>144</volume>, <fpage>86</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1016/j.buildenv.2018.07.051</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chalapathy</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Chawla</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Deep learning for anomaly detection: A survey</source>. <comment>
<italic>arXiv preprint arXiv:1901.03407</italic>
</comment>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chan</surname>
<given-names>T. K.</given-names>
</name>
<name>
<surname>Chin</surname>
<given-names>C. S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A comprehensive review of polyphonic sound event detection</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>103339</fpage>&#x2013;<lpage>103373</lpage>. <pub-id pub-id-type="doi">10.1109/access.2020.2999388</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021a</year>). <article-title>Deep learning for sensor-based human activity recognition: Overview, challenges, and opportunities</article-title>. <source>ACM Comput. Surv.</source> <volume>54</volume>, <fpage>1</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1145/3447744</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021b</year>). <article-title>Deep learning based multimodal complex human activity recognition using wearable devices</article-title>. <source>Appl. Intell.</source> <volume>51</volume>, <fpage>4029</fpage>&#x2013;<lpage>4042</lpage>. <pub-id pub-id-type="doi">10.1007/s10489-020-02005-7</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nugent</surname>
<given-names>C. D.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Human activity recognition and behaviour analysis</source>. <publisher-loc>New York City, NY)</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Z.-Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.-H.</given-names>
</name>
<name>
<surname>Lantz</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.-Q.</given-names>
</name>
</person-group> (<year>2007a</year>). &#x201c;<article-title>Hand gesture recognition research based on surface EMG sensors and 2D-accelerometers</article-title>,&#x201d; in <source>IEEE international symposium on wearable computers</source>, <fpage>11</fpage>&#x2013;<lpage>14</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Z.-Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.-H.</given-names>
</name>
<name>
<surname>Lantz</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.-Q.</given-names>
</name>
</person-group> (<year>2007b</year>). &#x201c;<article-title>Multiple hand gesture recognition based on surface EMG signal</article-title>,&#x201d; in <source>IEEE international conference on bioinformatics and biomedical engineering</source>, <fpage>506</fpage>&#x2013;<lpage>509</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xue</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>A deep learning approach to human activity recognition based on single accelerometer</article-title>,&#x201d; in <source>IEEE international conference on systems, man, and cybernetics</source>, <fpage>1488</fpage>&#x2013;<lpage>1492</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chung</surname>
<given-names>P.-C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C.-D.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>A daily behavior enabled hidden Markov model for human behavior understanding</article-title>. <source>Pattern Recognit.</source> <volume>41</volume>, <fpage>1572</fpage>&#x2013;<lpage>1580</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2007.10.022</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cleland</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Kikhia</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Nugent</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Boytsov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hallberg</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Synnes</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Optimal placement of accelerometers for the detection of everyday activities</article-title>. <source>Sensors</source> <volume>13</volume>, <fpage>9183</fpage>&#x2013;<lpage>9200</lpage>. <pub-id pub-id-type="doi">10.3390/s130709183</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cornacchia</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ozcan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Velipasalar</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A survey on activity detection and classification using wearable sensors</article-title>. <source>IEEE Sensors J.</source> <volume>17</volume>, <fpage>386</fpage>&#x2013;<lpage>403</lpage>. <pub-id pub-id-type="doi">10.1109/jsen.2016.2628346</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>C&#xf4;t&#xe9;-Allard</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Fall</surname>
<given-names>C. L.</given-names>
</name>
<name>
<surname>Drouin</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Campeau-Lecours</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gosselin</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Glette</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Deep learning for electromyographic hand gesture signal classification using transfer learning</article-title>. <source>IEEE Trans. Neural Syst. Rehabilitation Eng.</source> <volume>27</volume>, <fpage>760</fpage>&#x2013;<lpage>771</lpage>. <pub-id pub-id-type="doi">10.1109/tnsre.2019.2896269</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dash</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ferrari</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Malik</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Automatic speech activity recognition from MEG signals using seq2seq learning</article-title>,&#x201d; in <source>IEEE engineering in medicine and biology society international conference on neural engineering</source>, <fpage>340</fpage>&#x2013;<lpage>343</lpage>.</citation>
</ref>
<ref id="B40">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Datta</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Banerjee</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Konar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tibarewala</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Electrooculogram based cognitive context recognition</article-title>,&#x201d; in <source>IEEE international conference on electronics, communication and instrumentation</source>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Socher</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.-J.</given-names>
</name>
<name>
<surname>Kai</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.-F.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Imagenet: A large-scale hierarchical image database</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition</source>, <fpage>248</fpage>&#x2013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shangguan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). &#x201c;<article-title>Femo: A platform for free-weight exercise monitoring with RFIDs</article-title>,&#x201d; in <source>ACM conference on embedded networked sensor systems</source>, <fpage>141</fpage>&#x2013;<lpage>154</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dirgov&#xe1; Lupt&#xe1;kov&#xe1;</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Kubov&#x10d;&#xed;k</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Posp&#xed;chal</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Wearable sensor-based human activity recognition with transformer model</article-title>. <source>Sensors</source> <volume>22</volume>, <fpage>1911</fpage>. <pub-id pub-id-type="doi">10.3390/s22051911</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>Recognizing interaction activities using dynamic Bayesian network</article-title>,&#x201d; in <source>IEEE international conference on pattern recognition</source> <volume>1</volume>, <fpage>618</fpage>&#x2013;<lpage>621</lpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Elbasiony</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gomaa</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>A survey on human activity recognition based on temporal signals of portable inertial sensors</article-title>,&#x201d; in <source>International conference on advanced machine learning technologies and applications</source> (<publisher-name>Springer</publisher-name>), <fpage>734</fpage>&#x2013;<lpage>745</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="web">
<collab>EMOTIV</collab> (<year>2021</year>). <article-title>Brainwear&#xae; - wireless EEG technology</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://emotiv.com/">http://emotiv.com/</ext-link>
</comment> (<comment>Accessed June 26, 2021)</comment>.</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Faisal</surname>
<given-names>A. I.</given-names>
</name>
<name>
<surname>Majumder</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mondal</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Cowan</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Naseh</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Deen</surname>
<given-names>M. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Monitoring methods of human body joints: State-of-the-art and research challenges</article-title>. <source>Sensors</source> <volume>19</volume>, <fpage>2629</fpage>. <pub-id pub-id-type="doi">10.3390/s19112629</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>When RFID meets deep learning: Exploring cognitive intelligence for activity identification</article-title>. <source>IEEE Wirel. Commun.</source> <volume>26</volume>, <fpage>19</fpage>&#x2013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1109/mwc.2019.1800405</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Fathi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Farhadi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rehg</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Understanding egocentric activities</article-title>,&#x201d; in <source>IEEE international conference on computer vision</source>, <fpage>407</fpage>&#x2013;<lpage>414</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ferrari</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Micucci</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Mobilio</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Napoletano</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>On the personalization of classification models for human activity recognition</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>32066</fpage>&#x2013;<lpage>32079</lpage>. <pub-id pub-id-type="doi">10.1109/access.2020.2973425</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fortin-Simard</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bilodeau</surname>
<given-names>J.-S.</given-names>
</name>
<name>
<surname>Bouchard</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Gaboury</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bouchard</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bouzouane</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Exploiting passive RFID technology for activity recognition in smart homes</article-title>. <source>IEEE Intell. Syst.</source> <volume>30</volume>, <fpage>7</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/mis.2015.18</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fortune</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Heard</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Adams</surname>
<given-names>J. A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Real-time speech workload estimation for intelligent human-machine systems</article-title>. <source>Hum. Factors Ergonomics Soc. Annu. Meet.</source> <volume>64</volume>, <fpage>334</fpage>&#x2013;<lpage>338</lpage>. <pub-id pub-id-type="doi">10.1177/1071181320641076</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Frank</surname>
<given-names>A. E.</given-names>
</name>
<name>
<surname>Kubota</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Riek</surname>
<given-names>L. D.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Wearable activity recognition for robust human-robot teaming in safety-critical environments via hybrid neural networks</article-title>,&#x201d; in <source>IEEE/RSJ international conference on intelligent robots and systems</source>, <fpage>449</fpage>&#x2013;<lpage>454</lpage>.</citation>
</ref>
<ref id="B54">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mishchenko</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Shah</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Matsoukas</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Vitaladevuni</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Towards data-efficient modeling for wake word spotting</article-title>,&#x201d; in <source>IEEE international conference on acoustics, speech and signal processing</source>, <fpage>7479</fpage>&#x2013;<lpage>7483</lpage>.</citation>
</ref>
<ref id="B55">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Garcia-Hernando</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Baek</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>T.-K.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>First-person hand action benchmark with RGB-D videos and 3d hand pose annotations</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition</source>, <fpage>409</fpage>&#x2013;<lpage>419</lpage>.</citation>
</ref>
<ref id="B56">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gemmeke</surname>
<given-names>J. F.</given-names>
</name>
<name>
<surname>Ellis</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Freedman</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Jansen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lawrence</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Moore</surname>
<given-names>R. C.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). &#x201c;<article-title>Audio set: An ontology and human-labeled dataset for audio events</article-title>,&#x201d; in <source>IEEE international conference on acoustics, speech and signal processing</source>, <fpage>776</fpage>&#x2013;<lpage>780</lpage>.</citation>
</ref>
<ref id="B57">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gjoreski</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Bizjak</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gjoreski</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gams</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Comparing deep and classical machine learning methods for human activity recognition using wrist accelerometer</article-title>,&#x201d; in <source>International joint conference on artificial intelligence, workshop on deep learning for artificial intelligence</source>, <volume>10</volume>, <fpage>970</fpage>.</citation>
</ref>
<ref id="B58">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gjoreski</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kozina</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gams</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lu&#x161;trek</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>RAReFall&#x2014;Real-time activity recognition and fall detection system</article-title>,&#x201d; in <source>IEEE international conference on pervasive computing and communication workshops</source>, <fpage>145</fpage>&#x2013;<lpage>147</lpage>.</citation>
</ref>
<ref id="B59">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Courville</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Deep learning</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grana</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Aguilar-Moreno</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>De Lope Asiain</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Araquistain</surname>
<given-names>I. B.</given-names>
</name>
<name>
<surname>Garmendia</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Improved activity recognition combining inertial motion sensors and electroencephalogram signals</article-title>. <source>Int. J. Neural Syst.</source> <volume>30</volume>, <fpage>2050053</fpage>. <pub-id pub-id-type="doi">10.1142/s0129065720500537</pub-id>
</citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Granger</surname>
<given-names>C. W.</given-names>
</name>
</person-group> (<year>1969</year>). <article-title>Investigating causal relations by econometric models and cross-spectral methods</article-title>. <source>Econ. J. Econ. Soc.</source> <volume>37</volume>, <fpage>424</fpage>&#x2013;<lpage>438</lpage>. <pub-id pub-id-type="doi">10.2307/1912791</pub-id>
</citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Granger</surname>
<given-names>C. W.</given-names>
</name>
</person-group> (<year>1980</year>). <article-title>Testing for causality: A personal viewpoint</article-title>. <source>J. Econ. Dyn. Control</source> <volume>2</volume>, <fpage>329</fpage>&#x2013;<lpage>352</lpage>. <pub-id pub-id-type="doi">10.1016/0165-1889(80)90069-x</pub-id>
</citation>
</ref>
<ref id="B63">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Abdulbaqi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Marsic</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>Multimodal attention network for trauma activity recognition from spoken language and environmental sound</article-title>,&#x201d; in <source>IEEE international conference on healthcare informatics</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B64">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Haubrick</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Robust audio sensing with multi-sound classification</article-title>,&#x201d; in <source>IEEE international conference on pervasive computing and communications</source>, <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B65">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deep residual learning for image recognition</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition</source>, <fpage>770</fpage>&#x2013;<lpage>778</lpage>.</citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heard</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fortune</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Adams</surname>
<given-names>J. A.</given-names>
</name>
</person-group> (<year>2019a</year>). <article-title>Speech workload estimation for human-machine interaction</article-title>. <source>Hum. Factors Ergonomics Soc. Annu. Meet.</source> <volume>63</volume>, <fpage>277</fpage>&#x2013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.1177/1071181319631018</pub-id>
</citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heard</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Harriott</surname>
<given-names>C. E.</given-names>
</name>
<name>
<surname>Adams</surname>
<given-names>J. A.</given-names>
</name>
</person-group> (<year>2018a</year>). <article-title>A survey of workload assessment algorithms</article-title>. <source>IEEE Trans. Human-Machine Syst.</source> <volume>48</volume>, <fpage>434</fpage>&#x2013;<lpage>451</lpage>. <pub-id pub-id-type="doi">10.1109/thms.2017.2782483</pub-id>
</citation>
</ref>
<ref id="B68">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Heard</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Heald</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Harriott</surname>
<given-names>C. E.</given-names>
</name>
<name>
<surname>Adams</surname>
<given-names>J. A.</given-names>
</name>
</person-group> (<year>2018b</year>). &#x201c;<article-title>A diagnostic human workload assessment algorithm for human-robot teams</article-title>,&#x201d; in <source>ACM/IEEE international conference on human-robot interaction</source>, <fpage>123</fpage>&#x2013;<lpage>124</lpage>.</citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heard</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Paris</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Scully</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>McNaughton</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ehrenfeld</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Coco</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019b</year>). <article-title>Automatic clinical procedure detection for emergency services</article-title>. <source>IEEE Eng. Med. Biol. Soc.</source> <volume>2019</volume>, <fpage>337</fpage>&#x2013;<lpage>340</lpage>. <pub-id pub-id-type="doi">10.1109/EMBC.2019.8856281</pub-id>
</citation>
</ref>
<ref id="B70">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hershey</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chaudhuri</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ellis</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Gemmeke</surname>
<given-names>J. F.</given-names>
</name>
<name>
<surname>Jansen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Moore</surname>
<given-names>R. C.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). &#x201c;<article-title>Cnn architectures for large-scale audio classification</article-title>,&#x201d; in <source>IEEE international conference on acoustics, speech and signal processing</source>, <fpage>131</fpage>&#x2013;<lpage>135</lpage>.</citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hevesi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ward</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Amiraslanov</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Pirkl</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lukowicz</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Wearable eye tracking for multisensor physical activity recognition</article-title>. <source>Int. J. Adv. Intelligent Syst.</source> <volume>10</volume>, <fpage>103</fpage>&#x2013;<lpage>116</lpage>.</citation>
</ref>
<ref id="B72">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hsiao</surname>
<given-names>C.-P.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Do</surname>
<given-names>E. Y.-L.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Tactile teacher: Sensing finger tapping in piano playing</article-title>,&#x201d; in <source>ACM international conference on tangible, embedded, and embodied interaction</source>, <fpage>257</fpage>&#x2013;<lpage>260</lpage>.</citation>
</ref>
<ref id="B73">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>RTagCare: Deep human activity recognition powered by passive computational RFID sensors</article-title>,&#x201d; in <source>IEEE asia-pacific network operations and management symposium</source>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ignatov</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Real-time human activity recognition from accelerometer data using convolutional neural networks</article-title>. <source>Appl. Soft Comput.</source> <volume>62</volume>, <fpage>915</fpage>&#x2013;<lpage>922</lpage>. <pub-id pub-id-type="doi">10.1016/j.asoc.2017.09.027</pub-id>
</citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Inoue</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Inoue</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nishida</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep recurrent neural network for mobile human activity recognition with high throughput</article-title>. <source>Artif. Life Robotics</source> <volume>23</volume>, <fpage>173</fpage>&#x2013;<lpage>185</lpage>. <pub-id pub-id-type="doi">10.1007/s10015-017-0422-x</pub-id>
</citation>
</ref>
<ref id="B76">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ishimaru</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hoshika</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kunze</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Dengel</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Towards reading trackers in the wild: Detecting reading activities by EOG glasses and deep neural networks</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing and ACM international symposium on wearable computers</source>, <fpage>704</fpage>&#x2013;<lpage>711</lpage>.</citation>
</ref>
<ref id="B77">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ishimaru</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kunze</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Weppner</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dengel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lukowicz</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2014a</year>). &#x201c;<article-title>In the blink of an eye: Combining head motion and eye blink frequency for activity recognition with Google glass</article-title>,&#x201d; in <source>ACM augmented human international conference</source>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B78">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ishimaru</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kunze</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Uema</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Inami</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tanaka</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2014b</year>). &#x201c;<article-title>Smarter eyewear: Using commercial EOG glasses for activity recognition</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing</source> (<publisher-name>Adjunct Publication</publisher-name>), <fpage>239</fpage>&#x2013;<lpage>242</lpage>.</citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Islam</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Sakamoto</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yamada</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Vargo</surname>
<given-names>A. W.</given-names>
</name>
<name>
<surname>Iwata</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Iwamura</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Self-supervised learning for reading activity classification</article-title>. <source>Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.</source> <volume>5</volume>, <fpage>1</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1145/3478088</pub-id>
</citation>
</ref>
<ref id="B80">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Iwamoto</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Shinoda</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>Finger ring device for tactile sensing and human machine interface</article-title>,&#x201d; in <source>Annual conference of society of instrument and control engineers</source>, <fpage>2132</fpage>&#x2013;<lpage>2136</lpage>.</citation>
</ref>
<ref id="B81">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jagannath</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sarcevic</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kamireddi</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Marsic</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Assessing the feasibility of speech-based activity recognition in dynamic medical settings</article-title>,&#x201d; in <source>ACM conference on human factors in computing systems</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B82">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jagannath</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sarcevic</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Marsic</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>An analysis of speech as a modality for activity recognition during complex medical teamwork</article-title>,&#x201d; in <source>ACM/EAI international conference on pervasive computing technologies for healthcare. 2018</source>, <fpage>88</fpage>&#x2013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1145/3240925.3240941</pub-id>
</citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jalal</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kamal</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A depth video-based human detection and activity recognition using multi-features and embedded hidden Markov models for health care monitoring systems</article-title>. <source>Int. J. Interact. Multimedia Artif. Intell.</source> <volume>4</volume>, <fpage>54</fpage>. <pub-id pub-id-type="doi">10.9781/ijimai.2017.447</pub-id>
</citation>
</ref>
<ref id="B84">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jia</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Human daily activity recognition by fusing accelerometer and multi-lead ECG data</article-title>,&#x201d; in <source>IEEE international conference on signal processing, communication and computing</source>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jing</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A recognition method for one-stroke finger gestures using a MEMS 3D accelerometer</article-title>. <source>IEICE Trans. Inf. Syst.</source> <volume>94</volume>, <fpage>1062</fpage>&#x2013;<lpage>1072</lpage>. <pub-id pub-id-type="doi">10.1587/transinf.e94.d.1062</pub-id>
</citation>
</ref>
<ref id="B86">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ju</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Human hand motion analysis with multisensory information</article-title>. <source>IEEE/ASME Trans. Mechatronics</source> <volume>19</volume>, <fpage>456</fpage>&#x2013;<lpage>466</lpage>. <pub-id pub-id-type="doi">10.1109/tmech.2013.2240312</pub-id>
</citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kaczmarek</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ma&#x144;kowski</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tomczy&#x144;ski</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>putEMG&#x2014;a surface electromyography hand gesture recognition dataset</article-title>. <source>Sensors</source> <volume>19</volume>, <fpage>3548</fpage>. <pub-id pub-id-type="doi">10.3390/s19163548</pub-id>
</citation>
</ref>
<ref id="B88">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kawazoe</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Di Luca</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Visell</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Tactile echoes: A wearable system for tactile augmentation of objects</article-title>,&#x201d; in <source>IEEE world haptics conference</source>, <fpage>359</fpage>&#x2013;<lpage>364</lpage>.</citation>
</ref>
<ref id="B89">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ke</surname>
<given-names>S.-R.</given-names>
</name>
<name>
<surname>Thuc</surname>
<given-names>H. L. U.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>Y.-J.</given-names>
</name>
<name>
<surname>Hwang</surname>
<given-names>J.-N.</given-names>
</name>
<name>
<surname>Yoo</surname>
<given-names>J.-H.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>K.-H.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>A review on video-based human activity recognition</article-title>. <source>Computers</source> <volume>2</volume>, <fpage>88</fpage>&#x2013;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.3390/computers2020088</pub-id>
</citation>
</ref>
<ref id="B90">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kelton</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ahn</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Balasubramanian</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Das</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Samaras</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>Reading detection in real-time</article-title>,&#x201d; in <source>ACM symposium on eye tracking research and applications</source>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B91">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kher</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Pawar</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Thakar</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Combining accelerometer data with Gabor energy feature vectors for body movements classification in ambulatory ECG signals</article-title>,&#x201d; in <source>International conference on biomedical engineering and informatics</source>, <fpage>413</fpage>&#x2013;<lpage>417</lpage>.</citation>
</ref>
<ref id="B92">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>Y.-J.</given-names>
</name>
<name>
<surname>Kang</surname>
<given-names>B.-N.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Hidden Markov model ensemble for activity recognition using tri-axis accelerometer</article-title>,&#x201d; in <source>IEEE international conference on systems, man, and cybernetics</source>, <fpage>3036</fpage>&#x2013;<lpage>3041</lpage>.</citation>
</ref>
<ref id="B93">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kolekar</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Dash</surname>
<given-names>D. P.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Hidden Markov model based human activity recognition using shape and optical flow based features</article-title>,&#x201d; in <source>IEEE region 10 conference</source>, <fpage>393</fpage>&#x2013;<lpage>397</lpage>.</citation>
</ref>
<ref id="B94">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kollmorgen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Holmqvist</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2007</year>). <source>Automatically detecting reading in eye tracking data</source>, <volume>144</volume>. <publisher-name>Lund University Cognitive Studies</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</citation>
</ref>
<ref id="B95">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Koskimaki</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Huikari</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Siirtola</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Laurinen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Roning</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Activity recognition using a wrist-worn inertial measurement unit: A case study for industrial assembly lines</article-title>,&#x201d; in <source>IEEE mediterranean conference on control and automation</source>, <fpage>401</fpage>&#x2013;<lpage>405</lpage>.</citation>
</ref>
<ref id="B96">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Koskim&#xe4;ki</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Siirtola</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>R&#xf6;ning</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>MyoGym: Introducing an open gym data set for activity recognition collected using Myo armband</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing and ACM international symposium on wearable computers</source>, <fpage>537</fpage>&#x2013;<lpage>546</lpage>.</citation>
</ref>
<ref id="B97">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kubota</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Iqbal</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Shah</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Riek</surname>
<given-names>L. D.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Activity recognition in manufacturing: The roles of motion capture and sEMG&#x2b;inertial wearables in detecting fine vs. gross motion</article-title>,&#x201d; in <source>IEEE international conference on Robotics and automation</source>, <fpage>6533</fpage>&#x2013;<lpage>6539</lpage>.</citation>
</ref>
<ref id="B98">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>John</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Human activity recognition using optical flow based feature set</article-title>,&#x201d; in <source>IEEE international Carnahan conference on security technology</source>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B99">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kunze</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kawaichi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yoshimura</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2013a</year>). &#x201c;<article-title>Towards inferring language expertise using eye tracking</article-title>,&#x201d; in <source>ACM extended abstracts on human factors in computing systems</source>, <fpage>217</fpage>&#x2013;<lpage>222</lpage>.</citation>
</ref>
<ref id="B100">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kunze</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Shiga</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ishimaru</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2013b</year>). &#x201c;<article-title>Reading activity recognition using an off-the-shelf EEG &#x2013; detecting reading activities and distinguishing genres of documents</article-title>,&#x201d; in <source>IEEE international conference on document analysis and recognition</source>, <fpage>96</fpage>&#x2013;<lpage>100</lpage>.</citation>
</ref>
<ref id="B101">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kunze</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Utsumi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Shiga</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bulling</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2013c</year>). &#x201c;<article-title>I know what you are reading: Recognition of document types using mobile eye tracking</article-title>,&#x201d; in <source>ACM international symposium on wearable computers</source>, <fpage>113</fpage>&#x2013;<lpage>116</lpage>.</citation>
</ref>
<ref id="B102">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kwapisz</surname>
<given-names>J. R.</given-names>
</name>
<name>
<surname>Weiss</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Moore</surname>
<given-names>S. A.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Activity recognition using cell phone accelerometers</article-title>. <source>ACM SigKDD Explor. Newsl.</source> <volume>12</volume>, <fpage>74</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1145/1964897.1964918</pub-id>
</citation>
</ref>
<ref id="B103">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kwon</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Suh</surname>
<given-names>S. C.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>K. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A survey of deep learning-based network anomaly detection</article-title>. <source>Clust. Comput.</source> <volume>22</volume>, <fpage>949</fpage>&#x2013;<lpage>961</lpage>. <pub-id pub-id-type="doi">10.1007/s10586-017-1117-8</pub-id>
</citation>
</ref>
<ref id="B104">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ladjailia</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bouchrika</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Merouani</surname>
<given-names>H. F.</given-names>
</name>
<name>
<surname>Harrati</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Mahfouf</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Human activity recognition via optical flow: Decomposing activities into basic actions</article-title>. <source>Neural Comput. Appl.</source> <volume>32</volume>, <fpage>16387</fpage>&#x2013;<lpage>16400</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-018-3951-x</pub-id>
</citation>
</ref>
<ref id="B105">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lagodzinski</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Shirahama</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Grzegorzek</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Codebook-based electrooculography data analysis towards cognitive activity recognition</article-title>. <source>Comput. Biol. Med.</source> <volume>95</volume>, <fpage>277</fpage>&#x2013;<lpage>287</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2017.10.026</pub-id>
</citation>
</ref>
<ref id="B106">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lan</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Heit</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Scargill</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gorlatova</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>GazeGraph: Graph-based few-shot cognitive context sensing from human visual behavior</article-title>,&#x201d; in <source>Embedded networked sensor systems</source>, <fpage>422</fpage>&#x2013;<lpage>435</lpage>.</citation>
</ref>
<ref id="B107">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Landsmann</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Augereau</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Kise</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Classification of reading and not reading behavior based on eye movement analysis</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing and ACM international symposium on wearable computers</source>, <fpage>109</fpage>&#x2013;<lpage>112</lpage>.</citation>
</ref>
<ref id="B108">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Laput</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ahuja</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Goel</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Ubicoustics: Plug-and-play acoustic activity recognition</article-title>,&#x201d; in <source>ACM symposium on user interface software and technology</source>, <fpage>213</fpage>&#x2013;<lpage>224</lpage>.</citation>
</ref>
<ref id="B109">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Laput</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Sensing fine-grained hand activity with smartwatches</article-title>,&#x201d; in <source>ACM conference on human factors in computing systems</source>.</citation>
</ref>
<ref id="B110">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Laput</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sample</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>EM-Sense: Touch recognition of uninstrumented, electrical and electromechanical objects</article-title>,&#x201d; in <source>ACM symposium on user interface software and technology</source> (<publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>157</fpage>&#x2013;<lpage>166</lpage>.</citation>
</ref>
<ref id="B111">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lara</surname>
<given-names>O. D.</given-names>
</name>
<name>
<surname>Labrador</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>A survey on human activity recognition using wearable sensors</article-title>. <source>IEEE Commun. Surv. Tutorials</source> <volume>15</volume>, <fpage>1192</fpage>&#x2013;<lpage>1209</lpage>. <pub-id pub-id-type="doi">10.1109/surv.2012.110112.00192</pub-id>
</citation>
</ref>
<ref id="B112">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lara</surname>
<given-names>O. D.</given-names>
</name>
<name>
<surname>P&#xe9;rez</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Labrador</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Posada</surname>
<given-names>J. D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Centinela: A human activity recognition system based on acceleration and vital sign data</article-title>. <source>Pervasive Mob. Comput.</source> <volume>8</volume>, <fpage>717</fpage>&#x2013;<lpage>729</lpage>. <pub-id pub-id-type="doi">10.1016/j.pmcj.2011.06.004</pub-id>
</citation>
</ref>
<ref id="B113">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>K.-F.</given-names>
</name>
<name>
<surname>Hon</surname>
<given-names>H.-W.</given-names>
</name>
<name>
<surname>Reddy</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>An overview of the SPHINX speech recognition system</article-title>. <source>IEEE Trans. Acoust. Speech, Signal Process.</source> <volume>38</volume>, <fpage>35</fpage>&#x2013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1109/29.45616</pub-id>
</citation>
</ref>
<ref id="B114">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Nicholls</surname>
<given-names>H. R.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Review article tactile sensing for mechatronics&#x2014;A state of the art survey</article-title>. <source>Mechatronics</source> <volume>9</volume>, <fpage>1</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1016/s0957-4158(98)00045-2</pub-id>
</citation>
</ref>
<ref id="B115">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>S.-M.</given-names>
</name>
<name>
<surname>Yoon</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Cho</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Human activity recognition from accelerometer data using convolutional neural network</article-title>,&#x201d; in <source>IEEE international conference on big data and smart computing</source>, <fpage>131</fpage>&#x2013;<lpage>134</lpage>.</citation>
</ref>
<ref id="B116">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>Y.-S.</given-names>
</name>
<name>
<surname>Cho</surname>
<given-names>S.-B.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Activity recognition using hierarchical hidden Markov models on a smartphone with 3d accelerometer</article-title>,&#x201d; in <source>International conference on hybrid artificial intelligence systems</source> (<publisher-name>Springer</publisher-name>), <fpage>460</fpage>&#x2013;<lpage>467</lpage>.</citation>
</ref>
<ref id="B117">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Paris</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Pinson</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Coco</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Heard</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2020a</year>). <article-title>Emergency clinical procedure detection with deep learning</article-title>. <source>IEEE Eng. Med. Biol. Soc.</source> <volume>2020</volume>, <fpage>158</fpage>&#x2013;<lpage>163</lpage>. <pub-id pub-id-type="doi">10.1109/EMBC44109.2020.9175575</pub-id>
</citation>
</ref>
<ref id="B118">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rozgi&#x107;</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Thatte</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Emken</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Annavaram</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>Multimodal physical activity recognition by fusing temporal and cepstral information</article-title>. <source>IEEE Trans. Neural Syst. Rehabilitation Eng.</source> <volume>18</volume>, <fpage>369</fpage>&#x2013;<lpage>380</lpage>. <pub-id pub-id-type="doi">10.1109/tnsre.2010.2053217</pub-id>
</citation>
</ref>
<ref id="B119">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Younes</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020b</year>). &#x201c;<article-title>Activitygan: Generative adversarial networks for data augmentation in sensor-based human activity recognition</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing</source>, <fpage>249</fpage>&#x2013;<lpage>254</lpage>.</citation>
</ref>
<ref id="B120">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Johannaman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Webman</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2016a</year>). &#x201c;<article-title>Activity recognition for medical teamwork based on passive RFID</article-title>,&#x201d; in <source>IEEE international conference on RFID</source>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/RFID.2016.7488002</pub-id>
</citation>
</ref>
<ref id="B121">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Marsic</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Sarcevic</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Burd</surname>
<given-names>R. S.</given-names>
</name>
</person-group> (<year>2016b</year>). <article-title>Deep learning for RFID-based activity recognition</article-title>. <source>ACM Conf. Embed. Netw. Sens. Syst.</source> <volume>2016</volume>, <fpage>164</fpage>&#x2013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1145/2994551.2994569</pub-id>
</citation>
</ref>
<ref id="B122">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Thomaz</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Audio-based activities of daily living (ADL) recognition with large-scale acoustic embeddings from online videos</article-title>. <source>ACM Interact. Mob. Wearable Ubiquitous Technol.</source> <volume>3</volume>, <fpage>1</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1145/3314404</pub-id>
</citation>
</ref>
<ref id="B123">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Liao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Recognizing complex activities by a temporal causal network-based model</article-title>,&#x201d; in <source>Joint European conference on machine learning and knowledge discovery in databases</source> (<publisher-name>Springer</publisher-name>), <fpage>341</fpage>&#x2013;<lpage>357</lpage>.</citation>
</ref>
<ref id="B124">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Liao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fox</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kautz</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>Hierarchical conditional random fields for GPS-based activity recognition</article-title>,&#x201d; in <source>Robotics research</source> (<publisher-name>Springer</publisher-name>), <fpage>487</fpage>&#x2013;<lpage>506</lpage>.</citation>
</ref>
<ref id="B125">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lillo</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Niebles</surname>
<given-names>J. C.</given-names>
</name>
<name>
<surname>Soto</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Sparse composition of body poses and atomic actions for human activity recognition in RGB-D videos</article-title>. <source>Image Vis. Comput.</source> <volume>59</volume>, <fpage>63</fpage>&#x2013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1016/j.imavis.2016.11.004</pub-id>
</citation>
</ref>
<ref id="B126">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Wireless sensing for human activity: A survey</article-title>. <source>IEEE Commun. Surv. Tutorials</source> <volume>22</volume>, <fpage>1629</fpage>&#x2013;<lpage>1645</lpage>. <pub-id pub-id-type="doi">10.1109/comst.2019.2934489</pub-id>
</citation>
</ref>
<ref id="B127">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Z.-G.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Towards complex activity recognition using a Bayesian network-based probabilistic generative framework</article-title>. <source>Pattern Recognit.</source> <volume>68</volume>, <fpage>295</fpage>&#x2013;<lpage>309</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2017.02.028</pub-id>
</citation>
</ref>
<ref id="B128">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Rajan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ramasarma</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Bonato</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S. I.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>The use of a finger-worn accelerometer for monitoring of hand use in ambulatory settings</article-title>. <source>IEEE J. Biomed. Health Inf.</source> <volume>23</volume>, <fpage>599</fpage>&#x2013;<lpage>606</lpage>. <pub-id pub-id-type="doi">10.1109/jbhi.2018.2821136</pub-id>
</citation>
</ref>
<ref id="B129">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>B.-Y.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>X.-P.</given-names>
</name>
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A dual model approach to EOG-based human activity recognition</article-title>. <source>Biomed. Signal Process. Control</source> <volume>45</volume>, <fpage>50</fpage>&#x2013;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.1016/j.bspc.2018.05.011</pub-id>
</citation>
</ref>
<ref id="B130">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kitani</surname>
<given-names>K. M.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Going deeper into first-person activity recognition</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition</source>, <fpage>1894</fpage>&#x2013;<lpage>1903</lpage>.</citation>
</ref>
<ref id="B131">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mannini</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Intille</surname>
<given-names>S. S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Classifier personalization for activity recognition using wrist accelerometers</article-title>. <source>IEEE J. Biomed. health Inf.</source> <volume>23</volume>, <fpage>1585</fpage>&#x2013;<lpage>1594</lpage>. <pub-id pub-id-type="doi">10.1109/jbhi.2018.2869779</pub-id>
</citation>
</ref>
<ref id="B132">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mannini</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sabatini</surname>
<given-names>A. M.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>On-line classification of human activity and estimation of walk-run speed from acceleration data using support vector machines</article-title>. <source>IEEE Eng. Med. Biol. Soc.</source> <volume>2011</volume>, <fpage>3302</fpage>&#x2013;<lpage>3305</lpage>. <pub-id pub-id-type="doi">10.1109/IEMBS.2011.6090896</pub-id>
</citation>
</ref>
<ref id="B133">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marquart</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Cabrall</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>de Winter</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Review of eye-related measures of drivers&#x2019; mental workload</article-title>. <source>Procedia Manuf.</source> <volume>3</volume>, <fpage>2854</fpage>&#x2013;<lpage>2861</lpage>. <pub-id pub-id-type="doi">10.1016/j.promfg.2015.07.783</pub-id>
</citation>
</ref>
<ref id="B134">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Martinez</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Pissaloux</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Carbone</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Towards activity recognition from eye-movements using contextual temporal learning</article-title>. <source>Integr. Computer-Aided Eng.</source> <volume>24</volume>, <fpage>1</fpage>&#x2013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.3233/ica-160520</pub-id>
</citation>
</ref>
<ref id="B135">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Masana</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ruiz</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Serrat</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>van de Weijer</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lopez</surname>
<given-names>A. M.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Metric learning for novelty and anomaly detection</source>. <comment>
<italic>arXiv preprint arXiv:1808.05492</italic>
</comment>.</citation>
</ref>
<ref id="B136">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Matsuo</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yamada</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ueno</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Naito</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>An attention-based activity recognition for egocentric video</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition workshops</source>, <fpage>565</fpage>&#x2013;<lpage>570</lpage>.</citation>
</ref>
<ref id="B137">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mayol</surname>
<given-names>W. W.</given-names>
</name>
<name>
<surname>Murray</surname>
<given-names>D. W.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Wearable hand activity recognition for event summarization</article-title>,&#x201d; in <source>IEEE international symposium on wearable computers</source>, <fpage>122</fpage>&#x2013;<lpage>129</lpage>.</citation>
</ref>
<ref id="B138">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mekruksavanich</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Jitpattanakul</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sitthithakerngkiet</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Youplao</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Yupapin</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Resnet-se: Channel attention-based deep residual network for complex activity recognition using wrist-worn wearable sensors</article-title>. <source>IEEE Access</source> <volume>10</volume>, <fpage>51142</fpage>&#x2013;<lpage>51154</lpage>. <pub-id pub-id-type="doi">10.1109/access.2022.3174124</pub-id>
</citation>
</ref>
<ref id="B139">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mesaros</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Heittola</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Diment</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Elizalde</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Shah</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Vincent</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). &#x201c;<article-title>DCASE 2017 challenge setup: Tasks, datasets and baseline system</article-title>,&#x201d; in <source>Detection and classification of acoustic scenes and events workshops</source>.</citation>
</ref>
<ref id="B140">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meyer</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Frank</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Schlebusch</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kasneci</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>U-HAR: A convolutional approach to human activity recognition combining head and eye movements for context-aware smart glasses</article-title>. <source>ACM Human-Computer Interact.</source> <volume>6</volume>, <fpage>1</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1145/3530884</pub-id>
</citation>
</ref>
<ref id="B141">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Min</surname>
<given-names>C.-H.</given-names>
</name>
<name>
<surname>Ince</surname>
<given-names>N. F.</given-names>
</name>
<name>
<surname>Tewfik</surname>
<given-names>A. H.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Early morning activity detection using acoustics and wearable wireless sensors</article-title>,&#x201d; in <source>IEEE European signal processing conference</source>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B142">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Min</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ince</surname>
<given-names>N. F.</given-names>
</name>
<name>
<surname>Tewfik</surname>
<given-names>A. H.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>Generalization capability of a wearable early morning activity detection system</article-title>,&#x201d; in <source>IEEE European signal processing conference</source>, <fpage>1556</fpage>&#x2013;<lpage>1560</lpage>.</citation>
</ref>
<ref id="B143">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Min</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cho</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Activity recognition based on wearable sensors using selection/fusion hybrid ensemble</article-title>,&#x201d; in <source>IEEE international conference on systems, man, and cybernetics</source>, <fpage>1319</fpage>&#x2013;<lpage>1324</lpage>.</citation>
</ref>
<ref id="B144">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mostow</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>K.-M.</given-names>
</name>
<name>
<surname>Nelson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Toward exploiting EEG input in a reading tutor</article-title>,&#x201d; in <source>International conference on artificial intelligence in education</source> (<publisher-name>Springer</publisher-name>), <fpage>230</fpage>&#x2013;<lpage>237</lpage>.</citation>
</ref>
<ref id="B145">
<citation citation-type="web">
<collab>MUSE</collab> (<year>2021</year>). <article-title>Interaxon inc., MuseTM</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://choosemuse.com/muse-2/">https://choosemuse.com/muse-2/</ext-link>
</comment> (<comment>Accessed June 26, 2021)</comment>.</citation>
</ref>
<ref id="B146">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nandy</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Saha</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chowdhury</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Novel features for intensive human activity recognition based on wearable and smartphone sensors</article-title>. <source>Microsyst. Technol.</source> <volume>26</volume>, <fpage>1889</fpage>&#x2013;<lpage>1903</lpage>. <pub-id pub-id-type="doi">10.1007/s00542-019-04738-z</pub-id>
</citation>
</ref>
<ref id="B147">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Neili Boualia</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Essoukri Ben Amara</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Deep full-body HPE for activity recognition from RGB frames only</article-title>,&#x201d; in <source>Informatics (MDPI)</source>, <volume>8</volume>, <fpage>2</fpage>. <pub-id pub-id-type="doi">10.3390/informatics8010002</pub-id>
</citation>
</ref>
<ref id="B148">
<citation citation-type="web">
<collab>Neurosky</collab> (<year>2021</year>). <article-title>Neurosky, biosensors, Neurosky co</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://neurosky.com/">http://neurosky.com/</ext-link>
</comment> (<comment>Accessed June 26, 2021)</comment>.</citation>
</ref>
<ref id="B149">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nweke</surname>
<given-names>H. F.</given-names>
</name>
<name>
<surname>Teh</surname>
<given-names>Y. W.</given-names>
</name>
<name>
<surname>Al-Garadi</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Alo</surname>
<given-names>U. R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep learning algorithms for human activity recognition using mobile and wearable sensor networks: State of the art and research challenges</article-title>. <source>Expert Syst. Appl.</source> <volume>105</volume>, <fpage>233</fpage>&#x2013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2018.03.056</pub-id>
</citation>
</ref>
<ref id="B150">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ord&#xf3;&#xf1;ez</surname>
<given-names>F. J.</given-names>
</name>
<name>
<surname>Roggen</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition</article-title>. <source>Sensors</source> <volume>16</volume>, <fpage>115</fpage>. <pub-id pub-id-type="doi">10.3390/s16010115</pub-id>
</citation>
</ref>
<ref id="B151">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Orr</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Nugent</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>A multi agent approach to facilitate the identification of interleaved activities</article-title>,&#x201d; in <source>International conference on digital health</source>, <fpage>126</fpage>&#x2013;<lpage>130</lpage>.</citation>
</ref>
<ref id="B152">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ozioko</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Taube</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Hersh</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Dahiya</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>SmartFingerBraille: A tactile sensing and actuation based communication glove for deafblind people</article-title>,&#x201d; in <source>IEEE international symposium on industrial electronics</source>, <fpage>2014</fpage>&#x2013;<lpage>2018</lpage>.</citation>
</ref>
<ref id="B153">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Hengel</surname>
<given-names>A. V. D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep learning for anomaly detection: A review</article-title>. <source>ACM Comput. Surv.</source> <volume>54</volume>, <fpage>1</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1145/3439950</pub-id>
</citation>
</ref>
<ref id="B154">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Parascandolo</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Huttunen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Virtanen</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Recurrent neural networks for polyphonic sound event detection in real life recordings</article-title>,&#x201d; in <source>IEEE international conference on acoustics, speech and signal processing</source>, <fpage>6440</fpage>&#x2013;<lpage>6444</lpage>.</citation>
</ref>
<ref id="B155">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pareek</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Thakkar</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A survey on video-based human action recognition: Recent updates, datasets, challenges, and applications</article-title>. <source>Artif. Intell. Rev.</source> <volume>54</volume>, <fpage>2259</fpage>&#x2013;<lpage>2322</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-020-09904-8</pub-id>
</citation>
</ref>
<ref id="B156">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Park</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>S.-Y.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Youn</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>The role of heart-rate variability parameters in activity recognition and energy-expenditure estimation using wearable sensors</article-title>. <source>Sensors</source> <volume>17</volume>, <fpage>1698</fpage>. <pub-id pub-id-type="doi">10.3390/s17071698</pub-id>
</citation>
</ref>
<ref id="B157">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Parkka</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ermes</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Korpipaa</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Mantyjarvi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Peltola</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Korhonen</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Activity classification using realistic data from wearable sensors</article-title>. <source>IEEE Trans. Inf. Technol. Biomed.</source> <volume>10</volume>, <fpage>119</fpage>&#x2013;<lpage>128</lpage>. <pub-id pub-id-type="doi">10.1109/titb.2005.856863</pub-id>
</citation>
</ref>
<ref id="B158">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pawar</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Anantakrishnan</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Chaudhuri</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Duttagupta</surname>
<given-names>S. P.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Transition detection in body movement activities for wearable ECG</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>54</volume>, <fpage>1149</fpage>&#x2013;<lpage>1152</lpage>. <pub-id pub-id-type="doi">10.1109/tbme.2007.891950</pub-id>
</citation>
</ref>
<ref id="B159">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pawar</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chaudhuri</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Duttagupta</surname>
<given-names>S. P.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>Analysis of ambulatory ECG signal</article-title>,&#x201d; in <source>IEEE engineering in medicine and biology society</source>, <fpage>3094</fpage>&#x2013;<lpage>3097</lpage>.</citation>
</ref>
<ref id="B160">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Aroma: A deep multi-task learning based simple and complex human activity recognition method using wearable sensors</article-title>. <source>ACM Interact. Mob. Wearable Ubiquitous Technol.</source> <volume>2</volume>, <fpage>1</fpage>&#x2013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1145/3214277</pub-id>
</citation>
</ref>
<ref id="B161">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pennington</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Socher</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Manning</surname>
<given-names>C. D.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Glove: Global vectors for word representation</article-title>,&#x201d; in <source>Conference on empirical methods in natural language processing</source>, <fpage>1532</fpage>&#x2013;<lpage>1543</lpage>.</citation>
</ref>
<ref id="B162">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Perera</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>V. M.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Deep transfer learning for multiple class novelty detection</article-title>,&#x201d; in <source>IEEE/CVF conference on computer vision and pattern recognition</source>, <fpage>11544</fpage>&#x2013;<lpage>11552</lpage>.</citation>
</ref>
<ref id="B163">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pirsiavash</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ramanan</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Detecting activities of daily living in first-person camera views</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition</source>, <fpage>2847</fpage>&#x2013;<lpage>2854</lpage>.</citation>
</ref>
<ref id="B164">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Porter</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Troscianko</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gilchrist</surname>
<given-names>I. D.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Effort during visual search and counting: Insights from pupillometry</article-title>. <source>Q. J. Exp. Psychol.</source> <volume>60</volume>, <fpage>211</fpage>&#x2013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1080/17470210600673818</pub-id>
</citation>
</ref>
<ref id="B165">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Povey</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ghoshal</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Boulianne</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Burget</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Glembek</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Goel</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). &#x201c;<article-title>The Kaldi speech recognition toolkit</article-title>,&#x201d; in <source>IEEE workshop on automatic speech recognition and understanding</source> (<publisher-name>IEEE Signal Processing Society), CFP11SRW-USB</publisher-name>).</citation>
</ref>
<ref id="B166">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pratap</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Hannun</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kahn</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Synnaeve</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). &#x201c;<article-title>Wav2letter&#x2b;&#x2b;: A fast open-source speech recognition system</article-title>,&#x201d; in <source>IEEE international conference on acoustics, speech and signal processing</source>, <fpage>6460</fpage>&#x2013;<lpage>6464</lpage>.</citation>
</ref>
<ref id="B167">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Privitera</surname>
<given-names>C. M.</given-names>
</name>
<name>
<surname>Renninger</surname>
<given-names>L. W.</given-names>
</name>
<name>
<surname>Carney</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Klein</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Aguilar</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Pupil dilation during visual target detection</article-title>. <source>J. Vis.</source> <volume>10</volume>, <fpage>3</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1167/10.10.3</pub-id>
</citation>
</ref>
<ref id="B168">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Rahimian</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Zabihi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Asif</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mohammadi</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Hybrid deep neural networks for sparse surface EMG-based hand gesture recognition</article-title>,&#x201d; in <source>IEEE asilomar conference on signals, systems, and computers</source>, <fpage>371</fpage>&#x2013;<lpage>374</lpage>.</citation>
</ref>
<ref id="B169">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ramanujam</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Perumal</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Padmavathi</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Human activity recognition with smartphone and wearable sensors using deep learning techniques: A review</article-title>. <source>IEEE Sensors J.</source> <volume>21</volume>, <fpage>13029</fpage>&#x2013;<lpage>13040</lpage>. <pub-id pub-id-type="doi">10.1109/jsen.2021.3069927</pub-id>
</citation>
</ref>
<ref id="B170">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ravi</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Dandekar</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Mysore</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Littman</surname>
<given-names>M. L.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Activity recognition from accelerometer data</article-title>,&#x201d; in <source>AAAI conference on innovative applications of artificial intelligence</source>, <fpage>1541</fpage>&#x2013;<lpage>1546</lpage>.</citation>
</ref>
<ref id="B171">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Reddy</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mun</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Burke</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Estrin</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hansen</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Srivastava</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Using mobile phones to determine transportation modes</article-title>. <source>ACM Trans. Sens. Netw.</source> <volume>6</volume>, <fpage>1</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1145/1689239.1689243</pub-id>
</citation>
</ref>
<ref id="B172">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Redmon</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Divvala</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Girshick</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Farhadi</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>You only look once: Unified, real-time object detection</article-title>,&#x201d; in <source>IEEE conference on computer vision and pattern recognition</source>, <fpage>779</fpage>&#x2013;<lpage>788</lpage>.</citation>
</ref>
<ref id="B173">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Reiss</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Stricker</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Introducing a new benchmarked dataset for activity monitoring</article-title>,&#x201d; in <source>International symposium on wearable computers</source> (<publisher-name>IEEE</publisher-name>), <fpage>108</fpage>&#x2013;<lpage>109</lpage>.</citation>
</ref>
<ref id="B174">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Riboni</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bettini</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Cosar: Hybrid reasoning for context-aware activity recognition</article-title>. <source>Personal Ubiquitous Comput.</source> <volume>15</volume>, <fpage>271</fpage>&#x2013;<lpage>289</lpage>. <pub-id pub-id-type="doi">10.1007/s00779-010-0331-7</pub-id>
</citation>
</ref>
<ref id="B175">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Roggen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Calatroni</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rossi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Holleczek</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>F&#xf6;rster</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Tr&#xf6;ster</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). &#x201c;<article-title>Collecting complex activity datasets in highly rich networked sensor environments</article-title>,&#x201d; in <source>IEEE international conference on networked sensing systems</source>, <fpage>233</fpage>&#x2013;<lpage>240</lpage>.</citation>
</ref>
<ref id="B176">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saez</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Baldominos</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Isasi</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A comparison study of classifier algorithms for cross-person physical activity recognition</article-title>. <source>Sensors</source> <volume>17</volume>, <fpage>66</fpage>. <pub-id pub-id-type="doi">10.3390/s17010066</pub-id>
</citation>
</ref>
<ref id="B177">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Safyan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Qayyum</surname>
<given-names>Z. U.</given-names>
</name>
<name>
<surname>Sarwar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Garc&#xed;a-Castro</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ahmed</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Ontology-driven semantic unified modelling for concurrent activity recognition (OSCAR)</article-title>. <source>Multimedia Tools Appl.</source> <volume>78</volume>, <fpage>2073</fpage>&#x2013;<lpage>2104</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-018-6318-5</pub-id>
</citation>
</ref>
<ref id="B178">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saguna</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zaslavsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chakraborty</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Complex activity recognition using context-driven activity theory and activity signatures</article-title>. <source>ACM Trans. Computer-Human Interact.</source> <volume>20</volume>, <fpage>1</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1145/2490832</pub-id>
</citation>
</ref>
<ref id="B179">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salamon</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bello</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deep convolutional neural networks and data augmentation for environmental sound classification</article-title>. <source>IEEE Signal Process. Lett.</source> <volume>24</volume>, <fpage>279</fpage>&#x2013;<lpage>283</lpage>. <pub-id pub-id-type="doi">10.1109/lsp.2017.2657381</pub-id>
</citation>
</ref>
<ref id="B180">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salehzadeh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Calitz</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Greyling</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Human activity recognition using deep electroencephalography learning</article-title>. <source>Biomed. Signal Process. Control</source> <volume>62</volume>, <fpage>102094</fpage>. <pub-id pub-id-type="doi">10.1016/j.bspc.2020.102094</pub-id>
</citation>
</ref>
<ref id="B181">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salvador</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chan</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Toward accurate dynamic time warping in linear time and space</article-title>. <source>Intell. Data Anal.</source> <volume>11</volume>, <fpage>561</fpage>&#x2013;<lpage>580</lpage>. <pub-id pub-id-type="doi">10.3233/ida-2007-11508</pub-id>
</citation>
</ref>
<ref id="B182">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sarkar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Reddy</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Dorgan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fidopiastis</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Giering</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Wearable EEG-based activity recognition in phm-related service environment via deep learning</article-title>. <source>Int. J. Prognostics Health Manag.</source> <volume>7</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.36001/ijphm.2016.v7i4.2459</pub-id>
</citation>
</ref>
<ref id="B183">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sathiyanarayanan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rajan</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Myo armband for physiotherapy healthcare: A case study using gesture recognition application</article-title>,&#x201d; in <source>IEEE international conference on communication systems and networks</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B184">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Scheme</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Englehart</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Electromyogram pattern recognition for control of powered upper-limb prostheses: State of the art and challenges for clinical use</article-title>. <source>J. Rehabilitation Res. Dev.</source> <volume>48</volume>, <fpage>643</fpage>&#x2013;<lpage>659</lpage>. <pub-id pub-id-type="doi">10.1682/jrrd.2010.09.0177</pub-id>
</citation>
</ref>
<ref id="B185">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Scherh&#xe4;ufl</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pichler</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Stelzer</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>UHF RFID localization based on phase evaluation of passive tag arrays</article-title>. <source>IEEE Trans. Instrum. Meas.</source> <volume>64</volume>, <fpage>913</fpage>&#x2013;<lpage>922</lpage>. <pub-id pub-id-type="doi">10.1109/tim.2014.2363578</pub-id>
</citation>
</ref>
<ref id="B186">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schirrmeister</surname>
<given-names>R. T.</given-names>
</name>
<name>
<surname>Springenberg</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Fiederer</surname>
<given-names>L. D. J.</given-names>
</name>
<name>
<surname>Glasstetter</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Eggensperger</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Tangermann</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Deep learning with convolutional neural networks for EEG decoding and visualization</article-title>. <source>Hum. Brain Mapp.</source> <volume>38</volume>, <fpage>5391</fpage>&#x2013;<lpage>5420</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.23730</pub-id>
</citation>
</ref>
<ref id="B187">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Schuldt</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Laptev</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Caputo</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2004</year>). &#x201c;<article-title>Recognizing human actions: A local SVM approach</article-title>,&#x201d; in <source>IEEE international conference on pattern recognition</source>, <volume>3</volume>, <fpage>32</fpage>&#x2013;<lpage>36</lpage>.</citation>
</ref>
<ref id="B188">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sha</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Lian</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W. J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>An explicable keystroke recognition algorithm for customizable ring-type keyboards</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>22933</fpage>&#x2013;<lpage>22944</lpage>. <pub-id pub-id-type="doi">10.1109/access.2020.2968495</pub-id>
</citation>
</ref>
<ref id="B189">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shakya</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Comparative study of machine learning and deep learning architecture for human activity recognition using accelerometer data</article-title>. <source>Int. J. Mach. Learn. Comput.</source> <volume>8</volume>, <fpage>577</fpage>&#x2013;<lpage>582</lpage>.</citation>
</ref>
<ref id="B190">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Simonyan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zisserman</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2014</year>). <source>Very deep convolutional networks for large-scale image recognition</source>. <comment>
<italic>arXiv preprint arXiv:1409.1556</italic>
</comment>.</citation>
</ref>
<ref id="B191">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Srivastava</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Newn</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Velloso</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Combining low and mid-level gaze features for desktop activity recognition</article-title>. <source>ACM Interact. Mob. Wearable Ubiquitous Technol.</source> <volume>2</volume>, <fpage>1</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1145/3287067</pub-id>
</citation>
</ref>
<ref id="B192">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Steil</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bulling</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Discovery of everyday human activities from long-term visual behaviour using topic models</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing</source>, <fpage>75</fpage>&#x2013;<lpage>85</lpage>.</citation>
</ref>
<ref id="B193">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Stork</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Spinello</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Silva</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Arras</surname>
<given-names>K. O.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Audio-based human activity recognition using non-markovian ensemble voting</article-title>,&#x201d; in <source>IEEE international symposium on robot and human interactive communication</source>, <fpage>509</fpage>&#x2013;<lpage>514</lpage>.</citation>
</ref>
<ref id="B194">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Learning semantics-preserving attention and contextual interaction for group activity recognition</article-title>. <source>IEEE Trans. Image Process.</source> <volume>28</volume>, <fpage>4997</fpage>&#x2013;<lpage>5012</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2019.2914577</pub-id>
</citation>
</ref>
<ref id="B195">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tapia</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Intille</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>Haskell</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Larson</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wright</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>King</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). &#x201c;<article-title>Real-time recognition of physical activities and their intensities using wireless accelerometers and a heart rate monitor</article-title>,&#x201d; in <source>IEEE international symposium on wearable computers</source>, <fpage>37</fpage>&#x2013;<lpage>40</lpage>.</citation>
</ref>
<ref id="B196">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thu</surname>
<given-names>N. T. H.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>D. S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Hihar: A hierarchical hybrid deep learning architecture for wearable sensor-based human activity recognition</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>145271</fpage>&#x2013;<lpage>145281</lpage>. <pub-id pub-id-type="doi">10.1109/access.2021.3122298</pub-id>
</citation>
</ref>
<ref id="B197">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Triboan</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Semantic segmentation of real-time sensor data stream for complex activity recognition</article-title>. <source>Personal Ubiquitous Comput.</source> <volume>21</volume>, <fpage>411</fpage>&#x2013;<lpage>425</lpage>. <pub-id pub-id-type="doi">10.1007/s00779-017-1005-5</pub-id>
</citation>
</ref>
<ref id="B198">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Trigili</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Grazi</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Crea</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Accogli</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Carpaneto</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Micera</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Detection of movement onset using EMG signals for upper-limb exoskeletons in reaching tasks</article-title>. <source>J. Neuroengineering Rehabilitation</source> <volume>16</volume>, <fpage>45</fpage>&#x2013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1186/s12984-019-0512-1</pub-id>
</citation>
</ref>
<ref id="B199">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tsai</surname>
<given-names>T.-H.</given-names>
</name>
<name>
<surname>Hao</surname>
<given-names>P.-C.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Customized wake-up word with key word spotting using convolutional neural network</article-title>,&#x201d; in <source>IEEE international SoC design conference</source>, <fpage>136</fpage>&#x2013;<lpage>137</lpage>.</citation>
</ref>
<ref id="B200">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ullah</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Muhammad</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Del Ser</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Baik</surname>
<given-names>S. W.</given-names>
</name>
<name>
<surname>de Albuquerque</surname>
<given-names>V. H. C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Activity recognition using temporal optical flow convolutional features and multilayer lstm</article-title>. <source>IEEE Trans. Industrial Electron.</source> <volume>66</volume>, <fpage>9692</fpage>&#x2013;<lpage>9702</lpage>. <pub-id pub-id-type="doi">10.1109/tie.2018.2881943</pub-id>
</citation>
</ref>
<ref id="B201">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vail</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Veloso</surname>
<given-names>M. M.</given-names>
</name>
<name>
<surname>Lafferty</surname>
<given-names>J. D.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>Conditional random fields for activity recognition</article-title>,&#x201d; in <source>International joint conference on autonomous agents and multiagent systems</source>, <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B202">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vepakomma</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>De</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Das</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Bhansali</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>A-Wristocracy: Deep learning on wrist-worn sensing for recognition of user complex activities</article-title>,&#x201d; in <source>IEEE international conference on wearable and implantable body sensor networks</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B203">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wahn</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Ferris</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Hairston</surname>
<given-names>W. D.</given-names>
</name>
<name>
<surname>K&#xf6;nig</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Pupil sizes scale with attentional load and task experience in a multiple object tracking task</article-title>. <source>PLOS ONE</source> <volume>11</volume>, <fpage>01680877</fpage>&#x2013;<lpage>e168115</lpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0168087</pub-id>
</citation>
</ref>
<ref id="B204">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>E. J.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>T.-J.</given-names>
</name>
<name>
<surname>Mariakakis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Goel</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>S. N.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Magnifisense: Inferring device interaction using wrist-worn passive magneto-inductive sensors</article-title>,&#x201d; in <source>ACM international joint conference on pervasive and ubiquitous computing</source>, <fpage>15</fpage>&#x2013;<lpage>26</lpage>.</citation>
</ref>
<ref id="B205">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Sensorygans: An effective generative adversarial framework for sensor-based human activity recognition</article-title>,&#x201d; in <source>International joint conference on neural networks</source> (<publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B206">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Deep learning for sensor-based activity recognition: A survey</article-title>. <source>Pattern Recognit. Lett.</source> <volume>119</volume>, <fpage>3</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2018.02.010</pub-id>
</citation>
</ref>
<ref id="B207">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>An incremental learning method based on probabilistic neural networks and adjustable fuzzy clustering for human activity recognition by using wearable sensors</article-title>. <source>IEEE Trans. Inf. Technol. Biomed.</source> <volume>16</volume>, <fpage>691</fpage>&#x2013;<lpage>699</lpage>. <pub-id pub-id-type="doi">10.1109/titb.2012.2196440</pub-id>
</citation>
</ref>
<ref id="B208">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weiss</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Yoneda</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Hayajneh</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Smartphone and smartwatch-based biometrics using activities of daily living</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>133190</fpage>&#x2013;<lpage>133202</lpage>. <pub-id pub-id-type="doi">10.1109/access.2019.2940729</pub-id>
</citation>
</ref>
<ref id="B209">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Weng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Xiang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). &#x201c;<article-title>A low power and high accuracy mems sensor based activity recognition algorithm</article-title>,&#x201d; in <source>IEEE international conference on bioinformatics and biomedicine</source>, <fpage>33</fpage>&#x2013;<lpage>38</lpage>.</citation>
</ref>
<ref id="B210">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Blumenstein</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Recent advances in video-based human action recognition using deep learning: A review</article-title>,&#x201d; in <source>International joint conference on neural networks</source> (<publisher-name>IEEE</publisher-name>), <fpage>2865</fpage>&#x2013;<lpage>2872</lpage>.</citation>
</ref>
<ref id="B211">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Osuntogun</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Choudhury</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Philipose</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rehg</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>A scalable approach to activity recognition based on object use</article-title>,&#x201d; in <source>IEEE international conference on computer vision</source>, <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B212">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chai</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Duan</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Innohar: A deep neural network for complex human activity recognition</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>9893</fpage>&#x2013;<lpage>9902</lpage>. <pub-id pub-id-type="doi">10.1109/access.2018.2890675</pub-id>
</citation>
</ref>
<ref id="B213">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yatani</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Truong</surname>
<given-names>K. N.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Bodyscope: A wearable acoustic sensor for activity recognition</article-title>,&#x201d; in <source>ACM conference on ubiquitous computing</source>, <fpage>341</fpage>&#x2013;<lpage>350</lpage>.</citation>
</ref>
<ref id="B214">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>K.-M.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>J. N.</given-names>
</name>
<name>
<surname>Mostow</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Toward unobtrusive measurement of reading comprehension using low-cost EEG</article-title>,&#x201d; in <source>International conference on learning analytics and knowledge</source>, <fpage>54</fpage>&#x2013;<lpage>58</lpage>.</citation>
</ref>
<ref id="B215">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.-B.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.-X.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lei</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>J.-X.</given-names>
</name>
<etal/>
</person-group> (<year>2019a</year>). <article-title>A comprehensive survey of vision-based human action recognition methods</article-title>. <source>Sensors</source> <volume>19</volume>, <fpage>1005</fpage>. <pub-id pub-id-type="doi">10.3390/s19051005</pub-id>
</citation>
</ref>
<ref id="B216">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Tham</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Detection of activities for daily life surveillance: Eating and drinking</article-title>,&#x201d; in <source>IEEE international conference on eHealth networking, applications and services</source>, <fpage>171</fpage>&#x2013;<lpage>176</lpage>.</citation>
</ref>
<ref id="B217">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Shahabi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2022a</year>). <article-title>Deep learning in human activity recognition with wearable sensors: A review on advances</article-title>. <source>Sensors</source> <volume>22</volume>, <fpage>1476</fpage>. <pub-id pub-id-type="doi">10.3390/s22041476</pub-id>
</citation>
</ref>
<ref id="B218">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lantz</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A framework for hand gesture recognition based on accelerometer and EMG sensors</article-title>. <source>IEEE Trans. Syst. Man, Cybernetics-Part A Syst. Humans</source> <volume>41</volume>, <fpage>1064</fpage>&#x2013;<lpage>1076</lpage>. <pub-id pub-id-type="doi">10.1109/tsmca.2011.2116004</pub-id>
</citation>
</ref>
<ref id="B219">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.-H.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.-H.</given-names>
</name>
<name>
<surname>Lantz</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.-Q.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Hand gesture recognition and virtual game control based on 3D accelerometer and EMG sensors</article-title>,&#x201d; in <source>ACM international conference on intelligent user interfaces</source>, <fpage>401</fpage>&#x2013;<lpage>406</lpage>.</citation>
</ref>
<ref id="B220">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019b</year>). &#x201c;<article-title>Know your mind: Adaptive cognitive activity recognition with reinforced cnn</article-title>,&#x201d; in <source>IEEE international conference on data mining</source>, <fpage>896</fpage>&#x2013;<lpage>905</lpage>.</citation>
</ref>
<ref id="B221">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Harrison</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Tomo: Wearable, low-cost electrical impedance tomography for hand gesture recognition</article-title>,&#x201d; in <source>ACM symposium on user interface software and technology</source>, <fpage>167</fpage>&#x2013;<lpage>173</lpage>.</citation>
</ref>
<ref id="B222">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2022b</year>). <article-title>IF-ConvTransformer: A framework for human activity recognition using imu fusion and convtransformer</article-title>. <source>ACM Interact. Mob. Wearable Ubiquitous Technol.</source> <volume>6</volume>, <fpage>1</fpage>&#x2013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1145/3534584</pub-id>
</citation>
</ref>
<ref id="B223">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Swears</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Larios</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Modeling temporal interactions with interval temporal Bayesian networks for complex activity recognition</article-title>. <source>IEEE Trans. Pattern Analysis Mach. Intell.</source> <volume>35</volume>, <fpage>2468</fpage>&#x2013;<lpage>2483</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2013.33</pub-id>
</citation>
</ref>
<ref id="B224">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Sarcevic</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Constructing awareness through speech, gesture, gaze and movement during a time-critical medical task</article-title>,&#x201d; in <source>European conference on computer supported cooperative work</source> (<publisher-name>Springer</publisher-name>), <fpage>163</fpage>&#x2013;<lpage>182</lpage>.</citation>
</ref>
<ref id="B225">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Chevalier</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep residual bidir-LSTM for human activity recognition using wearable sensors</article-title>. <source>Math. Problems Eng.</source> <volume>2018</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1155/2018/7316954</pub-id>
</citation>
</ref>
<ref id="B226">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Jing</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Threshold selection and adjustment for online segmentation of one-stroke finger gestures using single tri-axial accelerometer</article-title>. <source>Multimedia Tools Appl.</source> <volume>74</volume>, <fpage>9387</fpage>&#x2013;<lpage>9406</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-014-2111-2</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>