<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2022.759255</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Single Shot Corrective CNN for Anatomically Correct 3D Hand Pose Estimation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Isaac</surname> <given-names>Joseph H. R.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/987273/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Manivannan</surname> <given-names>Muniyandi</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1141358/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ravindran</surname> <given-names>Balaraman</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/549562/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Computer Science and Engineering, Indian Institute of Technology Madras</institution>, <addr-line>Chennai</addr-line>, <country>India</country></aff>
<aff id="aff2"><sup>2</sup><institution>Touch Lab, Department of Applied Mechanics, Indian Institute of Technology Madras</institution>, <addr-line>Chennai</addr-line>, <country>India</country></aff>
<aff id="aff3"><sup>3</sup><institution>Robert Bosch Center for Data Science and Artificial Intelligence (RBC-DSAI), Department of Computer Science and Engineering, Indian Institute of Technology Madras</institution>, <addr-line>Chennai</addr-line>, <country>India</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Mohan Sridharan, University of Birmingham, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Chiranjoy Chattopadhyay, Indian Institute of Technology Jodhpur, India; Kalidas Yeturu, Indian Institute of Technology Tirupati, India</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Joseph H. R. Isaac <email>joeisaac&#x00040;cse.iitm.ac.in</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Machine Learning and Artificial Intelligence, a section of the journal Frontiers in Artificial Intelligence</p></fn></author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>02</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>5</volume>
<elocation-id>759255</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>01</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Isaac, Manivannan and Ravindran.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Isaac, Manivannan and Ravindran</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Hand pose estimation in 3D from depth images is a highly complex task. Current state-of-the-art 3D hand pose estimators focus only on the accuracy of the model as measured by how closely it matches the ground truth hand pose but overlook the resulting hand pose&#x00027;s anatomical correctness. In this paper, we present the Single Shot Corrective CNN (SSC-CNN) to tackle the problem of enforcing anatomical correctness at the architecture level. In contrast to previous works which use post-facto pose filters, SSC-CNN predicts the hand pose that conforms to the human hand&#x00027;s biomechanical bounds and rules in a single forward pass. The model was trained and tested on the HANDS2017 and MSRA datasets. Experiments show that our proposed model shows comparable accuracy to the state-of-the-art models as measured by the ground truth pose. However, the previous methods have high anatomical errors, whereas our model is free from such errors. Experiments show that our proposed model shows zero anatomical errors along with comparable accuracy to the state-of-the-art models as measured by the ground truth pose. The previous methods have high anatomical errors, whereas our model is free from such errors. Surprisingly even the ground truth provided in the existing datasets suffers from anatomical errors, and therefore Anatomical Error Free (AEF) versions of the datasets, namely AEF-HANDS2017 and AEF-MSRA, were created.</p></abstract>
<kwd-group>
<kwd>3D hand pose estimation</kwd>
<kwd>biomechanical constraints</kwd>
<kwd>anatomically correct tracking</kwd>
<kwd>single shot corrective CNN</kwd>
<kwd>depth based hand tracking</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="3"/>
<equation-count count="4"/>
<ref-count count="63"/>
<page-count count="11"/>
<word-count count="8538"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Hand pose estimation in 3D is the task of predicting the pose of the hand in 3D space provided the depth (or 2D) image of the hand. It is used in many fields such as human-computer interactions (Naik et al., <xref ref-type="bibr" rid="B34">2006</xref>; Yeo et al., <xref ref-type="bibr" rid="B60">2015</xref>; Lyubanenko et al., <xref ref-type="bibr" rid="B29">2017</xref>), gesture recognition (Fang et al., <xref ref-type="bibr" rid="B15">2007</xref>), Virtual Reality (VR), and Augmented Reality (AR) (Cameron et al., <xref ref-type="bibr" rid="B3">2011</xref>; Lee et al., <xref ref-type="bibr" rid="B26">2015</xref>, <xref ref-type="bibr" rid="B25">2019</xref>; Ferche et al., <xref ref-type="bibr" rid="B16">2016</xref>). With the advent of deep learning in computer vision, commercial systems such as Oculus&#x02122; and LeapMotion&#x02122; are shifting from marker-based tracking methods to purely vision-based hand tracking. The shift obliviates the need to wear cumbersome equipment, which affects the user experience. However, marker-less pose estimation is a challenging task as there are several factors such as the complexity of the hand poses, background noise, and occlusions.</p>
<p>A key problem overlooked by several state-of-the-art models is the realism of the output hand pose. Current state-of-the-art models focus on the accuracy as per the closeness to the ground truth pose of the model rather than the overall anatomical correctness of the model and report low errors in benchmark tests such as the ICVL (Tang et al., <xref ref-type="bibr" rid="B51">2016</xref>), NYU (Tompson et al., <xref ref-type="bibr" rid="B53">2014</xref>), MSRA (Sun et al., <xref ref-type="bibr" rid="B49">2015</xref>), BigHand2.2M (Yuan et al., <xref ref-type="bibr" rid="B62">2017</xref>) and HANDS2017 (Yuan et al., <xref ref-type="bibr" rid="B62">2017</xref>). It is possible to train a model to match almost all hand joints when tested; however, when the error is caused by a finger bent in the opposite direction (as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>), it can affect the user experience. The error can also affect the human system, leading to false information and mismatch in the motor cortex and the visual system (Pelphrey et al., <xref ref-type="bibr" rid="B37">2005</xref>). Hence, in this work, we focus on improving the realism of the predicted hand pose and the validity of the dataset. The main metric used for comparison in this paper is the anatomical error of the hand pose, which is computed by using the joint angles measured for each joint in the hand pose after prediction. These joint angles were compared with the true biomechanical bounds of the hand (discussed in Section 3.1.1). The absolute error between the true bound and predicted joint angle was then calculated for every joint and added together. This value is denoted as the <bold>anatomical error</bold>, and its unit is in degrees. The mean anatomical error of the joints was also reported, and this process was repeated for every hand pose prediction. During the experiments, we observe that the ground truth of the dataset itself contains many anatomical errors in many instances. We address this issue by proposing a new corrected ground truth that conforms to the anatomical bounds of a true human hand.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Example of an anatomically incorrect pose. Although most of the joints match the original pose, since two joints are in abnormal angles, the whole pose is considered implausible.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0001.tif"/>
</fig>
<p>Earlier approaches to incorporate anatomical information usually take the form of a hand pose filter applied post-facto after the prediction of the pose to correct for anatomical errors (Chen Chen et al., <xref ref-type="bibr" rid="B5">2013</xref>; Tompson et al., <xref ref-type="bibr" rid="B53">2014</xref>; Aristidou, <xref ref-type="bibr" rid="B1">2018</xref>). Post-processing often leads to significant computational overhead. We present a novel approach that we call the Single Shot Corrective CNN (SSC-CNN) that provides a highly accurate hand pose estimation with no anatomical errors by applying corrective functions in the forward pass of the neural network using three separate networks. The term &#x0201C;Single Shot&#x0201D; implies that the model will process the hand-pose and ensure the anatomical correctness in a single forward pass of the network. This ensures that the initial prediction from the network is free of anatomical errors and prevents the need for any correction using a post-processing function.</p>
<p>In summary, our paper provides the following contributions: (1) A novel approach by incorporating biomechanical filter functions in the model architecture, (2) a hand pose estimator that guarantees zero anatomical error while maintaining low deviation from the ground truth pose, (3) we show that these anatomical rules and bounds were not maintained when creating the HANDS2017 and the MSRA hand datasets, and (4) an Anatomical Error Free (AEF) version of the datasets called AEF-HANDS2017 and AEF-MSRA was created. In Section 2 we discuss recent hand pose estimation methods. The proposed architecture is described in Section 3 along with the biomechanical constraints. The experiments to compare the proposed model with the state-of-the-art pose estimators are described in Section 4, and its results are shown in Section 5. We conclude the paper along with our future works in Section 6.</p>
</sec>
<sec id="s2">
<title>2. Related Works</title>
<p>This section discusses hand pose estimation methods that use deep learning algorithms and hand pose estimators with biomechanics-related features such as anatomical bounds.</p>
<sec>
<title>2.1. Pose Estimation With Deep Learning</title>
<p>Hand pose estimation using deep learning algorithms can be classified into discriminative and model-based methods. The former category directly regresses the joint locations of the hand using deep networks such as CNNs (Ge et al., <xref ref-type="bibr" rid="B18">2017</xref>; Guo et al., <xref ref-type="bibr" rid="B20">2017</xref>; Simon et al., <xref ref-type="bibr" rid="B45">2017</xref>; Malik et al., <xref ref-type="bibr" rid="B30">2018a</xref>; Moon et al., <xref ref-type="bibr" rid="B33">2018</xref>; Rad et al., <xref ref-type="bibr" rid="B40">2018</xref>; Cai et al., <xref ref-type="bibr" rid="B2">2019</xref>; Chen et al., <xref ref-type="bibr" rid="B8">2019</xref>; Poier et al., <xref ref-type="bibr" rid="B38">2019</xref>; Xiong et al., <xref ref-type="bibr" rid="B58">2019</xref>). The latter category abstracts a model of the human hand and fits the model with minimum error (such as the mean distance between ground truth and predicted hand pose joints) on the input data (Vollmer et al., <xref ref-type="bibr" rid="B54">1999</xref>; Taylor et al., <xref ref-type="bibr" rid="B52">2016</xref>; Oberweger and Lepetit, <xref ref-type="bibr" rid="B35">2017</xref>; Ge et al., <xref ref-type="bibr" rid="B19">2018b</xref>; Malik et al., <xref ref-type="bibr" rid="B31">2018b</xref>). Directly regressing the joint locations achieves high accuracy poses but suffers from issues such as the hand&#x00027;s structural properties. Works such as Li and Lee (<xref ref-type="bibr" rid="B27">2019</xref>) and Xiong et al. (<xref ref-type="bibr" rid="B58">2019</xref>) used cost functions taking only the joint locations of the hands into account and no structural properties of the hand. Moon et al. (<xref ref-type="bibr" rid="B33">2018</xref>) proposed the V2V Posenet, which converts the 2D depth image into a 3D voxelized grid and then predicts the joint positions of the hand. The cost function of the V2V algorithm used the joint locations alone for training and did not consider biomechanical constraints such as the joint angles.</p>
</sec>
<sec>
<title>2.2. Pose Estimation With Biomechanical Constraints</title>
<p>Biomechanical constraints are well studied in earlier works to enable anatomically correct hand poses using structural limits of the hands (Ryf and Weymann, <xref ref-type="bibr" rid="B43">1995</xref>; Cobos et al., <xref ref-type="bibr" rid="B11">2008</xref>; Chen Chen et al., <xref ref-type="bibr" rid="B5">2013</xref>; Melax et al., <xref ref-type="bibr" rid="B32">2013</xref>; Sridhar et al., <xref ref-type="bibr" rid="B47">2013</xref>; Xu and Cheng, <xref ref-type="bibr" rid="B59">2013</xref>; Tompson et al., <xref ref-type="bibr" rid="B53">2014</xref>; Poier et al., <xref ref-type="bibr" rid="B39">2015</xref>; Dibra et al., <xref ref-type="bibr" rid="B13">2017</xref>; Aristidou, <xref ref-type="bibr" rid="B1">2018</xref>; Wan et al., <xref ref-type="bibr" rid="B55">2019</xref>; Spurr et al., <xref ref-type="bibr" rid="B46">2020</xref>). Some works such as Cai et al. (<xref ref-type="bibr" rid="B2">2019</xref>) used refinement models to adjust the poses with limits and rules. However, most of these works (Ryf and Weymann, <xref ref-type="bibr" rid="B43">1995</xref>; Cobos et al., <xref ref-type="bibr" rid="B11">2008</xref>; Chen Chen et al., <xref ref-type="bibr" rid="B5">2013</xref>; Melax et al., <xref ref-type="bibr" rid="B32">2013</xref>; Sridhar et al., <xref ref-type="bibr" rid="B47">2013</xref>; Xu and Cheng, <xref ref-type="bibr" rid="B59">2013</xref>; Tompson et al., <xref ref-type="bibr" rid="B53">2014</xref>; Aristidou, <xref ref-type="bibr" rid="B1">2018</xref>; Li et al., <xref ref-type="bibr" rid="B28">2021</xref>) apply the rules and bounds after estimating the pose of the hand using post-processing methods such as inverse kinematics and bound penalization. Recent works used biomechanical constraints for hand pose estimation using 2D images in the neural network&#x00027;s cost function to penalize the joints. Malik et al. (<xref ref-type="bibr" rid="B31">2018b</xref>) incorporated structural properties of the hand such as the finger lengths and inter-finger joint distances to provide an accurate estimation of the hand pose. The drawback of this method is that the joints&#x00027; angles are not considered for estimating the pose. Hence the resulting hand pose can still output a pose in which the joint angles can exceed the human joint bounds. Works such as Sun et al. (<xref ref-type="bibr" rid="B48">2017</xref>) and Zhou et al. (<xref ref-type="bibr" rid="B63">2017</xref>) successfully implemented bone length-based constraints on human pose estimation but only on the whole body and not for the intricate parts of the hand such as finger length constraints. The model designed by Spurr et al. (<xref ref-type="bibr" rid="B46">2020</xref>) achieved better accuracy when tested on 2D datasets; however, the model was weakly supervised, and bound constraints were soft. Hence there are poses where the joint angles exceed the anatomical bounds. Li et al. (<xref ref-type="bibr" rid="B28">2021</xref>) used a model-based iterative approach by first applying the PoseNet (Choi et al., <xref ref-type="bibr" rid="B10">2020</xref>) and then computing the motion parameters. The drawback of this approach is that it depends on the PoseNet for recovering the primary joint positions and fails to operate if PoseNet fails to predict the pose. Moreover, the resulting search space of the earlier networks still includes implausible hand poses as these models only rely on the training dataset to learn the kinematic rules. We encoded the biomechanical rules as a closed-form expression that does not require any form of training. SSC-CNN&#x00027;s search space is hence much smaller than the aforementioned models. In our approach, the hand joint locations and their respective angles are predicted, and the bounds were implicitly applied to the model such that the joint angle always lies between them. Also, as pointed out in Section 5, many datasets themselves are not free from anatomical errors due to errors during annotation, and hence learning kinematic structures based on the dataset alone might lead to absorbing those errors into our model. To the best of our knowledge, our work is the first to propose incorporating anatomical constraints implicitly in the neural architecture.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Proposed Framework</title>
<p>This paper&#x00027;s primary goal is to present a framework that provides hand poses that conform to the hand&#x00027;s biomechanical rules and bounds. This goal is achieved by applying the rules implicitly into the forward pass of the neural network. The code for this model is publicly available<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> and the overall architecture is shown in <bold>Figure 3</bold>.</p>
<sec>
<title>3.1. Biomechanical Structure of the Hand</title>
<p>The human hand comprises 27 bones with 39 active muscles, which enable complex tasks such as grasping and pointing (Schwarz and Taylor, <xref ref-type="bibr" rid="B44">1955</xref>; Ross and Lamperti, <xref ref-type="bibr" rid="B42">2006</xref>; Kehr and Graftiaux, <xref ref-type="bibr" rid="B24">2017</xref>). The hand can be simplified to 21 joint locations, each with its degree of freedom and range of motion, and is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The fingers of the hand are labeled as the thumb, index, middle, ring, and pinky finger. The key joints for the movements of the hand are (1) Carpometacarpal (CMC) joint, (2) Metacarpophalangeal (MCP) joint, (3) Distal interphalangeal (DIP) joint, and (4) Proximal interphalangeal (PIP) joint.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Structure of the human hand. The various joint names are: Metacarpophalangeal (MCP) joint, Distal interphalangeal (DIP) joint, Proximal interphalangeal (PIP) joint, and Carpometacarpal (CMC) joint.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0002.tif"/>
</fig>
<p>The wrist joint is the root of the hand and is simplified to 6 degrees of freedom (DoF) as it is the result of the chain of movements from the shoulder to the arm. The CMC joint is connected to the wrist joint, and consists of 3 DoFs as per (Chim, <xref ref-type="bibr" rid="B9">2017</xref>): (1) abduction/adduction, (2) flexion/extension, and (3) rotation. One MCP joint is connected to the CMC joint, while the other 4 MCP joints are connected to the wrist. The thumb MCP joint&#x00027;s function is slightly different from the other MCP joints as the thumb MCP joint has only 1 DoF (flexion/extension), whereas the other MCP joints have 2 DoFs each. The remaining joints are the interphalangeal (IP) joints comprised of two types, namely distal and proximal (DIP and PIP) joints. The IP joints are 1 DoF each for flexion and extension alone. The thumb has one DIP joint and does not have a PIP joint.</p>
<sec>
<title>3.1.1. Biomechanical Bounds</title>
<p>According to the works of Ross and Lamperti (<xref ref-type="bibr" rid="B42">2006</xref>) and Hochschild (<xref ref-type="bibr" rid="B22">2015</xref>), the angular bounds of the joints are consolidated and these rules and bounds are all incorporated in the SSC-CNN architecture:</p>
<list list-type="bullet">
<list-item><p>CMC joint: 45&#x000B0; abduction and 0&#x000B0; adduction, 20&#x000B0; flexion and 45&#x000B0; extension, and 10&#x000B0; of rotation.</p></list-item>
<list-item><p>Thumb MCP: flexion 80&#x000B0; and extension 0&#x000B0;. Other 4 MCPS: flexion 90&#x000B0; and extension 40&#x000B0;, as well as abduction 15&#x000B0; and adduction 15&#x000B0;.</p></list-item>
<list-item><p>PIP joints: flexion 130&#x000B0; and extension 0&#x000B0;</p></list-item>
<list-item><p>DIP joints: flexion 90&#x000B0; and extension 30&#x000B0;</p></list-item>
</list>
</sec>
</sec>
<sec>
<title>3.2. SSC-CNN Architecture</title>
<p>The Resnet50 model (He et al., <xref ref-type="bibr" rid="B21">2016</xref>) is used as a backbone for the SSC-CNN (architecture shown in <xref ref-type="fig" rid="F3">Figure 3</xref>). The layers up to &#x0201C;<bold>conv4_block6_out</bold>&#x0201D; are used (after 6 block computations of the Resnet50), and the weights were transferred from the model trained on the ImageNet dataset (Deng et al., <xref ref-type="bibr" rid="B12">2009</xref>). An input image of size 176 &#x000D7; 176 &#x000D7; 3 is provided to the pre-trained Resnet50 model. The output of this layer is then fed to a convolutional layer (1024 filters of size 3 &#x000D7; 3 with ReLU activation) and then a max-pooling layer (2 &#x000D7; 2). The size of the features at this time is 4 &#x000D7; 4 &#x000D7; 1024 which is sent through another convolutional layer and max-pooling with the same configuration as before and then flattened to a 1024-dimensional vector. The compressed set of features is passed to a single dense layer of size 512 using ReLU activation which is called the <italic>common dense layer</italic>. This is then sent to three individual networks for regressing the hand&#x00027;s various characteristics, which then predicts the pose of the hand using an assembler. The three individual networks are called: (1) PalmPoseNet, (2) AngleNet, and (3) LengthNet.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Framework architecture.The proposed architecture uses part of the Resnet50 model for feature extraction. The features then pass through two sets of convolutional layers and max pooling layers and then flattens to a common dense layer. This layer is then fed as input to three sub-nets: (1) The PalmPoseNet, which outputs an 18-dimensional vector corresponding to the 3D positions of the palm joints (root joint, MCPs and CMC), (2) the AngleNet, which outputs a 20-dimensional vector corresponding to the joint angles of the hand and (3) the LengthNet which outputs an 18-dimensional vector corresponding to the length of each finger segment of the hand. The features are then concatenated and sent as input to the assembly which then provides the joint locations as output.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0003.tif"/>
</fig>
<sec>
<title>3.2.1. PalmPoseNet</title>
<p>The PalmPoseNet predicts the joint locations of the root joint, the CMC joint of the thumb, and 4 MCP joints (the thumb MCP is excluded as the root joint is the CMC joint). These joints do not have any strong biomechanical bounds and are dependent on the user&#x00027;s palm-size and structure. Hence to make the model robust, these points are directly regressed by the PalmPoseNet. The 512 features from the common dense layer are taken as input to three dense layers which has 256 nodes each using the sigmoid activation function. The features then pass to a final dense layer with 18 nodes which also uses a sigmoid activation function and these 18 points correspond to the 3D location of the six joints.</p>
</sec>
<sec>
<title>3.2.2. AngleNet</title>
<p>The AngleNet provides the angle of each joint of the fingers. As there are five fingers, including the thumb, and each finger has four angles associated with it (as explained in Section 3.1), there are a total of 20 angles that are regressed by the AngleNet. The 512 features from the common dense layer are taken as input to three dense layers which has 256 nodes each using the sigmoid activation function. The features then pass to a final dense layer with 20 nodes that uses a sigmoid activation function. These 20 features are then used in the composition of the hand pose as described in Section 3.3.</p>
</sec>
<sec>
<title>3.2.3. LengthNet</title>
<p>The LengthNet provides the length of the individual segments of the fingers of the hand, such as the length of the part between the thumb CMC to the thumb MCP and the thumb MCP to thumb IP. The 512 features from the common dense layer are taken as input to three dense layers which has 256 nodes each using the sigmoid activation function. The features then pass to a dense layer with 15 nodes that uses a sigmoid activation function. These 15 features are relative values to calculate the segments&#x00027; lengths which are then used in the composition of the hand pose as described in Section 3.3.</p>
</sec>
</sec>
<sec>
<title>3.3. Assembly of the Pose</title>
<p>The assembly is a non-trainable portion of the architecture responsible for constructing the resulting hand pose based on the values from the previous individual networks. A sample process flow of the assembly for one finger (the thumb) is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, and this process repeats for each finger.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Overall sample process of the assembly to create a thumb of the hand pose. The initialization step shown in <bold>(A)</bold> uses the 3D positions of the joints directly regressed from the PalmPoseNet and then uses the root joint with the CMC joint to extrapolate a 3D line. The three joints are then placed on this line, and the length of each segment between the joints is taken from the LengthNet. The tip joint is first rotated using the axis, which passes the adjacent joint and is parallel to the plane created by the root joint, CMC joint, and the index MCP joint. The second and third joints rotate similar to the previous joint as shown in <bold>(B)</bold> using their axes of rotation. The last rotation uses an axis perpendicular to the line that passes the CMC and the current third joint position <bold>(C)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0004.tif"/>
</fig>
<p>The first step is to take the 3D positions from the PalmPoseNet and use these points as the reference for the fingers. Taking the thumb as an example sequence, the next step is to use the root joint and the CMC joint as line points and extrapolate the line beyond the CMC joint for placing the thumb joints as shown in <xref ref-type="fig" rid="F4">Figure 4A</xref>. The length of each segment between the joints is taken from the LengthNet.</p>
<p>The LengthNet output vector is from a sigmoid function and hence ranges from 0 to 1. These values are multiplied with a hyper-parameter (&#x003B3;) which is the longest possible length of a finger segment. Works from Sunil (<xref ref-type="bibr" rid="B50">2004</xref>) and Chan Jee and Yun (<xref ref-type="bibr" rid="B4">2016</xref>) include studies where the individual parts of the hand are measured. Using this data, we set &#x003B3; &#x0003D; 80 mm which is the maximum length of an average finger segment as per these studies. Using a sigmoid-based output for the individual lengths provides finer control for the model during training.</p>
<p>After extrapolating the thumb joints, the next step is to rotate the joints to their corresponding angles, as shown in <xref ref-type="fig" rid="F4">Figure 4B</xref>. The values from AngleNet are used for setting the angles of rotation. These values also range from 0 to 1 as the sigmoid activation function is used. Each value is then multiplied according to the biomechanical range of the joint. This ensures that the range of the angle does not overshoot or undershoot the range of the joint and is shown in equation 1.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0002A;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mtext>Upper</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mtext>Lower</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mtext>Lower</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B8;<sup><italic>i</italic></sup> is the <italic>i</italic>-th joint angle, <italic>A</italic> is a vector from the AngleNet, <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext>Lower</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the lower bound of &#x003B8;<sup><italic>i</italic></sup> and <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mstyle class="text"><mml:mtext>Upper</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the upper bound of &#x003B8;<sup><italic>i</italic></sup>. For example, the thumb IP ranges from &#x02212;30&#x000B0; (considering extension as negative) to 80&#x000B0; (flexion as positive) and if <italic>A</italic><sup><italic>i</italic></sup> &#x0003D; 0.2 then &#x003B8;<sup><italic>i</italic></sup> &#x0003D; (0.2&#x0002A;(80&#x02212;(&#x02212;30)) &#x0002B; (&#x02212;30) &#x0003D; &#x02212;8. This value lies in the range [&#x02212;30, 80].</p>
<p>To rotate the joint by an angle, a reference plane is required. The root joint is always used as one point of the reference plane, while two adjacent MCP joints will be used as the other two points for each finger. The plane used for the thumb rotation is formed by the root joint, CMC joint, and the index MCP joint. Similarly, for the index finger, the index MCP and the middle MCP is used. For the last finger, i.e., the pinky, the pinky MCP and the ring MCP are used, and the rotation signs are inverted.</p>
<p>The finger&#x00027;s tip is the first joint to be rotated, as shown in <xref ref-type="fig" rid="F4">Figure 4B</xref>. To rotate the joint, an axis of rotation must be calculated. This axis is created using a vector from the adjacent joint, which lies on the reference plane and is perpendicular to the line from the first joint to the adjacent joint. After the first joint rotation, the next joint in the chain is rotated using an axis vector constructed in a similar fashion to the first joint and originating from the next adjacent joint. The second joint rotation is also applied on the first joint using the same axis of rotation. The chain then continues for the third joint using the axis originating from the next adjacent joint in the line where the first and second joint also rotates. After the three rotations are performed, the last rotation takes place with an axis originating from the last joint (CMC in case of thumb and MCP for other fingers) and is projected perpendicular to the line from the current joint to the previous joint. The three joints are then rotated around this axis as shown in <xref ref-type="fig" rid="F4">Figure 4C</xref>. This whole process is repeated for each finger, resulting in the overall pose of the hand.</p>
</sec>
<sec>
<title>3.4. Loss Function of SSC-CNN</title>
<p>As the assembly module is non-trainable, the loss function is calculated using the 53-dimensional vector after the concatenation phase. The assembly process is invertible and hence the joint locations of the ground truth is converted to the target 53-dimensional vector and the loss function is calculated as shown in equation 2.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:mo stretchy="true">&#x02225;</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">GT</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo stretchy="true">&#x02225;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>15</mml:mn></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>15</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:mo stretchy="true">|</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle class="mbox"><mml:mtext>GT</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="true">|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>20</mml:mn></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>20</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:mo stretchy="true">|</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">GT</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="true">|</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>P</italic> is the vector of joint locations from the PalmPoseNet, <italic>L</italic> is the vector of lengths derived from the LengthNet and A is the vector of angles derived from the AngleNet. <italic>P</italic><sub>GT</sub>, <italic>L</italic><sub>GT</sub>, <italic>A</italic><sub>GT</sub> are the ground truth vectors which are derived by using the reverse assembly process. The loss <italic>L</italic> consists of three parts, (1) the mean euclidean distance between the ground truth and predicted PalmPoseNet joint locations, (2) the mean absolute difference between the ground truth and predicted lengths and (3) the mean absolute difference between the ground truth and predicted angles. Hence the gradients are computed on the pre-final output that comes before the assembly phase and not on the assembled pose (joint locations) of the hand.</p>
</sec>
<sec>
<title>3.5. Dataset Used</title>
<p>The proposed framework was tested on two popular datasets, namely the MSRA (Sun et al., <xref ref-type="bibr" rid="B49">2015</xref>) and HANDS2017 (Yuan et al., <xref ref-type="bibr" rid="B62">2017</xref>) datasets. These datasets were used as they use the true joint locations such as the MCP joint and CMC joint locations compared to the edge centers used by the NYU (Tompson et al., <xref ref-type="bibr" rid="B53">2014</xref>) dataset. The MSRA dataset comprises about 70,000 images, and HANDS2017 has more than 900,000 train images and 250,000 test images. The SSC-CNN was trained on these dataset&#x00027;s training sets and tested their respective test sets using the same architecture. No anatomical corrections were made on the datasets&#x00027; ground truth during training to maintain consistency with other state-of-the-art models during comparison.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments Performed</title>
<p>As this framework focuses primarily on the anatomical correctness of the pose instead of the supposed accuracy as reported by other papers, we performed comparative tests of anatomical correctness on other state-of-the-art models and datasets along with the accuracy metrics.</p>
<sec>
<title>4.1. Error of Model After External Correction</title>
<p>To study the change in the accuracy of the model when correcting the anatomical error of the model, a corrector module was designed based on our earlier work (Isaac et al., <xref ref-type="bibr" rid="B23">2021</xref>) so that it can take the hand poses of the current state-of-the-art models as input and correct the anatomical errors of the model. This module is plugged into each test model and used to correct the anatomical error, and the correction&#x00027;s strength is adjusted using a factor &#x003B1;.</p>
<sec>
<title>4.1.1. Corrector Module Construction</title>
<p>The module utilizes the bounds explained in Section 3.1.1 and corrects the pose of the hand according to the bounds. The first step is to calculate the joint angles from the 3D joint locations provided by the estimator. To keep the origin of rotations and measurements consistent between hand poses, the hand pose is temporarily aligned to the XY plane using affine 3D transformations. The code for the corrector module is publicly available<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>.</p>
<p>The second step is to calculate the deviation of each joint from its limit. Considering the current joint angle of a particular joint as &#x003B8;<sub>c</sub> &#x0003D; [&#x003B8;<sub><italic>x</italic></sub>, &#x003B8;<sub><italic>y</italic></sub>, &#x003B8;<sub><italic>z</italic></sub>], where &#x003B8;<sub><italic>x</italic></sub>, &#x003B8;<sub><italic>y</italic></sub>, and &#x003B8;<sub><italic>z</italic></sub> are the individual Euler angles to each axis, the anatomical error of the particular joint is derived in equation 3.</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M5"><mml:mrow><mml:msubsup><mml:mi>&#x003B5;</mml:mi><mml:mtext>d</mml:mtext><mml:mi>&#x003B8;</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>upper</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>upper</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>lower</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>lower</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtable><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mtext>where&#x000A0;</mml:mtext><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:math></disp-formula>
<p>The third step is to correct the joint&#x00027;s angle using the error derived from equation 3. The correction&#x00027;s strength is adjusted using a factor &#x003B1; and is shown in equation 4.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M6"><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>d</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mtext>new</mml:mtext><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x0002A;</mml:mo><mml:msubsup><mml:mi>&#x003B5;</mml:mi><mml:mtext>d</mml:mtext><mml:mi>&#x003B8;</mml:mi></mml:msubsup></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>upper</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x0002A;</mml:mo><mml:msubsup><mml:mi>&#x003B5;</mml:mi><mml:mtext>d</mml:mtext><mml:mi>&#x003B8;</mml:mi></mml:msubsup></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mrow><mml:mtext>lower</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>&#x003B8;</mml:mi><mml:mtext>d</mml:mtext></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>where <italic>d</italic> &#x0003D; <italic>x, y, z</italic> and &#x003B1; &#x02208; [0, 1]. If &#x003B1; &#x0003D; 0, then there is no correction and the resultant angle is the original angle. If &#x003B1; &#x0003D; 1, then the angle is 100% corrected based on the hand&#x00027;s biomechanical rules.</p>
</sec>
</sec>
<sec>
<title>4.2. Ground Truth Validation</title>
<p>To validate the anatomical correctness of the ground truth, the anatomical error of the ground truth labels is calculated, and the ground truth is also compared with itself after external correction using the corrector module described in Section 4.1.1.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Experiment Results and Discussion</title>
<p><xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref> shows the comparison of the models using the MSRA hand gesture (Sun et al., <xref ref-type="bibr" rid="B49">2015</xref>) and HANDS2017 (Yuan et al., <xref ref-type="bibr" rid="B62">2017</xref>) datasets, respectively. A qualitative comparison is shown in <xref ref-type="fig" rid="F7">Figure 7</xref> between a pose from SSC-CNN and another state-of-the-art model. For the first graph of the two sets (<xref ref-type="fig" rid="F5">Figures 5A</xref>, <xref ref-type="fig" rid="F6">6A</xref>), the x-axis shows the maximum allowed mean anatomical error (calculated as mean per joint per hand), and the y-axis denotes the percentage of frames of the dataset, which is up to the specified mean anatomical error. For context, the steeper the curve is in the graph, the better the model in terms of anatomical correctness. Our model has no anatomical errors and hence the steepest line in both datasets. The second part of the set (<xref ref-type="fig" rid="F5">Figures 5B</xref>, <xref ref-type="fig" rid="F6">6B</xref>) shows the total anatomical error (calculated as mean per hand and not per joint to show the difference) of the model per hand frame using the correction module set at each value of &#x003B1; at steps of 0.1. The third graph (<xref ref-type="fig" rid="F5">Figures 5C</xref>, <xref ref-type="fig" rid="F6">6C</xref>) represents the 3D joint error which is the mean Euclidean distance from the predicted joint to the ground truth joint. The ground truth used in the test is not anatomically corrected and is the original ground truth. The lowermost line seen in both graphs is the dataset&#x00027;s ground truth compared with itself after anatomical correction. As seen in the graphs, the ground truth itself has high anatomical errors, and a likely cause of this anatomical discrepancy is the method used in creating the datasets.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparison of the anatomical errors and the 3D joint errors of various state-of-the-art models along with our proposed model using the MSRA hand dataset <bold>(A,C)</bold>. The ground truth is also shown for comparison as it has high anatomical errors <bold>(B)</bold>. The error can be due to the noises during the recording of the lab.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Comparison of the anatomical errors and the 3D joint errors of various state-of-the-art models along with our proposed model using the HANDS hand dataset <bold>(A,C)</bold>. The ground truth is also shown for comparison as it has high anatomical errors <bold>(B)</bold>. This can be due to the noises during the recording of the labels.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Qualitative comparison between the pose from the MSRA dataset using <bold>(A)</bold> SSC-CNN and <bold>(B)</bold> the same pose from Wang et al. (<xref ref-type="bibr" rid="B57">2018</xref>). The circle shows the part which is anatomically wrong in <bold>(B)</bold> while its correctly shown in <bold>(A)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0007.tif"/>
</fig>
<p>As shown in the works of Oberweger et al. (<xref ref-type="bibr" rid="B36">2016</xref>), the ground-truth curated in the HANDS and MSRA datasets are not exact representations of the real hand poses. The MSRA dataset uses a combination of the author&#x00027;s hand pose estimator as a reference with manual editing, which is tedious and prone to human errors as seen in <xref ref-type="table" rid="T1">Table 1</xref>. The HANDS2017 dataset was recorded using the Ascension Trakstar&#x02122; <xref ref-type="fn" rid="fn0003"><sup>3</sup></xref>bwhich is reported to have an accuracy of &#x000B1;1.4 mm and is attached on top of the finger during recording. As shown in <xref ref-type="fig" rid="F8">Figure 8</xref>, if the sensor is placed on top of the finger during the recording of poses, the joint&#x00027;s actual position will be at an offset from the recorded position of the hand. Hence the ground truth may not always be the actual position of the hand for many frames. With anatomically incorrect models, the error to the ground truth (non-corrected) can tend to 0. However, our algorithm puts emphasis on the anatomic correctness over the closeness to the ground truth. Hence this resulted in a relatively higher 3D joint error of 9.48 mm using the HANDS2017 dataset and 11.42 mm using the MSRA dataset as compared to the state-of-the-art models. However, our model shows comparable results when using the correction module as these models have very high anatomical errors, and correcting these errors increases the 3D joint location error. To help the community for future hand tracking related works, we also provide our corrector module publicly available to correct the ground truth of the HANDS2017 and MSRA datasets to create an Anatomical Error-Free (AEF) version of those datasets.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Anatomical Errors (AE) of the HANDS2017 dataset and MSRA dataset ground truth and the 3DJE of the corrected ground truth to the non-corrected version.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="center"><bold>AE with no correction (<bold>&#x000B0;</bold>)</bold></th>
<th valign="top" align="center"><bold>AE with correction (<bold>&#x000B0;</bold>)</bold></th>
<th valign="top" align="center"><bold>3DJE after correction (mm)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">HANDS2017</td>
<td valign="top" align="center">131</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">2.38</td>
</tr>
<tr>
<td valign="top" align="left">MSRA</td>
<td valign="top" align="center">116</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1.89</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Illustration to show the error in the true position of the hand when measuring using a sensor placed on top of the finger. There is a small gap since the sensor placement is superficial, and the true position of joints that lie inside the hand will have large errors.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-759255-g0008.tif"/>
</fig>
<p>To study the effect of the sub-networks, ablation studies were performed by removing the subnetworks and regressing all the joints of the hand directly. The resultant hand poses did not conform to the biomechanical rules and had joints rotated by abnormal angles as well as abnormally long finger segments at times. This behavior shows that the subnetworks ensure that the hand pose conforms to the angle bounds and proper finger lengths.</p>
<p><xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref> contain the 3D error of the hand pose estimators before and after application of the corrector module. The anatomical error before using the module is also shown in the tables. As seen in the table, our model predicts poses with no anatomical errors and has the same 3DJE as these bounds are implicitly coded in the model&#x00027;s architecture, and the resulting hand poses always conform to these bounds. Whereas other models have large anatomical errors and deviates after correction.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>3D Joint Errors (3DJE) and Anatomical Errors (AE) derived from 40,000 images from the HANDS2017 dataset with all the models.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>3DJE with no correction (mm)</bold></th>
<th valign="top" align="center"><bold>AE with no correction (<bold>&#x000B0;</bold>)</bold></th>
<th valign="top" align="center"><bold>3DJE with correction (mm)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>SSC-CNN (Ours)</bold></td>
<td valign="top" align="center">9.48</td>
<td valign="top" align="center"><bold>0</bold></td>
<td valign="top" align="center"><bold>9.48</bold></td>
</tr>
<tr>
<td valign="top" align="left">A2J (Xiong et al., <xref ref-type="bibr" rid="B58">2019</xref>)</td>
<td valign="top" align="center"><bold>8.65</bold></td>
<td valign="top" align="center">125</td>
<td valign="top" align="center">9.73</td>
</tr>
<tr>
<td valign="top" align="left">V2V Posenet (Moon et al., <xref ref-type="bibr" rid="B33">2018</xref>)</td>
<td valign="top" align="center">10.42</td>
<td valign="top" align="center">135</td>
<td valign="top" align="center">11.27</td>
</tr>
<tr>
<td valign="top" align="left"><underline>Ground truth</underline></td>
<td valign="top" align="center"><underline>0</underline></td>
<td valign="top" align="center"><underline>131</underline></td>
<td valign="top" align="center"><underline>2.38</underline></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The first column contains the name of the model, the second column has the errors of the model with no external correction, the third column has the Anatomical Error (AE) of the model without correction and the fourth column has the errors of the model after correction. Bold numbers denote the smallest value present in the column excluding the underlined values which is from the ground truth of the dataset</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>3D Joint Errors (3DJE) and Anatomical Errors (AE) derived from 20,000 images from the MSRA dataset with all the models.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>3DJE with no correction (mm)</bold></th>
<th valign="top" align="center"><bold>AE with no correction (<bold>&#x000B0;</bold>)</bold></th>
<th valign="top" align="center"><bold>3DJE with correction (mm)</bold></th>
<th valign="top" align="center"><bold>3DJE with correction compared to corrected ground truth (mm)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>SSC-CNN (Ours)</bold></td>
<td valign="top" align="center">11.42</td>
<td valign="top" align="center"><bold>0</bold></td>
<td valign="top" align="center">11.42</td>
<td valign="top" align="center">11.32</td>
</tr>
<tr>
<td valign="top" align="left">SHPR Net (Chen et al., <xref ref-type="bibr" rid="B7">2018</xref>)</td>
<td valign="top" align="center">7.86</td>
<td valign="top" align="center">98</td>
<td valign="top" align="center">8.56</td>
<td valign="top" align="center">7.92</td>
</tr>
<tr>
<td valign="top" align="left">3DCNN (Ge et al., <xref ref-type="bibr" rid="B18">2017</xref>)</td>
<td valign="top" align="center">9.48</td>
<td valign="top" align="center">85</td>
<td valign="top" align="center">10.05</td>
<td valign="top" align="center">9.55</td>
</tr>
<tr>
<td valign="top" align="left">DenseReg (Wan et al., <xref ref-type="bibr" rid="B56">2018</xref>)</td>
<td valign="top" align="center">7.73</td>
<td valign="top" align="center">107</td>
<td valign="top" align="center"><bold>8.37</bold></td>
<td valign="top" align="center">7.80</td>
</tr>
<tr>
<td valign="top" align="left">HandPointNet (Ge et al., <xref ref-type="bibr" rid="B17">2018a</xref>)</td>
<td valign="top" align="center">8.31</td>
<td valign="top" align="center">100</td>
<td valign="top" align="center">9.14</td>
<td valign="top" align="center">8.55</td>
</tr>
<tr>
<td valign="top" align="left">V2V Posenet (Moon et al., <xref ref-type="bibr" rid="B33">2018</xref>)</td>
<td valign="top" align="center"><bold>7.59</bold></td>
<td valign="top" align="center">118</td>
<td valign="top" align="center">15.37</td>
<td valign="top" align="center">14.84</td>
</tr>
<tr>
<td valign="top" align="left">CrossInfoNet (Du et al., <xref ref-type="bibr" rid="B14">2019</xref>)</td>
<td valign="top" align="center">7.96</td>
<td valign="top" align="center">103</td>
<td valign="top" align="center">8.41</td>
<td valign="top" align="center"><bold>7.75</bold></td>
</tr>
<tr>
<td valign="top" align="left">Point-to-Point (Ge et al., <xref ref-type="bibr" rid="B19">2018b</xref>)</td>
<td valign="top" align="center">7.71</td>
<td valign="top" align="center">95</td>
<td valign="top" align="center">8.51</td>
<td valign="top" align="center">7.91</td>
</tr>
<tr>
<td valign="top" align="left">REN 9x6x6 (Wang et al., <xref ref-type="bibr" rid="B57">2018</xref>)</td>
<td valign="top" align="center">9.79</td>
<td valign="top" align="center">91</td>
<td valign="top" align="center">10.20</td>
<td valign="top" align="center">9.66</td>
</tr>
<tr>
<td valign="top" align="left">Pose REN (Chen et al., <xref ref-type="bibr" rid="B6">2020</xref>)</td>
<td valign="top" align="center">8.65</td>
<td valign="top" align="center">96</td>
<td valign="top" align="center">9.19</td>
<td valign="top" align="center">8.59</td>
</tr>
<tr>
<td valign="top" align="left"><underline>Ground truth</underline></td>
<td valign="top" align="center"><underline>0</underline></td>
<td valign="top" align="center"><underline>116</underline></td>
<td valign="top" align="center"><underline>1.89</underline></td>
<td valign="top" align="center"><underline>0</underline></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The first column contains the name of the model, the second column has the errors of the model with no external correction, the third column has the Anatomical Error (AE) of the model without correction, the fourth column has the errors of the model after correction and the last column has the errors of the model when compared to a correction version of the ground truth. Bold numbers denote the smallest value present in the column excluding the underlined values which is from the ground truth of the dataset</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s6">
<title>6. Conclusion, Limitations, and Future Works</title>
<p>We proposed a novel framework called the SSC-CNN for 3D hand pose estimation with biomechanical constraints. The network has biomechanical rules and bounds encoded in the architecture level such that the resulting hand poses always lie inside the biomechanical bounds and rules of the human hand, and no post-processing is required to correct the poses. Our framework was compared to several state-of-the-art models with two datasets. Experiments have shown that the SSC-CNN has comparable results but with no anatomical errors, whereas the state-of-the-art models have very high anatomical errors. The ground truth of the datasets also has anatomical errors, and anatomically error-free versions were created.</p>
<p>Our framework has a limitation in which the training phase requires data pre-processing to derive the joint angles as these angles were not available in the datasets used. Another limitation is that our hand pose estimator does not take the velocity of the joint movements into consideration when correcting them. The angular velocity of the joints also has biomechanical constraints, and these will be incorporated in future works for the model. Although the model is highly robust for varying palm sizes, extreme cases like estimating the hand poses of children may result in inaccurate poses as the dataset used for training does not cover young children&#x00027;s hands and can be investigated in a future work.</p>
<p>Future works also include using synthetic datasets such as the MANO hands (Romero et al., <xref ref-type="bibr" rid="B41">2017</xref>) so that the ground truth will be assured of the hands&#x00027; true location along with children&#x00027;s hand poses. Using these synthetic datasets, we can also compare the spectrum of poses covered by the currently available datasets and hence cover a broader spectrum of poses for training. Analyzing the history of the hands&#x00027; motion using methods such as recurrent neural networks (Yoo et al., <xref ref-type="bibr" rid="B61">2020</xref>) instead of processing only one instance of the hand can avoid erratic motions during self-occlusions and will be investigated in another study for adding the feature to the SSC-CNN. The history can include the velocity and acceleration of the joint motions, which also have biomechanical bounds and further enhance the pose realism during hand motion tracking.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article and further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>JI is the primary author who coded the SSC-CNN and performed the tests using the HANDS2017 and MSRA datasets. JI wrote the article with guidance from MM and BR. JI supervised the full research flow from concept discussion to code implementation. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aristidou</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Hand tracking with physiological constraints</article-title>. <source>Vis. Comput</source>. <volume>34</volume>, <fpage>213</fpage>&#x02013;<lpage>228</lpage>. <pub-id pub-id-type="doi">10.1007/s00371-016-1327-8</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>Y.</given-names></name> <name><surname>Ge</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Cai</surname> <given-names>J.</given-names></name> <name><surname>Cham</surname> <given-names>T.-J.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Exploiting spatial-temporal relationships for 3d pose estimation via graph convolutional networks</article-title>, in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2272</fpage>&#x02013;<lpage>2281</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cameron</surname> <given-names>C. R.</given-names></name> <name><surname>DiValentin</surname> <given-names>L. W.</given-names></name> <name><surname>Manaktala</surname> <given-names>R.</given-names></name> <name><surname>McElhaney</surname> <given-names>A. C.</given-names></name> <name><surname>Nostrand</surname> <given-names>C. H.</given-names></name> <name><surname>Quinlan</surname> <given-names>O. J.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Hand tracking and visualization in a virtual reality simulation</article-title>, in <source>2011 IEEE Systems and Information Engineering Design Symposium</source> (<publisher-loc>Charlottesville, VA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>127</fpage>&#x02013;<lpage>132</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan Jee</surname> <given-names>S.</given-names></name> <name><surname>Yun</surname> <given-names>M. H.</given-names></name></person-group> (<year>2016</year>). <article-title>An anthropometric survey of korean hand and hand shape types</article-title>. <source>Int. J. Ind. Ergon</source>. <volume>53</volume>, <fpage>10</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.ergon.2015.10.004</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen Chen</surname> <given-names>F.</given-names></name> <name><surname>Appendino</surname> <given-names>S.</given-names></name> <name><surname>Battezzato</surname> <given-names>A.</given-names></name> <name><surname>Favetto</surname> <given-names>A.</given-names></name> <name><surname>Mousavi</surname> <given-names>M.</given-names></name> <name><surname>Pescarmona</surname> <given-names>F.</given-names></name></person-group> (<year>2013</year>). <article-title>Constraint study for a hand exoskeleton: human hand kinematics and dynamics</article-title>. <source>J. Rob</source>. <volume>2013</volume>, <fpage>910961</fpage>. <pub-id pub-id-type="doi">10.1155/2013/910961</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>Pose guided structured region ensemble network for cascaded hand pose estimation</article-title>. <source>Neurocomputing</source> <volume>395</volume>, <fpage>138</fpage>&#x02013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2018.06.097</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Kim</surname> <given-names>T.-K.</given-names></name> <name><surname>Ji</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>Shpr-net: deep semantic hand pose regression from point clouds</article-title>. <source>IEEE Access</source> <volume>6</volume>, <fpage>43425</fpage>&#x02013;<lpage>43439</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2018.2863540</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Tu</surname> <given-names>Z.</given-names></name> <name><surname>Ge</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>D.</given-names></name> <name><surname>Chen</surname> <given-names>R.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>So-handnet: self-organizing network for 3d hand pose estimation with semi-supervised learning</article-title>, in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6961</fpage>&#x02013;<lpage>6970</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chim</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>Hand and wrist anatomy and biomechanics: a comprehensive guide</article-title>. <source>Plast Reconstr. Surg</source>. <volume>140</volume>, <fpage>865</fpage>. <pub-id pub-id-type="doi">10.1097/PRS.0000000000003745</pub-id><pub-id pub-id-type="pmid">28953744</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Choi</surname> <given-names>H.</given-names></name> <name><surname>Moon</surname> <given-names>G.</given-names></name> <name><surname>Lee</surname> <given-names>K. M.</given-names></name></person-group> (<year>2020</year>). <article-title>Pose2mesh: graph convolutional network for 3d human pose and mesh recovery from a 2d human pose</article-title>, in <source>European Conference on Computer Vision</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>769</fpage>&#x02013;<lpage>787</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cobos</surname> <given-names>S.</given-names></name> <name><surname>Ferre</surname> <given-names>M.</given-names></name> <name><surname>Sanch&#x000E9;z-Ur&#x000E1;n</surname> <given-names>M. A.</given-names></name> <name><surname>Ortego</surname> <given-names>J.</given-names></name> <name><surname>Pe na</surname> <given-names>C.</given-names></name></person-group> (<year>2008</year>). <article-title>Efficient human hand kinematics for manipulation tasks</article-title>, in <source>2008 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS</source> (<publisher-loc>Nice</publisher-loc>: <publisher-name>IEEE</publisher-name>).</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.-J.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>Imagenet: a large-scale hierarchical image database</article-title>, in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>248</fpage>&#x02013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dibra</surname> <given-names>E.</given-names></name> <name><surname>Wolf</surname> <given-names>T.</given-names></name> <name><surname>Oztireli</surname> <given-names>C.</given-names></name> <name><surname>Gross</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>How to refine 3d hand pose estimation from unlabelled depth data?</article-title> in <source>2017 International Conference on 3D Vision (3DV)</source> (<publisher-loc>Qingdao</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>135</fpage>&#x02013;<lpage>144</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>K.</given-names></name> <name><surname>Lin</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Ma</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Crossinfonet: multi-task information sharing based hand pose estimation</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>9896</fpage>&#x02013;<lpage>9905</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name></person-group> (<year>2007</year>). <article-title>A real-time hand gesture recognition method</article-title>, in <source>2007 IEEE International Conference on Multimedia and Expo</source> (<publisher-loc>Beijing</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>995</fpage>&#x02013;<lpage>998</lpage>. <pub-id pub-id-type="pmid">32325709</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferche</surname> <given-names>O.</given-names></name> <name><surname>Moldoveanu</surname> <given-names>A.</given-names></name> <name><surname>Moldoveanu</surname> <given-names>F.</given-names></name></person-group> <article-title>Evaluating lightweight optical hand tracking for Virtual Reality rehabilitation</article-title>. <source>Romanian J. Hum. Comput. Interact</source>. <volume>9</volume>, <fpage>85</fpage>&#x02013;<lpage>102</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ge</surname> <given-names>L.</given-names></name> <name><surname>Cai</surname> <given-names>Y.</given-names></name> <name><surname>Weng</surname> <given-names>J.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name></person-group> (<year>2018a</year>). <article-title>Hand pointnet: 3d hand pose estimation using point sets</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>8417</fpage>&#x02013;<lpage>8426</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ge</surname> <given-names>L.</given-names></name> <name><surname>Liang</surname> <given-names>H.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name> <name><surname>Thalmann</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>3D convolutional neural networks for efficient and robust hand pose estimation from single depth images</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1991</fpage>&#x02013;<lpage>2000</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ge</surname> <given-names>L.</given-names></name> <name><surname>Ren</surname> <given-names>Z.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name></person-group> (<year>2018b</year>). <article-title>Point-to-point regression pointnet for 3d hand pose estimation</article-title>, in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>475</fpage>&#x02013;<lpage>491</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Qiao</surname> <given-names>F.</given-names></name> <name><surname>Yang</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>Region ensemble network: Improving convolutional network for hand pose estimation</article-title>, in <source>2017 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Beijing</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4512</fpage>&#x02013;<lpage>4516</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>. <pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hochschild</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <source>Functional Anatomy for Physical Therapists</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Thieme</publisher-name>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Isaac</surname> <given-names>J. H. R.</given-names></name> <name><surname>Manivannan</surname> <given-names>M.</given-names></name> <name><surname>Ravindran</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Corrective filter based on kinematics of human hand for pose estimation</article-title>. <source>Front. Virt. Reality</source> <volume>2</volume>, <fpage>92</fpage>. <pub-id pub-id-type="doi">10.3389/frvir.2021.663618</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kehr</surname> <given-names>P.</given-names></name> <name><surname>Graftiaux</surname> <given-names>A. G.</given-names></name></person-group> (<year>2017</year>). <article-title>B. Hirt, H. Seyhan, M. Wagner, r. Zumhasch: hand and wrist anatomy and biomechanics: a comprehensive guide</article-title>. <source>Eur. J. Orthopaedic Surg. Traumatol</source>. <volume>27</volume>, <fpage>1029</fpage>&#x02013;<lpage>1029</lpage>. <pub-id pub-id-type="doi">10.1007/s00590-017-1991-z</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Sinclair</surname> <given-names>M.</given-names></name> <name><surname>Gonzalez-Franco</surname> <given-names>M.</given-names></name> <name><surname>Ofek</surname> <given-names>E.</given-names></name> <name><surname>Holz</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>Torc: a virtual reality controller for in-hand high-dexterity finger interaction</article-title>, in <source>Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems</source>, <volume>Glasgow</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>P.-W.</given-names></name> <name><surname>Wang</surname> <given-names>H.-Y.</given-names></name> <name><surname>Tung</surname> <given-names>Y.-C.</given-names></name> <name><surname>Lin</surname> <given-names>J.-W.</given-names></name> <name><surname>Valstar</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>Transection: hand-based interaction for playing a game within a virtual reality game</article-title>, in <source>Proceedings of the 33rd Annual ACM Conference Extended Abstracts on Human Factors in Computing Systems</source>, <volume>Seoul</volume>, <fpage>73</fpage>&#x02013;<lpage>76</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Point-to-pose voting based hand pose estimation using residual permutation equivariant layer</article-title>, in <source>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2019-June</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>11919</fpage>&#x02013;<lpage>11928</lpage>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Nie</surname> <given-names>Y.</given-names></name> <name><surname>Mao</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>3d hand reconstruction from a single image based on biomechanical constraints</article-title>. <source>Vis. Comput</source>. <volume>37</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1007/s00371-021-02250-y</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lyubanenko</surname> <given-names>V.</given-names></name> <name><surname>Kuronen</surname> <given-names>T.</given-names></name> <name><surname>Eerola</surname> <given-names>T.</given-names></name> <name><surname>Lensu</surname> <given-names>L.</given-names></name> <name><surname>K&#x000E4;lvi&#x000E4;inen</surname> <given-names>H.</given-names></name> <name><surname>H&#x000E4;kkinen</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Multi-camera finger tracking and 3d trajectory reconstruction for hci studies</article-title>, in <source>International Conference on Advanced Concepts for Intelligent Vision Systems</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>63</fpage>&#x02013;<lpage>74</lpage>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malik</surname> <given-names>J.</given-names></name> <name><surname>Elhayek</surname> <given-names>A.</given-names></name> <name><surname>Ahmed</surname> <given-names>S.</given-names></name> <name><surname>Shafait</surname> <given-names>F.</given-names></name> <name><surname>Malik</surname> <given-names>M. I.</given-names></name> <name><surname>Stricker</surname> <given-names>D.</given-names></name></person-group> (<year>2018a</year>). <article-title>3dairsig: a framework for enabling in-air signatures using a multi-modal depth sensor</article-title>. <source>Sensors</source> <volume>18</volume>, <fpage>3872</fpage>. <pub-id pub-id-type="doi">10.3390/s18113872</pub-id><pub-id pub-id-type="pmid">30423837</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Malik</surname> <given-names>J.</given-names></name> <name><surname>Elhayek</surname> <given-names>A.</given-names></name> <name><surname>Stricker</surname> <given-names>D.</given-names></name></person-group> (<year>2018b</year>). <article-title>Structure-aware 3d hand pose regression from a single depth image</article-title>, in <source>International Conference on Virtual Reality and Augmented Reality</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>3</fpage>&#x02013;<lpage>17</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Melax</surname> <given-names>S.</given-names></name> <name><surname>Keselman</surname> <given-names>L.</given-names></name> <name><surname>Orsten</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Dynamics based 3D skeletal hand tracking</article-title>, in <source>Proceedings of the ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games</source>, <publisher-loc>Orlando, FL</publisher-loc>, <fpage>184</fpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Moon</surname> <given-names>G.</given-names></name> <name><surname>Yong Chang</surname> <given-names>J.</given-names></name> <name><surname>Mu Lee</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>V2v-posenet: Voxel-to-voxel prediction network for accurate 3d hand and human pose estimation from a single depth map</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <publisher-loc>Salt Lake City, UT</publisher-loc>, <fpage>5079</fpage>&#x02013;<lpage>5088</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Naik</surname> <given-names>G. R.</given-names></name> <name><surname>Kumar</surname> <given-names>D. K.</given-names></name> <name><surname>Singh</surname> <given-names>V. P.</given-names></name> <name><surname>Palaniswami</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>Hand gestures for hci using ica of emg</article-title>, in <source>ACM International Conference Proceeding Series, Vol</source>. <volume>237</volume>, <publisher-loc>Darlinghurst</publisher-loc>, <fpage>67</fpage>&#x02013;<lpage>72</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oberweger</surname> <given-names>M.</given-names></name> <name><surname>Lepetit</surname> <given-names>V.</given-names></name></person-group> (<year>2017</year>). <article-title>Deepprior&#x0002B;&#x0002B;: Improving fast and accurate 3d hand pose estimation</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision Workshops</source>, <publisher-loc>Honolulu, HI</publisher-loc>, <fpage>585</fpage>&#x02013;<lpage>594</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Oberweger</surname> <given-names>M.</given-names></name> <name><surname>Riegler</surname> <given-names>G.</given-names></name> <name><surname>Wohlhart</surname> <given-names>P.</given-names></name> <name><surname>Lepetit</surname> <given-names>V.</given-names></name></person-group> (<year>2016</year>). <article-title>Efficiently creating 3d training data for fine hand pose estimation</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4957</fpage>&#x02013;<lpage>4965</lpage>.</citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pelphrey</surname> <given-names>K. A.</given-names></name> <name><surname>Morris</surname> <given-names>J. P.</given-names></name> <name><surname>Michelich</surname> <given-names>C. R.</given-names></name> <name><surname>Allison</surname> <given-names>T.</given-names></name> <name><surname>McCarthy</surname> <given-names>G.</given-names></name></person-group> (<year>2005</year>). <article-title>Functional anatomy of biological motion perception in posterior temporal cortex: an fmri study of eye, mouth and hand movements</article-title>. <source>Cereb. Cortex</source> <volume>15</volume>, <fpage>1866</fpage>&#x02013;<lpage>1876</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhi064</pub-id><pub-id pub-id-type="pmid">15746001</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poier</surname> <given-names>G.</given-names></name> <name><surname>Opitz</surname> <given-names>M.</given-names></name> <name><surname>Schinagl</surname> <given-names>D.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>Murauer: Mapping unlabeled real data for label austerity</article-title>, in <source>2019 IEEE Winter Conference on Applications of Computer Vision (WACV)</source> (<publisher-loc>Waikoloa, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1393</fpage>&#x02013;<lpage>1402</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poier</surname> <given-names>G.</given-names></name> <name><surname>Roditakis</surname> <given-names>K.</given-names></name> <name><surname>Schulter</surname> <given-names>S.</given-names></name> <name><surname>Michel</surname> <given-names>D.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name> <name><surname>Argyros</surname> <given-names>A. A.</given-names></name></person-group> (<year>2015</year>). <article-title>Hybrid one-shot 3d hand pose estimation by exploiting uncertainties</article-title>, in <source>Proceedings of the British Machine Vision Conference 2015, BMVC 2015, Swansea, UK, September 7&#x02013;10, 2015</source>, eds <person-group person-group-type="editor"><name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Jones</surname> <given-names>M. W.</given-names></name> <name><surname>Tam</surname> <given-names>G. K. L.</given-names></name></person-group> (<publisher-name>BMVA Press</publisher-name>), <fpage>182</fpage>.1&#x02013;182.14.</citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rad</surname> <given-names>M.</given-names></name> <name><surname>Oberweger</surname> <given-names>M.</given-names></name> <name><surname>Lepetit</surname> <given-names>V.</given-names></name></person-group> (<year>2018</year>). <article-title>Feature mapping for learning fast and accurate 3d pose inference from synthetic images</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4663</fpage>&#x02013;<lpage>4672</lpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Romero</surname> <given-names>J.</given-names></name> <name><surname>Tzionas</surname> <given-names>D.</given-names></name> <name><surname>Black</surname> <given-names>M. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Embodied hands: modeling and capturing hands and bodies together</article-title>. <source>ACM Trans. Graph</source>. <volume>36</volume>, <fpage>1</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1145/3130800.3130883</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ross</surname> <given-names>L. M.</given-names></name> <name><surname>Lamperti</surname> <given-names>E. D.</given-names></name></person-group> (<year>2006</year>). <source>Thieme Atlas of Anatomy: General Anatomy and Musculoskeletal System</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Thieme</publisher-name>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ryf</surname> <given-names>C.</given-names></name> <name><surname>Weymann</surname> <given-names>A.</given-names></name></person-group> (<year>1995</year>). <article-title>The neutral zero method&#x02013;a principle of measuring joint function</article-title>. <source>Injury</source> <volume>26</volume>, <fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1016/0020-1383(95)90116-7</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwarz</surname> <given-names>R. J.</given-names></name> <name><surname>Taylor</surname> <given-names>C.</given-names></name></person-group> (<year>1955</year>). <article-title>The anatomy and mechanics of the human hand</article-title>. <source>Artif. Limbs</source> <volume>2</volume>, <fpage>22</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="pmid">13249858</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Simon</surname> <given-names>T.</given-names></name> <name><surname>Joo</surname> <given-names>H.</given-names></name> <name><surname>Matthews</surname> <given-names>I.</given-names></name> <name><surname>Sheikh</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Hand keypoint detection in single images using multiview bootstrapping</article-title>. <source>In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1145</fpage>&#x02013;<lpage>1153</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Spurr</surname> <given-names>A.</given-names></name> <name><surname>Iqbal</surname> <given-names>U.</given-names></name> <name><surname>Molchanov</surname> <given-names>P.</given-names></name> <name><surname>Hilliges</surname> <given-names>O.</given-names></name> <name><surname>Kautz</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Weakly supervised 3D hand pose estimation via biomechanical constraints</article-title>, in <source>Computer Vision-ECCV 2020</source>, eds <person-group person-group-type="editor"><name><surname>Vedaldi</surname> <given-names>A.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name> <name><surname>Brox</surname> <given-names>T.</given-names></name> <name><surname>Frahm</surname> <given-names>J.-M.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>211</fpage>&#x02013;<lpage>228</lpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sridhar</surname> <given-names>S.</given-names></name> <name><surname>Oulasvirta</surname> <given-names>A.</given-names></name> <name><surname>Theobalt</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>Interactive markerless articulated hand motion tracking using RGB and depth data</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2456</fpage>&#x02013;<lpage>2463</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Shang</surname> <given-names>J.</given-names></name> <name><surname>Liang</surname> <given-names>S.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Compositional human pose regression</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2602</fpage>&#x02013;<lpage>2611</lpage>.</citation>
</ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name> <name><surname>Liang</surname> <given-names>S.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Cascaded hand pose regression</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Boston, MA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>824</fpage>&#x02013;<lpage>832</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sunil</surname> <given-names>T.</given-names></name></person-group> (<year>2004</year>). <article-title>Clinical indicators of normal thumb length in adults1 1no benefits in any form have been received or will be received by a commercial party related directly or indirectly to the subject of this article</article-title>. <source>J. Hand. Surg. Am</source>. <volume>29</volume>, <fpage>489</fpage>&#x02013;<lpage>493</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhsa.2003.12.016</pub-id><pub-id pub-id-type="pmid">15140494</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>D.</given-names></name> <name><surname>Chang</surname> <given-names>H. J.</given-names></name> <name><surname>Tejani</surname> <given-names>A.</given-names></name> <name><surname>Kim</surname> <given-names>T.-K.</given-names></name></person-group> (<year>2016</year>). <article-title>Latent regression forest: structured estimation of 3d hand poses</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>39</volume>, <fpage>1374</fpage>&#x02013;<lpage>1387</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2599170</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taylor</surname> <given-names>J.</given-names></name> <name><surname>Bordeaux</surname> <given-names>L.</given-names></name> <name><surname>Cashman</surname> <given-names>T.</given-names></name> <name><surname>Corish</surname> <given-names>B.</given-names></name> <name><surname>Keskin</surname> <given-names>C.</given-names></name> <name><surname>Sharp</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences</article-title>. <source>ACM Trans. Graph</source>. <volume>35</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1145/2897824.2925965</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tompson</surname> <given-names>J.</given-names></name> <name><surname>Stein</surname> <given-names>M.</given-names></name> <name><surname>Lecun</surname> <given-names>Y.</given-names></name> <name><surname>Perlin</surname> <given-names>K.</given-names></name></person-group> (<year>2014</year>). <article-title>Real-time continuous pose recovery of human hands using convolutional networks</article-title>. <source>ACM Trans Graph</source>. <volume>33</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1145/2629500</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vollmer</surname> <given-names>J.</given-names></name> <name><surname>Mencl</surname> <given-names>R.</given-names></name> <name><surname>Mueller</surname> <given-names>H.</given-names></name></person-group> (<year>1999</year>). <source>Improved Laplacian Smoothing of Noisy Surface Meshes, Vo. 18-3</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>Wiley Online Library</publisher-name>.</citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wan</surname> <given-names>C.</given-names></name> <name><surname>Probst</surname> <given-names>T.</given-names></name> <name><surname>Gool</surname> <given-names>L. V.</given-names></name> <name><surname>Yao</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Self-supervised 3d hand pose estimation through training by fitting</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>10853</fpage>&#x02013;<lpage>10862</lpage>.</citation>
</ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wan</surname> <given-names>C.</given-names></name> <name><surname>Probst</surname> <given-names>T.</given-names></name> <name><surname>Van Gool</surname> <given-names>L.</given-names></name> <name><surname>Yao</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Dense 3d regression for hand pose estimation.</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5147</fpage>&#x02013;<lpage>5156</lpage>.</citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>Region ensemble network: towards good practices for deep 3d hand pose estimation</article-title>. <source>J. Vis. Commun. Image Represent</source>. <volume>55</volume>, <fpage>404</fpage>&#x02013;<lpage>414</lpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2018.04.005</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiong</surname> <given-names>F.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name> <name><surname>Xiao</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Yu</surname> <given-names>T.</given-names></name> <name><surname>Zhou</surname> <given-names>J. T.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>A2J: anchor-to-joint regression network for 3D articulated pose estimation from a single depth image</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision 2019-October</source> (<publisher-loc>Seoul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>793</fpage>&#x02013;<lpage>802</lpage>.</citation>
</ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Cheng</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>Efficient hand pose estimation from a single depth image</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3456</fpage>&#x02013;<lpage>3462</lpage>. <pub-id pub-id-type="pmid">30932832</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeo</surname> <given-names>H.-S.</given-names></name> <name><surname>Lee</surname> <given-names>B.-G.</given-names></name> <name><surname>Lim</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>Hand tracking and gesture recognition system for human-computer interaction using low-cost hardware</article-title>. <source>Multimed Tools Appl</source>. <volume>74</volume>, <fpage>2687</fpage>&#x02013;<lpage>2715</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-013-1501-1</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoo</surname> <given-names>C.-H.</given-names></name> <name><surname>Ji</surname> <given-names>S.</given-names></name> <name><surname>Shin</surname> <given-names>Y.-G.</given-names></name> <name><surname>Kim</surname> <given-names>S.-W.</given-names></name> <name><surname>Ko</surname> <given-names>S.-J.</given-names></name></person-group> (<year>2020</year>). <article-title>Fast and accurate 3d hand pose estimation via recurrent neural network for capturing hand articulations</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>114010</fpage>&#x02013;<lpage>114019</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3001637</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>S.</given-names></name> <name><surname>Ye</surname> <given-names>Q.</given-names></name> <name><surname>Stenger</surname> <given-names>B.</given-names></name> <name><surname>Jain</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>T.-K.</given-names></name></person-group> (<year>2017</year>). <article-title>Bighand2. 2m benchmark: hand pose dataset and state of the art analysis</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4866</fpage>&#x02013;<lpage>4874</lpage>.</citation>
</ref>
<ref id="B63">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>Q.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Xue</surname> <given-names>X.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Towards 3d human pose estimation in the wild: a weakly-supervised approach</article-title>, in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>398</fpage>&#x02013;<lpage>407</lpage>.</citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/RBC-DSAI-IITM/SSCCNN">https://github.com/RBC-DSAI-IITM/SSCCNN</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/RBC-DSAI-IITM/SSCCNN">https://github.com/RBC-DSAI-IITM/SSCCNN</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup><ext-link ext-link-type="uri" xlink:href="https://tracklab.com.au/products/brands/ndi/ascension-trakstar/">https://tracklab.com.au/products/brands/ndi/ascension-trakstar/</ext-link>, Accessed March 2021.</p></fn>
</fn-group>
</back>
</article>