<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2025.1482281</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Latent space improved masked reconstruction model for human skeleton-based action recognition</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Chen</surname> <given-names>Enqing</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Xueting</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2819372/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Guo</surname> <given-names>Xin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhu</surname> <given-names>Ying</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Dong</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Electrical and Information Engineering, Zhengzhou University</institution>, <addr-line>Zhengzhou</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>State Grid Henan Electric Power Company Information and Communication Branch</institution>, <addr-line>Zhengzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Long Cheng, North China Electric Power University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Jihong Pei, Shenzhen University, China</p>
<p>Sibo Cheng, &#x000C9;cole des ponts ParisTech (ENPC), France</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Xin Guo <email>iexguo&#x00040;zzu.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>02</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>19</volume>
<elocation-id>1482281</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>08</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>01</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Chen, Wang, Guo, Zhu and Li.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Chen, Wang, Guo, Zhu and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Human skeleton-based action recognition is an important task in the field of computer vision. In recent years, masked autoencoder (MAE) has been used in various fields due to its powerful self-supervised learning ability and has achieved good results in masked data reconstruction tasks. However, in visual classification tasks such as action recognition, the limited ability of the encoder to learn features in the autoencoder structure results in poor classification performance. We propose to enhance the encoder&#x00027;s feature extraction ability in classification tasks by leveraging the latent space of variational autoencoder (VAE) and further replace it with the latent space of vector quantized variational autoencoder (VQVAE). The constructed models are called SkeletonMVAE and SkeletonMVQVAE, respectively. In SkeletonMVAE, we constrain the latent variables to represent features in the form of distributions. In SkeletonMVQVAE, we discretize the latent variables. These help the encoder learn deeper data structures and more discriminative and generalized feature representations. The experiment results on the NTU-60 and NTU-120 datasets demonstrate that our proposed method can effectively improve the classification accuracy of the encoder in classification tasks and its generalization ability in the case of few labeled data. SkeletonMVAE exhibits stronger classification ability, while SkeletonMVQVAE exhibits stronger generalization in situations with fewer labeled data.</p></abstract>
<kwd-group>
<kwd>human skeleton-based action recognition</kwd>
<kwd>variational autoencoder</kwd>
<kwd>vector quantized variational autoencoder</kwd>
<kwd>masked reconstruction model</kwd>
<kwd>self-supervised learning</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="10"/>
<equation-count count="6"/>
<ref-count count="43"/>
<page-count count="11"/>
<word-count count="7503"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Action recognition has consistently remained an active topic of research within the realm of computer vision. Compared with other data formats such as RGB (Simonyan and Zisserman, <xref ref-type="bibr" rid="B28">2014</xref>; Feichtenhofer et al., <xref ref-type="bibr" rid="B12">2019</xref>; Buch et al., <xref ref-type="bibr" rid="B2">2017</xref>; Varol et al., <xref ref-type="bibr" rid="B32">2017</xref>) and depth information (Cao et al., <xref ref-type="bibr" rid="B3">2017</xref>; Fang et al., <xref ref-type="bibr" rid="B11">2017</xref>; Xu et al., <xref ref-type="bibr" rid="B37">2020</xref>; Chen et al., <xref ref-type="bibr" rid="B4">2021</xref>), skeleton data (Duan et al., <xref ref-type="bibr" rid="B10">2022</xref>; Thoker et al., <xref ref-type="bibr" rid="B29">2021</xref>) eliminate the interference of redundant information such as background and lighting. It has the advantages of high order, lightweight, and high robustness. With the deepening of pose estimation (Cao et al., <xref ref-type="bibr" rid="B3">2017</xref>; Lu et al., <xref ref-type="bibr" rid="B24">2023</xref>) research, the extraction of human skeleton data has become more and more fast and accurate.</p>
<p>Nowadays, the mainstream human action recognition method is still fully supervised learning, which can be divided into recurrent neural network (RNN) (Du et al., <xref ref-type="bibr" rid="B9">2015</xref>; Li et al., <xref ref-type="bibr" rid="B20">2019</xref>), convolutional neural network (CNN) (Hou et al., <xref ref-type="bibr" rid="B15">2016</xref>; Wang et al., <xref ref-type="bibr" rid="B34">2016</xref>; Banerjee et al., <xref ref-type="bibr" rid="B1">2020</xref>), graph convolutional neural network (GCN) (Yan et al., <xref ref-type="bibr" rid="B39">2018</xref>; Liu Z. et al., <xref ref-type="bibr" rid="B23">2020</xref>), and transformer (Wang et al., <xref ref-type="bibr" rid="B35">2021</xref>; Zhang et al., <xref ref-type="bibr" rid="B42">2021</xref>). These methods require a large amount of labeled data. However, it is a costly and time-demanding task to collect and label data. In addition, these methods may lead to overfitting during the learning process. To alleviate these problems, some works use self-supervised methods to learn unlabeled data. They mainly hope to learn a universal feature representation by solving pretext tasks, and then use it for downstream tasks such as contrastive learning (Dong et al., <xref ref-type="bibr" rid="B8">2023</xref>; Lin et al., <xref ref-type="bibr" rid="B21">2023</xref>). Contrastive learning allows the model to learn the feature invariance of the same skeleton sequences from different views by constructing positive and negative pairs through data augmentation. However, these methods of comparative learning pay more attention to global features, ignore the context relationship between frames, and depend on the number of comparison pairs.</p>
<p>Recently, action recognition introduced a new self-supervised method, mask encoding, and proved its effectiveness. In this method, a part of the data is masked, and the model infers the semantic information of the masked part through the context of the visible part, which can effectively capture the contextual relationship by analyzing the global and local information in the data, thereby enhancing the model&#x00027;s ability to capture intricate patterns and relationships within the data. MAE (He et al., <xref ref-type="bibr" rid="B14">2022</xref>) has achieved success in the field of image. It can still effectively restore the original data by masking the image content with high probability. This excellent performance has attracted extensive research in different fields, and this concept has been applied to 3D human skeleton action recognition task. SkeletonMAE (Wu et al., <xref ref-type="bibr" rid="B36">2023</xref>) based on human 3D skeleton sequence follows the idea of MAE. It randomly masks some frames and skeleton joints, uses the encoder&#x02014;decoder structure to learn the relationship between unmasked skeleton joints, reconstructs the masked skeleton joints, and then uses the pre-trained encoder for human skeleton-based action recognition. However, the encoders of these methods can only encode limited features, and the extracted data information is not sufficient. To enhance the feature extraction ability and generalization ability of the encoder, we propose SkeletonMVAE, which inserts the potential space of the variational autoencoder (VAE) (Kingma and Welling, <xref ref-type="bibr" rid="B18">2013</xref>) behind the encoder of SkeletonMAE. The latent variables of the variational autoencoder (VAE) are expressed in the form of distribution, allowing the encoder to learn deeper data structures and data distributions. By constraining the latent variables to a normal distribution close to the standard, the encoder can encode more discriminative feature representations, which is more conducive to the classification task. Furthermore, we propose SkeletonMVQVAE, which replaces the latent space with the latent space of the vector quantization variational autoencoder (VQVAE) (Van Den Oord et al., <xref ref-type="bibr" rid="B31">2017</xref>). In the quantization process, the change of smaller action will also lead to a sharp change in the latent vector, so the latent vectors of the same category are forced to be expressed in a more compact and distinguishable form, which is beneficial to improve the accuracy of the encoder for classification tasks. The potential space of VAE and VQVAE will enable the encoder to capture the inherent uncertainty and variability in the data, so as to obtain a more robust, more expressive, and more generalized feature representation.</p>
<p>Specifically, in the pre-training stage, the input skeleton sequences are randomly masked in temporal and spacial dimensions, and then, the unmasked data are input into the network for the reconstruction of the masked part. Finally, the decoder is removed in the fine-tuning stage, and a simple output layer is added after the encoder to predict the skeleton data. In the experiment stage, we discuss the effects of masking rate, latent variable dimension, decoder dimension, and decoder depth on the recognition task and found the best combination. Experiment results show that our method is generalized and robust and effectively improves the accuracy of classification in downstream classification tasks.</p>
<p>In general, we have made the following contributions:</p>
<list list-type="simple">
<list-item><p>1) To improve the feature extraction ability of the encoder after the masked reconstruction task, we propose SkeletonMVAE and SkeletonMVQVAE, which insert the potential space of the variational autoencoder (VAE) and the vector quantization variational autoencoder (VQVAE) into SkeletonMAE, respectively. We discuss the differences between them.</p></list-item>
<list-item><p>2) We compare several mainstream models on the dataset NTU-60 and NTU-120. Experiments show that our models can effectively improve the accuracy of downstream classification tasks. SkeletonMVAE has obvious advantages.</p></list-item>
<list-item><p>3) We prove that our models still have good robustness and generalization ability under extremely few label data. SkeletonMVQVAE has more advantages in the case of fewer data labels.</p></list-item>
</list>
</sec>
<sec id="s2">
<title>2 Related work</title>
<sec>
<title>2.1 Contrastive learning</title>
<p>In the self-supervised learning, most methods use contrastive learning (Zhang et al., <xref ref-type="bibr" rid="B41">2022</xref>; Chen et al., <xref ref-type="bibr" rid="B5">2022</xref>) that aims to enable the model to differentiate between various inputs in the feature space, distinguishing between similarities and dissimilarities. Research in this area typically involves creating positive and negative pairs through data augmentation, extracting representations via an encoder, and computing the similarity between two samples. Positive samples exhibit high similarity, while negative samples demonstrate low similarity. Previous comparative learning used normal enhancement to construct similar positive sample pairs. AimCLR (Guo et al., <xref ref-type="bibr" rid="B13">2022</xref>) employs extreme data augmentation to obtain more diverse positive samples. CrosSCLR (Li et al., <xref ref-type="bibr" rid="B19">2021</xref>) dugs positive sample pairs from similar negative samples, uses multi-view mining positive samples to learn cross-view consistency, and extracts more comprehensive cross-view features. SkeAttnCLR (Hua et al., <xref ref-type="bibr" rid="B17">2023</xref>) focuses on the fact that human actions are often related to local body parts. Therefore, local salient features and non-salient features were proposed, and a large number of contrast pairs were generated to guide the model to learn the action representation of the human skeleton. The above contrastive learning usually has problems such as the need for a large number of comparison pairs and the lack of correlation between frames.</p>
</sec>
<sec>
<title>2.2 Masked encoding</title>
<p>Other self-supervised works such as BERT (Devlin et al., <xref ref-type="bibr" rid="B7">2018</xref>), MAE (He et al., <xref ref-type="bibr" rid="B14">2022</xref>), and SkeletonMAE (Wu et al., <xref ref-type="bibr" rid="B36">2023</xref>) leverage masked reconstruction as a pretext task, which well enhance the learning of contextual relationships in data time and space. In natural language processing, the famous model BERT (Devlin et al., <xref ref-type="bibr" rid="B7">2018</xref>) masks tokens representing sequential data and then predicted the masked tokens. It calculates the loss between the predicts results and the original data to capture the features of language sequences. Following the idea of BERT, in the field of image processing, MAE (He et al., <xref ref-type="bibr" rid="B14">2022</xref>) adopts an asymmetric encoder&#x02014;decoder structure to mask image patches and reconstructed them at the pixel level. Inspired by MAE, VideoMAE (Tong et al., <xref ref-type="bibr" rid="B30">2022</xref>) applies masking encoding to the field of RGB video. Because of the redundancy of time, it can also bring good performance with a very high masking ratio. MAR (Qing et al., <xref ref-type="bibr" rid="B26">2023</xref>) proposes &#x0201C;cell running masking&#x0201D; on the basis of VideoMAE to encourage the leakage of spatio-temporal information, hoping to use the redundancy of spatio-temporal to provide a detailed context for the encoder to reconstruct the missing patch. In the field of skeleton action recognition, SkeletonMAE (Wu et al., <xref ref-type="bibr" rid="B36">2023</xref>) masks joints at the frame and joint levels, only encodes the unmasked joints and predicts the masked ones. This integration of masked reconstruction with self-supervised learning has shown promising results in various classification tasks and its potential to improve feature representation and classification performance.</p>
</sec>
<sec>
<title>2.3 The feature extraction of VAE and VQVAE</title>
<p>VAE and VQVAE have always been regarded as excellent generative models. Cheng et al. (<xref ref-type="bibr" rid="B6">2023</xref>) used VQVAE to generate coherent and structured fire scenarios, and the generated data were used for training and predicting wildfires. Zhu et al. (<xref ref-type="bibr" rid="B43">2023</xref>) proposed DSCVAE to generate consistent and realistic samples for predicting drop coalescence based on process parameters, improving prediction accuracy. With the development of deep learning, VAE and VQVAE have been used for feature extraction in many fields. Yue et al. (<xref ref-type="bibr" rid="B40">2023</xref>) uses the variational autoencoder to extract the feature invariance of EEG signals and then classifies them through a one-dimensional convolutional network. To extract the semantic features between words, Xu et al. (<xref ref-type="bibr" rid="B38">2023</xref>) use vaE to reconstruct the feature space so that it conforms to the normal distribution. This method can effectively extract text features for text classification. In the field of speech emotion recognition, TACN (Liu J. et al., <xref ref-type="bibr" rid="B22">2020</xref>) proposes to use VQVAE to model speech signals and learn the intrinsic expression of datasets. Hsu et al. (<xref ref-type="bibr" rid="B16">2022</xref>) used VQVAE as the feature extraction module of the pre-training model to extract the spectral features of prosodic phrases. In the field of anomaly detection, LSGS (Wang et al., <xref ref-type="bibr" rid="B33">2023</xref>) uses VQVAE to extract image features and locates anomalies by reconstructing a more accurate image. They introduced VAE and VQVAE as feature extractors to improve the performance of the model. Therefore, in the field of human skeleton-based action recognition, we introduce VAE and VQVAE into the self-supervised method of masking reconstruction. We hope to recover the masked data through its good generation ability to improve the feature extraction ability of the original encoder.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3 Methods</title>
<p>In this section, based on SkeletonMAE, we propose to improve the potential space of the masked reconstruction model. We explore two potential spatial patterns: one is the continuous potential space of VAE, and the other is the discrete potential space of VQVAE. We first review the characteristics of the two potential spaces and then introduce the network structure in detail.</p>
<sec>
<title>3.1 The potential space of VAE and loss function</title>
<p>Given the skeleton joint dataset <inline-formula><mml:math id="M1"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which contains <italic>N</italic> samples. We make <italic>z</italic> obey the standard normal distribution, and the probability distribution of the reconstructed sample <italic>x</italic> of the decoder is <inline-formula><mml:math id="M2"><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:math></inline-formula>, where <italic>P</italic>(<italic>z</italic>) is the probability of sampling the encoded <italic>z</italic> from the standard normal distribution, and <italic>P</italic>(<italic>x</italic>|<italic>z</italic>) is the probability of the output sample <italic>x</italic> of the decoder when the encoded <italic>z</italic> is input. By maximizing <inline-formula><mml:math id="M3"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, the reconstructed data are similar to the original data. However, not all <italic>z</italic> is meaningful, so <italic>p</italic>(<italic>z</italic>|<italic>x</italic>) is introduced to obtain the <italic>z</italic> corresponding to the input <italic>x</italic>. The posterior distribution <italic>p</italic>(<italic>z</italic>|<italic>x</italic>) is difficult to obtain, so we can use the encoder to fit the distribution <italic>q</italic>(<italic>z</italic>|<italic>x</italic>) of any <italic>x</italic>, then</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x02003;&#x02003;&#x02003;</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x02003;&#x02003;&#x02003;</mml:mtext><mml:mo>&#x02265;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The right half of the above equation is the Evidence Lower Bound (ELBO). We express it as <italic>L</italic><sub><italic>b</italic></sub> and hope that it is as large as possible.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;</mml:mtext><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>As <italic>L</italic><sub><italic>b</italic></sub> increases, <italic>KL</italic>(<italic>q</italic>(<italic>z</italic>|<italic>x</italic>)||<italic>P</italic>(<italic>z</italic>)) decreases, and <inline-formula><mml:math id="M6"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0222B;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:munder><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:math></inline-formula> increases. Since <italic>P</italic>(<italic>z</italic>) obeys the standard Gaussian distribution and <italic>q</italic>(<italic>z</italic>|<italic>x</italic>) obeys the Gaussian distribution, our SkeletonMVAE reconstruction loss can be written as</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo>&#x0002A;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x02003;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C3;<sub><italic>i</italic></sub> represents the variance of the i-th sample in the latent space, &#x003BC;<sub><italic>i</italic></sub> represents the mean in the latent space, and <inline-formula><mml:math id="M8"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> is the reconstructed sample. Adjusting the value of &#x003B2; can affect the model&#x00027;s emphasis on reconstruction loss and KL divergence loss during training. We hope that the model is more inclined to focus on retaining the details and structural information of the data and pay more attention to retaining the specific characteristics of the input data when generating the data, so the value of &#x003B2; is set to 0.005. We add the first part of the loss function as a regularization term. When <italic>z</italic> is known to obey the standard normal distribution, the first part constrains <italic>p</italic>(<italic>z</italic>|<italic>x</italic>), that is, the latent variable is close to the standard normal distribution, which helps the encoder to learn a more compact and discriminative data representation. Moreover, the latent variable in the form of distribution rather than the single value in SkeletonMAE can enhance the robustness of the model to noise and abnormal data and has better generalization, which can adapt to datasets with different distributions. The second part aims to minimizing the disparity between the reconstructed data and the original data.</p>
</sec>
<sec>
<title>3.2 The potential space of VQVAE and loss function</title>
<p>Unlike the usual MAE, VQVAE do not directly use <italic>z</italic> as the input of the decoder but map it to a discrete vector <italic>z</italic><sub><italic>q</italic></sub> according to a set of codebooks. Through vector quantization technology, the continuous feature space is mapped to the discrete potential space, which helps to learn more meaningful feature representation and improve the model&#x00027;s ability to represent data. Because codebook is discrete, even if the input data <italic>x</italic> change slightly, the quantized latent variable <italic>z</italic><sub><italic>q</italic></sub> will change greatly (jump to another discrete vector). It forces the encoder to extract key information from the input data <italic>x</italic> for meaningful mapping in the discrete space. This mandatory information compression mechanism encourages encoders to learn more meaningful latent variable representations. In addition, due to its discreteness, redundant information is removed, and data of the same category will have a more compact representation, and the anti-interference ability is also enhanced.</p>
<p>Specifically, for skeleton sequence data, the encoder outputs a continuous vector <italic>z</italic>&#x02208;&#x0211D;<sup><italic>N</italic>&#x000D7;<italic>D</italic>&#x000D7;<italic>T</italic>&#x02032; &#x000D7; <italic>V</italic>&#x02032;</sup>. The network learns a codebook <italic>E</italic> &#x0003D; <italic>e</italic><sub>1</sub>, <italic>e</italic><sub>2</sub>, <italic>e</italic><sub>3</sub>...<italic>e</italic><sub><italic>K</italic></sub> (<italic>E</italic>&#x02208;&#x0211D;<sup><italic>D</italic>&#x000D7;<italic>K</italic></sup>), <italic>e</italic> is the D-dimensional vector in the codebook, and <italic>K</italic> is the size of the codebook. VQVAE completes the mapping between the continuous vector <italic>z</italic> and the codebook <italic>E</italic> through the nearest neighbor search.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">arg</mml:mo><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>j</italic> is the index of the codebook vector closest to <italic>z</italic><sub><italic>i</italic></sub>.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The continuous vector <italic>z</italic> is mapped to the discrete vector <inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msup><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x000D7;</mml:mo><mml:msup><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>The reconstruction loss is defined as</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M12"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>v</mml:mi><mml:mi>q</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo>&#x000A0;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x0007C;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>[</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy='false'>]</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02016;</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:msubsup><mml:mrow><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mrow><mml:mi>z</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>s</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x02016;</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:math></disp-formula>
<p>Among them, the first part is the reconstruction loss, which optimizes the encoder and decoder by reducing the error of the original sequence and the reconstructed sequence. The second part faces challenges due to the argmin operation on the feature vector during mapping, preventing gradient calculation. To train the latent space codebook, the L2 error between the encoder&#x00027;s output <italic>z</italic> and the latent space <italic>e</italic> is computed, with sg representing the stop gradient operation. In the third part, the L2 error between the encoder&#x00027;s output <italic>z</italic> and the corresponding potential space <italic>e</italic> is also calculated, but sg is applied to <italic>e</italic> to ensure the encoder&#x00027;s output aligns with the embedding space and avoids drastic changes (switching from one embedding vector to another). The &#x003B2; is the weight coefficient, and we set it to 0.25.</p>
</sec>
<sec>
<title>3.3 Model architecture</title>
<p>We propose to insert the potential space of VAE and VQVAE into SkeletonMAE to improve the feature extraction ability of the encoder. The model structure is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The same thing of the SkeletonMVAE and SkeletonMVQVAE is that both have encoder and decoder. The encoder is employed to extract the feature representation of the unmasked data, while the decoder reconstructs the masked data based on the latent variables obtained during encoding. The difference is that SkeletonMVAE adds the potential spatial structure of VAE after the encoder, while SkeletonMVQVAE adds the potential spatial structure of VQVAE. The potential spatial structure of VAE and VQVAE is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Masked reconstruction model structure.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-19-1482281-g0001.tif"/>
</fig>
<sec>
<title>3.3.1 Spatial-temporal masking strategy</title>
<p>Because of the randomness of data loss, we perform random masking in both temporal and spatial dimensions when we mask skeleton data. Given the skeleton sequence <italic>S</italic>&#x02208;&#x0211D;<sup><italic>N</italic>&#x000D7;<italic>C</italic>&#x000D7;<italic>T</italic>&#x000D7;<italic>J</italic></sup>. First, in the temporal dimensions, some frames are masked (i.e., deleted) according to the given frame masking rate <italic>M</italic><sub><italic>t</italic></sub>, and the masked skeleton sequence becomes <inline-formula><mml:math id="M13"><mml:mi>S</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>J</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Then, in the spatial dimensions, the joints on all frames are randomly masked according to the given joint masking rate <italic>M</italic><sub><italic>j</italic></sub>. Finally, the skeleton sequence input into the network is <inline-formula><mml:math id="M14"><mml:mi>S</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>T</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>J</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The above masking process is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. Gray represents the masked skeleton joints, and blue-green represents the unmasked skeleton joints.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Spatial-temporal random masking process.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-19-1482281-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.3.2 Encoder</title>
<p>As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, our model applies multiple STTFormer blocks to capture the relationships between different keypoints in consecutive frames, which is used to encode and represent the unmasked parts of the skeleton sequence. To preserve the location information, we introduce location embedding after masking.</p>
</sec>
<sec>
<title>3.3.3 SkeletonMVAE potential space</title>
<p>After the data pass through the encoder, the mean value &#x003BC; and the standard deviation &#x003C3; are output through the fully connected layer. The reparameterization technique is used to sample from the latent space to obtain the latent variables <italic>z</italic> &#x0003D; &#x003BC;&#x0002B;&#x003C3;&#x0002A;&#x003B5; (<inline-formula><mml:math id="M15"><mml:mi>&#x003B5;</mml:mi><mml:mo>&#x0007E;</mml:mo><mml:mrow><mml:mi mathvariant="script">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>). This process is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. As mentioned in Section 3.1, KL divergence, as a regularization term, limits the potential variables to approach the standard normal distribution. Regular potential space means that the Gaussian distribution parameters of the same category are basically the same after mapping to the potential space, and the adjacent points in the potential space are the same category after decoding. The encoder can learn the more compact and discriminative data representation, which is more conducive to the classification task.</p>
</sec>
<sec>
<title>3.3.4 SkeletonMVQVAE potential space</title>
<p>Different from SkeletonMVAE, in SkeletonMVQVAE, the data pass through the latent space of VQVAE after the encoder. As described in Section B, in the latent space of VQVAE, the feature vector <italic>z</italic> is mapped to the discrete latent vector <italic>z</italic><sub><italic>q</italic></sub> according to the discrete codebook. This process is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The discrete potential space makes the encoder more able to extract representative features and enhance its robustness to noise.</p>
</sec>
<sec>
<title>3.3.5 Decoder</title>
<p>As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, our model decoder also consists of multiple STTFormer blocks. Feature vectors are sampled from the latent space formed by the encoder. We complement the learnable token for the missing skeleton sequence. The decoder decodes the masked token based on the sampled feature vectors and position information. The reconstruction goal is to be consistent with the original skeleton sequences.</p>
</sec>
</sec>
<sec>
<title>3.4 Pre-train</title>
<p>The pre-training process of our model is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. We randomly mask skeleton joints at both temporal and spatial dimensions and then add the positional embedding to the skeleton sequences. The unmasked skeleton data are fed into the encoder, mapped to the latent space. The sampled unmasked skeleton data, along with the mask token, are input into the decoder for reconstruction. When VAE potential space is used, we use the reconstruction loss of <xref ref-type="disp-formula" rid="E3">Equation 3</xref> from Section A as the pre-training loss function. When VQVAE potential space is used, We use the reconstruction loss of <xref ref-type="disp-formula" rid="E6">Equation 6</xref> from Section B as the pre-training loss function. The learning of the feature representation is continuously improved by minimizing the disparity between the original data and the reconstructed data. We save the model with the minimum verification loss as the best model.</p>
</sec>
<sec>
<title>3.5 Fine-tune</title>
<p>To evaluate the representation learning ability of our model, we utilize only the encoder part of the pre-trained model and add a fully connected layer for classification. We load the pre-trained parameter weights onto all training data and perform end-to-end fine-tuning for downstream recognition tasks. Throughout the fine-tuning process, the cross-entropy loss is utilized as the loss function and save the model with the maxmum verification accuracy as the best model.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experiments</title>
<sec>
<title>4.1 Datasets</title>
<sec>
<title>4.1.1 NTU-RGB&#x0002B;D 60</title>
<p>The NTU RGB&#x0002B;D 60 dataset (Shahroudy et al., <xref ref-type="bibr" rid="B27">2016</xref>) contains 60 action classes with a total of 56,578 action sequences. Among them, there are 40 kinds of daily behavior actions, 9 kinds of health-related actions, and 11 kinds of mutual actions between two people. The dataset is partitioned into training and testing sets using two criteria. The first one is Cross-Subject, which divides the dataset into training set and test set based on different subject IDs. The training set and test set are completed by 20 different subjects, respectively, which are used to evaluate the performance of the model under different subjects of the same action. The second is Cross-View. The three cameras that capture the video are at the same height and different angles. The data of camera 1 are used in the test phase, and the data of camera 2 and camera 3 are used in the training phase, which can be used to evaluate whether the model can perform action recognition for skeletons at different angles.</p>
</sec>
<sec>
<title>4.1.2 NTU-RGB&#x0002B;D 120</title>
<p>The NTU RGB &#x0002B; D 120 dataset adds 60 action categories based on NTU RGB&#x0002B;D 60 and adds 32 settings. Each setting uses different camera heights and different distances from the subject. The dataset is also divided using two criteria. The rule of Cross-Subject is consistent with NTU RGB&#x0002B;D 60. According to the ID of the subject, 53 people are divided into the training set and the other 53 people are divided into the test set. The model considers both different subjects and different settings during the training and testing process, which is used to evaluate the generalization ability of the model in the real world. In addition, it also adopts another partitioning strategy: Cross-Setup, which divides the training set and the test set according to the IDs of 32 settings. The even ID is classified as the training set, and the odd ID is classified as the test set, which is used to evaluate the adaptability of the model under different perspectives on the setting of the same subject.</p>
</sec>
</sec>
<sec>
<title>4.2 Experiment settings</title>
<p>Our experiments are implemented under the framework of Pytorch (Paszke et al., <xref ref-type="bibr" rid="B25">2019</xref>), using a computing node on the supercomputing platform and four HYGON DCUs under one computing node. Both the pre-training model and fine-tuning model utilize the Adam optimizer. When VAE potential space is used, the base learning rate is set to 0.001. When VQVAE potential space is used, the base learning rate is set to 0.01. We set the weight decay to 0.0001. The pre-training epoch number is 200, and the fine-tuning epoch number is 200. Batch size is set to 64. We employ a step-wise learning rate strategy, adjusting the learning rate to one-tenth at epochs 60, 90, and 110.</p>
</sec>
<sec>
<title>4.3 Comparison with existing mainstream methods</title>
<p>The comparison of our SkeletonMVAE, SkeletonMVQVAE, and other mainstream models on the NTU-60 and NTU-120 datasets is shown in <xref ref-type="table" rid="T1">Table 1</xref>. On the NTU-60 dataset, our SkeletonMVAE achieved a 1.8% higher accuracy than SkeletonMAE under the X-sub protocol and a 0.2% higher accuracy under the X-view protocol, and our SkeletonMVQVAE achieved a 1.4% higher accuracy than SkeletonMAE under the X-sub protocol and a 0.4% higher accuracy under the X-view protocol. On the NTU-120 dataset, our SkeletonMVAE achieved a 3.8% higher accuracy than SkeletonMAE under the X-sub protocol and a 4.4% higher accuracy under the X-set protocol, and our SkeletonMVQVAE achieved a 3.2% higher accuracy than SkeletonMAE under the X-sub protocol and a 2.3% higher accuracy under the X-set protocol. We can see that our SkeletonMVAE fine-tuning results is not only outperform other classical methods on small datasets but also have the potential to perform even better on larger datasets. The accuracy of our SkeletonMVQVAE on NTU-60 and NTU-120 datasets is also improved to varying degrees compared to SkeletonMAE. In this way, in terms of improving accuracy, SkeletonMVAE is slightly better, and the potential space of VAE-style regularization is more conducive to the realization of classification tasks. Experimental results show the effectiveness of our proposed method.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Fine-tuned results on NTU-60 and NTU-120 datasets.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th/>
<th valign="top" align="center" colspan="2"><bold>NTU-60 (%)</bold></th>
<th valign="top" align="center" colspan="2"><bold>NTU-120 (%)</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Method</bold></td>
<td valign="top" align="center"><bold>Backbone</bold></td>
<td valign="top" align="center"><bold>X-sub</bold></td>
<td valign="top" align="center"><bold>X-view</bold></td>
<td valign="top" align="center"><bold>X-sub</bold></td>
<td valign="top" align="center"><bold>X-set</bold></td>
</tr>
<tr>
<td valign="top" align="left">SkeletonCLR (Hua et al., <xref ref-type="bibr" rid="B17">2023</xref>)</td>
<td valign="top" align="center">ST-GCN</td>
<td valign="top" align="center">82.2</td>
<td valign="top" align="center">88.9</td>
<td valign="top" align="center">73.6</td>
<td valign="top" align="center">75.3</td>
</tr>
<tr>
<td valign="top" align="left">CPM (Zhang et al., <xref ref-type="bibr" rid="B41">2022</xref>)</td>
<td valign="top" align="center">ST-GCN</td>
<td valign="top" align="center">84.8</td>
<td valign="top" align="center">91.1</td>
<td valign="top" align="center">78.4</td>
<td valign="top" align="center">78.9</td>
</tr>
<tr>
<td valign="top" align="left">CrosSCLR (Li et al., <xref ref-type="bibr" rid="B19">2021</xref>)</td>
<td valign="top" align="center">ST-GCN</td>
<td valign="top" align="center">86.2</td>
<td valign="top" align="center">92.5</td>
<td valign="top" align="center">80.5</td>
<td valign="top" align="center">80.4</td>
</tr>
<tr>
<td valign="top" align="left">AimCLR (Guo et al., <xref ref-type="bibr" rid="B13">2022</xref>)</td>
<td valign="top" align="center">ST-GCN</td>
<td valign="top" align="center">86.9</td>
<td valign="top" align="center">92.8</td>
<td valign="top" align="center">80.1</td>
<td valign="top" align="center">80.9</td>
</tr>
<tr>
<td valign="top" align="left">Hi-TRS (Chen et al., <xref ref-type="bibr" rid="B5">2022</xref>)</td>
<td valign="top" align="center">Transformer</td>
<td valign="top" align="center">86.0</td>
<td valign="top" align="center">93.0</td>
<td valign="top" align="center">80.6</td>
<td valign="top" align="center">81.6</td>
</tr>
<tr>
<td valign="top" align="left">AimCLR (Guo et al., <xref ref-type="bibr" rid="B13">2022</xref>)</td>
<td valign="top" align="center">STTFormer</td>
<td valign="top" align="center">83.9</td>
<td valign="top" align="center">90.4</td>
<td valign="top" align="center">74.6</td>
<td valign="top" align="center">77.2</td>
</tr>
<tr>
<td valign="top" align="left">CrosSCLR (Li et al., <xref ref-type="bibr" rid="B19">2021</xref>)</td>
<td valign="top" align="center">STTFormer</td>
<td valign="top" align="center">84.6</td>
<td valign="top" align="center">90.5</td>
<td valign="top" align="center">75.0</td>
<td valign="top" align="center">77.9</td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMAE (Wu et al., <xref ref-type="bibr" rid="B36">2023</xref>)</td>
<td valign="top" align="center">STTFormer</td>
<td valign="top" align="center">86.6</td>
<td valign="top" align="center">92.9</td>
<td valign="top" align="center">76.8</td>
<td valign="top" align="center">79.1</td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMVAE</td>
<td valign="top" align="center">STTFormer</td>
<td valign="top" align="center"><bold>88.4</bold></td>
<td valign="top" align="center">93.1</td>
<td valign="top" align="center"><bold>80.6</bold></td>
<td valign="top" align="center"><bold>83.5</bold></td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMVQVAE</td>
<td valign="top" align="center">STTFormer</td>
<td valign="top" align="center">88.0</td>
<td valign="top" align="center"><bold>93.3</bold></td>
<td valign="top" align="center">80.0</td>
<td valign="top" align="center">81.4</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values represent optimal values.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.4 Semi-supervised results</title>
<p>We randomly sample 5% and 10% of data from the training set for semi-supervised fine-tuning. The sampling rule is to sample the same proportion of data within each class. The semi-supervised results are presented in <xref ref-type="table" rid="T2">Table 2</xref>. The results show that the performance of our proposed method is significantly better than the compared methods. We compare our SkeletonMVAE and SkeletonMVQVAE with the SkeletonMAE on NTU-60 and NTU-120 datasets. SkeletonMVAE has different degrees of improvement (0.5% to 3.5%) than SkeletonMAE. SkeletonMVQVAE also has different degrees of improvement (0.4% to 5.3%) than SkeletonMAE. It can be seen that when the sampling ratio is 5%, SkeletonMVQVAE shows a greater advantage. When the sampling ratio is 10%, SkeletonMVAE and SkeletonMVQVAE perform basically the same. It can be concluded that the potential space of SkeletonMVQVAE discretization is more conducive to generalization in the case of less labeled data. The experimental results indicate that our SkeletonMVAE and SkeletonMVQVAE still exhibit generalization ability with a small amount of labeled data and the generalization is improved compared to SkeletonMAE.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Fine-tuned result comparison on the NTU-60 and NTU-120 datasets with fewer labeled data.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center" colspan="4"><bold>NTU-60 (%)</bold></th>
<th valign="top" align="center" colspan="4"><bold>NTU-120 (%)</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td valign="top" align="left"><bold>Method</bold></td>
<td valign="top" align="center" colspan="2"><bold>X-sub</bold></td>
<td valign="top" align="center" colspan="2"><bold>X-view</bold></td>
<td valign="top" align="center" colspan="2"><bold>X-sub</bold></td>
<td valign="top" align="center" colspan="2"><bold>X-set</bold></td>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="center"><bold>5%</bold></td>
<td valign="top" align="center"><bold>10%</bold></td>
<td valign="top" align="center"><bold>5%</bold></td>
<td valign="top" align="center"><bold>10%</bold></td>
<td valign="top" align="center"><bold>5%</bold></td>
<td valign="top" align="center"><bold>10%</bold></td>
<td valign="top" align="center"><bold>5%</bold></td>
<td valign="top" align="center"><bold>10%</bold></td>
</tr>
<tr>
<td valign="top" align="left">Hi-TRS (Chen et al., <xref ref-type="bibr" rid="B5">2022</xref>)</td>
<td valign="top" align="center">63.3</td>
<td valign="top" align="center">70.7</td>
<td valign="top" align="center">68.3</td>
<td valign="top" align="center">74.8</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">CrosSCLR (Li et al., <xref ref-type="bibr" rid="B19">2021</xref>)</td>
<td valign="top" align="center">63.5</td>
<td valign="top" align="center">71.0</td>
<td valign="top" align="center">66.9</td>
<td valign="top" align="center">75.1</td>
<td valign="top" align="center">50.2</td>
<td valign="top" align="center">58.5</td>
<td valign="top" align="center">50.4</td>
<td valign="top" align="center">60.6</td>
</tr>
<tr>
<td valign="top" align="left">AimCLR (Guo et al., <xref ref-type="bibr" rid="B13">2022</xref>)</td>
<td valign="top" align="center">63.9</td>
<td valign="top" align="center">70.2</td>
<td valign="top" align="center">67.5</td>
<td valign="top" align="center">76.2</td>
<td valign="top" align="center">49.0</td>
<td valign="top" align="center">58.6</td>
<td valign="top" align="center">51.8</td>
<td valign="top" align="center">60.5</td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMAE (Wu et al., <xref ref-type="bibr" rid="B36">2023</xref>)</td>
<td valign="top" align="center">64.4</td>
<td valign="top" align="center">73.0</td>
<td valign="top" align="center">68.8</td>
<td valign="top" align="center">76.9</td>
<td valign="top" align="center">50.4</td>
<td valign="top" align="center">61.8</td>
<td valign="top" align="center">52.0</td>
<td valign="top" align="center">62.5</td>
</tr>
<tr>
<td valign="top" align="left">CPM (Zhang et al., <xref ref-type="bibr" rid="B41">2022</xref>)</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">73.0</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">77.1</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMVAE</td>
<td valign="top" align="center">65.1</td>
<td valign="top" align="center"><bold>73.7</bold></td>
<td valign="top" align="center">67.1</td>
<td valign="top" align="center"><bold>78.1</bold></td>
<td valign="top" align="center">53.9</td>
<td valign="top" align="center"><bold>62.7</bold></td>
<td valign="top" align="center">53.0</td>
<td valign="top" align="center"><bold>64.6</bold></td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMVQVAE</td>
<td valign="top" align="center"><bold>66.2</bold></td>
<td valign="top" align="center">73.5</td>
<td valign="top" align="center"><bold>69.6</bold></td>
<td valign="top" align="center">77.9</td>
<td valign="top" align="center"><bold>55.7</bold></td>
<td valign="top" align="center">62.4</td>
<td valign="top" align="center"><bold>56.3</bold></td>
<td valign="top" align="center">64.3</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>The bold values represent optimal values.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.5 Ablation study</title>
<p>All the experiments in this section are carried out on the NTU-60 dataset, and more detail of the model we proposed is analyzed.</p>
<sec>
<title>4.5.1 Frame and joint masking ratio</title>
<p>According to practical experience, the missing information in actual data is typically random. Therefore, we adopted a random method to mask the joints in both temporal and spatial dimensions. In temporal dimension, frames are masked with probabilities of 0.4, 0.5, and 0.6, while in spatial dimension, joints are masked with probabilities of 0.4, 0.6, and 0.8. As shown in <xref ref-type="table" rid="T3">Table 3</xref>, under the X-sub partition standard of NTU-60, when the frame mask rate is 0.4 and the joint mask rate is 0.4, SkeletonMVAE and SkeletonMVQVAE have the best performance.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Ablation study on frame and joint masking ratio.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="center"><bold>Frame masking ratio</bold></th>
<th valign="top" align="center"><bold>Joint masking ratio</bold></th>
<th valign="top" align="center"><bold>SkeletonMVAE (%)</bold></th>
<th valign="top" align="center"><bold>SkeletonMVQVAE (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td valign="top" align="center">0.4</td>
<td valign="top" align="center">88.4</td>
<td valign="top" align="center">88.0</td>
</tr>
<tr>
<td valign="top" align="left">0.4</td>
<td valign="top" align="center">0.6</td>
<td valign="top" align="center">87.0</td>
<td valign="top" align="center">86.2</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">0.8</td>
<td valign="top" align="center">87.8</td>
<td valign="top" align="center">86.8</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">0.4</td>
<td valign="top" align="center">87.5</td>
<td valign="top" align="center">86.5</td>
</tr>
<tr>
<td valign="top" align="left">0.5</td>
<td valign="top" align="center">0.6</td>
<td valign="top" align="center">87.5</td>
<td valign="top" align="center">87.8</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">0.8</td>
<td valign="top" align="center">87.4</td>
<td valign="top" align="center">87.7</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">0.4</td>
<td valign="top" align="center">87.4</td>
<td valign="top" align="center">86.1</td>
</tr>
<tr>
<td valign="top" align="left">0.6</td>
<td valign="top" align="center">0.6</td>
<td valign="top" align="center">87.5</td>
<td valign="top" align="center">87.3</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">0.8</td>
<td valign="top" align="center">87.4</td>
<td valign="top" align="center">87.1</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.5.2 Latent variable dimension</title>
<p>The dimension of the latent variable in the VAE determines the number of features that the model can learn and represent, which influence model&#x00027;s capacity to learn and represent features and the quality of the generated data. We conduct experiments with different latent variable dimensions, and the results are presented in <xref ref-type="table" rid="T4">Table 4</xref>. It shows that the model performs best when the latent variable dimension is 25. The lower latent variable dimension can lead to information loss, while the higher latent variable dimension can increase the model&#x00027;s complexity, requiring more training data and time to achieve good performance. Considering the model&#x00027;s performance and available resources, we chose the latent variable dimension of 25.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Ablation study on SkeletonMVAE latent variable dimension.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="center"><bold>Latent variable dimension</bold></th>
<th valign="top" align="center"><bold>SkeletonMVAE (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">15</td>
<td valign="top" align="center">87.7</td>
</tr>
<tr>
<td valign="top" align="left">25</td>
<td valign="top" align="center">88.4</td>
</tr>
<tr>
<td valign="top" align="left">35</td>
<td valign="top" align="center">87.5</td>
</tr>
<tr>
<td valign="top" align="left">45</td>
<td valign="top" align="center">87.6</td>
</tr>
<tr>
<td valign="top" align="left">55</td>
<td valign="top" align="center">86.9</td>
</tr>
<tr>
<td valign="top" align="left">65</td>
<td valign="top" align="center">87.6</td>
</tr></tbody>
</table>
</table-wrap>
<p>We explore the impact of the codebook size of the latent space (i.e., the compactness of the latent vector) on the classification task. The influence of codebook size K on classification task is shown in <xref ref-type="table" rid="T5">Table 5</xref>. It is found that too large or too small K is not conducive to classification, because too large K has no way to learn compact representation, and too small K will lose information. When K is 128, the classification accuracy is 88.0%, which is the best result. The different latent space makes SkeletonMVQVAE greatly reduce the number of parameters of the model compared with SkeletonMVAE.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Ablation study on SkeletonMVQVAE latent variable dimension.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="center"><bold>Latent variable dimension</bold></th>
<th valign="top" align="center"><bold>SkeletonMVQVAE (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">100</td>
<td valign="top" align="center">86.8</td>
</tr>
<tr>
<td valign="top" align="left">128</td>
<td valign="top" align="center">88.0</td>
</tr>
<tr>
<td valign="top" align="left">256</td>
<td valign="top" align="center">86.3</td>
</tr>
<tr>
<td valign="top" align="left">512</td>
<td valign="top" align="center">87.7</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.5.3 Decoder embedding dimension</title>
<p>We perform ablation experiments on the decoder&#x00027;s embedding dimension, evaluating the model&#x00027;s performance across three different dimensions: 128, 256, and 512. The experimental results are shown in <xref ref-type="table" rid="T6">Table 6</xref>. The SkeletonMVAE and SkeletonMVQVAE achieve the best accuracy when the decoder embedding dimension is set to 256.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Ablation study on decoder embedding dimension.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="center"><bold>Dimension</bold></th>
<th valign="top" align="center"><bold>SkeletonMVAE (%)</bold></th>
<th valign="top" align="center"><bold>SkeletonMVQVAE (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">128</td>
<td valign="top" align="center">86.1</td>
<td valign="top" align="center">86.7</td>
</tr>
<tr>
<td valign="top" align="left">256</td>
<td valign="top" align="center">88.4</td>
<td valign="top" align="center">87.7</td>
</tr>
<tr>
<td valign="top" align="left">512</td>
<td valign="top" align="center">86.6</td>
<td valign="top" align="center">86.4</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.5.4 Decoder depth</title>
<p>We further conduct the ablation study on different decoder depths (i.e., the number of STTFormer blocks used and a full connection layer). The decoder depths are set to 5, 7, 9, and 11 layers. The results are shown in <xref ref-type="table" rid="T7">Table 7</xref>. Too deep or shallow depths both reduce the fine-tuning accuracy. Considering the fine-tuning accuracy and model parameters, we set the decoder depth of SkeletonMVAE to 9 layers, and the decoder depth of SkeletonMVQVAE to 7 layers.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Ablation study on decoder depth.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="center"><bold>Decoder depth</bold></th>
<th valign="top" align="center"><bold>SkeletonMVAE (%)</bold></th>
<th valign="top" align="center"><bold>SkeletonMVQVAE (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">87.9</td>
<td valign="top" align="center">87.4</td>
</tr>
<tr>
<td valign="top" align="left">7</td>
<td valign="top" align="center">87.5</td>
<td valign="top" align="center">88.0</td>
</tr>
<tr>
<td valign="top" align="left">9</td>
<td valign="top" align="center">88.4</td>
<td valign="top" align="center">87.7</td>
</tr>
<tr>
<td valign="top" align="left">11</td>
<td valign="top" align="center">87.9</td>
<td valign="top" align="center">87.1</td>
</tr></tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="T8">Table 8</xref>, compared with SkeletonMAE, the parameters of SkeletonMVAE increased by 14M, the FLOPs increased by 2.1G, and the training time increased by 3 h. The parameters of SkeletonMVQVAE decreased by 1M, the FLOPs decreased by 6.5G, and the training time decreased by 1 h. SkeletonMVQVAE achieves the best performance when the decoder depth is 7, and the training efficiency is also improved compared with SkeletonMVAE.</p>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>Number of model parameters, computational complexity, and pre-training time for SkeletonMVAE and SkeletonMVQVAE.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="center"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Parameter (M)</bold></th>
<th valign="top" align="center"><bold>FlOPs (G)</bold></th>
<th valign="top" align="center"><bold>Pre-training time (hours)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SkeletonMAE</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">42.8</td>
<td valign="top" align="center">39</td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMVAE</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">44.9</td>
<td valign="top" align="center">42</td>
</tr>
<tr>
<td valign="top" align="left">SkeletonMVQVAE</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">36.3</td>
<td valign="top" align="center">38</td>
</tr></tbody>
</table>
</table-wrap>
<p>We summarize the Skeletonmvae network structure as shown in <xref ref-type="table" rid="T9">Table 9</xref>. The dim(input) is the output dim of block8 &#x000D7; (1&#x02212;<italic>M</italic><sub><italic>t</italic></sub>) &#x000D7; <italic>T</italic>&#x000D7;(1&#x02212;<italic>M</italic><sub><italic>j</italic></sub>) &#x000D7; <italic>J</italic>. The dim(output) is <italic>C</italic>&#x000D7;(1&#x02212;<italic>M</italic><sub><italic>t</italic></sub>) &#x000D7; <italic>T</italic>&#x000D7;(1&#x02212;<italic>M</italic><sub><italic>j</italic></sub>) &#x000D7; <italic>J</italic>.</p>
<table-wrap position="float" id="T9">
<label>Table 9</label>
<caption><p>Structure of SkeletonMVAE.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>Layer name</bold></th>
<th valign="top" align="center"><bold>Input dim</bold></th>
<th valign="top" align="center"><bold>Output dim</bold></th>
<th valign="top" align="center"><bold>QKV dim</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" rowspan="9">Encoder</td>
<td valign="top" align="center">Input layer</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">64</td>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block1</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block2</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block3</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block4</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block5</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block6</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block7</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block8</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">vae(input)</td>
<td valign="top" align="center">Dim(input)</td>
<td valign="top" align="center">2&#x0002A;25</td>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="center">vae(output)</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">dim(output)</td>
<td/>
</tr>
<tr>
<td valign="top" align="left" rowspan="9">Decoder</td>
<td valign="top" align="center">Block1</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block2</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block3</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block4</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block5</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block6</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block7</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Block8</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">Output layer</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">3</td>
<td/>
</tr></tbody>
</table>
</table-wrap>
<p>In addition, we also list the model structure of SkeletonMVQVAE, as shown in <xref ref-type="table" rid="T10">Table 10</xref>. SkeletonMVQVAE shows better results on fewer decoder layers. We believe that it is because too strong decoder is not conducive to the encoder to extract features. In this way, SkeletonMVQVAE further reduces the parameters of the model.</p>
<table-wrap position="float" id="T10">
<label>Table 10</label>
<caption><p>Structure of SkeletonMVQVAE.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>Layer name</bold></th>
<th valign="top" align="center"><bold>Input dim</bold></th>
<th valign="top" align="center"><bold>Output dim</bold></th>
<th valign="top" align="center"><bold>QKV dim</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" rowspan="9">Encoder</td>
<td valign="top" align="center">Input layer</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">64</td>
<td/>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block1</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block2</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block3</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block4</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block5</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block6</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block7</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block8</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
<tr>
<td valign="top" align="left" rowspan="7">Decoder</td>
<td valign="top" align="center">block1</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block2</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block3</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">64</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block4</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block5</td>
<td valign="top" align="center">128</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">32</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">block6</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">16</td>
</tr>
 <tr>
<td/>
<td valign="top" align="center">output layer</td>
<td valign="top" align="center">64</td>
<td valign="top" align="center">3</td>
<td/>
</tr></tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5 Conclusion</title>
<p>The masked reconstruction model aims to improve the accuracy of the encoder in the downstream classification task by pre-training the reconstruction of the masked skeleton joints. The traditional masked reconstruction model uses the autoencoder structure but cannot learn richer potential information and data structure. To this end, we propose to improve the latent space based on the SkeletonMAE model and explore two different latent spaces. One is the VAE normal distribution regularization space, and the other is the VQVAE discrete latent space. We also performed pre-training of masked reconstruction and fine-tuning of downstream classification tasks on them and discussed the influence of two different latent spaces on downstream classification tasks, as well as the generalization ability of two different latent spaces. The experimental results show that the use of different latent spaces in pre-training can significantly improve the performance of downstream classification tasks in human skeleton-based action recognition. This shows that the choice of potential space plays a vital role in improving the overall effectiveness of the SkeletonMAE model. The shortcoming of this model is that the current focus is still on the field of action recognition, and the problem of cross-domain generalization needs further research to make it more conducive to practical application.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>EC: Writing &#x02013; review &#x00026; editing. XW: Software, Writing &#x02013; original draft. XG: Writing &#x02013; review &#x00026; editing. YZ: Writing &#x02013; review &#x00026; editing. DL: Writing &#x02013; review &#x00026; editing.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by the National Natural Science Foundation of China under Grant (62101503).</p>
</sec>
<ack><p>The author would like to express gratitude to the National Supercomputing Zhengzhou Center for providing computing resources. And we thank gpt-4o Mini for its touch up service.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Banerjee</surname> <given-names>A.</given-names></name> <name><surname>Singh</surname> <given-names>P. K.</given-names></name> <name><surname>Sarkar</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Fuzzy integral-based CNN classifier fusion for 3D skeleton action recognition</article-title>. <source>IEEE Trans. Circ. Syst. Video Technol</source>. <volume>31</volume>, <fpage>2206</fpage>&#x02013;<lpage>2216</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2020.3019293</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buch</surname> <given-names>S.</given-names></name> <name><surname>Escorcia</surname> <given-names>V.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Ghanem</surname> <given-names>B.</given-names></name> <name><surname>Carlos Niebles</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;SST: single-stream temporal action proposals,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, 2911&#x02013;2920. <pub-id pub-id-type="doi">10.1109/CVPR.2017.675</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Simon</surname> <given-names>T.</given-names></name> <name><surname>Wei</surname> <given-names>S.-E.</given-names></name> <name><surname>Sheikh</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). Realtime multi-person 2D pose estimation using part affinity fields,&#x0201D; in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, 7291&#x02013;7299. <pub-id pub-id-type="doi">10.1109/CVPR.2017.143</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Yuan</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Deng</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Channel-wise topology refinement graph convolution for skeleton-based action recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, 13359&#x02013;13368. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.01311</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name> <name><surname>Tian</surname> <given-names>Y.</given-names></name> <name><surname>Xia</surname> <given-names>Z.</given-names></name> <name><surname>Geng</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>&#x0201C;Hierarchically self-supervised transformer for human skeleton representation learning,&#x0201D;</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Springer</publisher-loc>), <fpage>185</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-031-19809-0_11</pub-id><pub-id pub-id-type="pmid">36904884</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>S.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name> <name><surname>Arcucci</surname> <given-names>R.</given-names></name></person-group> (<year>2023</year>). <article-title>A generative model for surrogates of spatial-temporal wildfire nowcasting</article-title>. <source>IEEE Trans. Emerg. Topics Comput. Intell</source>. <volume>7</volume>, <fpage>1420</fpage>&#x02013;<lpage>1430</lpage>. <pub-id pub-id-type="doi">10.1109/TETCI.2023.3298535</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Bert: pre-training of deep bidirectional transformers for language understanding</article-title>. <source>arXiv preprint arXiv:1810.04805</source>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Hierarchical contrast for unsupervised skeleton-based action representation learning,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, 525&#x02013;533. <pub-id pub-id-type="doi">10.1609/aaai.v37i1.25127</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Hierarchical recurrent neural network for skeleton based action recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, 1110&#x02013;1118. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298714</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name> <name><surname>Dai</surname> <given-names>B.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Revisiting skeleton-based action recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 2969&#x02013;2978. <pub-id pub-id-type="doi">10.1109/CVPR52688.2022.00298</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>H.-S.</given-names></name> <name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>Tai</surname> <given-names>Y.-W.</given-names></name> <name><surname>Lu</surname> <given-names>C.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;RMPE: regional multi-person pose estimation,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source>, 2334&#x02013;2343. <pub-id pub-id-type="doi">10.1109/ICCV.2017.256</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feichtenhofer</surname> <given-names>C.</given-names></name> <name><surname>Fan</surname> <given-names>H.</given-names></name> <name><surname>Malik</surname> <given-names>J.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Slowfast networks for video recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, 6202&#x02013;6211. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00630</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>T.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Ding</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Contrastive learning from extremely augmented skeleton sequences for self-supervised action recognition,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, 762&#x02013;770. <pub-id pub-id-type="doi">10.1609/aaai.v36i1.19957</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Doll&#x000E1;r</surname> <given-names>P.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Masked autoencoders are scalable vision learners,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 16000&#x02013;16009. <pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01553</pub-id><pub-id pub-id-type="pmid">38715952</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Skeleton optical spectra-based action recognition using convolutional neural networks</article-title>. <source>IEEE Trans. Circ. Syst. Video Technol</source>. <volume>28</volume>, <fpage>807</fpage>&#x02013;<lpage>811</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2016.2628339</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hsu</surname> <given-names>J.-H.</given-names></name> <name><surname>Wu</surname> <given-names>C.-H.</given-names></name> <name><surname>Yang</surname> <given-names>T.-H.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Using prosodic phrase-based VQVAE on audio albert for speech emotion recognition,&#x0201D;</article-title> in <source>2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>415</fpage>&#x02013;<lpage>419</lpage>. <pub-id pub-id-type="doi">10.23919/APSIPAASC55919.2022.9980239</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hua</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Zheng</surname> <given-names>C.</given-names></name> <name><surname>Lu</surname> <given-names>A.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Part aware contrastive learning for self-supervised action recognition</article-title>. <source>arXiv preprint arXiv:2305.00666</source>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Auto-encoding variational bayes</article-title>. <source>arXiv preprint arXiv:1312.6114</source>.<pub-id pub-id-type="pmid">32176273</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Ni</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;3D human action representation learning via cross-view consistency pursuit,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 4741&#x02013;4750. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00471</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Cook</surname> <given-names>C.</given-names></name> <name><surname>Gao</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Deep independently recurrent neural network (INDRNN)</article-title>. <source>arXiv preprint arXiv:1910.06251</source>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00572</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;ActionLet-dependent contrastive learning for unsupervised skeleton-based action recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 2363&#x02013;2372. <pub-id pub-id-type="doi">10.1109/CVPR52729.2023.00234</pub-id><pub-id pub-id-type="pmid">36149998</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Gao</surname> <given-names>Y.</given-names></name> <name><surname>Guo</surname> <given-names>L.</given-names></name> <name><surname>Dang</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Temporal attention convolutional network for speech emotion recognition with latent representation,&#x0201D;</article-title> in <source>INTERSPEECH</source>, 2337&#x02013;2341. <pub-id pub-id-type="doi">10.21437/Interspeech.2020-1520</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Ouyang</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Disentangling and unifying graph convolutions for skeleton-based action recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 143&#x02013;152. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00022</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>P.</given-names></name> <name><surname>Jiang</surname> <given-names>T.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Yang</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>RTMO: towards high-performance one-stage real-time multi-person pose estimation</article-title>. <source>arXiv preprint arXiv:2312.07526</source>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paszke</surname> <given-names>A.</given-names></name> <name><surname>Gross</surname> <given-names>S.</given-names></name> <name><surname>Massa</surname> <given-names>F.</given-names></name> <name><surname>Lerer</surname> <given-names>A.</given-names></name> <name><surname>Bradbury</surname> <given-names>J.</given-names></name> <name><surname>Chanan</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Pytorch: an imperative style, high-performance deep learning library,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, 32.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qing</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Huang</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Lv</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>MAR: Masked autoencoders for efficient action recognition</article-title>. <source>IEEE Trans. Multim</source>. <volume>26</volume>, <fpage>218</fpage>&#x02013;<lpage>233</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2023.3263288</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shahroudy</surname> <given-names>A.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Ng</surname> <given-names>T.-T.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;NTU RGB&#x0002B; D: a large scale dataset for 3D human activity analysis,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, 1010&#x02013;1019. <pub-id pub-id-type="doi">10.1109/CVPR.2016.115</pub-id><pub-id pub-id-type="pmid">31095476</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Two-stream convolutional networks for action recognition in videos,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, 27.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thoker</surname> <given-names>F. M.</given-names></name> <name><surname>Doughty</surname> <given-names>H.</given-names></name> <name><surname>Snoek</surname> <given-names>C. G.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Skeleton-contrastive 3D action representation learning,&#x0201D;</article-title> in <source>Proceedings of the 29th ACM International Conference on Multimedia</source>, 1655&#x02013;1663. <pub-id pub-id-type="doi">10.1145/3474085.3475307</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tong</surname> <given-names>Z.</given-names></name> <name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Videomae: masked autoencoders are data-efficient learners for self-supervised video pre-training,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, <fpage>10078</fpage>&#x02013;<lpage>10093</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Den Oord</surname> <given-names>A.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Neural discrete representation learning,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, 30.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varol</surname> <given-names>G.</given-names></name> <name><surname>Laptev</surname> <given-names>I.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name></person-group> (<year>2017</year>). <article-title>Long-term temporal convolutions for action recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>40</volume>, <fpage>1510</fpage>&#x02013;<lpage>1517</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2712608</pub-id><pub-id pub-id-type="pmid">28600238</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Luo</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Xia</surname> <given-names>S.-T.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>&#x0201C;Unsupervised anomaly detection with local-sensitive VQVAE and global-sensitive transformers,&#x0201D;</article-title> in <source>2023 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1080</fpage>&#x02013;<lpage>1084</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP49359.2023.10222596</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Hou</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Action recognition based on joint trajectory maps using convolutional neural networks,&#x0201D;</article-title> in <source>Proceedings of the 24th ACM International Conference on Multimedia</source>, 102&#x02013;106. <pub-id pub-id-type="doi">10.1145/2964284.2967191</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Peng</surname> <given-names>J.</given-names></name> <name><surname>Shi</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>Weng</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>IIP-transformer: Intra-inter-part transformer for skeleton-based action recognition</article-title>. <source>arXiv preprint arXiv:2110.13385</source>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Hua</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <name><surname>Lu</surname> <given-names>A.</given-names></name></person-group> (<year>2023</year>). <article-title>&#x0201C;Skeletonmae: spatial-temporal masked autoencoders for self-supervised skeleton action recognition,&#x0201D;</article-title> in <source>2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>224</fpage>&#x02013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1109/ICMEW59549.2023.00045</pub-id><pub-id pub-id-type="pmid">37856263</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Ni</surname> <given-names>B.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Deep kinematics analysis for monocular 3D human pose estimation,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, 899&#x02013;908. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00098</pub-id><pub-id pub-id-type="pmid">30587377</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>S.</given-names></name> <name><surname>Guo</surname> <given-names>C.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>G.</given-names></name> <name><surname>Xiong</surname> <given-names>N.</given-names></name></person-group> (<year>2023</year>). <article-title>CNN-VAE: an intelligent text representation algorithm</article-title>. <source>J. Supercomput</source>. <volume>79</volume>, <fpage>12266</fpage>&#x02013;<lpage>12291</lpage>. <pub-id pub-id-type="doi">10.1007/s11227-023-05139-w</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>S.</given-names></name> <name><surname>Xiong</surname> <given-names>Y.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Spatial temporal graph convolutional networks for skeleton-based action recognition,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, 32. <pub-id pub-id-type="doi">10.1609/aaai.v32i1.12328</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yue</surname> <given-names>Y.</given-names></name> <name><surname>Deng</surname> <given-names>J. D.</given-names></name> <name><surname>De Ridder</surname> <given-names>D.</given-names></name> <name><surname>Manning</surname> <given-names>P.</given-names></name> <name><surname>Adhia</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>Variational autoencoder learns better feature representations for EEG-based obesity classification</article-title>. <source>arXiv preprint arXiv:2302.00789</source>.</citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Hou</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Contrastive positive mining for unsupervised 3d action representation learning,&#x0201D;</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Springer</publisher-loc>), <fpage>36</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-031-19772-7_3</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Duan</surname> <given-names>L.</given-names></name> <name><surname>Gan</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;STST: spatial-temporal specialized transformer for skeleton-based action recognition,&#x0201D;</article-title> in <source>Proceedings of the 29th ACM International Conference on Multimedia</source>, 3229&#x02013;3237. <pub-id pub-id-type="doi">10.1145/3474085.3475473</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>K.</given-names></name> <name><surname>Cheng</surname> <given-names>S.</given-names></name> <name><surname>Kovalchuk</surname> <given-names>N.</given-names></name> <name><surname>Simmons</surname> <given-names>M.</given-names></name> <name><surname>Guo</surname> <given-names>Y.-K.</given-names></name> <name><surname>Matar</surname> <given-names>O. K.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Analyzing drop coalescence in microfluidic devices with a deep learning generative model</article-title>. <source>Phys. Chem. Chem. Phys</source>. <volume>25</volume>, <fpage>15744</fpage>&#x02013;<lpage>15755</lpage>. <pub-id pub-id-type="doi">10.1039/D2CP05975D</pub-id><pub-id pub-id-type="pmid">37232111</pub-id></citation></ref>
</ref-list>
</back>
</article>