<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2021.746985</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Micro-Expression Recognition Based on Pixel Residual Sum and Cropped Gaussian Pyramid</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhao</surname> <given-names>Yuan</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1277861/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chen</surname> <given-names>Zhuang</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1556355/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Luo</surname> <given-names>Song</given-names></name>
</contrib>
</contrib-group>
<aff><institution>School of Computer Science and Engineering, Chongqing University of Technology</institution>, <addr-line>Chongqing</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Zhen Cui, Nanjing University of Science and Technology, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Sofiane Boucenna, UMR8051 Equipes Traitement de l&#x00027;Information et Syst&#x000E8;mes (ETIS), France; Weisheng Li, Chongqing University of Posts and Telecommunications, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Yuan Zhao <email>1290889111&#x00040;qq.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>746985</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Zhao, Chen and Luo.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Zhao, Chen and Luo</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Facial micro-expression(ME) recognition has great significance for the progress of human society and could find a person&#x00027;s true feelings. Meanwhile, ME recognition faces a huge challenge, since it is difficult to detect and easy to be disturbed by the environment. In this article, we propose two novel preprocessing methods based on Pixel Residual Sum. These methods can preprocess video clips according to the unit pixel displacement of images, resist environmental interference, and be easy to extract subtle facial features. Furthermore, we propose a Cropped Gaussian Pyramid with Overlapping(CGPO) module, which divides images of different resolutions through Gaussian pyramids and crops different resolutions images into multiple overlapping subplots. Then, we use a convolutional neural networks of progressively increasing channels based on the depthwise convolution to extract preliminary features. Finally, we fuse preliminary features and make position embedding to get the last features. Our experiments show that the proposed methods and model have better performance than the well-known methods.</p></abstract>
<kwd-group>
<kwd>micro-expression recognition</kwd>
<kwd>deep learning</kwd>
<kwd>Gaussian pyramid</kwd>
<kwd>pixel residual sum</kwd>
<kwd>position embedding</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="4"/>
<equation-count count="11"/>
<ref-count count="49"/>
<page-count count="11"/>
<word-count count="6637"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Facial expression is a crucial channel for interpersonal socializing and can be used to convey inner emotions in daily life. Facial expression is divided into micro-expression(ME) and macro-expression. In past decades, macro-expression had a wide range of applications, and scholars have done a lot of research on macro-expression and facial recognition (Boucenna et al., <xref ref-type="bibr" rid="B1">2014</xref>; Liu et al., <xref ref-type="bibr" rid="B29">2018</xref>; Kim et al., <xref ref-type="bibr" rid="B17">2019</xref>; Xie et al., <xref ref-type="bibr" rid="B43">2019</xref>), but macro-expression is deceptive and can be easily hidden by human control. In contrast, ME will be unintentionally exposed as long as people intend to hide their true feeling. Hence, ME recognition has attracted much attention and has an extensive application prospect, such as clinical diagnosis, judiciary authorities, political elections, and national security.</p>
<p>ME has the following characteristics:</p>
<list list-type="bullet">
<list-item><p>ME is a very short facial expression and lasts between 1/25 and 1/3 (Yan et al., <xref ref-type="bibr" rid="B46">2013</xref>). As a result, untrained individuals have a weaker ability to recognize ME (Lies, <xref ref-type="bibr" rid="B24">1992</xref>).</p></list-item>
<list-item><p>ME is an unconscious and involuntary facial expression appearing when people disguise one&#x00027;s emotions and can be triggered in high-risk environments and show real or hidden emotions.</p></list-item>
<list-item><p>ME usually only appears in specific locations (Ekman and Friesen, <xref ref-type="bibr" rid="B6">1971</xref>; Ekman, <xref ref-type="bibr" rid="B5">2009b</xref>).</p></list-item>
<list-item><p>ME usually needs to be analyzed in the video clip, and macro-expression can be analyzed in the image.</p></list-item>
</list>
<p>Due to these characteristics, it is difficult to recognize the ME artificially. Therefore, Ekman and Paul tried a lot of efforts to improve the ability of individuals to recognize the ME, and they developed a tool for ME recognition in 2002 Micro Expression Training Tool (METT) (Ekman, <xref ref-type="bibr" rid="B4">2009a</xref>), which can effectively improve the individual&#x00027;s ability to recognize ME. However, the accuracy of relying on human recognition of ME is not high. According to reports, the accuracy of human-identified ME is only 47% (Frank et al., <xref ref-type="bibr" rid="B8">2009</xref>). Therefore, it is particularly important to recognize the ME through computer vision. With the development of technology, the rise of high-speed cameras and deep learning has made it possible to accurately recognize the ME. However, the current ME recognition is mainly faced with the following problems.</p>
<list list-type="bullet">
<list-item><p>How to extract the subtle feature of the human face?</p></list-item>
<list-item><p>How to overcome frame redundancy in the ME video?</p></list-item>
<list-item><p>How to have stronger universality and overcome environmental changes?</p></list-item>
</list>
<p>The structure of the study is as follows: In Section II, the pieces of literature related to ME recognition are reviewed in detail; In Section III, a preprocessing method and network framework for ME recognition are proposed; In Section IV, we describe the experimental settings and analyze the experimental results; Finally, Section V summarizes this study with remarks. The contributions of this study are as follows.</p>
<list list-type="bullet">
<list-item><p>We propose two more effective methods of preprocessing, which combine spatio-temporal dimensionality and can extract more robust features.</p></list-item>
<list-item><p>We design a module of Cropped Gaussian Pyramid with Overlapping(CGPO), which can use different scales information.</p></list-item>
<list-item><p>We design a network with feature fusion, and the network structure adopts a gradual way of increasing channels.</p></list-item>
</list>
</sec>
<sec id="s2">
<title>2. Related Work</title>
<sec>
<title>2.1. Handcrafted Features</title>
<p>Several years before, ME recognition was mainly based on traditionally handcrafted feature descriptors. These descriptors can be divided into geometric features and appearance features.</p>
<sec>
<title>2.1.1. Appearance-Based Features</title>
<p>For instance, Local Binary Pattern histograms from Three Orthogonal Planes (LBP-TOP) (Zhao and Pietikainen, <xref ref-type="bibr" rid="B48">2007</xref>), Spatiotemporal Completed Local Quantization Patterns (STCLQP) (Huang et al., <xref ref-type="bibr" rid="B14">2016</xref>), and LBP with Six Intersection Points (LBP-SIP) (Wang et al., <xref ref-type="bibr" rid="B40">2014</xref>) can be considered as methods based on appearance features. These methods led that the features, dimensions are relatively high with more redundant information.</p>
<p>The LBP-TOP, a development of the LBP in a three-dimensional space, is a typical LBP descriptor with spatial-temporal characteristics. The LBP-TOP operator extracts LBP features on the three orthogonal planes. Next, obtained results are stitched as the final LBP-TOP feature, since the video can be regarded as a cube in the three dimensions of x, y, and t. The LBP-TOP not only considers the spatial information but also considers the information in the video sequence. After obtaining the LBP-TOP features, Zhao et al. use Support Vector Machine(SVM) for spotting and classification. Zhao et al. made good use of LBP-TOP features, and used many tricks of conventional expression analysis. As an early work, the work has achieved good results and has established the basis for the subsequent ME recognition.</p>
<p>The LBP-TOP has great limitations for only considering the local appearance and movement characteristics. So, Huang et al. (<xref ref-type="bibr" rid="B14">2016</xref>) proposed STCLQP for the ME recognition. First, three significant information, including magnitude, orientation, and sign components, are extracted by STCLQP. Second, for each component in temporal and appearance domains, Huang et al. (<xref ref-type="bibr" rid="B14">2016</xref>) made dense and characteristic codebooks by developing productive codebook selection and vector quantization. Finally, in terms of this codebook, Huang et al. (<xref ref-type="bibr" rid="B14">2016</xref>) extracted and fused spatio-temporal features, included orientation components, magnitude, and sign. Compared with LBP-TOP, the STCLQP method considers more information. Although the recognition accuracy is improved, it will inevitably lead to higher dimensions.</p>
<p>Furthermore, Wang et al. (<xref ref-type="bibr" rid="B40">2014</xref>) proposed LBP-SIP volumetric descriptor, which is based on three intersecting lines passing through a central point. The superabundance of LBP-TOP patterns is diminished by LBP-SIP. Furthermore, LBP-SIP provides a more dense and weightless characterization and reduces computational complexity. It further promotes the improvement of the accuracy of the ME recognition and has become the baseline for many subsequent works.</p>
</sec>
<sec>
<title>2.1.2. Geometric-Based Features</title>
<p>Optical flow, a geometric-based feature, calculates the displacement of facial feature points or the optical flow of the action area. It can extract representative motion features that are robust for the diversity of facial textures. Furthermore, the data except for RGB channels can be enhanced by optical flow (Liu et al., <xref ref-type="bibr" rid="B28">2019</xref>).</p>
<p>Many works treat optical flow as a data preprocessing step. Liu et al. (<xref ref-type="bibr" rid="B30">2015</xref>) proposed an uncomplicated yet productive Main Directional Mean Optical-flow (MDMO) feature. On the ME video clips, an effective optical flow method is adopted. Meanwhile, Liu utilizes partial action units to divide the face into regions of interest (ROIs). MDMO is a normalized feature based on ROIs. It combines both spatial location and local statistic motion characteristics. MDMO has the advantage of small feature dimensions.</p>
<p>Some works (Liong et al., <xref ref-type="bibr" rid="B25">2019</xref>; Liu et al., <xref ref-type="bibr" rid="B28">2019</xref>; Zhou et al., <xref ref-type="bibr" rid="B49">2019</xref>) utilized optical flow information for ME recognition and have achieved good results. For instance, Liu et al. (<xref ref-type="bibr" rid="B28">2019</xref>) utilized two domain adaptation methods, which include expression magnification and reduction and adversarial training. Then, he preprocessed the raw images to capture the spatio-temporal optical flow from facial movements from onset frame (the first frame in the ME video) to apex frame (the most intense frame of action in the ME video), won the championship of 2019-the second facial Micro-expressions Grand Challenge (MEGC2019) (See et al., <xref ref-type="bibr" rid="B35">2019</xref>). Zhou et al. (<xref ref-type="bibr" rid="B49">2019</xref>) captured the TV-L1 optical flow (Zach et al., <xref ref-type="bibr" rid="B47">2007</xref>) of the onset frame and the mid-position frame, and then performs ME recognition through the Dual-Inception network. Instead of using apex frames, they use mid-position frames to cut down computation complexity. Furthermore, Liong et al. (<xref ref-type="bibr" rid="B25">2019</xref>) designed a STSTNet, which can be used to learn three features of optical flow, namely vertical optical flow, optical strain, and horizontal optical flow. These features are calculated by the onset frame and apex frame of ME video.</p>
<p>Optical flow has the advantage of small feature dimensions and the ability to capture subtle muscle movements. However, the optical flow has higher requirements on light and is easily affected by the external environment. In addition, these works only use the optical flow information of the apex frame and onset frame and lose the motion information of other frames.</p>
</sec>
</sec>
<sec>
<title>2.2. Deep Neural Networks</title>
<p>Deep learning (LeCun et al., <xref ref-type="bibr" rid="B21">2015</xref>) is universally used in various industries. Especially during the immediate past, the works on deep learning in the ME recognition field has gradually increased. In the field of deep learning, the features preprocessed by the optical flow method and LBP can be used as the input of convolution neural network (CNN). Then, CNN is usually used for feature extractors. For instance, Xia et al. (<xref ref-type="bibr" rid="B41">2019</xref>) proposed spatio-temporal recurrent convolutional networks based on optical flow, which extracts the optical flow information from the onset frame until the apex frame and inputs it into recurrent convolutional networks.</p>
<p>Furthermore, some works also use Long Short-term Memory (LSTM) to directly input ME video clips. One early work (Khor et al., <xref ref-type="bibr" rid="B16">2018</xref>) proposed an Enriched Long-term Recurrent Convolutional Network (ELRCN). First, every ME frame is encoded into a feature vector by CNN modules. Then, ELRCN uses an LSTM module to pass the feature vector and predicts ME at last. ELRCN uses the feature that the information can be retained for a long time in the gating unit to detect ME in the video, and achieve good performance. Therefore, the combination of LSTM and CNN have greater advantages in recognizing ME in videos. However, due to the small changes in the ME video clips, there is frame redundancy, leading to greater computational complexity.</p>
<p>In conclusion, compared with traditional manual features for ME recognition, deep learning technology can extract features from ME videos and classify them with higher accuracy. However, due to frame redundancy in ME videos, the speed of the deep learning training model is greatly affected. Therefore, we propose two new ME video preprocessing methods to overcome frame redundancy in ME video and improve the recognition of ME classes.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Method</title>
<sec>
<title>3.1. Preprocessing</title>
<p>As we discussed above, it is an inevitable stage to extract a discriminative and efficient feature. Therefore, this study proposes two methods based on the residual sum of image pixels to extract salient features: (1) Absolute Residual Sum (ARS) and (2) Relative Residual Sum (RRS). These methods take the frames in the ME clip at fixed intervals and consider the regional pixel displacement between frames. It not only avoids the redundancy of the ME clip but also makes full use of the ME information. The pixel-level displacement difference sum, named RS, can explain the tiny movement of the object. ARS and RRS preprocessing procedure are shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Preprocessing Flow chart (&#x000A9;Xiaolan Fu).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-746985-g0001.tif"/>
</fig>
<sec>
<title>3.1.1. Absolute Residual Sum</title>
<p>Preprocessing is divided into five stages.</p>
<sec>
<title>3.1.1.1. Select Video Clip</title>
<p>He et al. proposed MDMD, which used a reciprocal change from the onset frame to the offset frame to spotting ME (He et al., <xref ref-type="bibr" rid="B11">2020</xref>). Therefore, we only recognize the ME from the onset frame to the apex frame. First, we select a video clip from the ME video and calculate its start and end. We select the partial video clips from the ME video clip. The onset frame is taken as the start by Equation (1), and select the end by Equation (2).</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where T(<italic>x</italic>) represents the frame sequence of x in the video.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mn>10</mml:mn><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;</mml:mtext></mml:mtd><mml:mtd><mml:mo>&#x0003C;</mml:mo><mml:mn>10</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>min</italic>(<italic>x, y</italic>) represents the smaller values of <italic>x</italic> and <italic>y</italic>.</p>
</sec>
<sec>
<title>3.1.1.2. Detect Feature Point</title>
<p>The dlib library is utilized to spotting facial feature points.</p>
</sec>
<sec>
<title>3.1.1.3. Cropping</title>
<p>Cropping the face through the face feature points.</p>
</sec>
<sec>
<title>3.1.1.4. Select Five Frames</title>
<p>Notice that, ME data is very redundant. Useful information must be mined from the data. A few other works (Li et al., <xref ref-type="bibr" rid="B22">2013</xref>; Le Ngo et al., <xref ref-type="bibr" rid="B18">2015</xref>, <xref ref-type="bibr" rid="B20">2016</xref>) have proposed many methods to reduce frame redundancy in ME videos by using partial frames. Therefore, we require mining crucial frames from ME video clip. We define crucial frames as key-frames and define frames except for the key-frames as transition frames. Furthermore, we make two assumptions for getting rid of transition frames: (1) Transition frames are highly similar to the key-frames, and deletion does not affect the recognition accuracy. (2) Transition frames are continuously distributed, centered on key-frames.</p>
<p>Hence, we choose appropriate intervals by Equation (3) and select five key-frames as elements in &#x1D53D; according to Equation (4).</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">&#x02308;</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">&#x02309;</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x1D53D;</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mo>&#x0002A;</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where &#x02308;<italic>x</italic>&#x02309; is taking the smallest integer not less than <italic>x</italic> for some scalar, and <italic>N</italic><sub><italic>key</italic></sub> represents the number of key frames. <italic>N</italic><sub><italic>key</italic></sub> is set to five in the paper.</p>
</sec>
<sec>
<title>3.1.1.5. Generate Redisual Sum Image</title>
<p>Liu et al. (<xref ref-type="bibr" rid="B28">2019</xref>) took the motion difference between the onset frame and each frame to calibrate the apex frame, because the intensity relationship of ME can be indicated by the motion difference. Therefore, we cumulate the motion difference for calculating the variation trend of a single pixel. For the key frame in &#x1D53D;, Equation (5) is used to calculate the ARS.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mo>&#x1D53D;</mml:mo></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mi>%</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mn>256</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>Q</italic><sub><italic>f</italic></sub>(<italic>x, y, z</italic>) represents the pixel value of the three-channel image (<italic>x, y, z</italic>) of the <italic>f</italic><sub><italic>th</italic></sub> frame and <italic>ares</italic>(<italic>x, y, z</italic>) represents the pixel value of the generated ARS image.</p>
</sec>
</sec>
<sec>
<title>3.1.2. Relative Residual Sum</title>
<p>As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, the steps before the fifth step are the same as ARS. In the fifth step, we use Equation (6) to calculate the sum of residuals between frames. Then, we use Equation (7) to transform the range of sum to between <italic>gmin</italic> and <italic>gmax</italic>. In this experiment, <italic>gmin</italic> &#x0003D; 0 and <italic>gmax</italic> &#x0003D; 255.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mo>&#x1D53D;</mml:mo></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E7"><label>(7)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x0002A;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Where <italic>max</italic>(<italic>x, y</italic>), <italic>diff</italic>(<italic>x, y, z</italic>), and <italic>rres</italic>(<italic>x, y, z</italic>) represent the greater values of <italic>x</italic> and <italic>y</italic>, the sum of the displacement of the video frame at the three-channel image (<italic>x, y, z</italic>), and the pixel value of the generated RRS image, respectively.</p>
</sec>
</sec>
<sec>
<title>3.2. Framework</title>
<p>CropNet, based on the depthwise convolution (Sandler et al., <xref ref-type="bibr" rid="B34">2018</xref>), is used as a classification model. CropNet takes advantage of CGPO. The architecture of the CropNet is shown in <bold>Figure 3</bold>. Conv, BN, and FC in the figure represent Convolutional Layer, Batch Normalization Layer, and Fully Connected Layer, respectively.</p>
<sec>
<title>3.2.1. Image Augmentation</title>
<p>The number of network parameters is approximately 7.6M. Image augment is essential as the network framework is slightly large. According to the characteristics of the human face, we performed the following four data augmentation in turn. (1) The image brightness, contrast, and saturation are randomly changed to [20%, 180%] of the original image brightness, and the hue offset of the image is changed to [&#x02212;0.5, 0.5] of the original image. (2) The picture is converted to grayscale with a probability of 20%. (3) Flipping the image horizontally with a 50% probability. (4) Rotating the image randomly clockwise [&#x02212;15,15] degrees. The image augment module is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Image augment module.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-746985-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.2.2. Cropped Gaussian Pyramid With Overlapping</title>
<p>Different facial areas have different importance in the production of ME. Therefore, we propose a CGPO module, which divides ME video frames with different resolutions of the image into 10 overlapping subplots. It can separate the mouth, the eyes, the nose, etc. The introduction of overlapping mechanisms can reduce the risk of separating important parts of the face. The CGPO module is shown in <xref ref-type="fig" rid="F3">Figure 3</xref> CGPO, and its processing flow is as follows.</p>
<list list-type="bullet">
<list-item><p>First, we require 320 &#x000D7; 320 resolution of the image input and down-sample it to get an image with a resolution of 160 &#x000D7; 160.</p></list-item>
<list-item><p>Second, for each image with different scale resolution, we divide them into several 160 &#x000D7; 160 images and introduce the overlap factor &#x003B1;. &#x003B1; is used to control the size of the overlap when crop images with different precision. In this study, &#x003B1; is 0.3.</p></list-item>
<list-item><p>Finally, after going through the above process, images are fed CNN based with the depthwise convolution.</p></list-item>
</list>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The architecture of the network model. The numbers on convolution and MB6 block represent the number of output channels. MB6 refers to MobileNetV2 (Sandler et al., <xref ref-type="bibr" rid="B34">2018</xref>)&#x00027;s inverted bottlenecks with an expansion ratio of 6.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-746985-g0003.tif"/>
</fig>
</sec>
<sec>
<title>3.2.3. Feature Extraction</title>
<p>Han et al. (<xref ref-type="bibr" rid="B10">2020</xref>) designed ReXNet, which has achieved very good results in the ImageNet Challenge. Therefore, we use the ReXNet feature extraction module as the extractor. A network of progressively increasing channels are leveraged on the extracting feature, as shown in <xref ref-type="fig" rid="F3">Figure 3</xref> feature extractor.</p>
<p>Due to the difficulties in data collection and identification of ME, there are few ME datasets. It is difficult to apply deep learning in ME recognition. Therefore, we train this module on the ImageNet datasets (Deng et al., <xref ref-type="bibr" rid="B3">2009</xref>) and then apply it to the ME recognition through the transfer learning method (Pan and Yang, <xref ref-type="bibr" rid="B32">2009</xref>).</p>
</sec>
<sec>
<title>3.2.4. Feature Fusion and Classifier</title>
<p>Feature Fusion and Classifier are shown in <xref ref-type="fig" rid="F3">Figure 3</xref> Classifier. The features extracted in the previous module go through the Convolutional Layer, Batch Normalization Layer, Adaptive Pooling Layer, and Fully Connected Layer, in turn, and become a feature vector <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>i</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mn>24</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, where i represents the order of segmented images. Since the CGOP module segmented a total of 10 images, we could obtain 10 feature vectors {<italic><bold>z</bold></italic><sub><bold>0</bold></sub>, <italic><bold>z</bold></italic><sub><bold>1</bold></sub>&#x022EF;&#x022EF;&#x022EF;<italic><bold>z</bold></italic><sub><bold>9</bold></sub>}.</p>
<p>However, because the position information after image cropping becomes blurred, the model has a hard time learning about correlations between images. We combine the location information with the feature to make the features more explanatory. Therefore, for feature vectors {<italic><bold>z</bold></italic><sub><bold>1</bold></sub>, <italic><bold>z</bold></italic><sub><bold>2</bold></sub>&#x022EF;&#x022EF;&#x022EF;<italic><bold>z</bold></italic><sub><bold>9</bold></sub>} of segmented images, we introduce trainable position embedding vectors {<italic><bold>p</bold></italic><sub><bold>1</bold></sub>, <italic><bold>p</bold></italic><sub><bold>2</bold></sub>&#x022EF;&#x022EF;&#x022EF;<italic><bold>p</bold></italic><sub><bold>9</bold></sub>} to learn the position information of the image, where <italic>p</italic><sub><italic>i</italic></sub> has the same dimension as <italic><bold>z</bold></italic><sub><italic><bold>i</bold></italic></sub>. The position embedding vectors are initialized to random values that follow a normal distribution. The mean of the random values is 0 and the variance is 0.2. As shown in Equation (8), we calculate the new feature vectors <inline-formula><mml:math id="M9"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mn>2</mml:mn></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x022EF;</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mn>9</mml:mn></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>i</mml:mi></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x02032;</mml:mi></mml:mstyle></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>i</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo>&#x02295;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>p</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x02003;&#x000A0;</mml:mtext><mml:mn>0</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>10</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Finally, we mix <inline-formula><mml:math id="M11"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>0</mml:mn></mml:mstyle></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x02032;</mml:mi></mml:mstyle></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>2</mml:mn></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x02032;</mml:mi></mml:mstyle></mml:mrow></mml:msup><mml:mo>&#x022EF;</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>9</mml:mn></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> by splicing and classifying ME.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Experiment</title>
<sec>
<title>4.1. Datasets</title>
<p>Due to the characteristics of ME and its difficulty in triggering and collecting, the dataset is very scarce. As far as we know, there are three spontaneous datasets generally utilized for ME recognition: SMIC-HS (Li et al., <xref ref-type="bibr" rid="B22">2013</xref>), SAMM (Davison et al., <xref ref-type="bibr" rid="B2">2016</xref>), and CASME II (Yan et al., <xref ref-type="bibr" rid="B44">2014a</xref>). The details of these three spontaneous datasets are shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Micro-expression (ME) datasets.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Datasets</bold></th>
<th valign="top" align="left"><bold>CASME II</bold></th>
<th valign="top" align="left"><bold>SMIC-HS</bold></th>
<th valign="top" align="left"><bold>SAMM</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Particpants</td>
<td valign="top" align="left">26</td>
<td valign="top" align="left">16</td>
<td valign="top" align="left">29</td>
</tr>
<tr>
<td valign="top" align="left">Samples</td>
<td valign="top" align="left">255</td>
<td valign="top" align="left">157</td>
<td valign="top" align="left">159</td>
</tr>
<tr>
<td valign="top" align="left">Resolution</td>
<td valign="top" align="left">640&#x0002A;480</td>
<td valign="top" align="left">640&#x0002A;480</td>
<td valign="top" align="left">960&#x0002A;650</td>
</tr>
<tr>
<td valign="top" align="left">Frame rate(fps)</td>
<td valign="top" align="left">200</td>
<td valign="top" align="left">100</td>
<td valign="top" align="left">200</td>
</tr>
<tr>
<td valign="top" align="left">FACS coded</td>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">x</td>
<td valign="top" align="left">&#x02713;</td>
</tr>
<tr>
<td valign="top" align="left">APEX index</td>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">x</td>
<td valign="top" align="left">&#x02713;</td>
</tr>
<tr>
<td valign="top" align="left">Emotion</td>
<td valign="top" align="left">Other(99) Disgust(63) Surprise(28)<break/> Repression(27) Sadness(4) Happiness(32)<break/> Fear(2)</td>
<td valign="top" align="left">Negative(66)<break/> Positive(51)<break/> Surprise(40)</td>
<td valign="top" align="left">Other(26) Happiness(26) Disgust(9)<break/> Surprise(15) Sadness(6) Anger(57)<break/> Fear(8) Contempt(12)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.2. Experiment Settings</title>
<p>All experiments for this study were all carried out on Ubuntu 16.04 and Python 3.6.2 with Pytorch 1.6 on NVIDIA GTX Titan RTX GPU (24 GB). The label smoothing loss function (Lukasik et al., <xref ref-type="bibr" rid="B31">2020</xref>) is leveraged as the loss function. It can better generalize the network and ultimately produce, more accurate predictions on invisible data. AdamP (Heo et al., <xref ref-type="bibr" rid="B12">2021</xref>) is used as an optimizer. We use UF1 (commonly referred to as the macro average F1 score), UAR (commonly referred to as balanced accuracy), and Accuracy as our evaluation standard.</p>
<list list-type="bullet">
<list-item><p><bold>UF1</bold> score can equally emphasize in a rare class. So, it is a suitable indicator in a multi-class evaluation. The calculation formula for UF1 is as follows:
<disp-formula id="E9"><label>(9)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>U</mml:mi><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x0002A;</mml:mo><mml:mi>T</mml:mi><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x0002A;</mml:mo><mml:mi>T</mml:mi><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>F</mml:mi><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Where <italic>C</italic> represents the number of classes and <italic>FP</italic><sub><italic>i</italic></sub>, <italic>TP</italic><sub><italic>i</italic></sub>, and <italic>FN</italic><sub><italic>i</italic></sub> represent the false positive, the true positive, and the false negative for the <italic>i</italic><sub><italic>th</italic></sub> class, respectively.</p></list-item>
<list-item><p><bold>UAR</bold> is a more appropriate indicator instead of the standard accuracy indicator that may be partial to larger classes. The calculation formula for UAR is as follows:
<disp-formula id="E10"><label>(10)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>U</mml:mi><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Where <italic>N</italic><sub><italic>i</italic></sub> represents the number of <italic>i</italic><sub><italic>th</italic></sub> class.</p></list-item>
<list-item><p><bold>Accuracy</bold> is commonly used as a CASME II experiment in five classes. The calculation formula for Accuracy is as follows:
<disp-formula id="E11"><label>(11)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></list-item>
</list>
</sec>
<sec>
<title>4.3. Experiment With Five Classes of ME in the CASME II</title>
<p>We choose the CASME II as the evaluation dataset. Only five classes (Others, Disgust, Happiness, Repression, and Surprise) are considered, since the fear and sadness samples are very scarce. In this experiment, Leave-One-Subject-Out (LOSO) cross validation is utilized for evaluation protocol. LOSO cross validation refers to using the samples of one subject as the test set, and the rest as the training set in each fold. It can prevent the test set and the training set from having the same sample, thereby avoiding data leakage. Recognition Accuracy can be calculated by the LOSO cross validation evaluation protocol. In the same evaluation standard, we compare with a variety of methods. The result is shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Comparison of ME recognition performance in CASME II (5 classes).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">LBP-Top&#x0002B;AdaBoost (Le Ngo et al., <xref ref-type="bibr" rid="B19">2014</xref>)</td>
<td valign="top" align="center">0.437</td>
</tr>
<tr>
<td valign="top" align="left">STCLQP (Huang and Zhao, <xref ref-type="bibr" rid="B13">2017</xref>)</td>
<td valign="top" align="center">0.583</td>
</tr>
<tr>
<td valign="top" align="left">ELRCN (Khor et al., <xref ref-type="bibr" rid="B16">2018</xref>)</td>
<td valign="top" align="center">0.524</td>
</tr>
<tr>
<td valign="top" align="left">DSSN (Khor et al., <xref ref-type="bibr" rid="B15">2019</xref>)</td>
<td valign="top" align="center">0.707</td>
</tr>
<tr>
<td valign="top" align="left">TSCNN-I (Song et al., <xref ref-type="bibr" rid="B36">2019</xref>)</td>
<td valign="top" align="center">0.740</td>
</tr>
<tr>
<td valign="top" align="left">SSSN (Khor et al., <xref ref-type="bibr" rid="B15">2019</xref>)</td>
<td valign="top" align="center">0.711</td>
</tr>
<tr>
<td valign="top" align="left">TSCNN-II (Song et al., <xref ref-type="bibr" rid="B36">2019</xref>)</td>
<td valign="top" align="center">0.810</td>
</tr>
<tr>
<td valign="top" align="left">Bi-WOOF (apex and onset) (Liong et al., <xref ref-type="bibr" rid="B26">2018</xref>)</td>
<td valign="top" align="center">0.578</td>
</tr>
<tr>
<td valign="top" align="left">Su et al. (Su et al., <xref ref-type="bibr" rid="B37">2021</xref>)</td>
<td valign="top" align="center">0.727</td>
</tr>
<tr>
<td valign="top" align="left"><bold>RRS&#x0002B;CropNet(ours)</bold></td>
<td valign="top" align="center">0.790</td>
</tr>
<tr>
<td valign="top" align="left"><bold>ARS&#x0002B;CropNet(ours)</bold></td>
<td valign="top" align="center"><bold>0.862</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The confusion matrix obtained by applying the ARS and the CropNet is shown in <xref ref-type="fig" rid="F4">Figure 4B</xref>. Through the confusion matrix, the overall recognition rate is very high. The proposed method has great performance for all classes.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>(A)</bold> is the confusion matrix of composite datasets (SMIC-HS, CASME II, and SAMM) in the absolute residual sum (ARS) and the CropNet. <bold>(B)</bold> is the confusion matrix of CASME II in the ARS and the CropNet.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-746985-g0004.tif"/>
</fig>
</sec>
<sec>
<title>4.4. Composite Datasets Evaluation (CDE)</title>
<p>Composite datasets evaluation is a very effective evaluation method in cross-database ME recognition. In this experiment, we use the MEGC2019 standard. According to MEGC2019 standards, we combined all samples of the datasets (SAMM, CASME II, and SMIC-HS) into a composite dataset by unifying the number of ME class. ME are divided into three classes: negative, surprised, and positive. Disgust, contempt, fear, sadness, and anger is regarded as the negative class. Surprise is still regarded as surprise class. Happiness is regarded as the positive class. LOSO cross validation is utilized to split the training set and test set. <xref ref-type="table" rid="T3">Table 3</xref> compares the performance of proposed methods against a number of recent study. The methods in <xref ref-type="table" rid="T3">Table 3</xref> were all compared in the same datasets and at the same evaluation standard. The confusion matrix obtained by applying the ARS and the CropNet is shown in <xref ref-type="fig" rid="F4">Figure 4A</xref>. It shows that three classes have similar performance, and the proposed method also has a good fit for unbalanced data.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Comparison of ME recognition performance composite datasets.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;" colspan="2"><bold>Composite</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;" colspan="2"><bold>SMIC-HS</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;" colspan="2"><bold>CASME II</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;" colspan="2"><bold>SAMM</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>UF1</bold></th>
<th valign="top" align="center"><bold>UAR</bold></th>
<th valign="top" align="center"><bold>UF1</bold></th>
<th valign="top" align="center"><bold>UAR</bold></th>
<th valign="top" align="center"><bold>UF1</bold></th>
<th valign="top" align="center"><bold>UAR</bold></th>
<th valign="top" align="center"><bold>UF1</bold></th>
<th valign="top" align="center"><bold>UAR</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">LBP-TOP (Zhao and Pietikainen, <xref ref-type="bibr" rid="B48">2007</xref>)</td>
<td valign="top" align="center">0.588</td>
<td valign="top" align="center">0.578</td>
<td valign="top" align="center">0.200</td>
<td valign="top" align="center">0.528</td>
<td valign="top" align="center">0.702</td>
<td valign="top" align="center">0.742</td>
<td valign="top" align="center">0.395</td>
<td valign="top" align="center">0.410</td>
</tr>
<tr>
<td valign="top" align="left">Bi-WOOF (Liong et al., <xref ref-type="bibr" rid="B26">2018</xref>)</td>
<td valign="top" align="center">0.629</td>
<td valign="top" align="center">0.622</td>
<td valign="top" align="center">0.572</td>
<td valign="top" align="center">0.582</td>
<td valign="top" align="center">0.780</td>
<td valign="top" align="center">0.802</td>
<td valign="top" align="center">0.521</td>
<td valign="top" align="center">0.512</td>
</tr>
<tr>
<td valign="top" align="left">CapsuleNet (Van Quang et al., <xref ref-type="bibr" rid="B39">2019</xref>)</td>
<td valign="top" align="center">0.652</td>
<td valign="top" align="center">0.650</td>
<td valign="top" align="center">0.582</td>
<td valign="top" align="center">0.587</td>
<td valign="top" align="center">0.706</td>
<td valign="top" align="center">0.701</td>
<td valign="top" align="center">0.620</td>
<td valign="top" align="center">0.598</td>
</tr>
<tr>
<td valign="top" align="left">OFF-ApexNet (Gan et al., <xref ref-type="bibr" rid="B9">2019</xref>)</td>
<td valign="top" align="center">0.719</td>
<td valign="top" align="center">0.709</td>
<td valign="top" align="center">0.681</td>
<td valign="top" align="center">0.669</td>
<td valign="top" align="center">0.876</td>
<td valign="top" align="center">0.868</td>
<td valign="top" align="center">0.540</td>
<td valign="top" align="center">0.539</td>
</tr>
<tr>
<td valign="top" align="left">Dual-Inception (Zhou et al., <xref ref-type="bibr" rid="B49">2019</xref>)</td>
<td valign="top" align="center">0.732</td>
<td valign="top" align="center">0.727</td>
<td valign="top" align="center">0.664</td>
<td valign="top" align="center">0.672</td>
<td valign="top" align="center">0.862</td>
<td valign="top" align="center">0.856</td>
<td valign="top" align="center">0.586</td>
<td valign="top" align="center">0.566</td>
</tr>
<tr>
<td valign="top" align="left">STSTNet (Liong et al., <xref ref-type="bibr" rid="B25">2019</xref>)</td>
<td valign="top" align="center">0.735</td>
<td valign="top" align="center">0.760</td>
<td valign="top" align="center">0.680</td>
<td valign="top" align="center">0.701</td>
<td valign="top" align="center">0.838</td>
<td valign="top" align="center">0.868</td>
<td valign="top" align="center">0.658</td>
<td valign="top" align="center">0.681</td>
</tr>
<tr>
<td valign="top" align="left">ELTRCN (Khor et al., <xref ref-type="bibr" rid="B16">2018</xref>)</td>
<td valign="top" align="center">0.788</td>
<td valign="top" align="center">0.782</td>
<td valign="top" align="center">0.746</td>
<td valign="top" align="center">0.753</td>
<td valign="top" align="center">0.829</td>
<td valign="top" align="center">0.820</td>
<td valign="top" align="center">0.775</td>
<td valign="top" align="center">0.715</td>
</tr>
<tr>
<td valign="top" align="left">RCN-S (Xia et al., <xref ref-type="bibr" rid="B42">2020</xref>)</td>
<td valign="top" align="center">0.746</td>
<td valign="top" align="center">0.710</td>
<td valign="top" align="center">0.651</td>
<td valign="top" align="center">0.657</td>
<td valign="top" align="center">0.836</td>
<td valign="top" align="center">0.791</td>
<td valign="top" align="center">0.764</td>
<td valign="top" align="center">0.656</td>
</tr>
<tr>
<td valign="top" align="left">STSTNet&#x0002B;GA (Liu et al., <xref ref-type="bibr" rid="B27">2021</xref>)</td>
<td valign="top" align="center">0.836</td>
<td valign="top" align="center">0.836</td>
<td valign="top" align="center">0.814</td>
<td valign="top" align="center">0.812</td>
<td valign="top" align="center">0.882</td>
<td valign="top" align="center">0.891</td>
<td valign="top" align="center">0.800</td>
<td valign="top" align="center">0.790</td>
</tr>
<tr>
<td valign="top" align="left"><bold>RRS&#x0002B;CropNet(ours)</bold></td>
<td valign="top" align="center">0.875</td>
<td valign="top" align="center">0.877</td>
<td valign="top" align="center">0.813</td>
<td valign="top" align="center">0.819</td>
<td valign="top" align="center">0.972</td>
<td valign="top" align="center">0.969</td>
<td valign="top" align="center">0.842</td>
<td valign="top" align="center">0.827</td>
</tr>
<tr>
<td valign="top" align="left"><bold>ARS&#x0002B;CropNet(ours)</bold></td>
<td valign="top" align="center"><bold>0.911</bold></td>
<td valign="top" align="center"><bold>0.904</bold></td>
<td valign="top" align="center"><bold>0.855</bold></td>
<td valign="top" align="center"><bold>0.851</bold></td>
<td valign="top" align="center"><bold>0.974</bold></td>
<td valign="top" align="center"><bold>0.979</bold></td>
<td valign="top" align="center"><bold>0.912</bold></td>
<td valign="top" align="center"><bold>0.893</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Note that, the apex frame spotting is indispensable for ME recognition since the apex frame of the SMIC-HS dataset is not labeled. In recent years, there are a lot of apex frames spotting works (Yan et al., <xref ref-type="bibr" rid="B45">2014b</xref>; Li et al., <xref ref-type="bibr" rid="B23">2018</xref>; Peng et al., <xref ref-type="bibr" rid="B33">2019</xref>; Zhou et al., <xref ref-type="bibr" rid="B49">2019</xref>). In fact, apex frame spotting is a very difficult work. Therefore, this experiment considers a trade-off between efficiency and effectiveness. The middle frame of the video in the SMIC-HS dataset is used as the apex frame.</p>
</sec>
<sec>
<title>4.5. Ablation Experiments</title>
<p>We performed two ablation experiments on the CASME II dataset to verify the effectiveness of the module.</p>
<list list-type="bullet">
<list-item><p>We performed ablation experiments on preprocessing methods for comparing the effectiveness of the four preprocessing methods ARS, RRS, Farneback optical flow (Farneb&#x000E4;ck, <xref ref-type="bibr" rid="B7">2003</xref>), and TV-L1 optical flow.</p></list-item>
<list-item><p>We performed ablation experiments on model architect for verifying the effect of the GCOP module.</p></list-item>
</list>
<p>As shown in <xref ref-type="table" rid="T4">Table 4</xref>, ARS stands out among the four preprocessing methods. It can extract more reliable spatio-temporal features and improve the UF1 value of ME recognition. RRS also achieves very good results. There are significant differences between these two methods. RRS pays more attention to areas with greater displacement by relative displacement change between unit pixels, while is not too sensitive to small displacement areas. ARS considers the trade-off between displacement regions of different scales, which can focus on both small displacement areas and large displacement areas. Therefore, subtle displacement can be captured. At the same time, for areas with frequent displacement, ARS ignores the displacement of unit pixels and pays attention to regional displacement. But in our experimental environment, Farneback optical flow and TV-L1 optical flow are far less effective than the proposed methods in this study.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Ablation experiments in CASME II (5 classes).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Ablation module</bold></th>
<th valign="top" align="center"><bold>Ablation method</bold></th>
<th valign="top" align="center"><bold>UF1</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">paper method</td>
<td valign="top" align="center"><bold>CropNet&#x0002B;ARS</bold></td>
<td valign="top" align="center"><bold>0.863</bold></td>
<td valign="top" align="center"><bold>0.862</bold></td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Preprocessing Method</td>
<td valign="top" align="center">CropNet&#x0002B;RRS</td>
<td valign="top" align="center">0.803</td>
<td valign="top" align="center">0.790</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">CropNet&#x0002B;Optical FLow(Farneback)</td>
<td valign="top" align="center">0.661</td>
<td valign="top" align="center">0.625</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">CropNet&#x0002B;Optical FLow(TV-L1)</td>
<td valign="top" align="center">0.697</td>
<td valign="top" align="center">0.669</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Model architect</td>
<td valign="top" align="center">CropNet without GCOP &#x0002B;ARS</td>
<td valign="top" align="center">0.841</td>
<td valign="top" align="center">0.813</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The Cropped Gaussian Pyramid with Overlapping module focuses on different areas of the face, extracts features for each area, and then stitches the obtained features to classify them. Through the ablation experiment in <xref ref-type="table" rid="T4">Table 4</xref>, it is easy to find the efficiency of the CGPO module and the ARS method.</p>
<p>Furthermore, we conducted hyperparameter&#x00027;s ablation experiments in MEGC2019 composite datasets for verifying the effectiveness of the hyperparameters <italic>N</italic><sub><italic>key</italic></sub>. The experimental results are shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, which can be concluded that there is greater universality when <italic>N</italic><sub><italic>key</italic></sub> is set to five. Therefore, in all experiments, we only select five key-frames at equal intervals in the ME video clip.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><italic>N</italic><sub><italic>key</italic></sub> hyperparameter&#x00027;s ablation experiments.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-746985-g0005.tif"/>
</fig>
</sec>
<sec>
<title>4.6. Visualization Experiments</title>
<p>We use T-SNE (Van der Maaten and Hinton, <xref ref-type="bibr" rid="B38">2008</xref>) to visualize the preprocessed image for better comparing the effects of the proposed preprocessing methods. <xref ref-type="fig" rid="F6">Figure 6</xref> shows the feature distribution of images preprocessed by various methods. In this experiment, we use three classes (negative, positive, and surprised) of CASME II.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>(A&#x02013;D)</bold> represent preprocessing images by ARS, relative residual sum (RRS), farneback optical flow and TV-L1 optical flow, respectively.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-15-746985-g0006.tif"/>
</fig>
<p>The features extracted using Farneback optical flow and TV-L1 optical flow are disorganized, but the image extracted by residual sum methods can already distinguish many features. For example, surprise ME is easy to distinguish from other expressions. After preprocessing by the residual sum method, the features become initially orderly, but some of the ME are still mixed together. Therefore, further extraction of features through CNN can enhance the validity of features.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>In this study, we propose two novel preprocessing methods to solve ME recognition tasks with spatial-temporal feature extraction. These methods use the displacement residual sum of the unit pixels of the ME clip to extract a subtle motion feature. Through our experiment, it responds well to environmental change and subtle displacement. In addition, we propose a CGPO module, which divides the image into partial overlapping pictures of different precision and extracts features from different pictures. Hence, the model can focus on each facial local area, and then recognize the subtle movements of specific locations. Furthermore, we design CropNet which have a gradual way of increasing channels, features fusion module, and position embedding function.</p>
<p>In the experiment, we test the proposed two preprocessing methods and the designed network on the mixed dataset of MEGC2019 and five classes of ME on CASME II. The traditional manual method based on optical flow is labor-expensive and time-consuming, while the RRS and ARS preprocessing methods greatly improve the situation of frame redundancy and improve the recognition accuracy of each ME. In addition, the CGPO module can separate key parts of a person&#x00027;s face for more subtle feature extraction. In general, the method proposed in the study has better performance than the well-known method.</p>
<p>However, the proposed model does not belong to an end-to-end model, because it must go through the preprocessing method, which takes a certain amount of time to detect key points, align faces, crop, and calculate RRS and ARS. Therefore, in the future improvement, we will improve the method and model in this study into an end-to-end model.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The data analyzed in this study is subject to the following licenses/restrictions: this paper involves three databases (CASMEII, SMIC and SAMM). As each database involves human facial expressions, you need to apply for access. Requests to access these datasets should be directed to SMIC: <email>Xiaobai.Li&#x00040;oulu.fi</email>, SAMM: <email>M.Yap&#x00040;mmu.ac.uk</email>, CASMEII: <email>eagan-ywj&#x00040;foxmail.com</email>.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>YZ led the method design and experiment implementation. YZ and SL wrote sections of the manuscript. SL and ZC provided theoretical guidance, result review, and paper revision. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This publication of this paper was supported by the National Natural Science Foundation of China (no. 61872051), the Scientific and Technological Research Program of Chongqing Municipal Education Commission of China (no. KJ1600932), and the Graduate Innovation Fund of Chongqing University of Technology (no. clgycx20203123).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boucenna</surname> <given-names>S.</given-names></name> <name><surname>Gaussier</surname> <given-names>P.</given-names></name> <name><surname>Andry</surname> <given-names>P.</given-names></name> <name><surname>Hafemeister</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>A robot learns the facial expressions recognition and face/non-face discrimination through an imitation game</article-title>. <source>Int. J. Soc. Rob</source>. <volume>6</volume>, <fpage>633</fpage>&#x02013;<lpage>652</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-014-0245-z</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Davison</surname> <given-names>A. K.</given-names></name> <name><surname>Lansley</surname> <given-names>C.</given-names></name> <name><surname>Costen</surname> <given-names>N.</given-names></name> <name><surname>Tan</surname> <given-names>K.</given-names></name> <name><surname>Yap</surname> <given-names>M. H.</given-names></name></person-group> (<year>2016</year>). <article-title>Samm: a spontaneous micro-facial movement dataset</article-title>. <source>IEEE Trans. Affect. Comput</source>. <volume>9</volume>, <fpage>116</fpage>&#x02013;<lpage>129</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2016.2573832</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Li</surname> <given-names>L.-J.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Fei-Fei</surname> <given-names>L.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;Imagenet: a large-scale hierarchical image database,&#x0201D;</article-title> in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>248</fpage>&#x02013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ekman</surname> <given-names>P.</given-names></name></person-group> (<year>2009a</year>). <article-title>Lie catching and microexpressions</article-title>. <source>Philos. Decept</source>. <volume>1</volume>, <fpage>5</fpage>. <pub-id pub-id-type="doi">10.1093/acprof:oso/9780195327939.003.0008</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ekman</surname> <given-names>P.</given-names></name></person-group> (<year>2009b</year>). <source>Telling Lies: Clues to Deceit in the Marketplace, Politics, and Marriage (Revised Edition)</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>WW Norton &#x00026; Company</publisher-name>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ekman</surname> <given-names>P.</given-names></name> <name><surname>Friesen</surname> <given-names>W. V.</given-names></name></person-group> (<year>1971</year>). <article-title>Constants across cultures in the face and emotion</article-title>. <source>J. Pers. Soc. Psychol</source>. <volume>17</volume>, <fpage>124</fpage>. <pub-id pub-id-type="doi">10.1037/h0030377</pub-id><pub-id pub-id-type="pmid">5542557</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Farneb&#x000E4;ck</surname> <given-names>G.</given-names></name></person-group> (<year>2003</year>). <article-title>&#x0201C;Two-frame motion estimation based on polynomial expansion,&#x0201D;</article-title> in <source>Scandinavian Conference on Image Analysis</source> (<publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>).</citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Frank</surname> <given-names>M.</given-names></name> <name><surname>Herbasz</surname> <given-names>M.</given-names></name> <name><surname>Sinuk</surname> <given-names>K.</given-names></name> <name><surname>Keller</surname> <given-names>A.</given-names></name> <name><surname>Nolan</surname> <given-names>C.</given-names></name></person-group> (<year>2009</year>). <article-title>&#x0201C;I see how you feel: training laypeople and professionals to recognize fleeting emotions,&#x0201D;</article-title> in <source>The Annual Meeting of the International Communication Association</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Sheraton New Yorkpages</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>35</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gan</surname> <given-names>Y.</given-names></name> <name><surname>Liong</surname> <given-names>S.-T.</given-names></name> <name><surname>Yau</surname> <given-names>W.-C.</given-names></name> <name><surname>Huang</surname> <given-names>Y.-C.</given-names></name> <name><surname>Tan</surname> <given-names>L.-K.</given-names></name></person-group> (<year>2019</year>). <article-title>Off-apexnet on micro-expression recognition system</article-title>. <source>Signal Proc. Image Commun</source>. <volume>74</volume>, <fpage>129</fpage>&#x02013;<lpage>139</lpage>. <pub-id pub-id-type="doi">10.1016/j.image.2019.02.005</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>D.</given-names></name> <name><surname>Yun</surname> <given-names>S.</given-names></name> <name><surname>Heo</surname> <given-names>B.</given-names></name> <name><surname>Yoo</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Rexnet: Diminishing representational bottleneck on convolutional neural network</article-title>. <source>arXiv preprint</source> arXiv:2007.00992.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>S. J.</given-names></name> <name><surname>Li</surname> <given-names>J</given-names></name> <name><surname>Yap</surname> <given-names>H. M.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Spotting macro-and micro-expression intervals in long video sequences,&#x0201D;</article-title> in <source>15th IEEE International Conference on Automatic Face and Gesture Recognition</source> (<publisher-loc>Buenos Aires</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>742</fpage>&#x02013;<lpage>748</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heo</surname> <given-names>B.</given-names></name> <name><surname>Chun</surname> <given-names>S.</given-names></name> <name><surname>Oh</surname> <given-names>S. J.</given-names></name> <name><surname>Han</surname> <given-names>D.</given-names></name> <name><surname>Yun</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Adamp: slowing down the slowdown for momentum optimizers on scale-invariant weights,&#x0201D;</article-title> in <source>International Conference on Learning Representations, Vol</source>. <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Spontaneous facial micro-expression analysis using spatiotemporal local radon-based binary pattern,&#x0201D;</article-title> in <source>2017 International Conference on the Frontiers and Advances in Data Science (FADS)</source> (<publisher-loc>Xi&#x00027;an</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>159</fpage>&#x02013;<lpage>164</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Pietik&#x000E4;inen</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Spontaneous facial micro-expression analysis using spatiotemporal completed local quantized patterns</article-title>. <source>Neurocomputing</source> <volume>175</volume>, <fpage>564</fpage>&#x02013;<lpage>578</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2015.10.096</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Khor</surname> <given-names>H.-Q.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Liong</surname> <given-names>S.-T.</given-names></name> <name><surname>Phan</surname> <given-names>R. C.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Dual-stream shallow networks for facial micro-expression recognition,&#x0201D;</article-title> in <source>2019 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Taipei</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>36</fpage>&#x02013;<lpage>40</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Khor</surname> <given-names>H.-Q.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Phan</surname> <given-names>R. C. W.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Enriched long-term recurrent convolutional network for facial micro-expression recognition,&#x0201D;</article-title> in <source>2018 13th IEEE International Conference on Automatic Face &#x00026;Gesture Recognition (FG 2018)</source> (<publisher-loc>Xi&#x00027;an</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>667</fpage>&#x02013;<lpage>674</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.-H.</given-names></name> <name><surname>Kim</surname> <given-names>B.-G.</given-names></name> <name><surname>Roy</surname> <given-names>P. P.</given-names></name> <name><surname>Jeong</surname> <given-names>D.-M.</given-names></name></person-group> (<year>2019</year>). <article-title>Efficient facial expression recognition algorithm based on hierarchical deep neural network structure</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>41273</fpage>&#x02013;<lpage>41285</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2907327</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Le Ngo</surname> <given-names>A. C.</given-names></name> <name><surname>Liong</surname> <given-names>S.-T.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Phan</surname> <given-names>R. C.-W.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Are subtle expressions too sparse to recognize?&#x0201D;</article-title> in <source>2015 IEEE International Conference on Digital Signal Processing (DSP)</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1246</fpage>&#x02013;<lpage>1250</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Le Ngo</surname> <given-names>A. C.</given-names></name> <name><surname>Phan</surname> <given-names>R. C.-W.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Spontaneous subtle expression recognition: Imbalanced databases and solutions,&#x0201D;</article-title> in <source>Asian Conference on Computer Vision</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>33</fpage>&#x02013;<lpage>48</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le Ngo</surname> <given-names>A. C.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Phan</surname> <given-names>R. C.-W.</given-names></name></person-group> (<year>2016</year>). <article-title>Sparsity in dynamics of spontaneous subtle emotions: analysis and application</article-title>. <source>IEEE Trans. Affect. Comput</source>. <volume>8</volume>, <fpage>396</fpage>&#x02013;<lpage>411</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2016.2523996</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Pfister</surname> <given-names>T.</given-names></name> <name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Pietik&#x000E4;inen</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;A spontaneous micro-expression database: Inducement, collection and baseline,&#x0201D;</article-title> in <source>2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (fg)</source> (<publisher-loc>Shanghai</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Can micro-expression be recognized based on single apex frame?&#x0201D;</article-title> in <source>2018 25th IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Athens</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3094</fpage>&#x02013;<lpage>3098</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lies</surname> <given-names>T.</given-names></name></person-group> (<year>1992</year>). <source>Clues to Deceit in the Marketplace, Politics, and Marriage</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Norton</publisher-name>.</citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liong</surname> <given-names>S.-T.</given-names></name> <name><surname>Gan</surname> <given-names>Y.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Khor</surname> <given-names>H.-Q.</given-names></name> <name><surname>Huang</surname> <given-names>Y.-C.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Shallow triple stream three-dimensional cnn (ststnet) for micro-expression recognition,&#x0201D;</article-title> in <source>2019 14th IEEE International Conference on Automatic Face &#x00026;Gesture Recognition (FG 2019)</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liong</surname> <given-names>S. T.</given-names></name> <name><surname>See</surname> <given-names>J. S. Y.</given-names></name> <name><surname>Wong</surname> <given-names>K. S.</given-names></name> <name><surname>Phan</surname> <given-names>R. C. W.</given-names></name></person-group> (<year>2018</year>). <article-title>Less is more: micro-expression recognition from video using apex frame</article-title>. <source>Signal Proc. Image Commun</source>. <volume>62</volume>:<fpage>82</fpage>&#x02013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1016/j.image.2017.11.006</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>K.-H.</given-names></name> <name><surname>Jin</surname> <given-names>Q.-S.</given-names></name> <name><surname>Xu</surname> <given-names>H.-C.</given-names></name> <name><surname>Gan</surname> <given-names>Y.-S.</given-names></name> <name><surname>Liong</surname> <given-names>S.-T.</given-names></name></person-group> (<year>2021</year>). <article-title>Micro-expression recognition using advanced genetic algorithm</article-title>. <source>Signal Proc. Image Commun</source>. <volume>93</volume>:<fpage>116153</fpage>. <pub-id pub-id-type="doi">10.1016/j.image.2021.116153</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Du</surname> <given-names>H.</given-names></name> <name><surname>Zheng</surname> <given-names>L.</given-names></name> <name><surname>Gedeon</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;A neural micro-expression recognizer,&#x0201D;</article-title> in <source>2019 14th IEEE International Conference on Automatic Face &#x00026;Gesture Recognition (FG 2019)</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Yuan</surname> <given-names>X.</given-names></name> <name><surname>Gong</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>Z.</given-names></name> <name><surname>Fang</surname> <given-names>F.</given-names></name> <name><surname>Luo</surname> <given-names>Z.</given-names></name></person-group> (<year>2018</year>). <article-title>Conditional convolution neural network enhanced random forest for facial expression recognition</article-title>. <source>Pattern Recognit</source>. <volume>84</volume>, <fpage>251</fpage>&#x02013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2018.07.016</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.-J.</given-names></name> <name><surname>Zhang</surname> <given-names>J.-K.</given-names></name> <name><surname>Yan</surname> <given-names>W.-J.</given-names></name> <name><surname>Wang</surname> <given-names>S.-J.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Fu</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>A main directional mean optical flow feature for spontaneous micro-expression recognition</article-title>. <source>IEEE Trans. Affect. Comput</source>. <volume>7</volume>, <fpage>299</fpage>&#x02013;<lpage>310</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2015.2485205</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lukasik</surname> <given-names>M.</given-names></name> <name><surname>Bhojanapalli</surname> <given-names>S.</given-names></name> <name><surname>Menon</surname> <given-names>A.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Does label smoothing mitigate label noise?&#x0201D;</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>PMLR</publisher-loc>), <fpage>6448</fpage>&#x02013;<lpage>6458</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name></person-group> (<year>2009</year>). <article-title>A survey on transfer learning</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <volume>22</volume>, <fpage>1345</fpage>&#x02013;<lpage>1359</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2009.191</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Bi</surname> <given-names>T.</given-names></name> <name><surname>Shi</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;A novel apex-time network for cross-dataset micro-expression recognition,&#x0201D;</article-title> in <source>2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII)</source> (<publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sandler</surname> <given-names>M.</given-names></name> <name><surname>Howard</surname> <given-names>A.</given-names></name> <name><surname>Zhu</surname> <given-names>M.</given-names></name> <name><surname>Zhmoginov</surname> <given-names>A.</given-names></name> <name><surname>Chen</surname> <given-names>L.-C.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Mobilenetv2: inverted residuals and linear bottlenecks,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4510</fpage>&#x02013;<lpage>4520</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Yap</surname> <given-names>M. H.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>S.-J.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Megc 2019-the second facial micro-expressions grand challenge,&#x0201D;</article-title> in <source>2019 14th IEEE International Conference on Automatic Face &#x00026;Gesture Recognition (FG 2019)</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Zong</surname> <given-names>Y.</given-names></name> <name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Recognizing spontaneous micro-expression using a three-stream convolutional neural network</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>184537</fpage>&#x02013;<lpage>184551</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2960629</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Su</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Key facial components guided micro-expression recognition based on first amp; second-order motion,&#x0201D;</article-title> in <source>2021 IEEE International Conference on Multimedia and Expo (ICME)</source> (<publisher-loc>Shenzhen</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Visualizing data using t-sne</article-title>. <source>J. Mach. Learn. Res</source>. <volume>9</volume>, <fpage>2579</fpage>&#x02013;<lpage>2605</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Van Quang</surname> <given-names>N.</given-names></name> <name><surname>Chun</surname> <given-names>J.</given-names></name> <name><surname>Tokuyama</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Capsulenet for micro-expression recognition,&#x0201D;</article-title> in <source>2019 14th IEEE International Conference on Automatic Face &#x00026;Gesture Recognition (FG 2019)</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>See</surname> <given-names>J.</given-names></name> <name><surname>Phan</surname> <given-names>R. C.-W.</given-names></name> <name><surname>Oh</surname> <given-names>Y.-H.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Lbp with six intersection points: reducing redundant information in lbp-top for micro-expression recognition,&#x0201D;</article-title> in <source>Asian Conference on Computer Vision</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>525</fpage>&#x02013;<lpage>537</lpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xia</surname> <given-names>Z.</given-names></name> <name><surname>Hong</surname> <given-names>X.</given-names></name> <name><surname>Gao</surname> <given-names>X.</given-names></name> <name><surname>Feng</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>Spatiotemporal recurrent convolutional networks for recognizing spontaneous micro-expressions</article-title>. <source>IEEE Trans. Multimedia</source> <volume>22</volume>, <fpage>626</fpage>&#x02013;<lpage>640</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2019.2931351</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xia</surname> <given-names>Z.</given-names></name> <name><surname>Peng</surname> <given-names>W.</given-names></name> <name><surname>Khor</surname> <given-names>H.-Q.</given-names></name> <name><surname>Feng</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>Revealing the invisible with model and data shrinking for composite-database micro-expression recognition</article-title>. <source>IEEE Trans. Image Proc</source>. <volume>29</volume>, <fpage>8590</fpage>&#x02013;<lpage>8605</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2020.3018222</pub-id><pub-id pub-id-type="pmid">32845838</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>Hu</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Deep multi-path convolutional neural network joint with salient region attention for facial expression recognition</article-title>. <source>Pattern Recognit</source>. <volume>92</volume>, <fpage>177</fpage>&#x02013;<lpage>191</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2019.03.019</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>W.-J.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>S.-J.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Y.-J.</given-names></name> <name><surname>Chen</surname> <given-names>Y.-H.</given-names></name> <etal/></person-group>. (<year>2014a</year>). <article-title>Casme ii: an improved spontaneous micro-expression database and the baseline evaluation</article-title>. <source>PLoS ONE</source> <volume>9</volume>:<fpage>e86041</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0086041</pub-id><pub-id pub-id-type="pmid">24475068</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>W.-J.</given-names></name> <name><surname>Wang</surname> <given-names>S.-J.</given-names></name> <name><surname>Chen</surname> <given-names>Y.-H.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Fu</surname> <given-names>X.</given-names></name></person-group> (<year>2014b</year>). <article-title>&#x0201C;Quantifying micro-expressions with constraint local model and local binary pattern,&#x0201D;</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Zurich</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>296</fpage>&#x02013;<lpage>305</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>W.-J.</given-names></name> <name><surname>Wu</surname> <given-names>Q.</given-names></name> <name><surname>Liang</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>Y.-H.</given-names></name> <name><surname>Fu</surname> <given-names>X.</given-names></name></person-group> (<year>2013</year>). <article-title>How fast are the leaked facial expressions: the duration of micro-expressions</article-title>. <source>J. Nonverbal. Behav</source>. <volume>37</volume>, <fpage>217</fpage>&#x02013;<lpage>230</lpage>. <pub-id pub-id-type="doi">10.1007/s10919-013-0159-8</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zach</surname> <given-names>C.</given-names></name> <name><surname>Pock</surname> <given-names>T.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;A duality based approach for realtime tv-l 1 optical flow,&#x0201D;</article-title> in <source>Joint Pattern Recognition Symposium</source> (<publisher-loc>Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>214</fpage>&#x02013;<lpage>223</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Pietikainen</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <article-title>Dynamic texture recognition using local binary patterns with an application to facial expressions</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>29</volume>, <fpage>915</fpage>&#x02013;<lpage>928</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2007.1110</pub-id><pub-id pub-id-type="pmid">17431293</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>L.</given-names></name> <name><surname>Mao</surname> <given-names>Q.</given-names></name> <name><surname>Xue</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Dual-inception network for cross-database micro-expression recognition,&#x0201D;</article-title> in <source>2019 14th IEEE International Conference on Automatic Face &#x00026;Gesture Recognition (FG 2019)</source> (<publisher-loc>Lille</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
</ref-list>
 
</back>
</article>