<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2021.738167</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Hierarchical Spatiotemporal Electroencephalogram Feature Learning and Emotion Recognition With Attention-Based Antagonism Neural Network</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Pengwei</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Min</surname> <given-names>Chongdan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1443701/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Kangjia</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Xue</surname> <given-names>Wen</given-names></name>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Chen</surname> <given-names>Jingxia</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/969853/overview"/>
</contrib>
</contrib-group>
<aff><institution>School of Electronic Information and Artificial Intelligence, Shaanxi University of Science and Technology</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jane Zhen Liang, Shenzhen University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Ke Liu, Chongqing University of Posts and Telecommunications, China; Ying Yang, Facebook, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Jingxia Chen, <email>chenjx_sust@foxmail.com</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Perception Science, a section of the journal Frontiers in Neuroscience</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>738167</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2021 Zhang, Min, Zhang, Xue and Chen.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Zhang, Min, Zhang, Xue and Chen</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Inspired by the neuroscience research results that the human brain can produce dynamic responses to different emotions, a new electroencephalogram (EEG)-based human emotion classification model was proposed, named R2G-ST-BiLSTM, which uses a hierarchical neural network model to learn more discriminative spatiotemporal EEG features from local to global brain regions. First, the bidirectional long- and short-term memory (BiLSTM) network is used to obtain the internal spatial relationship of EEG signals on different channels within and between regions of the brain. Considering the different effects of various cerebral regions on emotions, the regional attention mechanism is introduced in the R2G-ST-BiLSTM model to determine the weight of different brain regions, which could enhance or weaken the contribution of each brain area to emotion recognition. Then a hierarchical BiLSTM network is again used to learn the spatiotemporal EEG features from regional to global brain areas, which are then input into an emotion classifier. Especially, we introduce a domain discriminator to work together with the classifier to reduce the domain offset between the training and testing data. Finally, we make experiments on the EEG data of the DEAP and SEED datasets to test and compare the performance of the models. It is proven that our method achieves higher accuracy than those of the state-of-the-art methods. Our method provides a good way to develop affective brain&#x2013;computer interface applications.</p>
</abstract>
<kwd-group>
<kwd>EEG</kwd>
<kwd>emotion recognition</kwd>
<kwd>spatiotemporal features</kwd>
<kwd>attention</kwd>
<kwd>antagonism neural network</kwd>
<kwd>BiLSTM</kwd>
</kwd-group>
<contract-num rid="cn001">61806118</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="10"/>
<equation-count count="26"/>
<ref-count count="43"/>
<page-count count="17"/>
<word-count count="12430"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>Emotion plays an important role in human life (<xref ref-type="bibr" rid="B30">Picard and Picard, 1997</xref>). Positive emotions may help improve the efficiency of our daily work, while negative emotions may affect our decision making, attention, and even health (<xref ref-type="bibr" rid="B30">Picard and Picard, 1997</xref>). Although it is easier for us to recognize emotions of other people from their facial expression or voice, it is still difficult for machines to do that (<xref ref-type="bibr" rid="B25">Li et al., 2019</xref>). In the past few years, emotion recognition by computer has attracted more and more researchers, and it has become a hot research topic in the field of affective computing and pattern recognition (<xref ref-type="bibr" rid="B31">Purnamasari et al., 2017</xref>). The emotion recognition methods can be based on speech signals, facial expression images, and physiological signals (<xref ref-type="bibr" rid="B5">Chen et al., 2019a</xref>). In recent years, EEG-based emotion recognition algorithms have been increasingly focused on by researchers.</p>
<p>While researching emotion recognition with EEG, we usually face two difficulties. One is how to obtain a discriminative feature representation method from original EEG signals, and the other is how to build an effective model to better improve the performance of emotion classification. Technically, EEG features can be extracted from the time domain, frequency domain, and time&#x2013;frequency domain (<xref ref-type="bibr" rid="B20">Jenke et al., 2014</xref>). For example, <xref ref-type="bibr" rid="B39">Zhang and Lee (2010)</xref> regarded the amplitude difference of symmetric electrodes in the time domain as the EEG feature of emotion recognition. <xref ref-type="bibr" rid="B26">Lin et al. (2010)</xref> studied the relationship between emotional state and brain activity, and extracted power spectral density, differential asymmetric power, and reasonable asymmetric power separately as features of EEG signals. <xref ref-type="bibr" rid="B12">Duan et al. (2013)</xref> extracted features by calculating the correlation coefficient between the features of each frequency band and their emotional labels. In the aspect of models, <xref ref-type="bibr" rid="B15">Garcia-Martinez et al. (2019)</xref> summarized the research results of applying nonlinear methods to EEG signal analysis in recent years. <xref ref-type="bibr" rid="B24">Li et al. (2018b)</xref> proposed a graph-regularized sparse linear regression model to make emotion classification and achieved better recognition performance. <xref ref-type="bibr" rid="B41">Zheng and Lu (2015)</xref> studied the key frequency bands and key brain regions of EEG signals, and proposed to use group sparse canonical correlation analysis algorithm (<xref ref-type="bibr" rid="B40">Zheng, 2017</xref>) for multichannel EEG-based emotion recognition.</p>
<p>With the development of artificial intelligence, deep learning has become very popular, and emotion classification based on deep learning has also continuously improved the performance of emotion recognition and, thus, has gradually become the dominant method. <xref ref-type="bibr" rid="B1">Alhagry et al. (2017)</xref> proposed an end-to-end LSTM-RNN network to learn the time dependence of EEG signals. <xref ref-type="bibr" rid="B23">Li et al. (2018a)</xref> considered the area shift of EEG data and used deep neural network to learn the difference between left and right hemispheres to narrow the distribution shift. <xref ref-type="bibr" rid="B34">Song et al. (2018)</xref> established a graph relationship based on multichannel EEG data, adjacency matrix to build the internal relationship between different EEG channels, and then used dynamical graph convolution network to extract features for emotion classification. <xref ref-type="bibr" rid="B32">Salama et al. (2018)</xref> used a three-dimensional convolutional neural network (3D-CNN) to recognize emotions from multichannel EEG data. The author of this paper has also proposed a deep CNN model (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>) to learn high-level discriminative feature representations from the combined features of the EEG signal in the time-frequency domain. In <xref ref-type="bibr" rid="B4">Chen et al. (2019b)</xref>, a hierarchical bidirectional LSTM model based on attention mechanism was proposed to reduce the influence of long-term instability of EEG sequences on emotion recognition.</p>
<p>Although many EEG emotion recognition methods have emerged recently, there are still some problems that needs to be further studied. One of the problems is how to obtain effective high-level features from the original EEG signals automatically. Most researchers often extract some time or frequency statistical EEG features manually combined with classic machine learning algorithms to make emotion classification. However, feature engineering needs to consume a lot of computation resources and time. It is expected to automatically learn more prominent spatiotemporal features with less feature engineering. The second question is which brain area contributes more to human emotion recognition, and how to use the distribution information of different brain areas to improve recognition performance. The latest researches (<xref ref-type="bibr" rid="B13">Etkin et al., 2011</xref>; <xref ref-type="bibr" rid="B27">Lindquist and Barrett, 2012</xref>) have shown that human emotions are closely related to multiple areas of the cerebral cortex, such as the orbitofrontal cortex, ventromedial prefrontal cortex, amygdala, and so on. Therefore, the contribution of EEG signals associated with each brain area is different. If the spatial information of different brain regions can be used, it is expected to provide help in understanding human emotions (<xref ref-type="bibr" rid="B18">Heller and Nitscke, 1997</xref>; <xref ref-type="bibr" rid="B10">Davidson, 2000</xref>; <xref ref-type="bibr" rid="B28">Lindquist et al., 2012</xref>). The third question is how to enhance the emotion recognition performance by using time series information in each brain area, as EEG signals are dynamic time series carrying important emotion dynamics, which is effective to identify human emotions.</p>
<p>Literature (<xref ref-type="bibr" rid="B26">Lin et al., 2010</xref>; <xref ref-type="bibr" rid="B39">Zhang and Lee, 2010</xref>; <xref ref-type="bibr" rid="B12">Duan et al., 2013</xref>; <xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>; <xref ref-type="bibr" rid="B40">Zheng, 2017</xref>; <xref ref-type="bibr" rid="B24">Li et al., 2018b</xref>; <xref ref-type="bibr" rid="B15">Garcia-Martinez et al., 2019</xref>) has proven that EEG signals in different brain regions have different contributions to emotion recognition. Literature (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>; <xref ref-type="bibr" rid="B23">Li et al., 2018a</xref>; <xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>; <xref ref-type="bibr" rid="B34">Song et al., 2018</xref>; <xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>) found that either deep CNN model or the bidirectional long- and short-term memory (BiLSTM) model combined with attention mechanism could hierarchically extract deep temporal and spatial context of EEG signals. Inspired by these two aspects and neuroscience research basis (<xref ref-type="bibr" rid="B18">Heller and Nitscke, 1997</xref>; <xref ref-type="bibr" rid="B10">Davidson, 2000</xref>; <xref ref-type="bibr" rid="B13">Etkin et al., 2011</xref>; <xref ref-type="bibr" rid="B27">Lindquist and Barrett, 2012</xref>; <xref ref-type="bibr" rid="B28">Lindquist et al., 2012</xref>), this paper proposes a new emotion computing model called R2G-ST-BiLSTM, which is used to solve the above three main problems. Its core idea is to extract the EEG spatial temporal dynamics associated with human emotions from local and global brain areas. Specifically, the R2G-ST-BiLSTM model contains two two-layer neural networks, in space and time domain, respectively, and features are learned hierarchically from region to global (R2G) to grasp more discriminative spatiotemporal EEG features related to human emotions. The proposed R2G-ST-BiLSTM model consists of three parts:</p>
<list list-type="simple">
<list-item>
<label>(1)</label>
<p>Feature learning module. It uses the bidirectional long- and short-term memory (BiLSTM) network to learn the hierarchical spatiotemporal EEG characteristics within and between each brain region. In order to better judge the effect of different brain regions on emotion recognition, this paper introduces the regional attention mechanism to learn a set of weights, which represent the contributions of different brain regions.</p>
</list-item>
<list-item>
<label>(2)</label>
<p>Emotion classifier. The purpose of this module is to predict emotion category based on EEG spatiotemporal features obtained by feature learning module. At the same time, it also guides the whole neural network to learn more discriminative EEG features for emotion classification.</p>
</list-item>
<list-item>
<label>(3)</label>
<p>Domain discriminator. This module aims to decrease the domain offset between the training EEG data and the testing EEG data through introducing a discriminator, so that the hierarchical feature learning module can produce EEG features with more emotional discrimination and stronger domain adaptability.</p>
</list-item>
</list>
<p>Through collaborative work of the above three modules, the R2G-ST-BiLSTM model can learn EEG features with better discrimination ability and domain robustness simultaneously, thus, further improving human emotion recognition performance. Overall, there are three main contributions in our work:</p>
<list list-type="simple">
<list-item>
<label>&#x2022;</label>
<p>Inspired by neuroscience, we propose a new hierarchical spatiotemporal EEG feature learning model, which obtains spatiotemporal emotional information from EEG data within and between each cerebral region.</p>
</list-item>
<list-item>
<label>&#x2022;</label>
<p>Proposes an attention weighted model to estimate the contribution of each cerebral region to the different affections of humans. The influence of the most dependent cerebral region is enhanced by the learned weight, and the impact of the less dependent region was reduced as well.</p>
</list-item>
<list-item>
<label>&#x2022;</label>
<p>Proposes a domain discriminator to work on antagonism with the classifier to improve the adaptability of the R2G-ST-BiLSTM model.</p>
</list-item>
</list>
</sec>
<sec id="S2">
<title>Method Based on the R2G-ST-BiLSTM Model</title>
<p>Traditional one-way LSTM network (<xref ref-type="bibr" rid="B19">Hochreiter and Schmidhuber, 1997</xref>) has a special structure that is different from the simple recurrent neural network (RNN) (<xref ref-type="bibr" rid="B17">Graves et al., 2013</xref>) and is more capable of dealing with the frequent dependence of the sample sequence. Its special &#x201C;gate&#x201D; structures enable LSTM to retain significant data information and forget unnecessary redundant information (<xref ref-type="bibr" rid="B37">Yan et al., 2017</xref>). However, one disadvantage of this network is that it only uses the context-related information that happened before. The BiLSTM network can process data by using separate hidden layers in two directions (<xref ref-type="bibr" rid="B2">Bottou, 2010</xref>). Because the BiLSTM network can obtain long-term contextual information in both forward and backward orientation, it is better than the traditional one-way LSTM network for modeling time series. Because EEG data related to each channel in each brain region are in time series with the same dimension, therefore, BiLSTM can be used to extract the deep spatiotemporal context features of EEG data from the local brain regions to the global brain.</p>
<p>In this section, we will introduce the framework of the R2G-ST-BiLSTM model in detail and explain the specific application of EEG signals for emotion recognition methods and procedures. <xref ref-type="fig" rid="F1">Figure 1</xref> shows the framework of the R2G-ST-BiLSTM model. It consists of three main modules, which are feature extractors, classifiers, and discriminators.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The model of the R2G-ST-BiLSTM network. Feature learning is processing from regional brain to global brain, respectively, in the spatial and temporal flows. Spatial flow learns the relationship between brain regions in different layers, while temporal flow learns the emotion-related EEG dynamic from the time series of each brain region.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g001.tif"/>
</fig>
<sec id="S2.SS1">
<title>Spatial Feature Extraction</title>
<p>First, we divide the EEG sequence into several equal-length segments. Then a set of manual features is extracted from the EEG segments corresponding to each electrode. For example, the differential entropy feature (DE feature) is extracted from &#x03B4;(1&#x223C;4 hz), &#x03B8;(5&#x223C;8 hz), &#x03B1;(9&#x223C;14 hz), &#x03B2;(15&#x223C;3 0hz), and &#x03B3;(31&#x223C;50 hz) (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>). In addition, to capture dynamic time information from input EEG sequence, every five adjacent EEG segments make up one EEG sample, and each EEG sample is represented by a tensor of its manual feature.</p>
<p>Let <italic>S</italic> = [<italic>s</italic><sub>1</sub>, <italic>s</italic><sub>2</sub>, <italic>s</italic><sub><italic>T</italic>&#x2212;1</sub>, <italic>s</italic><sub><italic>T</italic></sub>]&#x03F5;&#x211D;<sup><italic>d</italic> &#x00D7; <italic>n</italic> &#x00D7; <italic>T</italic></sup> represent an EEG sample, where <italic>s<sub>i</sub></italic> represents the feature data extracted from the divided <italic>i</italic>-th segment of EEG, shown in the bottom blue rectangle of <xref ref-type="fig" rid="F1">Figure 1</xref>, <italic>d</italic> is the number of EEG features per channel, <italic>n</italic> is the number of channels, and <italic>T</italic> is the number of segments per EEG sample. <xref ref-type="fig" rid="F1">Figure 1</xref> shows that when extracting spatial features, each sample includes a regional feature extraction layer and a global feature extraction layer to gradually learn high-level semantic features from local to global.</p>
<p><xref ref-type="fig" rid="F2">Figure 2</xref> shows the specific feature learning process of the EEG data <italic>s<sub>i</sub></italic>. At first, the channels of <italic>s<sub>i</sub></italic> are grouped into different areas according to the spatial position of the brain electrodes. The number of electrodes in each brain area varies due to the different functions of each brain area, thereby generating a set of regional manual feature vectors in each brain region. Then these manual feature vectors are input into the equal number of BiLSTM networks to learn the local abstract features of each region. After learning the regional deep features, the region attention mechanism is introduced to learn a set of weights that represent the significance of each region. Finally, at the top of <xref ref-type="fig" rid="F1">Figure 1</xref>, the extracted weighted feature of each region is input into another set of BiLSTM networks to further extract the global emotional semantic features.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Regional to global spatial feature learning process consisting of regional feature learning layer, dynamic weight layer, and global feature learning layer.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g002.tif"/>
</fig>
<p>(1) <italic>Regional feature extraction layer.</italic> Let <italic>x</italic><sub><italic>ij</italic></sub> represent the manual feature vector of the <italic>j-</italic>th EEG channel, so <italic>s</italic><sub><italic>i</italic></sub> = [<italic>x</italic><sub><italic>i</italic>1</sub>, <italic>x</italic><sub><italic>i</italic>2</sub>, &#x2026;, <italic>x</italic><sub><italic>in</italic></sub>]&#x03F5;&#x211D;<sup><italic>d</italic>&#x00D7;<italic>n</italic></sup>. Then according to the related electrodes, <italic>n</italic> channels of <italic>s<sub>i</sub></italic> are divided into different groups: each group of channels belongs to a cerebral area, and each area is expressed as: brain area 1: <inline-formula><mml:math id="INEQ3"><mml:mrow><mml:mrow><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="7.5pt">,</mml:mo></mml:mrow></mml:math></inline-formula>brain area 2: <inline-formula><mml:math id="INEQ4"><mml:mrow><mml:mrow><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="9.2pt">,</mml:mo></mml:mrow></mml:math></inline-formula>brain area n: <inline-formula><mml:math id="INEQ5"><mml:mrow><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>, where <italic>N</italic> is the quantity of cerebral areas, <italic>n<sub>j</sub></italic> is the quantity of the <italic>j-</italic>th cerebral area of the channels, and <italic>n</italic><sub>1</sub> + + <italic>n</italic><sub><italic>N</italic></sub> = <italic>n</italic>. Furthermore, we adjust the column order of <italic>s<sub>i</sub></italic>, which is represented as a new matrix <inline-formula><mml:math id="INEQ7"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>s</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>. The submatrix <inline-formula><mml:math id="INEQ8"><mml:mrow><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> represents cerebral area <italic>j</italic>, and per column of <inline-formula><mml:math id="INEQ9"><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:math></inline-formula> corresponds to an EEG channel of this area. The spatial relationship of the brain area can be modeled by a BiLSTM working on the <inline-formula><mml:math id="INEQ10"><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:math></inline-formula> matrix to extract the advanced features of each region, which process is expressed as:</p>
<disp-formula id="S2.E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mrow><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:msubsup><mml:mo rspace="5.8pt" stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>d</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:mrow><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>R</mml:mi><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:msubsup><mml:mo rspace="5.8pt" stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>d</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where &#x2131;(&#x22C5;) represents the BiLSTM operation, <inline-formula><mml:math id="INEQ12"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mi>j</mml:mi></mml:msubsup><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mi>d</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> represents hidden vectors output by the <italic>k</italic>th forward and backward hidden units of the BiLSTM, and <italic>d<sub>r</sub></italic> represents the dimensions of the output state vectors of each hidden unit in the BiLSTM. At last, the state vector outputs by the last hidden unit of each BiLSTM are connected as the local deep features of all regions, which are expressed as follows:</p>
<disp-formula id="S2.E3"><label>(3)</label><mml:math id="M3"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>d</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>For simplicity, each BiLSTM model in this part is initialized and fit jointly, and the hyperparameters are shared with each other.</p>
<p>(2) <italic>Attention-based brain region weighting layer.</italic> Neuroscience-related research shows that different brain areas respond to different types of emotions. Therefore, EEG signals from diverse brain areas have different contributions to emotion classification. To emphasize the role of the different brain area electrodes in EEG emotion recognition, we introduce a weighting layer based on attention mechanism. Expressed by <italic>W</italic> = {<italic>w</italic><sub><italic>ij</italic></sub>}, it can characterize the significance of channels in different areas. After that, the local deep features of all areas are expressed by <inline-formula><mml:math id="INEQ14"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup></mml:math></inline-formula> as follows:</p>
<disp-formula id="S2.E4"><label>(4)</label><mml:math id="M4"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup><mml:mi>W</mml:mi></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E5"><label>(5)</label><mml:math id="M5"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>W</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>U</mml:mi></mml:mpadded><mml:mrow><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msup><mml:mi>b</mml:mi><mml:mi>r</mml:mi></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E6"><label>(6)</label><mml:math id="M6"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mfrac><mml:mrow><mml:mpadded width="+3.3pt"><mml:mtext>exp</mml:mtext></mml:mpadded><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mrow><mml:mpadded width="+3.3pt"><mml:mtext>exp</mml:mtext></mml:mpadded><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>U</italic> and <italic>V</italic> are learnable transpositional matrices, <italic>b<sup>r</sup></italic> represents the deviation, and <italic>e</italic> represents an N-dimensional vector whose elements are 1, that is, <italic>e</italic> = [1, 1, &#x2026;, 1]<sup><italic>T</italic></sup>. The matrix <italic>W</italic> is normalized across the columns, so that its values are limited to non-negative by formula (6). The larger the <italic>w</italic><sub><italic>ij</italic></sub> value obtained, the more important the <italic>j-</italic>th brain area is for emotion recognition.</p>
<p>(3) <italic>Global feature extraction layer.</italic> To further capture the potential global structural information on the basis of the learned local deep feature <inline-formula><mml:math id="INEQ17"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup></mml:math></inline-formula>, we use another BiLSTM network with <italic>N</italic> hidden units to extract global spatial features.</p>
<disp-formula id="S2.E7"><label>(7)</label><mml:math id="M7"><mml:mrow><mml:mrow><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>r</mml:mi></mml:msubsup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>d</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where, <inline-formula><mml:math id="INEQ18"><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> represents the hidden vector output by the <italic>k-</italic>th forward and backward hidden unit of the BiLSTM network, and <italic>d<sub>g</sub></italic> is the dimension of the output state vector of each hidden unit. Next, input the vector sequence <inline-formula><mml:math id="INEQ19"><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula>,&#x2026;,<inline-formula><mml:math id="INEQ20"><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> into a fully connected layer to learn a new compressed feature vector with the following formulas:</p>
<disp-formula id="S2.E8"><label>(8)</label><mml:math id="M8"><mml:mrow><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>j</mml:mi><mml:mi>K</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mi>g</mml:mi></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo rspace="7.5pt">,</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E9"><label>(9)</label><mml:math id="M9"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>K</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="INEQ21"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>P</mml:mi><mml:mi>g</mml:mi></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:msub><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>K</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>N</mml:mi></mml:mpadded><mml:mo>=</mml:mo><mml:mrow><mml:mi/><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> denotes a projection matrix, <italic>b<sup>g</sup></italic> denotes deviation, <italic>K</italic> denotes the length of the compressed sequence, and &#x03C3;(&#x22C5;) is a nonlinear function. Thus, the global deep feature <inline-formula><mml:math id="INEQ24"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> related to the manual feature matrix <italic>s<sub>i</sub></italic> of the <italic>i-</italic>th EEG segment is finally obtained.</p>
</sec>
<sec id="S2.SS2">
<title>Temporal Feature Extraction</title>
<p>Let <inline-formula><mml:math id="INEQ25"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> represent the state vector output by the <italic>j</italic>-th brain region of the <italic>i</italic>-th EEG manual feature matrix <italic>s<sub>i</sub></italic> through the last hidden unit of the BiLSTM network, then the time series of each brain region feature can be expressed as:</p>
<disp-formula id="S2.E10"><label>(10)</label><mml:math id="M10"><mml:mrow><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x225C;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mn>1</mml:mn><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mn>2</mml:mn><mml:mn>1</mml:mn></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mi>T</mml:mi><mml:mn>1</mml:mn></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E11"><label>(11)</label><mml:math id="M11"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>r</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msup><mml:mo>&#x225C;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mn>2</mml:mn><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>In this way, the columns of the feature matrix <inline-formula><mml:math id="INEQ26"><mml:mrow><mml:msup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>r</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> constitute the time series of the feature vectors related to the <italic>j</italic>-th brain region. Therefore, a BiLSTM network can be applied again to learn the temporal context between these eigenvector sequence:</p>
<disp-formula id="S2.E12"><label>(12)</label><mml:math id="M12"><mml:mtable><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">~</mml:mo></mml:mover><mml:mrow><mml:mi>r</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">]</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mn>11</mml:mn><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mspace width="2.5em"/><mml:mrow><mml:mrow><mml:mi/><mml:mo lspace="0pt" rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="INEQ27"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mi>Z</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></inline-formula> represents the regional temporal feature matrix related to the <italic>j</italic>-th brain region, and <italic>d</italic><sub><italic>rt</italic></sub> is the dimension of each hidden unit state vector in the regional temporal BiLSTM network. Take the output <inline-formula><mml:math id="INEQ28"><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of the last hidden unit of the BiLSTM network in each brain area as the learned temporal feature of this brain area, and then get the final temporal feature <italic>z<sup>rt</sup></italic> of all brain areas, which is expressed as:</p>
<disp-formula id="S2.E13"><label>(13)</label><mml:math id="M13"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>In addition, to explore the time context on the basis of matrix <inline-formula><mml:math id="INEQ30"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula>, we convert the columns of <inline-formula><mml:math id="INEQ31"><mml:msubsup><mml:mover accent="true"><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> to a new sequence, which is represented by <inline-formula><mml:math id="INEQ32"><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula>:</p>
<disp-formula id="S2.E14"><label>(14)</label><mml:math id="M14"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>K</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>Set up <inline-formula><mml:math id="INEQ33"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>Z</mml:mi><mml:mi>g</mml:mi></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mn>1</mml:mn><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>T</mml:mi><mml:mi>g</mml:mi></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>. Then a BiLSTM network with <italic>T</italic> hidden units is used to learn the global temporal feature <italic>Z<sup>gt</sup></italic>:</p>
<disp-formula id="S2.E15"><label>(15)</label><mml:math id="M15"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mi>g</mml:mi></mml:msup><mml:mo rspace="5.8pt" stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:msup><mml:mi>&#x211D;</mml:mi><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>d</italic><sub><italic>gt</italic></sub> denotes the size of the hidden state vector of the global temporal BiLSTM network, and the output <inline-formula><mml:math id="INEQ35"><mml:msubsup><mml:mi>z</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> of the last hidden unit is taken as the learned global temporal feature. Finally, by concatenating <italic>z<sup>rt</sup></italic> with <inline-formula><mml:math id="INEQ37"><mml:msubsup><mml:mi>z</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, the optimal feature vector <italic>z<sup>rg</sup></italic> of the EEG sample <italic>S</italic> (composed of <italic>T</italic> EEG fragments) is obtained, which contains complex temporal context information, and its expression is:</p>
<disp-formula id="S2.E16"><label>(16)</label><mml:math id="M16"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
</sec>
<sec id="S2.SS3">
<title>Classifier and Discriminator</title>
<p>For the final eigenvector <italic>z<sup>rg</sup></italic> input to this layer, a simple linear transformation method can be used to recognize the human emotional type of the input EEG data <italic>S</italic> as the following formula:</p>
<disp-formula id="S2.E17"><label>(17)</label><mml:math id="M17"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>Y</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mi>Q</mml:mi><mml:msup><mml:mi>z</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>b</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>Q</italic> and <italic>b<sub>c</sub></italic>, respectively, denotes the projection matrix and deviation. <italic>C</italic> is number of emotional categories. The element of the transformation result <italic>Y</italic> is input into a softmax function to predict the emotion category:</p>
<disp-formula id="S2.E18"><label>(18)</label><mml:math id="M18"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:mi>S</mml:mi><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>max</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>i</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>C</mml:mi></mml:msubsup><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mfrac><mml:mo stretchy="false">|</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>P</italic>(<italic>c</italic>|<italic>X</italic>) represents the probability the input EEG data <italic>S</italic> is predicted to be the emotion of type <italic>c</italic>.</p>
<p>Supposing the training set of the model is composed of <italic>M</italic> EEG data, which is expressed by matrix <inline-formula><mml:math id="INEQ41"><mml:mrow><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>i</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>M</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>. The loss function of the emotion classifier can be expressed as:</p>
<disp-formula id="S2.E19"><label>(19)</label><mml:math id="M19"><mml:mrow><mml:mrow><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>S</mml:mi><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:msubsup><mml:mi mathvariant="normal">&#x2026;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>i</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:munderover><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>c</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>C</mml:mi></mml:munderover></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03C6;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo rspace="5.8pt" stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>l<sub>i</sub></italic> represents the real label of the <inline-formula><mml:math id="INEQ42"><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup></mml:math></inline-formula> sample, and &#x03B8;<sub><italic>f</italic></sub> and &#x03B8;<sub><italic>c</italic></sub> represent the learning parameters. &#x03C6;(<italic>l</italic><sub><italic>i</italic></sub>, <italic>c</italic>) is expressed as:</p>
<disp-formula id="S2.E20"><label>(20)</label><mml:math id="M20"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03C6;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable displaystyle="true" rowspacing="0pt"><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo rspace="7.5pt">,</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mpadded width="+5pt"><mml:mi>f</mml:mi></mml:mpadded><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>l</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mpadded></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mrow><mml:mn>&#x2005;0</mml:mn><mml:mo rspace="7.5pt">,</mml:mo><mml:mrow><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mi/></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>From formulas (19) and (20), it can be concluded that by minimizing the loss function <inline-formula><mml:math id="INEQ46"><mml:mrow><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mn>2</mml:mn><mml:mi>S</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>, the emotion category of each training sample can be correctly predicted to the maximum extent.</p>
<p>Let <italic>S</italic><sub><italic>test</italic></sub> represent a test sample, and the emotion label of <italic>S</italic><sub><italic>test</italic></sub> is determined by the formula:</p>
<disp-formula id="S2.E21"><label>(21)</label><mml:math id="M21"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:mi>c</mml:mi></mml:munder><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mi mathvariant="normal">P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">c</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>c</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo stretchy="false">}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>l</italic><sub><italic>test</italic></sub> represents the predicted label of the test sample <italic>S</italic><sub><italic>test</italic></sub>.</p>
<p>When performing prediction, the EEG samples for the training and testing data may be from various subjects and even different experiments. Based on this, the recognition model learned by using the training data may not have a high recognition accuracy for the test data. To optimize the generalization ability of the model, a discriminator is introduced to work collaboratively with the classifier to learn features with strong emotion discrimination and domain invariance.</p>
<p>Specifically, suppose that <inline-formula><mml:math id="INEQ47"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>D</mml:mi><mml:mi>S</mml:mi></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:msub><mml:mi>M</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mi>S</mml:mi></mml:msubsup><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the dataset of the source domain, and <inline-formula><mml:math id="INEQ48"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mi>D</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:msub><mml:mi>M</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mi>T</mml:mi></mml:msubsup><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the dataset of the target domain, where <italic>M</italic><sub>1</sub> and <italic>M</italic><sub>2</sub> are their sample number. To alleviate the domain difference, the loss function of the discriminator is defined as:</p>
<disp-formula id="S2.E22"><label>(22)</label><mml:math id="M22"><mml:mrow><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>i</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:munderover><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>|</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:munderover><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo rspace="7.5pt">.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>Here, <inline-formula><mml:math id="INEQ49"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>|</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> is the probability that EEG sample <inline-formula><mml:math id="INEQ50"><mml:msubsup><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>S</mml:mi></mml:msubsup></mml:math></inline-formula> is classified into the source domain, <inline-formula><mml:math id="INEQ51"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> is the probability that EEG sample <inline-formula><mml:math id="INEQ52"><mml:msubsup><mml:mi>S</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup></mml:math></inline-formula> is classified into the target domain, and &#x03B8;<sub><italic>d</italic></sub> is the parameter. The discriminator enables this model to learn the domain-invariant features gradually.</p>
</sec>
<sec id="S2.SS4">
<title>Optimization of the Bidirectional Long- and Short-Term Memory Neural Network From Region to Global Brain Model</title>
<p>The previous description indicates that through minimizing formula (19) and maximizing formula (22), domain difference can be reduced and better domain invariant characteristics can be learned. Therefore, we redefine the total loss function of R2G-ST-BiLSTM model as:</p>
<disp-formula id="S2.E23"><label>(23)</label><mml:math id="M23"><mml:mrow><mml:mrow><mml:mrow><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>To optimize our model, we need to find the best parameters that minimize the new loss function &#x2112;(<italic>S<sup>S</sup></italic>, <italic>S<sup>T</sup></italic>|&#x03B8;<sub><italic>f</italic></sub>, &#x03B8;<sub><italic>c</italic></sub>, &#x03B8;<sub><italic>d</italic></sub>). By minimizing &#x2112;<sub><italic>c</italic></sub>(<italic>S<sup>S</sup></italic>; &#x03B8;<sub><italic>f</italic></sub>, &#x03B8;<sub><italic>c</italic></sub>) and maximizing &#x2112;<sub><italic>d</italic></sub>(<italic>S<sup>S</sup></italic>, <italic>S<sup>T</sup></italic>; &#x03B8;<sub><italic>f</italic></sub>, &#x03B8;<sub><italic>d</italic></sub>) synchronously and iteratively, the optimal parameters of &#x2112;(<italic>S<sup>S</sup></italic>, <italic>S<sup>T</sup></italic>|&#x03B8;<sub><italic>f</italic></sub>, &#x03B8;<sub><italic>c</italic></sub>, &#x03B8;<sub><italic>d</italic></sub>) can be obtained. Specifically, the stochastic gradient descent (SGD) algorithm (<xref ref-type="bibr" rid="B38">Yu et al., 2015</xref>) is used to find the optimal model parameters:</p>
<disp-formula id="S2.E24"><label>(24)</label><mml:math id="M24"><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo rspace="5.8pt" stretchy="false">|</mml:mo><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mpadded width="+3.3pt"><mml:mi>g</mml:mi></mml:mpadded><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:munder><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mi>d</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E25"><label>(25)</label><mml:math id="M25"><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mover accent="true"><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mi>d</mml:mi></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mpadded width="+3.3pt"><mml:mi>g</mml:mi></mml:mpadded><mml:mrow><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:munder><mml:mrow><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mi>f</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>The feature extractor can learn to obtain emotional discriminative features by minimizing the loss function &#x2112;<sub><italic>c</italic></sub>. Meanwhile, it extracts domain invariant features by maximizing the loss function &#x2112;<sub><italic>d</italic></sub>. When obtaining the optimal parameters of the R2G-ST-BiLSTM model, we also introduced a gradient reverse layer (GRL) (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>), which performs gradient sign reversal when performing backward propagation operation and enables the discriminator to transform the maximization problem into a minimization problem, so that SGD can be used for parameter optimization. The parameter updating can be expressed as:</p>
<disp-formula id="S2.E26"><label>(26)</label><mml:math id="M26"><mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>&#x2190;</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi><mml:mfrac><mml:mrow><mml:mo>&#x2202;</mml:mo><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>&#x2202;</mml:mo><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>&#x2190;</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi><mml:mfrac><mml:mrow><mml:mo>&#x2202;</mml:mo><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi class="ltx_font_mathcaligraphic">&#x2112;</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>&#x2202;</mml:mo><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B8;</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where &#x03B1; is the learning rate.</p>
</sec>
<sec id="S2.SS5">
<title>Configuration and Training of Bidirectional Long- and Short-Term Memory Neural Network From Region to Global Brain Model</title>
<p>The proposed model is implemented in TensorFlow framework on a NVIDIA Titan &#x00D7; Pascal GPU-equipped work station and trained from scratch in a fully supervised manner. When training the whole model, we define a search space to find the optimal model parameters. The search space includes the hidden_layers (one to three layers), hidden_size (32, 64, 128, and 256), batch_size (30, 60, 80, and 120), learning_rate (0.1, 0.01, 0.001, and 0.0001), dropout (0.5, 0.6, and 0.7), and epochs (100, 200, 300, and 500). The search space was defined to balance the trade-off between a deeper architecture and limited training samples. For simplicity, each BiLSTM model is initialized and fit jointly, their hyperparameters are shared with each other, the hidden_size of the single-layer perception network used to learn the attention weight of each brain region is 128, the hidden_size of the full connection layer for learning the compressed global brain feature is 64, all hidden layers use ReLU activation function for faster approximation, all BiLSTM models are trained using SGD with AdaGrad optimizer, and the maximum training iteration was set to be 10,000. For searching each hyper parameter, we only adjust one hyperparameter in a defined search space and fix others each time. When observing that there is no growth trend of the accuracy on training and validation sets, we can judge to stop the training process in advance, as shown in <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F6">6</xref>. Finally, we select the best model that produces the highest accuracy on the validation dataset.</p>
<p>Through this fine-tuning process, the selected best hidden_layers is 2, the hidden_size of <italic>d<sub>r</sub></italic>, <italic>d<sub>g</sub></italic>, <italic>d</italic><sub><italic>rt</italic></sub>, and <italic>d</italic><sub><italic>gt</italic></sub> is consistently 64, learning rate is 0.001, batch-size is 120, and epochs is 200. All parameters and offsets are initialized with randomly assigned nonzero regularization float. For cross-subject experiment on DEAP dataset, the total number of parameters in the whole model is about 50,156, which is larger than the total number of training samples. To prevent the overfitting of the model, a dropout layer is added after the first full connection layer of each BiLSTM, and the selected optimal dropout is 0.7.</p>
</sec>
</sec>
<sec id="S3">
<title>Experiments and Results</title>
<sec id="S3.SS1">
<title>Dataset and Preprocessing</title>
<p>To evaluate the proposed method, we make extensive experiments on the DEAP (<xref ref-type="bibr" rid="B22">Koelstra et al., 2012</xref>) dataset, which come from the Queen Mary University of London and is publicly and freely available for research on emotion recognition. This dataset records EEG, EMG, ECG, and other types of physiological signals induced by 32 subjects watching 40 music videos with different emotional tendencies. The emotion labels are evaluated with 1&#x2013;9 consecutive values in four emotional dimensions of arousal, valence, preference, and dominance. In our research, we just take the EEG signals of each subject in 32 channels and 60 s from the DEAP dataset for study. The electrodes are positioned according to the 10&#x2013;20 system. The sampling frequency is reduced to 128 Hz. For other artifacts, a 4- to 45-Hz bandpass filter is used for data filtering, and then blind source separation is used to remove the electro-oculogram (EOG) interference.</p>
<p>According to the spatial distribution of EEG electrodes, 32 electrodes are divided into 12 regions, that is, the number of brain regions <italic>N</italic> is 12, and each region contains at least two electrodes. We divided the 32 electrodes into 12 clusters or brain regions, where the electrodes of the same color belong to the same region, as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. The electrodes contained in each brain region and the size of the corresponding manual feature set are listed in <xref ref-type="table" rid="T1">Table 1</xref>. In the DEAP database, there are 32 subjects, and each subject takes a 40-trial EEG data acquisition experiment. Each experiment collects 60 &#x00D7; 128 = 7,680 EEG records and emotional labels induced by watching videos for 60 s.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>The 32 electrodes are divided into 12 clusters, among which the electrodes of the same color belong to the same brain area.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g003.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Electroencephalogram electrodes and data size associated with each brain area.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Brain region</td>
<td valign="top" align="center" colspan="2">DEAP dataset<hr/></td>
<td valign="top" align="center" colspan="2">SEED dataset<hr/></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">Electrode name</td>
<td valign="top" align="center">Data size (d &#x00D7; n<sub><italic>j</italic></sub>)</td>
<td valign="top" align="left">Electrode name</td>
<td valign="top" align="center">Data size (d &#x00D7; n<sub><italic>j</italic></sub>)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Pre-frontal</td>
<td valign="top" align="left">FP1, FP2, AF3, AF4</td>
<td valign="top" align="center">4 &#x00D7; 4</td>
<td valign="top" align="left">AF3, FP1, FPZ, FP2, AF4</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Frontal</td>
<td valign="top" align="left">F3, FZ, F4</td>
<td valign="top" align="center">4 &#x00D7; 3</td>
<td valign="top" align="left">F3, F1, FZ, F2, F4</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Bilateral frontal</td>
<td valign="top" align="left">F7, F8</td>
<td valign="top" align="center">4 &#x00D7; 2</td>
<td valign="top" align="left">F7, F5, F6, F8</td>
<td valign="top" align="center">4 &#x00D7; 4</td>
</tr>
<tr>
<td valign="top" align="left">Left temporal</td>
<td valign="top" align="left">FC5, T7, CP5</td>
<td valign="top" align="center">4 &#x00D7; 3</td>
<td valign="top" align="left">FT7, FC5, T7, C5, TP7, CP5</td>
<td valign="top" align="center">4 &#x00D7; 6</td>
</tr>
<tr>
<td valign="top" align="left">Right temporal</td>
<td valign="top" align="left">FC6, T8, CP6</td>
<td valign="top" align="center">4 &#x00D7; 3</td>
<td valign="top" align="left">FT8, FC6, T8, C6, TP8, CP6</td>
<td valign="top" align="center">4 &#x00D7; 6</td>
</tr>
<tr>
<td valign="top" align="left">Frontal central</td>
<td valign="top" align="left">FC1, FC2</td>
<td valign="top" align="center">4 &#x00D7; 2</td>
<td valign="top" align="left">FC3, FC1, FCZ, FC2, FC4</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Central</td>
<td valign="top" align="left">C3, CZ, C4</td>
<td valign="top" align="center">4 &#x00D7; 3</td>
<td valign="top" align="left">C3, C1, CZ, C2, C4</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Central parietal</td>
<td valign="top" align="left">CP1, CP2</td>
<td valign="top" align="center">4 &#x00D7; 2</td>
<td valign="top" align="left">CP3, CP1, CPZ, CP2, CP4</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Bilateral parietal</td>
<td valign="top" align="left">P7, P8</td>
<td valign="top" align="center">4 &#x00D7; 2</td>
<td valign="top" align="left">P7, P5, P6, P8</td>
<td valign="top" align="center">4 &#x00D7; 4</td>
</tr>
<tr>
<td valign="top" align="left">Parietal</td>
<td valign="top" align="left">P3, PZ, P4</td>
<td valign="top" align="center">4 &#x00D7; 3</td>
<td valign="top" align="left">P3, P1, PZ, P2, P4</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Parietal occipital</td>
<td valign="top" align="left">PO3, PO4</td>
<td valign="top" align="center">4 &#x00D7; 2</td>
<td valign="top" align="left">PO5, PO3, POZ, PO4, PO6</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
<tr>
<td valign="top" align="left">Occipital</td>
<td valign="top" align="left">O1, OZ, O2</td>
<td valign="top" align="center">4 &#x00D7; 3</td>
<td valign="top" align="left">CB1, O1, OZ, O2, CB2</td>
<td valign="top" align="center">4 &#x00D7; 5</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To balance the samples of three kinds of emotion labels in DEAP, the values 4 and 7 are used as the threshold to distinguish the positive, neutral, and negative emotion labels. As a result, for the total 32 subjects of the DEAP dataset, the number of positive, neutral, and negative trials are 373, 540, and 367, respectively. The proportion of samples in positive, neutral, and negative class is about 29%, 42%, and 29%, respectively. In this way, 40 trials were collected for each subject including 2,400-s EEG records, which is segmented according to 1 s, including 2,400 EEG segments. Each segment corresponds to three types of emotional labels: positive, neutral, and negative, of which there are about 800 segments of each type of emotional label. In this way, each subject has a total of 40 trials &#x00D7; 60 s = 2,400 s of EEG records, which were segmented by 1 s and contained a total of 2,400 EEG segments. Each segment corresponds to three types of emotions: positive, neutral, or negative tags. Then all segments are divided into 480 EEG samples according to T = 5. That means each sample contains five EEG segments, and DE features of four bands are extracted from 32 electrodes of each segment, so that each EEG sample is expressed by a manual feature tensor of 4 &#x00D7; 32 &#x00D7; 5, and the size of each EEG dataset is 480 &#x00D7; 4 &#x00D7; 32 &#x00D7; 5. The size of the 32-subject EEG dataset is 15,360 &#x00D7; 4 &#x00D7; 32 &#x00D7; 5.</p>
<p>To further prove the performance of our proposed model and make the conclusion more convincing, we also conducted serial comparison experiments on the SEED dataset (<xref ref-type="bibr" rid="B43">Zheng and Lu, 2017</xref>). The dataset collected EEG records related to emotional stimulation from 64 channels of 15 subjects (7 men and 8 women). The emotional labels fed back by the subjects were divided into positive, neutral, and negative. The dataset has been preprocessed, and DE features for each subject were extracted. On the SEED dataset, we also used trial-wise randomization method to construct cross validation sets for within-subject experiments and used the same leave-one-subject-out (LOSO) method as that used on the DEAP dataset to construct cross validation sets for cross-subject experiments. As for brain area division, to facilitate comparison, we removed the PO7 and PO8 electrodes and divided the remaining 62 electrodes into 12 brain areas. <xref ref-type="table" rid="T1">Table 1</xref> shows the detailed brain area division method on the DEAP and SEED datasets.</p>
</sec>
<sec id="S3.SS2">
<title>Benchmark Methods</title>
<p>For comparison, we use the following benchmark methods to perform within-subject and cross-subject emotion classification experiments on the same dataset.</p>
<p>The three traditional learning methods are the following: support vector machine (SVM) (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>), bagging tree (BT) (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>), and random forest (RF) (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>).</p>
<p>The Seven deep learning methods are the following: deep confidence network (DBN) (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>), deep LSTM recurrent neural network (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>), 2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>), 3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>), hierarchical bidirectional GRU network based on attention mechanism (H-ATT-BGRU) (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>), domain adaptive neural network (DANN) (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>), and cascaded convolutional recurrent neural network (Casc-CNN-LSTM) (Chen et al.).</p>
<p>In order to horizontally compare the advantages of the proposed model, the input features of the benchmark models are also DE features extracted from four bands of EEG data in the DEAP and SEED datasets, which are consistent with those of our proposed model. The feature extraction method is the same as that stated in the experiment part of section &#x201C;Within-Subject Experiment of Electroencephalogram Emotion Recognition.&#x201D; <italic>Classifier and discriminator</italic>. However, the specific format of the input EEG features needs to be reshaped according to the interface of each model. Some key implementation details of these 10 benchmark models are listed in <xref ref-type="table" rid="T2">Table 2</xref>. The selection of model hyperparameters is also the result of fine-tuning experiments in the same search space mentioned in section &#x201C;Discussion About Several Variants of the Proposed Model.<italic>&#x201D; Configuration and training of the bidirectional long- and short-term memory neural network from region to global brain model</italic> of part II.</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>Implementation details of 10 benchmark models.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Benchmark models</td>
<td valign="top" align="center" colspan="2">Input data size<hr/></td>
<td valign="top" align="left">Implementation details</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">DEAP dataset</td>
<td valign="top" align="left">SEED dataset</td>
<td/>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Support vector machine (SVM) (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="left">[32 &#x00D7; 5, sample_size]</td>
<td valign="top" align="left">[Features, sample_size]</td>
<td valign="top" align="left">Kernel = &#x2018;rbf&#x2019;, gamma = 8, c = 0.05</td>
</tr>
<tr>
<td valign="top" align="left">Bagging tree (BT) (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="left">[32 &#x00D7; 5, sample_size]</td>
<td valign="top" align="left">[64 &#x00D7; 5, sample_size]</td>
<td valign="top" align="left">Method: bag, nLearn:100, weak learner: tree, type: classification</td>
</tr>
<tr>
<td valign="top" align="left">Random forest (RF) (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="left">[32 &#x00D7; 5, sample_size]</td>
<td valign="top" align="left">[64 &#x00D7; 5, sample_size]</td>
<td valign="top" align="left">n_estimators = 50, max_depth = 10, max_features = 8, min_samples_split = 20, min_samples_leaf = 10, oob_score = true, random_sate = 10</td>
</tr>
<tr>
<td valign="top" align="left">Deep confidence network (DBN) (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="left">[Batch_size, feature_size]: [60, 32 &#x00D7; 5]</td>
<td valign="top" align="left">[Batch_size, feature_size]: [60, 64 &#x00D7; 5]</td>
<td valign="top" align="left">hidden_layers = 3, hidden_size = 64, batch_size = 60, learning_rate = 0.04, dropout = 0.5, epochs = 200</td>
</tr>
<tr>
<td valign="top" align="left">Long- and short-term memory (LSTM) (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="left">[Batch_size, seq_len, channels]: [120, 5, 32]</td>
<td valign="top" align="left">[Batch_size, seq_len, channels]: [120, 5, 64]</td>
<td valign="top" align="left">hidden_layers = 2, hidden_size = 64, seq_len = 5, batch_size = 120, learning_rate = 0.03, dropout = 0.5, num_directions = 2, epochs = 100 &#x00D7;</td>
</tr>
<tr>
<td valign="top" align="left">Two-dimensional convolutional neural network (2D-CNN) (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
<td valign="top" align="left">[Batch_size, seq_len &#x00D7; band_size, channels]: [60, 5 &#x00D7; 4, 32]</td>
<td valign="top" align="left">[Batch_size, seq_len &#x00D7; band_size, channels]: [60, 5 &#x00D7; 4, 64]</td>
<td valign="top" align="left">conv_layers = 2, max_pool_layers = 2, full_conn_layers = 2 (hidden_size = 128), conv_kernels = [32, 64], kernel_zize = [5 &#x00D7; 5, 3 &#x00D7; 3], pool_size = (2,2), batch_size = 60, learning_rate = 0.05, dropout = 0.7, epochs = 300, padding = 0, stride = 1</td>
</tr>
<tr>
<td valign="top" align="left">Three-dimensional convolutional neural network (3D-CNN) (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="left">[Batch_size, band_size, seq_len, channels]: [80, 4, 5, 32]</td>
<td valign="top" align="left">[Batch_size, band_size, seq_len, channels]: [80, 4, 5, 64]</td>
<td valign="top" align="left">conv_layers = 2, max_pool_layers = 1, full_conn_layers = 1 (hidden_size = 128), conv_kernels = [8, 16], kernel_zize = [3 &#x00D7; 3 &#x00D7; 7, 2 &#x00D7; 2 &#x00D7; 5], pool_size = (2,2), batch_size = 80, learning_rate = 0.01, dropout = 0.6, epochs = 200, padding = 0, stride = 1</td>
</tr>
<tr>
<td valign="top" align="left">Hierarchical bidirectional GRU network based on attention mechanism (H-ATT-BGRU) (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="left">[Batch_size, band_size, seq_len, channels]: [60, 4, 5, 32]</td>
<td valign="top" align="left">[Batch_size, band_size, seq_len, channels]: [60, 4, 5, 64]</td>
<td valign="top" align="left">hidden_layers = 2, hidden_size = 64, seq_len = 5, batch_size = 60, learning_rate = 0.06, dropout = 0.4, num_directions = 2, epochs = 400</td>
</tr>
<tr>
<td valign="top" align="left">Domain adaptive neural network (DANN) (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="left">[Batch_size, feature_size]: [30, 32 &#x00D7; 5]</td>
<td valign="top" align="left">[Batch_size, feature_size]: [30, 64 &#x00D7; 5]</td>
<td valign="top" align="left">hidden_layers = 2, hidden_size = 128, batch_size = 30, L2-weight-regularization = 0.003, learning_rate = 0.02, dropout = 0.5, epochs = 500, momentum = 0.05, MMD regularization constant &#x03B3; = 10e3</td>
</tr>
<tr>
<td valign="top" align="left">convolutional recurrent neural network (Casc-CNN-LSTM) (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="left">[Batch_size, seq_len &#x00D7; band_size, channels]: [80, 5 &#x00D7; 4, 32]</td>
<td valign="top" align="left">[Batch_size, seq_len &#x00D7; band_size, channels]: [80, 5 &#x00D7; 4, 64]</td>
<td valign="top" align="left">CNN: conv_layers = 3, max_pool_layers = 3, full_conn_layers = 2 (hidden_size = 256), conv_kernels = [32, 64, 128], kernel_zize = [3 &#x00D7; 3, 3 &#x00D7; 3, 3 &#x00D7; 3], pool_size = (2,2), batch_size = 80, learning_rate = 0.05, dropout = 0.5, epochs = 500, padding = 0, stride = 1 LSTM: hidden_layers = 2, hidden_size = 128, seq_len = 256, batch_size = 80, learning_rate = 0.05, dropout = 0.5, num_directions = 2, epochs = 500</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="S3.SS3">
<title>Within-Subject Experiment of Electroencephalogram Emotion Recognition</title>
<p>We apply within-subject EEG emotion recognition method like that in literature (<xref ref-type="bibr" rid="B23">Li et al., 2018a</xref>) to evaluate our proposed model. To make the experiment result convincing, we use trial-wise randomization to construct the validation dataset. Specifically, we first picked out the subjects with a relatively balanced number of three types of trials. These 13 selected subjects include sub05, sub10, sub12, sub13, sub15, sub21, sub22, sub24, sub25, sub26, sub28, sub29, sub32. For each of these selected subjects, we randomly selected all segments of about 10% of the trials from each type as the test set, then randomly selected all segments of about another 10% of the trials from the rest of each type as the validation set, and at last take all segments of the remaining 80% of the trials as the training set. In this division process, we will make sure all segments belonging to one trial is allocated either as the training set, test set, or validation set to avoid &#x201C;data leakage.&#x201D; Then the proposed R2G-ST-BiLSTM model is used for feature learning and emotion classification. The learning process on the DEAP dataset is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Learning process of R2G-ST-BiLSTM model in within-subject experiment on DEAP dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g004.tif"/>
</fig>
<p>We use the average classification accuracy (ACC) and standard deviation (STD) of all subjects to evaluate the model performance. For comparison, we also use the abovementioned benchmark methods to make experiments on the equal dataset. We use paired t-test against the benchmark methods to show the difference between them. <italic>T</italic>-test is a test method for the difference between two mean values of small samples (sample size less than 30). It uses t-distribution theory to infer the probability of difference, to judge whether the difference is significant. The significance of the classification performance of the proposed method against each benchmark method is calculated with paired <italic>t</italic>-test. For all the paired <italic>t</italic>-tests, we used Bonferroni criteria (<xref ref-type="bibr" rid="B16">Genovese et al., 2002</xref>) and the implementation method (<xref ref-type="bibr" rid="B36">Weisstein, 2004</xref>) to make p-value correction for multiple hypothesis testing to limit false discovery rate (FDR). The results of within-subject experiment on the DEAP and SEED datasets are shown, respectively, in <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref>. The <italic>p-</italic>Value indicates the corrected results of paired <italic>t</italic>-test. A value of <italic>p</italic> &#x003C; 0.05 means the difference is significant.</p>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>The results of within-subject experiment on the DEAP dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">SVM (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="center">BT (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="center">RF (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="center">DBN (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="center">LSTM (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="center">2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Average classification accuracy (ACC)(%)/standard deviation (STD)</td>
<td valign="top" align="center">80.72/7.67</td>
<td valign="top" align="center">84.65/8.93</td>
<td valign="top" align="center">78.87/11.32</td>
<td valign="top" align="center">82.83/9.54</td>
<td valign="top" align="center">84.51/10.06</td>
<td valign="top" align="center">85.63/8.72</td>
</tr>
<tr>
<td valign="top" align="left"><italic>p-</italic>Value</td>
<td valign="top" align="center">0.0005</td>
<td valign="top" align="center">0.0004</td>
<td valign="top" align="center">0.0006</td>
<td valign="top" align="center">0.0008</td>
<td valign="top" align="center">0.0023</td>
<td valign="top" align="center">0.0019</td>
</tr>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="center">H-ATT-BGRU (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="center">DANN (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="center">Casc-CNN-LSTM (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="center">R2G-ST-BiLSTM</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">ACC (%)/STD</td>
<td valign="top" align="center">87.21/10.57</td>
<td valign="top" align="center">87.89/8.94</td>
<td valign="top" align="center">88.54/9.26</td>
<td valign="top" align="center">93.95/7.88</td>
<td valign="top" align="center">94.69/9.81</td>
<td/>
</tr>
<tr>
<td valign="top" align="left"><italic>p-</italic>Value</td>
<td valign="top" align="center">0.0052</td>
<td valign="top" align="center">0.0066</td>
<td valign="top" align="center">0.0074</td>
<td valign="top" align="center">0.0089</td>
<td/>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T4">
<label>TABLE 4</label>
<caption><p>The results of within-subject experiment on the SEED dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">SVM (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="center">BT (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="center">RF (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="center">DBN (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="center">LSTM (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="center">2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ACC (%)/STD</td>
<td valign="top" align="center">80.14/9.27</td>
<td valign="top" align="center">83.72/8.68</td>
<td valign="top" align="center">77.95/9.32</td>
<td valign="top" align="center">82.58/11.26</td>
<td valign="top" align="center">83.92/9.44</td>
<td valign="top" align="center">84.64/7.98</td>
</tr>
<tr>
<td valign="top" align="left"><italic>p-</italic>Value</td>
<td valign="top" align="center">0.0006</td>
<td valign="top" align="center">0.0002</td>
<td valign="top" align="center">0.0004</td>
<td valign="top" align="center">0.0007</td>
<td valign="top" align="center">0.0009</td>
<td valign="top" align="center">0.0008</td>
</tr>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="center">H-ATT-GRU (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="center">DANN (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="center">Casc-CNN-LSTM (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="center">R2G-ST-BiLSTM</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">ACC/STD</td>
<td valign="top" align="center">87.31/11.14</td>
<td valign="top" align="center">86.38/9.56</td>
<td valign="top" align="center">88.96/10.45</td>
<td valign="top" align="center">92.72/9.33</td>
<td valign="top" align="center">93.57/8.52</td>
<td/>
</tr>
<tr>
<td valign="top" align="left"><italic>p-</italic>Value</td>
<td valign="top" align="center">0.0021</td>
<td valign="top" align="center">0.0035</td>
<td valign="top" align="center">0.0052</td>
<td valign="top" align="center">0.0098</td>
<td/>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be seen from <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref> that the average accuracy of the R2G-ST-BiLSTM method achieves 94.69% on DEAP and 93.57% on SEED, which is best among the above methods. From a statistical point of view, the performance of the proposed model is significantly better than the benchmark models. This result is largely because our R2G-ST-BiLSTM model explores both temporal and spatial context information of the different brain regions of EEG signals.</p>
<p>According to the experimental result of our proposed model on DEAP, we draw a confusion matrix for the three categories of emotions in <xref ref-type="fig" rid="F5">Figure 5</xref>. It is found that compared with neutral emotions, positive and negative affections are less likely to be confused.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Confusion matrix of R2G-ST-BiLSTM model for within-subject experiment on DEAP dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g005.tif"/>
</fig>
<p>We also used a method-like reference (<xref ref-type="bibr" rid="B40">Zheng, 2017</xref>) to conduct some additional experiments to test the classification performance of different frequency bands of EEG data. Specifically, the DE features are extracted from four frequency bands &#x03B8;, &#x03B1;, &#x03B2;, and &#x03B3; related to the original signal, and then the EEG emotion recognition experiment is performed based on these DE features of the four bands. We can see the experimental results on DEAP and SEED datasets in <xref ref-type="table" rid="T5">Table 5</xref>, which indicate that on both datasets, the recognition performance in the higher frequency bands of &#x03B2; and &#x03B3; is better than those in the lower frequency bands of &#x03B8; and &#x03B1;. This result is consistent with the neurophysiology research in literature (<xref ref-type="bibr" rid="B29">Mauss and Robinson, 2009</xref>).</p>
<table-wrap position="float" id="T5">
<label>TABLE 5</label>
<caption><p>The results of four frequency bands in within-subject experiment.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Methods</td>
<td valign="top" align="center" colspan="8">The results of ACC (%) (STD)<hr/></td>
</tr>
<tr>
<td/>
<td valign="top" align="center" colspan="4">DEAP<hr/></td>
<td valign="top" align="center" colspan="4">SEED<hr/></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">&#x03B8;</td>
<td valign="top" align="center">&#x03B1;</td>
<td valign="top" align="center">&#x03B2;</td>
<td valign="top" align="center">&#x03B3;</td>
<td valign="top" align="center">&#x03B8;</td>
<td valign="top" align="center">&#x03B1;</td>
<td valign="top" align="center">&#x03B2;</td>
<td valign="top" align="center">&#x03B3;</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="center">60.90 (8.76)</td>
<td valign="top" align="center">62.16 (10.49)</td>
<td valign="top" align="center">72.75 (7.87)</td>
<td valign="top" align="center">74.28 (11.13)</td>
<td valign="top" align="center">57.64 (9.93)</td>
<td valign="top" align="center">63.19 (7.55)</td>
<td valign="top" align="center">76.85 (10.19)</td>
<td valign="top" align="center">72.26 (12.31)</td>
</tr>
<tr>
<td valign="top" align="left">BT (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="center">65.44 (7.82)</td>
<td valign="top" align="center">67.31 (9.65)</td>
<td valign="top" align="center">77.52 (10.04)</td>
<td valign="top" align="center">76.06 (8.28)</td>
<td valign="top" align="center">61.65 (11.42)</td>
<td valign="top" align="center">62.74 (8.16)</td>
<td valign="top" align="center">75.58 (9.37)</td>
<td valign="top" align="center">78.83 (6.96)</td>
</tr>
<tr>
<td valign="top" align="left">RF (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="center">62.38 (12.52)</td>
<td valign="top" align="center">64.54 (7.39)</td>
<td valign="top" align="center">72.06 (9.77)</td>
<td valign="top" align="center">71.87 (10.45)</td>
<td valign="top" align="center">62.79 (7.58)</td>
<td valign="top" align="center">62.85 (12.32)</td>
<td valign="top" align="center">69.10 (11.14)</td>
<td valign="top" align="center">71.72 (10.56)</td>
</tr>
<tr>
<td valign="top" align="left">DBN (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="center">61.19 (8.97)</td>
<td valign="top" align="center">62.74 (11.25)</td>
<td valign="top" align="center">73.73 (6.80)</td>
<td valign="top" align="center">75.05 (7.74)</td>
<td valign="top" align="center">58.32 (9.44)</td>
<td valign="top" align="center">62.56 (10.28)</td>
<td valign="top" align="center">70.47 (13.51)</td>
<td valign="top" align="center">74.29 (8.84)</td>
</tr>
<tr>
<td valign="top" align="left">LSTM (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="center">64.98 (9.15)</td>
<td valign="top" align="center">77.66 (10.57)</td>
<td valign="top" align="center">79.13 (7.22)</td>
<td valign="top" align="center">80.29 (8.83)</td>
<td valign="top" align="center">60.57 (11.89)</td>
<td valign="top" align="center">70.14 (12.67)</td>
<td valign="top" align="center">76.35 (10.23)</td>
<td valign="top" align="center">78.81 (9.50)</td>
</tr>
<tr>
<td valign="top" align="left">2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
<td valign="top" align="center">65.73 (8.89)</td>
<td valign="top" align="center">68.45 (6.35)</td>
<td valign="top" align="center">79.96 (9.84)</td>
<td valign="top" align="center">81.42 (10.77)</td>
<td valign="top" align="center">67.22 (6.73)</td>
<td valign="top" align="center">69.36 (8.65)</td>
<td valign="top" align="center">77.24 (11.38)</td>
<td valign="top" align="center">80.58 (12.72)</td>
</tr>
<tr>
<td valign="top" align="left">3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="center">65.26 (7.34)</td>
<td valign="top" align="center">70.17 (9.89)</td>
<td valign="top" align="center">82.51 (11.43)</td>
<td valign="top" align="center">83.68 (12.05)</td>
<td valign="top" align="center">64.54 (7.42)</td>
<td valign="top" align="center">71.09 (12.16)</td>
<td valign="top" align="center">78.67 (9.93)</td>
<td valign="top" align="center">82.11 (10.64)</td>
</tr>
<tr>
<td valign="top" align="left">H-ATT-BiGRU (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="center">66.27 (8.11)</td>
<td valign="top" align="center">68.58 (6.92)</td>
<td valign="top" align="center">81.96 (11.15)</td>
<td valign="top" align="center">84.25 (9.32)</td>
<td valign="top" align="center">65.03 (9.19)</td>
<td valign="top" align="center">67.15 (11.54)</td>
<td valign="top" align="center">81.58 (10.26)</td>
<td valign="top" align="center">85.34 (8.81)</td>
</tr>
<tr>
<td valign="top" align="left">DANN (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="center">68.39 (12.56)</td>
<td valign="top" align="center">70.87 (10.75)</td>
<td valign="top" align="center">85.73 (11.62)</td>
<td valign="top" align="center">86.92 (9.19)</td>
<td valign="top" align="center">67.56 (11.04)</td>
<td valign="top" align="center">72.42 (7.75)</td>
<td valign="top" align="center">79.96 (8.42)</td>
<td valign="top" align="center">85.47 (9.73)</td>
</tr>
<tr>
<td valign="top" align="left">Casc-CNN-LSTM (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="center">70.07 (7.44)</td>
<td valign="top" align="center">73.25 (8.81)</td>
<td valign="top" align="center">88.54 (9.69)</td>
<td valign="top" align="center">89.18 (11.23)</td>
<td valign="top" align="center">69.21 (8.12)</td>
<td valign="top" align="center">75.88 (9.93)</td>
<td valign="top" align="center">85.25 (10.36)</td>
<td valign="top" align="center">89.53 (7.39)</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM</td>
<td valign="top" align="center"><bold>71.46 (10.73)</bold></td>
<td valign="top" align="center"><bold>75.82 (9.55)</bold></td>
<td valign="top" align="center"><bold>90.57 (7.36)</bold></td>
<td valign="top" align="center"><bold>91.38 (8.92)</bold></td>
<td valign="top" align="center"><bold>71.35 (8.28)</bold></td>
<td valign="top" align="center"><bold>87.14 (6.67)</bold></td>
<td valign="top" align="center"><bold>86.72 (9.81)</bold></td>
<td valign="top" align="center"><bold>90.86 (11.92)</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Bold values represent the better results obtained by the proposed method, highlighting the comparison.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="S3.SS4">
<title>Cross-Subject Experiment of Electroencephalogram Emotion Recognition</title>
<p>In this section, we use the cross-subject and the leave-one-subject-out (LOSO) cross-validation strategy similar to that in <xref ref-type="bibr" rid="B42">Zheng and Lu (2016)</xref>; <xref ref-type="bibr" rid="B23">Li et al. (2018a)</xref> to evaluate the proposed method, in which the training and testing data are selected from different subjects. The EEG data of one subject is selected as test data, and the EEG data of all the rest of the subjects are used as training data. After each subject is rounded, the average prediction accuracy and standard deviation are calculated as the results. To better compare the performance of the proposed method, we use the abovementioned methods as benchmark. The comparison results of various methods on DEAP and SEED are illustrated in <xref ref-type="table" rid="T6">Tables 6</xref>, <xref ref-type="table" rid="T7">7</xref>, respectively. On both datasets, our R2G-ST-BiLSTM model also performs better. The learning process on the DEAP dataset is shown in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<table-wrap position="float" id="T6">
<label>TABLE 6</label>
<caption><p>The results of cross-subject experiment on the DEAP dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">SVM (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="center">BT (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="center">RF (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="center">DBN (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="center">LSTM (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="center">2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ACC/STD</td>
<td valign="top" align="center">56.32/10.25</td>
<td valign="top" align="center">58.49/8.76</td>
<td valign="top" align="center">51.74/11.13</td>
<td valign="top" align="center">59.01/7.88</td>
<td valign="top" align="center">64.66/11.40</td>
<td valign="top" align="center">65.25/9.37</td>
</tr>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="center">H-ATT-BGRU (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="center">DANN (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="center">Casc-CNN-LSTM (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="center">R2G-ST-BiLSTM</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">ACC/STD</td>
<td valign="top" align="center">68.13/14.07</td>
<td valign="top" align="center">77.82/10.12</td>
<td valign="top" align="center">75.24/8.59</td>
<td valign="top" align="center">82.36/7.15</td>
<td valign="top" align="center">84.51/9.26</td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T7">
<label>TABLE 7</label>
<caption><p>The results of cross-subject experiment on the SEED dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">CM (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="center">BT (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="center">RF (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="center">DBN (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="center">LSTM (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="center">2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ACC/STD</td>
<td valign="top" align="center">56.73/16.29</td>
<td valign="top" align="center">51.23/14.82</td>
<td valign="top" align="center">69.00/10.89</td>
<td valign="top" align="center">61.28/14.62</td>
<td valign="top" align="center">63.54/15.47</td>
<td valign="top" align="center">71.31/14.09</td>
</tr>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="center">H-ATT-BGRU (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="center">DANN (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="center">Casc-CNN-LSTM (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="center">R2G-ST-BiLSTM</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">ACC/STD</td>
<td valign="top" align="center">69.13/13.07</td>
<td valign="top" align="center">76.31/15.89</td>
<td valign="top" align="center">79.95/9.02</td>
<td valign="top" align="center">83.28/9.60</td>
<td valign="top" align="center">85.49/7.96</td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>Learning process of R2G-ST-BiLSTM model in cross-subject experiment on DEAP dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g006.tif"/>
</fig>
<p>We also draw a confusion matrix in <xref ref-type="fig" rid="F7">Figure 7</xref> according to the results of our model on the DEAP dataset, which shows that positive emotion is easier to be recognized than the negative and neutral emotions.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Confusion matrix of R2G-ST-BiLSTM model for cross-subject experiment on DEAP dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g007.tif"/>
</fig>
<p>We also compared the influence of the different frequency bands on cross-subject emotion recognition. The experimental results on DEAP and SEED datasets are shown in <xref ref-type="table" rid="T8">Table 8</xref>, from which it can be seen that, on both datasets, the classification performance in the higher frequency bands of &#x03B2; and &#x03B3; are better than those in the lower frequency bands of &#x03B8; and &#x03B1;, and the R2G-ST-BiLSTM method achieves the best performance on the four frequency bands.</p>
<table-wrap position="float" id="T8">
<label>TABLE 8</label>
<caption><p>The results of four frequency bands in cross-subject experiment.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Methods</td>
<td valign="top" align="center" colspan="8">The results (%) of ACC (STD)</td>
</tr>
<tr>
<td/>
<td valign="top" align="center" colspan="4">DEAP</td>
<td valign="top" align="center" colspan="4">SEED</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">&#x03B8;</td>
<td valign="top" align="center">&#x03B1;</td>
<td valign="top" align="center">&#x03B2;</td>
<td valign="top" align="center">&#x03B3;</td>
<td valign="top" align="center">&#x03B8;</td>
<td valign="top" align="center">&#x03B1;</td>
<td valign="top" align="center">&#x03B2;</td>
<td valign="top" align="center">&#x03B3;</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM (<xref ref-type="bibr" rid="B35">Suykens and Vandewalle, 1999</xref>)</td>
<td valign="top" align="center">41.91 (8.39)</td>
<td valign="top" align="center">44.73 (7.56)</td>
<td valign="top" align="center">48.66 (10.21)</td>
<td valign="top" align="center">51.32 (9.08)</td>
<td valign="top" align="center">40.62 (9.83)</td>
<td valign="top" align="center">42.05 (12.65)</td>
<td valign="top" align="center">47.97 (12.47)</td>
<td valign="top" align="center">50.06 (10.48)</td>
</tr>
<tr>
<td valign="top" align="left">BT (<xref ref-type="bibr" rid="B8">Chuang et al., 2012</xref>)</td>
<td valign="top" align="center">45.17 (9.38)</td>
<td valign="top" align="center">48.29 (14.77)</td>
<td valign="top" align="center">53.95 (8.54)</td>
<td valign="top" align="center">54.48 (7.43)</td>
<td valign="top" align="center">45.98 (9.70)</td>
<td valign="top" align="center">48.63 (10.28)</td>
<td valign="top" align="center">49.79 (12.41)</td>
<td valign="top" align="center">54.07 (6.87)</td>
</tr>
<tr>
<td valign="top" align="left">RF (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>)</td>
<td valign="top" align="center">41.50 (8.57)</td>
<td valign="top" align="center">41.86 (4.52)</td>
<td valign="top" align="center">47.31 (12.02)</td>
<td valign="top" align="center">47.72 (10.05)</td>
<td valign="top" align="center">40.07 (6.50)</td>
<td valign="top" align="center">42.09 (13.34)</td>
<td valign="top" align="center">48.29 (12.77)</td>
<td valign="top" align="center">48.98 (12.82)</td>
</tr>
<tr>
<td valign="top" align="left">DBN (<xref ref-type="bibr" rid="B41">Zheng and Lu, 2015</xref>)</td>
<td valign="top" align="center">44.36 (11.82)</td>
<td valign="top" align="center">46.15 (8.98)</td>
<td valign="top" align="center">55.94 (6.01)</td>
<td valign="top" align="center">56.81 (9.27)</td>
<td valign="top" align="center">45.76 (10.98)</td>
<td valign="top" align="center">48.43 (9.75)</td>
<td valign="top" align="center">56.66 (6.58)</td>
<td valign="top" align="center">56.62 (6.84)</td>
</tr>
<tr>
<td valign="top" align="left">LSTM (<xref ref-type="bibr" rid="B1">Alhagry et al., 2017</xref>)</td>
<td valign="top" align="center">47.92 (6.45)</td>
<td valign="top" align="center">48.69 (10.40)</td>
<td valign="top" align="center">59.02 (7.83)</td>
<td valign="top" align="center">59.16 (11.62)</td>
<td valign="top" align="center">48.63 (10.29)</td>
<td valign="top" align="center">51.59 (11.83)</td>
<td valign="top" align="center">62.13 (7.73)</td>
<td valign="top" align="center">59.37 (10.75)</td>
</tr>
<tr>
<td valign="top" align="left">2D-CNN (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>)</td>
<td valign="top" align="center">48.33 (7.61)</td>
<td valign="top" align="center">49.74 (13.26)</td>
<td valign="top" align="center">62.18 (9.90)</td>
<td valign="top" align="center">62.09 (11.31)</td>
<td valign="top" align="center">48.36 (10.31)</td>
<td valign="top" align="center">50.60 (8.30)</td>
<td valign="top" align="center">62.04 (6.74)</td>
<td valign="top" align="center">62.19 (7.62)</td>
</tr>
<tr>
<td valign="top" align="left">3D-CNN (<xref ref-type="bibr" rid="B32">Salama et al., 2018</xref>)</td>
<td valign="top" align="center">51.81 (9.79)</td>
<td valign="top" align="center">53.46 (9.84)</td>
<td valign="top" align="center">65.15 (11.32)</td>
<td valign="top" align="center">64.97 (8.46)</td>
<td valign="top" align="center">52.60 (11.84)</td>
<td valign="top" align="center">54.95 (10.45)</td>
<td valign="top" align="center">64.47 (13.69)</td>
<td valign="top" align="center">64.47 (14.69)</td>
</tr>
<tr>
<td valign="top" align="left">H-ATT-BiGRU (<xref ref-type="bibr" rid="B4">Chen et al., 2019b</xref>)</td>
<td valign="top" align="center">63.44 (12.50)</td>
<td valign="top" align="center">61.52 (7.07)</td>
<td valign="top" align="center">70.39 (12.14)</td>
<td valign="top" align="center">72.63 (5.28)</td>
<td valign="top" align="center">64.47 (14.96)</td>
<td valign="top" align="center">59,81 (12.43)</td>
<td valign="top" align="center">71.03 (10.48)</td>
<td valign="top" align="center">73.55 (8.80)</td>
</tr>
<tr>
<td valign="top" align="left">DANN (<xref ref-type="bibr" rid="B14">Ganin et al., 2016</xref>)</td>
<td valign="top" align="center">56.98 (5.33)</td>
<td valign="top" align="center">58.06 (11.80)</td>
<td valign="top" align="center">67.70 (8.65)</td>
<td valign="top" align="center">70.46 (12.17)</td>
<td valign="top" align="center">55.47 (9.80)</td>
<td valign="top" align="center">56.72 (10.79)</td>
<td valign="top" align="center">67.14 (7.17)</td>
<td valign="top" align="center">71.03 (10.14)</td>
</tr>
<tr>
<td valign="top" align="left">Casc-CNN-LSTM (<xref ref-type="bibr" rid="B7">Chen et al., 2020</xref>)</td>
<td valign="top" align="center">61.27 (8.02)</td>
<td valign="top" align="center">62.83 (6.56)</td>
<td valign="top" align="center">73.59 (10.54)</td>
<td valign="top" align="center">73.55 (8.69)</td>
<td valign="top" align="center">62.04 (6.64)</td>
<td valign="top" align="center">63.31 (11.96)</td>
<td valign="top" align="center">73.25 (9,12)</td>
<td valign="top" align="center">74.29 (7.98)</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM</td>
<td valign="top" align="center"><bold>64.03 (14.41)</bold></td>
<td valign="top" align="center"><bold>66.26 (5.99)</bold></td>
<td valign="top" align="center"><bold>74.64 (9.38)</bold></td>
<td valign="top" align="center"><bold>75.02 (10.10)</bold></td>
<td valign="top" align="center"><bold>66.14 (8.10)</bold></td>
<td valign="top" align="center"><bold>67.14 (7.05)</bold></td>
<td valign="top" align="center"><bold>74.85 (8.02)</bold></td>
<td valign="top" align="center"><bold>75.89 (8.15)</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Bold values represent the better results obtained by the proposed method, highlighting the comparison.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>To prove the influence of the different brain regions on emotion recognition, we visualize the weight distribution of brain regions based on the weighting matrix W defined in formula (5) and learned in our cross-subject experiment on DEAP, where the sum of each row of W matrix represents the contribution of corresponding brain region. <xref ref-type="fig" rid="F8">Figure 8</xref> shows the weighted map of the brain areas, where the darker the color of the region, the more significant contribution of the corresponding brain region. It can be seen from <xref ref-type="fig" rid="F6">Figure 6</xref> that EEG signals in the frontal lobe are very important for human emotion recognition, which is consistent with the results of the cognitive observations of biological psychology in the literature (<xref ref-type="bibr" rid="B9">Coan and Allen, 2004</xref>).</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption><p>Weighted map of brain areas.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-738167-g008.tif"/>
</fig>
</sec>
<sec id="S3.SS5">
<title>Discussion About Several Variants of the Proposed Model</title>
<p>Various experiments on the DEAP dataset demonstrates that the proposed R2G-ST-BILSM model is more effective than the other methods, which is largely due to our R2G-ST-BiLSM model utilizing both regional weighting layer and regional to global time layer. In order to confirm that, we obtained the following three simplified models by removing some layers from the R2G-ST-BiLSTM network, and use them to make within-subject and cross-subject experiments on the DEAP dataset. These three simplified models are described as follows:</p>
<list list-type="simple">
<list-item>
<label>(1)</label>
<p>R2G-ST-BiLSTM-V1&#x2014;removes the dynamic regional weighting layer and regional temporal feature learning layers;</p>
</list-item>
<list-item>
<label>(2)</label>
<p>R2G-ST-BiLSTM-V2&#x2014;only uses global temporal feature as the final input feature to classify;</p>
</list-item>
<list-item>
<label>(3)</label>
<p>R2G-ST-BiLSTM-V3&#x2014;does not change the original structure of the R2G-ST-BiLSTM, except that the weight of each brain region is set to 1, which means all brain regions are of the same importance to emotion classification.</p>
</list-item>
</list>
<p><xref ref-type="table" rid="T9">Table 9</xref> demonstrates the comparison outcome of the above four variant models. The comparison relationship is as follows:</p>
<table-wrap position="float" id="T9">
<label>TABLE 9</label>
<caption><p>Comparison results of four models on the DEAP dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Methods</td>
<td valign="top" align="center">Within-subject experiment<hr/></td>
<td valign="top" align="center">Cross-subject experiment<hr/></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ACC (%)</td>
<td valign="top" align="center">ACC (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM-V1</td>
<td valign="top" align="center">90.43</td>
<td valign="top" align="center">80.32</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM-V2</td>
<td valign="top" align="center">91.58</td>
<td valign="top" align="center">81.15</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM-V3</td>
<td valign="top" align="center">93.72</td>
<td valign="top" align="center">83.96</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM</td>
<td valign="top" align="center"><bold>94.69</bold></td>
<td valign="top" align="center"><bold>84.51</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Bold values represent the better results obtained by the proposed method, highlighting the comparison.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>R2G-ST-BiLSTM-V1 &#x003C; R2G-ST-BiLSTM-V2 &#x003C; R2G-ST-BiLSTM-V3 &#x003C; R2G-ST-BiLSTM, (27)</p>
<p>The significance of the regional weighting layer and the regional temporal feature learning layer has been proven by the above comparisons, which shows that these two parts play important roles in enhancing the capability of our R2G-ST-BiLSTM model.</p>
<p>To further discuss whether the different components of R2G-ST-BiLSTM are necessary to outperform other models, we modified it according to the following methods to obtain its several variants:</p>
<list list-type="simple">
<list-item>
<label>(1)</label>
<p>R2G-ST-CNN-V1: replaces all BiLSTM modules used for learning spatial and temporal features of local and global brain regions with two-layer 2D-CNN modules.</p>
</list-item>
<list-item>
<label>(2)</label>
<p>R2G-ST-CNN-V2: only the BiLSTM modules used for learning temporal features of local and global brain regions are replaced with two-layer 2D-CNN modules.</p>
</list-item>
<list-item>
<label>(3)</label>
<p>R2G-ST-CNN-V3: only the BiLSTM modules used for learning spatial features of local and global brain regions are replaced with two-layer 2D-CNN modules.</p>
</list-item>
<list-item>
<label>(4)</label>
<p>R2G-ST-BiLSTM-V4: only remove the domain discriminator from the proposed model.</p>
</list-item>
</list>
<p>The structure and parameter configuration of the 2D-CNN here are consistent with those in literature (<xref ref-type="bibr" rid="B6">Chen et al., 2019c</xref>). These four variant models are used to make within-subject and cross-subject emotion classification experiments on DEAP. The comparison results are shown in <xref ref-type="table" rid="T10">Table 10</xref>.</p>
<table-wrap position="float" id="T10">
<label>TABLE 10</label>
<caption><p>Comparison results of five models on the DEAP dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Methods</td>
<td valign="top" align="center">Within-subject experiment<hr/></td>
<td valign="top" align="center">Cross-subject experiment<hr/></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ACC (%)</td>
<td valign="top" align="center">ACC (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">R2G-ST-CNN-V1</td>
<td valign="top" align="center">85.26</td>
<td valign="top" align="center">75.64</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-CNN-V2</td>
<td valign="top" align="center">88.75</td>
<td valign="top" align="center">80.72</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-CNN-V3</td>
<td valign="top" align="center">90.83</td>
<td valign="top" align="center">81.47</td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM</td>
<td valign="top" align="center"><bold>94.69</bold></td>
<td valign="top" align="center"><bold>84.51</bold></td>
</tr>
<tr>
<td valign="top" align="left">R2G-ST-BiLSTM-V4</td>
<td valign="top" align="center">92.14</td>
<td valign="top" align="center">78.39</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Bold values represent the better results obtained by the proposed method, highlighting the comparison.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>It can be seen from <xref ref-type="table" rid="T10">Table 10</xref> that the classification performance of the proposed model is significantly better than that of the four variant models. Specifically, the proposed model outperforms the R2G-ST-CNN-V2, which indicates that the BiLSTM components can extract more discriminative time-context features from EEG sequences than 2D-CNN. The performance of our proposed model is better than that of R2G-ST-CNN-V3, which shows that BiLSTM components can better cooperate with the attention mechanism of brain regions and extract more spatial context-dependent features than 2D-CNN. The proposed model significantly outperforms R2G-ST-CNN-V1, which further proves that BiLSTM has obvious advantages over 2D-CNN in learning deep temporal and spatial features in our proposed hierarchical framework. The proposed model significantly outperforms R2G-ST-CNN-V4, especially in cross-subject experiment, which illustrates that the domain discriminator is indeed helpful to extract more discriminative EEG features with small differences between subjects and, therefore, improve the adaptability of the model. In general, the components of BiLSTM and the domain discriminator play very important roles on the whole performance of the proposed model and are necessary to outperform other models.</p>
</sec>
</sec>
<sec id="S4" sec-type="discussion">
<title>Discussion</title>
<p>Although the proposed model has achieved high classification accuracy, there are still some limitations to study and overcome in the future.</p>
<p>At first, the model is complex and lacks interpretability. The model proposed in this paper is a combined hierarchical deep neural network composed of multiple bidirectional LSTM models with attention mechanism. Although the principle and learning process of the model is clear, the decision making and intermediate process made by the model are difficult to understand and interpreted. It is hard to explain the correlation and the interaction among input data, learned features, and output class. At present, researchers have put forward some specific deep model interpretation methods including activation maximization, gradient-based interpretation, class activation mapping (CAM), and so on. The interpretation result of the activation maximization is more accurate and can help people understand the internal working logic of DNN, but the data containing some noise generated in the optimization process makes it difficult to interpret the input (<xref ref-type="bibr" rid="B11">Dong et al., 2017</xref>). The gradient-based interpretation methods include deconvolution, guided backpropagation, integrated gradients, and smooth gradients, which aim to use backpropagation to calculate the gradient of specific output relative to input to derive feature importance. This gradient information can only be used to locate important features, but not to quantify the contribution of each feature to the classification results. The CAM method (<xref ref-type="bibr" rid="B21">Jorg, 2019</xref>) can locate the objects from the learned features by the excellent ability of the last convolution layer in CNN, which could only provide coarse-grained interpretation results for various CNN models. Additionally, there are some model-agnostic (MA) explanations, such as LIME and knowledge distillation, and causal interpretable method. Although many methods have been proposed in the interpretability research for deep models, there are still many problems to be solved, such as the lack of unified indicators for evaluating interpretation methods, the balance between model accuracy and interpretability, and the balance between data privacy protection and model interpretability, which will be one of our future research directions to improve the performance of the model.</p>
<p>Second, the complex model and limited amount of data make the model prone to overfitting. We use the EEG data of the DEAP and SEED datasets, which include 32 and 15 subjects, respectively, to make our experiments. In cross-subject experiments on the DEAP dataset, the number of the model parameters reaches about 50,156, which exceeds the number of training samples at 15,360. Compared with the complexity of the model, the training dataset is small, which makes the model prone to overfitting. At present, researchers usually use methods such as expanding dataset, removing features, regularization, and terminating training in advance to prevent model overfitting (<xref ref-type="bibr" rid="B33">Sanjar et al., 2020</xref>). Data enhancement is a way to increase training data, which can be realized by flipping, translation, rotation, scaling, and generation methods. Removing features is to reduce the complexity of the model by removing some layers or some neurons from it. Through monitoring the performance of each training iteration and when the loss on the verification set tends to increase, we could stop the training process to prevent the model from overfitting. The regularization method reduces the complexity of the model by punishing the loss function with L1 or L2 paradigm. In our work, we use L2 regularization and dropout method to suppress the overfitting problem, but we still face the challenge of insufficient data. In the future, we will design experiments or ask some medical institutions or hospitals to collect more EEG data for the study. We will also explore to use the generated antagonism network (GAN) to generate a large number of artificial EEG data to make up for this deficiency.</p>
<p>Third, the proposed model is so complex that it needs to consume a lot of computation resources and time to train the model, and it is hard to quickly verify and improve the model, as well as make real-time prediction. In the future, we will try to further simplify the structure of the model without changing its performance, and make deep research on accelerating the speed of model training and real-time application.</p>
</sec>
<sec id="S5" sec-type="conclusion">
<title>Conclusion</title>
<p>Based on the discovery of neuroscience that each region of the human brain can produce different dynamic responses to emotions, we suggest a new hierarchical EEG feature learning method by using attention mechanism and bidirectional LSTM neural network from region to global brain. A large number of experiments and verification are carried out on the DEAP and SEED datasets. The results show that the proposed R2G-ST-BiLSTM model achieves the best performance in subject-dependent and subject-independent EEG-based emotion recognition. Through experiments on several variants of the model, we compare and analyze the impact of different components of the model on its overall performance, and summarize the following advantages of the proposed model:</p>
<list list-type="simple">
<list-item>
<label>(1)</label>
<p>The BiLSTM networks are used to hierarchically learn the deep spatial correlation features within and cross each brain region. The attention mechanism is combined to weigh the contribution of each brain region to the emotion classification, which could enhance the influence of the brain region with more contribution and reduce the influence of the brain region with less contribution.</p>
</list-item>
<list-item>
<label>(2)</label>
<p>The BiLSTM networks are used to hierarchically learn the deep temporal correlation features from the EEG time sequence of each local brain region and global brain. The learned deep temporal and spatial features are connected to make the features more discriminative.</p>
</list-item>
<list-item>
<label>(3)</label>
<p>By introducing the domain discriminator, the feature difference between different subjects is reduced, and the robustness and adaptability of the model are improved.</p>
</list-item>
</list>
<p>Although the proposed model shows some advantages, there are still some problems to be solved. For example, the model is more complicated, which costs much time and computing resource for training. The whole proposed model still works as a black box, and it is difficult to explain the physical meaning represented by the learned abstract features. The complex model and limited amount of data make the model prone to overfitting. Therefore, in the future, we will further study how to improve the interpretability of the proposed model, simplify the structure of the model, and further improve the robustness and domain adaptability of the model.</p>
</sec>
<sec id="S6" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="S7">
<title>Author Contributions</title>
<p>PZ designed and implemented the R2G-ST-BiLSTM model and participated in drafted the manuscript. CM carried out the cross-subject EEG-based emotion classification experiments with R2G-ST-BiLSTM model, analyzed the experimental results, and drafted the manuscript. KZ carried out the within-subject experiments and analyzed the experimental results. WX was responsible for literature review, EEG data preprocessing and manual feature extraction. JC conceived of the study, participated in its design, and revised and proofread the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="pudiscl1" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="S8" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by the National Natural Science Foundation of China under the Project Agreement No. 61806118 and the Research Startup Fund Project of Shaanxi University of Science and Technology under Project No. 2020bj-30.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alhagry</surname> <given-names>S.</given-names></name> <name><surname>Aly</surname> <given-names>A.</given-names></name> <name><surname>Reda</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Emotion recognition based on EEG using LSTM recurrent neural network.</article-title> <source><italic>Int. J. Adv. Comput. Sci. Appl.</italic></source> <volume>8</volume> <fpage>134</fpage>&#x2013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.14569/IJACSA.2017.081046</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bottou</surname> <given-names>L.</given-names></name></person-group> (<year>2010</year>). &#x201C;<article-title>Large-scale machine learning with stochastic gradient descent</article-title>,&#x201D; in <source><italic>Proceedings of COMPSTAT&#x2019;2010</italic></source> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>177</fpage>&#x2013;<lpage>186</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-7908-2604-3_16</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L.</given-names></name></person-group> (<year>2001</year>). <article-title>Random forests.</article-title> <source><italic>Mach. Learn.</italic></source> <volume>45</volume> <fpage>5</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/A:1010933404324</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J. X.</given-names></name> <name><surname>Jiang</surname> <given-names>D. M.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. N.</given-names></name></person-group> (<year>2019b</year>). <article-title>A hierarchical bidirectional GRU Model with attention for EEG-based emotion classification.</article-title> <source><italic>IEEE Access</italic></source> <volume>7</volume> <fpage>118530</fpage>&#x2013;<lpage>118540</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2936817</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J. X.</given-names></name> <name><surname>Jiang</surname> <given-names>D. M.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. N.</given-names></name></person-group> (<year>2019a</year>). <article-title>A common spatial pattern and wavelet packet decomposition combined method for EEG-based emotion recognition.</article-title> <source><italic>J. Adv. Computat. Intell. Intell. Informat.</italic></source> <volume>23</volume> <fpage>274</fpage>&#x2013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.20965/jaciii.2019.p0274</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J. X.</given-names></name> <name><surname>Zhang</surname> <given-names>P. W.</given-names></name> <name><surname>Mao</surname> <given-names>Z. J.</given-names></name> <name><surname>Huang</surname> <given-names>Y. F.</given-names></name> <name><surname>Jiang</surname> <given-names>D. M.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. N.</given-names></name></person-group> (<year>2019c</year>). <article-title>Accurate EEG-based emotion recognition on combined features using deep conutional neural networks.</article-title> <source><italic>IEEE Access</italic></source> <volume>7</volume> <fpage>4107</fpage>&#x2013;<lpage>4115</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2908285</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J. X.</given-names></name> <name><surname>Jiang</surname> <given-names>D. M.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. N.</given-names></name> <name><surname>Zhang</surname> <given-names>P. W.</given-names></name></person-group> (<year>2020</year>). <article-title>Emotion recognition from spatiotemporal EEG representations with hybrid conutional recurrent neural networks via wearable multi-channel headset [J].</article-title> <source><italic>Comput. Commun.</italic></source> <volume>154</volume>, <fpage>58</fpage>&#x2013;<lpage>65</lpage>.</citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chuang</surname> <given-names>S. W.</given-names></name> <name><surname>Ko</surname> <given-names>L. W.</given-names></name> <name><surname>Lin</surname> <given-names>Y. P.</given-names></name> <name><surname>Huang</surname> <given-names>R. S.</given-names></name> <name><surname>Jung</surname> <given-names>T. P.</given-names></name> <name><surname>Lin</surname> <given-names>C. T.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Co-modulatory spectral changes in independent brain processes are correlated with task performance.</article-title> <source><italic>Neuroimage</italic></source> <volume>62</volume> <fpage>1469</fpage>&#x2013;<lpage>1477</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2012.05.035</pub-id> <pub-id pub-id-type="pmid">22634852</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coan</surname> <given-names>J. A.</given-names></name> <name><surname>Allen</surname> <given-names>J. J.</given-names></name></person-group> (<year>2004</year>). <article-title>Frontal EEG asymmetry as a moderator and mediator of emotion.</article-title> <source><italic>Biol. Psychol.</italic></source> <volume>67</volume> <fpage>7</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1016/j.biopsycho.2004.03.002</pub-id> <pub-id pub-id-type="pmid">15130524</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Davidson</surname> <given-names>R. J.</given-names></name></person-group> (<year>2000</year>). <article-title>Affective style, psychopathology, and resilience: brain mechanisms and plasticity.</article-title> <source><italic>Am. Psychol.</italic></source> <volume>55</volume>:<issue>1196</issue>. <pub-id pub-id-type="doi">10.1037/0003-066X.55.11.1196</pub-id> <pub-id pub-id-type="pmid">11280935</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>Y.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Bao</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>Towards interpretable deep neural networks by leveraging adversarial examples[EB/OL].</article-title> <source><italic>ArXiv</italic></source> <comment>[Preprint]. ArXiv:1901.09035</comment>,</citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>R. N.</given-names></name> <name><surname>Zhu</surname> <given-names>J. Y.</given-names></name> <name><surname>Lu</surname> <given-names>B. L.</given-names></name></person-group> (<year>2013</year>). &#x201C;<article-title>Differential entropy feature for EEG-based emotion classification</article-title>,&#x201D; in <source><italic>Proceedings of the International IEEE EMBS Conference on Neural Engineering</italic></source> (<publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>81</fpage>&#x2013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1109/NER.2013.6695876</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Etkin</surname> <given-names>A.</given-names></name> <name><surname>Egner</surname> <given-names>T.</given-names></name> <name><surname>Kalisch</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Emotional processing in anterior cingulate and medial prefrontal cortex.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>15</volume> <fpage>85</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2010.11.004</pub-id> <pub-id pub-id-type="pmid">21167765</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ganin</surname> <given-names>Y.</given-names></name> <name><surname>Ustinova</surname> <given-names>E.</given-names></name> <name><surname>Ajakan</surname> <given-names>H.</given-names></name> <name><surname>Germain</surname> <given-names>P.</given-names></name> <name><surname>Larochelle</surname> <given-names>H.</given-names></name> <name><surname>Laviolette</surname> <given-names>F.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Domain-adversarial training of neural networks.</article-title> <source><italic>J. Mach. Learn. Res.</italic></source> <volume>17</volume> <fpage>1</fpage>&#x2013;<lpage>35</lpage>.</citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garcia-Martinez</surname> <given-names>B.</given-names></name> <name><surname>Martinez-Rodrigo</surname> <given-names>A.</given-names></name> <name><surname>Alcaraz</surname> <given-names>R.</given-names></name> <name><surname>Fernandez-Caballero</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>A review on nonlinear methods using electroencephalographic recordings for emotion recognition [J].</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>12</volume>, <fpage>801</fpage>&#x2013;<lpage>820</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2018.2890636</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Genovese</surname> <given-names>C. R.</given-names></name> <name><surname>Lazar</surname> <given-names>N. A.</given-names></name> <name><surname>Nichols</surname> <given-names>T.</given-names></name></person-group> (<year>2002</year>). <article-title>Thresholding of statistical maps in functional neuroimaging using the false discovery rate.</article-title> <source><italic>Neuroimage</italic></source> <volume>15</volume> <fpage>870</fpage>&#x2013;<lpage>878</lpage>. <pub-id pub-id-type="doi">10.1006/nimg.2001.1037</pub-id> <pub-id pub-id-type="pmid">11906227</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Graves</surname> <given-names>A.</given-names></name> <name><surname>Mohamed</surname> <given-names>A. R.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). &#x201C;<article-title>Speech recognition with deep recurrent neural networks</article-title>,&#x201D; in <source><italic>Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing</italic></source>, <volume>Vol. 38</volume> <publisher-loc>Vancouver, BC</publisher-loc>, <fpage>6645</fpage>&#x2013;<lpage>6649</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.2013.6638947</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heller</surname> <given-names>W.</given-names></name> <name><surname>Nitscke</surname> <given-names>J. B.</given-names></name></person-group> (<year>1997</year>). <article-title>Regional brain activity in emotion: a framework for understanding cognition in depression.</article-title> <source><italic>Cogn. Emot.</italic></source> <volume>11</volume> <fpage>637</fpage>&#x2013;<lpage>661</lpage>. <pub-id pub-id-type="doi">10.1080/026999397379845a</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory.</article-title> <source><italic>Neural Computat.</italic></source><volume>9</volume> <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id> <pub-id pub-id-type="pmid">9377276</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jenke</surname> <given-names>R.</given-names></name> <name><surname>Peer</surname> <given-names>A.</given-names></name> <name><surname>Buss</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Feature extraction and selection for emotion recognition from EEG [J].</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>5</volume>, <fpage>327</fpage>&#x2013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2014.2339834</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jorg</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Interpretable and fine-grained visual explanations for conutional neural network</article-title>,&#x201D; in <source><italic>Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <fpage>9097</fpage>&#x2013;<lpage>9107</lpage>.</citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koelstra</surname> <given-names>S.</given-names></name> <name><surname>Muhl</surname> <given-names>C.</given-names></name> <name><surname>Soleymani</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>J. S.</given-names></name> <name><surname>Yazdani</surname> <given-names>A.</given-names></name> <name><surname>Ebrahimi</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>DEAP: a database for emotion analysis using physiological signals.</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>3</volume> <fpage>18</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1109/T-AFFC.2011.15</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>T.</given-names></name> <name><surname>Zong</surname> <given-names>Y.</given-names></name></person-group> (<year>2018a</year>). &#x201C;<article-title>A novel neural network model based on cerebral hemispheric asymmetry for EEG emotion recognition</article-title>,&#x201D; in <source><italic>Proceedings of the 27th International Joint Conference on Artificial Intelligence (IJCAI), Sweden</italic>,</source> (<publisher-loc>Palo Alto, CA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>), <fpage>1561</fpage>&#x2013;<lpage>1567</lpage>. <pub-id pub-id-type="doi">10.24963/ijcai.2018/216</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name> <name><surname>Zong</surname> <given-names>Y.</given-names></name> <name><surname>Ge</surname> <given-names>S.</given-names></name></person-group> (<year>2018b</year>). <article-title>EEG emotion recognition based on graph regularized sparse linear regression.</article-title> <source><italic>Neural Process. Lett.</italic></source> <volume>49</volume> <fpage>555</fpage>&#x2013;<lpage>571</lpage>. <pub-id pub-id-type="doi">10.1007/s11063-018-9829-1</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Zong</surname> <given-names>Y.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>From regional to global brain: a novel hierarchical spatial-temporal neural network model for EEG emotion recognition</article-title>,&#x201D; in <source><italic>Proceedings of the IEEE Transactions on Affective Computing.</italic></source> <pub-id pub-id-type="doi">10.1109/TAFFC.2019.2922912</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>Y. P.</given-names></name> <name><surname>Wang</surname> <given-names>C. H.</given-names></name> <name><surname>Jung</surname> <given-names>T. P.</given-names></name> <name><surname>Wu</surname> <given-names>T. L.</given-names></name> <name><surname>Jeng</surname> <given-names>S. K.</given-names></name> <name><surname>Duann</surname> <given-names>J. R.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>EEG-based emotion recognition in music listening.</article-title> <source><italic>IEEE Trans. Biomed. Eng.</italic></source> <volume>57</volume> <fpage>1798</fpage>&#x2013;<lpage>1806</lpage>. <pub-id pub-id-type="doi">10.1109/TBME.2010.2048568</pub-id> <pub-id pub-id-type="pmid">20442037</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lindquist</surname> <given-names>K. A.</given-names></name> <name><surname>Barrett</surname> <given-names>L. F.</given-names></name></person-group> (<year>2012</year>). <article-title>A functional architecture of the human brain: emerging insights from the science of emotion.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>16</volume> <fpage>533</fpage>&#x2013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2012.09.005</pub-id> <pub-id pub-id-type="pmid">23036719</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lindquist</surname> <given-names>K. A.</given-names></name> <name><surname>Wager</surname> <given-names>T. D.</given-names></name> <name><surname>Kober</surname> <given-names>H.</given-names></name> <name><surname>Bliss-Moreau</surname> <given-names>E.</given-names></name> <name><surname>Barrett</surname> <given-names>L. F.</given-names></name></person-group> (<year>2012</year>). <article-title>The brain basis of emotion: a meta-analytic review.</article-title> <source><italic>Behav. Brain Sci.</italic></source> <volume>35</volume> <fpage>121</fpage>&#x2013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.1017/S0140525X11000446</pub-id> <pub-id pub-id-type="pmid">22617651</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mauss</surname> <given-names>I. B.</given-names></name> <name><surname>Robinson</surname> <given-names>M. D.</given-names></name></person-group> (<year>2009</year>). <article-title>Measures of emotion: a review.</article-title> <source><italic>Cogn. Emot.</italic></source> <volume>23</volume> <fpage>209</fpage>&#x2013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.1080/02699930802204677</pub-id> <pub-id pub-id-type="pmid">19809584</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Picard</surname> <given-names>R. W.</given-names></name> <name><surname>Picard</surname> <given-names>R.</given-names></name></person-group> (<year>1997</year>). <source><italic>Affective Computing</italic></source>, <volume>Vol. 252</volume>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>MIT press</publisher-name>. <pub-id pub-id-type="doi">10.1037/e526112012-054</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Purnamasari</surname> <given-names>P. D.</given-names></name> <name><surname>Ratna</surname> <given-names>A. A. P.</given-names></name> <name><surname>Kusumoputro</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>Development of filtered bispectrum for EEG signal feature extraction in automatic emotion recognition using artificial neural networks.</article-title> <source><italic>Algorithms</italic></source> <volume>10</volume>:<issue>63</issue>. <pub-id pub-id-type="doi">10.3390/a10020063</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salama</surname> <given-names>E. S.</given-names></name> <name><surname>El-Khoribi</surname> <given-names>R. A.</given-names></name> <name><surname>Shoman</surname> <given-names>M. E.</given-names></name> <name><surname>Shalaby</surname> <given-names>M. A.</given-names></name></person-group> (<year>2018</year>). <article-title>EEG based emotion recognition using 3D conutional neural networks.</article-title> <source><italic>Int. J. Adv. Comput. Sci. Appl.</italic></source> <volume>9</volume> <fpage>329</fpage>&#x2013;<lpage>337</lpage>. <pub-id pub-id-type="doi">10.14569/IJACSA.2018.090843</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanjar</surname> <given-names>K.</given-names></name> <name><surname>Rehman</surname> <given-names>A.</given-names></name> <name><surname>Paul</surname> <given-names>A.</given-names></name> <name><surname>JeongHong</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Weight dropout for preventing neural networks from overfitting</article-title>,&#x201D; in <source><italic>Proceedings of the 2020 8th International Conference on Orange Technology (ICOT) Daegu, South Korea.</italic></source> <publisher-name>IEEE</publisher-name>, <volume>2020</volume>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1109/ICOT51877.2020.9468799</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>T.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Song</surname> <given-names>P.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name></person-group> (<year>2018</year>). <article-title>EEG emotion recognition using dynamical graph conutional neural networks [J].</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>11</volume>, <fpage>532</fpage>&#x2013;<lpage>541</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2018.2817622</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Suykens</surname> <given-names>J. A.</given-names></name> <name><surname>Vandewalle</surname> <given-names>J.</given-names></name></person-group> (<year>1999</year>). <article-title>Least squares support vector machine classifiers.</article-title> <source><italic>Neural Process. Lett.</italic></source> <volume>9</volume> <fpage>293</fpage>&#x2013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.1023/A:1018628609742</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weisstein</surname> <given-names>E. W.</given-names></name></person-group> (<year>2004</year>). <source><italic>Bonferroni Correction[J].</italic></source> Available online at: <ext-link ext-link-type="uri" xlink:href="https://mne.tools/dev/generated/mne.stats.bonferroni_correction.html">https://mne.tools/dev/generated/mne.stats.bonferroni_correction.html</ext-link>.</citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>X.</given-names></name> <name><surname>Zheng</surname> <given-names>W. L.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Lu</surname> <given-names>B. L.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Investigating gender differences of brain areas in emotion recognition using LSTM neural network</article-title>,&#x201D; in <source><italic>Proceedings of the Neural Information Processing</italic></source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>820</fpage>&#x2013;<lpage>829</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-70093-9_87</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Ramanarayanan</surname> <given-names>V.</given-names></name> <name><surname>Suendermann-Oeft</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Zechner</surname> <given-names>K.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name><etal/></person-group> (<year>2015</year>). &#x201C;<article-title>Using bidirectional LSTM recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speech</article-title>,&#x201D; in <source><italic>Proceedings of the Automatic Speech Recognition and Understanding (ASRU)</italic></source> (<publisher-loc>Scottsdale, AZ</publisher-loc>), <fpage>338</fpage>&#x2013;<lpage>345</lpage>. <pub-id pub-id-type="doi">10.1109/ASRU.2015.7404814</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Lee</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>A hierarchical positive and negative emotion understanding system based on integrated analysis of visual and brain signals.</article-title> <source><italic>Neurocomputing</italic></source> <volume>73</volume> <fpage>3264</fpage>&#x2013;<lpage>3272</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2010.04.001</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>Multichannel EEG-based emotion recognition via group sparse canonical correlation analysis.</article-title> <source><italic>IEEE Trans. Cogn. Dev. Syst.</italic></source> <volume>9</volume> <fpage>281</fpage>&#x2013;<lpage>290</lpage>. <pub-id pub-id-type="doi">10.1109/TCDS.2016.2587290</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>W. L.</given-names></name> <name><surname>Lu</surname> <given-names>B. L.</given-names></name></person-group> (<year>2015</year>). <article-title>Investigating critical frequency bands and channels for EEG-based emotion recognition with deep neural networks.</article-title> <source><italic>IEEE Trans. Auton. Ment. Dev.</italic></source> <volume>7</volume> <fpage>162</fpage>&#x2013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1109/TAMD.2015.2431497</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>W.-L.</given-names></name> <name><surname>Lu</surname> <given-names>B.-L.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Personalizing EEG-based affective models with transfer learning</article-title>,&#x201D; in <source><italic>Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI)</italic></source> (<publisher-loc>Palo Alto, CA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>), <fpage>2732</fpage>&#x2013;<lpage>2738</lpage>.</citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>W. L.</given-names></name> <name><surname>Lu</surname> <given-names>B. L.</given-names></name></person-group> (<year>2017</year>). <article-title>A multimodal approach to estimating vigilance using EEG and forehead EOG.</article-title> <source><italic>J. Neural Eng.</italic></source> <volume>14</volume>:<issue>026017</issue>. <pub-id pub-id-type="doi">10.1088/1741-2552/aa5a98</pub-id> <pub-id pub-id-type="pmid">28102833</pub-id></citation></ref>
</ref-list>
</back>
</article>
