<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2022.872311</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>The multiscale 3D convolutional network for emotion recognition based on electroencephalogram</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Su</surname> <given-names>Yun</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1661646/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Zhixuan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Xuan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Bingtao</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1289934/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ma</surname> <given-names>Huifang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Computer Science and Engineering, Northwest Normal University</institution>, <addr-line>Lanzhou</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Electronic and Information Engineering, Lanzhou Jiaotong University</institution>, <addr-line>Lanzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Varun Bajaj, PDPM Indian Institute of Information Technology, Design and Manufacturing, India</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Chang Li, Hefei University of Technology, China; Yu Zhang, Lehigh University, United States</p></fn>
<corresp id="c001">&#x002A;Correspondence: Yun Su, <email>suyun@nwnu.edu.cn</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Neuroprosthetics, a section of the journal Frontiers in Neuroscience</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>15</day>
<month>08</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>872311</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>07</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Su, Zhang, Li, Zhang and Ma.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Su, Zhang, Li, Zhang and Ma</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Emotion recognition based on EEG (electroencephalogram) has become a research hotspot in the field of brain-computer interfaces (BCI). Compared with traditional machine learning, the convolutional neural network model has substantial advantages in automatic feature extraction in EEG-based emotion recognition. Motivated by the studies that multiple smaller scale kernels could increase non-linear expression than a larger scale, we propose a 3D convolutional neural network model with multiscale convolutional kernels to recognize emotional states based on EEG signals. We select more suitable time window data to carry out the emotion recognition of four classes (low valence vs. low arousal, low valence vs. high arousal, high valence vs. low arousal, and high valence vs. high arousal). The results using EEG signals in the DEAP and SEED-IV datasets show accuracies for our proposed emotion recognition network model (ERN) of 95.67 and 89.55%, respectively. The experimental results demonstrate that the proposed approach is potentially useful for enhancing emotional experience in BCI.</p>
</abstract>
<kwd-group>
<kwd>BCI</kwd>
<kwd>emotion recognition</kwd>
<kwd>EEG</kwd>
<kwd>3D CNN</kwd>
<kwd>spatiotemporal features</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<contract-sponsor id="cn002">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="8"/>
<equation-count count="8"/>
<ref-count count="47"/>
<page-count count="14"/>
<word-count count="9376"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>Automatic EEG-based emotion recognition has been a research hotspot in the field of BCI and human&#x2013;computer interaction over the past decade. Efficient emotion recognition methods based on EEG can prompt BCI to build a harmonious human-computer interaction environment, which can promote a natural, convenient, and friendly experience as communication between people (<xref ref-type="bibr" rid="B43">Zhang et al., 2015</xref>; <xref ref-type="bibr" rid="B32">Xie et al., 2019</xref>).</p>
<p>In general, there are two ways to recognize emotion, i.e., through non-physiological and physiological signals. Non-physiological signals (such as facial expressions, speech, gestures, etc.) can be artificially controlled (<xref ref-type="bibr" rid="B35">Yao, 2014</xref>). However, physiological signals (such as electroencephalogram (EEG), electrocardiograph (ECG), electromyography (EMG), magnetoencephalography (MEG), functional near-infrared spectroscopy (fNIRS) etc.) can show reliable and natural emotions without subjective control (<xref ref-type="bibr" rid="B4">Cai et al., 2018</xref>). EEG and MEG are the types of electrical signal produced by the brain that provides very useful information relating to emotional activity of the brain. Moreover, EEG and MEG have good temporal resolution and both are non-invasive. Some researchers (<xref ref-type="bibr" rid="B33">Xing et al., 2019</xref>; <xref ref-type="bibr" rid="B27">Sun et al., 2020</xref>) found that the MEG signal and fNIRS signal can realize emotional recognition, and could obtain the accuracy of high binary classification. However, MEG mainly reflects the inner structure of the brain. Nevertheless, EEG is used to check the function of the brain and mainly through brain waves reflecting the mood of the brain. As a result, EEG-based emotion recognition methods have become popular in current research.</p>
<p>Currently, traditional machine learning methods (<xref ref-type="bibr" rid="B38">Yoon and Chuang, 2013</xref>; <xref ref-type="bibr" rid="B14">Li et al., 2018</xref>) can effectively recognize emotions but require manual feature extraction and only consider the independence of a single feature in time or space. The two-dimensional convolutional neural network (2D-CNN) of deep learning can solve these problems. However, emotion recognition requires taking into account not only the time dependence between data points but also the spatial relevance between different electrodes of EEG signals (<xref ref-type="bibr" rid="B36">Yea-Hoon et al., 2018</xref>). In contrast to 2D-CNN (<xref ref-type="bibr" rid="B18">Mei and Xu, 2017</xref>), the emotion recognition method based on three-dimensional convolutional neural networks (3D-CNN) can meet these needs (<xref ref-type="bibr" rid="B20">Salama et al., 2018</xref>; <xref ref-type="bibr" rid="B44">Zhao et al., 2020</xref>). The 3D-CNN models can automatically extract spatiotemporal features. The existing emotion recognition model has achieved high accuracy, while most researchers believe that multiple smaller-scale kernels have the ability to increase non-linear expression more than a larger kernel. Therefore, how to define the convolutional kernel size in the convolutional network is still an interesting topic in emotion recognition research.</p>
<p>In this paper, we propose a four-class emotion recognition method based on a multiscale convolutional kernel 3D network, in which EEG-based emotional states can be efficiently recognized. First, we located the spatial position of the EEG signal electrode according to the 10&#x2013;20 system diagram, the positional relationship between the positioning electrodes, and retained the spatial information of the EEG. The emotional recognition model based on three-dimensional EEG is generally used in the size of a consistent convolutional kernel. However, this paper attempts to use different smaller sizes of convolutional kernels, which are expected to increase the non-linear features of the data and the amount of data available. In addition, we join the double linear convolutional structure to the emotion recognition network model (ERN), and the EEG data are analyzed in parallel, thereby obtaining efficient recognition results.</p>
<p>According to existing studies, most researchers have applied two lengths of time windows, i.e., 1 s (1 s) and 2 s (2 s). To find the most suitable time window length for the ERN model, we compare the classification performance of 1 and 2 s time window lengths. The experimental results have shown that multiscale convolutional kernels with suitable time window lengths are more effective, which can improve the accuracy for emotion recognition. In summary, the main contributions of this paper are as follows. (1) We optimized the original dataset, designed a repositioned electrode topology, and constructed a 3D dataset for the model. (2) We enriched the emotion recognition method based on EEG and constructed a multiscale convolutional kernel 3D-CNN model to achieve more efficient emotion recognition performance.</p>
<p>The paper is structured as follows: related research is discussed, followed by consideration of the methodology adopted in our work. Experimentation and evaluation are addressed, and a discussion with results derived from the experimentation is presented. The paper closes with concluding observations and consideration of future work.</p>
</sec>
<sec id="S2">
<title>Related work</title>
<p>Currently, a variety of traditional machine learning-based emotion recognition methods have been documented in the literature, which also confirms the effectiveness and accuracy of traditional machine learning on emotion classification. However, machine learning-based emotion recognition methods require the specifically detailed design of classification models and manual extraction of temporal or spatial emotion features of EEG signals. For example, the traditional classifiers used in the literature include the support vector machine (SVM), and k-nearest neighbors (KNN) (<xref ref-type="bibr" rid="B7">Jenke et al., 2014</xref>). However, the use of traditional machine learning requires the manual extraction of relevant emotion features, limited to the temporal or frequency domains, with domain knowledge barriers and timeliness problems.</p>
<p>Emotion recognition methods based on deep learning can solve these problems. The deep network can extract different types of features at the same time, with the advantages of automatic detection features, and solve the dependence of artificial feature extraction. For example, <xref ref-type="bibr" rid="B15">Lin et al. (2017)</xref> proposed an emotion recognition method based on CNN, converting EEG data from the signal format into an image format containing time domain and frequency domain information, combined with the characteristics of other physiological signals, which were input into the pretrained AlexNet (<xref ref-type="bibr" rid="B10">Krizhevsky et al., 2012</xref>) network model for emotion recognition. <xref ref-type="bibr" rid="B11">Kwon et al. (2018)</xref> also proposed a sentiment classification method for extracting features based on the 2D CNN model, which were preprocessed before convoluting EEG signals by wavelet transformation considering both the time and frequency domains to improve the recognition performance.</p>
<p>In addition to CNN model, there are other methods effectively identify emotions. <xref ref-type="bibr" rid="B16">Liu et al. (2018)</xref> proposed that the Residual Network-50 (ResNet-50) model can automatically learn deep semantic EEG information and classify the new features of the fusion of linear-frequency cepstral coefficients (LFCC). <xref ref-type="bibr" rid="B33">Xing et al. (2019)</xref> used a stack auto-encoder (SAE) to build and solve the linear EEG mixing model and the emotion timing model based on the long short-term memory recurrent neural network (LSTM-RNN). <xref ref-type="bibr" rid="B17">Liu et al. (2020)</xref> proposed a capsule network based on the multi-level feature boot, which can recognize multi-electrode EEG emotions. In the same year, Tao (<xref ref-type="bibr" rid="B29">Tao et al., 2020</xref>) proposed a convolutional recursive neural network (ACRNN) based on the attention mechanism, which can extract more discriminant features from the EEG signal, and improve the accuracy of emotional identification.</p>
<p>However, these models ignore the spatial structure of the EEG and the variations and distortion of the electrodes in each dimension. With the study of deep learning neural networks, methods for extracting spatiotemporal features have been proposed. <xref ref-type="bibr" rid="B34">Yang et al. (2018)</xref> implemented a hybrid neural network integrating a CNN and a recurrent neural network (RNN) so that network models could extract and integrate spatiotemporal features. <xref ref-type="bibr" rid="B31">Wang et al. (2018)</xref> proposed a simple and efficient preprocessing method that converts multiple electrodes of EEG signals into electrode topological maps containing topological location information. <xref ref-type="bibr" rid="B2">An et al. (2021)</xref> proposed an EEG emotion recognition algorithm based on 3D feature fusion and a convolutional auto-encoder (CAE). <xref ref-type="bibr" rid="B44">Zhao et al. (2020)</xref> proposed a 3D-CNN model to automatically extract the spatiotemporal features of EEG signals, introduced the preprocessing method for baseline signal and electrode topology relocation, compared the performance of the 2D convolutional kernel and 3D convolutional kernel in detail, and showed that the 3D-CNN model was more advantageous.</p>
<p>Different convolutional network recognition models set different convolutional kernel sizes, while most researchers (<xref ref-type="bibr" rid="B28">Szegedy et al., 2016</xref>) believe that multiple smaller scale kernels have the ability to increase non-linear expression more than a larger kernel. The different sizes of kernels in the network can increase the number of features that can be used and improve model performance. In this paper, we propose a four-class emotion recognition method based on a multiscale convolutional kernel 3D network, in which EEG-based emotional states can be efficiently recognized. The experimental results on the DEAP and SEED-IV datasets show that the proposed model has preferable performance than the other existing models in terms of recognition accuracy.</p>
</sec>
<sec id="S3" sec-type="materials|methods">
<title>Materials and methods</title>
<p>The experimental process proposed in this paper is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. In <xref ref-type="fig" rid="F1">Figure 1</xref>, first, the original EEG signals are preprocessed. Then, the preprocessed data are converted from the 2D form to the 3D format and divided into two kinds of datasets: training data and testing data. Finally, we input preprocessed data and evaluated our ERN model by the recognition results of the four classes of labels.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Overview of the proposed approach.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g001.tif"/>
</fig>
<sec id="S3.SS1">
<title>Processing</title>
<p>To improve the recognition accuracy, it is necessary to preprocess the raw EEG data. First, the brain electrode data are selected and subsampled, and the noise artifacts are removed using a bandpass frequency filter of 4.0&#x2013;45.0 Hz. A preprocessing method for the baseline EEG signal will affect the recognition result (<xref ref-type="bibr" rid="B34">Yang et al., 2018</xref>). Therefore, the specific practice of the data baseline signal processing is as follows. First, the first 3 s of baseline signals are extracted from all electrodes c of a single subject, and then cut into a fragment of the N-section fixed length L, thereby obtaining the N &#x00D7; C &#x00D7; L matrix. Then, calculate the average of this N &#x00D7; C &#x00D7; L matrix, obtain the Z matrix, and the structure of the Z matrix is C &#x00D7; L. The last 60 s of the signals are divided into M fragments to obtain the matrix of M &#x00D7; C &#x00D7; L. Then the matrix of M &#x00D7; C &#x00D7; L subtracts the average matrix Z. N and M are the number of data segments, C is the number of electrodes, and L is the data length. This calculation step can obtain all the data of a single subject and be repeated, and we can obtain all the pretreatment data.</p>
<p>The 32 electrodes of the EEG signals in the dataset are repositioned to a 2D electrode topology based on the International 10&#x2013;20 System Diagram to acquire the spatial information of the EEG. During the recognition of emotional type based on EEG, <xref ref-type="bibr" rid="B47">Zhong et al. (2020)</xref> confirmed that both the position of signals acquisition and the interaction of EEG electrode position are conducive to improving the accuracy of emotion recognition based on EEG signals. Therefore, to retain the spatial information of the EEG, we located the spatial position of the EEG signal electrode according to the 10&#x2013;20 system diagram (the positional relationship between the positioning electrodes), and retained the spatial information of the EEG. Based on the topological location information of the vulnerable electrodes during the original EEG emotion analysis experiment, Zhong and An (<xref ref-type="bibr" rid="B47">Zhong et al., 2020</xref>; <xref ref-type="bibr" rid="B2">An et al., 2021</xref>) proposed a solution by repositioning the 32 electrodes of the EEG signals in the dataset to the 2D electrode topology based on the international 10&#x2013;20 system diagram, thereby preserving spatial information among electrodes. In this paper, 1-dimensional (1D) electrodes of the obtained dataset are repositioned into a 2D electrode topology.</p>
<p>As shown in <xref ref-type="fig" rid="F2">Figure 2A</xref>, we choose the 32-electrode EEG data of the dataset, which is located in the International 10&#x2013;20 system diagram (<xref ref-type="bibr" rid="B22">Sharbrough et al., 1991</xref>). According to the farthest distance between the two electrodes in <xref ref-type="fig" rid="F2">Figure 2A</xref>, we set the size of the two-dimensional matrix and the size of the two-dimensional matrix is 9 &#x00D7; 9. Then, the selected 32 EEG signals are mapped to the 9 &#x00D7; 9 matrix. In <xref ref-type="fig" rid="F2">Figure 2B</xref>, the positioning of each electrode is located based on the positional relationship between the various electrodes in <xref ref-type="fig" rid="F2">Figure 2A</xref>. The blank position is represented as a topological position of the unselected physiological signal. Therefore, unused topologies are set to zero in the 9 &#x00D7; 9 matrix, and the matrix is normalized.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Brain electrode distribution diagram. <bold>(A)</bold> Brain electrode distribution in 10&#x2013;20 system diagram. <bold>(B)</bold> The corresponding matrix.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g002.tif"/>
</fig>
</sec>
<sec id="S3.SS2">
<title>Model</title>
<p>We design the multiscale convolutional kernel 3D-CNN model based on the final obtained 3D EEG dataset. The reasons for the use of this network are as follows: First, the advantage of the convolutional network (<xref ref-type="bibr" rid="B41">Zeiler and Fergus, 2014b</xref>) is that it can calculate the eigenvalue rather than the original value with no need for the accurate mathematical expression between inputs and outputs. Second, the convolutional network can avoid the problem of gradient loss when reverse propagation occurs in the BP neural network. The CNN can be trained in parallel, which reduces the complexity of the network. In particular, the network can directly input multi-dimensional data directly, which avoids the complexity of data reconstruction during feature extraction and classification. The flexibility of the three-dimensional convolutional kernel is higher than that of the two-dimensional convolutional kernel, which helps to learn the advanced representation of learning information (<xref ref-type="bibr" rid="B30">Tran et al., 2015</xref>). The controllable range of the three-dimensional convolutional kernel is expanded to the spatial domain, which can utilize the interaction between the electrodes and increase the identification ability of the model.</p>
<p>The detailed architecture of the ERN model is shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. The architecture of this 3D-CNN model consists of three convolutional layers, with the first convolutional layer implemented in parallel with the second convolutional layer in the model. The kernel size is 3 &#x00D7; 3 &#x00D7; 4 in the first convolutional and the third convolutional layers, where spatiotemporal features are generated by the local spatial topology of 3 &#x00D7; 3 and fragments of the temporally sampled point 4. The kernel size is 3 &#x00D7; 3 &#x00D7; 5 in the second convolutional layer and combines advanced spatial features via a local spatial topology of 3 &#x00D7; 3 and a temporal sampling point of 5. The first convolutional layer and the second convolutional layer are used to calculate the feature images, and the obtained feature images are superimposed to obtain a new feature image. Multiple 3 &#x00D7; 3 kernels have more non-linear functions than a larger convolutional kernel, which increases the non-linear expression and makes the judgment function more efficient. Selecting more small convolutional kernels favors more accurate emotion recognition. Under the conditions of ensuring the same kernel, the depth of the network is improved, the parameters of the model are reduced, and the effect of the neural network is improved to some extent.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>The framework of the ERN model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g003.tif"/>
</fig>
<p>We provide a detailed description in <xref ref-type="fig" rid="F4">Figure 4</xref> to visualize the feature representation learned by the hidden layers of the ERN model. Taking 3 &#x00D7; 3 &#x00D7; 5 as an example, the step size of the convolutional kernel is 1, the format of the input data is 9 &#x00D7; 9 &#x00D7; 128, and the resulting feature format is 9 &#x00D7; 9 &#x00D7; 64. The figure demonstrates the detailed process of the convolutional operation. Each convolutional kernel is used to compute features and moves with a fixed step on the dataset. We can calculate the size of the features according to the equation (1). <italic>I<sub>n</sub></italic>(<italic>n = 1,2,3</italic>) is defined as the size of the input convolutional layer data, <italic>O<sub>n</sub></italic>(<italic>n = 1,2,3</italic>) is the size of the output convolutional feature, <italic>K<sub>n</sub></italic>(<italic>n = 1,2,3</italic>) is the convolutional kernel size, n is one of the three dimensions, <italic>N</italic> is the number of convolutional kernels, S is the moving step of the convolutional kernel, and <italic>P</italic> is padding value; this article set <italic>P</italic> is 1. The size of the three dimensions of the feature shown in <xref ref-type="fig" rid="F4">Figure 4</xref> can be calculated separately by the equation (1). Equation (1) can calculate the size of the 3D features map.</p>
<disp-formula id="S3.E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:msub><mml:mi>O</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>P</mml:mi></mml:mrow></mml:mrow><mml:mi>S</mml:mi></mml:mfrac><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow></mml:math></disp-formula>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>The feature representation learned in ERN. (Overview of the proposed approach detailed convolutional operations).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g004.tif"/>
</fig>
<p>A 3D maximum pooling layer is set behind each convolutional layer, and the kernel size is 1 &#x00D7; 1 &#x00D7; 2. The maximum pooling layer (<xref ref-type="bibr" rid="B21">Scherer et al., 2010</xref>) is used to extract the features more efficiently here; it can reduce the quantity of data on the time dimension and improve the robustness of the extracted features and provides a better generalization. The first maximum pooling layer and the second maximum pooling layer further squeeze the extracted spatiotemporal features to generate advanced high-level spatiotemporal features. The last pooling layer is followed by a fully connected layer, and the Softmax layer is deployed as the output. In the experiment, the input data size of the model is 9 &#x00D7; 9 &#x00D7; 128, where 9 &#x00D7; 9 is the 2D electrode topology and 128 is the number of continuous-time sampling points for one treatment. The number of feature maps for the last convolutional layer is 64, passing the 64 feature maps to the fully connected layer, which maps the input as vectors. The N in its output represented the number of labels in the task. The empty is set to zero in each convolutional layer to prevent the loss of information from the input data, and the ReLU activation function is used after each convolutional layer.</p>
</sec>
</sec>
<sec id="S4">
<title>Experiment</title>
<p>We test the model in the public databases DEAP (<xref ref-type="bibr" rid="B9">Koelstra et al., 2012</xref>) and SEED-IV (<xref ref-type="bibr" rid="B3">BCMI, 1994</xref>). We use the PyTorch framework (<xref ref-type="bibr" rid="B26">Chaudhary et al., 2020</xref>) to implement this model and deploy it on the GeForce RTX 3060. The learning rate is set to 0.001 with the Adam AdaDelta Optimizer, and the probability of the dropout operation is set to 0.6. We use 10-fold cross-validation to evaluate the performance of the ERN model. The average accuracy of the 10-fold validation processes is taken as the final result.</p>
<sec id="S4.SS1">
<title>Processing</title>
<sec id="S4.SS1.SSS1">
<title>DEAP dataset</title>
<p>The multimodal DEAP dataset is an open multimodal standardized dataset used to study the analysis of human emotional states. The dataset includes the 32 electrodes of EEG signals and the 8 electrodes of peripheral physiological signals when subjects watch music videos. After watching a video, subjects scored each video based on four psychological scales of arousal, valence, liking, and dominance. We select 32 electrodes of EEG signal data from the dataset for the analysis of the human emotional state.</p>
<p>The preprocessing step is as follows: First, the data are downsampled from 512 to 128 Hz, and then a bandpass frequency filter of 4.0&#x2013;45.0 Hz to remove noise artifacts. At this time, the processed dataset data structure is 40 &#x00D7; 32 &#x00D7; 8,064 (video number &#x00D7; EEG electrode number &#x00D7; signal data), of which 8,064 signal data contained 384 baseline signals. The DEAP dataset is divided into two parts, as shown in <xref ref-type="table" rid="T1">Table 1</xref>. The data matrix refers to the EEG of 40 electrodes observed when each subject watched music videos. The label matrix refers to the four types of labels after each subject watches videos: arousal, valence, dominance, and liking.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>The DEAP dataset and SEED-IV dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Matrix name</td>
<td valign="top" align="left">Matrix structure representation</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="2">DEAP Dataset</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Data</italic><sub>40 &#x00D7; 40 &#x00D7; 8,064</sub></td>
<td valign="top" align="left"><italic>Data</italic><sub><italic>video</italic> &#x00D7; <italic>channel</italic> &#x00D7; <italic>fixedpointintime</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left"><italic>Label</italic><sub>40 &#x00D7; 4</sub></td>
<td valign="top" align="left"><italic>Label</italic><sub><italic>video</italic> &#x00D7; <italic>value</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left" colspan="2">SEED-IV Dataset</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Data</italic><sub>15 &#x00D7; 62 &#x00D7; &#x002A;</sub></td>
<td valign="top" align="left"><italic>Data</italic><sub><italic>video</italic> &#x00D7; <italic>channel</italic> &#x00D7; <italic>fixedpointintime</italic></sub></td>
</tr>
<tr>
<td valign="top" align="left"><italic>Label</italic><sub>15 &#x00D7; 1</sub></td>
<td valign="top" align="left"><italic>Label</italic><sub><italic>video</italic> &#x00D7; <italic>value</italic></sub></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In this dataset, each video of each stimulus is 60 s so that the first 3 s is the baseline signals of the unstimulated, and the last 60 s is the signals of the stimulus in the 63 s signals of each stimulus trial. Therefore, we need to carry out baseline signal processing for each trial signal after preprocessing. The processing step is as follows: For each trial signal (32 &#x00D7; 8,064), cut the baseline signal (32 &#x00D7; 384) of the first 3 s to 3 segments (32 &#x00D7; 128) and calculate the mean value of the baseline signals (32 &#x00D7; 128). Then, the signal data of the last 60 s are cut into 60 segments (32 &#x00D7; 128), and the mean of the baseline signal is substracted and merged to obtain the processed signal (32 &#x00D7; 7,680). Next, each electrode needs to be repositioned to the two-dimensional topological location to learn the spatial properties of the data. 128, 384, 7,680, and 8,064 are the time points and 32 stands for the number of electrodes. To extract the spatiotemporal features, the EEG data are mapped into a 9 &#x00D7; 9 matrix based on the International 10&#x2013;20 system diagram. Finally, the matrix is cut into fragments with a length of 1 s (9 &#x00D7; 9 &#x00D7; &#x002A;), and the 3D electrode topology was obtained (7,680 9 &#x00D7; 9 &#x00D7; &#x002A;), of which the symbol &#x201C;&#x002A;&#x201D; represents the size of the time window and 7,680 represents the number of matrixes.</p>
<p>Specifically, for the selected labels, the valence describes the degree of pleasure associated with the stimulus, represented by continuous values ranging from 1 (negative) to 5 (neutral) to 9 (positive). Arousal represents the degree of waking to the stimulus with the same range, with 1 and 9 indicating negative and positive, respectively. As shown in <xref ref-type="table" rid="T2">Table 2</xref>, we set the distribution of 4 label values based on EEG arousal and valence markers: low valence vs. low arousal (LVLA), low valence vs. high arousal (LVHA), high valence vs. low arousal (HVLA), and high valence vs. high arousal (HVHA). As shown in <xref ref-type="table" rid="T2">Table 2</xref>, we set the values of the four types of labels based on arousal and valence and set 5 as the threshold. After processing, the label structure is 40 &#x00D7; 1 (number of videos &#x00D7; label value).</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>Label values in the DEAP.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Label</td>
<td valign="top" align="center">LVLA</td>
<td valign="top" align="center">LVHA</td>
<td valign="top" align="center">HVLA</td>
<td valign="top" align="center">HVHA</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Valence</td>
<td valign="top" align="center">&#x2264; 5</td>
<td valign="top" align="center">&#x2264;5</td>
<td valign="top" align="center">&#x003E; 5</td>
<td valign="top" align="center">&#x003E;5</td>
</tr>
<tr>
<td valign="top" align="left">Arousal</td>
<td valign="top" align="center">&#x2264; 5</td>
<td valign="top" align="center">&#x003E; 5</td>
<td valign="top" align="center">&#x2264; 5</td>
<td valign="top" align="center">&#x003E; 5</td>
</tr>
<tr>
<td valign="top" align="left">Value</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="S4.SS1.SSS2">
<title>SEED-IV dataset</title>
<p>We also use the SEED-IV dataset as a standardized dataset to study the model recognition performance of this paper, which is a well-formed multimodal dataset for emotion recognition. In the SEED-IV dataset, a total of 15 subjects participated in the experiment. For each subject, the test, respectively, was performed on three different days and each test contained 24 trials. In each trial, his or her EEG signals are saved when the subject watches each film clip.</p>
<p>The preprocessing step is as follows: First, the same 32-electrode EEG data as DEAP in the SEED-IV dataset were selected to analyze the human emotional state. Then, the data are downsampled from 1,000 Hz to 128 Hz using a bandpass frequency filter of 4.0&#x2013;45.0 Hz to remove noise artifacts. The SEED-IV dataset is divided into two parts, as shown in <xref ref-type="table" rid="T1">Table 1</xref>. The data matrix refers to the physiological data of 62 electrodes observed that include the EEG signal and peripheral physiological signal. In the data matrix, the length of the movie clips resulted in the different lengths of the EEG data in each trial. The label data matrix refers to the four types of labels played when a subject watches the film clips: happy, sad, neutral, and fear.</p>
<p>In the SEED-IV dataset, the length of the data varies in each trial. Therefore, we need to select the data length suitable for the model after preprocessing in each stage of each subject and obtain 15 matrixes, each with 32 rows and 128 columns (32 &#x00D7; 128), and 15 is the number of movie clips. In each 32 &#x00D7; 128, the 32 is the number of EEG electrodes and 128 is the data of 1 s. Then, 32 EEG electrodes are mapped into the matrix with 9 rows and 9 columns, and 15 3D-matrixes (9 &#x00D7; 9 &#x00D7; 128) are obtained. Finally, we combine the data from three stages of 15 subjects and obtain 675 (15 &#x00D7; 3 &#x00D7; 15, subject number &#x00D7; stage number &#x00D7; video number) segments of the total dataset (9 &#x00D7; 9 &#x00D7; 128). After processing, the label structure is 15 &#x00D7; 1 (number of videos &#x00D7; label value).</p>
</sec>
</sec>
<sec id="S4.SS2">
<title>Optimal time window in model</title>
<p>We test and select the time window length more suitable for the experiment to obtain the best recognition result of the model and apply the 1 s time window. To improve the accuracy of the data in the experimental input model, one of the solutions of this paper is the application of the time window. The available window size is not necessarily fixed, it can constantly expand until certain conditions are met, it can be constantly reduced until a minimum window to meet the conditions is found, and it can be a fixed size.</p>
<p>In the literature, people (<xref ref-type="bibr" rid="B44">Zhao et al., 2020</xref>) confirmed that the average classification accuracy of a 1 s period based on EEG was superior to other periods, and selected the 1 s length as the most appropriate time window length. However, someone (<xref ref-type="bibr" rid="B5">Candra et al., 2015</xref>) believed that a length of 2 s was the most appropriate time window length. To address this problem, we compare all EEG classification performances at two different time window lengths: 1 and 2 s. In <xref ref-type="fig" rid="F5">Figure 5</xref>, the left side of the figure shows the recognition result trend of 1 s time window data, and the right side of the figure shows the recognition result trend of 2 s time window data. The classification accuracy of the 1 s time window is better than that of the 2 s time window with the same number of iterations, and no overfitting occurs in the first 500 epochs. Therefore, we compare the results of the different time windows, select the time window that is more suitable for the model, and improve the accuracy of the experiment.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Accuracy comparison of different time window lengths. <bold>(A)</bold> Comparison of different time windows in the DEAP without overlapped time window. <bold>(B)</bold> Comparison of different time windows in the SEED-IV without overlapped time window. <bold>(C)</bold> The recognition of overlapped window in the DEAP.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g005.tif"/>
</fig>
<p>Since we have chosen the more suitable size of the time window for the ERN model, and the next step is the analysis of whether the ERN model needs to set the overlapping window. Taking the DEAP dataset as an example, we set 0.5 s overlapping windows to process the dataset. Depending on the identification result after using the overlapping window, there may be multiple problems. First, the baseline processing method requires each second of data to subtract the mean of the baseline data. This method can disconnect the time continuity of data and eliminate the advantage of high time resolution. Second, the amount of data doubled after using the 0.5 s overlapping time window. About another data set SEED-IV, the data length is 128 after the process of downsampling. The pre-processed data in the SEED-IV have high time continuity, so there are no setup 0.5 s overlapping windows to process this dataset. The overlapping window processing generates a large amount of available data. But in the same number of iteration, the recognition efficiency by overlapping window processing is much lower as shown in <xref ref-type="fig" rid="F5">Figure 5C</xref>. To achieve out high-efficiency recognition accuracy, we have not used overlapping windows.</p>
</sec>
</sec>
<sec id="S5" sec-type="results|discussions">
<title>Results and discussions</title>
<sec id="S5.SS1">
<title>Accuracy</title>
<p>To estimate the accuracy of the emotion recognition for the ERN model, we conducted a comparative analysis that compared the results obtained from our proposed approach applied over the DEAP and SEED-IV datasets with other methods. All of the classification procedures were conducted under 10-fold cross-validation (<xref ref-type="bibr" rid="B39">Yu et al., 2015</xref>), and the average value of 10 accuracies, <italic>F</italic><sub>1</sub> values and AUC (the area under the ROC curve) values were calculated as the evaluation standard of model accuracy. In the process of training, the optimal weight parameters are obtained and set by random training data to avoid the overfitting problem caused by optimization.</p>
<p><xref ref-type="table" rid="T3">Table 3</xref> shows the 10-fold cross-validation accuracy in the DEAP and SEED-IV datasets. The 2D-CNN only utilizes high features of EEG time resolution in EEG and neglects space information. In contrast to 2D-CNN (<xref ref-type="bibr" rid="B18">Mei and Xu, 2017</xref>), the emotion recognition method based on three-dimensional convolutional neural networks (3D-CNN) can meet the need (<xref ref-type="bibr" rid="B20">Salama et al., 2018</xref>; <xref ref-type="bibr" rid="B44">Zhao et al., 2020</xref>) for spatial information. Not only does the 3D-CNN extract the time characteristics based on the 1 s time window, but it can also obtain the spatial feature between the electrodes. The previous literature report and the ERN model experiments in this article show that our model has the ability to improve the accuracy of EEG-based emotion recognition.</p>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>Ten-fold cross-validation accuracy in the DEAP and SEED-IV datasets.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Fold ID</td>
<td valign="top" align="center">DEAP</td>
<td valign="top" align="center">SEED-IV</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Fold 1</td>
<td valign="top" align="center">94.78%</td>
<td valign="top" align="center">88.56%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 2</td>
<td valign="top" align="center">95.83%</td>
<td valign="top" align="center">85.07%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 3</td>
<td valign="top" align="center">95.42%</td>
<td valign="top" align="center">92.54%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 4</td>
<td valign="top" align="center">97.08%</td>
<td valign="top" align="center">94.78%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 5</td>
<td valign="top" align="center">93.87%</td>
<td valign="top" align="center">94.04%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 6</td>
<td valign="top" align="center">96.67%</td>
<td valign="top" align="center">82.29%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 7</td>
<td valign="top" align="center">96.78%</td>
<td valign="top" align="center">85.46%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 8</td>
<td valign="top" align="center">93.75%</td>
<td valign="top" align="center">89.48%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 9</td>
<td valign="top" align="center">95.85%</td>
<td valign="top" align="center">90.78%</td>
</tr>
<tr>
<td valign="top" align="left">Fold 10</td>
<td valign="top" align="center">96.67%</td>
<td valign="top" align="center">92.47%</td>
</tr>
<tr>
<td valign="top" align="left">Mean</td>
<td valign="top" align="center">95.67%</td>
<td valign="top" align="center">89.55%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In <xref ref-type="table" rid="T4">Table 4</xref>, the average accuracy of our model for four emotion classes is up to 95.67% in DEAP and 89.55% in SEED-IV datasets, which is higher than previously reported models where 93.72 and 87.71% accuracies were achieved. It is known that the model based on deep learning proposed by <xref ref-type="bibr" rid="B19">Qiu et al. (2018)</xref> and <xref ref-type="bibr" rid="B44">Zhao et al. (2020)</xref> currently has the best performance in the DEAP dataset and SEED-IV dataset. However, the results of the ERN model are approximately 1.95 and 1.84% higher than those models. Compared with the CNN model (<xref ref-type="bibr" rid="B18">Mei and Xu, 2017</xref>; <xref ref-type="bibr" rid="B20">Salama et al., 2018</xref>; <xref ref-type="bibr" rid="B29">Tao et al., 2020</xref>; <xref ref-type="bibr" rid="B44">Zhao et al., 2020</xref>; <xref ref-type="bibr" rid="B37">Yin et al., 2021</xref>) and compared with other methods (<xref ref-type="bibr" rid="B40">Zangeneh et al., 2019</xref>; <xref ref-type="bibr" rid="B25">Song et al., 2020</xref>; <xref ref-type="bibr" rid="B2">An et al., 2021</xref>; <xref ref-type="bibr" rid="B13">Li S. et al., 2021</xref>) in the DEAP dataset, our model adopts a simpler and more efficient structure and has the best performance. Our model has higher speed efficiency and better identification performance than other models (<xref ref-type="bibr" rid="B19">Qiu et al., 2018</xref>; <xref ref-type="bibr" rid="B45">Zheng et al., 2019</xref>; <xref ref-type="bibr" rid="B1">Acharya et al., 2020</xref>) in the SEED-IV dataset.</p>
<table-wrap position="float" id="T4">
<label>TABLE 4</label>
<caption><p>Comparison of ERN model with previous studies.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Research</td>
<td valign="top" align="center">Year</td>
<td valign="top" align="center">Method</td>
<td valign="top" align="center">Accuracy</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="4">DEAP dataset</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B18">Mei and Xu</xref></td>
<td valign="top" align="center">2017</td>
<td valign="top" align="center">2D-CNN</td>
<td valign="top" align="center">73.10%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B20">Salama et al.</xref></td>
<td valign="top" align="center">2018</td>
<td valign="top" align="center">3D-CNN</td>
<td valign="top" align="center">88.49%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B40">Zangeneh et al.</xref></td>
<td valign="top" align="center">2019</td>
<td valign="top" align="center">HcF+KNN+MSVM</td>
<td valign="top" align="center">86.01%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B25">Song et al.</xref></td>
<td valign="top" align="center">2020</td>
<td valign="top" align="center">DGCNN</td>
<td valign="top" align="center">90.4%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B44">Zhao et al.</xref></td>
<td valign="top" align="center">2020</td>
<td valign="top" align="center">3D-CNN</td>
<td valign="top" align="center">93.53%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B29">Tao et al.</xref></td>
<td valign="top" align="center">2020</td>
<td valign="top" align="center">ACRNN</td>
<td valign="top" align="center">93.72%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B13">Li S. et al.</xref></td>
<td valign="top" align="center">2021</td>
<td valign="top" align="center">The binary gray wolf optimization algorithm+SVM</td>
<td valign="top" align="center">90.48%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B37">Yin et al.</xref></td>
<td valign="top" align="center">2021</td>
<td valign="top" align="center">GCNN+LSTM</td>
<td valign="top" align="center">90.53%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B2">An et al.</xref></td>
<td valign="top" align="center">2021</td>
<td valign="top" align="center">3D Feature Fusion+CAE</td>
<td valign="top" align="center">90.76%</td>
</tr>
<tr>
<td valign="top" align="left">Our model</td>
<td valign="top" align="center">2021</td>
<td valign="top" align="center">3D-CNN</td>
<td valign="top" align="center">95.67%</td>
</tr>
<tr>
<td valign="top" align="left" colspan="4">SEED-IV Dataset</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B45">Zheng et al.</xref></td>
<td valign="top" align="center">2015</td>
<td valign="top" align="center">DBN</td>
<td valign="top" align="center">86.08%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B19">Qiu et al.</xref></td>
<td valign="top" align="center">2018</td>
<td valign="top" align="center">CAN</td>
<td valign="top" align="center">87.71%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B45">Zheng et al.</xref></td>
<td valign="top" align="center">2019</td>
<td valign="top" align="center">EmotionMeter</td>
<td valign="top" align="center">85.11%</td>
</tr>
<tr>
<td valign="top" align="left"><xref ref-type="bibr" rid="B1">Acharya et al.</xref></td>
<td valign="top" align="center">2020</td>
<td valign="top" align="center">LSTM</td>
<td valign="top" align="center">87.22%</td>
</tr>
<tr>
<td valign="top" align="left">Our model</td>
<td valign="top" align="center">2021</td>
<td valign="top" align="center">3D-CNN</td>
<td valign="top" align="center">89.55%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For classification, cross-validation is not effective protection against overfitting or overhyping. It would be better to use techniques such as lockboxes, blind analyses, pre-registrations, or nested cross-validation to limit overhyping. We use the lockbox (<xref ref-type="bibr" rid="B6">Hosseini et al., 2020</xref>) approach to determine whether overhyping has occurred in the CNN model. The lockbox approach is a new technique that can be used to determine whether overhyping has occurred. The lockbox is accessed just one time to generate an unbiased estimate of the model&#x2019;s performance. In the DEAP and SEED-IV datasets, 90% of the data are set aside in the hyperparameter optimization set and the remaining 10% of the data are set aside in a lockbox. With the 10-fold cross-validation approach, the hyperparameters in the ERN model can be iteratively modified on the hyperparameter optimization set. When the average accuracy in the model is good enough, the model is tested on the lockbox data.</p>
<p>As shown in <xref ref-type="table" rid="T5">Table 5</xref>, the training result on the hyperparameter optimization set and the testing result on the lockbox set are 98.59 and 95.67% in the DEAP, 93.05 and 89.55% in the SEED-IV. According to the identification result, we can obtain the following conclusions. The theta band is in the state of sleep and a less responsive emotional state, so the recognition rate of emotion is lower than that of the other three waveforms. The excitement state of alpha, beta, gamma waveforms increased successively, so the recognition accuracy of emotion is higher.</p>
<table-wrap position="float" id="T5">
<label>TABLE 5</label>
<caption><p>Comparison of ERN model with different bands.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Modality</td>
<td valign="top" align="center">Train result (%)</td>
<td valign="top" align="center">Test result (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" colspan="3">EEG signals (DEAP)</td>
</tr>
<tr>
<td valign="top" align="left">Theta</td>
<td valign="top" align="center">87.81</td>
<td valign="top" align="center">84.40</td>
</tr>
<tr>
<td valign="top" align="left">Alpha</td>
<td valign="top" align="center">92.47</td>
<td valign="top" align="center">88.74</td>
</tr>
<tr>
<td valign="top" align="left">Beta</td>
<td valign="top" align="center">97.01</td>
<td valign="top" align="center">90.91</td>
</tr>
<tr>
<td valign="top" align="left">Gamma</td>
<td valign="top" align="center">94.39</td>
<td valign="top" align="center">88.31</td>
</tr>
<tr>
<td valign="top" align="left">EEG</td>
<td valign="top" align="center">98.59</td>
<td valign="top" align="center">95.67</td>
</tr>
<tr>
<td valign="top" align="left" colspan="3">EEG signals (SEED-IV)</td>
</tr>
<tr>
<td valign="top" align="left">Theta</td>
<td valign="top" align="center">83.08</td>
<td valign="top" align="center">82.09</td>
</tr>
<tr>
<td valign="top" align="left">Alpha</td>
<td valign="top" align="center">90.48</td>
<td valign="top" align="center">86.57</td>
</tr>
<tr>
<td valign="top" align="left">Beta</td>
<td valign="top" align="center">85.65</td>
<td valign="top" align="center">85.07</td>
</tr>
<tr>
<td valign="top" align="left">Gamma</td>
<td valign="top" align="center">92.45</td>
<td valign="top" align="center">88.06</td>
</tr>
<tr>
<td valign="top" align="left">EEG</td>
<td valign="top" align="center">93.05</td>
<td valign="top" align="center">89.55</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="S5.SS2">
<title><italic>F1</italic> and <italic>ROC</italic></title>
<p>We also calculated the <italic>F</italic><sub>1</sub> value and <italic>ROC</italic> value to analyze the performance of the ERN. In Formulas (2) and (3), <italic>F</italic><sub>1</sub> is the unweighted average of multiple categories of <italic>F</italic><sub>1<sub><italic>m</italic></sub></sub> (<italic>m</italic> = 0,&#x2026;, C, C = 3), where <italic>m</italic> means four classes of emotion: LVLA, LVHA, HVLA and HVHA. <italic>F</italic><sub>1<sub><italic>m</italic></sub></sub> is calculated from <italic>Precision</italic><sub><italic>m</italic></sub> and <italic>Recall</italic><sub><italic>m</italic></sub>, and <italic>Precision</italic><sub><italic>m</italic></sub> and <italic>Recall</italic><sub><italic>m</italic></sub> are calculated from <italic>FN</italic><sub><italic>m</italic></sub>, <italic>TP</italic><sub><italic>m</italic></sub>, <italic>TN</italic><sub><italic>m</italic></sub> and <italic>FP</italic><sub><italic>m</italic></sub>. In Formulas (4) and (5), <italic>FN</italic><sub><italic>m</italic></sub> and <italic>TP</italic><sub><italic>m</italic></sub>, represent the number of (incorrectly) recognized samples of a certain category, <italic>FP</italic><sub><italic>m</italic></sub> and <italic>TN</italic><sub><italic>m</italic></sub>, represent the number of (incorrectly) recognized samples of other categories except the <italic>m</italic>-th emotion category. According to Formula (6) and Formula (7), the true positive rate (<italic>TPR</italic>) and false positive rate (<italic>FPR</italic>) are calculated by <italic>FN</italic><sub><italic>m</italic></sub>, <italic>TP</italic><sub><italic>m</italic></sub>, <italic>TN</italic><sub><italic>m</italic></sub> and <italic>FP</italic><sub><italic>m</italic></sub>, and the <italic>ROC</italic> curves of the model are calculated to obtain <italic>AUC</italic><sub><italic>m</italic></sub> (the area under the <italic>ROC</italic> curve) of each category.</p>
<disp-formula id="S5.E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>C</mml:mi></mml:mfrac><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>C</mml:mi></mml:munderover><mml:msub><mml:mi>F</mml:mi><mml:msub><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:msub></mml:msub></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S5.E3"><label>(3)</label><mml:math id="M3"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:msub><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:msub></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>P</mml:mi></mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi></mml:mrow><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<disp-formula id="S5.E4"><label>(4)</label><mml:math id="M4"><mml:mrow><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo mathvariant="italic" separator="true">&#x2003;&#x2002;&#x2006;</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="S5.E5"><label>(5)</label><mml:math id="M5"><mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mpadded width="+8.3pt"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mpadded></mml:mrow></mml:math></disp-formula>
<disp-formula id="S5.E6"><label>(6)</label><mml:math id="M6"><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<disp-formula id="S5.E7"><label>(7)</label><mml:math id="M7"><mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>F</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>After several iterations, the values of the four classes are shown in <xref ref-type="table" rid="T6">Table 6</xref>. In the DEAP dataset, the LVLA, LVHA, HVLA, and HVHA values are 97.51, 98.62, 98.03, and 98.76, respectively. In the SEED-IV dataset, the LVLA, LVHA, HVLA, and HVHA values are 99.92, 98.97, 99.94, and 99.85, respectively. As shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, the solid red line is the average ROC curve of the four categories, and the average AUC is 98.23 for DEAP and 99.33 for SEED-IV.</p>
<table-wrap position="float" id="T6">
<label>TABLE 6</label>
<caption><p><italic>F</italic><sub>1</sub> values and AUC values of the four-class.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Dataset</td>
<td valign="top" align="center"><italic>m</italic></td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">Mean</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">DEAP</td>
<td valign="top" align="center"><italic>F</italic><sub>1<sub><italic>m</italic></sub></sub></td>
<td valign="top" align="center">94.82</td>
<td valign="top" align="center">95.06</td>
<td valign="top" align="center">94.52</td>
<td valign="top" align="center">96.742</td>
<td valign="top" align="center">95.29</td>
</tr>
<tr>
<td/>
<td valign="top" align="center"><italic>AUC</italic><sub><italic>m</italic></sub></td>
<td valign="top" align="center">97.51</td>
<td valign="top" align="center">98.62</td>
<td valign="top" align="center">98.03</td>
<td valign="top" align="center">98.76</td>
<td valign="top" align="center">98.23</td>
</tr>
<tr>
<td valign="top" align="left">SEED-IV</td>
<td valign="top" align="center"><italic>F</italic><sub>1><sub><italic>m</italic></sub></sub></td>
<td valign="top" align="center">95.652</td>
<td valign="top" align="center">96.97</td>
<td valign="top" align="center">92.683</td>
<td valign="top" align="center">85.714</td>
<td valign="top" align="center">92.75</td>
</tr>
<tr>
<td/>
<td valign="top" align="center"><italic>AUC</italic><sub><italic>m</italic></sub></td>
<td valign="top" align="center">99.92</td>
<td valign="top" align="center">98.97</td>
<td valign="top" align="center">99.94</td>
<td valign="top" align="center">99.85</td>
<td valign="top" align="center">99.67</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>Average ROC of the four-classes. (The ROC of the DEAP in the left, the ROC of the SEED-IV in the right).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g006.tif"/>
</fig>
<p>In the operation process, mass test data and the large threshold distance between the two samples result in the ROC curve not being smooth in <xref ref-type="fig" rid="F6">Figure 6</xref>. The higher values of the four categories indicate that the model constructed in this paper has better performance, among which the data of the fourth category (HVHA) have higher identifiability. Each column of the confusion matrix represents the predictive category, and the total number of each column indicates the number of data predicted for this category. Each line represents the real category of the data, and the total number of each line of data represents the number of data instances of the category in <xref ref-type="fig" rid="F7">Figure 7</xref>. The matrix verifies that our model is stronger than others in predicting complex labels.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>The confusion matrixes of datasets. (The confusion matrix of the DEAP in the left, the confusion matrix of the SEED-IV in the right).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872311-g007.tif"/>
</fig>
<p>From our experimental results shown in <xref ref-type="table" rid="T4">Tables 4</xref>&#x2013;<xref ref-type="table" rid="T6">6</xref>, <xref ref-type="fig" rid="F6">Figures 6</xref>, <xref ref-type="fig" rid="F7">7</xref>, it can be seen that our results are superior to those of previous studies reported in the literature over the same EEG datasets from DEAP and SEED-IV. There are 3 possible reasons. (1) We conduct a simple and efficient preprocessing method, including data baseline signal processing, EEG electrode topological mapping, and 1 s time window length selection. (2) A 3D convolutional structure is a necessary technique to study emotional recognition based on EEG signals, as this structure can identify space information of the electrodes to quickly extract spatiotemporal features. (3) The multiscale convolution kernel not only reduces the computation of the model but also improves the identification ability of the model because multiple smaller scale kernels have the ability to increase non-linear expression more than a larger kernel. These all enhance the operating efficiency and improve the recognition performance of the ERN model.</p>
</sec>
<sec id="S5.SS3">
<title>Fisher</title>
<p>The parietal lobe of the human brain and non-human primate brain have been associated with attention based on the evidence of clinical and physiological (<xref ref-type="bibr" rid="B8">Joseph, 1990</xref>). Studies in literature show that the lateral Intraparietal area (LIP) plays an independent role in target selection and visual attention generation (<xref ref-type="bibr" rid="B23">Simon et al., 2002</xref>). These findings can be validated by the distribution of Fisher&#x2019;s selected EEG electrodes. We score the 32 electrodes of the applied dataset using the <italic>Fisher</italic> (<xref ref-type="bibr" rid="B12">Li R. et al., 2021</xref>) scorer and rank the 32 electrodes from highest to lowest to obtain the top 8, top 16 and top 24. In Formula (8), &#x03BC;<sub><italic>m</italic>,<italic>v</italic></sub> is the average of each electrode of the <italic>m</italic>-th class, &#x03BC;<sub><italic>v</italic></sub> is the average of each electrode, &#x03C3;<sub><italic>m</italic>,<italic>v</italic></sub> is the variance of each electrode of the <italic>m</italic>-th class and <italic>n<sub>m</sub></italic> is the number of samples of the <italic>m</italic>-th class (<italic>m</italic> = 0,&#x2026;,M, M = 3).</p>
<disp-formula id="S5.E8"><label>(8)</label><mml:math id="M8"><mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03BC;</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03BC;</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>M</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:msubsup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>We can sort the impact sequence of different electrodes based on emotions by the recognition results. According to the output results of the ERN model, we can study the recognition performance of the model for many EEG electrodes. <xref ref-type="table" rid="T7">Table 7</xref> shows that the emotion recognition results of different quantities of electrodes in the sort. Experimental results showed that there is a 4.7% difference between the data of the first 8 electrodes and the whole dataset in DEAP and a 5.97% difference between the data of the first 8 electrodes and the whole dataset in SEED-IV. This indicates that the model can quickly obtain high recognition performance with less EEG electrodes data. We can use the ERN model for portable emotional identification (<xref ref-type="bibr" rid="B4">Cai et al., 2018</xref>) based on EEG and the model meets simple and fast needs. According to the identification results of the first 8, 16 and 24 electrodes in the table, most of the effective EEG electrodes are distributed in the frontal and parietal lobes (such as, O2, PO4, AF3, F3, F7, FC5, FC1, and so on). Therefore, the frontal and parietal lobes have a large effect on emotional identification.</p>
<table-wrap position="float" id="T7">
<label>TABLE 7</label>
<caption><p>Evaluation of recognition results based on <italic>Fisher.</italic></p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Dataset</td>
<td valign="top" align="center">The number of electrodes</td>
<td valign="top" align="left">The name of electrodes</td>
<td valign="top" align="center">Accuracy</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">DEAP</td>
<td valign="top" align="center">8</td>
<td valign="top" align="left">O2\PO4\AF3\F3\F7\FC5\FC1\C3</td>
<td valign="top" align="center">90.31%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">16</td>
<td valign="top" align="left">O2\PO4\AF3\F3\F7\FC5\FC1\C3\T7\CP5\CP1\P3\P7\PO3\O1\Oz</td>
<td valign="top" align="center">92.07%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">24</td>
<td valign="top" align="left">O2\PO4\AF3\F3\F7\FC5\FC1\C3\T7\CP5\CP1\P3\P7\PO3\O1\Oz\<break/> Pz\Fp2\AF4\Fz\F4\F8\FC6\FC2</td>
<td valign="top" align="center">93.03%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">32</td>
<td valign="top" align="left">O2\PO4\AF3\F3\F7\FC5\FC1\C3\T7\CP5\CP1\P3\P7\PO3\O1\Oz\Pz\Fp2\AF4\Fz\F4\F8\FC6\FC2\Cz\C4\T8\CP6\CP2\P4\P8\Fp1</td>
<td valign="top" align="center">95.67%</td>
</tr>
<tr>
<td valign="top" align="left">SEED-IV</td>
<td valign="top" align="center">8</td>
<td valign="top" align="left">Oz\O2\FP2\AF3\AF4\F7\F3\Fz</td>
<td valign="top" align="center">83.58%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">16</td>
<td valign="top" align="left">Oz\O2\FP2\AF3\AF4\F7\F3\Fz\F4\F8\FC5\FC1\FC2\FC6\T7\C3</td>
<td valign="top" align="center">85.07%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">24</td>
<td valign="top" align="left">Oz\O2\FP2\AF3\AF4\F7\F3\Fz\F4\F8\FC5\FC1\FC2\FC6\T7\C3\Cz\C4\T8\CP5\CP1\CP2\CP6\P7</td>
<td valign="top" align="center">86.57%</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">32</td>
<td valign="top" align="left">Fp1\FP2\AF3\AF4\F7\F3\Fz\F4\F8\FC5\FC1\FC2\FC6\T7\C3\Cz\C4\T8\CP5\CP1\CP2\CP6\P7\P3\PZ\P4\P8\PO3\PO4\O1\Oz\O2</td>
<td valign="top" align="center">89.55%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="S5.SS4">
<title>Ablation experiments</title>
<p>The generic dimension and volume type of the kernel have 3 &#x00D7; 3, 5 &#x00D7; 5, and 7 &#x00D7; 7, where a plurality of 3 &#x00D7; 3 stacked approximately a 5 &#x00D7; 5 or 7 &#x00D7; 7. Because the activation function is set after the convolutional layer, <xref ref-type="bibr" rid="B10">Krizhevsky et al. (2012)</xref> believes that the recognition capability of the model can be controlled by the volume of kernel. Multiscale small kernel subscriptions have diversely increased the network capacity so that the decision function is more distinguished for different categories.</p>
<p>We assume that the size of the 3D convolutional kernel is M &#x00D7; N &#x00D7; K. When using a 3D kernel, it can be divided into three steps: First, complete the convolution of the M &#x00D7; 1 &#x00D7; 1 content; second, complete the convolution of the 1 &#x00D7; N &#x00D7; 1 content; finally, complete the convolution of the 1 &#x00D7; 1 &#x00D7; K content. The total convolution process can increase the non-linear expression of the model because the local convolution of each small step is completed and pass through the non-linear function.</p>
<p>We can clearly see that the entire process has three non-linear transformations. Therefore, the non-linear characteristics of the results eventually increase, making the decision function more decisive and helping the model increase the accuracy of emotion recognition. The size of conv3 &#x00D7; 3 (the size of convolutional kernel is 3 &#x00D7; 3) compared to conv5 &#x00D7; 5 and conv7 &#x00D7; 7 significantly reduces the number of parameters. <xref ref-type="bibr" rid="B24">Simonyan and Zisserman (2014)</xref> replaces a conv7 &#x00D7; 7 with three conv3 &#x00D7; 3, which is considered to further decompose the characteristics mentioned by the 7 &#x00D7; 7 larger volume kernels. The Regularization of the multiscale small kernel can improve the model performance.</p>
<p>In addition to the small conv3 &#x00D7; 3, there is the small conv2 &#x00D7; 2. However, <xref ref-type="bibr" rid="B42">Zeiler and Fergus (2014a)</xref> studies conv2 &#x00D7; 2 and is unable to find the central point of the convolution, which causes the characteristics of the padding process to constantly offset. As the number of layers deepens, conv2 &#x00D7; 2 makes the distance of feature offset increasingly obvious. Thus, this paper expects to apply the multiscale convolutional kernel to increase the feature amount of the model calculation and improve the model recognition.</p>
<p>According to the above research results, this paper takes the types of convolutional kernels of conv3 &#x00D7; 3. The same dataset is carried out by different types of kernels and each classification result is compared. Conv3 &#x00D7; 3&#x002A; represents the multiscale kernel, and conv3 &#x00D7; 3 represents the same-scale kernel. <xref ref-type="table" rid="T8">Table 8</xref> shows that the multiscale small kernel subscriptions diversely increase the network capacity so that the decision function is more distinguished for different categories.</p>
<table-wrap position="float" id="T8">
<label>TABLE 8</label>
<caption><p>Comparison of different kernels.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Kernel Size</td>
<td valign="top" align="center">DEAP (%)</td>
<td valign="top" align="center">SEED-IV (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Conv3 &#x00D7; 3<xref ref-type="table-fn" rid="t8fns1">&#x002A;</xref></td>
<td valign="top" align="center">95.67</td>
<td valign="top" align="center">89.88</td>
</tr>
<tr>
<td valign="top" align="left">Conv3 &#x00D7; 3</td>
<td valign="top" align="center">92.88</td>
<td valign="top" align="center">85.13</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="t8fns1"><p>&#x002A;represents the multiscale kernel.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="S6" sec-type="conclusion">
<title>Conclusion</title>
<p>In this paper, we have presented the ERN model which uses the multiscale 3D-CNN to recognize emotions based on EEG. We obtain the optimal parameters through random training data and design experiments to compare the promotion classification performance at different time windows. Then, based on the 1 s time window dataset with better classification performance, an effective multiscale convolutional kernel 3D-CNN model based on EEG signals is implemented to simultaneously extract spatial and temporal features, and achieves higher accuracy of emotion recognition.</p>
<p>In the comparative analysis using the EEG signals in the DEAP and SEED-IV datasets, we have demonstrated the superior accuracy, <italic>F</italic><sub>1</sub>, and AUC values of emotion recognition for the ERN model based on multiscale 3D-CNN. From the experimental results, we show that this model can achieve higher performance, which helps to efficiently recognize the emotional state of the subjects so that BCI technology can quickly and accurately convert the neural electrical signal into commands that can be identified by the computer, greatly improving human-machine interaction.</p>
<p>The limitations of the model include the exploratory interpretability of the convolutional model. The calculation process in the convolution network is similar to a black box, and it is especially difficult to understand how the method works on feature learning. If the feature representation learned by the hidden layers can be visualized, it will be more conducive to the optimization of the convolutional network.</p>
<p>While we have achieved superior accuracy when compared to alternative methods as discussed in this paper using EEG data, we consider that there are further potential improvements. Our projected future directions for research include addressing subject-independent emotion recognition (the model can be trained using data acquired from a limited number of participants and can be applied to a subject who has never experienced the system prior to the experiment.) used our model to conduct the fusion of multimodal signals such as EEG and EMG studies and the investigation of other methods to improve our model with respect to the recognition accuracy. In addition, we can also use the model to challenge other tasks, such as emotion recognition based on multimodal physiological signal fusion, which could facilitate the performance of real-time emotion recognition and enhance emotional experience in the field of BCI.</p>
</sec>
<sec id="S7" sec-type="data-availability">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="http://www.eecs.qmul.ac.uk/mmv/datasets/deap/">http://www.eecs.qmul.ac.uk/mmv/datasets/deap/</ext-link> and <ext-link ext-link-type="uri" xlink:href="https://bcmi.sjtu.edu.cn/~seed/index.html">https://bcmi.sjtu.edu.cn/~seed/index.html</ext-link>.</p>
</sec>
<sec id="S8">
<title>Author contributions</title>
<p>YS and ZZ designed this project, carried out most of the experiments and data analysis, and revised the manuscript. All authors analyzed the results and presented the discussion and conclusion, contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="S9" sec-type="funding-information">
<title>Funding</title>
<p>This research was supported in part by the National Natural Science Foundation of China (NSFC) (Grant nos. 61862058 and 61962034), the Gansu Provincial Science and Technology Department (Grant no. 20JR10RA076), and the Cultivation plan of major Scientific Research Projects of Northwest Normal University (NWNU-LKZD2021-06).</p>
</sec>
<sec id="S10" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="S11" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Acharya</surname> <given-names>D.</given-names></name> <name><surname>Goel</surname> <given-names>S.</given-names></name> <name><surname>Bhardwaj</surname> <given-names>H.</given-names></name> <name><surname>Sakalle</surname> <given-names>A.</given-names></name> <name><surname>Bhardwaj</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>A long short term memory deep learning network for the classification of negative emotions using EEG Signals</article-title>,&#x201D; in <source><italic>Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN)</italic></source>, <publisher-loc>Glasgow, UK</publisher-loc>. <pub-id pub-id-type="doi">10.1109/IJCNN48605.2020.9207280</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>An</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>S.</given-names></name> <name><surname>Duan</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Xie</surname> <given-names>C.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Electroencephalogram emotion recognition based on 3D feature fusion and convolutional autoencoder.</article-title> <source><italic>Front. Comput. Neurosci.</italic></source> <volume>15</volume>: <issue>743426</issue>.</citation></ref>
<ref id="B3"><citation citation-type="journal"><collab>BCMI</collab> (<year>1994</year>). <source><italic>Center for Brain-like Computing and Machine Intelligence(BCMI) laboratory, China.</italic></source> <publisher-loc>Shanghai</publisher-loc>: <publisher-name>Shanghai Jiao Tong University</publisher-name>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Hu</surname> <given-names>B.</given-names></name></person-group> (<year>2018</year>). <article-title>A case-based reasoning model for depression based on three-electrode EEG data.</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>11</volume> <fpage>383</fpage>&#x2013;<lpage>392</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2018.2801289</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Candra</surname> <given-names>H.</given-names></name> <name><surname>Yuwono</surname> <given-names>M.</given-names></name> <name><surname>Chai</surname> <given-names>R.</given-names></name> <name><surname>Handojoseno</surname> <given-names>A.</given-names></name> <name><surname>Su</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Investigation of window size in classification of EEG-emotion signal with wavelet entropy and support vector machine.</article-title> <source><italic>Annu. Int. Conf. IEEE Eng. Med. Biol. Soc.</italic></source> <volume>2015</volume> <fpage>7250</fpage>&#x2013;<lpage>7253</lpage>. <pub-id pub-id-type="doi">10.1109/EMBC.2015.7320065</pub-id> <pub-id pub-id-type="pmid">26737965</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hosseini</surname> <given-names>M.</given-names></name> <name><surname>Powell</surname> <given-names>M.</given-names></name> <name><surname>Collins</surname> <given-names>J.</given-names></name> <name><surname>Mahan</surname> <given-names>H.</given-names></name> <name><surname>Michael</surname> <given-names>P.</given-names></name> <name><surname>John</surname> <given-names>C.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>I tried a bunch of things: The dangers of unexpectedoverfitting in classification of brain data.</article-title> <source><italic>Neurosci. Biobehav. Rev.</italic></source> <volume>119</volume> <fpage>456</fpage>&#x2013;<lpage>467</lpage>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2020.09.036</pub-id> <pub-id pub-id-type="pmid">33035522</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jenke</surname> <given-names>R.</given-names></name> <name><surname>Peer</surname> <given-names>A.</given-names></name> <name><surname>Buss</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Feature extraction and selection for emotion recognition from EEG.</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>5</volume> <fpage>327</fpage>&#x2013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2019.2901673</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joseph</surname> <given-names>R.</given-names></name></person-group> (<year>1990</year>). <source><italic>The parietal lobes.</italic></source> <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koelstra</surname> <given-names>S.</given-names></name> <name><surname>M&#x00FC;hl</surname> <given-names>C.</given-names></name> <name><surname>Soleymani</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>J.-S.</given-names></name> <name><surname>Yazdani</surname> <given-names>A.</given-names></name> <name><surname>Ebrahimi</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>DEAP: a database for emotion analysis using physiological signals.</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>3</volume> <fpage>18</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1109/T-AFFC.2011.15</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <source><italic>ImageNet classification with deep convolutional neural networks.</italic></source> <publisher-loc>Red Hook; NY</publisher-loc>: <publisher-name>NIPS. Curran Associates Inc</publisher-name>.</citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kwon</surname> <given-names>Y. H.</given-names></name> <name><surname>Shin</surname> <given-names>S. B.</given-names></name> <name><surname>Kim</surname> <given-names>S. D.</given-names></name></person-group> (<year>2018</year>). <article-title>Electroencephalography based fusion two-dimensional (2D)-convolution neural networks (CNN) model for emotion recognition system.</article-title> <source><italic>Sensors</italic></source> <volume>18</volume>:<issue>1383</issue>. <pub-id pub-id-type="doi">10.3390/s18051383</pub-id> <pub-id pub-id-type="pmid">29710869</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>R.</given-names></name> <name><surname>Jared</surname> <given-names>J. S.</given-names></name> <name><surname>Ahmed</surname> <given-names>H.</given-names></name> <name><surname>IIyevsky</surname> <given-names>T. V.</given-names></name> <name><surname>Wilbur</surname> <given-names>R. B.</given-names></name> <name><surname>Bharadwaj</surname> <given-names>H. M.</given-names></name></person-group> (<year>2021</year>). <article-title>The perils and pitfalls of block design for EEG classification experiments.</article-title> <source><italic>IEEE Trans. Pattern Anal. Mach. Intell.</italic></source> <volume>43</volume>:<fpage>316</fpage>&#x2013;<lpage>333</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2020.2973153</pub-id> <pub-id pub-id-type="pmid">33211652</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Lyu</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Gong</surname> <given-names>A.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Identification of emotion using electroencephalogram by tunable Q-factor wavelet transform and binary gray wolf optimization.</article-title> <source><italic>Front. Comput. Neurosci.</italic></source> <volume>15</volume>:<issue>732763</issue>. <pub-id pub-id-type="doi">10.3389/fncom.2021.732763</pub-id> <pub-id pub-id-type="pmid">34566614</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Lei</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Cui</surname> <given-names>W.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>T. W.</given-names></name><etal/></person-group> (<year>2018</year>). <article-title>Boosted convolutional neural networks for motor imagery EEG decoding with multiwavelet-based time-frequency conditional granger causality analysis.</article-title> <source><italic>arXiv</italic></source> [<comment>Preprint</comment>]. <pub-id pub-id-type="doi">10.48550/arXiv.1810.10353</pub-id> <pub-id pub-id-type="pmid">35895330</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>W. Q.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Sun</surname> <given-names>S. Q.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep convolutional neural network for emotion recognition using EEG and peripheral physiological signal.</article-title> <source><italic>Lecture Notes Comp. Sci.</italic></source> <volume>10667</volume> <fpage>385</fpage>&#x2013;<lpage>394</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-71589-6_33</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>N. J.</given-names></name> <name><surname>Fang</surname> <given-names>Y. C.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Hou</surname> <given-names>L. M.</given-names></name> <name><surname>Yang</surname> <given-names>F. L.</given-names></name> <name><surname>Guo</surname> <given-names>Y. K.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Multiple feature fusion for automatic emotion recognition using EEG signals</article-title>,&#x201D; in <source><italic>Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</italic></source>, <publisher-loc>Piscataway, NJ</publisher-loc>. <pub-id pub-id-type="doi">10.1109/ICASSP.2018.8462518</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Ding</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <name><surname>Song</surname> <given-names>R.</given-names></name> <name><surname>Wan</surname> <given-names>F.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Multi-channel EEG-based emotion recognition via a multi-level features guided capsule network.</article-title> <source><italic>Comput. Biol. Med.</italic></source> <volume>123</volume>:<issue>103927</issue>. <pub-id pub-id-type="doi">10.1016/J.COMPBIOMED.2020.103927</pub-id> <pub-id pub-id-type="pmid">32768036</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mei</surname> <given-names>M.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>EEG-Based Emotion Classification Using Convolutional Neural Networks</article-title>,&#x201D; in <source><italic>Proceedings of the International Conference on Security, Pattern Analysis, and Cybernetics (SPAC)</italic></source>, <publisher-loc>Chengdu</publisher-loc>.</citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiu</surname> <given-names>J. L.</given-names></name> <name><surname>Li</surname> <given-names>X. Y.</given-names></name> <name><surname>Hu</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Correlated Attention Networks for Multimodal Emotion Recognition</article-title>,&#x201D; in <source><italic>Proceedings of the 2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)</italic></source>, <publisher-loc>Piscataway, NJ</publisher-loc>.</citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salama</surname> <given-names>E.</given-names></name> <name><surname>El-Khoribi</surname> <given-names>R.</given-names></name> <name><surname>Shoman</surname> <given-names>M.</given-names></name> <name><surname>Wahby</surname> <given-names>S. M.</given-names></name></person-group> (<year>2018</year>). <article-title>EEG-based emotion recognition using 3D convolutional neural networks.</article-title> <source><italic>Int. J. Adv. Comput. Sci. Appl.</italic></source> <volume>9</volume> <fpage>329</fpage>&#x2013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.14569/IJACSA.2018.090843</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scherer</surname> <given-names>D.</given-names></name> <name><surname>Andreas</surname> <given-names>M.</given-names></name> <name><surname>Behnke</surname> <given-names>S.</given-names></name></person-group> (<year>2010</year>). &#x201C;<article-title>Evaluation of pooling operations in convolutional architectures for object recognition</article-title>,&#x201D; in <source><italic>Artificial Neural Networks &#x2013; ICANN 2010. ICANN 2010. Lecture Notes in Computer Science</italic></source>, <role>eds</role> <person-group person-group-type="editor"><name><surname>Diamantaras</surname> <given-names>K.</given-names></name> <name><surname>Duch</surname> <given-names>W.</given-names></name> <name><surname>Iliadis</surname> <given-names>L. S.</given-names></name></person-group>, (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>92</fpage>&#x2013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-15825-4_10</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharbrough</surname> <given-names>F.</given-names></name> <name><surname>Chatrian</surname> <given-names>G. E.</given-names></name> <name><surname>L&#x00FC;ders</surname> <given-names>H.</given-names></name> <name><surname>Nuwer</surname> <given-names>M.</given-names></name> <name><surname>Picton</surname> <given-names>T. W.</given-names></name><etal/></person-group> (<year>1991</year>). <article-title>American electroencephalographic society guidelines for standard electrode position nomenclature.</article-title> <source><italic>J. Clin. Neurophysiol.</italic></source> <volume>8</volume> <fpage>200</fpage>&#x2013;<lpage>202</lpage>.</citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simon</surname> <given-names>O.</given-names></name> <name><surname>Mangin</surname> <given-names>J. F.</given-names></name> <name><surname>Cohen</surname> <given-names>L.</given-names></name> <name><surname>Le Bihan</surname> <given-names>D.</given-names></name> <name><surname>Dehaene</surname> <given-names>S.</given-names></name></person-group> (<year>2002</year>). <article-title>Topographical layout of hand, eye, calculation, and language-related areas in the human parietal lobe.</article-title> <source><italic>Neuron</italic></source> <volume>33</volume> <fpage>475</fpage>&#x2013;<lpage>487</lpage>. <pub-id pub-id-type="doi">10.1016/S0896-6273(02)00575-5</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Very deep convolutional networks for large-scale image recognition.</article-title> <source><italic>arXiv</italic></source> [<comment>Preprint</comment>]. <pub-id pub-id-type="doi">10.48550/arXiv.1409.1556</pub-id> <pub-id pub-id-type="pmid">35895330</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>T.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Song</surname> <given-names>P.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>EEG emotion recognition using dynamical graph convolutional neural networks.</article-title> <source><italic>IEEE Trans. Affect. Comput.</italic></source> <volume>11</volume> <fpage>532</fpage>&#x2013;<lpage>541</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2018.2817622</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chaudhary</surname> <given-names>A.</given-names></name> <name><surname>Chouhan</surname> <given-names>K. S.</given-names></name> <name><surname>Gajrani</surname> <given-names>J.</given-names></name> <name><surname>Sharma</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Deep learning with PyTorch</article-title>,&#x201D; in <source><italic>Machine learning and deep learning in real-time applications</italic></source> (<publisher-loc>Hershey, PA</publisher-loc>: <publisher-name>IGI Global</publisher-name>), <fpage>61</fpage>&#x2013;<lpage>95</lpage>.</citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Ayaz</surname> <given-names>H.</given-names></name> <name><surname>Akansu</surname> <given-names>A. N.</given-names></name></person-group> (<year>2020</year>). <article-title>Multimodal affective state assessment using fnirs + eeg and spontaneous facial expression.</article-title> <source><italic>Brain Sci.</italic></source> <volume>10</volume>:<issue>85</issue>. <pub-id pub-id-type="doi">10.3390/brainsci10020085</pub-id> <pub-id pub-id-type="pmid">32041316</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Vanhoucke</surname> <given-names>V.</given-names></name> <name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Shlens</surname> <given-names>J.</given-names></name> <name><surname>Wojna</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Rethinking the Inception Architecture for Computer Vision</article-title>,&#x201D; in <source><italic>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Chengdu</publisher-loc>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.308</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tao</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Song</surname> <given-names>R.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Wan</surname> <given-names>F.</given-names></name><etal/></person-group> (<year>2020</year>). &#x201C;<article-title>EEG-based Emotion Recognition via Channelwise Attention and Self Attention</article-title>,&#x201D; in <source><italic>Proceedings of the IEEE Transactions on Affective Computing</italic></source>, <publisher-loc>Chengdu</publisher-loc>. <pub-id pub-id-type="doi">10.1109/TAFFC.2020.3025777</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tran</surname> <given-names>D.</given-names></name> <name><surname>Bourdev</surname> <given-names>L.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name> <name><surname>Torresani</surname> <given-names>L.</given-names></name> <name><surname>Paluri</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Learning Spatiotemporal Features with 3D Convolutional Networks</article-title>,&#x201D; in <source><italic>Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV)</italic></source>, <publisher-loc>Chengdu</publisher-loc>. <pub-id pub-id-type="doi">10.1109/ICCV.2015.510</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>Z.</given-names></name> <name><surname>Mccane</surname> <given-names>B.</given-names></name> <name><surname>Neo</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>EmotioNet: A 3-D Convolutional Neural Network for EEG-based Emotion Recognition</article-title>,&#x201D; in <source><italic>Proceedings of the International Joint Conference on Neural Networks</italic></source>, <publisher-loc>Chengdu</publisher-loc>. <pub-id pub-id-type="doi">10.1109/IJCNN.2018.8489715</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>Y.</given-names></name> <name><surname>Liang</surname> <given-names>R.</given-names></name> <name><surname>Liang</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>C.</given-names></name> <name><surname>Schuller</surname> <given-names>B.</given-names></name></person-group> (<year>2019</year>). <article-title>Speech emotion classification using attention-based LSTM.</article-title> <source><italic>IEEE ACM Trans. Audio Speech Lang</italic>. <italic>Proc.</italic></source> <volume>27</volume> <fpage>1675</fpage>&#x2013;<lpage>1685</lpage>. <pub-id pub-id-type="doi">10.1109/TASLP.2019.2925934</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xing</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Xu</surname> <given-names>T.</given-names></name> <name><surname>Shu</surname> <given-names>L.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>SAE+LSTM: a new framework for emotion recognition from multi-channel EEG.</article-title> <source><italic>Front. Neurorobot.</italic></source> <volume>13</volume>: 37. <pub-id pub-id-type="doi">10.3389/fnbot.2019.00037</pub-id> <pub-id pub-id-type="pmid">31244638</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>Q.</given-names></name> <name><surname>Ming</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Emotion Recognition from Multi-ChannelEEG through Parallel Convolutional Recurrent Neural Network</article-title>,&#x201D; in <source><italic>Proceedings of the International Joint Conference on Neural Networks</italic></source>, <publisher-loc>Piscataway, NJ.</publisher-loc> <pub-id pub-id-type="doi">10.1109/IJCNN.2018.8489331</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>Q.</given-names></name></person-group> (<year>2014</year>). <source><italic>Multi-Sensory Emotion Recognition with Speech and Facial Expression.</italic></source> <comment>Ph.D. thesis</comment>. <publisher-loc>Mississippi</publisher-loc>: <publisher-name>University of Southern Mississippi</publisher-name>.</citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yea-Hoon</surname> <given-names>K.</given-names></name> <name><surname>Sae-Byuk</surname> <given-names>S.</given-names></name> <name><surname>Shin-Dug</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>Electroencephalography based fusion two-dimensional (2D)-convolution neural networks (CNN) model for emotion recognition system.</article-title> <source><italic>Sensors</italic></source> <volume>18</volume>:<issue>1383</issue>. <pub-id pub-id-type="doi">10.3390/s18051383</pub-id> <pub-id pub-id-type="pmid">29710869</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yin</surname> <given-names>Y. Q.</given-names></name> <name><surname>Zheng</surname> <given-names>X. W.</given-names></name> <name><surname>Hu</surname> <given-names>B.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Cui</surname> <given-names>X. C.</given-names></name></person-group> (<year>2021</year>). <article-title>EEG emotion recognition using fusion model of graph convolutional neural networks and LSTM.</article-title> <source><italic>Appl. Soft Comput.</italic></source> <volume>100</volume>:<issue>106954</issue>. <pub-id pub-id-type="doi">10.1016/j.asoc.2020.106954</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoon</surname> <given-names>H. J.</given-names></name> <name><surname>Chuang</surname> <given-names>S. Y.</given-names></name></person-group> (<year>2013</year>). <article-title>EEG-based emotion estimation using Bayesian weighted-log-posterior function and perceptron convergence algorithm.</article-title> <source><italic>Comp. Biol. Med.</italic></source> <volume>43</volume> <fpage>2230</fpage>&#x2013;<lpage>2237</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2013.10.017</pub-id> <pub-id pub-id-type="pmid">24290940</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Yoshimoto</surname> <given-names>J.</given-names></name> <name><surname>Shigeru</surname> <given-names>T.</given-names></name> <name><surname>Masahiro</surname> <given-names>T.</given-names></name> <name><surname>Shinpei</surname> <given-names>Y.</given-names></name> <name><surname>Yasumasa</surname> <given-names>O.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Nested 10-fold cross-validation.</article-title> <source><italic>PLoS One</italic></source> <volume>10</volume>:<issue>e0123524</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0123524.g002</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zangeneh</surname> <given-names>S. M.</given-names></name> <name><surname>Maghooli</surname> <given-names>K.</given-names></name> <name><surname>Setarehdan</surname> <given-names>S. K.</given-names></name> <name><surname>Motie</surname> <given-names>N. A.</given-names></name></person-group> (<year>2019</year>). <article-title>A novel eeg-based approach to classify emotions through phase space dynamics.</article-title> <source><italic>Signal Image Video Process.</italic></source> <volume>13</volume> <fpage>1149</fpage>&#x2013;<lpage>1156</lpage>. <pub-id pub-id-type="doi">10.1007/s11760-019-01455-y</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeiler</surname> <given-names>M. D.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name></person-group> (<year>2014b</year>). &#x201C;<article-title>Visualizing and Understanding Convolutional Networks</article-title>,&#x201D; in <source><italic>Computer Vision &#x2013; ECCV 2014. Lecture Notes in Computer Science</italic></source>, <role>eds</role> <person-group person-group-type="editor"><name><surname>Fleet</surname> <given-names>D.</given-names></name> <name><surname>Pajdla</surname> <given-names>T.</given-names></name> <name><surname>Schiele</surname> <given-names>B.</given-names></name> <name><surname>Tuytelaars</surname> <given-names>T.</given-names></name></person-group> (<publisher-loc>Zurich</publisher-loc>: <publisher-name>ECCV</publisher-name>). <pub-id pub-id-type="doi">10.1007/978-3-319-10590-1_53</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeiler</surname> <given-names>M. D.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name></person-group> (<year>2014a</year>). &#x201C;<article-title>Architecture of Convolutional Neural Networks (CNNs) demystified</article-title>,&#x201D; in <source><italic>Proceedings of the 12th European conference on computer vision</italic></source>, <publisher-loc>Chapel Hill, NC</publisher-loc>.</citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Bont</surname> <given-names>C. D.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Emotional Engagement for Human-Computer Interaction in Exhibition Design</article-title>,&#x201D; in <source><italic>Human-Computer Interaction: Design and Evaluation(HCI). Lecture Notes in Computer Science</italic></source>, <role>eds</role> <person-group person-group-type="editor"><name><surname>Kurosu</surname> <given-names>M.</given-names></name></person-group> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <pub-id pub-id-type="doi">10.1007/978-3-319-20901-2_51</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>D.</given-names></name> <name><surname>Cao</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>A 3D Convolutional Neural Network for Emotion Recognition based on EEG Signals</article-title>,&#x201D; in <source><italic>Proceedings of the International Joint Conference on Neural Networks (IJCNN)</italic></source>, <publisher-loc>Piscataway, NJ.</publisher-loc> <pub-id pub-id-type="doi">10.1109/IJCNN48605.2020.9207420</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>W. L.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Lu</surname> <given-names>Y.</given-names></name> <name><surname>Lu</surname> <given-names>B. L.</given-names></name> <name><surname>Cichocki</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>EmotionMeter: A multimodal framework for recognizing human emotions.</article-title> <source><italic>IEEE Trans. Cybernet.</italic></source> <volume>49</volume> <fpage>1110</fpage>&#x2013;<lpage>1122</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2018.2797176</pub-id> <pub-id pub-id-type="pmid">29994384</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>W. L.</given-names></name> <name><surname>Lu</surname> <given-names>B. L.</given-names></name></person-group> (<year>2015</year>). <article-title>Investigating critical frequency bands and channels for EEG-Based emotion recognition with deep neural networks.</article-title> <source><italic>IEEE Trans. Autonom. Ment. Dev.</italic></source> <volume>7</volume> <fpage>162</fpage>&#x2013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1109/TAMD.2015.2431497</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>Q.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Cai</surname> <given-names>D.</given-names></name> <name><surname>Xiao</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Electroencephalogram access for emotion recognition based on a deep hybrid network.</article-title> <source><italic>Front. Hum. Neurosci.</italic></source> <volume>14</volume>:<issue>589001</issue>. <pub-id pub-id-type="doi">10.3389/fnhum.2020.589001</pub-id> <pub-id pub-id-type="pmid">33390918</pub-id></citation></ref>
</ref-list>
</back>
</article>
