<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neuroinform.</journal-id>
<journal-title>Frontiers in Neuroinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neuroinform.</abbrev-journal-title>
<issn pub-type="epub">1662-5196</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fninf.2022.844667</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>On the Minimal Amount of EEG Data Required for Learning Distinctive Human Features for Task-Dependent Biometric Applications</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>G&#x000F3;mez-Tapia</surname> <given-names>Carlos</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1607875/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bozic</surname> <given-names>Bojan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1094742/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Longo</surname> <given-names>Luca</given-names></name>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/589949/overview"/>
</contrib>
</contrib-group>
<aff><institution>Artificial Intelligence and Cognitive Load Research Lab, Applied Intelligence Research Centre, School of Computer Science, Technological University Dublin</institution>, <addr-line>Dublin</addr-line>, <country>Ireland</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Giuseppe Placidi, University of L&#x00027;Aquila, Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Maciej Szymkowski, Bialystok University of Technology, Poland; Vinay Kumar, Thapar Institute of Engineering &#x00026; Technology, India</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Carlos G&#x000F3;mez-Tapia <email>carlos.g.tapia&#x00040;myTUDublin.ie</email></corresp>
<corresp id="c002">Luca Longo <email>luca.longo&#x00040;tudublin.ie</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>844667</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 G&#x000F3;mez-Tapia, Bozic and Longo.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>G&#x000F3;mez-Tapia, Bozic and Longo</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Biometrics is the process of measuring and analyzing human characteristics to verify a given person&#x00027;s identity. Most real-world applications rely on unique human traits such as fingerprints or iris. However, among these unique human characteristics for biometrics, the use of Electroencephalogram (EEG) stands out given its high inter-subject variability. Recent advances in Deep Learning and a deeper understanding of EEG processing methods have led to the development of models that accurately discriminate unique individuals. However, it is still uncertain how much EEG data is required to train such models. This work aims at determining the minimal amount of training data required to develop a robust EEG-based biometric model (&#x0002B;95% and &#x0002B;99% testing accuracies) from a subject for a task-dependent task. This goal is achieved by performing and analyzing 11,780 combinations of training sizes, by employing various neural network-based learning techniques of increasing complexity, and feature extraction methods on the affective EEG-based DEAP dataset. Findings suggest that if Power Spectral Density or Wavelet Energy features are extracted from the artifact-free EEG signal, 1 and 3 s of data per subject is enough to achieve &#x0002B;95% and &#x0002B;99% accuracy, respectively. These findings contributes to the body of knowledge by paving a way for the application of EEG to real-world ecological biometric applications and by demonstrating methods to learn the minimal amount of data required for such applications.</p></abstract>
<kwd-group>
<kwd>biometrics</kwd>
<kwd>EEG</kwd>
<kwd>feature extraction</kwd>
<kwd>machine learning</kwd>
<kwd>deep learning</kwd>
<kwd>graph neural networks</kwd>
</kwd-group>
<contract-num rid="cn001">8/CRT/6183</contract-num>
<contract-sponsor id="cn001">Science Foundation Ireland<named-content content-type="fundref-id">10.13039/501100001602</named-content></contract-sponsor>
<counts>
<fig-count count="6"/>
<table-count count="2"/>
<equation-count count="5"/>
<ref-count count="57"/>
<page-count count="13"/>
<word-count count="10087"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Among the different techniques used for measuring brain activity, Electroencephalography (EEG) is a method used for measuring voltage fluctuation inside the electrical field generated by a subject&#x00027;s brain. The sampling process involves placing surface electrodes onto a subject&#x00027;s scalp, capturing brain activity at different regions. The resulting EEG signals are non-stationary, meaning frequency values along time are not constant but variable, making them unpredictable. Recent advances in Deep Learning (DL) (Scarselli et al., <xref ref-type="bibr" rid="B38">2008</xref>; Goodfellow et al., <xref ref-type="bibr" rid="B17">2014</xref>; Vaswani et al., <xref ref-type="bibr" rid="B48">2017</xref>) have allowed for new and improved applications of EEG signals. These applications include but are not limited to: EEG-based biometric systems (DelPozo-Banos et al., <xref ref-type="bibr" rid="B16">2015</xref>), Brain-Computer Interfaces (BCIs) (Vaid et al., <xref ref-type="bibr" rid="B47">2015</xref>), emotion recognition systems (Jenke et al., <xref ref-type="bibr" rid="B20">2014</xref>), alertness analysis (Subasi, <xref ref-type="bibr" rid="B43">2005</xref>) and medical diagnosis and progression assessment for neurological diseases such as Alzheimer&#x00027;s (Cassani et al., <xref ref-type="bibr" rid="B9">2018</xref>) or epilepsy (Acharya et al., <xref ref-type="bibr" rid="B2">2015</xref>). These systems are often composed of an automatic feature extraction layer that extracts representative features from the raw EEG signal, and a classifier for fitting a specific target feature. For example, Wilaiprasitporn et al. (<xref ref-type="bibr" rid="B54">2019</xref>) used a combination of Convolutional Neural Networks along with different type of temporal layers to learn distinctive features, and a fully connected layer for the purpose of classification. Similarly, Ullah et al. (<xref ref-type="bibr" rid="B46">2018</xref>) used a model composed of several 1-D convolutional layers for feature extraction and denoising to obtain representative features subsequently handled by a fully connected layer with a majority voting scheme.</p>
<p>In detail, this research focuses on biometrics, a process consisting of measuring and analyzing unique physical characteristics from humans to authenticate their unique identity. Among the methods employed for such a task, EEG has proven to be a robust biometrics method (Jayarathne et al., <xref ref-type="bibr" rid="B19">2017</xref>). As pointed by Campisi and La Rocca (<xref ref-type="bibr" rid="B6">2014</xref>), EEG poses several advantages when compared to traditional exposed biometric modalities such as fingerprint or iris scanning. Namely: The advantage of being more robust against spoofing attacks (Revett, <xref ref-type="bibr" rid="B36">2012</xref>), given the technological complexity of synthetically generating brain signals replicating the response for a specific individual. The advantage of being universal, meaning the sampling process is valid for anyone with no brain pathological conditions. Lastly, EEG poses the advantage of the aliveness detection problem not being present, meaning we can assume that the user is alive to produce brain signals sensor can read. This assumption does not necessarily hold in other biometric methods. There are two main applications for EEG in biometric systems: Person identification (PI), identifying an individual from a group of known subjects, and person authentication (PA), accepting or denying the identity of a particular individual. Both tasks&#x00027; objective is to extract features with discriminating properties from EEG signals that a statistical model can later classify. Depending on the task subjects perform while their EEG recordings are taking place, it is possible to divide biometric systems into two groups, namely task-dependent (Wilaiprasitporn et al., <xref ref-type="bibr" rid="B54">2019</xref>) and task-independent (Kong et al., <xref ref-type="bibr" rid="B23">2018</xref>). Training a task-dependent system often means presenting the model data where all subjects perform a particular task. Whereas task-independent systems use data sampled from various tasks, making them more robust at handling unseen tasks at inference time. This work is centered around using EEG signals applied to task-dependent PI and is devoted to answering the following research question:</p>
<list list-type="bullet">
<list-item><p><bold>RQ:</bold> What is the minimal amount of data required to train an affective EEG-based person identification model for biometric applications with and without an explicit feature extraction method?</p></list-item>
</list>
<p>The structure of this document is as follows: Section 2 provides a brief literature review on the different methods used for EEG-Based Person Identification. Section 3 explains the design and methodology, with reproducibility and replicability in mind. Section 4 discusses the results and findings and finally Section 5 summarizes these informing possible future directions.</p>
</sec>
<sec id="s2">
<title>2. Related Work</title>
<p>There exists evidence suggesting that different subjects produce different EEG responses (Marcel and Mill&#x000E1;n, <xref ref-type="bibr" rid="B26">2007</xref>), with studies dating as back as the 1970&#x00027;s suggesting the uniqueness in EEG responses depends on each subject genetic information (Vogel, <xref ref-type="bibr" rid="B52">1970</xref>; Anokhin et al., <xref ref-type="bibr" rid="B4">1992</xref>). This unique property from EEG data makes it suitable for biometrics applications. However, building such applications has proven challenging to build and deploy given the non-stationary nature of EEG signals and the high quantity of noise generated during recording by, for example, muscle movements, eye blinking, electrode displacement. Additionally, there has been evidence showing that the EEG responses from an individual may vary depending on their emotions. For example, EEG responses may vary if the participant is emotionally attached to the watched person (Koelstra et al., <xref ref-type="bibr" rid="B21">2009</xref>). EEG analysis often requires the extraction of high-level features fed into statistical models that automatically learn to find hidden patterns within the features. There are several different feature extraction techniques present in the literature. The feature extraction methods that are applied most often to this domain are Power Spectral Density (PSD) (Riera et al., <xref ref-type="bibr" rid="B37">2007</xref>), Auto-Regressive (AR) model coefficients (Mohammadi et al., <xref ref-type="bibr" rid="B27">2006</xref>; Zivot and Wang, <xref ref-type="bibr" rid="B57">2006</xref>), Discrete Wavelet transform (DWT) (Guo et al., <xref ref-type="bibr" rid="B18">2009</xref>; Murugappan et al., <xref ref-type="bibr" rid="B30">2009</xref>), and Hilbert-Huang Transform (HHT) (Li et al., <xref ref-type="bibr" rid="B24">2009</xref>). Among these techniques, wavelet-based feature extraction stands up given its performance at characterizing non-stationary data.</p>
<p>The available literature on person identification from EEG features dates as far as the 90&#x00027;s [see (Poulos et al., <xref ref-type="bibr" rid="B33">1999a</xref>,<xref ref-type="bibr" rid="B34">b</xref>,<xref ref-type="bibr" rid="B35">c</xref>)]. Scholars used sets of features obtained from an AR model and a set of different Machine Learning (ML) classifiers obtaining accuracy scores ranging from 80 to 100% at classifying unique individuals from a pool of four total subjects. Mohammadi et al. (<xref ref-type="bibr" rid="B27">2006</xref>) trained a competitive neural network model with features obtained from an AR. In this case, scholars performed experiments training models for single-channel and multi-channel settings with up to three channels. They collected 24 s of EEG readings for 10 participants and used 15 s to train a model and a total of 24 s for testing. Results demonstrated that it was possible to obtain accuracies ranging from 80% up to 100% by using this technique on a larger pool of subjects. Brigham and Kumar (<xref ref-type="bibr" rid="B5">2010</xref>) used coefficients obtained for each channel from an AR model along with a Support Vector Machine (SVM) for PI. The scholars obtained testing classification accuracies close to 99% at identifying subjects from a pool of 120 subjects, suggesting that it is possible to obtain separable features from each individual. Shedeed (<xref ref-type="bibr" rid="B40">2011</xref>) used a combination of features obtained from applying the Discrete Fourier Transform and the Discrete Wavelet Transform followed by a voting scheme to retain the most discriminating features, which are input to a multi-layer perceptron classifier. However, their dataset was composed solely of three subjects, casting doubt on the generalization ability of their model. Thomas and Vinod (<xref ref-type="bibr" rid="B45">2018</xref>) used PSD features obtained from a single channel from the gamma (&#x003B3;, 30&#x02013;50 Hz) band along with simple correlation-based matching operations on a dataset composed of 109 subjects with an error rate of 0.0196. These results suggest that it is possible to apply a non-parametric operation over the raw EEG signal that maps input data onto a set of discriminative features for each subject.</p>
<p>Recent advances in DL have allowed the development of models with more expressing capabilities. Unlike approaches using hand-crafted features, DL models are often composed of two parts: a parametric feature extraction block and a classifier. The feature extraction block aims to learn representative features from the raw input signal automatically. The classifier uses the extracted features as input and predicts a target value such as a specific subject identity. Automatic feature extraction layers use spatial layers to reduce the dimensionality of the input data while trying to maintain informative features. This process usually consists of employing a combination of convolutional and pooling layers (Das et al., <xref ref-type="bibr" rid="B14">2017</xref>, <xref ref-type="bibr" rid="B15">2018</xref>). The input dimensionality of the data is usually not very large, given that there is a limited number of electrodes. The small input size leads to convolutional models with a relatively small number of layers, often ranging from 1 to 5 (Maiorana, <xref ref-type="bibr" rid="B25">2020</xref>). &#x000D6;zdenizci et al. (<xref ref-type="bibr" rid="B31">2019</xref>) used a dataset composed of 128 subjects with readings obtained from different sessions and proposed an adversarial network from an invariant representation-learning perspective. Their goal was to create features that would be invariant between different recording sessions. Their approach consisted of an encoder convolutional layer to encode the raw signals and two classification layers. One of these layers is the identifier that predicts the correct person ID. The other one is the adversary layer that tries to predict the recording session ID. The assumption was that the encoder obtained features that the identifier could analyse but the adversary network could not, which led to session-invariant features. Additionally to the spatial layers, some approaches consider temporal layers to include temporal information in the features. Recurrent Neural Networks often include a combination of spatial and temporal layers (Chen et al., <xref ref-type="bibr" rid="B11">2020</xref>). Das et al. (<xref ref-type="bibr" rid="B13">2019</xref>) proposed a model based on a combination of spatial layers (Convolutional) and temporal layers by employing Long Short Term Memory (LSTM). Results pointed to accuracies close to 100% on a pool of 109 subjects using the raw data from recordings of 64 channels. Wilaiprasitporn et al. (<xref ref-type="bibr" rid="B54">2019</xref>) proposed an alternative method using Gated Recurrent Units instead of the LSTM, achieving 99.9&#x02013;100% test accuracy on the DEAP dataset (Koelstra et al., <xref ref-type="bibr" rid="B22">2011</xref>) using all available channels. They also achieved 99.17% test accuracy, reducing the number of channels from 32 to 5, using their Frontal and Parietal (F3, F4, Fz, F7, and F8) configuration. Chen et al. (<xref ref-type="bibr" rid="B11">2020</xref>) demonstrated that it was possible to train models that automatically learn feature extraction techniques from the raw signal using spatial and temporal layers. These rich features are then injected into a classifier, where the final target feature is the subject ID. They achieved 96% test accuracy on a pool of 157 subjects sampled from four different experiments, demonstrating their model&#x00027;s ability to generalize.</p>
<p>The recent surge in interest in Graph Neural Network (GNN) architectures has also affected EEG-based applications. Applying GNNs to domains where elements and relationships exist has proven to be very effective. These applications include but are not limited to traffic flow prediction, point cloud classification, and text classification (Cui et al., <xref ref-type="bibr" rid="B12">2019</xref>; Yao et al., <xref ref-type="bibr" rid="B55">2019</xref>; Shi and Rajkumar, <xref ref-type="bibr" rid="B41">2020</xref>). These systems also pose the advantage of implicitly stating the relationships between the different elements, making them potentially more useful for explainability purposes (Pope et al., <xref ref-type="bibr" rid="B32">2019</xref>; Vilone and Longo, <xref ref-type="bibr" rid="B49">2020</xref>, <xref ref-type="bibr" rid="B50">2021a</xref>,<xref ref-type="bibr" rid="B51">b</xref>). The use of graphs for EEG-based applications remains underexplored. Zhong et al. (<xref ref-type="bibr" rid="B56">2020</xref>) used a Graph Convolutional Network model for EEG-based emotion recognition obtaining results comparable to other non-graph-based methods. Wang et al. (<xref ref-type="bibr" rid="B53">2020</xref>) used a graph to map brain functional connectivity for Participant Classification without making use of message-passing Graph Neural Networks. Their work consists of creating brain connectivity networks using multi-channel settings, node centrality, and global network metrics to compute features for the different nodes composing the graph. They prove that generated networks have significant inter-individual distinctiveness, making them suitable for biometric applications.</p>
<p>As these studies suggest, it is possible to obtain meaningful distinguishable patterns of behavior, as gathered by EEG signals, of a specific user, for biometric applications, with current approaches demonstrating error rates close to 0% (Wilaiprasitporn et al., <xref ref-type="bibr" rid="B54">2019</xref>; Seha and Hatzinakos, <xref ref-type="bibr" rid="B39">2021</xref>). However, these approaches rely on datasets composed of several minutes of recordings for each participant, sampling significant EEG windows or using a high number of training samples. Neglecting that sampling data for each participant is an expensive process that could burden real-world scenarios. Carri&#x000F3;n-Ojeda et al. (<xref ref-type="bibr" rid="B8">2019</xref>) performed a similar study to find out the smallest EEG window size for a biometric system to perform subject identification. The researchers used a fixed amount of training samples per participant, leaving open the question of the minimum total time in seconds required for a PI system to be trained. Similarly, Seha and Hatzinakos (<xref ref-type="bibr" rid="B39">2021</xref>) developed a PI model with session-invariant features. Researchers emphasize using smaller windows than previous approaches while obtaining better results. However, they also used most of their available data for training the model, still not focusing on the total training dataset size but solely on the size of the epochs. These reasons motivated the focus on this work to provide hindsight on the minimal amount of affective-based EEG data required from each subject to train task-dependent biometric models. This minimal amount is a combination of the number of samples used for training and the window size for each sample. An exhaustive experiment have been designed, and described in the next section.</p>
</sec>
<sec id="s3">
<title>3. Design and Methodology</title>
<p>This research work is devoted to understanding the minimal EEG data required to identify unique discriminative person-specific features. In order to tackle this goal, the experiment includes a combination of different training data sizes, which are defined based on the size of the sampled EEG windows and the number of training samples available per participant. In detail, these experiments include the training of models by employing several training data sizes along with different machine learning techniques, including a <italic>Logistic Regressor</italic> (LR), <italic>Multi-Layer Perceptron</italic> (MLP), <italic>1-D Convolutional Neural Network</italic> (CNN), and a Graph Convolutional Network (<italic>Graphconv</italic>) and a set of feature extraction methods, namely Power Spectral Density (PSD) and Wavelet Energy (WE). <xref ref-type="fig" rid="F1">Figure 1</xref> depicts the experimental pipeline for a given configuration. Each experimental configuration is used 10 times with different, randomly sampled data to obtain results distributions. Each pipeline aims to train and evaluate a classification model that can perform Person Identification given a sub-sampled window of EEG recordings. Training a discriminative model with data for a pool of <italic>p</italic> participants that automatically learns to differentiate input features depending on the participant from whom the features have been sub-sampled. More formally, the goal is to approximate a parametric function <italic>f</italic><sub>&#x003C3;</sub> that maps an input feature matrix <bold>X</bold> to a vector of predicted probabilities <bold>p</bold> with one entry per known participant. <bold><italic>p</italic><sub><italic>i</italic></sub></bold> represents the probability of the input features belonging to the subject with ID <italic>i</italic>. All together we can represent it as <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>p</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> where <italic>n</italic> indicates the number of EEG channels used, <italic>d</italic> depends on the feature extraction method and denotes the number of obtained features per channel and <italic>p</italic> is the total number of participants. Models are trained and validated on a randomly sub-sampled portion of the complete dataset, where there are <italic>tr</italic><sub><italic>s</italic></sub> training samples and 100 validation samples for each user. The evaluation process for a trained model consists of making predictions on the remaining unseen data and measuring classification accuracy. In detail, classification is correct when the highest value for a predicted probability vector <bold><italic>p</italic><sub><italic>s</italic></sub></bold> matches the ID of the subject the recordings were sampled from.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Single experiment pipeline. (1) Read raw data, remove baseline recording and pre-process. (2) Sample (60 &#x000D7; 128)/<italic>w</italic> EEG windows with size <italic>w</italic> for each video. (3) Extract features from all windows individually and normalize. (4) Split dataset into train, validation and testing(<italic>tr</italic><sub><italic>s</italic></sub>/ 100/ <italic>s</italic> &#x02212; <italic>tr</italic><sub><italic>s</italic></sub> &#x02212; 100). (5) Train chosen model using train and validation datasets. (6) Evaluate trained model on the testing dataset. Record experiment parameters, training metrics, and testing accuracy.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fninf-16-844667-g0001.tif"/>
</fig>
<p>We hypothesize that by running a set of experimental configurations and evaluating the results, it is possible to provide an answer to the research question stated in Section 1. The overall experiment, consisting of the different configuration settings, is performed to test the following hypothesis:</p>
<list list-type="bullet">
<list-item><p><bold>H1:</bold> IF different affective EEG-based person identification models are trained using a variable dataset size (number of training samples and window sizes) and pre-process data extracting wavelet energy or power spectral density</p>
<p>THEN it exists a minimum amount of data that can lead these models to achieve 95% and 99% of accuracy, and this amount is significantly lower than the amount associated to the same models trained without the pre-processing methods.</p>
</list-item>
</list>
<p>In the context of training person identification models, the minimum required amount of data is defined based on the total number of train EEG windows used (<italic>tr</italic><sub><italic>s</italic></sub>) and the size for each window in seconds (<italic>w</italic>/128). The minimal amount of data and the best learning approach and feature extraction combination can be determined by comparing the result distributions and picking the experimental configuration able to obtain &#x0002B;95% and &#x0002B;99% accuracy on unseen data with the least amount of total training data. Moreover, performing a Mann&#x02013;Whitney <italic>U</italic>-test over the results distributions where training data time is the same but the window size differs would answer whether it is better to have fewer oversized windows or many smaller ones.</p>
<sec>
<title>3.1. Data Acquisition and Pre-processing</title>
<p>The experiment uses the dataset for emotion analysis using EEG, physiological, and video signals (DEAP) (Koelstra et al., <xref ref-type="bibr" rid="B22">2011</xref>). It consists of a set of EEG signals and physiological and face information from 32 different participants, recorded as they watched 40 one-minute music videos meant to trigger different kinds of emotions, making this dataset task-dependent in general and affective-based in particular. The DEAP dataset poses the advantage of being sampled under naturalistic, ecological conditions. Instead of evoking a certain kind of response in a short time window, the DEAP dataset provides a continuous stream of EEG data that better resembles a real-world environment.</p>
<p>Thirty-two electrodes (<italic>n</italic> &#x0003D; 32) were positioned using the international 10-20 system. The current research study only considers the EEG signals. The DEAP dataset has a total of 1,280 segments (32 participants &#x000D7; 40 videos), where each segment consists of 32 channel readings, each with 3 s of baseline signal followed by the main recordings 60 s. The pre-processing stage consists of applying a set of pre-processing techniques to the original EEG recordings, including:</p>
<list list-type="bullet">
<list-item><p>Downsample to 128 Hz</p></list-item>
<list-item><p>High-pass filter to 4.0 Hz</p></list-item>
<list-item><p>Low-pass filter to 45.0 Hz.</p></list-item>
</list>
<p>A band-pass filter is applied in order to extract the main frequencies of EEG human waves (Abhang et al., <xref ref-type="bibr" rid="B1">2016</xref>) (Theta [4&#x02013;8 Hz], Alpha [8&#x02013;12 Hz], Beta [12&#x02013;35 Hz], Gamma), dismissing the Delta (0.5&#x02013;4 Hz) band due to it being associated with sleep activity (Amzica and Steriade, <xref ref-type="bibr" rid="B3">1998</xref>). Non-parametric down-sampling is used to compress the data size. The first 3 s representing baseline recordings are removed from the rest of the signal and a sliding, non-overlapping window of varying size <italic>w</italic> is applied in order to split each segment into (60&#x0002A;128)/<italic>w</italic> windows (see <xref ref-type="fig" rid="F2">Figure 2B</xref>). After the pre-processing steps, the total number of samples <italic>s</italic> is equal to 32 (participants) &#x000D7; 40 (videos) &#x000D7; (60&#x0002A;128)/<italic>w</italic> (windows per segment). Each pre-processed window <inline-formula><mml:math id="M3"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> consists of readings for <italic>n</italic> channels with <italic>w</italic> readings for each of them. The DEAP dataset performed its sampling with the following electrodes: [Fp1, AF3, F7, F3, FC1, FC5, T7, C3, CP1, CP5, P7, P3, Pz, PO3, O1, Oz, O2, PO4, P4, P8, CP6, CP2, C4, T8, FC6, FC2, F4, F8, AF4, Fp2, Fz, Czz].</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The top part of this figure shows the hyperparameters (in red) that can be modified to build the train, validation, and test datasets starting from the pre-processed segments. <bold>(A)</bold> Complete dataset visualized as segments. There exists one segment per video. The shape of each segment is defined based on the total number of channels (<italic>n</italic>), the number of subjects(<italic>p</italic>), and the total number of data points per video with baseline removed (60 s at 128 Hz). Panel <bold>(B)</bold> shows how the EEG windows are formed based on the window size (<italic>w</italic>). <bold>(C)</bold> Transforms pre-processed windows <inline-formula><mml:math id="M2"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> using the feature extraction method that is set as hyperparameter. <bold>(D)</bold> Randomly samples <italic>tr</italic><sub><italic>s</italic></sub>, 100 and total samples(s) - <italic>tr</italic><sub><italic>s</italic></sub> - 100 for train, validation and testing, respectively, per participant. The bottom part of this figure shows the considered learning approaches, with non-linear approaches having a <italic>hc</italic> hyperparameter for deciding the number of hidden channels. <bold>(E)</bold> Subject ID prediction &#x00177; is calculated by choosing the index of the maximum value in the predicted probability vector <bold>p</bold>. Training loss is computed using CELF (Equation 5). Evaluation is calculated as the number of correct predictions divided by the total number of predictions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fninf-16-844667-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.2. Feature Extraction</title>
<p>All windows sub-sampled from the original segments are input to the feature extraction block. This block extracts a set of discriminating normalized features <bold>X</bold> from a pre-processed window <inline-formula><mml:math id="M4"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> (see <xref ref-type="fig" rid="F2">Figure 2C</xref>). This research work considers two feature extraction methods: Power Spectral Density (PSD) and Wavelet Energy (WE). The rationale for selecting these two techniques is that wavelet-based methods are often applied to characterize non-stationary signals, hence making it suitable for EEG data (Guo et al., <xref ref-type="bibr" rid="B18">2009</xref>; Chai et al., <xref ref-type="bibr" rid="B10">2017</xref>). On the other hand, using power spectral density to characterize EEG data is widely used in literature, given their high expressive power at characterizing all kinds of signals (Carrier et al., <xref ref-type="bibr" rid="B7">2001</xref>). PSD features are based on the Fourier Transform (FT), whereas WE features are computed using the Wavelet Transform (WT). These transforms are invertible frequency decompositions of the original signal. The FT is often more suitable for stationary signals. In contrast, the WT better captures the characteristics of non-stationary data (Sifuzzaman et al., <xref ref-type="bibr" rid="B42">2009</xref>). The main difference between FT and WT is that FT is localized within the frequency domain, and WT is both within frequency and time domains. The WT includes information about the point in time the frequencies occurs, whereas the FT does not. Whether time information makes a difference for biometric purposes is subject to study.</p>
<p>To investigate their respective characterization power, raw signals are used as baseline, as depicted in <xref ref-type="fig" rid="F3">Figure 3A</xref>. Raw signal features represent normalized signal power over time. Applying normalization over the input signal results in feature matrix <inline-formula><mml:math id="M5"><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow></mml:msub></mml:math></inline-formula> with <italic>d</italic> &#x0003D; <italic>w</italic>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Feature extraction methods for a given sample with window size <italic>w</italic> = 128. <bold>(A)</bold> Raw signal features consist of z-score normalized power over time. <bold>(B)</bold> Power spectral density features are obtained by averaging power spectral density from each frequency band ([4&#x02013;8, 8&#x02013;16, 16&#x02013;32, 32&#x02013;64 Hz]) then z-score normalizing sample-wise. <bold>(C)</bold> Discrete wavelet transformation (DWT) is applied to the input signal and then recursively over the approximation coefficients. It is possible to extract the wavelet energy feature vectors using the wavelet energy formula (Equation 1) over detail coefficients [2&#x02013;5]. The final feature matrix is built based on these wavelet energy feature vectors with z-score normalization applied to them.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fninf-16-844667-g0003.tif"/>
</fig>
<sec>
<title>3.2.1. Power Spectral Density (PSD) vs. Wavelet Energy (WE)</title>
<p>The PSD represents signal power over frequencies. Spectral density estimation is performed over the input data <inline-formula><mml:math id="M6"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> in order to transform the time domain signal into the frequency domain. This domain transformation is performed using Welch&#x00027;s method and aims at extracting the four desired frequency bands ([4&#x02013;8, 8&#x02013;16, 16&#x02013;32, 32&#x02013;64 Hz]). Using PSD, it is possible to extract any desired frequency band. This specific band choice has the objective of matching the ones obtained by performing the wavelet decomposition for obtaining the WE feature. The final feature for every frequency band is obtained by averaging the PSD from each frequency band. Computing this feature for each channel and each frequency band results in feature matrix <bold>X</bold> &#x02208; &#x0211D;<sup><italic>n</italic> &#x000D7; <italic>d</italic></sup> with <italic>d</italic> &#x0003D; 4. Each row in the feature matrix represents a channel, and each column represents a frequency band as depicted in <xref ref-type="fig" rid="F3">Figure 3B</xref>.</p>
<p>Wavelet energy feature is computed by splitting a pre-processed window <inline-formula><mml:math id="M7"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> into different frequency ranges using a series of DWTs (see <xref ref-type="fig" rid="F3">Figure 3C</xref>). The wavelet transform requires the selection of the mother wavelet and the decomposition level (N). The mother wavelet acts as a filter that is applied recursively to the input signal. In order to extract highly representative features, the mother wavelet will ideally have similar characteristics as the input signal. The &#x0201C;db4&#x0201D; mother wavelet has proven to be the most effective for EEG data (Subasi, <xref ref-type="bibr" rid="B44">2007</xref>) and hence was used for all experiments. Given pre-processed signals are filtered from 4 to 45 Hz, setting the maximum decomposition level equal to 5 avoids adding noise to the features since choosing a higher decomposition level would include frequencies lower than 4 Hz. <xref ref-type="table" rid="T1">Table 1</xref> displays the frequency range obtained from each decomposition level. Each decomposition level consists of two filters and two down samples. These filters produce an approximation coefficient containing low-frequency information and a detail coefficient containing high-frequency information. By performing the wavelet energy computations on the coefficients ranging from D2 to D5 and disregarding the D1 coefficient, we expect to reduce the total noise added onto the features, as pointed by Mohammadi et al. (<xref ref-type="bibr" rid="B28">2017</xref>). These coefficients for each EEG window by applying the DWT recursively over the signal&#x00027;s approximation coefficients at different decomposition levels. The shape of the coefficients will vary for each level due to the down sample applied at each step. With all the coefficients computed, it is possible to extract the wavelet energy feature using (1)</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>d</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>e</italic><sub><italic>j</italic></sub> is the wavelet energy feature for a certain channel at decomposition level <italic>j</italic>. <italic>k</italic><sub><italic>j</italic></sub> denotes the number of wavelet coefficients for decomposition level <italic>j</italic> and <bold><italic>d</italic><sub><italic>j</italic></sub></bold>(<italic>i</italic>) is the value of the detail coefficient at point <italic>i</italic> for decomposition level <italic>j</italic>. Calculating the wavelet energy feature for all channels results in feature matrix <bold>X</bold> &#x02208; &#x0211D;<sup><italic>n</italic> &#x000D7; <italic>d</italic></sup> with <italic>d</italic> &#x0003D; 4 where, similarly to the power spectral density features, each row represents a channel obtained from an electrode, and each column represents a particular frequency band.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Wavelet signal frequencies for different decomposition levels.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Decomposed signal</bold></th>
<th valign="top" align="center"><bold>Frequencies (Hz)</bold></th>
<th valign="top" align="center"><bold>Decomposition level</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">D1 (noise)</td>
<td valign="top" align="center">64&#x02013;128</td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">D2</td>
<td valign="top" align="center">32&#x02013;64</td>
<td valign="top" align="center">2</td>
</tr>
<tr>
<td valign="top" align="left">D3</td>
<td valign="top" align="center">16&#x02013;32</td>
<td valign="top" align="center">3</td>
</tr>
<tr>
<td valign="top" align="left">D4</td>
<td valign="top" align="center">8&#x02013;16</td>
<td valign="top" align="center">4</td>
</tr>
<tr>
<td valign="top" align="left">D5</td>
<td valign="top" align="center">4&#x02013;8</td>
<td valign="top" align="center">5</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec>
<title>3.3. Learning Approaches</title>
<p>Experiment settings include four different learning approaches, all of them sharing an identical goal: mapping an input sample <bold>X</bold> to a vector of predicted probabilities <bold>p</bold> indicating the probability of the input sample belonging to each unique individual. There is a linear model and three non-linear models inside the four different approaches. The <italic>Logistic Regressor</italic> (LR) is linear, meaning the transformation between input and output features does not include a non-linear activation function. It hence is a direct mapping between input and output. The MLP acts as a baseline for non-linear approaches. The CNN model consists of a 1-D Convolutional Layer followed by an MLP (CNN) to analyse whether merging the different frequency bands into one could reduce trainable parameters without sacrificing predictive performance. Finally, the <italic>GraphConv</italic> model consists of a Graph Convolutional layer followed by an MLP. This latter approach aims at modeling spatial dependencies in 3-D space among the different channels. All of the non-linear approaches have an additional hyperparameter <italic>h</italic> ranging from 64 up to 2048 that defines the number of hidden channels for different layers as summarized in <xref ref-type="table" rid="T2">Table 2</xref>. <xref ref-type="fig" rid="F2">Figure 2</xref> displays a simplified visualization of every considered learning approach.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Configuration for each experimental phase.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Setting name</bold></th>
<th valign="top" align="left"><bold>Values main exp</bold>.</th>
<th valign="top" align="left"><bold>Values raw exp</bold>.</th>
<th valign="top" align="left"><bold>Description</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Feature extraction method</td>
<td valign="top" align="left">raw, psd, wav</td>
<td valign="top" align="left">raw</td>
<td valign="top" align="left">Feature to extract from each raw sample to build the dataset</td>
</tr>
<tr>
<td valign="top" align="left">EEG window size (<italic>w</italic>)</td>
<td valign="top" align="left">0.25, 0.5, 1, 1.5, 2</td>
<td valign="top" align="left">0.5</td>
<td valign="top" align="left">Number of seconds per sub-sampled EEG window</td>
</tr>
<tr>
<td valign="top" align="left">Number of train samples (<italic>tr</italic><sub><italic>s</italic></sub>)</td>
<td valign="top" align="left">1, 2, 4, 8</td>
<td valign="top" align="left">16, 32, 64, 128, 256, 512</td>
<td valign="top" align="left">Number of train samples per participant</td>
</tr>
<tr>
<td valign="top" align="left">Model</td>
<td valign="top" align="left">MLP, CNN, GraphConv, LR</td>
<td valign="top" align="left">MLP, CNN, GraphConv, LR</td>
<td valign="top" align="left">Model architecture</td>
</tr>
<tr>
<td valign="top" align="left">Hidden channels</td>
<td valign="top" align="left">64, 128, 256, 512, 1,024, 2,048</td>
<td valign="top" align="left">512, 1,024</td>
<td valign="top" align="left">Number of hidden channels within non-linear models</td>
</tr>
<tr>
<td valign="top" align="left">Dropout rate</td>
<td valign="top" align="left">0.25</td>
<td valign="top" align="left">0.5</td>
<td valign="top" align="left">Probability for neurons before a Dropout layer not updating their weights after a backward pass</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The <italic>logistic regressor</italic> consists of a single linear layer mapping the input sample <bold>X</bold> to a vector with a length equal to the number of participants. Input feature matrix gets flattened into a vector (with length <italic>n</italic> &#x000D7; <italic>d</italic>), input to the linear layer. Finally, a softmax activation function gets applied over the output of the linear layer in order to convert its output into a vector of probabilities <bold>p</bold> with their combined values adding up to 1. The LR is a linear model meaning the output will be a linear input transformation. High testing accuracies would indicate that the classes (unique subjects in this case) are linearly separable for the input features.</p>
<p>The MLP has a single hidden layer with a varying number of hidden channels <italic>h</italic>. The input feature matrix is flattened and fed into the hidden layer, followed by a ReLU non-linear activation function. Dropout is applied after the hidden layer, essentially deactivating at random some of the neuron weight updates to reduce overfitting and improve the final model&#x00027;s generalization capabilities. A softmax activation function is used after the output layer to obtain the desired probability vector <bold>p</bold>. As opposed to the LR, the MLP is non-linear, meaning it can apply learnable non-linear transformations to the input, which could be a benefit in some cases and is subject to study.</p>
<p>A 1-D Convolutional Neural Network merges all features into a single value by combining the different features before feeding this into the final classifier, an MLP with varying hidden channels <italic>h</italic>. The convolutional layer reduces the trainable parameter count by minimizing the shape of the input features from <italic>d</italic> to 1. The objective of the CNN is to map the input feature matrix <bold>X</bold> &#x02208; &#x0211D;<sup><italic>n</italic> &#x000D7; <italic>d</italic></sup> into a smaller, more informative feature vector <bold>x</bold> &#x02208; &#x0211D;<sup><italic>n</italic></sup>. A ReLU function gets applied to the CNN&#x00027;s output and after the first layer of the MLP classifier. Similar to the MLP, dropout gets applied after the hidden layer of the MLP. A Softmax function acts as a final activation after the output layer in order to obtain the probability vector <bold>p</bold>.</p>
<p>Graph Neural Network refers to a learning approach that directly operates on graphs and can perform convolution-like operations efficiently over the input data by employing a message-passing algorithm. It is possible to model complex relationships between individual elements by modeling complex structures as graphs. A <italic>graph</italic> is defined as <inline-formula><mml:math id="M9"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">G</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">E</mml:mi></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> where <inline-formula><mml:math id="M10"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:math></inline-formula> represents the set of nodes that form the graph and <inline-formula><mml:math id="M11"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">E</mml:mi></mml:mrow></mml:math></inline-formula> refers to the set of edges connecting said nodes. <inline-formula><mml:math id="M12"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">V</mml:mi></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> can be represented as a feature matrix <bold>X</bold> &#x02208; &#x0211D;<sup><italic>n</italic> &#x000D7; <italic>d</italic></sup> with <italic>n</italic> being the total number of nodes and <italic>d</italic> being the node feature dimensionality. An undirected edge set <inline-formula><mml:math id="M13"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">E</mml:mi></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> can be viewed as a symmetric weighted adjacency matrix <bold>A</bold> &#x02208; &#x0211D;<sup><italic>n</italic> &#x000D7; <italic>n</italic></sup>. <bold><italic>A</italic><sub><italic>ij</italic></sub></bold> represents the importance of the relationship between node <italic>i</italic> and node <italic>j</italic>. It is possible to model a sample <bold>X</bold> into a graph by representing each electrode as a node, where each initial node feature <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is directly obtained from each sample feature matrix <bold>X</bold>. The adjacency matrix <bold>A</bold> is constructed based on electrode 3-D positional information (see <xref ref-type="fig" rid="F4">Figure 4</xref>), and it remains constant for all samples given they follow the same 10-20 electrode placement scheme.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Pipeline for computing graph adjacency matrix from electrode 3-D position. <bold>(A)</bold> Electrode positional information in 3-D space. <bold>(B)</bold> Distance matrix representing Euclidean distance between electrodes. <bold>(C)</bold> Local connections using (2) with &#x003B4; &#x0003D; 5. <bold>(D)</bold> Local connection matrix with added self-loops. <bold>(E)</bold> Global connections to local connection matrix with self-loops using (3).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fninf-16-844667-g0004.tif"/>
</fig>
<p>The first step to building the adjacency matrix is to compute local connections among electrodes. The purpose of local connectivity is to allow nodes to spread information along with their neighborhood, which should theoretically help model spatial relationships inside the graph. This process consists of constructing a distance matrix based on electrode positions and then applying the local connection formula (2); with &#x003B4; &#x0003D; 5 denoting a calibration constant used to keep around 20% of the possible connections to maximize the efficiency of the network topology according to Zhong et al. (<xref ref-type="bibr" rid="B56">2020</xref>). <italic>d</italic><sub><italic>ij</italic></sub> represents the Euclidean distance between electrode <italic>i</italic> and electrode <italic>j</italic>.</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>A</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign="left" style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mstyle mathvariant="bold"><mml:mtext>min</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mstyle></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">if&#x000A0;</mml:mtext><mml:mstyle class="math"><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003E;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>1</mml:mn><mml:mtext class="textrm" mathvariant="normal"></mml:mtext></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">otherwise</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The last step is to add global connections, which aim at modeling EEG asymmetry, which are connections between distant channels in opposing parts of the brain. {<italic>GC</italic>} is the set of nodes that require a global connection between them ({(FP1, FP2), (AF3, AF4), (F5, F6), (FC5, FC6), (C5, C6), (CP5, CP6), (P5, P6), (PO5, PO6),(O1, O2)}) given their superior performance at modeling EEG asymmetry as stated by Zhong et al. (<xref ref-type="bibr" rid="B56">2020</xref>).</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>A</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign="left" style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">if&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>C</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">otherwise</mml:mtext></mml:mtd></mml:mtr><mml:mtr></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The graph layer used in this research work is similar to that presented in Morris et al. (<xref ref-type="bibr" rid="B29">2019</xref>), where they proposed a higher-order graph convolutional network. This architecture proposes a node update function in which nodes are updated based on their features and the features from neighboring nodes. The graph adjacency matrix defines node neighborhoods by specifying relevant connections among the different nodes that compose the graph. Node features are updated at every layer using (4), where <inline-formula><mml:math id="M17"><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes node i&#x00027;s feature vector at time <italic>t</italic>, &#x00398;<sub>1</sub> and &#x00398;<sub>2</sub> are weight matrices that are updated using loss backpropagation and <inline-formula><mml:math id="M18"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> indicates node i&#x00027;s set of neighboring nodes. The number of updates performed on every node is equal to the total number of graph convolutional layers that is included onto the architecture, either with shared or un-shared weights. Having <italic>k</italic> convolutional layer implies considering <italic>k-hop</italic> neighborhoods around each node. This experiment only considers a single graph convolutional layer due to increased performance in preliminary experiments.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo>&#x00398;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x00398;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The <italic>GraphConv</italic> architecture&#x00027;s tail is the same as the MLP but uses the features obtained from its graph convolutional layer as input features instead of inputting the window directly. One message-passing step gets applied to the input feature matrix <bold>X</bold> and this resulting feature matrix is then input to the MLP, which makes the final prediction for subject ID.</p>
</sec>
<sec>
<title>3.4. Training, Validation, and Testing</title>
<p>The training process consists of two phases: the feature extraction phase and the no-feature extraction or the raw phase. The objective of the feature extraction phase is to find out what is the minimum amount of data required to fit a Person Identification model. The raw phase is similar but aims to answer how much data is needed to train a PI model if a learned feature extraction method replaces the feature extraction step. <xref ref-type="table" rid="T2">Table 2</xref> displays the different experimental variants . Each experiment was run 10 times with a Montecarlo sampling strategy, sub-sampling several training and validation samples from randomly selected videos at random timesteps for every participant (i.e., sample second 1 of video 23 for all participants). Combining the different experimental settings leads to 11,400 experiments for the feature extraction phase and 380 for the raw phase.</p>
<p>The first step is building the dataset for each desired configuration, according to the employed feature extraction method and the EEG window size. The output of such as feature extraction step is a set of feature matrices <inline-formula><mml:math id="M20"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>X</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The target feature is equal to the subject ID. Each training sample also has an associated adjacency matrix (A) that Graph-based models use and non-graph models disregard. The number of train samples per participant (<italic>tr</italic><sub><italic>s</italic></sub>) is a hyperparameter for each experiment. Setting <italic>tr</italic><sub><italic>s</italic></sub> &#x0003D; 1 means having just one training example per participant, making this problem one-shot learning (see <xref ref-type="fig" rid="F2">Figure 2D</xref>). The total number of samples per participant depends on each experimental setting, and it ranges from 1,200 (<italic>w</italic> = 2) to 9,600 (<italic>w</italic> = 0.25). The number of validation samples was kept constant at 100 for all experiments. The remaining data, unseen during training, is used for model evaluation. The number of validation samples was selected so that it is possible to keep it constant for all experiments. This was for ensuring that there are enough validation samples to avoid overfitting, and enough unseen testing samples to evaluate the models when the window size is largest (<italic>w</italic> = 2). Every learning approach is presented with the same training and validation data to make the model comparison fairer. Additionally, the size for each window is an experimental setting. Whether the model performs better with smaller EEG windows or bigger ones is also subject to this study. We hypothesize that having more, smaller samples would lead the model to achieve higher prediction performance with the same total amount of time per subject. Including more samples introduces the model to examples of different mental states.</p>
<p>All approaches were trained using the Cross Entropy Loss function (CEL, Equation 5)</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Given that PI is a supervised classification problem and every architecture output are the logits of the last layer passed through a SoftMax function, along with the Adam optimizer with a fixed learning rate of 0.0005 as it empirically demonstrated superior performance in preliminary experiments. The batch size was kept constant at 32 since this is the minimum amount the training dataset could have if <italic>tr</italic><sub><italic>s</italic></sub> &#x0003D; 1. The dropout rate did not seem to make much difference in the results. It was also kept constant at 0.25 in the main experimental phase and raised to 0.5 in the raw phase to improve the model&#x00027;s generalization capabilities. The values for the dropout probabilities were also decided based on the preliminary experiments. An early stopping patience limit was set equal to 30 to avoid wasting resources during training. The whole training took place using 4 Tesla P100 GPUs running Pytorch and Pytorch Geometric. All experiments run in 2 weeks, each taking around 8&#x02013;10 min to complete on average.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Results and Discussion</title>
<p>This research work&#x00027;s main objective is to provide insight into the minimal amount of data required for training an accurate Person Identification model. Each section provides results for an experimental phase. The first phase aims at testing the hypotheses that consider an explicit feature extraction block. This first phase also aims to test the hypothesis of having more minor windows over fewer bigger ones. The second phase tests the hypothesis with no explicit feature extraction block. Results are presented in these two first sections and interpreted in the subsequent discussion section.</p>
<sec>
<title>4.1. Feature Extraction Phase Results</title>
<p>The first step is to analyse whether it is better to have more minor EEG windows with more available examples or bigger windows with fewer examples. Answering this question is possible by performing a Mann&#x02013;Whitney <italic>U</italic>-test over the result distributions to compare testing accuracy for different window sizes and the number of train samples. Results show that there was a significant difference in test accuracy between experiment configurations that used (2 &#x000D7; 0.25 s) as opposed to (1 &#x000D7; 0.5 s) [<italic>t</italic> = 2.618, <italic>p</italic> &#x0003C; 0.01]. Similarly, experiments using (4 &#x000D7; 0.25 s) also performed significantly better than the ones using (1 &#x000D7; 1.0 s) [<italic>t</italic> = 3.584, <italic>p</italic> &#x0003C; 0.01]. For the 2-s combination, smaller windows again performed better, using (8 &#x000D7; 0.25 s) obtained better results than (1 &#x000D7; 2.0 s) [<italic>t</italic> = 4.629, <italic>p</italic> &#x0003C; 0.01]. Same was the case for (8 &#x000D7; 0.5 s) vs. (2 &#x000D7; 2.0) [<italic>t</italic> = 2.705, <italic>p</italic> &#x0003C; 0.01]. <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure 1</xref> displays the complete results for the independent <italic>U</italic>-Test.</p>
<p><xref ref-type="supplementary-material" rid="SM1">Supplementary Tables 1</xref>&#x02013;<xref ref-type="supplementary-material" rid="SM1">3</xref> display the results for the main experimental phase, listing the mean and standard deviation of the test accuracy distributions for every model. These results include a row for every learning approach and a column indicating the number of seconds of data available for each architecture during training. <xref ref-type="fig" rid="F5">Figure 5</xref> provides a visualization of these three tables for easier understanding. From these results, it is possible to observe that the LR model achieved very high performance using pre-processed features, meaning these features have a very high discriminating power that makes them linearly separable. We observe little difference in performance between the same models with different hidden channels. The MLP was able to get outstanding performance, superior to the LR but requiring more trainable parameters. The CNN achieved lower mean performance than the other models while having a significantly higher significant standard deviation. Finally, <italic>GraphConv</italic> had a slightly worse performance than the LR and the MLP whilst superior to the CNN. All models achieved &#x0002B;99% test accuracy using 8 s of data for training.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Results visualization for the feature extraction phase. Plots represent test accuracy against available training data for each participant (in seconds). Each row represents one learning approach with a legend accompanying non-linear approaches and displaying the number of hidden channels. Each column represents the employed feature extraction method, considering the raw signal as a baseline for comparison.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fninf-16-844667-g0005.tif"/>
</fig>
</sec>
<sec>
<title>4.2. No Feature Extraction (Raw) Phase</title>
<p><xref ref-type="table" rid="T2">Table 2</xref> displays the results for the raw phase. <xref ref-type="fig" rid="F6">Figure 6</xref> provides a visualization of these results for easier understanding. Results suggest that the Graph Convolutional model handles raw data better than the other models. We attribute <italic>GraphConv&#x00027;s</italic> success to its ability to extract compressed, representative features in its upper layers from the raw data alone. Similarly, the CNN model also creates features from raw data automatically. However, results demonstrate that features learned by the CNN contain less expressive information than the <italic>GraphConv</italic>, at least when constraining the training dataset size. Results also imply that the automatic feature extraction method is better than feeding the raw data into a classifier directly, given that both LR and MLP performed poorly when the data was not processed.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Results visualization for the no-feature extraction (raw) phase. Graphs represent accuracy against available training data for each participant (in seconds). There exists a legend accompanying non-linear approaches and displaying the number of hidden channels.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fninf-16-844667-g0006.tif"/>
</fig>
<p>In this case, the minimal amount of data required for the best performing model to achieve 95% is 32 s per participant, with 64 EEG windows with a size of 0.5 s. As for 99% test accuracy, the model requires at least 128 s with 256 windows of size 0.5 s. In this case, it was only possible to achieve accuracy higher than 95% employing the <italic>GraphConv</italic> learning approach, indicating that this model&#x00027;s learned feature extraction step is superior to other learning approaches.</p>
</sec>
<sec>
<title>4.3. Discussion</title>
<p>Findings suggest that having more training examples per participant leads to better accuracy on test data than having fewer, bigger windows. We assume these results are because training the model with data from different trials introduces different mental states for every subject, enhancing the model&#x00027;s ability to recognize discriminative user features and improving predictive performance.</p>
<p>The minimal amount of data required to train a Person Identification model is 1 and 3 s to obtain 95% and 99% test accuracy, respectively. The model that was able to achieve 95% test accuracy and 99% with this minimal amount of data was the MLP. The results for the LR were similar to the MLP while requiring fewer parameters to be trained. Both feature extraction methods performed similarly, with PSD obtaining slightly better performance. Using raw data alone and avoiding the feature extraction step, it was possible to train a Graph convolutional model with 32 and 128 s of total data to obtain the target accuracies. These results confirm that deep learning approaches require more data than machine learning to allow robust training. We attribute the <italic>GraphConv</italic> model&#x00027;s success to its ability to learn complex feature representation automatically from raw data alone. <italic>GraphConv</italic> introduces spatial relationships along the different electrodes dismissed in other learning approaches. By performing a single graph convolution, the model could obtain rich features that were input to the MLP. Compared to the MLP alone, <italic>GraphConv</italic> shows superior performance. This superior performance indicates that the message-passing operation along the graph adds discriminative information onto the features, beneficial to the overall model predictive performance.</p>
<p>There exist limitations to the experiments, specifically related to the DEAP dataset. Firstly, as stated in Maiorana (<xref ref-type="bibr" rid="B25">2020</xref>), having single-session recordings might lead the model to learn session-specific exogenous conditions instead of personal biometric traits. However, the DEAP dataset could be considered multi-session. It consists of several independent recordings with a break at the half-mark, where electrodes are recalibrated. These facts make us believe the non-stationary nature of EEG data is preserved. Furthermore, exogenous conditions are minimized due to the experiment taking place for a long time duration. Secondly, the dataset is affective-based and sampled under naturalistic conditions. These characteristics are inherent to this dataset and should be addressed if comparing our results against any other work.</p>
</sec>
</sec>
<sec id="s5">
<title>5. Conclusions and Future Work</title>
<p>Biometric systems using EEG data have seen a rise in popularity with recent advances in ML and DL, along with a deeper understanding of feature extraction methods. These techniques have allowed researchers to obtain low error rates in a wide variety of settings including using recordings sampled from different sessions and having a larger pool of unique subjects. However, there exists little work exploring how much data is needed to train accurate person identification models. This work focuses on answering this question. For this purpose a set of experiments was run comparing different train data sizes, learning approaches and feature extraction methods. The predictive performance of the trained models is then measured based on their accuracy for the unseen, test dataset. This work demonstrates that, in the context of affective EEG-based person identification, having just 1 s of data per participant is enough for training a model to achieve &#x0002B;95% test accuracy. Similarly, 3 s of data suffices to train a model with &#x0002B;99% test accuracy, if a suitable feature extraction method is chosen. The use of raw data is also explored comparing three simple models against a Graph Neural Network which is able to achieve &#x0002B;95% test accuracy using 32 s and &#x0002B;99% test accuracy using 128 s. EEG-based biometrics poses several advantages over traditional biometric systems, the main one it being secure, however, it also poses several disadvantages. Sampling EEG data is not as straightforward as scanning an iris or a fingerprint, several electrodes have to be attached to each subject and their signals must be recorded for a certain amount of time. Data collection is an expensive process. If real-world EEG-based biometric applications were to be implemented, data for every subject would have to be gathered and this data collection process should be as efficient as possible. Current research work focuses mainly on the method but fails to address this fact. These results pave the way for EEG biometrics used in real world applications, the objective being providing insight on how the data quantity changes the performance of the predictive models. Our findings suggest that in order to train an affective EEG-based biometric model, sampling data for a few seconds for every subject would be enough, as opposed to other methods found in literature which use several minutes from each subject to train their respective models. Each subject would not need to sit in a room with electrodes attached to its head for more than a few seconds. As a future direction, the number of electrodes attached to every subject to measure brain signals could also be studied, making the sampling process for this kind of application more straightforward and hence more efficient. Furthermore, task-independent methods considering a wider variety of tasks and session invariant features considering the non-stationary nature of EEG signals on multi-session datasets could also be subject to exploration.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://www.eecs.qmul.ac.uk/mmv/datasets/deap/">https://www.eecs.qmul.ac.uk/mmv/datasets/deap/</ext-link>.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>The main research work was developed by CG-T. LL and BB provided great insight and feedback throughout the project as well as revising the different article drafts that were developed, providing suggestions for potential improvements. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This publication has emanated from research supported in part by a grant from SFI Centre for Research Training in Machine Learning at Technological University Dublin under Grant number 18/CRT/6183. For the purpose of OpenAccess, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission. We thank Giuliano Anselmi from IBM for granting us access to computing resources and helping us configure the IBM powerstations our models were trained on.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fninf.2022.844667/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fninf.2022.844667/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>

<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abhang</surname> <given-names>P. A.</given-names></name> <name><surname>Gawali</surname> <given-names>B. W.</given-names></name> <name><surname>Mehrotra</surname> <given-names>S. C.</given-names></name></person-group> (<year>2016</year>). <source>Introduction to EEG-and Speech-Based Emotion Recognition</source>. <edition>1st ed</edition>. <publisher-name>Academic Press</publisher-name>. <pub-id pub-id-type="doi">10.1016/B978-0-12-804490-2.00007-5</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Acharya</surname> <given-names>U. R.</given-names></name> <name><surname>Fujita</surname> <given-names>H.</given-names></name> <name><surname>Sudarshan</surname> <given-names>V. K.</given-names></name> <name><surname>Bhat</surname> <given-names>S.</given-names></name> <name><surname>Koh</surname> <given-names>J. E.</given-names></name></person-group> (<year>2015</year>). <article-title>Application of entropies for automated diagnosis of epilepsy using EEG signals: a review</article-title>. <source>Knowledge Based Syst</source>. <volume>88</volume>, <fpage>85</fpage>&#x02013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2015.08.004</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amzica</surname> <given-names>F.</given-names></name> <name><surname>Steriade</surname> <given-names>M.</given-names></name></person-group> (<year>1998</year>). <article-title>Electrophysiological correlates of sleep delta waves</article-title>. <source>Electroencephalogr. Clin. Neurophysiol</source>. <volume>107</volume>, <fpage>69</fpage>&#x02013;<lpage>83</lpage>. <pub-id pub-id-type="doi">10.1016/S0013-4694(98)00051-0</pub-id><pub-id pub-id-type="pmid">9751278</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anokhin</surname> <given-names>A.</given-names></name> <name><surname>Steinlein</surname> <given-names>O.</given-names></name> <name><surname>Fischer</surname> <given-names>C.</given-names></name> <name><surname>Mao</surname> <given-names>Y.</given-names></name> <name><surname>Vogt</surname> <given-names>P.</given-names></name> <name><surname>Schalt</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>1992</year>). <article-title>A genetic study of the human low-voltage electroencephalogram</article-title>. <source>Hum. Genet</source>. <volume>90</volume>, <fpage>99</fpage>&#x02013;<lpage>112</lpage>. <pub-id pub-id-type="doi">10.1007/BF00210751</pub-id><pub-id pub-id-type="pmid">1427795</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brigham</surname> <given-names>K.</given-names></name> <name><surname>Kumar</surname> <given-names>B. V.</given-names></name></person-group> (<year>2010</year>). <article-title>Subject identification from electroencephalogram (EEG) signals during imagined speech,</article-title> in <source>2010 Fourth IEEE International Conference on Biometrics: Theory, Applications and Systems (BTAS)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/BTAS.2010.5634515</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campisi</surname> <given-names>P.</given-names></name> <name><surname>La Rocca</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>Brain waves for automatic biometric-based user recognition</article-title>. <source>IEEE Trans. Inform. Forensics Secur</source>. <volume>9</volume>, <fpage>782</fpage>&#x02013;<lpage>800</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2014.2308640</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carrier</surname> <given-names>J.</given-names></name> <name><surname>Land</surname> <given-names>S.</given-names></name> <name><surname>Buysse</surname> <given-names>D. J.</given-names></name> <name><surname>Kupfer</surname> <given-names>D. J.</given-names></name> <name><surname>Monk</surname> <given-names>T. H.</given-names></name></person-group> (<year>2001</year>). <article-title>The effects of age and gender on sleep EEG power spectral density in the middle years of life (ages 20-60 years old)</article-title>. <source>Psychophysiology</source> <volume>38</volume>, <fpage>232</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1111/1469-8986.3820232</pub-id><pub-id pub-id-type="pmid">11347869</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Carri&#x000F3;n-Ojeda</surname> <given-names>D.</given-names></name> <name><surname>Mejia-Vallejo</surname> <given-names>H.</given-names></name> <name><surname>Fonseca-Delgado</surname> <given-names>R.</given-names></name> <name><surname>G&#x000F3;mez-Gil</surname> <given-names>P.</given-names></name> <name><surname>Ramirez-Cort&#x000E9;s</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>A method for studying how much time of EEG recording is needed to have a good user identification,</article-title> in <source>2019 IEEE Latin American Conference on Computational Intelligence (LA-CCI)</source> (<publisher-loc>Guayaquil</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/LA-CCI47412.2019.9037054</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cassani</surname> <given-names>R.</given-names></name> <name><surname>Estarellas</surname> <given-names>M.</given-names></name> <name><surname>San-Martin</surname> <given-names>R.</given-names></name> <name><surname>Fraga</surname> <given-names>F. J.</given-names></name> <name><surname>Falk</surname> <given-names>T. H.</given-names></name></person-group> (<year>2018</year>). <article-title>Systematic review on resting-state EEG for Alzheimer&#x00027;s disease diagnosis and progression assessment</article-title>. <source>Dis. Mark</source>. 2018, 5174815. <pub-id pub-id-type="doi">10.1155/2018/5174815</pub-id><pub-id pub-id-type="pmid">30405860</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chai</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>D.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>A fast, efficient domain adaptation technique for cross-domain electroencephalography (EEG)-based emotion recognition</article-title>. <source>Sensors</source> <volume>17</volume>, <fpage>1014</fpage>. <pub-id pub-id-type="doi">10.3390/s17051014</pub-id><pub-id pub-id-type="pmid">28467371</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Mao</surname> <given-names>Z.</given-names></name> <name><surname>Yao</surname> <given-names>W.</given-names></name> <name><surname>Huang</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>EEG-based biometric identification with convolutional neural network</article-title>. <source>Multimedia Tools Appl</source>. <volume>79</volume>, <fpage>10655</fpage>&#x02013;<lpage>10675</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-019-7258-4</pub-id><pub-id pub-id-type="pmid">31281339</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cui</surname> <given-names>Z.</given-names></name> <name><surname>Henrickson</surname> <given-names>K.</given-names></name> <name><surname>Ke</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Traffic graph convolutional recurrent neural network: a deep learning framework for network-scale traffic learning and forecasting</article-title>. <source>IEEE Trans. Intell. Transport. Syst</source>. <volume>21</volume>, <fpage>4883</fpage>&#x02013;<lpage>4894</lpage>. <pub-id pub-id-type="doi">10.1109/TITS.2019.2950416</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Das</surname> <given-names>B. B.</given-names></name> <name><surname>Kumar</surname> <given-names>P.</given-names></name> <name><surname>Kar</surname> <given-names>D.</given-names></name> <name><surname>Ram</surname> <given-names>S. K.</given-names></name> <name><surname>Babu</surname> <given-names>K. S.</given-names></name> <name><surname>Mohapatra</surname> <given-names>R. K.</given-names></name></person-group> (<year>2019</year>). <article-title>A spatio-temporal model for EEG-based person identification</article-title>. <source>Multimedia Tools Appl</source>. <volume>78</volume>, <fpage>28157</fpage>&#x02013;<lpage>28177</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-019-07905-6</pub-id><pub-id pub-id-type="pmid">22174850</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Das</surname> <given-names>R.</given-names></name> <name><surname>Maiorana</surname> <given-names>E.</given-names></name> <name><surname>Campisi</surname> <given-names>P.</given-names></name></person-group> (<year>2017</year>). <article-title>Visually evoked potential for EEG biometrics using convolutional neural network,</article-title> in <source>2017 25th European Signal Processing Conference (EUSIPCO)</source> (<publisher-loc>Rome</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>951</fpage>&#x02013;<lpage>955</lpage>. <pub-id pub-id-type="doi">10.23919/EUSIPCO.2017.8081348</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Das</surname> <given-names>R.</given-names></name> <name><surname>Maiorana</surname> <given-names>E.</given-names></name> <name><surname>Campisi</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Motor imagery for EEG biometrics using convolutional neural network,</article-title> in <source>2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source> (<publisher-loc>Rome</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2062</fpage>&#x02013;<lpage>2066</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.2018.8461909</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>DelPozo-Banos</surname> <given-names>M.</given-names></name> <name><surname>Travieso</surname> <given-names>C. M.</given-names></name> <name><surname>Weidemann</surname> <given-names>C. T.</given-names></name> <name><surname>Alonso</surname> <given-names>J. B.</given-names></name></person-group> (<year>2015</year>). <article-title>EEG biometric identification: a thorough exploration of the time-frequency domain</article-title>. <source>J. Neural Eng</source>. 12, 056019. <pub-id pub-id-type="doi">10.1088/1741-2560/12/5/056019</pub-id><pub-id pub-id-type="pmid">26394698</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Goodfellow</surname> <given-names>I.</given-names></name> <name><surname>Pouget-Abadie</surname> <given-names>J.</given-names></name> <name><surname>Mirza</surname> <given-names>M.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Warde-Farley</surname> <given-names>D.</given-names></name> <name><surname>Ozair</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Generative adversarial nets,</article-title> in <source>Advances in Neural Information Processing Systems 27</source> (Curran Associates, Inc.). Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2014/file/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf">https://proceedings.neurips.cc/paper/2014/file/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf</ext-link></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>L.</given-names></name> <name><surname>Rivero</surname> <given-names>D.</given-names></name> <name><surname>Seoane</surname> <given-names>J. A.</given-names></name> <name><surname>Pazos</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Classification of EEG signals using relative wavelet energy and artificial neural networks,</article-title> in <source>Proceedings of the first ACM/SIGEVO Summit on Genetic and Evolutionary Computation</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>177</fpage>&#x02013;<lpage>184</lpage>. <pub-id pub-id-type="doi">10.1145/1543834.1543860</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jayarathne</surname> <given-names>I.</given-names></name> <name><surname>Cohen</surname> <given-names>M.</given-names></name> <name><surname>Amarakeerthi</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Survey of EEG-based biometric authentication,</article-title> in <source>2017 IEEE 8th International Conference on Awareness Science and Technology (iCAST)</source> (<publisher-loc>Fukushima</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>324</fpage>&#x02013;<lpage>329</lpage>. <pub-id pub-id-type="doi">10.1109/ICAwST.2017.8256471</pub-id><pub-id pub-id-type="pmid">26364201</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jenke</surname> <given-names>R.</given-names></name> <name><surname>Peer</surname> <given-names>A.</given-names></name> <name><surname>Buss</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Feature extraction and selection for emotion recognition from EEG</article-title>. <source>IEEE Trans. Affect. Comput</source>. <volume>5</volume>, <fpage>327</fpage>&#x02013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2014.2339834</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Koelstra</surname> <given-names>S.</given-names></name> <name><surname>M&#x000FC;hl</surname> <given-names>C.</given-names></name> <name><surname>Patras</surname> <given-names>I.</given-names></name></person-group> (<year>2009</year>). <article-title>EEG analysis for implicit tagging of video data,</article-title> in <source>2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops</source> (<publisher-loc>Queen Mary</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/ACII.2009.5349482</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koelstra</surname> <given-names>S.</given-names></name> <name><surname>Muhl</surname> <given-names>C.</given-names></name> <name><surname>Soleymani</surname> <given-names>M.</given-names></name> <name><surname>Lee</surname> <given-names>J.-S.</given-names></name> <name><surname>Yazdani</surname> <given-names>A.</given-names></name> <name><surname>Ebrahimi</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Deap: A database for emotion analysis; using physiological signals</article-title>. <source>IEEE Trans. Affect. Comput</source>. <volume>3</volume>, <fpage>18</fpage>&#x02013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1109/T-AFFC.2011.15</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kong</surname> <given-names>X.</given-names></name> <name><surname>Kong</surname> <given-names>W.</given-names></name> <name><surname>Fan</surname> <given-names>Q.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name> <name><surname>Cichocki</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Task-independent EEG identification via low-rank matrix decomposition,</article-title> in <source>2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)</source> (<publisher-loc>Tokyo</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>412</fpage>&#x02013;<lpage>419</lpage>. <pub-id pub-id-type="doi">10.1109/BIBM.2018.8621531</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Yingle</surname> <given-names>F.</given-names></name> <name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Qinye</surname> <given-names>T.</given-names></name></person-group> (<year>2009</year>). <article-title>Sleep stage classification based on EEG Hilbert-Huang transform,</article-title> in <source>2009 4th IEEE Conference on Industrial Electronics and Applications</source> (<publisher-loc>Hangzhou</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3676</fpage>&#x02013;<lpage>3681</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maiorana</surname> <given-names>E..</given-names></name></person-group> (<year>2020</year>). <article-title>Deep learning for EEG-based biometric recognition</article-title>. <source>Neurocomputing</source> <volume>410</volume>, <fpage>374</fpage>&#x02013;<lpage>386</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.06.009</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marcel</surname> <given-names>S.</given-names></name> <name><surname>Mill&#x000E1;n</surname> <given-names>J. R.</given-names></name></person-group> (<year>2007</year>). <article-title>Person authentication using brainwaves (EEG) and maximum a posteriori model adaptation</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>29</volume>, <fpage>743</fpage>&#x02013;<lpage>752</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2007.1012</pub-id><pub-id pub-id-type="pmid">17299229</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mohammadi</surname> <given-names>G.</given-names></name> <name><surname>Shoushtari</surname> <given-names>P.</given-names></name> <name><surname>Molaee Ardekani</surname> <given-names>B.</given-names></name> <name><surname>Shamsollahi</surname> <given-names>M. B.</given-names></name></person-group> (<year>2006</year>). <article-title>Person identification by using AR model for EEG signals,</article-title> in <source>Proceeding of World Academy of Science, Engineering and Technology</source> (<publisher-loc>Prague</publisher-loc>), <fpage>281</fpage>&#x02013;<lpage>285</lpage>.<pub-id pub-id-type="pmid">27389803</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohammadi</surname> <given-names>Z.</given-names></name> <name><surname>Frounchi</surname> <given-names>J.</given-names></name> <name><surname>Amiri</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Wavelet-based emotion recognition system using EEG signal</article-title>. <source>Neural Comput. Appl</source>. <volume>28</volume>, <fpage>1985</fpage>&#x02013;<lpage>1990</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-015-2149-8</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Morris</surname> <given-names>C.</given-names></name> <name><surname>Ritzert</surname> <given-names>M.</given-names></name> <name><surname>Fey</surname> <given-names>M.</given-names></name> <name><surname>Hamilton</surname> <given-names>W. L.</given-names></name> <name><surname>Lenssen</surname> <given-names>J. E.</given-names></name> <name><surname>Rattan</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Weisfeiler and Leman go neural: higher-order graph neural networks</article-title>. <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, 4602&#x02013;4609. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.33014602</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Murugappan</surname> <given-names>M.</given-names></name> <name><surname>Nagarajan</surname> <given-names>R.</given-names></name> <name><surname>Yaacob</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Comparison of different wavelet features from EEG signals for classifying human emotions,</article-title> in <source>2009 IEEE Symposium on Industrial Electronics</source> &#x00026; <italic>Applications</italic> (<publisher-loc>Perlis; Jejawi;Kangar</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>836</fpage>&#x02013;<lpage>841</lpage>. <pub-id pub-id-type="doi">10.1109/ISIEA.2009.5356339</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>&#x000D6;zdenizci</surname> <given-names>O.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Koike-Akino</surname> <given-names>T.</given-names></name> <name><surname>Erdo&#x0011F;mu&#x0015F;</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Adversarial deep learning in EEG biometrics</article-title>. <source>IEEE Signal Process. Lett</source>. <volume>26</volume>, <fpage>710</fpage>&#x02013;<lpage>714</lpage>. <pub-id pub-id-type="doi">10.1109/LSP.2019.2906826</pub-id><pub-id pub-id-type="pmid">31814690</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pope</surname> <given-names>P. E.</given-names></name> <name><surname>Kolouri</surname> <given-names>S.</given-names></name> <name><surname>Rostami</surname> <given-names>M.</given-names></name> <name><surname>Martin</surname> <given-names>C. E.</given-names></name> <name><surname>Hoffmann</surname> <given-names>H.</given-names></name></person-group> (<year>2019</year>). <article-title>Explainability methods for graph convolutional neural networks,</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Los Alamitos, CA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>), <fpage>10764</fpage>&#x02013;<lpage>10773</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01103</pub-id><pub-id pub-id-type="pmid">34817762</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poulos</surname> <given-names>M.</given-names></name> <name><surname>Rangoussi</surname> <given-names>M.</given-names></name> <name><surname>Alexandris</surname> <given-names>N.</given-names></name></person-group> (<year>1999a</year>). <article-title>Neura network based person identification using EEG features,</article-title> in <source>1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings</source> (<publisher-loc>Paphos</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1117</fpage>&#x02013;<lpage>1120</lpage>. <pub-id pub-id-type="doi">10.1109/ICASSP.1999.759940</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poulos</surname> <given-names>M.</given-names></name> <name><surname>Rangoussi</surname> <given-names>M.</given-names></name> <name><surname>Chrissikopoulos</surname> <given-names>V.</given-names></name> <name><surname>Evangelou</surname> <given-names>A.</given-names></name></person-group> (<year>1999b</year>). <article-title>Parametric person identification from the EEG using computational geometry,</article-title> in <source>ICECS&#x00027;99. Proceedings of ICECS&#x00027;99. 6th IEEE International Conference on Electronics, Circuits and Systems</source> (<publisher-loc>Paphos</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1005</fpage>&#x02013;<lpage>1008</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Poulos</surname> <given-names>M.</given-names></name> <name><surname>Rangoussi</surname> <given-names>M.</given-names></name> <name><surname>Chrissikopoulos</surname> <given-names>V.</given-names></name> <name><surname>Evangelou</surname> <given-names>A.</given-names></name></person-group> (<year>1999c</year>). <article-title>Person identification based on parametric processing of the EEG,</article-title> in <source>ICECS&#x00027;99. Proceedings of ICECS&#x00027;99. 6th IEEE International Conference on Electronics, Circuits and Systems</source>. (<publisher-loc>Paphos</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>283</fpage>&#x02013;<lpage>286</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Revett</surname> <given-names>K..</given-names></name></person-group> (<year>2012</year>). <article-title>Cognitive biometrics: a novel approach to person authentication</article-title>. <source>Int. J. Cogn. Biometr</source>. <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1504/IJCB.2012.046516</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riera</surname> <given-names>A.</given-names></name> <name><surname>Soria-Frisch</surname> <given-names>A.</given-names></name> <name><surname>Caparrini</surname> <given-names>M.</given-names></name> <name><surname>Grau</surname> <given-names>C.</given-names></name> <name><surname>Ruffini</surname> <given-names>G.</given-names></name></person-group> (<year>2007</year>). <article-title>Unobtrusive biometric system based on electroencephalogram analysis</article-title>. <source>EURASIP J. Adv. Signal Process</source>. <volume>2008</volume>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1155/2008/143728</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scarselli</surname> <given-names>F.</given-names></name> <name><surname>Gori</surname> <given-names>M.</given-names></name> <name><surname>Tsoi</surname> <given-names>A. C.</given-names></name> <name><surname>Hagenbuchner</surname> <given-names>M.</given-names></name> <name><surname>Monfardini</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>The graph neural network model</article-title>. <source>IEEE Trans. Neural Netw</source>. <volume>20</volume>, <fpage>61</fpage>&#x02013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1109/TNN.2008.2005605</pub-id><pub-id pub-id-type="pmid">19068426</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Seha</surname> <given-names>S. N. A.</given-names></name> <name><surname>Hatzinakos</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>Longitudinal assessment of EEG biometrics under auditory stimulation: a deep learning approach,</article-title> in <source>2021 29th European Signal Processing Conference (EUSIPCO)</source> (<publisher-loc>Toronto, ON</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1386</fpage>&#x02013;<lpage>1390</lpage>. <pub-id pub-id-type="doi">10.23919/EUSIPCO54536.2021.9616098</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shedeed</surname> <given-names>H. A..</given-names></name></person-group> (<year>2011</year>). <article-title>A new method for person identification in a biometric security system based on brain EEG signal processing,</article-title> in <source>2011 World Congress on Information and Communication Technologies</source> (<publisher-loc>Cairo</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1205</fpage>&#x02013;<lpage>1210</lpage>. <pub-id pub-id-type="doi">10.1109/WICT.2011.6141420</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>W.</given-names></name> <name><surname>Rajkumar</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Point-GNN: graph neural network for 3D object detection in a point cloud,</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Pittsburgh, PA</publisher-loc>), <fpage>1711</fpage>&#x02013;<lpage>1719</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00178</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sifuzzaman</surname> <given-names>M.</given-names></name> <name><surname>Islam</surname> <given-names>M.</given-names></name> <name><surname>Ali</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Application of wavelet transform and its advantages compared to Fourier transform</article-title>. <source>J. Phys. Sci</source> 13.<pub-id pub-id-type="pmid">34093115</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Subasi</surname> <given-names>A..</given-names></name></person-group> (<year>2005</year>). <article-title>Automatic recognition of alertness level from EEG by using neural network and wavelet coefficients</article-title>. <source>Expert Syst. Appl</source>. <volume>28</volume>, <fpage>701</fpage>&#x02013;<lpage>711</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2004.12.027</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Subasi</surname> <given-names>A..</given-names></name></person-group> (<year>2007</year>). <article-title>EEG signal classification using wavelet feature extraction and a mixture of expert model</article-title>. <source>Expert Sys. Appl</source>. <volume>32</volume>, <fpage>1084</fpage>&#x02013;<lpage>1093</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2006.02.005</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thomas</surname> <given-names>K. P.</given-names></name> <name><surname>Vinod</surname> <given-names>A. P.</given-names></name></person-group> (<year>2018</year>). <article-title>Eeg-based biometric authentication using gamma band power during rest state</article-title>. <source>Circ. Syst. Signal Process</source>. <volume>37</volume>, <fpage>277</fpage>&#x02013;<lpage>289</lpage>. <pub-id pub-id-type="doi">10.1007/s00034-017-0551-4</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ullah</surname> <given-names>I.</given-names></name> <name><surname>Hussain</surname> <given-names>M.</given-names></name> <name><surname>Aboalsamh</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>An automated system for epilepsy detection using EEG brain signals based on deep learning approach</article-title>. <source>Expert Syst. Appl</source>. <volume>107</volume>, <fpage>61</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2018.04.021</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vaid</surname> <given-names>S.</given-names></name> <name><surname>Singh</surname> <given-names>P.</given-names></name> <name><surname>Kaur</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>EEG signal analysis for BCI interface: a review,</article-title> in <source>2015 Fifth International Conference on Advanced Computing</source> &#x00026; <italic>Communication Technologies</italic> (Chandigarh: IEEE), <fpage>143</fpage>&#x02013;<lpage>147</lpage>. <pub-id pub-id-type="doi">10.1109/ACCT.2015.72</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Attention is all you need,</article-title> in <source>Advances in Neural Information Processing Systems</source> (Long Beach, CA: Curran Associates, Inc.), <fpage>5998</fpage>&#x02013;<lpage>6008</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf">https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf</ext-link></citation>
</ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vilone</surname> <given-names>G.</given-names></name> <name><surname>Longo</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>Explainable artificial intelligence: a systematic review</article-title>. <source>arXiv [Preprint]</source>. arXiv:2006.00093. <pub-id pub-id-type="doi">10.48550/arXiv.2006.00093</pub-id><pub-id pub-id-type="pmid">35390650</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vilone</surname> <given-names>G.</given-names></name> <name><surname>Longo</surname> <given-names>L.</given-names></name></person-group> (<year>2021a</year>). <article-title>Classification of explainable artificial intelligence methods through their output formats</article-title>. <source>Mach. Learn. Knowledge Extract</source>. <volume>3</volume>, <fpage>615</fpage>&#x02013;<lpage>661</lpage>. <pub-id pub-id-type="doi">10.3390/make3030032</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vilone</surname> <given-names>G.</given-names></name> <name><surname>Longo</surname> <given-names>L.</given-names></name></person-group> (<year>2021b</year>). <article-title>Notions of explainability and evaluation approaches for explainable artificial intelligence</article-title>. <source>Inform. Fusion</source>. <volume>76</volume>, <fpage>89</fpage>&#x02013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2021.05.009</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vogel</surname> <given-names>F..</given-names></name></person-group> (<year>1970</year>). <article-title>The genetic basis of the normal human electroencephalogram (EEG)</article-title>. <source>Humangenetik</source> <volume>10</volume>, <fpage>91</fpage>&#x02013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1007/BF00295509</pub-id><pub-id pub-id-type="pmid">5528299</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Abbass</surname> <given-names>H. A.</given-names></name></person-group> (<year>2020</year>). <article-title>Brainprint: EEG biometric identification based on analyzing brain connectivity graphs</article-title>. <source>Pattern Recogn</source>. 105, 107381. <pub-id pub-id-type="doi">10.1016/j.patcog.2020.107381</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wilaiprasitporn</surname> <given-names>T.</given-names></name> <name><surname>Ditthapron</surname> <given-names>A.</given-names></name> <name><surname>Matchaparn</surname> <given-names>K.</given-names></name> <name><surname>Tongbuasirilai</surname> <given-names>T.</given-names></name> <name><surname>Banluesombatkul</surname> <given-names>N.</given-names></name> <name><surname>Chuangsuwanich</surname> <given-names>E.</given-names></name></person-group> (<year>2019</year>). <article-title>Affective EEG-based person identification using the deep learning approach</article-title>. <source>IEEE Trans. Cogn. Dev. Syst</source>. <volume>12</volume>, <fpage>486</fpage>&#x02013;<lpage>496</lpage>. <pub-id pub-id-type="doi">10.1109/TCDS.2019.2924648</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>L.</given-names></name> <name><surname>Mao</surname> <given-names>C.</given-names></name> <name><surname>Luo</surname> <given-names>Y.</given-names></name></person-group> (<year>2019</year>). <article-title>Graph convolutional networks for text classification,</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Chicago, IL</publisher-loc>), <fpage>7370</fpage>&#x02013;<lpage>7377</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.33017370</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>P.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Miao</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>EEG-based emotion recognition using regularized graph neural networks</article-title>. <source>IEEE Trans. Affect. Comput</source>. 1&#x02013;1. <pub-id pub-id-type="doi">10.1109/TAFFC.2020.2994159</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zivot</surname> <given-names>E.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Vector autoregressive models for multivariate time series,</article-title> in <source>Modeling Financial Time Series with S-Plustextregistered</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer New York</publisher-name>), <fpage>385</fpage>&#x02013;<lpage>429</lpage>. <pub-id pub-id-type="doi">10.1007/978-0-387-32348-0_11</pub-id><pub-id pub-id-type="pmid">24515888</pub-id></citation></ref>
</ref-list> 
</back>
</article>