<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="review-article" dtd-version="2.3">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2022.839619</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Personality Trait Recognition: A Survey</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhao</surname>
<given-names>Xiaoming</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Tang</surname>
<given-names>Zhiwei</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Zhang</surname>
<given-names>Shiqing</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1451107/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute of Intelligence Information Processing, Taizhou University</institution>, <addr-line>Taizhou, Zhejiang</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Faculty of Mechanical Engineering and Automation, Zhejiang Sci-Tech University</institution>, <addr-line>Hangzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by">
<p>Edited by: Kostas Karpouzis, Panteion University, Greece</p>
</fn>
<fn id="fn0002" fn-type="edited-by">
<p>Reviewed by: Erik Cambria, Nanyang Technological University, Singapore; Juan Sebastian Olier, Tilburg University, Netherlands</p>
</fn>
<corresp id="c001">&#x002A;Correspondence: Shiqing Zhang, <email>tzczsq@163.com</email></corresp>
<fn id="fn0003" fn-type="other">
<p>This article was submitted to Human-Media Interaction, a section of the journal Frontiers in Psychology</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>839619</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Zhao, Tang and Zhang.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Zhao, Tang and Zhang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Automatic personality trait recognition has attracted increasing interest in psychology, neuropsychology, and computer science, etc. Motivated by the great success of deep learning methods in various tasks, a variety of deep neural networks have increasingly been employed to learn high-level feature representations for automatic personality trait recognition. This paper systematically presents a comprehensive survey on existing personality trait recognition methods from a computational perspective. Initially, we provide available personality trait data sets in the literature. Then, we review the principles and recent advances of typical deep learning techniques, including deep belief networks (DBNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). Next, we describe the details of state-of-the-art personality trait recognition methods with specific focus on hand-crafted and deep learning-based feature extraction. These methods are analyzed and summarized in both single modality and multiple modalities, such as audio, visual, text, and physiological signals. Finally, we analyze the challenges and opportunities in this field and point out its future directions.</p>
</abstract>
<kwd-group>
<kwd>personality trait recognition</kwd>
<kwd>personality computing</kwd>
<kwd>deep learning</kwd>
<kwd>multimodal</kwd>
<kwd>survey</kwd>
</kwd-group>
<contract-num rid="cn2">LZ20F020002</contract-num>
<contract-num rid="cn2">LQ21F020002</contract-num>
<contract-num rid="cn2">61976149</contract-num>
<contract-sponsor id="cn1">Zhejiang Provincial National Science Foundation of China</contract-sponsor>
<contract-sponsor id="cn2">National Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="6"/>
<equation-count count="0"/>
<ref-count count="125"/>
<page-count count="18"/>
<word-count count="13602"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>In (<xref ref-type="bibr" rid="ref107">Vinciarelli and Mohammadi, 2014</xref>), the concept of personality can be defined as &#x201C;<italic>personality is a psychological construct aimed at explaining the wide variety of human behaviors in terms of a few, stable and measurable individual characteristics</italic>.&#x201D; In this case, personality can be characterized as a series of traits. The trait theory (<xref ref-type="bibr" rid="ref17">Costa and McCrae, 1998</xref>) aims to predict relatively stable measurable aspects in the people&#x2019;s daily lives on the basis of traits. It is used to measure human personality traits, that is, customary patterns of human behaviors, ideas, and emotions which are relatively kept steady over time. Some previous works explored the interaction between personality and computing by means of measuring the connection between traits and the used techniques (<xref ref-type="bibr" rid="ref35">Guadagno et al., 2008</xref>; <xref ref-type="bibr" rid="ref79">Qiu et al., 2012</xref>; <xref ref-type="bibr" rid="ref80">Quercia et al., 2012</xref>; <xref ref-type="bibr" rid="ref67">Liu et al., 2016</xref>; <xref ref-type="bibr" rid="ref53">Kim and Song, 2018</xref>; <xref ref-type="bibr" rid="ref69">Masuyama et al., 2018</xref>; <xref ref-type="bibr" rid="ref34">Goreis and Voracek, 2019</xref>; <xref ref-type="bibr" rid="ref62">Li et al., 2020a</xref>). The central idea behind these works is that users aim to externalize their personality by the way of using techniques. Accordingly, personality traits can be identified as predictive for users&#x2019; behaviors.</p>
<p>At present, various personality trait theories have been developed to categorize, interpret and understand human personality. The representative personality trait theories contain the Cattell Sixteen Personality Factor (16PF; <xref ref-type="bibr" rid="ref13">Cattell and Mead, 2008</xref>), the Hans Eysenck&#x2019;s psychoticism, extraversion and neuroticism (PEN; <xref ref-type="bibr" rid="ref25">Eysenck, 2012</xref>), Myers&#x2013;Briggs Type Indicator (MBTI; <xref ref-type="bibr" rid="ref29">Furnham and Differences, 1996</xref>), Big-Five (<xref ref-type="bibr" rid="ref70">McCrae and John, 1992</xref>), and so on. So far, the widely used measure for automatic personality trait recognition is the Big-Five personality traits. The Big-Five (<xref ref-type="bibr" rid="ref70">McCrae and John, 1992</xref>) model measures personality through five dipolar scales:</p>
<p>&#x201C;Extraversion&#x201D;: outgoing, energetic, talkative, active, assertive, etc.</p>
<p>&#x201C;Neuroticism&#x201D;: worrying, self-pitying, unstable, tense, anxious, etc.</p>
<p>&#x201C;Agreeableness&#x201D;: sympathetic, forgiving, generous, kind, appreciative, etc.</p>
<p>&#x201C;Conscientiousness&#x201D;: responsible, organized, reliable, efficient, planful, etc.</p>
<p>&#x201C;Openness&#x201D;: artistic, curious, imaginative, insightful, original, wide interests, etc.</p>
<p>In recent years, personality computing (<xref ref-type="bibr" rid="ref107">Vinciarelli and Mohammadi, 2014</xref>) has become a very active research subject that focuses on computational techniques related to human personality. It mainly addresses three fundamental problems: automatic personality trait recognition, perception, and synthesis. The first one aims at correctly identifying or predicting the actual (self-assessed) personality traits of human beings. This allows the construction of an apparent personality (or first impression) of an unacquainted individual. Automatic personality trait perception concentrates on analyzing the different subjective factors that affect the personality perception for a given individual. Automatic personality trait synthesis tries to realize the generation of artificial personalities through artificial agents and robots. This paper focuses on the first problem of personality computing, that is, automatic personality trait recognition, due to its potential applications to emotional and empathetic virtual agents in human&#x2013;computer interaction (HCI).</p>
<p>Most prior works focus on personality trait modeling and prediction from different cues, both behavioral and verbal. Therefore, automatic personality trait recognition takes into account multiple input modalities, such as audio, text, and visual cues. In 2015, the INTERSPEECH Speaker Trait Challenge (<xref ref-type="bibr" rid="ref87">Schuller et al., 2015</xref>) provided a unified test run for predicting the Big-Five personality traits, likability, and pathology of speakers, and meanwhile presented a performance comparison of computational models with the given data sets, and extracted features. In 2016, the well-known European Conference on Computer Vision (ECCV) released a benchmark open-domain personality data set, that is, Cha-Learn-2016, to organize a competition of personality recognition (<xref ref-type="bibr" rid="ref77">Ponce-L&#x00F3;pez et al., 2016</xref>).</p>
<p>Automatic personality trait recognition from social media contents has recently become a challenging issue and attracted much attention in the fields of artificial intelligence and computer vision, etc. So far, several surveys on personality trait recognition have been published in recent years. Specially, <xref ref-type="bibr" rid="ref107">Vinciarelli and Mohammadi (2014)</xref> provided the first review on personality computing, related to automatic personality trait recognition, perception, and synthesis. This review was organized from a more general point of view (personality computing). <xref ref-type="bibr" rid="ref50">Junior et al. (2019)</xref>, also presented a survey on vision-based personality trait analysis from visual data. This survey focused on the single visual modality. Moreover, these two surveys concentrate on classical methods, and recently emerged deep learning techniques (<xref ref-type="bibr" rid="ref46">Hinton et al., 2006</xref>) have seldom been reviewed. Very recently, <xref ref-type="bibr" rid="ref73">Mehta et al. (2020b)</xref> presented a brief review deep learning-based personality trait detection. Nevertheless, they did not provide a summary on personality trait databases and technical details on deep learning techniques. Therefore, this paper gives a comprehensive review for personality trait recognition from a computational perspective. In particular, we focus on reviewing the recent advances of existing both single and multimodal personality trait recognition methods between 2012 and 2022 with specific emphasis on hand-crafted and deep learning-based feature extraction. We aim at providing a newcomer to this field, a summary of the systematic framework, and main skills for deep personality trait recognition. We also examine state-of-the-art methods that have not been mentioned in prior surveys.</p>
<p>In this survey, we have searched the published literature between January 2012, and February 2022 through Scholar.google, ScienceDirect, IEEEXplore, ACM, Springer, PubMed, and Web of Science, on the basis of the following keywords: &#x201C;personality trait recognition,&#x201D; &#x201C;personality computing,&#x201D; &#x201C;deep learning,&#x201D; &#x201C;deep belief networks,&#x201D; &#x201C;convolutional neural networks,&#x201D; &#x201C;recurrent neural networks,&#x201D; &#x201C;long short-term memory,&#x201D; &#x201C;audio,&#x201D; &#x201C;visual,&#x201D; &#x201C;text,&#x201D; &#x201C;physiological signals,&#x201D; &#x201C;bimodal,&#x201D; &#x201C;trimodal,&#x201D; and &#x201C;multimodal.&#x201D; There is no any language restriction for the searching process. We designed and conducted this systematic survey by complying with the PRISMA statement (<xref ref-type="bibr" rid="ref85">Sarkis-Onofre et al., 2021</xref>) in an effort to improve the reporting of systematic reviews. Eligibility criteria of this survey contain the suitable depictions of different hand-crafted and deep learning-based feature extraction methods for personality trait recognition in both single modality and multiple modalities.</p>
<p>It is noted that a basic personality trait recognition system generally consists of two key parts: feature extraction and personality trait classification or prediction. Feature extraction can be divided into hand-crafted and deep learning-based methods. For personality trait classification or prediction, the common classifiers/regressors, such as Support Vector Machines (SVM) and linear regressors, are usually used. In this survey, we focus on the advances of feature extraction algorithms ranging from 2012 to 2022 in a basic personality trait recognition system. <xref rid="fig1" ref-type="fig">Figure 1</xref> shows the evolution of personality trait recognition with feature extraction algorithms and databases.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption><p>The evolution of personality trait recognition with feature extraction algorithms and databases. From 2012 to 2022, feature extraction algorithms have changed from hand-crafted to deep learning. Meanwhile, the developed databases have evolved from single modality (audio or visual) to multiple modalities (audio, visual, text, etc.).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g001.tif"/>
</fig>
<p>In this work, our contributions can be summarized as follows:</p>
<list list-type="order">
<list-item>
<p>We provide an up-to-date literature survey on deep personality trait analysis from a perspective of both single modality and multiple modalities. In particular, this work focuses on a systematical single and multimodal analysis of human personality. To the best of our knowledge, this is the first attempt to present a comprehensive review covering both single and multimodal personality trait analysis related to hand-crafted and deep learning-based feature extraction algorithms in this field.</p>
</list-item>
<list-item>
<p>We summarize existing personality trait data sets and review the typical deep learning techniques and its recent variants. We present the significant advances in single modality personality trait recognition related to audio, visual, text, etc., and multimodal personality trait recognition related to bimodal and trimodal modalities.</p>
</list-item>
<list-item>
<p>We analyze and discuss the challenges and opportunities faced to personality trait recognition and point out future directions in this field.</p>
</list-item>
</list>
<p>The remainder of this paper is organized as follows. Section &#x201C;Personality Trait Databases&#x201D; describes the available personality trait data sets. Several typical deep learning techniques and its recent variants are reviewed in detail in Section &#x201C;Review of Deep Learning Techniques.&#x201D; Section &#x201C;Review of Single Modality Personality Trait Recognition Techniques&#x201D; introduces the related techniques of single modality personality trait recognition. Section &#x201C;Multimodal Fusion for Personality Trait Recognition&#x201D; provides the details of multimodal fusion for personality trait recognition. Section &#x201C;Challenges and Opportunities&#x201D; discusses the challenges and opportunities in this field. Finally, the conclusions are given in Section &#x201C;Conclusion.&#x201D;</p>
</sec>
<sec id="sec2">
<title>Personality Trait Databases</title>
<p>To evaluate the performance of different methods, a variety of personality trait data sets, as shown in <xref rid="tab1" ref-type="table">Table 1</xref>, are collected for automatic personality trait recognition. These representative data sets are described as follows.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption><p>Comparisons of representative personality trait recognition databases.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Data set</th>
<th align="center" valign="top">Year</th>
<th align="left" valign="top">Brief description</th>
<th align="left" valign="top">Central issues</th>
<th align="left" valign="top">Labels</th>
<th align="left" valign="top">Modality</th>
<th align="left" valign="top">Environment</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">SSPNet (<xref ref-type="bibr" rid="ref74">Mohammadi and Vinciarelli, 2012</xref>)</td>
<td align="center" valign="top">2012</td>
<td align="left" valign="top">640 audio clips from 322 speakers</td>
<td align="left" valign="top">Personality trait assessment from speech</td>
<td align="left" valign="top">BFI-10 personality assessment questionnaire, Big-Five impressions</td>
<td align="left" valign="top">Audio</td>
<td align="left" valign="top">Uncontrolled</td>
</tr>
<tr>
<td align="left" valign="top">SEMAINE (<xref ref-type="bibr" rid="ref71">McKeown et al., 2012</xref>)</td>
<td align="center" valign="top">2012</td>
<td align="left" valign="top">959 conversations from 150 participants</td>
<td align="left" valign="top">Face-to-face conversations with sensitive artificial listener agents</td>
<td align="left" valign="top">Five affective dimensions and 27 associated categories</td>
<td align="left" valign="top">Audio-visual</td>
<td align="left" valign="top">Controlled</td>
</tr>
<tr>
<td align="left" valign="top">YouTube Vlogs (<xref ref-type="bibr" rid="ref10">Biel and Gatica-Perez, 2012</xref>)</td>
<td align="center" valign="top">2012</td>
<td align="left" valign="top">2,269 videos from 469 different vloggers</td>
<td align="left" valign="top">Conversational vlogs and apparent personality trait analysis</td>
<td align="left" valign="top">Big-Five impressions</td>
<td align="left" valign="top">Audio-visual</td>
<td align="left" valign="top">Uncontrolled</td>
</tr>
<tr>
<td align="left" valign="top">ELEA (<xref ref-type="bibr" rid="ref84">Sanchez-Cortes et al., 2013</xref>)</td>
<td align="center" valign="top">2013</td>
<td align="left" valign="top">40 meeting sessions with about 10&#x2009;h of recordings (148 participants)</td>
<td align="left" valign="top">Small group interactions and emergent leadership</td>
<td align="left" valign="top">Big-Five impressions</td>
<td align="left" valign="top">Audio-visual</td>
<td align="left" valign="top">Controlled</td>
</tr>
<tr>
<td align="left" valign="top">ChaLearn First Impression V1 (<xref ref-type="bibr" rid="ref77">Ponce-L&#x00F3;pez et al., 2016</xref>)</td>
<td align="center" valign="top">2016</td>
<td align="left" valign="top">10,000 videos from 2,762 YouTube users</td>
<td align="left" valign="top">Apparent personality trait analysis</td>
<td align="left" valign="top">Big-Five impressions</td>
<td align="left" valign="top">Audio-visual</td>
<td align="left" valign="top">Uncontrolled</td>
</tr>
<tr>
<td align="left" valign="top">ChaLearn First Impression V2 (<xref ref-type="bibr" rid="ref23">Escalante et al., 2017</xref>)</td>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">An extended version of [5], including the newly added hirability impressions and audio transcripts</td>
<td align="left" valign="top">Apparent personality trait and hirability impressions</td>
<td align="left" valign="top">Big-Five impressions, job interview variable, and transcripts</td>
<td align="left" valign="top">Multimodal</td>
<td align="left" valign="top">Uncontrolled</td>
</tr>
<tr>
<td align="left" valign="top">MHHRI (<xref ref-type="bibr" rid="ref14">Celiktutan et al., 2017</xref>)</td>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">12 interaction sessions (about 4&#x2009;h) from 18 participants</td>
<td align="left" valign="top">Personality and engagement during HHI and HCI</td>
<td align="left" valign="top">Self/acquaintance assessed Big-Five, and engagement</td>
<td align="left" valign="top">Multimodal</td>
<td align="left" valign="top">Controlled</td>
</tr>
<tr>
<td align="left" valign="top">UDIVA (<xref ref-type="bibr" rid="ref75">Palmero et al., 2021</xref>)</td>
<td align="center" valign="top">2021</td>
<td align="left" valign="top">188 dyadic sessions (90.5&#x2009;h) from 147 participants</td>
<td align="left" valign="top">Context-aware personality inference in dyadic scenarios</td>
<td align="left" valign="top">Big-Five scores, sociodemographics, mood, fatigue, relationship type</td>
<td align="left" valign="top">Multimodal</td>
<td align="left" valign="top">Controlled</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="sec3">
<title>SSPNet</title>
<p>The SSPNet (<xref ref-type="bibr" rid="ref74">Mohammadi and Vinciarelli, 2012</xref>) speaker personality corpus is the biggest up-to-date data set for the assessment of personality traits in speech signals. It contains 640 audio clips from 322 speakers with a sampling rate of 8&#x2009;kHz. These audio clips are randomly derived from the French news in Switzerland. Most of them are 10&#x2009;s long. In addition, 11 judges are invited to annotate every clip by means of filling out the BFI-10 personality evaluation questionnaire (<xref ref-type="bibr" rid="ref81">Rammstedt and John, 2007</xref>). A score is calculated for every Big-Five personality trait on the basis of the questionnaire. The judges are not familiar with French and thus could not be affected by linguistic cues.</p>
</sec>
<sec id="sec4">
<title>Emergent Leader</title>
<p>The Emergent LEAder (ELEA; <xref ref-type="bibr" rid="ref84">Sanchez-Cortes et al., 2013</xref>) data set comprises of 40 meeting sessions associated with about 10&#x2009;h of recordings. It consists of 28 four-person conferences as well as 12 three-person conferences in newly constructed groups, in which previously unacquainted persons are included. The mean age for 148 participants (48 women and 100 men) is 25.4&#x2009;years old. All the participants at the ELEA conferences are required to take part in a winter survival task, but are not assigned any special role. Audio recordings are collected by using a microphone, and the audio sampling rate is 16&#x2009;kHz. Video recordings are gathered with two setup settings: a static setting with six cameras, and a portable setting with two webcams. The video frame rates for these two settings are separately 25 fps and 30 fps, respectively.</p>
</sec>
<sec id="sec5">
<title>SEMAINE</title>
<p>The SEMAINE (<xref ref-type="bibr" rid="ref71">McKeown et al., 2012</xref>) audio-visual data set contains 150 participants (57 men and 93 women) with a mean age of 32.8&#x2009;years old. These participants are undergraduate and postgraduate students from eight different nations. The representative conversation duration for Solid SAL and Semi-automatic SAL is approximately 30&#x2009;min. A total of 959 conversations with individual SAL characters are collected, each of which lasts about 5&#x2009;min, although there are large individual differences. The Automatic SAL conversation lasts almost 1&#x2009;h with eight-character interaction per 3&#x2009;min. Participants interacted with both versions of the system and finished psychometric measures at an interval of 10&#x2013;15&#x2009;min.</p>
</sec>
<sec id="sec6">
<title>YouTube Vlogs</title>
<p>The YouTube Vlogs (<xref ref-type="bibr" rid="ref10">Biel and Gatica-Perez, 2012</xref>) data set comprises of 2,269 videos with a total of 150&#x2009;h. These videos, ranging from 1 to 6&#x2009;min in length, come from 469 different vloggers. It contains video metadata and viewer comments gathered in 2009 (<xref ref-type="bibr" rid="ref9">Biel and Gatica-Perez, 2010</xref>). The video samples are collected with keywords like &#x201C;vlogs&#x201D; and &#x201C;vlogging.&#x201D; Meanwhile, the recording setting is that a participant is talking to a camera displaying the participant&#x2019;s head and shoulder. The recording contents contain various topics, such as personal video blogs, film, product comments, and so on.</p>
</sec>
<sec id="sec7">
<title>ChaLearn First Impression V1-V2</title>
<p>The ChaLearn First Impression data set has been developed into two versions: the ChaLearn First Impression V1 (<xref ref-type="bibr" rid="ref77">Ponce-L&#x00F3;pez et al., 2016</xref>), and the ChaLearn First Impression V2 (<xref ref-type="bibr" rid="ref23">Escalante et al., 2017</xref>): The ChaLearn First Impression V1 contains 10,000 short video clips, collected from about 2,762 YouTube high-definition videos of persons who are facing and speaking to the camera in English. Each video has a resolution of 1,280&#x2009;&#x00D7;&#x2009;720, and a length of 15&#x2009;s. The involved persons have different genders, ages, nationalities, and races. In this case, the task of predicting apparent personality traits becomes more difficult and challenging. The ChaLearn First Impression V2 (<xref ref-type="bibr" rid="ref23">Escalante et al., 2017</xref>) is an extension of the ChaLearn First Impression V1 (<xref ref-type="bibr" rid="ref77">Ponce-L&#x00F3;pez et al., 2016</xref>). In this data set, the new variable of &#x201C;job interview&#x201D; is added for prediction. The manual transcriptions associated with the corresponding audio in videos are also provided.</p>
</sec>
<sec id="sec8">
<title>Multimodal Human&#x2013;Human&#x2013;Robot Interactions</title>
<p>The multimodal human&#x2013;human&#x2013;robot interactions (MHHRI; <xref ref-type="bibr" rid="ref14">Celiktutan et al., 2017</xref>) data set contains 18 participants (nine men and nine women), most of whom are graduate students and researchers. It includes 12 interaction conversations (about 4&#x2009;h). Each interactive conversation has 10&#x2013;15&#x2009;min and is recorded with several sensors. For recording first-person videos, two liquid image egocentric cameras are located on the participants&#x2019; forehead. For RGB-D recordings, two static Kinect depth sensors are placed opposite to each other for capturing the entire scene. For audio recordings, the microphone in the egocentric cameras is used. Additionally, participants are required to wear a Q-sensor with Affectiva for recording physiological signals, such as electrodermal activity (EDA).</p>
</sec>
<sec id="sec9">
<title>Understanding Dyadic Interactions From Video and Audio Signals</title>
<p>The understanding dyadic interactions from video and audio signals (UDIVA; <xref ref-type="bibr" rid="ref75">Palmero et al., 2021</xref>) data set, comprises of 90.5&#x2009;h of non-scripted face-to-face dyadic interactions between 147 participants (81 men and 66 women) from 4 to 84&#x2009;years old. Participants were distributed into 188 dyadic sessions. This data set was recorded by using multiple audio-visual and physiological sensors. The raw audio frame rate is 44.1&#x2009;kHz. Video recordings are collected from 6 HD tripod-mounted cameras with a resolution of 1,280&#x2009;&#x00D7;&#x2009;720. They adopted questionnaire based assessments, including sociodemographic, self- and peer-reported personality, internal state, and relationship profiling from participants.</p>
<p>From <xref rid="tab1" ref-type="table">Table 1</xref>, we can see that the representative personality trait recognition databases are developed from the single modality (audio), bimodality (audio-visual), and multiple modalities. For obtaining the ground-truth scores of personality traits on these databases, personality questionnaires are presented to the users for annotations. Nevertheless, such subjective annotations with personality questionnaires may affect the reliability of trained models on these databases.</p>
</sec>
</sec>
<sec id="sec10">
<title>Review of Deep Learning Techniques</title>
<p>In recent years, deep learning techniques have been an active research subject and obtained promising performance in various applications, such as object detection and classification, speech processing, natural language processing, and so on (<xref ref-type="bibr" rid="ref119">Yu and Deng, 2010</xref>; <xref ref-type="bibr" rid="ref58">LeCun et al., 2015</xref>; <xref ref-type="bibr" rid="ref86">Schmidhuber, 2015</xref>; <xref ref-type="bibr" rid="ref124">Zhao et al., 2015</xref>). In essence, deep learning methods aim to achieve high-level abstract representations by means of hierarchical architectures of multiple non-linear transformations. After implementing feature extraction with deep learning techniques, the Softmax (Sigmoid) function is usually for classification or prediction. In this section, we briefly review several representative deep learning methods and its recent variants, which can be potentially used for personality trait analysis.</p>
<sec id="sec11">
<title>Deep Belief Networks</title>
<p>Deep belief networks (DBNs; <xref ref-type="bibr" rid="ref46">Hinton et al., 2006</xref>) developed by Hinton et al. in 2006, are a generative model that aim to capture a high-order hierarchical feature representation of input data. The conventional DBN is a multilayered deep architecture, which is built by a sequence of superimposed restricted Boltzmann machines (RBMs; <xref ref-type="bibr" rid="ref26">Freund and Haussler, 1994</xref>). A RBM is a two-layer generative stochastic neural network consisting of a visual layer and a hidden layer. These two layers in a RBM constitute a bipartite graph without any lateral connection. Training a DBN needs two-stage steps: pretraining and fine-tuning. Pretraining is realized by means of an efficient layer-by-layer greedy learning strategy (<xref ref-type="bibr" rid="ref6">Bengio et al., 2007</xref>) in an unsupervised manner. During the pretraining process, a contrastive divergence (<xref ref-type="bibr" rid="ref45">Hinton, 2002</xref>; CD) algorithm is adopted to train RBMs in a DBN to enable the optimization of the weights and bias of DBN models. Then, fine-tuning is performed to update the network parameters by using the back propagation (BP) algorithm.</p>
<p>Several improved versions of DBNs are developed in recent years. <xref ref-type="bibr" rid="ref60">Lee et al. (2009)</xref>, proposed a convolutional deep belief network (CDBN) for full-sized images, in which multiple max-pooling based convolutional RBMs were stacked on the top of one another. <xref ref-type="bibr" rid="ref110">Wang et al. (2018)</xref> presented a growing DBN with transfer learning (TL-GDBN). TL-GDBN aimed to grow its network structure by means of transferring the learned feature representations from the original structure to the newly developed structure. Then, a partial least squares regression (PLSR)-based fine-tuning was implemented to update the network parameters instead of the traditional BP algorithm.</p>
</sec>
<sec id="sec12">
<title>Convolutional Neural Networks</title>
<p>Convolutional neural networks (CNNs) were originally proposed by <xref ref-type="bibr" rid="ref59">LeCun et al. (1998)</xref> in 1998, and initially developed as an advanced version of deep CNNs, such as AlexNet (<xref ref-type="bibr" rid="ref55">Krizhevsky et al., 2012</xref>) in 2012. The basic structure of CNNs comprises of convolutional layers, pooling layers, as well as fully connected (FC) layers. CNNs usually have multiple convolutional and pooling layers, in which pooling layers are frequently followed by convolutional layers. The convolutional layer adopts a number of learnable filters to perform convolution operation on the whole input image, thereby yielding the corresponding activation feature maps. The pooling layer is employed to reduce the spatial size of produced feature maps by using non-linear down-sampling methods for translation invariance. Two well-known used pooling strategies are average pooling and max-pooling. The FC layer, in which all neurons are fully connected, is often placed at the end of the CNN network. It is used to activate the previous layer for producing the final feature representations and classification.</p>
<p>In recent years, several advanced versions of deep CNNs have been presented in various applications. The representative deep CNN models include AlexNet (<xref ref-type="bibr" rid="ref55">Krizhevsky et al., 2012</xref>), VGGNet (<xref ref-type="bibr" rid="ref91">Simonyan and Zisserman, 2014</xref>), GoogleNet (<xref ref-type="bibr" rid="ref98">Szegedy et al., 2015</xref>), ResNet (<xref ref-type="bibr" rid="ref42">He et al., 2016</xref>), DenseNet (<xref ref-type="bibr" rid="ref48">Huang et al., 2017</xref>), and so on. In particular, DenseNet (<xref ref-type="bibr" rid="ref48">Huang et al., 2017</xref>), in which each layer is connected to each other layer in a feed-forward manner, has been proved that it beats most deep models on objection recognition tasks with less network parameters. <xref rid="tab2" ref-type="table">Table 2</xref> presents the comparisons of the configurations and characteristics of these typical deep CNNs, as described below.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption><p>Comparisons of deep CNN models and its configurations.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th align="center" valign="top">AlexNet</th>
<th align="center" valign="top">VGGNet</th>
<th align="center" valign="top">GoolgeNet</th>
<th align="center" valign="top">ResNet</th>
<th align="center" valign="top">DenseNet</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Year</td>
<td align="center" valign="top">2012</td>
<td align="center" valign="top">2015</td>
<td align="center" valign="top">2015</td>
<td align="center" valign="top">2016</td>
<td align="center" valign="top">2017</td>
</tr>
<tr>
<td align="left" valign="top">layers (Conv.&#x2009;+&#x2009;FC)</td>
<td align="center" valign="top">5&#x2009;+&#x2009;3</td>
<td align="center" valign="top">19&#x2009;+&#x2009;3</td>
<td align="center" valign="top">21&#x2009;+&#x2009;1</td>
<td align="center" valign="top">151&#x2009;+&#x2009;1</td>
<td align="center" valign="top">264&#x2009;+&#x2009;1</td>
</tr>
<tr>
<td align="left" valign="top">Conv. kernel</td>
<td align="center" valign="top">11,5,3</td>
<td align="center" valign="top">3</td>
<td align="center" valign="top">7,1,3,5</td>
<td align="center" valign="top">7,1,3,5</td>
<td align="center" valign="top">7,1,3</td>
</tr>
<tr>
<td align="left" valign="top">Dropout</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
</tr>
<tr>
<td align="left" valign="top">Inception</td>
<td align="center" valign="top">&#x00D7;</td>
<td align="center" valign="top">&#x00D7;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x00D7;</td>
<td align="center" valign="top">&#x00D7;</td>
</tr>
<tr>
<td align="left" valign="top">DA</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
</tr>
<tr>
<td align="left" valign="top">BN</td>
<td align="center" valign="top">&#x00D7;</td>
<td align="center" valign="top">&#x00D7;</td>
<td align="center" valign="top">&#x00D7;</td>
<td align="center" valign="top">&#x221A;</td>
<td align="center" valign="top">&#x221A;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Conv., convolution; DA, data augmentation; BN, batch normalization. The number of layers is the used maximum in deep models.</p>
</table-wrap-foot>
</table-wrap>
<p>Compared with the above-mentioned deep CNNs processing 2D images, the recently developed 3D-CNNs (<xref ref-type="bibr" rid="ref103">Tran et al., 2015</xref>) aim to learn temporal-spatio feature representations by using 3D convolution operations on large-scale video data sets. Some improved versions of 3D-CNNs are also recently proposed to reduce the computation complexity of 3D convolutions. <xref ref-type="bibr" rid="ref118">Yang et al. (2019)</xref> provided an asymmetric 3D-CNN on the basis of the proposed MicroNets, in which a set of local 3D convolutional networks were adopted so as to incorporate multiscale 3D convolution branches. <xref ref-type="bibr" rid="ref56">Kumawat and Raman (2019)</xref> proposed a LP-3DCNN in which a rectified local phase volume (ReLPV) block was used to replace the conventional 3D convolutional block. <xref ref-type="bibr" rid="ref15">Chen et al. (2020)</xref> developed a frequency domain compact 3D-CNN model, in which they utilized a set of learned optimal transformation with few network parameters to implement 3D convolution operations by converting the time domain into the frequency domain.</p>
</sec>
<sec id="sec13">
<title>Recurrent Neural Networks</title>
<p>Recurrent neural networks (RNNs; <xref ref-type="bibr" rid="ref22">Elman, 1990</xref>) are a single feed-forward neural network for capturing temporal information, and thus suitable to deal with sequence data. RNNs contain recurrent edges connecting adjacent time steps, thereby providing the concept of time in this model. In addition, RNNs share the same network parameters across all time steps. For training RNNs, the traditional back propagation through time (BPTT; <xref ref-type="bibr" rid="ref112">Werbos, 1990</xref>) was usually adopted.</p>
<p>Long short-term memory (LSTM; <xref ref-type="bibr" rid="ref47">Hochreiter and Schmidhuber, 1997</xref>), proposed by Hochreiter and Schmidhuber in 1997, is a relatively new recurrent network architecture, which is combined with a suitable gradient-based learning fashion. Specially, LSTMs aim to alleviate the gradient vanishing and exploding problems produced during the procedure of training RNNs. There are three types of gates in a LSTM cell unit: input gate, forget gate, and output gate. Input gate is used to control how much of the current input data is flowing into the memory unit of the network. Forget gate, as a key component of the LSTM cell unit, is used for controlling which information to keep and which to forget, and somehow avoiding the gradient loss and explosion problems. Output gate controls the effect of the memory cell on the current output value. On the basis of these three special gates, LSTMs have an ability of modeling long-term dependencies of sequence data, such as video sequences.</p>
<p>In recent years, a variant of LSTMs called gated recurrent unit (GRU; <xref ref-type="bibr" rid="ref16">Chung et al., 2014</xref>), was developed by Chung et al. in 2014. GRU makes every recurrent unit to adaptively model long-term dependencies of different time scales. Different from the LSTM unit, GRU does not have a separate memory cell inside the unit. In addition, combining CNNs with LSTMs becomes a research trend. In particular, <xref ref-type="bibr" rid="ref125">Zhao et al. (2019)</xref> proposed a Bayesian graph based a convolution LSTM for identifying skeleton-based actions. <xref ref-type="bibr" rid="ref123">Zhang et al. (2019)</xref> developed a multiscale deep convolutional LSTM for speech emotion classification.</p>
</sec>
</sec>
<sec id="sec14">
<title>Review of Single Modality Personality Trait Recognition Techniques</title>
<p>Automatic personality trait recognition aims to adopt computer science techniques to realize the modeling of personality trait recognition problems in cognitive science. It is one of the most important research subjects in the field of personality computing (<xref ref-type="bibr" rid="ref107">Vinciarelli and Mohammadi, 2014</xref>; <xref ref-type="bibr" rid="ref51">Junior et al., 2018</xref>). According to the types of input data, automatic personality trait recognition can be divided into: single modality and multiple modalities. In particular, it contains the single audio or visual personality trait recognition, and multimodal personality trait recognition, integrating multiple modal behavior data, such as audio, visual, and text information.</p>
<sec id="sec15">
<title>Audio-Based Personality Trait Recognition</title>
<p><xref rid="tab3" ref-type="table">Table 3</xref> presents a brief summary of existing literature related to audio-based personality trait recognition.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption><p>A brief summary of audio-based on personality trait recognition.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Year</th>
<th align="left" valign="top">Authors</th>
<th align="left" valign="top">Feature descriptions</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">2012</td>
<td align="left" valign="top">Mohammadi et al.</td>
<td align="left" valign="top">Pitch, formants, energy, and speaking rate</td>
</tr>
<tr>
<td align="left" valign="top">2016</td>
<td align="left" valign="top">An et al.</td>
<td align="left" valign="top">Interspeech-2013 ComParE feature set</td>
</tr>
<tr>
<td align="left" valign="top">2017</td>
<td align="left" valign="top">Su et al.</td>
<td align="left" valign="top">Wavelet-based multiresolution analysis and CNNs for feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2019</td>
<td align="left" valign="top">Hayat et al.</td>
<td align="left" valign="top">Fine-tuning the pretrained AudioSet for feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2020</td>
<td align="left" valign="top">Carbonneau et al.</td>
<td align="left" valign="top">Learning feature dictionary from the extracted patches in speech spectrograms</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The early-used audio features for automatic personality trait recognition are hand-crafted low-level descriptive (LLD) features, such as prosody (intensity, pitch), voice quality (formants), spectral features (Mel Frequency Cepstrum Coefficients, MFCCs), and so on. Specially, <xref ref-type="bibr" rid="ref74">Mohammadi and Vinciarelli (2012)</xref> utilized the LLD features, such as pitch, formants, energy, and speaking rate to detect personality traits in audio clips with less than 10&#x2009;s. They adopted Logistic Regression to identify whether an audio clip exceeded the average score for each of the Big-five personality traits. In (<xref ref-type="bibr" rid="ref2">An et al., 2016</xref>), 6,373 acoustic&#x2013;prosodic features like the Interspeech-2013 ComParE feature set (<xref ref-type="bibr" rid="ref88">Schuller et al., 2013</xref>) were extracted as an input of the SVM classifier for identifying the Big-Five personality traits. In (<xref ref-type="bibr" rid="ref12">Carbonneau et al., 2020</xref>), the authors learned a discriminating feature dictionary from the extracted patches in the speech spectrograms, followed by the SVM classifier for the classification of the Big-Five personality traits.</p>
<p>The recently used audio features for automatic personality trait recognition are deep audio features extracted by deep learning techniques. <xref ref-type="bibr" rid="ref92">Su et al. (2017)</xref> proposed to employ wavelet-based multiresolution analysis and CNNs for personality trait perception from speech signals. <xref rid="fig2" ref-type="fig">Figure 2</xref> presents the details of the used CNN scheme. The wavelet transform was adopted to decompose the original speech signals at different levels of resolution. Then, based on the extracted prosodic acoustic features, CNNs were leveraged to produce the profiles of the Big-Five Inventory-10 (BFI-10) for a quantitative measure, followed by artificial neural networks (ANNs) for personality trait recognition. <xref ref-type="bibr" rid="ref41">Hayat et al. (2019)</xref> fine-tuned a pretrained CNN model called AudioSet to learn an audio feature representation for predicting the Big-five personality trait scores of a speaker. They showed the advantages of CNN-based learned features over hand-crafted features.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption><p>The used CNN scheme for personality trait perception from speech signals (<xref ref-type="bibr" rid="ref92">Su et al., 2017</xref>).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g002.tif"/>
</fig>
</sec>
<sec id="sec16">
<title>Visual-Based Personality Trait Recognition</title>
<p>According to the type of vision-based input data, visual-based personality trait recognition can be categorized into two types: static images and dynamic video sequences. Visual feature extraction is the key step related to the input static images and dynamic video sequences for personality trait recognition. <xref rid="tab4" ref-type="table">Table 4</xref> provides a brief summary of existing literature related to visual-based (static images, and dynamic video sequences) personality trait recognition.</p>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption><p>A brief summary of visual-based on personality trait recognition.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Visual type</th>
<th align="center" valign="top">Year</th>
<th align="left" valign="top">Authors</th>
<th align="left" valign="top">Feature descriptions</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top" rowspan="6">Static images</td>
<td align="center" valign="top">2015</td>
<td align="left" valign="top">Guntuku et al.</td>
<td align="left" valign="top">LBP, GIST, aesthetic features</td>
</tr>
<tr>
<td align="center" valign="top">2016</td>
<td align="left" valign="top">Yan et al.</td>
<td align="left" valign="top">HOG, SIFT, LBP</td>
</tr>
<tr>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">Zhang et al.</td>
<td align="left" valign="top">Fine-tuning the pretrained VGG-face model for facial feature extraction</td>
</tr>
<tr>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">Segalin et al.</td>
<td align="left" valign="top">Fine-tuning the pretrained AlexNet and VGG-16 for aesthetic attributes</td>
</tr>
<tr>
<td align="center" valign="top">2020</td>
<td align="left" valign="top">Rodr&#x00ED;guez et al.</td>
<td align="left" valign="top">Trained a ResNet-50 to derive personality representations from the posted images</td>
</tr>
<tr>
<td align="center" valign="top">2021</td>
<td align="left" valign="top">Fu et al.</td>
<td align="left" valign="top">An improved ASM model for facial feature extraction, followed by a DBN</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="6">Dynamic video sequences</td>
<td align="center" valign="top">2012</td>
<td align="left" valign="top">Biel et al.</td>
<td align="left" valign="top">Facial activity statistics based on frame-by-frame estimation</td>
</tr>
<tr>
<td align="center" valign="top">2013</td>
<td align="left" valign="top">Aran et al.</td>
<td align="left" valign="top">Statistical information derived from the weighted motion energy images</td>
</tr>
<tr>
<td align="center" valign="top">2014</td>
<td align="left" valign="top">Teijeiro-Mosquera, et al.</td>
<td align="left" valign="top">Four sets of behavioral cues, such as statistic, THR, HMM, and WTA cues</td>
</tr>
<tr>
<td align="center" valign="top">2016</td>
<td align="left" valign="top">G&#x00FC;rpinar et al.</td>
<td align="left" valign="top">Fine-tuning the pretrained VGG-19 to extract deep facial and scene features</td>
</tr>
<tr>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">Ventura et al.</td>
<td align="left" valign="top">An extension of DAN for facial feature extraction in videos</td>
</tr>
<tr>
<td align="center" valign="top">2019</td>
<td align="left" valign="top">Beyan et al.</td>
<td align="left" valign="top">Deep visual activity-based features derived from key-dynamic images in videos</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="sec17">
<title>Static Images</title>
<p>As far as static image-based personality trait recognition is concerned, researchers have found that a facial image presents most of meaningful descriptive cues for personality trait recognition (<xref ref-type="bibr" rid="ref113">Willis and Todorov, 2006</xref>). Hence, the extracted visual features involve in the analysis of facial features for personality trait prediction. In (<xref ref-type="bibr" rid="ref38">Guntuku et al., 2015</xref>), the authors proposed to leverage several low-level features of facial images, such as color histograms, local binary patterns (LBP), global descriptor (GIST), and aesthetic features, to train the SVM classifier for detecting mid-level clues (gender, age). Then, they predicted the Big-five personality traits of users in self-portrait images with the lasso regressor. <xref ref-type="bibr" rid="ref117">Yan et al. (2016)</xref> investigated the connection between facial appearance and personality impression in the manner of trustworthy. They obtained middle-level cues through clustering methods from different low-level features, such as histogram of oriented gradients (HOG), scale-invariant feature transform (SIFT), LBP, and so on. Then, a SVM classifier was used to exploit the connection between facial appearance and personality impression.</p>
<p>In recent years, CNNs were also widely used for facial feature extraction on static image-based personality trait recognition tasks. <xref ref-type="bibr" rid="ref121">Zhang et al. (2017)</xref> presented an end-to-end CNN structure <italic>via</italic> fine-tuning a pretrained VGG-face model for feature learning so as to predict personality traits and intelligence jointly. They aimed to explore whether self-reported personality traits and intelligence can be jointly measured from facial images. <xref ref-type="bibr" rid="ref89">Segalin et al. (2017)</xref> explored the linking the Big-Five personality traits and preferred images in the Flickr social network through image understanding and a deep CNN framework. In particular, they fine-tuned the pretrained AlexNet and VGG-16 modal to capture the aesthetic attributes of the images characterizing the personality traits associated with those images. They changed the last layer of the AlexNet and VGG-16 model to adapt them to a binary classification problem. Experiments results showed that the characterization of each image can be locked within the CNN layers, thereby discovering entangled attributes, such as the aesthetic and semantic information for generalizing the patterns that identify a personality trait. <xref ref-type="bibr" rid="ref83">Rodr&#x00ED;guez et al. (2020)</xref> presented a personality trait analysis in social networks by using a weakly supervised learning method of shared images. They trained a ResNet-50 network to derive personality representations from the posted images in social networks, so as to infer whether the personality scores from the posted images are correlated to those scores obtained from text. For predicting personality traits, the images without manually labeling were used for training the ResNet-50 model. Experiment results indicate that people&#x2019;s personality is not only related to text, but also with the image content. <xref ref-type="bibr" rid="ref27">Fu and Zhang (2021)</xref> provided a personality trait recognition method by using active shape model (ASM) localization and DBNs. They employed an improved ASM model to extract facial features, followed by a DBN which was used to train and classify the students&#x2019; four personality traits.</p>
</sec>
<sec id="sec18">
<title>Dynamic Video Sequences</title>
<p>Dynamic video sequences consist of a series of video image frames, thereby providing temporal information and scene dynamics. This brings about certain useful and complementary cues for personality trait analysis (<xref ref-type="bibr" rid="ref50">Junior et al., 2019</xref>).</p>
<p>In (<xref ref-type="bibr" rid="ref11">Biel et al., 2012</xref>), the authors investigated the connection between facial expressions and personality of vloggers in conversation videos (vlogs) from a subset of existing YouTube vlog data set (<xref ref-type="bibr" rid="ref9">Biel and Gatica-Perez, 2010</xref>). They employed a computer expression recognition toolbox to identify the categories of facial expressions of vloggers. They finally adopted a SVM classifier to predict personality traits in conjunction with facial activity statistics on the basis of frame-by-frame estimation. The results indicate that extraversion has the highest utilization of activity cues. This is consistent with previous findings (<xref ref-type="bibr" rid="ref8">Biel et al., 2011</xref>; <xref ref-type="bibr" rid="ref10">Biel and Gatica-Perez, 2012</xref>). <xref ref-type="bibr" rid="ref3">Aran and Gatica-Perez (2013)</xref> adopted the social media contents from conversational videos for analyzing the specific trait of extraversion. To address this issue, they integrated the ridge regression with a SVM classifier on the basis of statistical information derived from the weighted motion energy images. In (<xref ref-type="bibr" rid="ref101">Teijeiro-Mosquera et al., 2014</xref>), the relations between facial expressions and personality impressions were investigated as an extended version of the used method (<xref ref-type="bibr" rid="ref11">Biel et al., 2012</xref>). To characterize face statistics, they derived four sets of behavioral cues, such as statistic-based cues, Threshold (THR) cues, Hidden Markov Models (HMM) cues, and Winner Takes All (WTA) cues. Their research indicates that when multiple facial expression clues are significantly correlated with a certain number of the Big-Five traits, they could only obviously predict the particular trait of extraversion.</p>
<p>In consideration of the tremendous progress in the areas of deep learning, CNNs and LSTMs are widely for personality trait analysis from dynamic video sequences. <xref ref-type="bibr" rid="ref40">G&#x00FC;rp&#x0131;nar et al. (2016)</xref> fine-tuned a pretrained VGG-19 network to extract deep facial and scene feature representations, as shown in <xref rid="fig3" ref-type="fig">Figure 3</xref>. Then, they were merged and fed into a kernel extreme learning machine (ELM) regressor for first impression estimation. <xref ref-type="bibr" rid="ref105">Ventura et al. (2017)</xref> adopted an extension of Descriptor Aggregation Networks (DAN) to investigate why CNN models performed well in automatically predicting first impressions. They used class activation maps (CAM) for visualization and provided a possible interpretation on understanding why CNN models succeeded in learning discriminative facial features related to personality traits of users. <xref rid="fig4" ref-type="fig">Figure 4</xref> shows the used CAM to interpret the CNN models in learning facial features. Experimental results indicate that: (1) face presents most of discriminative information for the inference of personality traits, (2) the internal representations of CNNs primarily focus on crucial facial regions including eyes, nose, and mouth, (3) some action units (AUs) provide a partial impact on the inference of facial traits. <xref ref-type="bibr" rid="ref7">Beyan et al. (2019)</xref> aimed to perceive personality traits by means of using deep visual activity (VA)-based features derived only from key-dynamic images in videos. In order to determine key-dynamic images in videos, they employed three key steps: construction of multiple dynamic images, long-term VA learning with CNN&#x2009;+&#x2009;LSTM, and spatio-temporal saliency detection.</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption><p>The flowchart of personality trait prediction by using deep facial and scene feature representations (<xref ref-type="bibr" rid="ref40">G&#x00FC;rp&#x0131;nar et al., 2016</xref>).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g003.tif"/>
</fig>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption><p>The used class activation maps (CAM) to interpret the CNN models in learning facial features (<xref ref-type="bibr" rid="ref105">Ventura et al., 2017</xref>).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g004.tif"/>
</fig>
</sec>
</sec>
<sec id="sec19">
<title>Other Modality-Based Personality Trait Recognition</title>
<p>In addition to the above-mentioned audio and visual modality, there are other single modalities, such as text, and physiological signals, etc., which can be applied for personality trait recognition. <xref rid="tab5" ref-type="table">Table 5</xref> gives a brief summary of personality trait recognition based on text and physiological signals.</p>
<table-wrap position="float" id="tab5">
<label>Table 5</label>
<caption><p>A brief summary of text and physiological-based personality trait recognition.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Input type</th>
<th align="center" valign="top">Year</th>
<th align="left" valign="top">Authors</th>
<th align="left" valign="top">Feature descriptions</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="8">
</td>
<td align="center" valign="top">2013</td>
<td align="left" valign="top">Bazelli et al.</td>
<td align="left" valign="top">Predicting the personality traits of Stack Overflow authors with LIWC</td>
</tr>
<tr>
<td align="center" valign="top">2016</td>
<td align="left" valign="top">Golbeck et al.</td>
<td align="left" valign="top">The Receptiviti API providing personality score predictions with LIWC</td>
</tr>
<tr>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">Majumder et al.</td>
<td align="left" valign="top">A CNN with injection of the document-level Mairesse features</td>
</tr>
<tr>
<td align="center" valign="top">2017</td>
<td align="left" valign="top">Hernandez et al.</td>
<td align="left" valign="top">RNNs and its variants, such as GRU, LSTM, and Bi-LSTM for text features</td>
</tr>
<tr>
<td align="center" valign="top">2018</td>
<td align="left" valign="top">Xue et al.</td>
<td align="left" valign="top">A hierarchical deep neural network for learning deep semantic features</td>
</tr>
<tr>
<td align="center" valign="top">2018</td>
<td align="left" valign="top">Sun et al.</td>
<td align="left" valign="top">A 2CLSTM integrating a Bi-LSTM with a CNN for feature extraction</td>
</tr>
<tr>
<td align="center" valign="top">2020</td>
<td align="left" valign="top">Mehta et al.</td>
<td align="left" valign="top">Psycholinguistic features were combined with BERT embeddings</td>
</tr>
<tr>
<td align="center" valign="top">2021</td>
<td align="left" valign="top">Ren et al.</td>
<td align="left" valign="top">A BERT for text feature extraction, followed by GRU, LSTM, and CNN</td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">Physiological signals</td>
<td align="center" valign="top">2014</td>
<td align="left" valign="top">Wache et al.</td>
<td align="left" valign="top">The measurements of ECG, EEG, GSR</td>
</tr>
<tr>
<td align="center" valign="top">2018</td>
<td align="left" valign="top">Subramanian et al.</td>
<td align="left" valign="top">The measurements of ECG, EEG, GSR and facial activity data</td>
</tr>
<tr>
<td align="center" valign="top">2020</td>
<td align="left" valign="top">Taib et al.</td>
<td align="left" valign="top">Adopting eye-tracking and skin conductivity sensors for capturing their physiological responses</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="sec20">
<title>Text-Based Personality Trait Recognition</title>
<p>The text modality can effectively display traces of the user&#x2019;s personality (<xref ref-type="bibr" rid="ref31">Golbeck et al., 2011</xref>). One of the early-used features from text is the popular linguistic inquiry and word count (LIWC; <xref ref-type="bibr" rid="ref76">Pennebaker et al., 2001</xref>), which is often used to extract lexical features. LIWC divides the words into a variety of psychologically buckets, such as function words (e.g., conjunctions and pronouns), affective words (e.g., amazing and cried), and so on. Then, the used frequency of different categories of words is counted in each bucket in purpose of predicting the personality traits of the writer. <xref ref-type="bibr" rid="ref5">Bazelli et al. (2013)</xref> predicted the personality traits of Stack Overflow authors by means of analyzing the community&#x2019;s questions and answers on the basis of LIWC. The recently developed Receptiviti API (<xref ref-type="bibr" rid="ref30">Golbeck, 2016</xref>) is a popular tool using LIWC for personality trait prediction from text in psychology studies.</p>
<p>In recent years, several deep learning techniques have been employed for text-based personality trait recognition. <xref ref-type="bibr" rid="ref68">Majumder et al. (2017)</xref> proposed a deep CNN method for document-level personality prediction from text, as depicted in <xref rid="fig5" ref-type="fig">Figure 5</xref>. The used CNN model consists of seven layers and aims to extract the monogram, bigram, and trigram features from text. <xref ref-type="bibr" rid="ref43">Hernandez and Scott (2017)</xref> aimed at learning temporal dependencies among sentences by feeding the input text data into simple RNNs and its variants, such as GRU, LSTM, and Bi-LSTM. It was found that LSTM achieved better performance compared to RNN, GRU, and Bi-LSTM on MBTI personality trait recognition tasks. <xref ref-type="bibr" rid="ref115">Xue et al. (2018)</xref> adopted a hierarchical deep neural network, including an attentive recurrent CNN structure and a variant of the inception structure, to learn deep semantic features from text posts of online social networks for the Big-five personality trait recognition. <xref ref-type="bibr" rid="ref96">Sun et al. (2018)</xref> presented a model called 2CLSTM, integrating a Bi-LSTM with a CNN, for predicting user&#x2019;s personality on the basis of structures of texts. <xref ref-type="bibr" rid="ref72">Mehta et al. (2020a)</xref> proposed a deep learning-based model in which conventional psycholinguistic features were combined with language model embeddings like Bidirectional Encoder Representation From Transformers (BERT; <xref ref-type="bibr" rid="ref20">Devlin et al., 2018</xref>) for personality trait prediction. <xref ref-type="bibr" rid="ref82">Ren et al. (2021)</xref> presented a multilabel personality prediction model <italic>via</italic> deep learning, which integrated semantic and emotional features from social media texts. They conducted sentence-level extraction of both semantic and emotion features by means of a BERT model and a SentiNet5 (<xref ref-type="bibr" rid="ref106">Vilares et al., 2018</xref>) dictionary model, respectively. Then, they fed these features into GRU, LSTM, and CNN for further feature extraction and classification. It was found that BERT+CNN performed best on MBTI and Big-Five personality trait classification tasks.</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption><p>The flowchart of CNN-based document-level personality prediction from text (<xref ref-type="bibr" rid="ref68">Majumder et al., 2017</xref>).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="sec21">
<title>Physiological Signal-Based Personality Trait Recognition</title>
<p>Since the user&#x2019;s physiological responses to affective stimuli are highly correlated with personality traits, numerous works have tried to carry out physiological signal-based personality recognition. <xref ref-type="bibr" rid="ref108">Wache (2014)</xref> investigated emotional states and personality traits on the basis of physiological responses to affective video clips. When watching 36 affective video clips, they utilized the measurements of Electrocardiogram (ECG), Galvanic Skin Response (GSR), Electroencephalogram (EEG) to characterize their Big-Five personality traits. Moreover, they also provided a multimodal database for implicit personality and affect classification by means of commercial physiological sensors (<xref ref-type="bibr" rid="ref94">Subramanian et al., 2016</xref>). <xref ref-type="bibr" rid="ref99">Taib et al. (2020)</xref> proposed a method of personality detection from physiological responses to affective image and video stimuli. They adopted eye-tracking and skin conductivity sensors for capturing their physiological responses.</p>
</sec>
</sec>
<sec id="sec22">
<title>Multimodal Fusion for Personality Trait Recognition</title>
<p>For multimodal fusion on personality trait recognition tasks, there are generally three types: feature-level fusion, decision-level fusion, and model-level fusion (<xref ref-type="bibr" rid="ref120">Zeng et al., 2008</xref>; <xref ref-type="bibr" rid="ref4">Atrey et al., 2010</xref>).</p>
<p>Feature-level fusion aims to directly concatenate the extracted features from multimodal modalities, into one feature set. Therefore, feature-level fusion is also called early fusion (EF). As the simplest way of implementing feature integration, feature-level fusion has relatively low cost and complexity. Moreover, it considers the correlation between modalities. However, integrating different time scale and metric level of features from multimodal modalities will significantly increase the dimensionality of the concatenated feature vector, resulting in the difficulty of training models.</p>
<p>In decision-level fusion, each modality is firstly modeled independently, and then these obtained results from single-modality are combined to produce final results by using a certain number of decision fusion rules. Decision-level fusion is thus called late fusion (LF). The commonly used decision fusion rules include &#x201C;Majority vote,&#x201D; &#x201C;Max,&#x201D; &#x201C;Sum,&#x201D; &#x201C;Min,&#x201D; &#x201C;Average,&#x201D; &#x201C;Product,&#x201D; etc. (<xref ref-type="bibr" rid="ref97">Sun et al., 2015</xref>). Since decision-level fusion considers different modalities as mutually independent, it can easily deal with asynchrony among modalities, resulting in the scalability with the number of modalities. Nevertheless, it fails to make use of the correlation between modalities at feature-level.</p>
<p>Model-level fusion aims to separately model each modality while taking into account the correlation between modalities. Therefore, it can consider the inter-correlation among different modalities and loose the demand of timing synchronization of these modalities.</p>
<p><xref rid="tab6" ref-type="table">Table 6</xref> shows a brief summary of multimodal fusion for personality trait recognition. In the following, we present an analysis of these multimodal fusion methods from two aspects: bimodal and trimodal modalities for personality trait recognition.</p>
<table-wrap position="float" id="tab6">
<label>Table 6</label>
<caption><p>A brief summary of multimodal fusion for personality trait recognition.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Year</th>
<th align="left" valign="top">Authors</th>
<th align="left" valign="top">Modalities</th>
<th align="left" valign="top">Fusion methods</th>
<th align="left" valign="top">Feature descriptions</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">2016</td>
<td align="left" valign="top">G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk et al.</td>
<td align="left" valign="top">Audio, visual</td>
<td align="left" valign="top">Feature-level</td>
<td align="left" valign="top">An deep residual network for audio and visual feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2016, 2017</td>
<td align="left" valign="top">Zhang et al.</td>
<td align="left" valign="top">Audio, visual</td>
<td align="left" valign="top">Decision-level</td>
<td align="left" valign="top">A DBR method integrating audio and visual (scene and face) modality</td>
</tr>
<tr>
<td align="left" valign="top">2016</td>
<td align="left" valign="top">G&#x00FC;rpinar et al.</td>
<td align="left" valign="top">Audio, visual</td>
<td align="left" valign="top">Score-level</td>
<td align="left" valign="top">Fine-tuning a pretrained VGG model to derive facial emotion and ambient features. The INTERSPEECH-2009 for audio feature set</td>
</tr>
<tr>
<td align="left" valign="top">2016</td>
<td align="left" valign="top">Subramaniam et al.</td>
<td align="left" valign="top">Audio, visual</td>
<td align="left" valign="top">Feature-level</td>
<td align="left" valign="top">A volumetric (3D) convolution network for visual feature extraction. The statistics of zero-crossing rate, energy, MFCCs for audio features</td>
</tr>
<tr>
<td align="left" valign="top">2021</td>
<td align="left" valign="top">Curto et al.</td>
<td align="left" valign="top">Audio, visual</td>
<td align="left" valign="top">Model-level</td>
<td align="left" valign="top">The pretrained VGGish for audio feature extraction, and the pretrained R(2&#x2009;+&#x2009;1)D for video feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2016</td>
<td align="left" valign="top">Xianyu et al.</td>
<td align="left" valign="top">Text, visual</td>
<td align="left" valign="top">Model-level</td>
<td align="left" valign="top">A heterogeneity entropy (HE) neural network (HENN) consisting of HE-DBN, HE-AE and common DBN for common feature representations among text, image and behavior statistical modalities</td>
</tr>
<tr>
<td align="left" valign="top">2019</td>
<td align="left" valign="top">Principi et al.</td>
<td align="left" valign="top">Audio, visual</td>
<td align="left" valign="top">Model-level/Feature-level</td>
<td align="left" valign="top">A multimodal deep learning model (ResNet-50 for visual modality and 14-layer 1D CNN for audio modality) for feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2020</td>
<td align="left" valign="top">Li et al.</td>
<td align="left" valign="top">Audio, visual, text</td>
<td align="left" valign="top">Feature-level</td>
<td align="left" valign="top">A deep CR-Net to predict the multimodal Big-Five personality traits based on video, audio, and text cues</td>
</tr>
<tr>
<td align="left" valign="top">2017</td>
<td align="left" valign="top">G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk et al.</td>
<td align="left" valign="top">Audio, visual, text</td>
<td align="left" valign="top">Feature-level</td>
<td align="left" valign="top">A deep residual networks for audio-visual feature extraction. A bag-of-words and a skip-thought vector model for text feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2017, 2018</td>
<td align="left" valign="top">Gorbova et al.</td>
<td align="left" valign="top">Audio, visual, text</td>
<td align="left" valign="top">Decision-level</td>
<td align="left" valign="top">Acoustic LLD features (MFCCs, ZCR, speaking rate), facial action unit features, as well as negative and positive word scores</td>
</tr>
<tr>
<td align="left" valign="top">2018</td>
<td align="left" valign="top">Kampman et al.</td>
<td align="left" valign="top">Audio, visual, text</td>
<td align="left" valign="top">Decision-level/Model-level</td>
<td align="left" valign="top">An trimodal deep CNN method for audio, visual, text feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2020</td>
<td align="left" valign="top">Escalante et al.</td>
<td align="left" valign="top">Audio, visual, text</td>
<td align="left" valign="top">Feature-level</td>
<td align="left" valign="top">A bag-of-words model and a skip-thought vector model for text feature extraction, and the ResNet18 model for audio-visual feature extraction</td>
</tr>
<tr>
<td align="left" valign="top">2022</td>
<td align="left" valign="top">Suman et al.</td>
<td align="left" valign="top">Audio, visual, text</td>
<td align="left" valign="top">Feature-level/Decision-level</td>
<td align="left" valign="top">A MTCNN and ResNet for facial and ambient feature extraction, respectively. A VGGish model for audio feature extraction and an <italic>n</italic>-gram CNN model for text feature extraction</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="sec23">
<title>Bimodal Modalities Based Personality Trait Recognition</title>
<p>For bimodal modalities based personality trait recognition, the widely used one is audio-visual modality. In order to effectively extract audio-visual feature representations of short video sequences, numerical studies have been conducted for audio-visual personality trait recognition.</p>
<p><xref ref-type="bibr" rid="ref37">G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk et al. (2016)</xref> developed an end-to-end audio-visual deep residual network for audio-visual apparent personality trait recognition. In detail, the audio data and visual data were firstly extracted from the video clip. Then, the whole audio data were fed into an audio deep residual network for feature learning. Note that the activities of the penultimate layer in the audio deep residual network were temporally pooled. Similarly, the whole visual data were fed into a visual deep residual network with a frame at a time. The activities of the penultimate layer in the visual deep residual network were spatiotemporally pooled. Finally, the pooled activities of the audio and visual stream were concatenated at feature-level as an input of a fully connected layer for personality trait prediction.</p>
<p>Zhang et al., developed a deep bimodal regression (DBR) method so as to capture rich information from the audio and visual modality in videos (<xref ref-type="bibr" rid="ref122">Zhang et al., 2016</xref>; <xref ref-type="bibr" rid="ref111">Wei et al., 2017</xref>). <xref rid="fig6" ref-type="fig">Figure 6</xref> shows the flowchart of the proposed DBR method audio-visual personality trait prediction. In particular, for visual feature extraction, they modified the traditional CNNs by means of discarding the fully connected layers. Additionally, they merged the average and max pooled features of the last convolutional layer into a whole feature vector, followed by the standard L2 normalization. For audio feature extraction, they extracted the logfbank features from the original audio utterances of videos. Then, they trained the linear regressor to produce the Big-Five trait values. To integrate the complementary cues from the audio-visual modality, they fused these predicted regression scores at decision-level.</p>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption><p>The flowchart of the proposed DBR method for audio-visual personality trait prediction (<xref ref-type="bibr" rid="ref111">Wei et al., 2017</xref>).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g006.tif"/>
</fig>
<p><xref ref-type="bibr" rid="ref39">G&#x00FC;rpinar et al. (2016)</xref> proposed a multimodal fusion method of audio and visual (scene and face) features for personality trait analysis. They fine-tuned a pretrained VGG model to derive facial emotion and ambient information from images. They also extracted local Gabor binary patterns from three orthogonal planes (LGBP-TOP) video descriptor as video features. The typical acoustic features, such as the INTERSPEECH-2009, 2010, 2012, and 2013 feature set in computational paralinguistics challenges, were employed. The kernel ELM was adopted for personality trait prediction on audio and visual (scene and face) modalities. Finally, a score-level method was leveraged to fuse the results of different modalities.</p>
<p><xref ref-type="bibr" rid="ref93">Subramaniam et al. (2016)</xref> employed two end-to-end deep learning models for audio-visual first impression analysis. They used a volumetric (3D) convolution network for visual feature extraction from face aligned images. For audio feature extraction, they obtained the statistics, such as mean and standard deviation of hand-crafted features like zero-crossing rate, energy, MFCCs, etc. Then, they concatenated the extracted audio and visual features at feature-level, followed by a multimodal LSTM network of temporal modeling for final personality trait prediction tasks.</p>
<p><xref ref-type="bibr" rid="ref114">Xianyu et al. (2016)</xref> proposed an unsupervised cross-modal feature learning method, called heterogeneity entropy (HE) neural network (HENN), for multimodal personality trait prediction. The proposed HENN consists of HE-DBN, HE-AE, and common DBN and is used to learn common feature representations among text, image, and behavior statistical modalities, and then map them into the user&#x2019;s personality. The input of HENN is hand-crafted features. In particular, a bag of textual word (BoTW; <xref ref-type="bibr" rid="ref61">Li et al., 2016</xref>) model was used to extract the text feature vector. Based on the extracted scale-invariant feature transform (SIFT; <xref ref-type="bibr" rid="ref18">Cruz-Mota et al., 2012</xref>) features of each image, a bag of visual word model was used to produce visual image features. The time series information related to sharing numbers and comment numbers in both text and image modalities were employed to compute behavior statistical parameters. These hand-crafted features were individually fed into three HE-DBNs for initial feature learning, and then HE-AE and common DBN were separately adopted to fuse these features produced with HE-DBNs at model-level for final Big-Five personality prediction.</p>
<p><xref ref-type="bibr" rid="ref78">Principi et al. (2019)</xref> developed a multimodal deep learning model combining the raw visual with audio streams to conduct the Big-Five personality trait prediction. For each video sample, different task-specific deep models, related to individual factor, such as facial expressions, attractiveness, age, gender, and ethnicity, were leveraged to estimate per-frame attribute. Then, these estimated results were concatenated at feature-level to produce a video-level attribute prediction by spatio-temporal aggregation methods. For visual feature extraction, they adopted a ResNet-50 network pretrained on the ImageNet data to produce high-level feature representations on each video frame. For audio feature extraction, a 14-layer 1D CNN like the ResNet-18 was used. They fused these modalities in two steps. First, they employed a FC layer for model-level fusion to learn the joint feature representations of the concatenated video-level attribute predictions. This model-level fusion step was also used to reduce the dimensionality of the concatenated video-level attribute predictions. Second, they combined such learned joint video-level attribute predictions with the extracted audio and visual features at feature-level, to perform final the Big-Five personality trait prediction.</p>
<p><xref ref-type="bibr" rid="ref19">Curto et al. (2021)</xref> developed the Dyadformer for modeling individual and interpersonal audio-visual features in dyadic interactions for personality trait prediction. The Dyadformer was a multimodal multisubject Transformer framework consisting of a set of attention encoder modules (self, cross-modal, and cross-subject) with Transformer layers. They employed the pretrained VGGish (<xref ref-type="bibr" rid="ref44">Hershey et al., 2017</xref>) model to produce a 128-dimensional embedding for each audio chunk. They leveraged the pretrained R(2&#x2009;+&#x2009;1)D (<xref ref-type="bibr" rid="ref104">Tran et al., 2018</xref>) model to generate a 512-dimensional embedding for each video chunk. They used cross-modal and cross-subject attentions for multimodal Transformer fusion in model-level.</p>
</sec>
<sec id="sec24">
<title>Trimodal Modalities Based Personality Trait Recognition</title>
<p><xref ref-type="bibr" rid="ref64">Li et al. (2020b)</xref> presented a deep classification&#x2013;regression network (CR-Net) to predict the multimodal Big-Five personality traits based on video, audio, and text cues and further applied to the job interview recommendation. For the visual input, they extracted the global scene cues and local face cues by using the ResNet-34 network. Considering audio-text inner correlations, they concatenated the extracted acoustic LLD and text-based skip-thought vectors at feature-level as inputs of the ResNet-34 network for audio-text feature learning. Finally, they merged all extracted features from visual, audio, and text modalities at feature-level and fed them into the CR-Net network to analyze the multimodal Big-Five personality traits.</p>
<p><xref ref-type="bibr" rid="ref36">G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk et al. (2017)</xref> presented a method of multimodal first impression analysis integrating audio, visual, and text (language) modalities, based on deep residual networks. They adopted two similar 17-layer deep residual networks for extracting audio-visual features. The used 17-layer deep residual networks consist of one convolutional layer and eight residual blocks of two convolutional layers. The pooled activities of audio-visual networks were concatenated as an input of a fully connected layer so as to learn the joint audio-visual feature representations. For text feature extraction, they utilized two language models, including a bag-of-words model and a skip-thought vector model, to produce the annotations as a function of the language data. Both of the language models contain an embedding layer, followed by a linear layer. Finally, they combined the extracted features from audio, visual, and text at feature-level for the multimodal Big-five personality trait analysis and job interview recommendation.</p>
<p><xref ref-type="bibr" rid="ref33">Gorbova et al. (2017</xref>, <xref ref-type="bibr" rid="ref32">2018)</xref> provided an automatic personality screening method on the basis of visual, audio, and text (lexical) cues from short video clips for predicting the Big-five personality traits. The extracted hand-crafted features contained acoustic LLD features (MFCCs, ZCR, speaking rate, etc.), facial action unit features, as well as negative and positive word scores. This system adopted the weighted average strategy to fuse the final obtained results from three modalities at decision-level. <xref rid="fig7" ref-type="fig">Figure 7</xref> shows the flowchart of integrating audio, vision, and language for first impression personality analysis (<xref ref-type="bibr" rid="ref32">Gorbova et al., 2018</xref>). In <xref rid="fig7" ref-type="fig">Figure 7</xref>, after extracted audio, visual, and lexical features from input video, three separate LSTM cells were used for modeling long dependency. Then, the hidden features in LSTMs were processed by a linear regressor. Finally, the obtained results were fed to an output layer for the Big-five personality trait analysis.</p>
<fig position="float" id="fig7">
<label>Figure 7</label>
<caption><p>The flowchart of Integrating audio, vision and language for first-Impression personality analysis (<xref ref-type="bibr" rid="ref32">Gorbova et al., 2018</xref>).</p></caption>
<graphic xlink:href="fpsyg-13-839619-g007.tif"/>
</fig>
<p><xref ref-type="bibr" rid="ref52">Kampman et al. (2018)</xref> presented an end-to-end trimodal deep learning architecture for predicting the Big-Five personality traits by means of integrating audio, visual, and text modalities. For audio channel, the raw audio waveform and its energy components with squared amplitude were fed into a CNN network with four convolutional layers and a global average pooling layer for audio feature extraction. For visual channel, based on a random frame image of a video, they fine-tuned the pretrained VGG-16 model for video feature extraction. For text channel, they adopted &#x201C;Word2vec&#x201D; word embedding from transcriptions as an input of a CNN network for text feature extraction. In this text CNN network, three different convolutional windows corresponding to three, four, and five words over the sentence were used. Finally, they fused audio, visual, and text modalities at both decision-level and model-level. For decision-level fusion, a voting scheme was used. For model-level fusion, by means of concatenating the output of FC layers of each CNN, they added another two FC layers on top to learn shared feature representations of input trimodal data.</p>
<p>Escalante et al. explored the explainability in first impressions analysis from video sequences at the first time. They provided a baseline method of integrating audio, visual, and text (audio transcripts) information (<xref ref-type="bibr" rid="ref24">Escalante et al., 2020</xref>). They used a variant of original 18-layer deep residual networks (ResNet-18) for audio and visual feature extraction, respectively. The feature-level fusion was adopted after the global average pooling layers of the audio-visual ResNet-18 models <italic>via</italic> concatenation of their obtained latent features. For text modality, two language models, such as a skip-thought vector model and a bag-of-words model, were employed for text feature extraction. Finally, a concatenation of audio, visual, text-based latent features was implemented at feature-level for multimodal first-impression analysis.</p>
<p><xref ref-type="bibr" rid="ref95">Suman et al. (2022)</xref> developed a deep learning-based multimodal personality prediction system integrating audio, visual, and text modalities. They extracted facial and ambient features from the visual modality by using Multi-task Cascaded Convolutional Neural Networks (MTCNN; <xref ref-type="bibr" rid="ref49">Jiang et al., 2018</xref>) and ResNet, individually. They extracted the audio features by using a VGGish (<xref ref-type="bibr" rid="ref44">Hershey et al., 2017</xref>) model, and the text features by using an <italic>n</italic>-gram CNN model, respectively. These extracted audio, visual, and text features were fed into a fully connected layer followed by a sigmoid function for the final personality trait prediction. It was concluded that the concatenation of audio, visual, and text features in feature-level fusion showed comparable performance with the averaging method in decision-level fusion.</p>
</sec>
</sec>
<sec id="sec25">
<title>Challenges and Opportunities</title>
<p>To date, although there are a number of literature related to multimodal personality trait prediction, showing its certain advance, a few challenges still exist in this area. In the following, we discuss these challenges and opportunities, and point out potential research directions in future.</p>
<sec id="sec26">
<title>Personality Trait Recognition Data Sets</title>
<p>Although researchers have developed a variety of relevant data sets for personality trait recognition, as shown in <xref rid="tab1" ref-type="table">Table 1</xref>, these data sets are relatively small. To date, the most popular multimodal data sets, such as the ChaLearn First Impression V1 (<xref ref-type="bibr" rid="ref77">Ponce-L&#x00F3;pez et al., 2016</xref>), and its enhanced version V2 (<xref ref-type="bibr" rid="ref23">Escalante et al., 2017</xref>), consist of 10,000 short video clips. Such data sets are definitely smaller, compared with the well-known ImageNet data set with a total of 14 million images used for training deep models. Considering that automatic personality trait recognition is a data-driven task associated with a deep neural network, a large amount of training data is required for training sufficiently deep models. Therefore, one major challenge for deep multimodal personality trait recognition is the lack of a large amount of training data on the basis of both quantity and quality.</p>
<p>In addition, owing to the difference of data collecting and annotating environment, data bias and inconsistent annotations usually exist among these different data sets. Most researchers conventionally verify the performance of their proposed methods within a specific data set, resulting in promising results. Such trained models based on intra-data set protocols commonly lack generalizability on unseen test data. Therefore, it is interesting to investigate the performance of multimodal personality trait recognition methods in cross-data set environment. To address this issue, deep domain adaption methods (<xref ref-type="bibr" rid="ref109">Wang et al., 2020</xref>; <xref ref-type="bibr" rid="ref57">Kurmi et al., 2021</xref>; <xref ref-type="bibr" rid="ref90">Shao and Zhong, 2021</xref>) may be an alternative. Note that the display of personality traits and the traits themself can be considered as context-dependent. This will also give a considerable challenge for the training models on personality trait recognition tasks.</p>
</sec>
<sec id="sec27">
<title>Integrating More Modalities</title>
<p>For multimodal personality trait recognition, bimodal modalities like audio-visual, or trimodal modalities like audio, visual, and text, are usually employed. Note that the user&#x2019;s physiological responses to affective stimuli are highly correlated with personality traits. However, few researchers explore the performance of integrating physiological signals with other modalities for multimodal personality trait recognition. This is because so far these are few multimodal personality trait recognition data sets, which incorporate physiological signals with other modalities. Hence, one may challenge is how to combine physiological signals and other modalities, such as audio, visual, and text clues, based on the corresponding developed multimodal data sets.</p>
<p>Besides, other behavior signals, such as head and body pose information, which is related to personality trait clues (<xref ref-type="bibr" rid="ref1">Alameda-Pineda et al., 2015</xref>), may present complementary information to further enhance the robustness of multimodal personality trait recognition. It is thus a promising research direction to integrate head and body clues with existing modalities, such as audio, visual, and text clues for multimodal personality trait recognition.</p>
</sec>
<sec id="sec28">
<title>Limitations of Deep Learning Techniques</title>
<p>So far, a variety of representative deep leaning methods have been successfully applied to learn high-level feature representations for automatic personality trait recognition. Moreover, these deep learning methods usually beat other methods adopting hand-crafted features. Nevertheless, these used deep learning techniques have a tremendous amount of network parameters, resulting in its large computation complexity. In this case, for real-time application sceneries it is often difficult to implement fast automatic personality trait prediction with these complicated deep models. To alleviate this issue, a deep model compression (<xref ref-type="bibr" rid="ref65">Liang et al., 2021a</xref>; <xref ref-type="bibr" rid="ref100">Tartaglione et al., 2021</xref>) may present a possible solution.</p>
<p>Although deep learning has become a state-of-the-art technique in term of the performance measure on various feature learning tasks, the black box problem still exists. In particular, it is unknown that what exactly are various internal representations learned by multiple hidden layers of a deep model. Owing to its multilayer non-linear structure, deep learning techniques are usually criticized to be non-transparent, and their prediction results are often not traceable by human beings. To alleviate this problem, directly visualizing the learned features has become the widely used way of understanding deep models (<xref ref-type="bibr" rid="ref24">Escalante et al., 2020</xref>). Nevertheless, such visualizing way does not really present the related theories to explain what exactly this algorithm is doing. Therefore, it is an important research direction to explore the explainability and interpretability of deep learning techniques (<xref ref-type="bibr" rid="ref102">Tjoa and Guan, 2020</xref>; <xref ref-type="bibr" rid="ref54">Krichmar et al., 2021</xref>; <xref ref-type="bibr" rid="ref66">Liang et al., 2021b</xref>; <xref ref-type="bibr" rid="ref116">Yan et al., 2021</xref>) from a theoretical perspective for automatic personality trait recognition.</p>
</sec>
<sec id="sec29">
<title>Investigating Other Trait Theories</title>
<p>It is noted that most researchers focus on personality trait analysis <italic>via</italic> the Big-Five personality model. This is because almost all of the current data sets were developed based on the Big-Five personality measures, as shown in <xref rid="tab1" ref-type="table">Table 1</xref>. However, very few literature concentrate on other personality measures, such as the MBTI, PEN, and 16PF, due to the lacking data resources. In particular, the MBTI personality measure, as the most popular administered personality test throughout the world, shows more difficulty in prediction than the Big-Five model (<xref ref-type="bibr" rid="ref29">Furnham and Differences, 1996</xref>; <xref ref-type="bibr" rid="ref28">Furnham, 2020</xref>). Therefore, it is an open issue to investigate the effect of other trait theories on personality trait prediction on the basis of correspondingly constructed data sets.</p>
</sec>
</sec>
<sec id="sec30" sec-type="conclusions">
<title>Conclusion</title>
<p>Due to the strong feature learning ability of deep learning, multiple recent works using deep learning have been developed for personality trait recognition associated with promising performance. This paper attempts to provide a comprehensive survey of existing personality trait recognition methods with specific focus on hand-crafted and deep learning-based feature extraction. These methods systematically review the topic from the single modality and multiple modalities. We also highlight numerous issues for future challenges and opportunities. Apparently, personality trait recognition is a very broad and multidisciplinary research issue. This survey only focuses on reviewing existing personality trait recognition methods from a computational perspective and does not take psychological studies into account on personality trait recognition.</p>
<p>In future, it is interesting to explore the application of personality trait recognition techniques to personality-aware recommendation systems (<xref ref-type="bibr" rid="ref21">Dhelim et al., 2021</xref>). In addition, since personality traits are usually strongly connected with emotions, it is an important direction to investigate a CNN-based multitask learning framework for emotion and personality detection (<xref ref-type="bibr" rid="ref63">Li et al., 2021</xref>).</p>
</sec>
<sec id="sec31">
<title>Author Contributions</title>
<p>XZ contributed to the writing and drafted this article. ZT contributed to the collection and analysis of existing literature. SZ contributed to the conception and design of this work and revised this article. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec32" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by Zhejiang Provincial National Science Foundation of China and National Science Foundation of China (NSFC) under Grant Nos. LZ20F020002, LQ21F020002, and 61976149.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec34" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alameda-Pineda</surname> <given-names>X.</given-names></name> <name><surname>Staiano</surname> <given-names>J.</given-names></name> <name><surname>Subramanian</surname> <given-names>R.</given-names></name> <name><surname>Batrinca</surname> <given-names>L.</given-names></name> <name><surname>Ricci</surname> <given-names>E.</given-names></name> <name><surname>Lepri</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Salsa: a novel dataset for multimodal group behavior analysis</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>38</volume>, <fpage>1707</fpage>&#x2013;<lpage>1720</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2496269</pub-id>, PMID: <pub-id pub-id-type="pmid">26540677</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="other"><person-group person-group-type="author"><name><surname>An</surname> <given-names>G.</given-names></name> <name><surname>Levitan</surname> <given-names>S. I.</given-names></name> <name><surname>Levitan</surname> <given-names>R.</given-names></name> <name><surname>Rosenberg</surname> <given-names>A.</given-names></name> <name><surname>Levine</surname> <given-names>M.</given-names></name> <name><surname>Hirschberg</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Automatically classifying self-rated personality scores from speech</article-title>&#x201D;, in <source>INTERSPEECH</source>, <fpage>1412</fpage>&#x2013;<lpage>1416</lpage>.</citation></ref>
<ref id="ref3"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Aran</surname> <given-names>O.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). &#x201C;<article-title>Cross-domain personality prediction: from video blogs to small group meetings</article-title>,&#x201D; in <source>Proceedings of the 15th ACM on International Conference on Multimodal Interaction</source>, <fpage>127</fpage>&#x2013;<lpage>130</lpage>.</citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atrey</surname> <given-names>P. K.</given-names></name> <name><surname>Hossain</surname> <given-names>M. A.</given-names></name> <name><surname>El Saddik</surname> <given-names>A.</given-names></name> <name><surname>Kankanhalli</surname> <given-names>M. S.</given-names></name></person-group> (<year>2010</year>). <article-title>Multimodal fusion for multimedia analysis: a survey</article-title>. <source>Multimedia Systems</source> <volume>16</volume>, <fpage>345</fpage>&#x2013;<lpage>379</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s00530-010-0182-0</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Bazelli</surname> <given-names>B.</given-names></name> <name><surname>Hindle</surname> <given-names>A.</given-names></name> <name><surname>Stroulia</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). &#x201C;<article-title>On the personality traits of StackOverflow users</article-title>,&#x201D; in <source>2013 IEEE International Conference on Software Maintenance</source>, <fpage>460</fpage>&#x2013;<lpage>463</lpage>.</citation></ref>
<ref id="ref6"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Lamblin</surname> <given-names>P.</given-names></name> <name><surname>Popovici</surname> <given-names>D.</given-names></name> <name><surname>Larochelle</surname> <given-names>H.</given-names></name></person-group> (<year>2007</year>). &#x201C;<article-title>Greedy layer-wise training of deep networks</article-title>,&#x201D; in <source>Advances in Neural Information Processing Systems</source>, <fpage>153</fpage>&#x2013;<lpage>160</lpage>.</citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beyan</surname> <given-names>C.</given-names></name> <name><surname>Zunino</surname> <given-names>A.</given-names></name> <name><surname>Shahid</surname> <given-names>M.</given-names></name> <name><surname>Murino</surname> <given-names>V.</given-names></name></person-group> (<year>2019</year>). <article-title>Personality traits classification using deep visual activity-based nonverbal features of key-dynamic images</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>12</volume>, <fpage>1084</fpage>&#x2013;<lpage>1099</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2019.2944614</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Biel</surname> <given-names>J.-I.</given-names></name> <name><surname>Aran</surname> <given-names>O.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2011</year>). &#x201C;<article-title>You are known by how you vlog: personality impressions and nonverbal behavior in youtube</article-title>,&#x201D; in <source>Proceedings of the International AAAI Conference on Web and Social Media.</source></citation></ref>
<ref id="ref9"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Biel</surname> <given-names>J.-I.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2010</year>). &#x201C;<article-title>Voices of vlogging</article-title>,&#x201D; in <source>Proceedings of the International AAAI Conference on Web and Social Media.</source></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Biel</surname> <given-names>J.-I.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>The youtube lens: Crowdsourced personality impressions and audiovisual analysis of vlogs</article-title>. <source>IEEE Trans. Multimedia</source> <volume>15</volume>, <fpage>41</fpage>&#x2013;<lpage>55</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TMM.2012.2225032</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Biel</surname> <given-names>J.-I.</given-names></name> <name><surname>Teijeiro-Mosquera</surname> <given-names>L.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). &#x201C;<article-title>Facetube: predicting personality from facial expressions of emotion in online conversational video</article-title>,&#x201D; in <source>Proceedings of the 14th ACM International Conference on Multimodal Interaction</source>, <fpage>53</fpage>&#x2013;<lpage>56</lpage>.</citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carbonneau</surname> <given-names>M.</given-names></name> <name><surname>Granger</surname> <given-names>E.</given-names></name> <name><surname>Attabi</surname> <given-names>Y.</given-names></name> <name><surname>Gagnon</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>Feature learning from spectrograms for assessment of personality traits</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>11</volume>, <fpage>25</fpage>&#x2013;<lpage>31</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2017.2763132</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Cattell</surname> <given-names>H. E.</given-names></name> <name><surname>Mead</surname> <given-names>A. D.</given-names></name></person-group> (<year>2008</year>). The Sixteen Personality Factor Questionnaire (16PF).</citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Celiktutan</surname> <given-names>O.</given-names></name> <name><surname>Skordos</surname> <given-names>E.</given-names></name> <name><surname>Gunes</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>Multimodal human-human-robot interactions (mhhri) dataset for studying personality and engagement</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>10</volume>, <fpage>484</fpage>&#x2013;<lpage>497</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2017.2737019</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Shu</surname> <given-names>H.</given-names></name> <name><surname>Tang</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Shi</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2020</year>). &#x201C;<article-title>Frequency domain compact 3d convolutional neural networks</article-title>,&#x201D; in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, <fpage>1641</fpage>&#x2013;<lpage>1650</lpage>.</citation></ref>
<ref id="ref16"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chung</surname> <given-names>J.</given-names></name> <name><surname>Gulcehre</surname> <given-names>C.</given-names></name> <name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). Empirical evaluation of gated recurrent neural networks on sequence modeling. <source>arXiv preprint arXiv.</source></citation></ref>
<ref id="ref17"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Costa</surname> <given-names>P. T.</given-names></name> <name><surname>McCrae</surname> <given-names>R. R.</given-names></name></person-group> (<year>1998</year>). &#x201C;<article-title>Trait theories of personality</article-title>,&#x201D; in <source>Advanced Personality. The Plenum Series in Social/Clinical Psychology.</source> eds. D. F. Barone, M. Hersen and V. B. van Hasselt (<publisher-name>Boston, MA: Springer</publisher-name>), <fpage>103</fpage>&#x2013;<lpage>121</lpage>.</citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cruz-Mota</surname> <given-names>J.</given-names></name> <name><surname>Bogdanova</surname> <given-names>I.</given-names></name> <name><surname>Paquier</surname> <given-names>B.</given-names></name> <name><surname>Bierlaire</surname> <given-names>M.</given-names></name> <name><surname>Thiran</surname> <given-names>J.-P.</given-names></name></person-group> (<year>2012</year>). <article-title>Scale invariant feature transform on the sphere: theory and applications</article-title>. <source>Int. J. Comput. Vis.</source> <volume>98</volume>, <fpage>217</fpage>&#x2013;<lpage>241</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11263-011-0505-4</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Curto</surname> <given-names>D.</given-names></name> <name><surname>Clap&#x00E9;s</surname> <given-names>A.</given-names></name> <name><surname>Selva</surname> <given-names>J.</given-names></name> <name><surname>Smeureanu</surname> <given-names>S.</given-names></name> <name><surname>Junior</surname> <given-names>J.</given-names></name> <name><surname>Jacques</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2021</year>). &#x201C;<article-title>Dyadformer: a multi-modal transformer for long-range modeling of dyadic interactions</article-title>,&#x201D; in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, <fpage>2177</fpage>&#x2013;<lpage>2188</lpage>.</citation></ref>
<ref id="ref20"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). Bert: pre-training of deep bidirectional transformers for language understanding. <source>arXiv preprint arXiv:1810.04805.</source></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dhelim</surname> <given-names>S.</given-names></name> <name><surname>Aung</surname> <given-names>N.</given-names></name> <name><surname>Bouras</surname> <given-names>M. A.</given-names></name> <name><surname>Ning</surname> <given-names>H.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2021</year>). <article-title>A survey on personality-aware recommendation systems</article-title>. <source>Artif. Intell. Rev.</source> <volume>55</volume>, <fpage>2409</fpage>&#x2013;<lpage>2454</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10462-021-10063-7</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elman</surname> <given-names>J. L.</given-names></name></person-group> (<year>1990</year>). <article-title>Finding structure in time</article-title>. <source>Cogn. Sci.</source> <volume>14</volume>, <fpage>179</fpage>&#x2013;<lpage>211</lpage>. doi: <pub-id pub-id-type="doi">10.1207/s15516709cog1402_1</pub-id></citation></ref>
<ref id="ref23"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Escalante</surname> <given-names>H. J.</given-names></name> <name><surname>Guyon</surname> <given-names>I.</given-names></name> <name><surname>Escalera</surname> <given-names>S.</given-names></name> <name><surname>Jacques</surname> <given-names>J.</given-names></name> <name><surname>Madadi</surname> <given-names>M.</given-names></name> <name><surname>Bar&#x00F3;</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2017</year>). &#x201C;<article-title>Design of an explainable machine learning challenge for video interviews</article-title>,&#x201D; in <source>2017 International Joint Conference on Neural Networks (IJCNN): IEEE</source>, <fpage>3688</fpage>&#x2013;<lpage>3695</lpage>.</citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Escalante</surname> <given-names>H. J.</given-names></name> <name><surname>Kaya</surname> <given-names>H.</given-names></name> <name><surname>Salah</surname> <given-names>A. A.</given-names></name> <name><surname>Escalera</surname> <given-names>S.</given-names></name> <name><surname>G&#x00FC;&#x00E7;</surname> <given-names>Y.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;</surname> <given-names>U.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Modeling, recognizing, and explaining apparent personality from videos</article-title>. <source>IEEE Trans. Affect. Comput.</source>, <fpage>1</fpage>&#x2013;<lpage>18</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2020.2973984</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Eysenck</surname> <given-names>H. J.</given-names></name></person-group> (<year>2012</year>). <source>A Model for Personality.</source> <publisher-name>New York: Springer Science &#x0026; Business Media</publisher-name>.</citation></ref>
<ref id="ref26"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Freund</surname> <given-names>Y.</given-names></name> <name><surname>Haussler</surname> <given-names>D.</given-names></name></person-group> (<year>1994</year>). <article-title>Unsupervised learning of distributions of binary vectors using two layer networks</article-title>.</citation></ref>
<ref id="ref27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Personality trait detection based on ASM localization and deep learning</article-title>. <source>Sci. Program.</source> <volume>2021</volume>, <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi: <pub-id pub-id-type="doi">10.1155/2021/5675917</pub-id></citation></ref>
<ref id="ref28"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Furnham</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Myers-Briggs type indicator (MBTI)</article-title>,&#x201D; in <source>Encyclopedia of personality and individual differences.</source> eds. <person-group person-group-type="editor"><name><surname>Zeigler-Hill</surname> <given-names>V.</given-names></name> <name><surname>Shackelford</surname> <given-names>T. K.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>3059</fpage>&#x2013;<lpage>3062</lpage>.</citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Furnham</surname> <given-names>A. J. P.</given-names></name> <name><surname>Differences</surname> <given-names>I.</given-names></name></person-group> (<year>1996</year>). <article-title>The big five versus the big four: the relationship between the Myers-Briggs type indicator (MBTI) and NEO-PI five factor model of personality</article-title>. <source>Personal. Individ. Differ.</source> <volume>21</volume>, <fpage>303</fpage>&#x2013;<lpage>307</lpage>. doi: <pub-id pub-id-type="doi">10.1016/0191-8869(96)00033-5</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Golbeck</surname> <given-names>J. A.</given-names></name></person-group> (<year>2016</year>). <article-title>Predicting personality from social media text</article-title>. <source>AIS Trans. Replic. Res.</source> <volume>2</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi: <pub-id pub-id-type="doi">10.17705/1atrr.00009</pub-id></citation></ref>
<ref id="ref31"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Golbeck</surname> <given-names>J.</given-names></name> <name><surname>Robles</surname> <given-names>C.</given-names></name> <name><surname>Turner</surname> <given-names>K.</given-names></name></person-group> (<year>2011</year>). &#x201C;<article-title>Predicting personality with social media</article-title>,&#x201D; in <source>CHI&#x2019;11 Extended Abstracts on Human Factors in Computing Systems.</source> <fpage>253</fpage>&#x2013;<lpage>262</lpage>.</citation></ref>
<ref id="ref32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gorbova</surname> <given-names>J.</given-names></name> <name><surname>Avots</surname> <given-names>E.</given-names></name> <name><surname>L&#x00FC;si</surname> <given-names>I.</given-names></name> <name><surname>Fishel</surname> <given-names>M.</given-names></name> <name><surname>Escalera</surname> <given-names>S.</given-names></name> <name><surname>Anbarjafari</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>Integrating vision and language for first-impression personality analysis</article-title>. <source>IEEE Multimedia</source> <volume>25</volume>, <fpage>24</fpage>&#x2013;<lpage>33</lpage>. doi: <pub-id pub-id-type="doi">10.1109/MMUL.2018.023121162</pub-id></citation></ref>
<ref id="ref33"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Gorbova</surname> <given-names>J.</given-names></name> <name><surname>Lusi</surname> <given-names>I.</given-names></name> <name><surname>Litvin</surname> <given-names>A.</given-names></name> <name><surname>Anbarjafari</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Automated screening of job candidate based on multimodal video processing</article-title>,&#x201D; in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</source>, <fpage>29</fpage>&#x2013;<lpage>35</lpage>.</citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goreis</surname> <given-names>A.</given-names></name> <name><surname>Voracek</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>A systematic review and meta-analysis of psychological research on conspiracy beliefs: field characteristics, measurement instruments, and associations with personality traits</article-title>. <source>Front. Psychol.</source> <volume>10</volume>:<fpage>205</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fpsyg.2019.00205</pub-id>, PMID: <pub-id pub-id-type="pmid">30853921</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guadagno</surname> <given-names>R. E.</given-names></name> <name><surname>Okdie</surname> <given-names>B. M.</given-names></name> <name><surname>Eno</surname> <given-names>C. A.</given-names></name></person-group> (<year>2008</year>). <article-title>Who blogs? Personality predictors of blogging</article-title>. <source>Comput. Hum. Behav.</source> <volume>24</volume>, <fpage>1993</fpage>&#x2013;<lpage>2004</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.chb.2007.09.001</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk</surname> <given-names>Y.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;</surname> <given-names>U.</given-names></name> <name><surname>Baro</surname> <given-names>X.</given-names></name> <name><surname>Escalante</surname> <given-names>H. J.</given-names></name> <name><surname>Guyon</surname> <given-names>I.</given-names></name> <name><surname>Escalera</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Multimodal first impression analysis with deep residual networks</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>9</volume>, <fpage>316</fpage>&#x2013;<lpage>329</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2017.2751469</pub-id></citation></ref>
<ref id="ref37"><citation citation-type="other"><person-group person-group-type="author"><name><surname>G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk</surname> <given-names>Y.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;</surname> <given-names>U.</given-names></name> <name><surname>van Gerven</surname> <given-names>M. A.</given-names></name> <name><surname>van Lier</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Deep impression: audiovisual deep residual networks for multimodal apparent personality trait recognition</article-title>,&#x201D; in <source>European Conference on Computer Vision (Springer)</source>, <fpage>349</fpage>&#x2013;<lpage>358</lpage>.</citation></ref>
<ref id="ref38"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Guntuku</surname> <given-names>S. C.</given-names></name> <name><surname>Qiu</surname> <given-names>L.</given-names></name> <name><surname>Roy</surname> <given-names>S.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Jakhetiya</surname> <given-names>V.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Do others perceive you as you want them to? Modeling personality based on selfies</article-title>,&#x201D; in <source>Proceedings of the 1st International Workshop on Affect &#x0026; Sentiment in Multimedia</source>, <fpage>21</fpage>&#x2013;<lpage>26</lpage>.</citation></ref>
<ref id="ref39"><citation citation-type="other"><person-group person-group-type="author"><name><surname>G&#x00FC;rpinar</surname> <given-names>F.</given-names></name> <name><surname>Kaya</surname> <given-names>H.</given-names></name> <name><surname>Salah</surname> <given-names>A. A.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Multimodal fusion of audio, scene, and face features for first impression estimation</article-title>,&#x201D; in <source>2016 23rd International Conference on Pattern Recognition (ICPR): IEEE</source>, <fpage>43</fpage>&#x2013;<lpage>48</lpage>.</citation></ref>
<ref id="ref40"><citation citation-type="other"><person-group person-group-type="author"><name><surname>G&#x00FC;rp&#x0131;nar</surname> <given-names>F.</given-names></name> <name><surname>Kaya</surname> <given-names>H.</given-names></name> <name><surname>Salah</surname> <given-names>A. A.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Combining deep facial and ambient features for first impression estimation</article-title>,&#x201D; in <source>European Conference on Computer Vision (Springer)</source>, <fpage>372</fpage>&#x2013;<lpage>385</lpage>.</citation></ref>
<ref id="ref41"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hayat</surname> <given-names>H.</given-names></name> <name><surname>Ventura</surname> <given-names>C.</given-names></name> <name><surname>Lapedriza</surname> <given-names>&#x00C0;.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>On the use of interpretable CNN for personality trait recognition from audio</article-title>,&#x201D; in <source>CCIA</source>, <fpage>135</fpage>&#x2013;<lpage>144</lpage>.</citation></ref>
<ref id="ref42"><citation citation-type="other"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Deep residual learning for image recognition</article-title>,&#x201D; in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>770</fpage>&#x2013;<lpage>778</lpage>.</citation></ref>
<ref id="ref43"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hernandez</surname> <given-names>R.</given-names></name> <name><surname>Scott</surname> <given-names>I.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Predicting Myers-Briggs type indicator with text</article-title>,&#x201D; in <source>31st Conference on Neural Information Processing Systems (NIPS)</source>, <fpage>4</fpage>&#x2013;<lpage>9</lpage>.</citation></ref>
<ref id="ref44"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Hershey</surname> <given-names>S.</given-names></name> <name><surname>Chaudhuri</surname> <given-names>S.</given-names></name> <name><surname>Ellis</surname> <given-names>D. P.</given-names></name> <name><surname>Gemmeke</surname> <given-names>J. F.</given-names></name> <name><surname>Jansen</surname> <given-names>A.</given-names></name> <name><surname>Moore</surname> <given-names>R. C.</given-names></name> <etal/></person-group>. (<year>2017</year>). &#x201C;<article-title>CNN architectures for large-scale audio classification</article-title>,&#x201D; in <source>2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP): IEEE</source>, <fpage>131</fpage>&#x2013;<lpage>135</lpage>.</citation></ref>
<ref id="ref45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2002</year>). <article-title>Training products of experts by minimizing contrastive divergence</article-title>. <source>Neural Comput.</source> <volume>14</volume>, <fpage>1771</fpage>&#x2013;<lpage>1800</lpage>. doi: <pub-id pub-id-type="doi">10.1162/089976602760128018</pub-id>, PMID: <pub-id pub-id-type="pmid">12180402</pub-id></citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Osindero</surname> <given-names>S.</given-names></name> <name><surname>Teh</surname> <given-names>Y.-W.</given-names></name></person-group> (<year>2006</year>). <article-title>A fast learning algorithm for deep belief nets</article-title>. <source>Neural Comput.</source> <volume>18</volume>, <fpage>1527</fpage>&#x2013;<lpage>1554</lpage>. doi: <pub-id pub-id-type="doi">10.1162/neco.2006.18.7.1527</pub-id>, PMID: <pub-id pub-id-type="pmid">16764513</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput.</source> <volume>9</volume>, <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>. doi: <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Van Der Maaten</surname> <given-names>L.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Densely connected convolutional networks</article-title>,&#x201D; in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>4700</fpage>&#x2013;<lpage>4708</lpage>.</citation></ref>
<ref id="ref49"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>B.</given-names></name> <name><surname>Ren</surname> <given-names>Q.</given-names></name> <name><surname>Dai</surname> <given-names>F.</given-names></name> <name><surname>Xiong</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Gui</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Multi-task cascaded convolutional neural networks for real-time dynamic face recognition method</article-title>,&#x201D; in <source>International Conference in Communications, Signal Processing, and Systems (Springer)</source>, <fpage>59</fpage>&#x2013;<lpage>66</lpage>.</citation></ref>
<ref id="ref50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Junior</surname> <given-names>J. C. S. J.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk</surname> <given-names>Y.</given-names></name> <name><surname>P&#x00E9;rez</surname> <given-names>M.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;</surname> <given-names>U.</given-names></name> <name><surname>Andujar</surname> <given-names>C.</given-names></name> <name><surname>Bar&#x00F3;</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>First impressions: a survey on vision-based apparent personality trait analysis</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>13</volume>, <fpage>75</fpage>&#x2013;<lpage>95</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2019.2930058</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Junior</surname> <given-names>J.</given-names></name> <name><surname>Jacques</surname> <given-names>C.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;t&#x00FC;rk</surname> <given-names>Y.</given-names></name> <name><surname>P&#x00E9;rez</surname> <given-names>M.</given-names></name> <name><surname>G&#x00FC;&#x00E7;l&#x00FC;</surname> <given-names>U.</given-names></name> <name><surname>Andujar</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2018</year>). First impressions: a survey on computer vision-based apparent personality trait analysis. <source>arXiv preprint arXiv:.08046.</source></citation></ref>
<ref id="ref52"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Kampman</surname> <given-names>O.</given-names></name> <name><surname>Barezi</surname> <given-names>E. J.</given-names></name> <name><surname>Bertero</surname> <given-names>D.</given-names></name> <name><surname>Fung</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Investigating audio, video, and text fusion methods for end-to-end automatic personality prediction</article-title>,&#x201D; in <source>Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>, <fpage>606</fpage>&#x2013;<lpage>611</lpage>.</citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>D. Y.</given-names></name> <name><surname>Song</surname> <given-names>H. Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Method of predicting human mobility patterns using deep learning</article-title>. <source>Neurocomputing</source> <volume>280</volume>, <fpage>56</fpage>&#x2013;<lpage>64</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2017.07.069</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krichmar</surname> <given-names>J. L.</given-names></name> <name><surname>Olds</surname> <given-names>J. L.</given-names></name> <name><surname>Sanchez-Andres</surname> <given-names>J. V.</given-names></name> <name><surname>Tang</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Explainable artificial intelligence and neuroscience: cross-disciplinary perspectives</article-title>. <source>Front. Neurorobot.</source> <volume>15</volume>:<fpage>731733</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fnbot.2021.731733</pub-id></citation></ref>
<ref id="ref55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>Adv. Neural Inf. Proces. Syst.</source> <volume>25</volume>, <fpage>1097</fpage>&#x2013;<lpage>1105</lpage>.</citation></ref>
<ref id="ref56"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Kumawat</surname> <given-names>S.</given-names></name> <name><surname>Raman</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Lp-3dcnn: unveiling local phase in 3d convolutional neural networks</article-title>,&#x201D; in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, <fpage>4903</fpage>&#x2013;<lpage>4912</lpage>.</citation></ref>
<ref id="ref57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kurmi</surname> <given-names>V. K.</given-names></name> <name><surname>Subramanian</surname> <given-names>V. K.</given-names></name> <name><surname>Namboodiri</surname> <given-names>V. P.</given-names></name></person-group> (<year>2021</year>). <article-title>Exploring dropout discriminator for domain adaptation</article-title>. <source>Neurocomputing</source> <volume>457</volume>, <fpage>168</fpage>&#x2013;<lpage>181</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2021.06.043</pub-id></citation></ref>
<ref id="ref58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x2013;<lpage>444</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature14539</pub-id></citation></ref>
<ref id="ref59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bottou</surname> <given-names>L.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Haffner</surname> <given-names>P.</given-names></name></person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x2013;<lpage>2324</lpage>. doi: <pub-id pub-id-type="doi">10.1109/5.726791</pub-id></citation></ref>
<ref id="ref60"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Grosse</surname> <given-names>R.</given-names></name> <name><surname>Ranganath</surname> <given-names>R.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2009</year>). &#x201C;<article-title>Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations</article-title>,&#x201D; in <source>Proceedings of the 26th Annual International Conference on Machine Learning</source>, <fpage>609</fpage>&#x2013;<lpage>616</lpage>.</citation></ref>
<ref id="ref61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Dong</surname> <given-names>P.</given-names></name> <name><surname>Xiao</surname> <given-names>B.</given-names></name> <name><surname>Zhou</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Object recognition based on the region of interest and optimal bag of words model</article-title>. <source>Neurocomputing</source> <volume>172</volume>, <fpage>271</fpage>&#x2013;<lpage>280</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2015.01.083</pub-id></citation></ref>
<ref id="ref62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <name><surname>Long</surname> <given-names>X.</given-names></name> <name><surname>Tang</surname> <given-names>L.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2020a</year>). <article-title>EEG responses to emotional videos can quantitatively predict big-five personality traits</article-title>. <source>Neurocomputing</source> <volume>415</volume>, <fpage>368</fpage>&#x2013;<lpage>381</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2020.07.123</pub-id></citation></ref>
<ref id="ref63"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Kazameini</surname> <given-names>A.</given-names></name> <name><surname>Mehta</surname> <given-names>Y.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2021</year>). <article-title>Multitask learning for emotion and personality detection</article-title>. <source>CoRR</source> abs/2101.02346.</citation></ref>
<ref id="ref64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Wan</surname> <given-names>J.</given-names></name> <name><surname>Miao</surname> <given-names>Q.</given-names></name> <name><surname>Escalera</surname> <given-names>S.</given-names></name> <name><surname>Fang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2020b</year>). <article-title>CR-net: a deep classification-regression network for multimodal apparent personality analysis</article-title>. <source>Int. J. Comput. Vis.</source> <volume>128</volume>, <fpage>2763</fpage>&#x2013;<lpage>2780</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11263-020-01309-y</pub-id></citation></ref>
<ref id="ref65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>T.</given-names></name> <name><surname>Glossner</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Shi</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2021a</year>). <article-title>Pruning and quantization for deep neural network acceleration: a survey</article-title>. <source>Neurocomputing</source> <volume>461</volume>, <fpage>370</fpage>&#x2013;<lpage>403</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2021.07.045</pub-id></citation></ref>
<ref id="ref66"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Yan</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Jiang</surname> <given-names>C.</given-names></name></person-group> (<year>2021b</year>). <article-title>Explaining the black-box model: a survey of local interpretation methods for deep neural networks</article-title>. <source>Neurocomputing</source> <volume>419</volume>, <fpage>168</fpage>&#x2013;<lpage>182</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2020.08.011</pub-id></citation></ref>
<ref id="ref67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Jiang</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>PT-LDA: A latent variable model to predict personality traits of social network users</article-title>. <source>Neurocomputing</source> <volume>210</volume>, <fpage>155</fpage>&#x2013;<lpage>163</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2015.10.144</pub-id></citation></ref>
<ref id="ref68"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Majumder</surname> <given-names>N.</given-names></name> <name><surname>Poria</surname> <given-names>S.</given-names></name> <name><surname>Gelbukh</surname> <given-names>A.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep learning-based document modeling for personality detection from text</article-title>. <source>IEEE Intell. Syst.</source> <volume>32</volume>, <fpage>74</fpage>&#x2013;<lpage>79</lpage>. doi: <pub-id pub-id-type="doi">10.1109/MIS.2017.23</pub-id></citation></ref>
<ref id="ref69"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Masuyama</surname> <given-names>N.</given-names></name> <name><surname>Loo</surname> <given-names>C. K.</given-names></name> <name><surname>Seera</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Personality affected robotic emotional model with associative memory for human-robot interaction</article-title>. <source>Neurocomputing</source> <volume>272</volume>, <fpage>213</fpage>&#x2013;<lpage>225</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2017.06.069</pub-id></citation></ref>
<ref id="ref70"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCrae</surname> <given-names>R. R.</given-names></name> <name><surname>John</surname> <given-names>O. P.</given-names></name></person-group> (<year>1992</year>). <article-title>An introduction to the five-factor model and its applications</article-title>. <source>J. Pers.</source> <volume>60</volume>, <fpage>175</fpage>&#x2013;<lpage>215</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-6494.1992.tb00970.x</pub-id>, PMID: <pub-id pub-id-type="pmid">1635039</pub-id></citation></ref>
<ref id="ref71"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McKeown</surname> <given-names>G.</given-names></name> <name><surname>Valstar</surname> <given-names>M.</given-names></name> <name><surname>Cowie</surname> <given-names>R.</given-names></name> <name><surname>Pantic</surname> <given-names>M.</given-names></name> <name><surname>Schroder</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>The SEMAINE database: annotated multimodal records of emotionally colored conversations between a person and a limited agent</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>3</volume>, <fpage>5</fpage>&#x2013;<lpage>17</lpage>. doi: <pub-id pub-id-type="doi">10.1109/T-AFFC.2011.20</pub-id></citation></ref>
<ref id="ref72"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Mehta</surname> <given-names>Y.</given-names></name> <name><surname>Fatehi</surname> <given-names>S.</given-names></name> <name><surname>Kazameini</surname> <given-names>A.</given-names></name> <name><surname>Stachl</surname> <given-names>C.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name> <name><surname>Eetemadi</surname> <given-names>S.</given-names></name></person-group> (<year>2020a</year>). &#x201C;<article-title>Bottom-up and top-down: predicting personality with psycholinguistic and language model features</article-title>,&#x201D; in <source>2020 IEEE International Conference on Data Mining (ICDM)</source>, <fpage>1184</fpage>&#x2013;<lpage>1189</lpage>.</citation></ref>
<ref id="ref73"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehta</surname> <given-names>Y.</given-names></name> <name><surname>Majumder</surname> <given-names>N.</given-names></name> <name><surname>Gelbukh</surname> <given-names>A.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2020b</year>). <article-title>Recent trends in deep learning based personality detection</article-title>. <source>Artif. Intell. Rev.</source> <volume>53</volume>, <fpage>2313</fpage>&#x2013;<lpage>2339</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10462-019-09770-z</pub-id></citation></ref>
<ref id="ref74"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohammadi</surname> <given-names>G.</given-names></name> <name><surname>Vinciarelli</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Automatic personality perception: prediction of trait attribution based on prosodic features</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>3</volume>, <fpage>273</fpage>&#x2013;<lpage>284</lpage>. doi: <pub-id pub-id-type="doi">10.1109/T-AFFC.2012.5</pub-id></citation></ref>
<ref id="ref75"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Palmero</surname> <given-names>C.</given-names></name> <name><surname>Selva</surname> <given-names>J.</given-names></name> <name><surname>Smeureanu</surname> <given-names>S.</given-names></name> <name><surname>Junior</surname> <given-names>J.</given-names></name> <name><surname>Jacques</surname> <given-names>C.</given-names></name> <name><surname>Clap&#x00E9;s</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). &#x201C;<article-title>Context-aware personality inference in dyadic scenarios: introducing the udiva dataset</article-title>,&#x201D; in <source>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>.</citation></ref>
<ref id="ref76"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Pennebaker</surname> <given-names>J. W.</given-names></name> <name><surname>Francis</surname> <given-names>M. E.</given-names></name> <name><surname>Booth</surname> <given-names>R. J.</given-names></name></person-group> (<year>2001</year>). <source>Linguistic Inquiry and Word Count: LIWC 2001.</source> <publisher-loc>Mahway</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates</publisher-name>.</citation></ref>
<ref id="ref77"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Ponce-L&#x00F3;pez</surname> <given-names>V.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Oliu</surname> <given-names>M.</given-names></name> <name><surname>Corneanu</surname> <given-names>C.</given-names></name> <name><surname>Clap&#x00E9;s</surname> <given-names>A.</given-names></name> <name><surname>Guyon</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2016</year>). &#x201C;<article-title>Chalearn lap 2016: first round challenge on first impressions-dataset and results</article-title>,&#x201D; in <source>European Conference on Computer Vision</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>400</fpage>&#x2013;<lpage>418</lpage>.</citation></ref>
<ref id="ref78"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Principi</surname> <given-names>R. D. P.</given-names></name> <name><surname>Palmero</surname> <given-names>C.</given-names></name> <name><surname>Junior</surname> <given-names>J. C.</given-names></name> <name><surname>Escalera</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>On the effect of observed subject biases in apparent personality analysis from audio-visual signals</article-title>. <source>IEEE Trans. Affect. Comput.</source> 12, <fpage>607</fpage>&#x2013;<lpage>621</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2019.2956030</pub-id></citation></ref>
<ref id="ref79"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiu</surname> <given-names>L.</given-names></name> <name><surname>Lin</surname> <given-names>H.</given-names></name> <name><surname>Ramsay</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>). <article-title>You are what you tweet: personality expression and perception on twitter</article-title>. <source>J. Res. Pers.</source> <volume>46</volume>, <fpage>710</fpage>&#x2013;<lpage>718</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jrp.2012.08.008</pub-id></citation></ref>
<ref id="ref80"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Quercia</surname> <given-names>D.</given-names></name> <name><surname>Las Casas</surname> <given-names>D.</given-names></name> <name><surname>Pesce</surname> <given-names>J. P.</given-names></name> <name><surname>Stillwell</surname> <given-names>D.</given-names></name> <name><surname>Kosinski</surname> <given-names>M.</given-names></name> <name><surname>Almeida</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2012</year>). &#x201C;<article-title>Facebook and privacy: The balancing act of personality, gender, and relationship currency</article-title>,&#x201D; in <source>Sixth International AAAI Conference on Weblogs and Social Media.</source></citation></ref>
<ref id="ref81"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rammstedt</surname> <given-names>B.</given-names></name> <name><surname>John</surname> <given-names>O. P.</given-names></name></person-group> (<year>2007</year>). <article-title>Measuring personality in one minute or less: a 10-item short version of the big five inventory in English and German</article-title>. <source>J. Res. Pers.</source> <volume>41</volume>, <fpage>203</fpage>&#x2013;<lpage>212</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jrp.2006.02.001</pub-id></citation></ref>
<ref id="ref82"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>Z.</given-names></name> <name><surname>Shen</surname> <given-names>Q.</given-names></name> <name><surname>Diao</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>A sentiment-aware deep learning approach for personality detection from text</article-title>. <source>Inf. Process. Manag.</source> <volume>58</volume>:<fpage>102532</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ipm.2021.102532</pub-id></citation></ref>
<ref id="ref83"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodr&#x00ED;guez</surname> <given-names>P.</given-names></name> <name><surname>Velazquez</surname> <given-names>D.</given-names></name> <name><surname>Cucurull</surname> <given-names>G.</given-names></name> <name><surname>Gonfaus</surname> <given-names>J. M.</given-names></name> <name><surname>Roca</surname> <given-names>F. X.</given-names></name> <name><surname>Ozawa</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Personality trait analysis in social networks based on weakly supervised learning of shared images</article-title>. <source>Appl. Sci.</source> <volume>10</volume>:<fpage>8170</fpage>. doi: <pub-id pub-id-type="doi">10.3390/app10228170</pub-id></citation></ref>
<ref id="ref84"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sanchez-Cortes</surname> <given-names>D.</given-names></name> <name><surname>Aran</surname> <given-names>O.</given-names></name> <name><surname>Jayagopi</surname> <given-names>D. B.</given-names></name> <name><surname>Schmid Mast</surname> <given-names>M.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>Emergent leaders through looking and speaking: from audio-visual data to multimodal recognition</article-title>. <source>J. Multimodal User Interfaces</source> <volume>7</volume>, <fpage>39</fpage>&#x2013;<lpage>53</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s12193-012-0101-0</pub-id></citation></ref>
<ref id="ref85"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sarkis-Onofre</surname> <given-names>R.</given-names></name> <name><surname>Catal&#x00E1;-L&#x00F3;pez</surname> <given-names>F.</given-names></name> <name><surname>Aromataris</surname> <given-names>E.</given-names></name> <name><surname>Lockwood</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>How to properly use the PRISMA statement</article-title>. <source>Syst. Rev.</source> <volume>10</volume>, <fpage>1</fpage>&#x2013;<lpage>3</lpage>. doi: <pub-id pub-id-type="doi">10.1186/s13643-021-01671-z</pub-id></citation></ref>
<ref id="ref86"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning in neural networks: an overview</article-title>. <source>Neural Netw.</source> <volume>61</volume>, <fpage>85</fpage>&#x2013;<lpage>117</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neunet.2014.09.003</pub-id>, PMID: <pub-id pub-id-type="pmid">25462637</pub-id></citation></ref>
<ref id="ref87"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schuller</surname> <given-names>B.</given-names></name> <name><surname>Steidl</surname> <given-names>S.</given-names></name> <name><surname>Batliner</surname> <given-names>A.</given-names></name> <name><surname>N&#x00F6;th</surname> <given-names>E.</given-names></name> <name><surname>Vinciarelli</surname> <given-names>A.</given-names></name> <name><surname>Burkhardt</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>A survey on perceived speaker traits: personality, likability, pathology, and the first challenge</article-title>. <source>Comput. Speech Lang.</source> <volume>29</volume>, <fpage>100</fpage>&#x2013;<lpage>131</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.csl.2014.08.003</pub-id></citation></ref>
<ref id="ref88"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Schuller</surname> <given-names>B.</given-names></name> <name><surname>Steidl</surname> <given-names>S.</given-names></name> <name><surname>Batliner</surname> <given-names>A.</given-names></name> <name><surname>Vinciarelli</surname> <given-names>A.</given-names></name> <name><surname>Scherer</surname> <given-names>K.</given-names></name> <name><surname>Ringeval</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2013</year>). &#x201C;<article-title>The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism</article-title>,&#x201D; in <source>Proceedings INTERSPEECH 2013, 14th Annual Conference of the International Speech Communication Association, Lyon, France.</source></citation></ref>
<ref id="ref89"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Segalin</surname> <given-names>C.</given-names></name> <name><surname>Cheng</surname> <given-names>D. S.</given-names></name> <name><surname>Cristani</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Social profiling through image understanding: personality inference using convolutional neural networks</article-title>. <source>Comput. Vis. Image Underst.</source> <volume>156</volume>, <fpage>34</fpage>&#x2013;<lpage>50</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cviu.2016.10.013</pub-id></citation></ref>
<ref id="ref90"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shao</surname> <given-names>H.</given-names></name> <name><surname>Zhong</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>One-shot cross-dataset palmprint recognition via adversarial domain adaptation</article-title>. <source>Neurocomputing</source> <volume>432</volume>, <fpage>288</fpage>&#x2013;<lpage>299</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2020.12.072</pub-id></citation></ref>
<ref id="ref91"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A. J.</given-names></name></person-group> (<year>2014</year>). Very deep convolutional networks for large-scale image recognition.</citation></ref>
<ref id="ref92"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Su</surname> <given-names>M.-H.</given-names></name> <name><surname>Wu</surname> <given-names>C.-H.</given-names></name> <name><surname>Huang</surname> <given-names>K.-Y.</given-names></name> <name><surname>Hong</surname> <given-names>Q.-B.</given-names></name> <name><surname>Wang</surname> <given-names>H.-M.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Personality trait perception from speech signals using multiresolution analysis and convolutional neural networks</article-title>,&#x201D; in <source>2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC): IEEE</source>, <fpage>1532</fpage>&#x2013;<lpage>1536</lpage>.</citation></ref>
<ref id="ref93"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Subramaniam</surname> <given-names>A.</given-names></name> <name><surname>Patel</surname> <given-names>V.</given-names></name> <name><surname>Mishra</surname> <given-names>A.</given-names></name> <name><surname>Balasubramanian</surname> <given-names>P.</given-names></name> <name><surname>Mittal</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Bi-modal first impressions recognition using temporally ordered deep audio and stochastic visual features</article-title>,&#x201D; in <source>European Conference on Computer Vision (Springer)</source>, <fpage>337</fpage>&#x2013;<lpage>348</lpage>.</citation></ref>
<ref id="ref94"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Subramanian</surname> <given-names>R.</given-names></name> <name><surname>Wache</surname> <given-names>J.</given-names></name> <name><surname>Abadi</surname> <given-names>M. K.</given-names></name> <name><surname>Vieriu</surname> <given-names>R. L.</given-names></name> <name><surname>Winkler</surname> <given-names>S.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>ASCERTAIN: emotion and personality recognition using commercial sensors</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>9</volume>, <fpage>147</fpage>&#x2013;<lpage>160</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2016.2625250</pub-id></citation></ref>
<ref id="ref95"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Suman</surname> <given-names>C.</given-names></name> <name><surname>Saha</surname> <given-names>S.</given-names></name> <name><surname>Gupta</surname> <given-names>A.</given-names></name> <name><surname>Pandey</surname> <given-names>S. K.</given-names></name> <name><surname>Bhattacharyya</surname> <given-names>P.</given-names></name></person-group> (<year>2022</year>). <article-title>A multi-modal personality prediction system</article-title>. <source>Knowl.-Based Syst.</source> <volume>236</volume>:<fpage>107715</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2021.107715</pub-id></citation></ref>
<ref id="ref96"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Cao</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name> <name><surname>Shen</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Who am I? Personality detection based on deep learning for texts</article-title>,&#x201D; in <source>2018 IEEE International Conference on Communications (ICC): IEEE)</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</citation></ref>
<ref id="ref97"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Z.</given-names></name> <name><surname>Song</surname> <given-names>Q.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>A novel ensemble method for classifying imbalanced data</article-title>. <source>Pattern Recogn.</source> <volume>48</volume>, <fpage>1623</fpage>&#x2013;<lpage>1637</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.patcog.2014.11.014</pub-id></citation></ref>
<ref id="ref98"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Sermanet</surname> <given-names>P.</given-names></name> <name><surname>Reed</surname> <given-names>S.</given-names></name> <name><surname>Anguelov</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2015</year>). &#x201C;<article-title>Going deeper with convolutions</article-title>,&#x201D; in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</citation></ref>
<ref id="ref99"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taib</surname> <given-names>R.</given-names></name> <name><surname>Berkovsky</surname> <given-names>S.</given-names></name> <name><surname>Koprinska</surname> <given-names>I.</given-names></name> <name><surname>Wang</surname> <given-names>E.</given-names></name> <name><surname>Zeng</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Personality sensing: detection of personality traits using physiological responses to image and video stimuli</article-title>. <source>ACM Trans. Interact. Intell. Syst.</source> <volume>10</volume>, <fpage>1</fpage>&#x2013;<lpage>32</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3357459</pub-id></citation></ref>
<ref id="ref100"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tartaglione</surname> <given-names>E.</given-names></name> <name><surname>Lathuili&#x00E8;re</surname> <given-names>S.</given-names></name> <name><surname>Fiandrotti</surname> <given-names>A.</given-names></name> <name><surname>Cagnazzo</surname> <given-names>M.</given-names></name> <name><surname>Grangetto</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>HEMP: high-order entropy minimization for neural network compression</article-title>. <source>Neurocomputing</source> <volume>461</volume>, <fpage>244</fpage>&#x2013;<lpage>253</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2021.07.022</pub-id></citation></ref>
<ref id="ref101"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teijeiro-Mosquera</surname> <given-names>L.</given-names></name> <name><surname>Biel</surname> <given-names>J.-I.</given-names></name> <name><surname>Alba-Castro</surname> <given-names>J. L.</given-names></name> <name><surname>Gatica-Perez</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>What your face vlogs about: expressions of emotion and big-five traits impressions in YouTube</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>6</volume>, <fpage>193</fpage>&#x2013;<lpage>205</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2014.2370044</pub-id></citation></ref>
<ref id="ref102"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Tjoa</surname> <given-names>E.</given-names></name> <name><surname>Guan</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>A survey on explainable artificial intelligence (xai): Toward medical xai</article-title>.&#x201D; in <source>IEEE Transactions on Neural Networks and Learning Systems.</source></citation></ref>
<ref id="ref103"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Tran</surname> <given-names>D.</given-names></name> <name><surname>Bourdev</surname> <given-names>L.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name> <name><surname>Torresani</surname> <given-names>L.</given-names></name> <name><surname>Paluri</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Learning spatiotemporal features with 3d convolutional networks</article-title>,&#x201D; in <source>Proceedings of the IEEE International Conference on Computer Vision</source>, <fpage>4489</fpage>&#x2013;<lpage>4497</lpage>.</citation></ref>
<ref id="ref104"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Tran</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Torresani</surname> <given-names>L.</given-names></name> <name><surname>Ray</surname> <given-names>J.</given-names></name> <name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Paluri</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>A closer look at spatiotemporal convolutions for action recognition</article-title>,&#x201D; in <source>Proceedings of the IEEE conference on Computer Vision and Pattern Recognition</source>, <fpage>6450</fpage>&#x2013;<lpage>6459</lpage>.</citation></ref>
<ref id="ref105"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Ventura</surname> <given-names>C.</given-names></name> <name><surname>Masip</surname> <given-names>D.</given-names></name> <name><surname>Lapedriza</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Interpreting CNN models for apparent personality trait regression</article-title>,&#x201D; in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</source>, <fpage>55</fpage>&#x2013;<lpage>63</lpage>.</citation></ref>
<ref id="ref106"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Vilares</surname> <given-names>D.</given-names></name> <name><surname>Peng</surname> <given-names>H.</given-names></name> <name><surname>Satapathy</surname> <given-names>R.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>BabelSenticNet: a commonsense reasoning framework for multilingual sentiment analysis</article-title>,&#x201D; in <source>2018 IEEE symposium series on computational intelligence (SSCI)</source>, <fpage>1292</fpage>&#x2013;<lpage>1298</lpage>.</citation></ref>
<ref id="ref107"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vinciarelli</surname> <given-names>A.</given-names></name> <name><surname>Mohammadi</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>A survey of personality computing</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>5</volume>, <fpage>273</fpage>&#x2013;<lpage>291</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2014.2330816</pub-id></citation></ref>
<ref id="ref108"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wache</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). &#x201C;<article-title>The secret language of our body: affect and personality recognition using physiological signals</article-title>,&#x201D; in <source>Proceedings of the 16th International Conference on Multimodal Interaction</source>, <fpage>389</fpage>&#x2013;<lpage>393</lpage>.</citation></ref>
<ref id="ref109"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Zou</surname> <given-names>Q.</given-names></name> <name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep domain adaptation with differential privacy</article-title>. <source>IEEE Trans. Inf. Forensic. Secur.</source> <volume>15</volume>, <fpage>3093</fpage>&#x2013;<lpage>3106</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIFS.2020.2983254</pub-id></citation></ref>
<ref id="ref110"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Qiao</surname> <given-names>J.</given-names></name> <name><surname>Bi</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Zhou</surname> <given-names>M.</given-names></name> <collab id="coll1">Electrical and Computer Engineering</collab></person-group> (<year>2018</year>). <article-title>TL-GDBN: growing deep belief network with transfer learning</article-title>. <source>IEEE Trans. Autom. Sci.</source> <volume>16</volume>, <fpage>874</fpage>&#x2013;<lpage>885</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TASE.2018.2865663</pub-id></citation></ref>
<ref id="ref111"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>X.-S.</given-names></name> <name><surname>Zhang</surname> <given-names>C.-L.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep bimodal regression of apparent personality traits from short video sequences</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>9</volume>, <fpage>303</fpage>&#x2013;<lpage>315</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2017.2762299</pub-id></citation></ref>
<ref id="ref112"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Werbos</surname> <given-names>P.</given-names></name></person-group> (<year>1990</year>). <article-title>Backpropagation through time: what it does and how to do it</article-title>. <source>Proc. IEEE</source> <volume>78</volume>, <fpage>1550</fpage>&#x2013;<lpage>1560</lpage>. doi: <pub-id pub-id-type="doi">10.1109/5.58337</pub-id></citation></ref>
<ref id="ref113"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Willis</surname> <given-names>J.</given-names></name> <name><surname>Todorov</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>First impressions: making up your mind after a 100-ms exposure to a face</article-title>. <source>Psychol. Sci.</source> <volume>17</volume>, <fpage>592</fpage>&#x2013;<lpage>598</lpage>. doi: <pub-id pub-id-type="doi">10.1111/j.1467-9280.2006.01750.x</pub-id></citation></ref>
<ref id="ref114"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Xianyu</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Cai</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Heterogeneity-entropy based unsupervised feature learning for personality prediction with cross-media data</article-title>,&#x201D; in <source>2016 IEEE international conference on multimedia and Expo (ICME)</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</citation></ref>
<ref id="ref115"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xue</surname> <given-names>D.</given-names></name> <name><surname>Wu</surname> <given-names>L.</given-names></name> <name><surname>Hong</surname> <given-names>Z.</given-names></name> <name><surname>Guo</surname> <given-names>S.</given-names></name> <name><surname>Gao</surname> <given-names>L.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Deep learning-based personality recognition from text posts of online social networks</article-title>. <source>Appl. Intell.</source> <volume>48</volume>, <fpage>4232</fpage>&#x2013;<lpage>4246</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10489-018-1212-4</pub-id></citation></ref>
<ref id="ref116"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>A.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Peng</surname> <given-names>L.</given-names></name> <name><surname>Yan</surname> <given-names>Q.</given-names></name> <name><surname>Hassan</surname> <given-names>M. U.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Effective detection of mobile malware behavior based on explainable deep neural network</article-title>. <source>Neurocomputing</source> <volume>453</volume>, <fpage>482</fpage>&#x2013;<lpage>492</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2020.09.082</pub-id></citation></ref>
<ref id="ref117"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Nie</surname> <given-names>J.</given-names></name> <name><surname>Huang</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Cao</surname> <given-names>Q.</given-names></name> <name><surname>Wei</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Exploring relationship between face and trustworthy impression using mid-level facial features</article-title>,&#x201D; in <source>International Conference on Multimedia Modeling (Springer)</source>, <fpage>540</fpage>&#x2013;<lpage>549</lpage>.</citation></ref>
<ref id="ref118"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Yuan</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Du</surname> <given-names>Y.</given-names></name> <name><surname>Xing</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Asymmetric 3d convolutional neural networks for action recognition</article-title>. <source>Pattern Recogn.</source> <volume>85</volume>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.patcog.2018.07.028</pub-id></citation></ref>
<ref id="ref119"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>D.</given-names></name> <name><surname>Deng</surname> <given-names>L.</given-names></name></person-group> (<year>2010</year>). <article-title>Deep learning and its applications to signal and information processing [exploratory dsp]</article-title>. <source>IEEE Signal Process. Mag.</source> <volume>28</volume>, <fpage>145</fpage>&#x2013;<lpage>154</lpage>. doi: <pub-id pub-id-type="doi">10.1109/MSP.2010.939038</pub-id></citation></ref>
<ref id="ref120"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>Z.</given-names></name> <name><surname>Pantic</surname> <given-names>M.</given-names></name> <name><surname>Roisman</surname> <given-names>G. I.</given-names></name> <name><surname>Huang</surname> <given-names>T. S.</given-names></name></person-group> (<year>2008</year>). <article-title>A survey of affect recognition methods: audio, visual, and spontaneous expressions</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>31</volume>, <fpage>39</fpage>&#x2013;<lpage>58</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2008.52</pub-id></citation></ref>
<ref id="ref121"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>T.</given-names></name> <name><surname>Qin</surname> <given-names>R.-Z.</given-names></name> <name><surname>Dong</surname> <given-names>Q.-L.</given-names></name> <name><surname>Gao</surname> <given-names>W.</given-names></name> <name><surname>Xu</surname> <given-names>H.-R.</given-names></name> <name><surname>Hu</surname> <given-names>Z.-Y.</given-names></name></person-group> (<year>2017</year>). <article-title>Physiognomy: personality traits prediction by learning</article-title>. <source>Int. J. Autom. Comput.</source> <volume>14</volume>, <fpage>386</fpage>&#x2013;<lpage>395</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11633-017-1085-8</pub-id></citation></ref>
<ref id="ref122"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.-L.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Wei</surname> <given-names>X.-S.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Deep bimodal regression for apparent personality analysis</article-title>,&#x201D; in <source>European Conference on Computer Vision (Springer)</source>, <fpage>311</fpage>&#x2013;<lpage>324</lpage>.</citation></ref>
<ref id="ref123"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Tian</surname> <given-names>Q.</given-names></name></person-group> (<year>2019</year>). <article-title>Spontaneous speech emotion recognition using multiscale deep convolutional LSTM</article-title>. <source>IEEE Trans. Affect. Comput.</source> doi: <pub-id pub-id-type="doi">10.1109/TAFFC.2019.2947464</pub-id></citation></ref>
<ref id="ref124"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Shi</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Facial expression recognition via deep learning</article-title>. <source>IETE Tech. Rev.</source> <volume>32</volume>, <fpage>347</fpage>&#x2013;<lpage>355</lpage>. doi: <pub-id pub-id-type="doi">10.1080/02564602.2015.1017542</pub-id></citation></ref>
<ref id="ref125"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Ji</surname> <given-names>Q.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Bayesian graph convolution lstm for skeleton based action recognition</article-title>,&#x201D; in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, <fpage>6882</fpage>&#x2013;<lpage>6892</lpage>.</citation></ref>
</ref-list>
</back>
</article>