<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2022.863926</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Exploring Factor Structures Using Variational Autoencoder in Personality Research</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Huang</surname> <given-names>Yufei</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1879361/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhang</surname> <given-names>Jianqiu</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/154261/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Medicine, University of Pittsburgh School of Medicine</institution>, <addr-line>Pittsburgh, PA</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>University of Pittsburgh Medical Center Hillman Cancer Center</institution>, <addr-line>Pittsburgh, PA</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Electrical and Computer Engineering, The University of Texas</institution>, <addr-line>San Antonio, TX</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Alexander Robitzsch, IPN - Leibniz Institute for Science and Mathematics Education, Germany</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: David Goretzko, Ludwig Maximilian University of Munich, Germany; Andrew Cutler, Boston University, United States; Steffen Nestler, University of M&#x000FC;nster, Germany</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Jianqiu Zhang <email>michelle.zhang&#x00040;utsa.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Quantitative Psychology and Measurement, a section of the journal Frontiers in Psychology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>05</day>
<month>08</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>863926</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Huang and Zhang.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Huang and Zhang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>An accurate personality model is crucial to many research fields. Most personality models have been constructed using linear factor analysis (LFA). In this paper, we investigate if an effective deep learning tool for factor extraction, the Variational Autoencoder (VAE), can be applied to explore the factor structure of a set of personality variables. To compare VAE with LFA, we applied VAE to an International Personality Item Pool (IPIP) Big 5 dataset and an IPIP HEXACO (Humility-Honesty, Emotionality, Extroversion, Agreeableness, Conscientiousness, Openness) dataset. We found that LFA tends to break factors into ever smaller, yet still significant fractions, when the number of assumed latent factors increases, leading to the need to organize personality variables at the factor level and then the facet level. On the other hand, the factor structure returned by VAE is very stable and VAE only adds noise-like factors after significant factors are found as the number of assumed latent factors increases. VAE reported more stable factors by elevating some facets in the HEXACO scale to the factor level. Since this is a data-driven process that exhausts all stable and significant factors that can be found, it is not necessary to further conduct facet level analysis and it is anticipated that VAE will have broad applications in exploratory factor analysis in personality research.</p></abstract>
<kwd-group>
<kwd>non-linear factor analysis</kwd>
<kwd>variational auto encoder (VAE)</kwd>
<kwd>personality trait</kwd>
<kwd>artificial intelligence</kwd>
<kwd>Big 5 personality factors</kwd>
<kwd>HEXACO model of personality</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<counts>
<fig-count count="15"/>
<table-count count="6"/>
<equation-count count="6"/>
<ref-count count="85"/>
<page-count count="21"/>
<word-count count="12491"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Linear Factor Analysis (LFA) has enabled the discovery of the most popular personality models, including notably the Big 5 model (Fiske, <xref ref-type="bibr" rid="B31">1949</xref>; Norman, <xref ref-type="bibr" rid="B61">1963</xref>; Costa and McCrae, <xref ref-type="bibr" rid="B24">1992</xref>; Goldberg, <xref ref-type="bibr" rid="B36">1992</xref>) and the HEXACO model (Lee and Ashton, <xref ref-type="bibr" rid="B54">2004</xref>, <xref ref-type="bibr" rid="B55">2005</xref>), which have been extensively utilized to study a wide array of topics, such as personality disorder (Saulsman and Page, <xref ref-type="bibr" rid="B70">2004</xref>; Widiger and Lowe, <xref ref-type="bibr" rid="B81">2007</xref>), academic success (Ziegler et al., <xref ref-type="bibr" rid="B85">2010</xref>; Carthy et al., <xref ref-type="bibr" rid="B18">2014</xref>), leadership (Judge and Bono, <xref ref-type="bibr" rid="B47">2000</xref>; Hassan et al., <xref ref-type="bibr" rid="B40">2016</xref>), relationship satisfaction (O&#x00027;Meara and South, <xref ref-type="bibr" rid="B62">2019</xref>), job performance (Barrick and Mount, <xref ref-type="bibr" rid="B11">1991</xref>), education outcomes (Noftle and Robins, <xref ref-type="bibr" rid="B60">2007</xref>), and health outcomes (Jerram and Coleman, <xref ref-type="bibr" rid="B46">1999</xref>).</p>
<p>The root of applying LFA for the construction of personality models can be traced back to Galton&#x00027;s lexical hypothesis of personality (Galton, <xref ref-type="bibr" rid="B32">1884</xref>), which assumed that significant individual character differences could be discovered in language. Allport and Odbert applied the lexical approach to investigate personality-related dictionary words. They found approximately 4,500 terms that were considered descriptive of personality traits.</p>
<p>Before high-performance computing was available, Cattell first applied the grouped centroid method in factor analysis (Cattell, <xref ref-type="bibr" rid="B20">1943</xref>) to the list of traits generated by Allport and Odbert. He selected 171 from the list and developed a set of 35 to 40 clusters of words. He eventually settled on 16 personality factors (Cattell et al., <xref ref-type="bibr" rid="B21">1970</xref>) and made his data available to other researchers. After the arrival of high-performance computers, later researchers consistently found a five-factor model (Tupes and Christal, <xref ref-type="bibr" rid="B77">1961</xref>; Norman, <xref ref-type="bibr" rid="B61">1963</xref>). Through the years, the terms used for the five-factor model had changed, and finally, Goldberg coined the term &#x0201C;Big 5&#x0201D; personality model consisting of openness, conscientiousness, extraversion, agreeableness, and neuroticism (Goldberg, <xref ref-type="bibr" rid="B34">1981</xref>).</p>
<p>To eliminate the doubt that LFA methods may heavily influence the discovery process of the Big 5 model, Goldberg tested five methods of factor extraction (principal components, principal factors, alpha-factoring, image-factoring, and maximum-likelihood procedures), each rotated by an orthogonal (varimax) and an oblique (oblimin) algorithm (Goldberg, <xref ref-type="bibr" rid="B35">1990</xref>). He found that procedural variations do not change the five-factor structure, and the factor scores across different methods are highly congruent. In addition, Goldberg and Saucier investigated the relationship between person-descriptive adjective clusters and the Big 5 traits. They concluded that mostly all personality-relevant clusters are not &#x0201C;beyond the Big 5&#x0201D; (Goldberg and Saucier, <xref ref-type="bibr" rid="B38">1998</xref>). Furthermore, it was shown that the Big 5 model is replicable across cultures (McCrae et al., <xref ref-type="bibr" rid="B58">1998</xref>). These findings have cemented the unrivaled popularity of the Big 5 model (Feher and Vernon, <xref ref-type="bibr" rid="B30">2021</xref>).</p>
<p>The other side of Big 5&#x00027;s sustained popularity is that few advances have been made in personality model development (Feher and Vernon, <xref ref-type="bibr" rid="B30">2021</xref>). One exception is the HEXACO model, which applied LFA and the lexical approach to several languages worldwide (Lee and Ashton, <xref ref-type="bibr" rid="B54">2004</xref>, <xref ref-type="bibr" rid="B55">2005</xref>; Ashton and Lee, <xref ref-type="bibr" rid="B3">2007</xref>). A sixth personality factor, Honesty-Humility, consistently showed up in cross-cultural studies, which address the fairness and modesty aspects of personality (Ashton and Lee, <xref ref-type="bibr" rid="B4">2008a</xref>,<xref ref-type="bibr" rid="B5">b</xref>). The underlying meaning of some of the factors (agreeableness and emotionality) differs slightly from the Big 5 model (Ashton et al., <xref ref-type="bibr" rid="B6">2014</xref>). The HEXACO model is highly correlated to most existing narrow trait models, such as the Dark Triad models outside of the Big 5 model (Lee and Ashton, <xref ref-type="bibr" rid="B55">2005</xref>; Ashton and Lee, <xref ref-type="bibr" rid="B3">2007</xref>; De Vries et al., <xref ref-type="bibr" rid="B26">2009</xref>).</p>
<p>A critical consideration in factor analysis is how many factors should be extracted. For example, in the IPIP HEXACO dataset that we tested, there could be 8 factors on the scree plot. When we required the eigenvalue to be greater than one, there were 37 factors. The large discrepancy in these criteria makes it impossible to know how many factors should be extracted without examining stability across multiple datasets. Although simulated data can be used to determine the threshold on eigenvalues as in parallel analysis (Horn, <xref ref-type="bibr" rid="B42">1965</xref>), since we do not know the actual distribution of the latent factors and the functions that transforms these factors into the measured personality variables, we cannot simulate a ground truth dataset with a known number of factors in VAE. We observed that as the number of assumed latent factors increases in LFA, bigger factors that contain many personality variables tend to break down into smaller yet significant factors and the fractioning process will not stop until a very large number of latent factors. Yet, these smaller factors are not stable. As a result, established personality models only report a small number of stable factors and only six in the case of HEXACO. However, it has been found that facet-level information must be incorporated in applications (Reynolds and Clark, <xref ref-type="bibr" rid="B66">2001</xref>; Samuel and Widiger, <xref ref-type="bibr" rid="B69">2008</xref>). This indicates that factor-level information is insufficient, yet facet-level research is not entirely data-driven (Goldberg, <xref ref-type="bibr" rid="B37">1999</xref>). As a result, personality scales get frequently revised which is costly for data collection and research.</p>
<p>We can view LFA as a type of unsupervised machine learning (ML) method (Chauhan and Singh, <xref ref-type="bibr" rid="B22">2018</xref>) and we can treat the Big 5 or HEXACO traits as latent generative factors that can be transformed to construct the observable personality variables. We can search in the broader context of unsupervised ML to look for a suitable tool for personality model construction.</p>
<p>In ML, the most recent advances have been driven by Deep Learning (DL) (LeCun et al., <xref ref-type="bibr" rid="B53">2015</xref>; Sengupta et al., <xref ref-type="bibr" rid="B72">2020</xref>). DL methods employ artificial neural networks capable of approximating every function under mild assumptions (Cybenko, <xref ref-type="bibr" rid="B25">1989</xref>; Hornik, <xref ref-type="bibr" rid="B43">1991</xref>). DL had enabled phenomenal technological advancements in computer vision (Krizhevsky et al., <xref ref-type="bibr" rid="B51">2017</xref>), natural language processing (Devlin et al., <xref ref-type="bibr" rid="B27">2018</xref>), autonomous vehicles (Sun et al., <xref ref-type="bibr" rid="B74">2020</xref>), personalization, and recommender systems (Jacobson et al., <xref ref-type="bibr" rid="B44">2016</xref>; Batmaz et al., <xref ref-type="bibr" rid="B12">2019</xref>; Bobadilla et al., <xref ref-type="bibr" rid="B15">2020</xref>), and live translation of languages (Castelvecchi, <xref ref-type="bibr" rid="B19">2016</xref>).</p>
<p>Given such a promise, we have also seen significant growth in applying DL methods for personality traits detection based on data gathered from social media platforms (Liu and Zhu, <xref ref-type="bibr" rid="B56">2016</xref>; Yu and Markov, <xref ref-type="bibr" rid="B82">2017</xref>; Kumar and Gavrilova, <xref ref-type="bibr" rid="B52">2019</xref>; Ahmad et al., <xref ref-type="bibr" rid="B1">2020</xref>; Salminen et al., <xref ref-type="bibr" rid="B68">2020</xref>), vision and language samples (Eddine Bekhouche et al., <xref ref-type="bibr" rid="B28">2017</xref>; Chhabra et al., <xref ref-type="bibr" rid="B23">2019</xref>; Rodriguez et al., <xref ref-type="bibr" rid="B67">2019</xref>; Kim et al., <xref ref-type="bibr" rid="B48">2020</xref>; Ren et al., <xref ref-type="bibr" rid="B65">2021</xref>), handwriting samples (Elngar et al., <xref ref-type="bibr" rid="B29">2020</xref>; Remaida et al., <xref ref-type="bibr" rid="B64">2020</xref>), and mobile-sensing data (Baumeister and Montag, <xref ref-type="bibr" rid="B13">2019</xref>; Spathis et al., <xref ref-type="bibr" rid="B73">2019</xref>). In most of these studies, factors from the Big 5 model are used as labels in the training datasets such that neural networks can be trained to predict the Big 5 traits (Azucar et al., <xref ref-type="bibr" rid="B9">2018</xref>; Bhavya et al., <xref ref-type="bibr" rid="B14">2020</xref>; Mehta et al., <xref ref-type="bibr" rid="B59">2020</xref>; Ren et al., <xref ref-type="bibr" rid="B65">2021</xref>).</p>
<p>Despite these advances, the extent to which DL methods are used for personality model construction has not been extensively conducted. It has motivated us to look for a DL-based non-linear factor analysis tool. In this regard, variational autoencoder (VAE) (Kingma and Welling, <xref ref-type="bibr" rid="B49">2013</xref>; Lopez-Alvis et al., <xref ref-type="bibr" rid="B57">2020</xref>) is a state-of-the-art DL method for unsupervised representation learning.</p>
<p>The first versions of autoencoders emerged over two decades ago (Bourlard and Kamp, <xref ref-type="bibr" rid="B16">1988</xref>; Zemel and Hinton, <xref ref-type="bibr" rid="B83">1993</xref>) and they were primarily used for dimensional reduction initially. They consist of an encoding artificial neural network, which outputs a latent representation of the input data, and a decoding neural network that tries to accurately reconstruct the input data from its latent representation. Very shallow versions of autoencoders (with a small number of middle layer nodes) can reproduce the results of principal component analysis (Baldi and Hornik, <xref ref-type="bibr" rid="B10">1989</xref>).</p>
<p>The VAE is motivated by the more general problem of &#x0201C;obtaining a joint distribution over all input variables through learning a generative model, which simulates how the data is generated in the real world&#x0201D; (Kingma and Welling, <xref ref-type="bibr" rid="B50">2019</xref>). It was designed to find a set of &#x0201C;disentangled, semantically meaningful, statistically independent and causal factors of variation in data,&#x0201D; as the original inventor of VAE described it. VAE differs from traditional autoencoders by imposing restrictions on the distribution of latent variables, which allows it to find independent latent variables (Kingma and Welling, <xref ref-type="bibr" rid="B49">2013</xref>). By taking the sampling step that treats the joint posterior distribution of the latent variables as independent, the algorithm is forced to converge to solutions, in which the latent variables are almost independent. Previous empirical evidence (Burgess et al., <xref ref-type="bibr" rid="B17">2018</xref>) shows that in image processing, these factors can often be tied to an &#x0201C;interpretable&#x0201D; factor. VAE and its variants (Ainsworth et al., <xref ref-type="bibr" rid="B2">2018</xref>; Zhou and Wei, <xref ref-type="bibr" rid="B84">2020</xref>) are more &#x0201C;interpretable&#x0201D; compared to common deep neural networks in this sense. Among various variants of VAE, we have employed the original VAE, which can be considered a special case of beta-VAE (Higgins et al., <xref ref-type="bibr" rid="B41">2016</xref>) because VAE performed the best on the tested datasets.</p>
<p>The VAE does not assume that the observed variables are linear combinations of latent factors plus unique factors as in LFA. Compared to PCA, it also drops the assumption that the generating function of the observed variables is linear. VAE only assumes that the latent variables are Gaussian and independent. In this sense, VAE is closer to PCA than LFA.</p>
<p>Given that the deep neural networks in VAE can be configured to simulate non-linear functions (Cybenko, <xref ref-type="bibr" rid="B25">1989</xref>; Hornik, <xref ref-type="bibr" rid="B43">1991</xref>), it has found applications in many areas that require non-linear modeling of the generative process. For example, it has been applied to non-linear channel equalization (Avi and Burshtein, <xref ref-type="bibr" rid="B8">2020</xref>), 3D mesh models transformation in computer animation (Tan et al., <xref ref-type="bibr" rid="B75">2018</xref>), and fault detection in complex non-linear process controls (Wang et al., <xref ref-type="bibr" rid="B80">2019</xref>). We anticipate that VAE can be applied to find latent and independent personality factors while assuming a non-linear underlying psychological process.</p>
<p>Urban and Bauer (<xref ref-type="bibr" rid="B78">2021</xref>) first introduced a deep learning-based variational inference (VI) algorithm that applies an importance-weighted autoencoder (IWAE) for exploratory item factor analysis (IFA) that is computationally efficient even in large datasets with many latent factors. IWAE can recover the 5-factor structure of the Big five model based on a large Big5 dataset. IWAE is very similar to our proposed VAE algorithm except that it sets the output layer to predict the log-likelihood probability of all possible responses on a Likert scale. In contrast, in VAE, we set the output layer to produce a continuous variable. Although it has been established that IWAE-like algorithms can be used for exploring the factor structures of a set of personality variables, however, there are still many unanswered questions. We need to develop new performance measures and factor extracting guidelines to compare the difference between VAE and LFA because VAE does not assume linear data models anymore. Specifically, we need to: (i) Select a stable set of factors across multiple VAE runs; (ii) Compare the accuracy of VAE generated personality models to LFA generated models; (iii) Develop a method for inspecting factor-personality variable association because we cannot rely on factor loadings as in LFA; and (iv) Study the stability of the VAE-generated model across different datasets and regions.</p>
<p>We hypothesize that VAE can do the following: (1) generate personality models that have higher correlations between the input and reconstructed personality variables than LFA, and (2) discover more stable factors than LFA.</p>
<p>We are aware of the limitations of self-reported data in generating useful personality models. However, this research is meant to establish the validity of using VAE as a replacement for LFA for exploratory factor analysis. Due to the scope and complexity involved in combining self-reports and observer reports, we plan to combine both types of data and construct useful personality models using VAE in future research.</p>
<sec>
<title>Datasets</title>
<p>In this study, we want to compare the performance of VAE-generated models to that of LFA-generated models. We selected two datasets collected based on the two most popular LFA constructed models, the Big 5 and the HEXACO models. Note that this is an initial study on the applicability of VAE to personality model analysis.</p>
</sec>
<sec>
<title>The International Personality Inventory Pool (IPIP) Big 5 Dataset</title>
<p>The IPIP Big 5 factor markers consist of a 50 or a 100-item inventory(Goldberg and Others 2001). We used the 50-item version consisting of 10 items for each of the Big 5 personality factors: Extraversion (E), Agreeableness (A), Conscientiousness (C), Neuroticism (N), and Openness/Intellect (I). Each item is given in a sentence form (e.g., &#x0201C;I am the life of the party&#x0201D;). Participants were requested to read each of the 50 items and then rate on a 5-point scale (from strongly disagree to strongly agree). The dataset was collected through an online questionnaire downloadable at the Open-Source Psychometrics Project (Goettfert and Kriner, n.d.). The dataset contains 19,719 samples, and we have used all samples in our study. The alpha reliability of the factors ranged from 0.80 to 0.89, and the mean and the standard deviation (SD) of the factors are consistent with previous publications (Costa and McCrae, <xref ref-type="bibr" rid="B24">1992</xref>; Goldberg, <xref ref-type="bibr" rid="B36">1992</xref>).</p>
</sec>
<sec>
<title>The IPIP HEXACO Dataset</title>
<p>We downloaded an IPIP HEXACO dataset collected from a questionnaire that measures 240 personality variables from the Open-Source Psychometrics Project website (Goettfert and Kriner, n.d.). The IPIP HEXACO inventory was constructed by correlating all 2036 IPIP items with the 24 HEXACO-Personality Inventory (PI) facet scales (Lee and Ashton, <xref ref-type="bibr" rid="B54">2004</xref>): Honesty-Humility (H) with facets: Sincerity (HSinc), Fairness (HFair), Greed (HGree), Avoidance (HAvoi), Modesty (HMode); Emotionality (E) with facets: Fearfulness (EFear), Anxiety (EAnxi), Dependence (EDepe), Sentimentality (ESent); Extraversion (X) with facets: Social Self-Esteem (XExper), Social Boldness (XSocB), Sociability (XSoci), Liveliness (XLive); Agreeableness (A) with facets Forgivingness (AForg), Gentleness (Agent), Flexibility (AFlex), Patience (APati); Conscientiousness (C) with facets: Organization (COrga), Diligence (CDili), Perfectionism (CPerf), Prudence (CPrud); Openness to Experience (O) with facets: Aesthetic Appreciation (OAesA), Inquisitiveness (OInqu), Creativity (OCrea), and Unconventionality (OUnco).</p>
<p>Within each of these 24 groups of IPIP items, the 10 personality variables showing the highest absolute correlations with their corresponding HEXACO-PI scale were selected. The resulting set of 24 IPIP&#x02014;HEXACO scales showed alpha reliabilities ranging from 0.73 to 0.88 with a mean of 0.81. Some personality variables were subsequently adjusted to reduce the correlation between Agreeableness and Honesty-Humility items (Ashton et al., <xref ref-type="bibr" rid="B7">2007</xref>).</p>
<p>The IPIP HEXACO dataset contained 22,786 samples. The 240 personality variables were rated on a seven-point scale (1 = strongly disagree, 2 = disagree, 3 = slightly disagree, 4 = neutral, 5 = slightly agree, 6 = agree, and 7 = strongly agree). We kept samples that answered 7 on both verification questions 1 and 2, which were administered at the beginning and the end of the test to ensure that questionnaire takers understood the test and answered all questions as accurately as possible. While lowering the threshold on the validation questions would admit more samples, it resulted in few performance changes. After this filtering process, a total of 18,779 samples were used in our analysis.</p>
</sec>
</sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<sec>
<title>Analytical Procedure</title>
<p>Our analytical procedure follows standard protocols in machine learning research. All code is made available on OSF (<ext-link ext-link-type="uri" xlink:href="https://osf.io/6b3w/">https://osf.io/6b3w/</ext-link>).</p>
</sec>
<sec>
<title>Data Preprocessing</title>
<p>For all datasets used in the studies, scores from each of the questionnaires are scaled by subtracting the mean and dividing by the SD to shift the distribution to have a mean of zero and a standard deviation of one. This pre-processing step is performed separately for the training and the testing dataset before further processing. Missing values are set to zero after scaling.</p>
</sec>
<sec>
<title>Training VAE Inference and Generative Models</title>
<p>The VAE is designed to learn interpretable non-linear generative factors. A VAE model comprises two independently parameterized components: an inference model (the encoder) that maps the inputs to a latent variable vector <italic><bold>z</bold></italic>, and a generative model (the decoder) that decodes the latent variable vector back into the original data space. These two components mirror each other with a shared bottleneck layer, with the fewest nodes representing the latent generative factors. There could be several hidden middle layers between the bottleneck and the input layers. An illustration of a VAE model is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>An example of a VAE with 100 hidden middle layer nodes and 8 bottleneck layer nodes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0001.tif"/>
</fig>
<p>The inference model estimates the posterior distribution of the latent factors in the bottleneck layer, which is assumed to be independent Gaussian. Consequently, the bottleneck layer consists of a vector of the means <inline-formula><mml:math id="M1"><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BC;</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula> and a vector of the standard deviations <inline-formula><mml:math id="M2"><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003C3;</mml:mi></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula> of the posterior distribution. Then, samples drawn from the posterior distribution <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>z</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>n</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are passed to the generative model to reconstruct the original input data.</p>
<p>Input data <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represent an <italic>M</italic> &#x000D7; 1 vector of personality variable scores from sample <italic>n</italic>, and <italic>x</italic><sub><italic>mn</italic></sub> denotes the <italic>m</italic>th personality variable. The <bold>x</bold><sub><italic>n</italic></sub> is assumed to follow a multivariate independent Gaussian distribution, and its mean and variances are modeled as a function of an <italic>d</italic>-dimension latent representation of personality traits <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> by an encoder neural network <italic>D</italic><sub>&#x003B8;</sub>, where <italic><bold>&#x003B8;</bold></italic> is a vector of encoder weights. Then the likelihood function of the input variable <bold>x</bold><sub><italic>n</italic></sub> can be defined as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x00398;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x00398; &#x0003D; {&#x003B8;,&#x003D5;} is the combined vector of encoder and decoder weights. Since <bold>x</bold><sub><italic>n</italic></sub> only depends on the encoder weights, the model in (1) omitted decoder weights in the list of dependent variables. In VAE, <bold>z</bold><sub><italic>n</italic></sub> is commonly assumed to follow a prior distribution, which is the multivariate standard normal, i.e., <inline-formula><mml:math id="M7"><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>I</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:math></inline-formula> with <bold>I</bold><sub><bold>d</bold></sub> being a <italic>d</italic> &#x000D7; <italic>d</italic> identity matrix.</p>
<p>The goal of training or inference is to compute the maximum likelihood estimate of &#x00398;</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M8"><mml:mrow><mml:msub><mml:mrow><mml:mover accent='true'><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>argmax</mml:mtext></mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0007C;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mrow><mml:mtext>argmax</mml:mtext></mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mi>&#x02112;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle></mml:mrow></mml:math></disp-formula>
<p>where <italic>p</italic>(<bold>x</bold><sub><italic>n</italic></sub>|<bold>&#x00398;</bold>) &#x0003D; &#x0222B;<italic>p</italic>(<bold>x</bold><sub><italic>n</italic></sub>|<bold>z</bold><sub><italic>n</italic></sub>, <bold>&#x00398;</bold>)<italic>p</italic>(<bold>z</bold><sub><italic>n</italic></sub>)<italic>d</italic><bold>z</bold><sub><italic>n</italic></sub> is the marginal likelihood which is analytically intractable but can be lower bounded by the evidence lower bound (ELBO) <italic>L</italic>(&#x00398;),</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M9"><mml:mrow><mml:mi>&#x02112;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0007C;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0007C;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>[</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0007C;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02225;</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>q</italic>(<bold>z</bold><sub><italic>n</italic></sub>|<bold>x</bold><sub><italic>n</italic></sub>, <bold>&#x00398;</bold>) is an approximate to the intractable posterior distribution <italic>p</italic>(<bold>z</bold><sub><italic>n</italic></sub>|<bold>x</bold><sub><italic>n</italic></sub>, <bold>&#x00398;</bold>), and <italic>D</italic><sub><italic>KL</italic></sub>[<italic>q</italic>(<bold>z</bold><sub><italic>n</italic></sub>|<bold>x</bold><sub><italic>n</italic></sub>, <bold>&#x00398;</bold>) ||<italic>p</italic>(<bold>z</bold><sub><italic>n</italic></sub>)] measures the Kullback-Leibler distance (Walters-Williams and Li, <xref ref-type="bibr" rid="B79">2010</xref>) between the approximated posterior distribution <italic>q</italic>(<bold>z</bold><sub><italic>n</italic></sub>|<bold>x</bold><sub><italic>n</italic></sub>, <bold>&#x00398;</bold>) and the prior distribution <italic>p</italic>(<bold>z</bold><sub><italic>n</italic></sub>). To make the variational inference tractable, <italic>q</italic>(<bold>z</bold><sub><italic>n</italic></sub>|<bold>x</bold><sub><italic>n</italic></sub>, <bold>&#x00398;</bold>) is assumed in most cases as a multivariate Gaussian,</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mi>&#x00398;</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>&#x000A0;</mml:mtext><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>whose means and variances <inline-formula><mml:math id="M12"><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>n</mml:mi></mml:mstyle></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula> are given by a decoder network <italic>E</italic><sub>&#x003D5;</sub> applied to <italic>x</italic><sub><italic>n</italic></sub> as</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:mi>z</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003D5; is the vector of the unknown decoder weights. Because of the approximation by <inline-formula><mml:math id="M14"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x00398;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> in Equation (2) and the introduction of the decoder network in Equation (4), the model parameters to be estimated become &#x00398; &#x0003D; {&#x003B8;,&#x003D5;}. Optimization of <inline-formula><mml:math id="M15"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x00398;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> in Equation (2) is computed by the stochastic gradient descent algorithm, where the gradient is calculated by backpropagation. Note that the first part in Equation (3) will be proportional to the mean square error (MSE) between the input <bold>x</bold><sub><italic>n</italic></sub> and the reconstructed scores <bold>v</bold><sub><italic>x<sub>n</sub></italic></sub> when we assume that the likelihood function of the personality variables follows an independent Gaussian distribution as in Equation (1). In VAE, the missing values are excluded when calculating the MSE.</p>
<p>VAE training is carried out with the TensorFlow machine learning module imported to Python (Jason, <xref ref-type="bibr" rid="B45">2016</xref>).</p>
</sec>
<sec>
<title>Model Accuracy Metric and VAE Model Selection</title>
<p>We calculate the Person Correlation between the input personality variables and the reconstructed ones (input-reconstruction correlation) as the performance metric. The R2 statistics can be calculated sample-wise, i.e., between <bold>x</bold><sub><italic>n</italic></sub> &#x0003D; [<italic>x</italic><sub>1<italic>n</italic></sub>, <italic>x</italic><sub>2<italic>n</italic></sub>, &#x022EF;<italic>x</italic><sub><italic>Mn</italic></sub>] and <bold>v</bold><sub><italic>n</italic></sub> &#x0003D; [<italic>v</italic><sub>1<italic>n</italic></sub>, <italic>v</italic><sub>2<italic>n</italic></sub>, &#x022EF;<italic>v</italic><sub><italic>Mn</italic></sub>] <italic>n</italic> &#x02208; (1, <italic>N</italic>), or variable-wise, i.e. between <inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>N</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>N</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> for all variable indices <italic>m</italic> &#x02208; (1, <italic>M</italic>). Note that the reconstructed variables are assumed to be the sum of input variables plus independent noise after the input is put through an encoder and decoder function. Other commonly used performance measures in LFA, such as the communality or the percentage of variance represented by the selected factors, are not applicable in the context of VAE because these measures assume a linear data model between the factors and the input data, while the encoder and decoder in VAE are not linear.</p>
<p>In the context of LFA, to calculate the Person Correlation between the input and the reconstructed personality variable scores, we can first reconstruct the score for the <italic>m</italic>th personality variable in the <italic>n</italic>th sample, <italic>x</italic><sub><italic>mn</italic></sub> based on the data model in LFA, which is a linear combination of <italic>d</italic> factors <inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>n</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of the <italic>n</italic>th sample:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003B5;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M19"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>l</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mtext>&#x000A0;</mml:mtext></mml:math></inline-formula> is the vector of factor loadings in the <italic>m</italic>th personality variable, &#x003B5;<sub><italic>m</italic></sub> is the item specific factor, and &#x003C3;<sub><italic>mn</italic></sub> is the observation noise for the <italic>m</italic>th personality variable in the <italic>n</italic>th sample. Then the reconstructed score becomes <italic>v</italic><sub><italic>mn</italic></sub> = <inline-formula><mml:math id="M20"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>l</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>n</mml:mtext></mml:mstyle></mml:mrow></mml:msub></mml:math></inline-formula> if we assume that both &#x003B5;<sub><italic>m</italic></sub> <italic>and &#x003C3;</italic><sub><italic>mn</italic></sub> have zero means. Then the correlation between the input and the reconstructed personality variables can be calculated either sample-wise or variable-wise for LFA.</p>
<p>The <italic>d</italic> latent factors <inline-formula><mml:math id="M21"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>z</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>n</mml:mtext></mml:mstyle></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are calculated both by the Thurstone and the Bartlett (Grice, <xref ref-type="bibr" rid="B39">2001</xref>) methods. We compared the results of reconstruction and found that the two methods did not make a significant difference on input-reconstruction correlations in the tested datasets.</p>
<p>Note that in the machine learning community, the commonly used measure is Coefficient of Determinate (R2 statistics), which is defined as one minus the total error variance divided by the total sample variance. In the context of VAE, it can be viewed like communality in LFA, which measures how much variance has been explained.</p>
<p>In VAE, the <italic>m</italic>th personality variable&#x00027;s loadings on the <italic>i</italic>th factor, <italic>l</italic><sub><italic>mi</italic></sub>, is estimated by calculating the correlation between the latent factor&#x00027;s mean vector <inline-formula><mml:math id="M22"><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo>&#x000A0;</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mi>i</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mrow><mml:mi>z</mml:mi><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula>, with the <italic>m</italic>th reconstructed input personality variable <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>v</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>N</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> over all <italic>N</italic> training samples. Note that this loading is calculated for the purpose of finding a rotation of the factors such that the latent factors from different VAE runs can be aligned. We used the reconstructed inputs for loading calculation because the effect of noise has been removed after the reconstruction process and <bold>&#x003BC;</bold><sub><bold>zi</bold></sub> represents the maximum posterior (MAP) estimation of the latent variables.</p>
<p>In LFA, the <italic>m</italic>th personality variable&#x00027;s loadings on the <italic>i</italic>th factor, <italic>l</italic><sub><italic>mi</italic></sub>, is estimated by calculating the correlation between the vector of latent factor <inline-formula><mml:math id="M24"><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>z</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>i</mml:mi></mml:mstyle></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>[</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>N</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> with the vector of the <italic>m</italic>th input personality variable <inline-formula><mml:math id="M25"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>1</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>N</mml:mi><mml:mtext>&#x000A0;</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> over all <italic>N</italic> training samples. The calculation is performed by the factor analyzer in Python. Since we do not compare VAE and LFA in factor loadings, the difference in the calculation procedure of these loading factors is not consequential for the interpretation of the results.</p>
</sec>
<sec>
<title>VAE Factor Stability Analysis Based on Congruence Scores Over Multiple Runs</title>
<p>For LFA, noise factors beyond the first 5 did not appear consistently in different studies (Goldberg, <xref ref-type="bibr" rid="B35">1990</xref>). Factors from different LFA runs could be rotated or perturbed due to variations, which can be attributed either to the analyzing methods or the datasets. Similarly, we also exclude noise factors and determine stable factors to be included in VAE constructed models.</p>
<p>The VAE employs a stochastic gradient descent algorithm that may not converge to the same solution across multiple runs. Such variations may introduce noise factors in addition to the ones introduced by the variations in the training datasets. To make sure that we only include stable factors in the final personality model, we extended the concept of congruence coefficient introduced by Goldberg for studying factor stability across different methods (Goldberg, <xref ref-type="bibr" rid="B35">1990</xref>) and defined the congruence score between any two factors <italic>i</italic> and <italic>j</italic> in run <italic>r1</italic> and run <italic>r2</italic> as: <inline-formula><mml:math id="M26"><mml:msubsup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mo>,</mml:mo><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M27"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>l</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the <italic>i</italic>th factor loadings on all M personality variables, and <italic>Corr()</italic> represents the Pearson correlation function.</p>
<p>In each VAE run, the latent factors are supposed to be independent and for each factor, only one factor from another run can be matched with it with a high congruence score. However, VAE often returns correlated factors within a run when we set the number of bottleneck layer nodes higher than the actual number of stable factors. In such cases, multiple factors from the same run may be clustered together with a given factor from another run. We observed that falsely matched factors generally have lower congruence scores than the true matching factors. To prevent the clustering algorithm from falsely matching factors, we set up a threshold and removed congruence scores below the threshold.</p>
<p>After the filtering step, the Leiden clustering algorithm (Traag et al., <xref ref-type="bibr" rid="B76">2019</xref>) was applied to cluster factors from different runs. The Leiden clustering algorithm was developed to improve the Louvain algorithm. The Leiden algorithm allows both splitting and merging. The Leiden algorithm guarantees that clusters are well-connected and the clusters it finds are not too far from optimal.</p>
<p>All matching factors from all runs will be reported as clusters by the Leiden clustering algorithm. Then, we manually validated the clusters returned by the Leiden algorithm by inspecting the personality variables associated with each factor. To ensure that factors from different runs can be aligned together, we performed varimax rotation on the factor loadings before they were used to calculate the congruence scores.</p>
<p>To apply the clustering algorithm, d factor loading vectors from all R runs are retrieved for calculating a congruence score/factor loading correlation matrix with (dR) <sup>&#x0002A;</sup> (dR) elements. Then, the elements in the congruence score matrix below the threshold will be set to zero before it is fed into the Leiden clustering algorithms.</p>
<p>In practice, the threshold on the congruence scores was gradually raised until factors from the same run cannot be clustered together. We define a factor as stable when it can be identified in each VAE run. The average of the congruence scores within a cluster is used to represent the overall factor congruence.</p>
<p>We determine the final number of stable factors by increasing the number of bottleneck layer nodes until no more stable factors emerge.</p>
</sec>
<sec>
<title>Determining Factor-Variable Associations in VAE</title>
<p>After VAE identified a set of latent factors, it is important to understand how these factors can be interpreted. In LFA, this is done by inspecting the factor loadings of personality variables. However, since factor loadings are calculated by assuming a linear relationship between the factors and personality variables, we cannot apply the same method for identifying factor-variable association in VAE. We propose to inspect the reduction of correlations between the input and reconstructed personality variables. To calculate the reduction, we first set the investigated latent factor to zero, while the rest of the factors are fed into the VAE decoder to reconstruct the personality variables across all samples. Then, we calculate the correlations between the input and reconstructed variables after muting the investigated factor and compare it to the correlations before muting the factor. Personality variables with input-reconstruction correlation reduction above a threshold are associated with a latent factor. Since this method does not make any assumption on the linearity of the encoding and the decoding process, it is suitable for inspecting factor-variable association in both VAE and LFA.</p>
</sec>
<sec>
<title>Cross-Regional Study</title>
<p>To see whether the VAE returned factor model is valid across different regions, we trained a VAE model based on 80% of the samples from North America in the HEXACO analysis. Then, we estimated the factor statistics, especially the mean, the standard deviation, and the zero-order correlations of the derived factors to see if the factor model varies across different regions.</p>
</sec>
<sec>
<title>Linear Factor Analysis (LFA)</title>
<p>The LFA was conducted using the python function FactorAnalyzer imported from the sklearn.decomposition module (Persson and Khojasteh, <xref ref-type="bibr" rid="B63">2021</xref>). Since previous research indicated that the selection of the LFA method would not make a significant difference in LFA (Goldberg, <xref ref-type="bibr" rid="B35">1990</xref>), we used the principal factor method included in the package and selected &#x02018;varimax&#x00027; as the rotation method. The same scaled and normalized input data were used for both LFA and VAE.</p>
</sec>
<sec>
<title>Testing and Training Dataset Separation</title>
<p>We separated each input dataset into a training and a testing dataset by an 80&#x02013;20% ratio. The training dataset was used for training the weights in the encoder and the decoder in VAE. It was also used to train the factor analyzer in LFA, which generated the factor loading matrix, a factor transformer (by default the Thurstone method) that can be used to calculate the factor scores, and other statistics, such as the eigenvalues of the data. We have also trained the weights for calculating the factor scores in the Bartlett method.</p>
<p>In VAE, the testing dataset was used to estimate the latent factors using the decoder, the reconstructed personality variables by using the encoder and the decoder, the input-reconstruction correlations, the congruence scores in factor stability analysis, and input-reconstruction correlation reduction in factor-variable association analysis. In LFA, the testing dataset was used to calculate the factor scores, the reconstructed personality variables, and the rest of the measures as in VAE.</p>
<p>In VAE training, the training dataset was further split into a training and validation dataset to ensure that the algorithm does not over-fit. The split ratio is 80&#x02013;20%.</p>
<p>Complete separation between the training and the testing datasets is ensured. This procedure reduces the risk of overfitting because all testing was conducted on samples not included in the model construction process. Ten VAE models were trained using 10 training datasets in all analyses.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<sec>
<title>Analysis of the IPIP Big 5 Dataset</title>
<sec>
<title>LFA Exploratory Factor Analysis</title>
<p>We first conducted a LFA exploratory analysis by plotting the eigenvalues of the dataset. The scree plot in <xref ref-type="fig" rid="F2">Figure 2</xref> shows that there are 7 factors with eigenvalues greater than 1. If we look for the factors left to the elbow point, there are 5 factors in the data. Past research found that there are 5 stable factors (Costa and McCrae, <xref ref-type="bibr" rid="B24">1992</xref>; Goldberg, <xref ref-type="bibr" rid="B36">1992</xref>) that emerge from run to run. We can see that the results from the eigenvalue analysis and the scree plot are different and fewer factors can be generalized.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Exploratory factor analysis of the IPIP Big 5 dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0002.tif"/>
</fig>
<p>We then analyzed the factor-variable associations by inspecting the input-reconstruction correlation reduction as we increased the number of factors in LFA. An example result is shown in <xref ref-type="fig" rid="F3">Figure 3</xref> when there are 10 factors. We can see that the first 4 four factors of Extroversion, Neuroticism, Agreeableness, and Conscientiousness stayed the same as the original scale. However, the Openness factor fractionated in to 3 smaller factors, which makes the total number of factors to be 7.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Input-reconstruction correlation reduction when assuming 10 Latent Factors in LFA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0003.tif"/>
</fig>
</sec>
<sec>
<title>VAE Analysis of the IPIP Big 5 Dataset</title>
<sec>
<title>Model Parameter Selection</title>
<p>Model accuracy is measured by the variable-wise input-reconstruction correlations. <xref ref-type="table" rid="T1">Table 1</xref> shows the results when using different numbers of hidden middle layers, different numbers of hidden middle layer nodes, and different numbers of bottleneck layer nodes. Note that the activation function of the last layer of the generative model is linear so that the output can have negative results. The activation function of bottleneck layer nodes must also be linear because they are assumed to be Gaussian variables. The Relu activation function was used for the rest of the layers.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>The Mean (Std) of variable-wise input-reconstruction correlations.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>&#x00023; Bottleneck nodes-1 layer</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>&#x00023; Bottleneck nodes-two layers</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>&#x00023; Mid-Layer Nodes</bold></th>
<th valign="top" align="center"><bold>8</bold></th>
<th valign="top" align="center"><bold>10</bold></th>
<th valign="top" align="center"><bold>8</bold></th>
<th valign="top" align="center"><bold>10</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">50</td>
<td valign="top" align="center">0.694 (0.066)</td>
<td valign="top" align="center">0.714 (0.062)</td>
<td valign="top" align="center">0.670 (0.076)</td>
<td valign="top" align="center">0.671 (0.075)</td>
</tr>
<tr>
<td valign="top" align="left">100</td>
<td valign="top" align="center">0.728 (0.061)</td>
<td valign="top" align="center">0.748 (0.057)</td>
<td valign="top" align="center">0.715 (0.058)</td>
<td valign="top" align="center">0.716 (0.059)</td>
</tr>
<tr>
<td valign="top" align="left">200</td>
<td valign="top" align="center">0.731 (0.058)</td>
<td valign="top" align="center">0.752 (0.050)</td>
<td valign="top" align="center">0.730 (0.057)</td>
<td valign="top" align="center">0.750 (0.048)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N = 3,944 samples from the testing dataset. When two middle layers are used, the number of nodes used in the second mid-layer is half of that in the first middle layer</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Inspecting how the mean of the input-reconstruction correlations changes as we increase the number of middle layer nodes from 50 to 200 in <xref ref-type="table" rid="T1">Table 1</xref>, we can see that increasing the number of middle layer nodes constantly improves the performance, although increasing it beyond 100 nodes offered little performance gain. Further increasing the number of nodes does not change the factor structure or improve the correlations between the input and reconstructed variables significantly. Meanwhile the computational cost will grow significantly. Also, as the model becomes more complex, the required number of samples for appropriately estimating the weights in the deep neural network will increase. Another point is that in <xref ref-type="table" rid="T1">Table 1</xref>, it is evident that using two middle layers did not improve the performance. These observations hold with different the number of bottleneck layer nodes. Consequently, it should be sufficient to double the number of input layer nodes for the single middle layer.</p>
<p>To further investigate the impact of the number of bottleneck layer nodes, in <xref ref-type="fig" rid="F4">Figure 4</xref>, we plot the means and the standard deviations of variable-wise input-reconstruction correlations. The number of latent factors was increased from 2 to 15. The standard deviations are indicated by the error bars in the plot. These results are evaluated based on the testing dataset with 3,944 samples.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Mean of input-reconstruction correlations in VAE and LFA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0004.tif"/>
</fig>
<p>The improvement on the mean of correlations significantly slows down after the 5th latent factor in LFA, consistent with previous results (Costa and McCrae, <xref ref-type="bibr" rid="B24">1992</xref>; Goldberg, <xref ref-type="bibr" rid="B36">1992</xref>). On the other hand, the performance of VAE keeps on improving as the number of bottleneck layer nodes increases and there is not an obvious &#x0201C;elbow&#x0201D; point as in LFA for determining the number of factors to be extracted.</p>
<p>In <xref ref-type="fig" rid="F5">Figure 5</xref>, we also compared R2 statistics in VAE to communality in LFA because it measures explained variance.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Mean of R2 statistics in VAE and communality in LFA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0005.tif"/>
</fig>
<p>We can see that in <xref ref-type="fig" rid="F5">Figure 5</xref>, R2 statistics in VAE outperforms communality in LFA and the curves reflect the same trend as in <xref ref-type="fig" rid="F4">Figure 4</xref>. Note that if the linear model in LFA completely holds, then communality should be the equivalent to R2. To plot (<xref ref-type="fig" rid="F5">Figure 5</xref>), we first calculated the two measures (R2 and communality) for each variable, and then, we calculated the mean and standard deviation of the two measures across all 50 IPIP Big5 variables.</p>
</sec>
<sec>
<title>Factor Stability Analysis</title>
<p>From the plot in <xref ref-type="fig" rid="F5">Figure 5</xref>, we cannot determine the number of factors that should be included in the personality model, and we further studied the stability of the discovered factors following the procedure outlined in the method section and summarized the results in <xref ref-type="fig" rid="F6">Figure 6</xref>, in which we used 10 bottleneck layer nodes and 100 middle layer nodes.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Clustering of factors from 10 VAE runs with 10 bottleneck layer nodes and 100 middle layer nodes in the Big 5 analysis.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0006.tif"/>
</fig>
<p>After varimax rotation, we gradually raised the threshold on the congruence scores until no two factors from the same run could be clustered together. The threshold was set to 0.90. We can see that 7 stable factor clusters emerged with average congruence scores greater than 0.98. For the first 4 factors, Extraversion, Agreeableness, Conscientiousness, and Neuroticism, the top 10 personality variables ranked by factor loadings are the same as those in the original Big 5 model. The average congruence scores of these factor clusters are very high, which are rounded to one when two decimal points are considered. The rest of the 3 factors have a slightly smaller congruence score. Since factor loadings are not very appropriate for exploring factor-variable association in VAE and these factors have relatively less obvious interpretations, we temporarily mark them as Imagination, Linguistic Intellect, and a factor X1. The X1 did not have high loadings on any items and so we did not assign a name to it.</p>
<p>We also investigated the case of using 9 bottleneck layer nodes and 100 middle hidden layer nodes. The stability analysis results are included in <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure 1</xref>.</p>
</sec>
<sec>
<title>Factor-Variable Associations</title>
<p>In <xref ref-type="fig" rid="F7">Figure 7</xref>, we plot the input-reconstruction correlation reduction after muting each factor when 10 bottleneck layer nodes are used.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Input-reconstruction correlation reduction in a Big5 VAE run (E, Extraversion; N, Neuroticism; A, Agreeableness; C, Conscientiousness; O, Openness. See <xref ref-type="table" rid="T2">Table 2</xref> for variable definitions).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0007.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Factor-variable association by inspecting input-reconstruction correlation reduction with 10 bottleneck layer nodes.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="left"><bold>Reduction</bold></th>
<th valign="top" align="left"><bold>Item</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Neuroticism</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">N6</td>
<td valign="top" align="left">0.54880679</td>
<td valign="top" align="left">I get upset easily.</td>
</tr>
<tr>
<td valign="top" align="left">N9</td>
<td valign="top" align="left">0.52292319</td>
<td valign="top" align="left">I get irritated easily.</td>
</tr>
<tr>
<td valign="top" align="left">N1</td>
<td valign="top" align="left">0.49218794</td>
<td valign="top" align="left">I get stressed out easily.</td>
</tr>
<tr>
<td valign="top" align="left">N8</td>
<td valign="top" align="left">0.46884062</td>
<td valign="top" align="left">I have frequent mood swings.</td>
</tr>
<tr>
<td valign="top" align="left">N7</td>
<td valign="top" align="left">0.43760175</td>
<td valign="top" align="left">I change my mood a lot.</td>
</tr>
<tr>
<td valign="top" align="left">N3</td>
<td valign="top" align="left">0.39057886</td>
<td valign="top" align="left">I worry about things.</td>
</tr>
<tr>
<td valign="top" align="left">N10</td>
<td valign="top" align="left">0.34668783</td>
<td valign="top" align="left">I often feel blue.</td>
</tr>
<tr>
<td valign="top" align="left">N5</td>
<td valign="top" align="left">0.3111992</td>
<td valign="top" align="left">I am easily disturbed.</td>
</tr>
<tr>
<td valign="top" align="left">N2</td>
<td valign="top" align="left">0.29470729</td>
<td valign="top" align="left">I am relaxed most of the time.</td>
</tr>
<tr>
<td valign="top" align="left">N4</td>
<td valign="top" align="left">0.17285815</td>
<td valign="top" align="left">I seldom feel blue.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Agreeableness</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">A4</td>
<td valign="top" align="left">0.67151324</td>
<td valign="top" align="left">I sympathize with others&#x00027; feelings.</td>
</tr>
<tr>
<td valign="top" align="left">A9</td>
<td valign="top" align="left">0.51527633</td>
<td valign="top" align="left">I feel others&#x00027; emotions.</td>
</tr>
<tr>
<td valign="top" align="left">A8</td>
<td valign="top" align="left">0.43419151</td>
<td valign="top" align="left">I take time out for others.</td>
</tr>
<tr>
<td valign="top" align="left">A5</td>
<td valign="top" align="left">0.43105851</td>
<td valign="top" align="left">I am not interested in other people&#x00027;s problems.</td>
</tr>
<tr>
<td valign="top" align="left">A6</td>
<td valign="top" align="left">0.41301325</td>
<td valign="top" align="left">I have a soft heart.</td>
</tr>
<tr>
<td valign="top" align="left">A7</td>
<td valign="top" align="left">0.31421132</td>
<td valign="top" align="left">I am not really interested in others.</td>
</tr>
<tr>
<td valign="top" align="left">A2</td>
<td valign="top" align="left">0.24600361</td>
<td valign="top" align="left">I am interested in people.</td>
</tr>
<tr>
<td valign="top" align="left">A1</td>
<td valign="top" align="left">0.18277531</td>
<td valign="top" align="left">I feel little concern for others.</td>
</tr>
<tr>
<td valign="top" align="left">A3</td>
<td valign="top" align="left">0.1748963</td>
<td valign="top" align="left">I insult people.</td>
</tr>
<tr>
<td valign="top" align="left">A10</td>
<td valign="top" align="left">0.15831119</td>
<td valign="top" align="left">I make people feel at ease.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Conscientiousness</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">C5</td>
<td valign="top" align="left">0.4724694</td>
<td valign="top" align="left">I get chores done right away.</td>
</tr>
<tr>
<td valign="top" align="left">C9</td>
<td valign="top" align="left">0.47059017</td>
<td valign="top" align="left">I follow a schedule.</td>
</tr>
<tr>
<td valign="top" align="left">C1</td>
<td valign="top" align="left">0.39228489</td>
<td valign="top" align="left">I am always prepared.</td>
</tr>
<tr>
<td valign="top" align="left">C6</td>
<td valign="top" align="left">0.37797469</td>
<td valign="top" align="left">I often forget to put things back in their proper place.</td>
</tr>
<tr>
<td valign="top" align="left">C7</td>
<td valign="top" align="left">0.361025</td>
<td valign="top" align="left">I like order.</td>
</tr>
<tr>
<td valign="top" align="left">C2</td>
<td valign="top" align="left">0.32991215</td>
<td valign="top" align="left">I leave my belongings around.</td>
</tr>
<tr>
<td valign="top" align="left">C4</td>
<td valign="top" align="left">0.28578476</td>
<td valign="top" align="left">I make a mess of things.</td>
</tr>
<tr>
<td valign="top" align="left">C8</td>
<td valign="top" align="left">0.23683019</td>
<td valign="top" align="left">I shirk my duties.</td>
</tr>
<tr>
<td valign="top" align="left">C10</td>
<td valign="top" align="left">0.22386922</td>
<td valign="top" align="left">I am exacting in my work.</td>
</tr>
<tr>
<td valign="top" align="left">C3</td>
<td valign="top" align="left">0.16418281</td>
<td valign="top" align="left">I pay attention to details.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">X1</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">A1</td>
<td valign="top" align="left">0.16598865</td>
<td valign="top" align="left">I feel little concern for others.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Not Stable</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">O4</td>
<td valign="top" align="left">0.17283142</td>
<td valign="top" align="left">I am not interested in abstract ideas.</td>
</tr>
<tr>
<td valign="top" align="left">N4</td>
<td valign="top" align="left">0.13201049</td>
<td valign="top" align="left">I seldom feel blue.</td>
</tr>
<tr>
<td valign="top" align="left">O2</td>
<td valign="top" align="left">0.10712201</td>
<td valign="top" align="left">I have difficulty understanding abstract ideas.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Not Stable</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">O9</td>
<td valign="top" align="left">0.13302234</td>
<td valign="top" align="left">I spend time reflecting on things.</td>
</tr>
<tr>
<td valign="top" align="left">A3</td>
<td valign="top" align="left">0.11271361</td>
<td valign="top" align="left">I insult people.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Linguistic Intellect</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">O1</td>
<td valign="top" align="left">0.53025199</td>
<td valign="top" align="left">I have a rich vocabulary.</td>
</tr>
<tr>
<td valign="top" align="left">O8</td>
<td valign="top" align="left">0.47849211</td>
<td valign="top" align="left">I use difficult words.</td>
</tr>
<tr>
<td valign="top" align="left">O7</td>
<td valign="top" align="left">0.32149661</td>
<td valign="top" align="left">I am quick to understand things.</td>
</tr>
<tr>
<td valign="top" align="left">O2</td>
<td valign="top" align="left">0.28731204</td>
<td valign="top" align="left">I have difficulty understanding abstract ideas.</td>
</tr>
<tr>
<td valign="top" align="left">O5</td>
<td valign="top" align="left">0.2294474</td>
<td valign="top" align="left">I have excellent ideas.</td>
</tr>
<tr>
<td valign="top" align="left">O10</td>
<td valign="top" align="left">0.21341601</td>
<td valign="top" align="left">I am full of ideas.</td>
</tr>
<tr>
<td valign="top" align="left">O4</td>
<td valign="top" align="left">0.17363996</td>
<td valign="top" align="left">I am not interested in abstract ideas.</td>
</tr>
<tr>
<td valign="top" align="left">O9</td>
<td valign="top" align="left">0.10652723</td>
<td valign="top" align="left">I spend time reflecting on things.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Imagination</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">O6</td>
<td valign="top" align="left">0.40875143</td>
<td valign="top" align="left">I do not have a good imagination.</td>
</tr>
<tr>
<td valign="top" align="left">O3</td>
<td valign="top" align="left">0.33323853</td>
<td valign="top" align="left">I have a vivid imagination.</td>
</tr>
<tr>
<td valign="top" align="left">O10</td>
<td valign="top" align="left">0.1164824</td>
<td valign="top" align="left">I am full of ideas.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Extraversion</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">E4</td>
<td valign="top" align="left">0.54024501</td>
<td valign="top" align="left">I keep in the background.</td>
</tr>
<tr>
<td valign="top" align="left">E2</td>
<td valign="top" align="left">0.52025504</td>
<td valign="top" align="left">I don&#x00027;t talk a lot.</td>
</tr>
<tr>
<td valign="top" align="left">E7</td>
<td valign="top" align="left">0.51419152</td>
<td valign="top" align="left">I talk to a lot of different people at parties.</td>
</tr>
<tr>
<td valign="top" align="left">E5</td>
<td valign="top" align="left">0.50472147</td>
<td valign="top" align="left">I start conversations.</td>
</tr>
<tr>
<td valign="top" align="left">E10</td>
<td valign="top" align="left">0.50394706</td>
<td valign="top" align="left">I am quiet around strangers.</td>
</tr>
<tr>
<td valign="top" align="left">E1</td>
<td valign="top" align="left">0.4833213</td>
<td valign="top" align="left">I am the life of the party.</td>
</tr>
<tr>
<td valign="top" align="left">E9</td>
<td valign="top" align="left">0.42410076</td>
<td valign="top" align="left">I don&#x00027;t mind being the center of attention.</td>
</tr>
<tr>
<td valign="top" align="left">E8</td>
<td valign="top" align="left">0.38028153</td>
<td valign="top" align="left">I don&#x00027;t like to draw attention to myself.</td>
</tr>
<tr>
<td valign="top" align="left">E6</td>
<td valign="top" align="left">0.34816174</td>
<td valign="top" align="left">I have little to say.</td>
</tr>
<tr>
<td valign="top" align="left">E3</td>
<td valign="top" align="left">0.33659072</td>
<td valign="top" align="left">I feel comfortable around people.</td>
</tr>
<tr>
<td valign="top" align="left">Factor</td>
<td valign="top" align="left">Not stable</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">C3</td>
<td valign="top" align="left">0.13559234</td>
<td valign="top" align="left">I pay attention to details.</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N = 3,944 testing samples; Only variables with correlation reduction higher than 0.1 are kept for each factor</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>From <xref ref-type="fig" rid="F7">Figure 7</xref>, we can see that the factor-variable associations are almost the same as in the LFA analysis. The openness factor fractionated into Imagination (F7 in <xref ref-type="fig" rid="F7">Figure 7</xref>), Linguistic Intellect (F6), and X1 (F3). X1 is a minor factor in the sense that its main items have small reductions. However, it withstood the test of stability, and it cannot be ignored as noise. Compared to <xref ref-type="fig" rid="F3">Figure 3</xref>, we can see that while LFA fractionated Openness into 3 factors, VAE fractionated it into two. LFA tends to fractionate more as we increase the number of assumed latent factors.</p>
<p>Inspecting the variables associated with each factor in <xref ref-type="table" rid="T2">Table 2</xref>, we can see that by using input-reconstruction correlation reduction, it is possible to identify factor-variable associations that mostly reproduce the results in the Big5 model.</p>
<p>We listed the correlations between the 7 factors in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Correlations between the 7 stable factors in the VAE IPIP Big 5 dataset analysis.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>(1)</bold></th>
<th valign="top" align="center"><bold>(2)</bold></th>
<th valign="top" align="center"><bold>(3)</bold></th>
<th valign="top" align="center"><bold>(4)</bold></th>
<th valign="top" align="center"><bold>(5)</bold></th>
<th valign="top" align="center"><bold>(6)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">(1) Conscientiousness</td>
<td valign="top" align="center">-</td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(2) Linguistic Intellect</td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center">-</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(3) Extraversion</td>
<td valign="top" align="center">&#x02212;0.05<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">-</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(4) Agreeableness</td>
<td valign="top" align="center"><bold>0.08<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.01</td>
<td valign="top" align="center"><bold>0.07<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">-</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(5) X1</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">&#x02212;0.05<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">-</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">(6) Neuroticism</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">&#x02212;0.05<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>0.06<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">(7) Imagination</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center"><bold>0.1<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center"><bold>0.08<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.02</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N = 3,944</italic>.</p>
<fn id="TN1"><label>&#x0002A;</label><p><italic>P &#x0003C; 0.05; Bold values and</italic></p></fn>
<fn id="TN2"><label>&#x0002A;&#x0002A;</label><p><italic>are indicate that P-value &#x0003C; 0.001</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>We have also calculated the sample-wise and variable-wise correlations between the reconstructed and the original inputs in both analyses over 3,944 testing samples and 50 personality variables. We used 7 hidden layer nodes for VAE and 7 factors for LFA for a fair comparison. We applied the Wilcoxon signed-rank test implemented in the SciPy package in python (Scipy Wilcoxon Signed-Rank Test Manual, n.d.) and calculated the <italic>p</italic>-values. The resulting statistics are shown in the box plots in <xref ref-type="fig" rid="F8">Figure 8</xref>. They show that the 7-factor VAE model performs better than the 7-factor LFA model. The variance of variable-wise correlation is significantly smaller in VAE than LFA.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Variable-wise and sample-wise correlation statistics in VAE and LFA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0008.tif"/>
</fig>
<p>In this study, we can see that VAE mostly replicated the factor structure of the Big 5 scale initially discovered by LFA.</p>
</sec>
</sec>
</sec>
<sec>
<title>Analysis of the IPIP HEXACO Dataset</title>
<p>In the VAE analysis of the IPIP Big 5 data, the inventory only contains 50 personality variables. We wanted to investigate if VAE can be used when more personality variables are on the questionnaire and to verify if the factor selection principle derived in the first VAE analysis can be generalized. For this purpose, we applied VAE to an IPIP HEXACO dataset with 240 personality variables. We selected the model structure according to the principles derived previously. Also, since the IPIP HEXACO dataset has many more items than the Big5 dataset, we anticipated more factors.</p>
<sec>
<title>LFA Analysis of the IPIP HEXACO Dataset</title>
<p>We first applied LFA analysis of the dataset and the resulting eigenvalues are plotted in the following scree plot:</p>
<p>From <xref ref-type="fig" rid="F9">Figure 9</xref>, we can see that there are about 8 factors before the elbow point. However, if we applied the rule of selecting factors with eigenvalues greater than one, there are over 37 factors with eigenvalues above 1. Yet in the literature, the reported number of factors through LFA analysis is 6.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Scree plot of the eigenvalues in the HEXACO dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0009.tif"/>
</fig>
<p>Since there is a big discrepancy in the number of factors that should be extracted according to various factor extraction criteria, we investigated the factor structure when different numbers of latent factors are assumed in LFA. We first plotted the input-reconstruction correlation reduction plot when we assumed 6 latent factors and each of them was muted in turn in <xref ref-type="fig" rid="F10">Figure 10</xref>. We can see that the plot reflects the standard HEXACO model structure because each factor is represented by a line with significant input-reconstruction correlation reduction over variables grouped for one factor and lower for other variables except some facets of Agreeableness and Humility-Honesty. Some of these facets are influenced by both the Agreeable and Humility-Honesty factor. We know that the factor level description is not sufficient to describe the complexity of the personality model. Personality variables associated with each factor are further divided into facets.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>HEXACO factor- variable association when 6 latent factors are assumed in LFA (see section Methods for the list of acronyms).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0010.tif"/>
</fig>
<p>We then investigated how the factor structure would change when we assume the number of latent factors to be 9 and 14 in <xref ref-type="fig" rid="F11">Figures 11</xref>, <xref ref-type="fig" rid="F12">12</xref>, respectively.</p>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p>HEXACO factor-variable association when 9 latent factors are assumed in LFA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0011.tif"/>
</fig>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p>HEXACO factor-personality variable association when 14 latent factors are assumed in LFA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0012.tif"/>
</fig>
<p>Inspecting the factor structure when we assume 6, 9, and 14 factors, respectively, we can see that the factor structure keeps on fractionating without stability. The factor structure also does not improve the representation of the model as the number of factors increases. For example, the Depression and the Anxiety facet in Emotionality is well-represented in the 6-factor structure in <xref ref-type="fig" rid="F10">Figure 10</xref>. However, when the number of factors is increased to 9, these two facets are influenced by several factors in which the factor plotted by the red dashed line (F7) is not representing any facet or factor (see the top half of <xref ref-type="fig" rid="F11">Figure 11</xref>). In the case of 14 factors in <xref ref-type="fig" rid="F12">Figure 12</xref>, we can see that except Extraversion and Agreeableness, all the rest of the factors are fractionated and there are 12 factors with a peak reduction greater than 0.2 and more than 5 variables. The structure is different from the 9-factor case, in which 9 factors have a peak reduction of greater than 0.2.</p>
<p>It is evident that when we apply LFA to factor analysis to a large set of personality variables, it is hard to determine the number of factors to be extracted based on eigenvalues or elbow points. The factor structure keeps on fractionating into smaller but still significant factors, and there is no definitive way to determine when to stop.</p>
</sec>
<sec>
<title>VAE Model Parameter Selection</title>
<p>We employed the model selection principles used in the VAE IPIP Big 5 dataset analysis. We set the number of nodes in the middle layer to 480, twice the number of input variables. Further increasing the number of nodes brought little improvement. A single middle layer between the input and the bottleneck layer was used and the activation function for the mid-layer was set to Relu. We gradually increased the number of bottleneck layer nodes to 12 and a maximum of 9 stable factors emerged in the analysis.</p>
</sec>
<sec>
<title>Factor Stability Analysis</title>
<p>The result of factor stability analysis on the 240 IPIP HEXACO variables is shown in <xref ref-type="fig" rid="F13">Figure 13</xref>.</p>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p>VAE analysis of the IPIP HEXACO dataset using 12 bottleneck layer nodes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0013.tif"/>
</fig>
<p>The filter threshold on the congruence scores was set to 0.875 such that no single factor is matched to two factors in another run. Nine stable factors with an average congruence score greater than 0.97 appeared across 10 VAE runs with 12 bottleneck layer nodes.</p>
<p>We then inspected the factor structure through the input-reconstruction correlation reduction plots when we set the number of bottleneck layer nodes to 9 and 14, respectively in <xref ref-type="fig" rid="F14">Figures 14</xref>, <xref ref-type="fig" rid="F15">15</xref>.</p>
<fig id="F14" position="float">
<label>Figure 14</label>
<caption><p>HEXACO factor-personality variable association when 9 latent factors are assumed in VAE.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0014.tif"/>
</fig>
<fig id="F15" position="float">
<label>Figure 15</label>
<caption><p>HEXACO factor-personality variable association when 14 latent factors are assumed in VAE.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-13-863926-g0015.tif"/>
</fig>
<p>Comparing <xref ref-type="fig" rid="F14">Figures 14</xref>, <xref ref-type="fig" rid="F15">15</xref>, we can see that the factor structures are mostly the same. The top panel consists of 4 factors: Agreeableness, Conscientiousness, Thrill-seeking (the Fearfulness facet of Emotionality), Emotionality (All facets of Emotionality except the Fearfulness facet). The bottom panel consists of 4 major factors: Machiavellianism (All Humility-Honesty facets except two Modesty variables), Creativity (The Creativity facet of Openness to Experience), Inquisitiveness (The Inquisitiveness facet of Openness to Experience), and Extraversion. Humility (Two Modesty facet variables in Humility-Honesty plus several Unconventionality facet variables in Openness to Experience) appeared as a stable factor in both our VAE stability analysis and the 14-factor analysis (dashed green line in <xref ref-type="fig" rid="F15">Figure 15</xref>).</p>
<p>Unlike the LFA analysis which keeps on fractionating as we increase the number of latent factors, VAE does not change the factor structure significantly and increasing the number of factors beyond the nine prominent factors added noise-like factors, such as F3, F5, F6, F7, F8 in <xref ref-type="fig" rid="F15">Figure 15</xref> (We define a factor as noise like factors when the maximum reduction on correlations is less than 0.1 and there are fewer than 5 items under the factor. Future research shall be conducted in terms of how to best chose these thresholds.). We can see that inspecting the limit on the number significant factors is a viable method for determining the number of factors to be extracted. We term this as the Significant Factor Limit Search (SFLS) method.</p>
<p>We conclude that by using VAE, we can determine the factor structure directly when a large pool of personality variables exists. Facet level analysis is not necessary since no significant factors can be discovered after the stable factors. Although we still cannot guarantee that the stability analysis or this SFLS method reports the true number of factors, at least in this case, the two types of analysis reported similar number of factors in HEXACO and Big 5. Unlike in LFA, the two types of analysis may return radically different number of factors especially when the set of analyzed personality variables are complex.</p>
<p>The zero-order correlations between the 9 discovered factors are listed in <xref ref-type="table" rid="T4">Table 4</xref>. The most significant correlations are highlighted in bold.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Correlations between rotated factors in the IPIP HEXACO VAE analysis.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>(1)</bold></th>
<th valign="top" align="center"><bold>(2)</bold></th>
<th valign="top" align="center"><bold>(3)</bold></th>
<th valign="top" align="center"><bold>(4)</bold></th>
<th valign="top" align="center"><bold>(5)</bold></th>
<th valign="top" align="center"><bold>(6)</bold></th>
<th valign="top" align="center"><bold>(7)</bold></th>
<th valign="top" align="center"><bold>(8)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">(1) Inquisitiveness</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(2) Agreeableness</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(3) Extraversion</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(4) Conscientiousness</td>
<td valign="top" align="center">&#x02013;<bold>0.14<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>0.06<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(5) Humility</td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.01</td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02013;<bold>0.06<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(6) Machiavellianism</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">&#x02013;<bold>0.09<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(7) Thrill-seeking</td>
<td valign="top" align="center">&#x02013;<bold>0.09<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">&#x02212;0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>0.07<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">(8) Emotionality</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center">0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>0.08<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">(9) Creativity</td>
<td valign="top" align="center">&#x02013;<bold>0.07<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center"><bold>0.06<xref ref-type="table-fn" rid="TN4"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.1</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">&#x02212;0.04<xref ref-type="table-fn" rid="TN3"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.03</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N = 3,756 testing samples</italic>;</p>
<fn id="TN3"><label>&#x0002A;</label><p><italic>P &#x0003C; 0.05; Bold values and</italic></p></fn>
<fn id="TN4"><label>&#x0002A;&#x0002A;</label><p><italic>are indicate that P-value &#x0003C; 0.001</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec>
<title>Cross-Regional Study</title>
<p>There is no significant deviation in mean, standard deviation. The mean and standard deviation statistics are included in the <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref>. We further studied the zero-order correlations between the 9 factors in different regions. Again, little variation was detected. We listed the zero-order correlations in West Europe and Asia in <xref ref-type="table" rid="T5">Tables 5</xref>, <xref ref-type="table" rid="T6">6</xref> as examples. Please see the <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref> for the rest of the regions.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Correlations between factors trained on North American samples and tested on West Europe samples.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>(1)</bold></th>
<th valign="top" align="center"><bold>(2)</bold></th>
<th valign="top" align="center"><bold>(3)</bold></th>
<th valign="top" align="center"><bold>(4)</bold></th>
<th valign="top" align="center"><bold>(5)</bold></th>
<th valign="top" align="center"><bold>(6)</bold></th>
<th valign="top" align="center"><bold>(7)</bold></th>
<th valign="top" align="center"><bold>(8)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">(1) Machiavellianism</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(2) Emotionality</td>
<td valign="top" align="center"><bold>0.08<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(3) Thrill-seeking</td>
<td valign="top" align="center">0.05<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(4) Conscientiousness</td>
<td valign="top" align="center">0.2</td>
<td valign="top" align="center"><bold>0.14<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.05<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(5) Inquisitiveness</td>
<td valign="top" align="center"><bold>&#x02212;0.12<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center"><bold>&#x02212;0.11<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.03</td>
<td valign="top" align="center"><bold>&#x02212;0.08<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(6) Creativity</td>
<td valign="top" align="center">0.1</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center"><bold>&#x02212;0.07<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(7) Extraversion</td>
<td valign="top" align="center"><bold>&#x02212;0.07<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.06<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.04</td>
<td valign="top" align="center">&#x02212;0.04<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>0.21<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center"><bold>&#x02212;0.18<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">(8) Agreeableness</td>
<td valign="top" align="center">0.05</td>
<td valign="top" align="center"><bold>0.14<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center"><bold>0.18<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.06<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.04<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>&#x02212;0.13<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">(9) Humility</td>
<td valign="top" align="center"><bold>&#x02212;0.08<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center"><bold>0.08<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center"><bold>&#x02212;0.08<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.04<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>0.11<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.04<xref ref-type="table-fn" rid="TN5"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>&#x02212;0.08<xref ref-type="table-fn" rid="TN6"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N = 2,607</italic>;</p>
<fn id="TN5"><label>&#x0002A;</label><p><italic>P &#x0003C; 0.05; Bold values and</italic></p></fn>
<fn id="TN6"><label>&#x0002A;&#x0002A;</label><p><italic>are indicate that P-value &#x0003C; 0.001</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Correlations between factors trained on North American samples and tested on Asian Samples.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>(1)</bold></th>
<th valign="top" align="center"><bold>(2)</bold></th>
<th valign="top" align="center"><bold>(3)</bold></th>
<th valign="top" align="center"><bold>(4)</bold></th>
<th valign="top" align="center"><bold>(5)</bold></th>
<th valign="top" align="center"><bold>(6)</bold></th>
<th valign="top" align="center"><bold>(7)</bold></th>
<th valign="top" align="center"><bold>(8)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">(1) Machiavellianism</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(2) Emotionality</td>
<td valign="top" align="center"><bold>0.15<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(3) Thrill-seeking</td>
<td valign="top" align="center">&#x02212;0.07<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(4) Conscientiousness</td>
<td valign="top" align="center"><bold>0.13<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center"><bold>0.12<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.01</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(5) Inquisitiveness</td>
<td valign="top" align="center">&#x02212;0.09<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center"><bold>&#x02212;0.12<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">&#x02212;0.05</td>
<td valign="top" align="center">1</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(6) Creativity</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.09<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.04</td>
<td valign="top" align="center">&#x02212;0.11<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">(7) Extraversion</td>
<td valign="top" align="center"><bold>&#x02212;0.17<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center"><bold>0.16<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.09<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">(8) Agreeableness</td>
<td valign="top" align="center">&#x02212;0.11<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.08<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center"><bold>0.21<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center"><bold>&#x02212;0.12<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">1</td>
</tr>
<tr>
<td valign="top" align="left">(9) Humility</td>
<td valign="top" align="center"><bold>&#x02212;0.12<xref ref-type="table-fn" rid="TN8"><sup>&#x0002A;&#x0002A;</sup></xref></bold></td>
<td valign="top" align="center">0.11<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.11<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.04</td>
<td valign="top" align="center">&#x02212;0.02</td>
<td valign="top" align="center">0.11<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">&#x02212;0.07<xref ref-type="table-fn" rid="TN7"><sup>&#x0002A;</sup></xref></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N = 850</italic>;</p>
<fn id="TN7"><label>&#x0002A;</label><p><italic>P &#x0003C; 0.05; Bold values and</italic></p></fn>
<fn id="TN8"><label>&#x0002A;&#x0002A;</label><p><italic>are indicate that P-value &#x0003C; 0.001</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Note that there are significantly larger correlations in the cross-regional study. This difference is caused by a smaller number of training samples, which affected the convergence of the VAE algorithm. The cross-region study was aimed at verifying the consistency of factor statistics across different regions, the results cannot be directly compared to the results obtained by using training samples from all regions.</p>
</sec>
</sec>
<sec id="s4">
<title>Discussions</title>
<p>In this paper, we first investigated how VAE can be applied in exploring the factor structure in personality variables and how it compares to LFA. Past research showed that the personality model discovered via LFA must organize the scales at the factor and the facet level. When the number of investigated personality variables is large, LFA returns unstable factor structures as the number of assumed latent factors grows. LAF must cut a smaller number of factors such that the factors are generalizable. Yet, it is well-known that factor level representation alone is not sufficient and facet level information cannot be ignored when personality models are used for predicting various behaviors. In contrast, VAE returns stable factor structures even if the number of assumed latent factors is greater than the number of stable factors. Consequently, VAE can be applied to explore factor structures in a one-shot process given a pool of personality variables. A follow-up analysis at the facet level is not necessary and consequently, a lot of <italic>ad-hoc</italic> decisions can be avoided in the process.</p>
<p>This difference is due to VAE&#x00027;s inherent ability in dealing with non-linear data models. The fact that LFA cannot exhaust all useful information at the factor level is an indication of the non-linearity in the underlying personality model.</p>
<p>Note that VAE and LFA returned similar factor structures in the Big5 dataset. This shows that the 50 variables in the IPIP Big 5 dataset are mostly linearly related to the latent factors. On the other hand, the factor structures returned by VAE and LFA are very different in the more complicated HEXACO dataset, which indicates that the 240 variables in HEXACO are less linearly organized by the latent factors.</p>
<p>We aim to show that VAE is a viable analytical tool for exploratory factor analysis in this paper. The datasets we analyzed were collected using the IPIP Big 5 and IPIP HEXACO questionnaires for the purpose of comparison with LFA. We fully understand that the resulting factor models may not be comprehensive. There could still exist factors that the questionnaires have not covered, and self-reported data will limit the usefulness of the calculated model. Now, we will be able to extend the application of VAE to include various types of data in the future for constructing useful personality models.</p>
<p>For proper personality model development, VAE can be applied in a two-stage process: In stage one, the aim is to discover all essential and stable factors and select personality variables to represent all the factors properly. Comparing the stable factors found in the IPIP Big5 and the IPIP HEXACO analyses, we can see that similar factors can be discovered and aligned together even if the measured personality variables are not entirely identical. This suggests that in the future, we may perform a meta-analysis of multiple datasets and then pool all the discovered factors and their associated personality variables together for final analysis. In stage two, enough sample data should be collected for the pooled personality variables to construct a comprehensive personality model. Furthermore, it is anticipated that observer report data can be easily incorporated in the VAE framework as we can simply extend the input vector to include observer reported variables.</p>
<sec>
<title>Limitations of VAE</title>
<p>The main limitation of effective VAE application is the number of samples available for analysis. For example, when we applied VAE on a reduced dataset with 2,099 samples, the reported mean (std) of the input-reconstruction correlations was 0.61 (0.08), which is significantly smaller than 0.65 (0.08) when all 18,779 samples are used in the HEXACO analysis. This indicates that VAE cannot converge to a lower cost function value without enough samples. VAE analysis requires a significantly larger number of samples than those required in LFA.</p>
<p>The required number of samples should also scale with the number of input variables. For example, in the IPIP HEXACO analysis, when we reduce the number of input variables to 79, the algorithm returns a much higher variable-wise mean correlation of 0.72 (0.08) than the case with 240 variables at 0.65 (0.08). However, if we reduce the number of input variables, some of the discovered factors will not be adequately represented by a group of personality variables just like the case of Big 5 dataset analysis. The pruning of the input items should be done carefully such that all factors are well represented.</p>
<p>We should balance the need to include more personality variables with the need to collect more samples, so that the VAE algorithm can converge properly. If the goal is to discover stable factors, then the average congruence score, the number of stable factors, and the reliability of the discovered factors can be used as guiding metrics in determining if enough samples have been included. One should evaluate if reducing the number of samples will affect the constructed model in all these statistics. If increasing the number of samples doesn&#x00027;t change the set of discovered stable factors, then the number of samples is enough.</p>
<p>If an accurate encoder is required, the guiding metric should be the mean and standard deviation of the input-reconstruction correlations. Enough samples should be included such that these measures can meet the requirement.</p>
<p>Another limitation of VAE is overfitting. To prevent this problem, besides a strict separation in the training testing dataset, we also monitor the loss function values calculated based on the training and the validation dataset. When the loss function value in the validation dataset does not decrease with the loss function in the training dataset, we stop the algorithm to prevent over-fitting.</p>
<p>While VAE-based factor analysis can be used to explore non-linear associations between latent factors and manifest personality variables, a common concern is the interpretability of the model. VAE does not have the equivalent measure of factor loading as in LFA. However, as we have demonstrated, the reduction of input-reconstruction correlations can be used as a substitute for examining factor-variable associations which returned similar results to that of factor loadings in LFA. Further research is needed to confirm its suitability as a measure for interpretating the discovered factors in VAE for practitioners to adopt it.</p>
</sec>
<sec>
<title>Future Research</title>
<p>In VAE, since we do not have a reliability measure as in LFA, we cannot determine what is adequate for representing a factor. In LFA, the convention in the field has been that each factor is represented by 5&#x02013;10 variables with high alpha reliability. However, it seems that these calculated measures are never enough and a link to the biological basis is required for the factors to be called &#x0201C;adequate&#x0201D; in the view of some researchers. Ultimately, we should test how well VAE constructed factor structures can be used for various behavior predictions in comparison to LFA.</p>
<p>Another area for extending the current research is to incorporate observer report data into personality model construction. The hurdle lies in the availability of large observer datasets since it is much more difficult and costly to collect observer ratings. Nevertheless, given the increasing availability of various data and the complicated data structure, VAE shall find broad applications in such areas.</p>
</sec>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s9">Supplementary Material</xref>, further inquiries can be directed to the corresponding author/s. All code is made available on OSF (<ext-link ext-link-type="uri" xlink:href="https://osf.io/6bf3w/">https://osf.io/6bf3w/</ext-link>).</p>
</sec>
<sec id="s6">
<title>Ethics Statement</title>
<p>Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. Written informed consent from the patients/participants was not required to participate in this study in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>JZ came up with the idea of investigating personality model construction using deep learning methods, implemented the algorithm, and wrote the manuscript. YH identified variational autoencoder as the appropriate tool for the task and advised on the application of VAE to the problem. Both authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack><p>We thank reviewers for their helpful comments that have greatly improved the presentation of the paper.</p>
</ack>
<sec sec-type="supplementary-material" id="s9">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyg.2022.863926/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpsyg.2022.863926/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.DOCX" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ahmad</surname> <given-names>H.</given-names></name> <name><surname>Arif</surname> <given-names>A.</given-names></name> <name><surname>Khattak</surname> <given-names>A. M.</given-names></name> <name><surname>Habib</surname> <given-names>A.</given-names></name> <name><surname>Asghar</surname> <given-names>M. Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Applying deep neural networks for predicting dark triad personality trait of online users</article-title>, in <source>2020 International Conference on Information Networking (ICOIN)</source>, p. <fpage>102</fpage>&#x02013;<lpage>5</lpage>. <publisher-name>IEEE</publisher-name>. <pub-id pub-id-type="doi">10.1109/ICOIN48656.2020.9016525</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ainsworth</surname> <given-names>S. K.</given-names></name> <name><surname>Foti</surname> <given-names>N. J.</given-names></name> <name><surname>Lee</surname> <given-names>A. K.</given-names></name></person-group> (<year>2018</year>). <article-title>Oi-VAE: output interpretable VAEs for nonlinear group factor analysis</article-title>, in <source>International Conference on Machine Learning</source>, p. <fpage>119</fpage>&#x02013;<lpage>28</lpage>. <publisher-name>PMLR</publisher-name>.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashton</surname> <given-names>M. C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name></person-group> (<year>2007</year>). <article-title>Empirical, theoretical, and practical advantages of the HEXACO model of personality structure</article-title>. <source>Personal. Soc. Psychol. Rev.</source> <volume>11</volume>, <fpage>150</fpage>&#x02013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1177/1088868306294907</pub-id><pub-id pub-id-type="pmid">18453460</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashton</surname> <given-names>M. C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name></person-group> (<year>2008a</year>). <article-title>The HEXACO model of personality structure and the importance of the H factor</article-title>. <source>Soc. Personal. Psychol. Compass</source> <volume>2</volume>, <fpage>1952</fpage>&#x02013;<lpage>1962</lpage>. <pub-id pub-id-type="doi">10.1111/j.1751-9004.2008.00134.x</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashton</surname> <given-names>M. C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name></person-group> (<year>2008b</year>). <article-title>The prediction of honesty&#x02013;humility-related criteria by the HEXACO and five-factor models of personality</article-title>. <source>J. Res. Personal.</source> <volume>42</volume>, <fpage>1216</fpage>&#x02013;<lpage>1228</lpage>. <pub-id pub-id-type="doi">10.1016/j.jrp.2008.03.006</pub-id><pub-id pub-id-type="pmid">21859221</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashton</surname> <given-names>M. C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>De Vries</surname> <given-names>R. E.</given-names></name></person-group> (<year>2014</year>). <article-title>The HEXACO honesty-humility, agreeableness, and emotionality factors: a review of research and theory</article-title>. <source>Personal. Soc. Psychol. Rev.</source> <volume>18</volume>, <fpage>139</fpage>&#x02013;<lpage>152</lpage>. <pub-id pub-id-type="doi">10.1177/1088868314523838</pub-id><pub-id pub-id-type="pmid">24577101</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashton</surname> <given-names>M. C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Goldberg</surname> <given-names>L. R.</given-names></name></person-group> (<year>2007</year>). <article-title>The IPIP&#x02013;HEXACO scales: an alternative, public-domain measure of the personality constructs in the HEXACO model</article-title>. <source>Personal. Indiv. Differ.</source> <volume>42</volume>, <fpage>1515</fpage>&#x02013;<lpage>1526</lpage>. <pub-id pub-id-type="doi">10.1016/j.paid.2006.10.027</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Avi</surname> <given-names>C.</given-names></name> <name><surname>Burshtein</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Unsupervised Linear and Nonlinear Channel Equalization and Decoding Using Variational Autoencoders</article-title>. <source>IEEE Transac. Cogn. Commun. Netw.</source> <volume>6</volume>, <fpage>1003</fpage>&#x02013;<lpage>1018</lpage>. <pub-id pub-id-type="doi">10.1109/TCCN.2020.2990773</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Azucar</surname> <given-names>D.</given-names></name> <name><surname>Marengo</surname> <given-names>D.</given-names></name> <name><surname>Settanni</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Predicting the big 5 personality traits from digital footprints on social media: a meta-analysis</article-title>. <source>Personal. Indiv. Differ.</source> <volume>124</volume>, <fpage>150</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1016/j.paid.2017.12.018</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baldi</surname> <given-names>P.</given-names></name> <name><surname>Hornik</surname> <given-names>K.</given-names></name></person-group> (<year>1989</year>). <article-title>Neural networks and principal component analysis: learning from examples without local minima</article-title>. <source>Neural Netw.</source> <volume>2</volume>, <fpage>53</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1016/0893-6080(89)90014-2</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barrick</surname> <given-names>M. R.</given-names></name> <name><surname>Mount</surname> <given-names>M. K.</given-names></name></person-group> (<year>1991</year>). <article-title>The big five personality dimensions and job performance: a meta-analysis</article-title>. <source>Personnel Psychol.</source> <volume>44</volume>, <fpage>1</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1111/j.1744-6570.1991.tb00688.x</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Batmaz</surname> <given-names>Z.</given-names></name> <name><surname>Yurekli</surname> <given-names>A.</given-names></name> <name><surname>Bilge</surname> <given-names>A.</given-names></name> <name><surname>Kaleli</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>A review on deep learning for recommender systems: challenges and remedies</article-title>. <source>Artif. Intell. Rev.</source> <volume>52</volume>, <fpage>1</fpage>&#x02013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-018-9654-y</pub-id><pub-id pub-id-type="pmid">34305252</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baumeister</surname> <given-names>H.</given-names></name> <name><surname>Montag</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <source>Digital Phenotyping and Mobile Sensing</source>. <publisher-loc>Basel</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bhavya</surname> <given-names>S.</given-names></name> <name><surname>Pillai</surname> <given-names>A. S.</given-names></name> <name><surname>Guazzaroni</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>Personality identification from social media using deep learning: a review</article-title>, in <source>Soft Computing for Problem Solving</source>. <publisher-name>Singapore</publisher-name>: <publisher-loc>Springer</publisher-loc>. p. <fpage>523</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-15-0184-5_45</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bobadilla</surname> <given-names>J.</given-names></name> <name><surname>Alonso</surname> <given-names>S.</given-names></name> <name><surname>Hernando</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep learning architecture for collaborative filtering recommender systems</article-title>. <source>NATO Adv. Sci. Inst. Series E.</source> <volume>10</volume>, <fpage>2441</fpage>. <pub-id pub-id-type="doi">10.3390/app10072441</pub-id><pub-id pub-id-type="pmid">34451118</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bourlard</surname> <given-names>H.</given-names></name> <name><surname>Kamp</surname> <given-names>Y.</given-names></name></person-group> (<year>1988</year>). <article-title>Auto-association by multilayer perceptrons and singular value decomposition</article-title>. <source>Biol. Cyber.</source> <volume>59</volume>, <fpage>291</fpage>&#x02013;<lpage>294</lpage>. <pub-id pub-id-type="doi">10.1007/BF00332918</pub-id><pub-id pub-id-type="pmid">3196773</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Burgess</surname> <given-names>C. P.</given-names></name> <name><surname>Higgins</surname> <given-names>I.</given-names></name> <name><surname>Pal</surname> <given-names>A.</given-names></name> <name><surname>Matthey</surname> <given-names>L.</given-names></name> <name><surname>Watters</surname> <given-names>N.</given-names></name> <name><surname>Desjardins</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>Understanding disentangling in &#x003B2;-VAE</article-title>. <source>arXiv[Preprint].arXiv:1804.03599</source>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carthy</surname> <given-names>A.</given-names></name> <name><surname>Gray</surname> <given-names>G.</given-names></name> <name><surname>McGuinness</surname> <given-names>C.</given-names></name> <name><surname>Owende</surname> <given-names>P.</given-names></name></person-group> (<year>2014</year>). <article-title>A review of psychometric data analysis and applications in modelling of academic achievement in tertiary education</article-title>. <source>J. Learn. Anal</source>. <volume>1</volume>, <fpage>57</fpage>&#x02013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.18608/jla.2014.11.5</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Castelvecchi</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep learning boosts google translate tool</article-title>. <source>Nature</source>. <pub-id pub-id-type="doi">10.1038/nature.2016.20696</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cattell</surname> <given-names>R. B.</given-names></name></person-group> (<year>1943</year>). <article-title>The description of personality: basic traits resolved into clusters</article-title>. <source>J. Abnormal Soc. Psychol.</source> <volume>38</volume>, <fpage>476</fpage>&#x02013;<lpage>506</lpage>. <pub-id pub-id-type="doi">10.1037/h0054116</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cattell</surname> <given-names>R. B.</given-names></name> <name><surname>Eber</surname> <given-names>H. W.</given-names></name> <name><surname>Tatsuoka</surname> <given-names>M. M.</given-names></name></person-group> (<year>1970</year>). <source>Handbook for the Sixteen Personality Factor Questionnaire (16 PF): In Clinical, Educational, Industrial, and Research Psychology, for Use with All Forms of the Test</source>. <publisher-name>Institute for Personality and Ability Testing</publisher-name>.</citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chauhan</surname> <given-names>N. K.</given-names></name> <name><surname>Singh</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>A review on conventional machine learning vs deep learning</article-title>, in <source>2018 International Conference on Computing, Power and Communication Technologies (GUCON)</source>. <publisher-loc>Greater Noida</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>347</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1109/GUCON.2018.8675097</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chhabra</surname> <given-names>G. S.</given-names></name> <name><surname>Sharma</surname> <given-names>A.</given-names></name> <name><surname>Krishnan</surname> <given-names>N. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Deep learning model for personality traits classification from text emphasis on data slicing</article-title>, in <source>IOP Conference Series: Materials Science and Engineering.</source> vol. <volume>495</volume>, p. <fpage>012007</fpage>. <pub-id pub-id-type="doi">10.1088/1757-899X/495/1/012007</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Costa</surname> <given-names>P. T.</given-names></name> <name><surname>McCrae</surname> <given-names>R. R.</given-names></name></person-group> (<year>1992</year>). <article-title>The five-factor model of personality and its relevance to personality disorders</article-title>. <source>J. Personality Diso.</source> <volume>6</volume>, <fpage>343</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1521/pedi.1992.6.4.343</pub-id><pub-id pub-id-type="pmid">25613659</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cybenko</surname> <given-names>G.</given-names></name></person-group> (<year>1989</year>). <article-title>Approximation by Superpositions of a Sigmoidal Function</article-title>. <source>Mathem. Control, Sign. Syst.</source> <volume>2</volume>, <fpage>303</fpage>&#x02013;<lpage>314</lpage>. <pub-id pub-id-type="doi">10.1007/BF02551274</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>De Vries</surname> <given-names>R. E.</given-names></name> <name><surname>de Vries</surname> <given-names>A.</given-names></name> <name><surname>Feij</surname> <given-names>J. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Sensation seeking, risk-taking, and the HEXACO model of personality</article-title>. <source>Personality Indiv. Differ.</source> <volume>47</volume>, <fpage>536</fpage>&#x02013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1016/j.paid.2009.05.029</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M. W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name></person-group> (<year>2018</year>). <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>. <source>arXiv[Preprint].arXiv:1810.04805</source>.<pub-id pub-id-type="pmid">35689168</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eddine Bekhouche</surname> <given-names>S.</given-names></name> <name><surname>Dornaika</surname> <given-names>F.</given-names></name> <name><surname>Ouafi</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Personality traits and job candidate screening via analyzing facial videos</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</source> (<publisher-loc>Honolulu, HI</publisher-loc>: <publisher-name>IEEE</publisher-name>), p. <fpage>10</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2017.211</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elngar</surname> <given-names>A. A.</given-names></name> <name><surname>Jain</surname> <given-names>N.</given-names></name> <name><surname>Sharma</surname> <given-names>D.</given-names></name> <name><surname>Negi</surname> <given-names>H.</given-names></name> <name><surname>Trehan</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>A deep learning based analysis of the big five personality traits from handwriting samples using image processing</article-title>. <source>J. Inf. Technol. Manag.</source> <volume>12</volume>, <fpage>3</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.22059/JITM.2020.78884</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feher</surname> <given-names>A.</given-names></name> <name><surname>Vernon</surname> <given-names>P. A.</given-names></name></person-group> (<year>2021</year>). <article-title>Looking beyond the big five: a selective review of alternatives to the big five model of personality</article-title>. <source>Personal. Indiv. Differ.</source> <volume>169</volume>, <fpage>110002</fpage>. <pub-id pub-id-type="doi">10.1016/j.paid.2020.110002</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fiske</surname> <given-names>D. W.</given-names></name></person-group> (<year>1949</year>). <article-title>Consistency of the Factorial Structures of Personality Ratings from Different Sources</article-title>. <source>J. Abnormal Soc. Psychol.</source> <volume>44</volume>, <fpage>329</fpage>&#x02013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1037/h0057198</pub-id><pub-id pub-id-type="pmid">18146776</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galton</surname> <given-names>F.</given-names></name></person-group> (<year>1884</year>). <article-title>Measurement of Character</article-title>. <source>Fortnightly Rev.</source> <volume>36</volume>, <fpage>179</fpage>&#x02013;<lpage>185</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="web"><person-group person-group-type="editor"><name><surname>Goettfert</surname> <given-names>S.</given-names></name> <name><surname>Kriner</surname> <given-names>F.</given-names></name></person-group> (n.d). <article-title>Open-Source Psychometrics Project</article-title>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://openpsychometrics.org/_rawdata">https://openpsychometrics.org/_rawdata</ext-link> (accessed August 10, 2021).</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>L. R.</given-names></name></person-group> (<year>1981</year>). <article-title>Language and individual differences: the search for universals in personality lexicons</article-title>. <source>Personal. Soc. Psychol. Rev.</source> <volume>2</volume>, <fpage>141</fpage>&#x02013;<lpage>165</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>L. R.</given-names></name></person-group> (<year>1990</year>). <article-title>An Alternative &#x02018;Description of Personality&#x00027;: The Big-Five Factor Structure</article-title>. <source>J. Personality Soc. Psychol.</source> <volume>59</volume>, <fpage>1216</fpage>&#x02013;<lpage>1229</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.59.6.1216</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>L. R.</given-names></name></person-group> (<year>1992</year>). <article-title>The development of markers for the big-five factor structure</article-title>. <source>Psychol. Assess.</source> <volume>4</volume>, <fpage>26</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1037/1040-3590.4.1.26</pub-id><pub-id pub-id-type="pmid">11575511</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>L. R.</given-names></name></person-group> (<year>1999</year>). <article-title>A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models</article-title>. <source>Personal. Psychol. Europe</source>. <volume>7</volume>, <fpage>7</fpage>&#x02013;<lpage>28</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>L. R.</given-names></name> <name><surname>Saucier</surname> <given-names>G.</given-names></name></person-group> (<year>1998</year>). <article-title>What is beyond the big five?</article-title> <source>J. Personality</source>. <volume>66</volume>, <fpage>495</fpage>&#x02013;<lpage>524</lpage>. <pub-id pub-id-type="doi">10.1111/1467-6494.00022</pub-id><pub-id pub-id-type="pmid">9728415</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grice</surname> <given-names>J. W.</given-names></name></person-group> (<year>2001</year>). <article-title>Computing and evaluating factor scores</article-title>. <source>Psychol. Method.</source> <volume>6</volume>, <fpage>430</fpage>&#x02013;<lpage>450</lpage>. <pub-id pub-id-type="doi">10.1037/1082-989X.6.4.430</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hassan</surname> <given-names>H.</given-names></name> <name><surname>Asad</surname> <given-names>S.</given-names></name> <name><surname>Hoshino</surname> <given-names>Y</given-names></name></person-group> (<year>2016</year>). <article-title>Determinants of leadership style in big five personality dimensions</article-title>. <source>Univ. J. Manage.</source> <volume>4</volume>, <fpage>161</fpage>&#x02013;<lpage>179</lpage>. <pub-id pub-id-type="doi">10.13189/ujm.2016.040402</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Higgins</surname> <given-names>I.</given-names></name> <name><surname>Matthey</surname> <given-names>L.</given-names></name> <name><surname>Pal</surname> <given-names>A.</given-names></name> <name><surname>Burgess</surname> <given-names>C.</given-names></name> <name><surname>Glorot</surname> <given-names>X.</given-names></name> <name><surname>Botvinick</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <source>Beta-VAE: learning basic visual concepts with a constrained variational framework</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://openreview.net/pdf?id=Sy2fzU9gl">https://openreview.net/pdf?id=Sy2fzU9gl</ext-link> (accessed December 20, 2013).</citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Horn</surname> <given-names>J. L.</given-names></name></person-group> (<year>1965</year>). <article-title>A rationale and test for the number of factors in factor analysis</article-title>. <source>Psychometrika</source> <volume>30</volume>, <fpage>179</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1007/BF02289447</pub-id><pub-id pub-id-type="pmid">14306381</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hornik</surname> <given-names>K.</given-names></name></person-group> (<year>1991</year>). <article-title>Approximation capabilities of multilayer feedforward networks</article-title>. <source>Neural Net.</source> <volume>4</volume>, <fpage>251</fpage>&#x02013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.1016/0893-6080(91)90009-T</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jacobson</surname> <given-names>K.</given-names></name> <name><surname>Murali</surname> <given-names>V.</given-names></name> <name><surname>Newett</surname> <given-names>E.</given-names></name> <name><surname>Whitman</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <article-title>Music personalization at spotify</article-title>, in <source>Proceedings of the 10th ACM Conference on Recommender Systems</source>. <publisher-loc>RecSys &#x00027;16. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>373</fpage>. <pub-id pub-id-type="doi">10.1145/2959100.2959120</pub-id><pub-id pub-id-type="pmid">30270951</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jason</surname> <given-names>B.</given-names></name></person-group> (<year>2016</year>). <source>Deep Learning With Python: Develop Deep Learning Models on Theano and TensorFlow Using Keras</source>. <publisher-name>Machine Learning Mastery</publisher-name>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jerram</surname> <given-names>K. L.</given-names></name> <name><surname>Coleman</surname> <given-names>P. G.</given-names></name></person-group> (<year>1999</year>). <article-title>The big five personality traits and reporting of health problems and health behaviour in old age</article-title>. <source>Br. J. Health Psychol.</source> <volume>4</volume>, <fpage>181</fpage>&#x02013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1348/135910799168560</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Judge</surname> <given-names>T. A.</given-names></name> <name><surname>Bono</surname> <given-names>J. E.</given-names></name></person-group> (<year>2000</year>). <article-title>Five-factor model of personality and transformational leadership</article-title>. <source>J. Appl. Psychol.</source> <volume>85</volume>, <fpage>751</fpage>&#x02013;<lpage>765</lpage>. <pub-id pub-id-type="doi">10.1037/0021-9010.85.5.751</pub-id><pub-id pub-id-type="pmid">11055147</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Park</surname> <given-names>E.</given-names></name></person-group> (<year>2020</year>). <article-title>A deep learning model for detecting mental illness from user content on social media</article-title>. <source>Sci. Rep.</source> <volume>10</volume>, <fpage>11846</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-68764-y</pub-id><pub-id pub-id-type="pmid">32678250</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Auto-encoding variational bayes</article-title>. <source>arXiv[Prepint].arXiv:1312.6114v.10.</source>.<pub-id pub-id-type="pmid">32176273</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>An introduction to variational autoencoders</article-title>. <source>arXiv[Preprint].arXiv:1906.02691</source>. <pub-id pub-id-type="doi">10.1561/9781680836233</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2017</year>). <article-title>ImageNet classification with deep convolutional neural networks</article-title>. <source>Commun ACM.</source> <volume>60</volume>, <fpage>84</fpage>&#x02013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1145/3065386</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>K. P.</given-names></name> <name><surname>Gavrilova</surname> <given-names>M. L.</given-names></name></person-group> (<year>2019</year>). <article-title>Personality traits classification on Twitter</article-title>, in <source>2019 16th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)</source>, p. <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/AVSS.2019.8909839</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Ashton</surname> <given-names>M. C.</given-names></name></person-group> (<year>2004</year>). <article-title>Psychometric properties of the HEXACO personality inventory</article-title>. <source>Multiv. Behav. Res.</source> <volume>39</volume>, <fpage>329</fpage>&#x02013;<lpage>358</lpage>. <pub-id pub-id-type="doi">10.1207/s15327906mbr3902_8</pub-id><pub-id pub-id-type="pmid">26804579</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Ashton</surname> <given-names>M. C.</given-names></name></person-group> (<year>2005</year>). <article-title>Psychopathy, machiavellianism, and narcissism in the five-factor model and the HEXACO model of personality structure</article-title>. <source>Personal. Indiv. Differ.</source> <volume>38</volume>, <fpage>1571</fpage>&#x02013;<lpage>1582</lpage>. <pub-id pub-id-type="doi">10.1016/j.paid.2004.09.016</pub-id><pub-id pub-id-type="pmid">32881539</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Zhu</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep learning for constructing microblog behavior representation to identify social media user&#x00027;s personality</article-title>. <source>PeerJ. Comput. Sci.</source> <volume>2</volume>, <fpage>e81</fpage>. <pub-id pub-id-type="doi">10.7717/peerj-cs.81</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lopez-Alvis</surname> <given-names>J.</given-names></name> <name><surname>Laloy</surname> <given-names>E.</given-names></name> <name><surname>Nguyen</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep generative models in inversion: a review and development of a new approach based on a variational autoencoder</article-title>. <source>arXiv[Preprint].arXiv:2008.12056.</source> <pub-id pub-id-type="doi">10.1016/j.cageo.2021.104762</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCrae</surname> <given-names>R. R.</given-names></name> <name><surname>Costa</surname> <given-names>P. T.</given-names></name> <name><surname>Del Pilar</surname> <given-names>G. H.</given-names></name> <name><surname>Rolland</surname> <given-names>J. P.</given-names></name> <name><surname>Parker</surname> <given-names>W. D.</given-names></name></person-group> (<year>1998</year>). <article-title>Cross-cultural assessment of the five-factor model: the revised NEO personality inventory</article-title>. <source>J. Cross-Cult. Psychol.</source> <volume>29</volume>, <fpage>171</fpage>&#x02013;<lpage>188</lpage>. <pub-id pub-id-type="doi">10.1177/0022022198291009</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehta</surname> <given-names>Y.</given-names></name> <name><surname>Majumder</surname> <given-names>N.</given-names></name> <name><surname>Gelbukh</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Recent trends in deep learning based personality detection</article-title>. <source>Artif. Intell. Rev.</source> <volume>53</volume>, <fpage>2313</fpage>&#x02013;<lpage>2339</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-019-09770-z</pub-id></citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noftle</surname> <given-names>E. E.</given-names></name> <name><surname>Robins</surname> <given-names>R. W.</given-names></name></person-group> (<year>2007</year>). <article-title>Personality predictors of academic outcomes: big five correlates of GPA and SAT scores</article-title>. <source>J. Personality Soc. Psychol.</source> <volume>93</volume>, <fpage>116</fpage>&#x02013;<lpage>130</lpage>. <pub-id pub-id-type="doi">10.1037/0022-3514.93.1.116</pub-id><pub-id pub-id-type="pmid">17605593</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Norman</surname> <given-names>W. T.</given-names></name></person-group> (<year>1963</year>). <article-title>Toward an adequate taxonomy of personality attributes: replicated factor structure in peer nomination personality ratings</article-title>. <source>J. Abnormal Soc. Psychol</source>. <pub-id pub-id-type="doi">10.1037/h0040291</pub-id><pub-id pub-id-type="pmid">13938947</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Meara</surname> <given-names>M. S.</given-names></name> <name><surname>South</surname> <given-names>S. C.</given-names></name></person-group> (<year>2019</year>). <article-title>Big five personality domains and relationship satisfaction: direct effects and correlated change over time</article-title>. <source>J. Personality</source> <volume>87</volume>, <fpage>1206</fpage>&#x02013;<lpage>1220</lpage>. <pub-id pub-id-type="doi">10.1111/jopy.12468</pub-id><pub-id pub-id-type="pmid">30776092</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Persson</surname> <given-names>I.</given-names></name> <name><surname>Khojasteh</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Python packages for exploratory factor analysis</article-title>. <source>Struct. Equat. Model</source>. <volume>28</volume>, <fpage>983</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1080/10705511.2021.1910037</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Remaida</surname> <given-names>A.</given-names></name> <name><surname>Moumen</surname> <given-names>A.</given-names></name> <name><surname>El Idrissi</surname> <given-names>Y. E. B.</given-names></name></person-group> (<year>2020</year>). <article-title>Handwriting personality recognition with machine learning: a comparative study</article-title>, in: <source>2020 IEEE 2nd International Conference on Electronics, Control, Optimization and Computer Science (ICECOCS)</source>, p. <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <publisher-loc>Kenitra</publisher-loc>: <publisher-name>IEEE</publisher-name>. <pub-id pub-id-type="doi">10.1109/ICECOCS50124.2020.9314529</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>Z.</given-names></name> <name><surname>Shen</surname> <given-names>Q.</given-names></name> <name><surname>Diao</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>A sentiment-aware deep learning approach for personality detection from text</article-title>. <source>Inf. Process. Manage.</source> <volume>58</volume>, <fpage>102532</fpage>. <pub-id pub-id-type="doi">10.1016/j.ipm.2021.102532</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reynolds</surname> <given-names>S. K.</given-names></name> <name><surname>Clark</surname> <given-names>L. A.</given-names></name></person-group> (<year>2001</year>). <article-title>Predicting dimensions of personality disorder from domains and facets of the five-factor model</article-title>. <source>J. Personal.</source> <volume>69</volume>, <fpage>199</fpage>&#x02013;<lpage>222</lpage>. <pub-id pub-id-type="doi">10.1111/1467-6494.00142</pub-id><pub-id pub-id-type="pmid">11339796</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>P.</given-names></name> <name><surname>Gonz&#x000E0;lez</surname> <given-names>J.</given-names></name> <name><surname>Gonfaus</surname> <given-names>J. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Integrating vision and language in social networks for identifying visual patterns of personality traits</article-title>. <source>Int. J. Soc. Sci. Humanit.</source> <volume>9</volume>, <fpage>6</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.18178/ijssh.2019.V9.981</pub-id></citation>
</ref>
<ref id="B68">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Salminen</surname> <given-names>J.</given-names></name> <name><surname>Rao</surname> <given-names>R. G.</given-names></name> <name><surname>Jung</surname> <given-names>S. G.</given-names></name> <name><surname>Chowdhury</surname> <given-names>S. A.</given-names></name> <name><surname>Jansen</surname> <given-names>B. J.</given-names></name></person-group> (<year>2020</year>). <article-title>Enriching social media personas with personality traits: a deep learning approach using the big five classes</article-title>, in <source>Artificial Intelligence in HCI</source>. <publisher-name>Springer International Publishing</publisher-name>. p. <fpage>101</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-50334-5_7</pub-id></citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samuel</surname> <given-names>D. B.</given-names></name> <name><surname>Widiger</surname> <given-names>T. A.</given-names></name></person-group> (<year>2008</year>). <article-title>A meta-analytic review of the relationships between the five-factor model and DSM-IV-TR personality disorders: a facet level analysis</article-title>. <source>Clin. Psychol. Rev.</source> <volume>28</volume>, <fpage>1326</fpage>&#x02013;<lpage>1342</lpage>. <pub-id pub-id-type="doi">10.1016/j.cpr.2008.07.002</pub-id><pub-id pub-id-type="pmid">18708274</pub-id></citation></ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saulsman</surname> <given-names>L. M.</given-names></name> <name><surname>Page</surname> <given-names>A. C.</given-names></name></person-group> (<year>2004</year>). <article-title>The five-factor model and personality disorder empirical literature: a meta-analytic review</article-title>. <source>Clin. Psychol. Rev.</source> <volume>23</volume>, <fpage>1055</fpage>&#x02013;<lpage>1085</lpage>. <pub-id pub-id-type="doi">10.1016/j.cpr.2002.09.001</pub-id><pub-id pub-id-type="pmid">14729423</pub-id></citation></ref>
<ref id="B71">
<citation citation-type="web"><person-group person-group-type="author"><collab>Scipy Wilcoxon Signed-Rank Test Manual</collab></person-group>. (n.d.). Available online at: <ext-link ext-link-type="uri" xlink:href="https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.wilcoxon.html">https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.wilcoxon.html</ext-link> (accessed November 10, 2021).</citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sengupta</surname> <given-names>S.</given-names></name> <name><surname>Basak</surname> <given-names>S.</given-names></name> <name><surname>Saikia</surname> <given-names>P.</given-names></name> <name><surname>Paul</surname> <given-names>S.</given-names></name> <name><surname>Tsalavoutis</surname> <given-names>V.</given-names></name> <name><surname>Atiah</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A review of deep learning with special emphasis on architectures, applications and recent trends</article-title>. <source>Knowl. Based Syst.</source> <volume>194</volume>, <fpage>105596</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2020.105596</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Spathis</surname> <given-names>D.</given-names></name> <name><surname>Servia-Rodriguez</surname> <given-names>S.</given-names></name> <name><surname>Farrahi</surname> <given-names>K.</given-names></name> <name><surname>Mascolo</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>Passive mobile sensing and psychological traits for large scale mood prediction</article-title>, in <source>Proceedings of the 13th EAI International Conference on Pervasive Computing Technologies for Healthcare</source>. <publisher-loc>PervasiveHealth&#x00027;19. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>272</fpage>&#x02013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1145/3329189.3329213</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>P.</given-names></name> <name><surname>Kretzschmar</surname> <given-names>H.</given-names></name> <name><surname>Dotiwalla</surname> <given-names>X.</given-names></name> <name><surname>Chouard</surname> <given-names>A.</given-names></name> <name><surname>Patnaik</surname> <given-names>V.</given-names></name> <name><surname>Tsui</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Scalability in perception for autonomous driving: waymo open dataset</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-name>IEEE</publisher-name>), p. <fpage>2446</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00252</pub-id></citation>
</ref>
<ref id="B75">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>Q.</given-names></name> <name><surname>Gao</surname> <given-names>L.</given-names></name> <name><surname>Lai</surname> <given-names>Y. K.</given-names></name> <name><surname>Xia</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Variational autoencoders for deforming 3d mesh models</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), p. <fpage>5841</fpage>&#x02013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00612</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Traag</surname> <given-names>V. A.</given-names></name> <name><surname>Waltman</surname> <given-names>L.</given-names></name> <name><surname>Van Eck</surname> <given-names>N. J.</given-names></name></person-group> (<year>2019</year>). <article-title>From Louvain to Leiden: guaranteeing well-connected communities</article-title>. <source>Sci. Rep.</source> <volume>9</volume>, <fpage>5233</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-019-41695-z</pub-id><pub-id pub-id-type="pmid">30914743</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tupes</surname> <given-names>E. C.</given-names></name> <name><surname>Christal</surname> <given-names>R. E.</given-names></name></person-group> (<year>1961</year>). <article-title>Recurrent personality factors based on trait ratings</article-title>. <source>J. Personal</source>. <volume>60</volume>, <fpage>225</fpage>&#x02013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-6494.1992.tb00973.x</pub-id><pub-id pub-id-type="pmid">1635043</pub-id></citation></ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Urban</surname> <given-names>C. J.</given-names></name> <name><surname>Bauer</surname> <given-names>D. J.</given-names></name></person-group> (<year>2021</year>). <article-title>A deep learning algorithm for high-dimensional exploratory item factor analysis</article-title>. <source>Psychometrika</source> <volume>86</volume>, <fpage>1</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-021-09748-3</pub-id><pub-id pub-id-type="pmid">33528784</pub-id></citation></ref>
<ref id="B79">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Walters-Williams</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name></person-group> (<year>2010</year>). <article-title>Comparative study of distance functions for nearest neighbors</article-title>, in <source>Advanced Techniques in Computing Sciences and Software Engineering</source>. <publisher-loc>Netherlands</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>79</fpage>&#x02013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1007/978-90-481-3660-5_14</pub-id></citation>
</ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Forbes</surname> <given-names>M. G.</given-names></name> <name><surname>Gopaluni</surname> <given-names>B.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Song</surname> <given-names>Z.</given-names></name></person-group> (<year>2019</year>). <article-title>Systematic development of a new variational autoencoder model based on uncertain data for monitoring nonlinear processes</article-title>. <source>IEEE Access</source>. <volume>7</volume>, <fpage>22554</fpage>&#x02013;<lpage>22565</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2894764</pub-id></citation>
</ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Widiger</surname> <given-names>T. A.</given-names></name> <name><surname>Lowe</surname> <given-names>J. R.</given-names></name></person-group> (<year>2007</year>). <article-title>Five-factor model assessment of personality disorder</article-title>. <source>J. Personality Assessm.</source> <volume>89</volume>, <fpage>16</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1080/00223890701356953</pub-id><pub-id pub-id-type="pmid">29323509</pub-id></citation></ref>
<ref id="B82">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>J.</given-names></name> <name><surname>Markov</surname> <given-names>K.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep learning based personality recognition from facebook status updates</article-title>, in <source>2017 IEEE 8th International Conference on Awareness Science and Technology (iCAST)</source>. <publisher-loc>Taichung</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>383</fpage>&#x02013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1109/ICAwST.2017.8256484</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Zemel</surname> <given-names>R. S.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>1993</year>). <article-title>Developing population codes for object instantiation parameters</article-title>, in <source>AAAI Fall Symposium Series: Machine Learning in Computer Vision Raleigh.</source> North Carolina USA. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.aaai.org/Papers/Symposia/Fall/1993/FS-93-04/FS93-04-022.pdf">https://www.aaai.org/Papers/Symposia/Fall/1993/FS-93-04/FS93-04-022.pdf</ext-link> (accessed December 20, 2013).</citation>
</ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Wei</surname> <given-names>X. X.</given-names></name></person-group> (<year>2020</year>). <article-title>Learning identifiable and interpretable latent models of high-dimensional neural activity using Pi-VAE</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>33</volume>, <fpage>7234</fpage>&#x02013;<lpage>7247</lpage>.</citation>
</ref>
<ref id="B85">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ziegler</surname> <given-names>M.</given-names></name> <name><surname>Danay</surname> <given-names>E.</given-names></name> <name><surname>Sch&#x000F6;lmerich</surname> <given-names>F.</given-names></name> <name><surname>B&#x000FC;hner</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Predicting academic success with the big 5 rated from different points of view: self-rated, other rated and faked</article-title>. <source>Eur. J. Personal.</source> <volume>24</volume>, <fpage>341</fpage>&#x02013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1002/per.753</pub-id></citation>
</ref>
</ref-list>
</back>
</article>