<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychol.</journal-id>
<journal-title>Frontiers in Psychology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychol.</abbrev-journal-title>
<issn pub-type="epub">1664-1078</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyg.2025.1640864</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A transformer-based embedding approach to developing short-form psychological measures</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Jung</surname> <given-names>Se-Jin</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/3084604/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Seo</surname> <given-names>Jang-Won</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>Department of Psychology, Jeonbuk National University</institution>, <addr-line>Jeonju</addr-line>, <country>Republic of Korea</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Emily Ho, Northwestern University, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Lihua Yao, Feinberg School of Medicine Northwestern University, United States</p>
<p>Anthony Raborn, Vector Psychometric Group LLC, United States</p>
<p>Bj&#x000F6;rn E. Hommel, Leipzig University, Germany</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Jang-Won Seo <email>jwseo&#x00040;jbnu.ac.kr</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>08</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>16</volume>
<elocation-id>1640864</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>06</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>07</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Jung and Seo.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Jung and Seo</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Developing short-form psychological measures is essential for reducing respondent burden, saving time, and conserving resources. However, existing short-form development approaches typically require full-scale administration and rely on factor analysis or machine learning techniques based on response data.</p>
</sec>
<sec>
<title>Methods</title>
<p>This study proposes a novel, data-independent method for item reduction using transformer-based semantic embeddings. Items from the International Personality Item Pool Big-Five Factor Markers (IPIP-50) were embedded using the sentence-t5-xxl model to generate dense semantic representations. These embeddings were clustered via K-means, and representative items were selected based on their proximity to cluster centroids.</p>
</sec>
<sec>
<title>Results</title>
<p>The resulting 30-item short form preserved the original five-factor structure and demonstrated strong psychometric properties. When compared with Classical Test Theory and a Genetic Algorithm, the proposed method achieved comparable levels of reliability, convergent validity, and predictive performance.</p>
</sec>
<sec>
<title>Discussion</title>
<p>These findings highlight the potential of transformer-based embedding approaches for efficient item reduction and item development. The results support the feasibility of a resource-efficient, linguistically grounded alternative to data-dependent reduction methods.</p>
</sec></abstract>
<kwd-group>
<kwd>short-form development</kwd>
<kwd>item reduction</kwd>
<kwd>transformer-based embedding</kwd>
<kwd>semantic clustering</kwd>
<kwd>psychological measures</kwd>
</kwd-group>
<contract-num rid="cn001">4199990714213</contract-num>
<contract-sponsor id="cn001">Jeonbuk National University<named-content content-type="fundref-id">https://doi.org/10.13039/501100015499</named-content></contract-sponsor>
<counts>
<fig-count count="1"/>
<table-count count="6"/>
<equation-count count="0"/>
<ref-count count="32"/>
<page-count count="9"/>
<word-count count="5913"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Quantitative Psychology and Measurement</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>In self-report psychological assessments, an excessive number of items can increase respondent fatigue, leading to inattentive or random responses, which may compromise the reliability and validity of the results (<xref ref-type="bibr" rid="B9">Herzog and Bachman, 1981</xref>). Furthermore, longer instruments require more administration time, posing practical limitations. Conversely, reducing the number of items without adequately representing the construct&#x00027;s structural dimensions may also result in decreased reliability and validity (<xref ref-type="bibr" rid="B29">Smith et al., 2000</xref>). Therefore, developing short forms that maintain both reliability and validity remains a critical objective in psychological measurement.</p>
<p>For these reasons, there has been sustained interest in reducing the number of items in psychological assessments without compromising their psychometric quality. One traditional approach is Classical Test Theory (CTT), which selects items based on their correlations with total scores and evaluates internal consistency using Cronbach&#x00027;s alpha (<xref ref-type="bibr" rid="B20">Lord and Novick, 2008</xref>).</p>
<p>While CTT offers advantages such as computational simplicity and ease of interpretation, it also has notable limitations. The method may retain redundant items due to high inter-item correlations and, because it is based on univariate analysis, may fail to capture the multidimensional nature of psychological constructs (<xref ref-type="bibr" rid="B28">Sijtsma, 2009</xref>).</p>
<p>Principal Component Analysis (PCA) has also been used as a traditional statistical method for item reduction. PCA identifies orthogonal linear combinations of observed variables that account for the maximum variance in the data (<xref ref-type="bibr" rid="B13">Jolliffe, 2002</xref>). Items with high loadings on the first few principal components are often retained to construct a reduced form. However, like CTT, PCA is heavily influenced by inter-item correlations. This reliance can lead to the removal of theoretically important items simply due to statistical redundancy, potentially compromising content validity (<xref ref-type="bibr" rid="B17">Kriegel et al., 2008</xref>; <xref ref-type="bibr" rid="B32">Zheng et al., 2020</xref>).</p>
<p>To address these limitations, Item Response Theory (IRT) has been proposed as a more refined method for item selection and scale construction. Compared to CTT, IRT offers several advantages, including the ability to estimate item parameters independently of the sample, provide item-level information, and model measurement precision across the latent trait continuum (<xref ref-type="bibr" rid="B7">Hambleton et al., 1991</xref>). These properties allow for more flexible and detailed assessments, especially in the development of adaptive or shortened forms.</p>
<p>However, IRT does not systematically explore all possible item combinations or automate the search for the most optimal item sets, leaving room for subjective decisions by researchers during the item reduction process (<xref ref-type="bibr" rid="B31">Yarkoni, 2010</xref>).</p>
<p>To overcome these limitations, recent studies have explored machine learning&#x02013;based approaches for item reduction. For example, <xref ref-type="bibr" rid="B31">Yarkoni (2010)</xref> applied a machine learning technique known as the Genetic Algorithm (GA) to select and reduce questionnaire items. GA is an optimization algorithm inspired by biological evolution, beginning with a randomly generated population and iteratively improving solutions across generations (<xref ref-type="bibr" rid="B10">Holland, 1992</xref>). This method is particularly well suited to identifying optimal solutions for complex problems (<xref ref-type="bibr" rid="B4">Goldberg et al., 1989</xref>). Using GA, Yarkoni successfully shortened existing questionnaires while maintaining internal consistency, thereby demonstrating the potential of machine learning techniques for item reduction in psychological assessment.</p>
<p>However, a major limitation of both traditional and modern item reduction methods is their reliance on response data, which necessitates prior administration of the full questionnaire. Additionally, methods such as the Genetic Algorithm (GA) are probabilistic in nature, meaning that their outcomes can vary across different runs, even under the same conditions (<xref ref-type="bibr" rid="B15">Katoch et al., 2021</xref>). With the exception of manual selection based on expert judgment, very few approaches allow for item reduction without first administering the questionnaire (<xref ref-type="bibr" rid="B12">Howard, 2018</xref>). In response to this limitation, the present study aimed to develop an objective method for item reduction that does not require prior data collection or pilot testing.</p>
<p>Recent advances in artificial intelligence have led to increased interest in Large Language Models (LLMs). LLMs are artificial intelligence systems trained on large-scale text data to understand and generate human-like language (<xref ref-type="bibr" rid="B1">Chang et al., 2024</xref>). To generate meaningful language, these models must first comprehend input text, which involves converting language into a numerical form&#x02014;a process known as embedding.</p>
<p>Embedding refers to the transformation of linguistic information into numerical vector representations. Through this process, computers can computationally process semantic relationships between words, contextual dependencies, and latent meaning structures. A key property of embedding is that words or texts with similar meanings are represented by vectors that are closer in the embedding space (<xref ref-type="bibr" rid="B25">Rodriguez and Spirling, 2022</xref>). By leveraging these semantic distances between vectors, LLMs can effectively interpret the linguistic meaning of text (<xref ref-type="bibr" rid="B25">Rodriguez and Spirling, 2022</xref>).</p>
<p>In recent scale development research, there has been growing interest in large language model (LLM)-based embedding approaches that explore the semantic structure of scale items. Within this emerging trend, studies have applied LLM embeddings such as BERT and SBERT combined with cosine similarity to quantify semantic consistency, reduce semantic redundancy while recovering factor structures, or reliably infer item correlation patterns (<xref ref-type="bibr" rid="B8">Hernandez and Nie, 2023</xref>; <xref ref-type="bibr" rid="B6">Guenole et al., 2024</xref>; <xref ref-type="bibr" rid="B11">Hommel and Arslan, 2024</xref>).</p>
<p>In the present study, we applied transformer-based embedding techniques to reduce questionnaire items without relying on prior response data. By numerically encoding the semantic content of each item and clustering them based on vector proximity, we aimed to generate a data-independent short form.</p>
<p>To evaluate this approach, we applied the method to a widely used personality assessment, the International Personality Item Pool 50-item Big-Five Factor Markers (IPIP-50; <xref ref-type="bibr" rid="B5">Goldberg et al., 2006</xref>). We then tested the validity of the reduced items, their semantic correspondence with the original items, and the method&#x00027;s effectiveness in comparison to other established item reduction techniques.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>2 Methods</title>
<sec>
<title>2.1 Samples</title>
<p>Data were obtained from an online administration of the International Personality Item Pool 50-item Big-Five Factor Markers (IPIP-50), which is publicly available through the Open-Source Psychometrics Project (<ext-link ext-link-type="uri" xlink:href="https://openpsychometrics.org/tests/IPIP-BFFM/">https://openpsychometrics.org/tests/IPIP-BFFM/</ext-link>). Developed based on the publicly available IPIP database, the scale was designed to be freely accessible and is widely used on online platforms. A total of 1,013,558 individual responses were collected.</p>
</sec>
<sec>
<title>2.2 Measures</title>
<sec>
<title>2.2.1 International Personality Item Pool 50-item Big-Five Factor Markers (IPIP-50)</title>
<p>The IPIP-50 is a self-report personality assessment designed to measure the Big Five personality traits (<xref ref-type="bibr" rid="B5">Goldberg et al., 2006</xref>). Grounded in the Big Five theory, the IPIP-50 assesses personality across five dimensions: Extraversion, Agreeableness, Conscientiousness, Emotional Stability, and Openness to Experience. The scale consists of 50 items, with 10 items corresponding to each of the five traits. Items are rated on a 5-point Likert scale indicating the extent to which the statement applies to the respondent.</p>
</sec>
</sec>
<sec>
<title>2.3 Data preprocessing</title>
<p>First, variables irrelevant to the present study (e.g., response time per item) were excluded. Next, cases with missing item responses were identified and removed, resulting in the exclusion of 140,907 responses. To further ensure data quality, Mahalanobis distance was applied to detect multivariate outliers (<xref ref-type="bibr" rid="B16">Kline, 2004</xref>), and 61,208 responses were excluded based on a significance threshold of <italic>p</italic> &#x0003C; 0.01. As a result of these preprocessing steps, a total of 813,226 valid responses remained for analysis.</p>
<p>In preparation for semantic embedding and clustering, reverse-scored items were rephrased into their positively worded counterparts (e.g., &#x0201C;I don&#x00027;t talk a lot&#x0201D; was reworded as &#x0201C;I talk a lot&#x0201D;), and the corresponding response scores were adjusted accordingly to reflect direct scoring.</p>
</sec>
<sec>
<title>2.4 Item reduction</title>
<p>To compare the proposed method with existing item reduction techniques, four approaches were applied to shorten the 50-item IPIP scale to 30 items: (1) a traditional method based on CTT, (2) a machine learning-based method using GA, (3) a factor analytic method using Principal Component Analysis (PCA), and (4) the transformer embedding-based method (TE) developed in the present study.</p>
<sec>
<title>2.4.1 CTT-based item reduction</title>
<p>In accordance with the principles of CTT, items were reduced by calculating the correlation between each item and the total score of its corresponding subscale. Items within each of the five factors were then ranked in descending order based on these correlations. The top six items from each factor were selected, resulting in a 30-item shortened version of the IPIP-50.</p>
</sec>
<sec>
<title>2.4.2 GA-based item reduction</title>
<p>The GA-based item reduction method was implemented using Python and R, based on the approach introduced by <xref ref-type="bibr" rid="B31">Yarkoni (2010)</xref>. First, the Graded Response Model (GRM; <xref ref-type="bibr" rid="B27">Samejima, 1969</xref>) was used to estimate item discrimination and threshold parameters, using the R package &#x0201C;mirt&#x0201D;. Second, a genetic algorithm was applied using the Distributed Evolutionary Algorithms in Python (DEAP) library. The fitness function was defined as the sum of four components: (a) the correlation between the reduced-form and full-scale total scores, (b) internal consistency measured by Cronbach&#x00027;s alpha, (c) average item information at &#x003B8; = 0, and (d) average item discrimination. which served as a basis for evaluating each item&#x00027;s contribution during optimization.</p>
</sec>
<sec>
<title>2.4.3 Principal Component Analysis (PCA) based item reduction</title>
<p>The content was reduced using Principal Component Analysis (PCA). For each subfactor, items with high loadings were identified based on the principal components that explained a substantial portion of the total variance (<xref ref-type="bibr" rid="B22">Moret et al., 2007</xref>; <xref ref-type="bibr" rid="B24">Porter et al., 2016</xref>). Six items were selected from each subfactor to construct a shortened version of the scale.</p>
</sec>
<sec>
<title>2.4.4 Transformer-based item reduction using sentence embeddings</title>
<p>For the transformer-based embedding approach, the model &#x0201C;sentence-t5-xxl&#x0201D; was selected. T5-based sentence embeddings have been reported to achieve over 10% higher correlation scores than BERT-based models on semantic textual similarity (STS) tasks. Furthermore, compared to traditional approaches such as TF-IDF, Word2Vec, and GloVe, T5-based models more accurately capture word order, semantic nuance, and abstract relationships, thereby yielding superior performance in clustering semantically similar items (<xref ref-type="bibr" rid="B23">Ni et al., 2021</xref>). Given its expected superiority in embedding quality, <italic>sentence-t5-xxl</italic>&#x02014;the largest parameterized model within the T5 model series&#x02014;was chosen for this study. It contains a total of 11 billion parameters and is based on the T5-XXL architecture, which consists of 24 transformer layers, 1,024-dimensional hidden states, 128 attention heads, and 65,536-dimensional feed-forward layers built upon Google&#x00027;s T5 (Text-to-Text Transfer Transformer) framework.</p>
<p>To group semantically similar items, Item texts were first embedded using the sentence-t5-xxl model to capture their semantic relationships in vector form. Since the resulting embeddings were high-dimensional, dimensionality reduction was performed using Uniform Manifold Approximation and Projection (UMAP; <xref ref-type="bibr" rid="B21">McInnes et al., 2018</xref>) to enhance the efficiency of clustering and the interpretability of the results. The reduced embeddings were clustered into five groups using K-means, a commonly used clustering algorithm that partitions data into k groups by minimizing the within-cluster variance (<xref ref-type="bibr" rid="B19">Lloyd, 1982</xref>). The centroid of each cluster was then computed, and to ensure both representativeness and content diversity, the six items closest to each centroid (based on Cosine distance) were selected. This procedure resulted in a final set of 30 items, preserving the semantic structure of the original item pool.</p>
</sec>
</sec>
<sec>
<title>2.5 Validation of the proposed method</title>
<sec>
<title>2.5.1 Semantic clustering and alignment evaluation</title>
<p>Each IPIP-50 item was embedded into a high-dimensional semantic vector using a transformer-based model. The embedded items were then grouped into five clusters using K-means clustering. To evaluate the alignment between the resulting semantic clusters and the original Big Five factor labels, the Hungarian algorithm was applied to identify the optimal one-to-one mapping (<xref ref-type="bibr" rid="B18">Kuhn, 1955</xref>). Classification accuracy was calculated by comparing the mapped cluster assignments with the original factor structure. In addition, to assess cluster separability and cohesion, silhouette scores were computed after reducing the embedding space. The silhouette score is a metric that reflects how closely items are grouped within a cluster and how distinctly each cluster is separated from the others (<xref ref-type="bibr" rid="B26">Rousseeuw, 1987</xref>).</p>
</sec>
<sec>
<title>2.5.2 Validation against the original scale</title>
<p>To evaluate the extent to which the reduced items retained the psychological properties and structure of the original instrument, we compared the shortened version developed in this study with the full version of the IPIP-50. Specifically, we examined the internal consistency (Cronbach&#x00027;s &#x003B1;) of the original scale and assessed the correlations between scores derived from the original and reduced item sets.</p>
<p>Additionally, these relationships were visualized using correlation matrices to analyze structural consistency, enabling us to evaluate how well the shortened item set preserved the psychometric properties of the original scale. Such visual analyses are commonly employed in scale validation to assess pattern similarity and structural integrity (<xref ref-type="bibr" rid="B3">Eisenbarth et al., 2015</xref>). The purpose of this analysis was to determine whether reliability and validity could be maintained despite the reduction in the number of items.</p>
</sec>
<sec>
<title>2.5.3 Comparison across item reduction methods</title>
<p>To benchmark the effectiveness of the proposed transformer-based method, we compared it with three alternative approaches: CTT, PCA and a GA. Internal consistency was calculated for each subscale and as an overall average, using the 30-item versions derived from each method.</p>
<p>We then assessed convergent validity by correlating each short form with the full IPIP-50 scale. Finally, predictive performance was evaluated by using each short form to predict item-level scores from the original scale.</p>
<p>Performance was measured using four regression metrics: Mean Absolute Error (MAE), which represents the average absolute difference between predicted and original scores (lower is better); Root Mean Squared Error (RMSE), which emphasizes larger errors by squaring the differences before averaging (lower is better); Coefficient of Determination (<italic>R</italic><sup>2</sup>), which indicates the proportion of variance in the dependent variable explained by the model (higher is better); and Mean Absolute Percentage Error (MAPE), which expresses prediction accuracy as a percentage of the original scores (lower is better).</p>
<p>Predictions were generated using a deep learning model implemented in PyTorch, consisting of three fully connected layers designed to predict continuous outcomes.</p>
</sec>
</sec>
<sec>
<title>2.6 Availability of code</title>
<p>The methods used in this study are documented on Github (<ext-link ext-link-type="uri" xlink:href="https://github.com/sdoublej/teshort/tree/master">https://github.com/sdoublej/teshort/tree/master</ext-link>), where we also provide tools that allow researchers to apply the item reduction procedure themselves. This is intended to enhance the reproducibility and practical utility of the proposed method.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3 Results</title>
<p>The results of clustering using the proposed transformer-based embedding method are presented as a confusion matrix in <xref ref-type="table" rid="T1">Table 1</xref>. Each item from the IPIP-50 was encoded into a semantic vector using a transformer-based model and subsequently grouped via K-means clustering. To evaluate the degree of alignment between the resulting semantic clusters and the original Big Five factor structure, the Hungarian algorithm was applied to find the optimal one-to-one mapping between the semantic clusters and the original factors. The resulting matching yielded an overall accuracy of 96%, and The mean of silhouette score was 0.49.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Confusion matrix between original Big Five factors and semantic clusters obtained via K-means clustering on transformer-based item embeddings.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Cluster Label</bold></th>
<th valign="top" align="center" colspan="5"><bold>Factor</bold></th>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>AGR</bold></th>
<th valign="top" align="center"><bold>CSN</bold></th>
<th valign="top" align="center"><bold>EST</bold></th>
<th valign="top" align="center"><bold>EXT</bold></th>
<th valign="top" align="center"><bold>OPN</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Label: 0</td>
<td/>
<td/>
<td/>
<td valign="top" align="center">9</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">Label: 1</td>
<td/>
<td/>
<td valign="top" align="center">10</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Label: 2</td>
<td/>
<td valign="top" align="center">10</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Label: 3</td>
<td/>
<td/>
<td/>
<td/>
<td valign="top" align="center">9</td>
</tr>
<tr>
<td valign="top" align="left">Label: 4</td>
<td valign="top" align="center">10</td>
<td/>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Each row represents a semantic cluster label obtained through K-means clustering on transformer-based item embeddings, while each column represents one of the original Big Five personality factors (AGR, Agreeableness; CSN, Conscientiousness; EST, Emotional Stability; EXT, Extraversion; OPN, Openness). For example, of the 10 items originally associated with Agreeableness (AGR), 9 were assigned to semantic cluster Label 1 and 1 item was assigned to Label 3, indicating a high degree of semantic coherence within this cluster.</p>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="T2">Table 2</xref> shows the result of item reduction using the proposed method, in which 30 items were selected from five semantic clusters. Items were clearly grouped by their corresponding Big Five subscales within each cluster, indicating strong alignment between semantic structure and the original factor structure. To further assess the quality of this semantic clustering, silhouette scores were examined for each selected item. The average silhouette score was 0.50, suggesting a reasonable degree of cohesion within clusters and separation between them (<xref ref-type="bibr" rid="B26">Rousseeuw, 1987</xref>).</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Items selected through transformer-based embedding for short form construction.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Item_id</bold></th>
<th valign="top" align="left"><bold>Item</bold></th>
<th valign="top" align="center"><bold>Semantic cluster label</bold></th>
<th valign="top" align="center"><bold>Original subscale label</bold></th>
<th valign="top" align="center"><bold>Silhouette score</bold></th>
<th valign="top" align="center"><bold>Distance</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">AGR10</td>
<td valign="top" align="left">I make people feel at ease.</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">AGR</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.40</td>
</tr>
<tr>
<td valign="top" align="left">AGR4</td>
<td valign="top" align="left">I sympathize with others&#x00027; feelings.</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">AGR</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.41</td>
</tr>
<tr>
<td valign="top" align="left">AGR5</td>
<td valign="top" align="left">I am interested in other people&#x00027;s problems.</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">AGR</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.30</td>
</tr>
<tr>
<td valign="top" align="left">AGR6</td>
<td valign="top" align="left">I have a soft heart.</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">AGR</td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.34</td>
</tr>
<tr>
<td valign="top" align="left">AGR8</td>
<td valign="top" align="left">I take time out for others.</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">AGR</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.44</td>
</tr>
<tr>
<td valign="top" align="left">AGR9</td>
<td valign="top" align="left">I feel others&#x00027; emotions.</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">AGR</td>
<td valign="top" align="center">0.24</td>
<td valign="top" align="center">0.36</td>
</tr>
<tr>
<td valign="top" align="left">CSN1</td>
<td valign="top" align="left">I am always prepared.</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">CSN</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.36</td>
</tr>
<tr>
<td valign="top" align="left">CSN10</td>
<td valign="top" align="left">I am exacting in my work.</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">CSN</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.26</td>
</tr>
<tr>
<td valign="top" align="left">CSN2</td>
<td valign="top" align="left">I keep my belongings organized.</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">CSN</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.34</td>
</tr>
<tr>
<td valign="top" align="left">CSN4</td>
<td valign="top" align="left">I keep things tidy.</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">CSN</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.50</td>
</tr>
<tr>
<td valign="top" align="left">CSN5</td>
<td valign="top" align="left">I get chores done right away.</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">CSN</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.42</td>
</tr>
<tr>
<td valign="top" align="left">CSN8</td>
<td valign="top" align="left">I take responsibility for my duties.</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">CSN</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.57</td>
</tr>
<tr>
<td valign="top" align="left">EST1</td>
<td valign="top" align="left">I get stressed out easily.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">EST</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.36</td>
</tr>
<tr>
<td valign="top" align="left">EST4</td>
<td valign="top" align="left">I often feel blue.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">EST</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.35</td>
</tr>
<tr>
<td valign="top" align="left">EST5</td>
<td valign="top" align="left">I am easily disturbed.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">EST</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.32</td>
</tr>
<tr>
<td valign="top" align="left">EST6</td>
<td valign="top" align="left">I get upset easily.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">EST</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.31</td>
</tr>
<tr>
<td valign="top" align="left">EST7</td>
<td valign="top" align="left">I change my mood a lot.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">EST</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.44</td>
</tr>
<tr>
<td valign="top" align="left">EST8</td>
<td valign="top" align="left">I have frequent mood swings.</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">EST</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.44</td>
</tr>
<tr>
<td valign="top" align="left">EXT1</td>
<td valign="top" align="left">I am the life of the party.</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">EXT</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.35</td>
</tr>
<tr>
<td valign="top" align="left">EXT10</td>
<td valign="top" align="left">I am talkative around strangers.</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">EXT</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.35</td>
</tr>
<tr>
<td valign="top" align="left">EXT4</td>
<td valign="top" align="left">I take the lead.</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">EXT</td>
<td valign="top" align="center">0.35</td>
<td valign="top" align="center">0.26</td>
</tr>
<tr>
<td valign="top" align="left">EXT6</td>
<td valign="top" align="left">I have a lot to say.</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">EXT</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.49</td>
</tr>
<tr>
<td valign="top" align="left">EXT7</td>
<td valign="top" align="left">I talk to a lot of different people at parties.</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">EXT</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.45</td>
</tr>
<tr>
<td valign="top" align="left">EXT9</td>
<td valign="top" align="left">I don&#x00027;t mind being the center of attention.</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">EXT</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.50</td>
</tr>
<tr>
<td valign="top" align="left">OPN1</td>
<td valign="top" align="left">I have a rich vocabulary.</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">OPN</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">0.71</td>
</tr>
<tr>
<td valign="top" align="left">OPN2</td>
<td valign="top" align="left">I understand abstract ideas easily.</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">OPN</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">0.69</td>
</tr>
<tr>
<td valign="top" align="left">OPN3</td>
<td valign="top" align="left">I have a vivid imagination.</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">OPN</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">0.49</td>
</tr>
<tr>
<td valign="top" align="left">OPN4</td>
<td valign="top" align="left">I am interested in abstract ideas.</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">OPN</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">0.72</td>
</tr>
<tr>
<td valign="top" align="left">OPN6</td>
<td valign="top" align="left">I have a good imagination.</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">OPN</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.67</td>
</tr>
<tr>
<td valign="top" align="left">OPN7</td>
<td valign="top" align="left">I am quick to understand things.</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">OPN</td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.60</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Item_id, The item_id indicates the original subscale and item number; Distance, Cosine distance between each item&#x00027;s embedding vector and the centroid of its assigned cluster. A smaller distance indicates closer semantic proximity to the cluster center, representing higher representativeness of that item within the group; Silhouette Score, A measure of how well an item fits within its assigned semantic cluster. Higher values indicate greater cohesion within the cluster and better separation from other clusters.</p>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="T3">Table 3</xref> presents the internal consistency and convergent validity of the reduced item set derived using the proposed method. Convergent correlations ranged from 0.95 to 0.98 (<italic>M</italic> = .0.96), and Cronbach&#x00027;s alpha values ranged from 0.73 to 0.85 (<italic>M</italic> = 0.79), all of which indicate acceptable reliability and strong alignment with the original scale. These results support the psychometric adequacy of the proposed transformer-based method in preserving both the factorial structure and measurement quality of the original instrument.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Convergent correlations of factor scores, total score, and Cronbach&#x00027;s alpha between the transformer-based short form and the original IPIP-50 (<italic>N</italic> = 874,434).</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Factor</bold></th>
<th valign="top" align="center"><bold>TE convergent correlations</bold></th>
<th valign="top" align="center"><bold>TE Cronbach&#x00027;s alpha</bold></th>
<th valign="top" align="center"><bold>Original version Cronbach&#x00027;s alpha</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">EXT</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.91</td>
</tr>
<tr>
<td valign="top" align="left">EST</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.88</td>
</tr>
<tr>
<td valign="top" align="left">AGR</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">CSN</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.83</td>
</tr>
<tr>
<td valign="top" align="left">OPN</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.81</td>
</tr>
<tr>
<td valign="top" align="left">Total</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.86</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>TE: Transformer&#x00027;s embedding.</p>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> visualizes the intercorrelations among the original Big Five factors and the correlations between the original and reduced scales. The similarity across panels suggests that the reduced form preserved the factor structure and inter-trait relationships.</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>The left panel shows the intercorrelation matrix among the original Big Five factor scores from the IPIP-50 scale. The right panel presents the correlation matrix between the original full-scale scores and the corresponding short-form scores derived using the proposed method. Diagonal values in the right matrix indicate convergent validity between the original and reduced scales. This visual similarity between the two matrices suggests that the proposed short form preserves the structural pattern of interrelationships among the original Big Five traits; AGR, Agreeableness; CSN, Conscientiousness; EST, Emotional Stability; EXT, Extraversion; OPN, Openness.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpsyg-16-1640864-g0001.tif">
<alt-text>Two heatmaps compare correlation patterns. The left heatmap shows intercorrelations among original scales: EXT, EST, AGR, CSN, OPN. The right heatmap shows correlations between original and short-form scales. Color gradients represent correlation strength from -1.00 (blue) to 1.00 (red).</alt-text>
</graphic>
</fig>
<p><xref ref-type="table" rid="T4">Table 4</xref> presents the convergent validity coefficients between each of the shortened item sets and the original IPIP-50 scale, calculated separately for each of the five personality factors, along with the overall mean convergent validity. The CTT-based and PCA-based short forms selected the same set of items, which may be due to the fact that both methods rely on inter-item correlations.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Convergent validity (r) by factor and mean convergent validity between shortened versions and the original IPIP-50.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Factor</bold></th>
<th valign="top" align="center" colspan="3"><bold>Item reduction method</bold></th>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>CTT/PCA</bold></th>
<th valign="top" align="center"><bold>Ga</bold></th>
<th valign="top" align="center"><bold>TE</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">EXT</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.98</td>
</tr>
<tr>
<td valign="top" align="left">EST</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.97</td>
</tr>
<tr>
<td valign="top" align="left">AGR</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.95</td>
</tr>
<tr>
<td valign="top" align="left">CSN</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.95</td>
</tr>
<tr>
<td valign="top" align="left">OPN</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.95</td>
</tr>
<tr>
<td valign="top" align="left">Mean</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">0.97</td>
<td valign="top" align="center">0.96</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>CTT/PCA, dataset reduced using Classical Test Theory&#x02013;based item selection and Principal Component Analysis; GA, dataset reduced using IRT-informed Genetic Algorithm; TE, dataset reduced using the transformer embedding&#x02013;based method developed in this study. The CTT- and PCA-based methods selected the same set of items.</p>
</table-wrap-foot>
</table-wrap>
<p>The CTT/PCA-based short form yielded an average convergent validity of 0.94, while the GA-based method produced an average of 0.97. The transformer embedding-based method developed in the present study demonstrated an average convergent validity of 0.96. These results indicate that all three item reduction methods maintained strong alignment with the original factor structure, with minor variations in strength across methods.</p>
<p><xref ref-type="table" rid="T5">Table 5</xref> displays the Cronbach&#x00027;s alpha coefficients for the item sets produced by each reduction method&#x02014;CTT, PCA, GA, and TE proposed in this study. The CTT/PCA-based short form yielded the highest average internal consistency (&#x003B1; = 0.83), followed by the TE-based method (&#x003B1; = 0.79), and the GA-based method (&#x003B1; = 0.78). These findings suggest that while all three methods produced reasonably reliable short forms, CTT/PCA resulted in the highest internal consistency among them.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Comparison of Cronbach&#x00027;s alpha by factor across shortened versions.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Factor</bold></th>
<th valign="top" align="center" colspan="4"><bold>Item reduction method</bold></th>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>Origin</bold></th>
<th valign="top" align="center"><bold>CTT/PCA</bold></th>
<th valign="top" align="center"><bold>Ga</bold></th>
<th valign="top" align="center"><bold>TE</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">EXT</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">EST</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">AGR</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.81</td>
</tr>
<tr>
<td valign="top" align="left">CSN</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.76</td>
</tr>
<tr>
<td valign="top" align="left">OPN</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.73</td>
</tr>
<tr>
<td valign="top" align="left">Mean</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.79</td>
</tr></tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="T6">Table 6</xref> presents the predictive performance of each item reduction method in estimating the original subscale scores. In summary, when predicting under the same conditions, the TE method showed predictive performance that was overall comparable to or better than that of the GA-based and CTT/PCA-based methods in terms of error reduction and explanatory power.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Prediction performance for original subscale scores using shortened versions from different item reduction methods.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Metric</bold></th>
<th valign="top" align="center" colspan="3"><bold>Item reduction method</bold></th>
</tr>
<tr style="background-color:#919498;color:#ffffff">
<th/>
<th valign="top" align="center"><bold>CTT and PCA</bold></th>
<th valign="top" align="center"><bold>GA</bold></th>
<th valign="top" align="center"><bold>TE</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">MAE</td>
<td valign="top" align="center">2.71</td>
<td valign="top" align="center">2.70</td>
<td valign="top" align="center">2.57</td>
</tr>
<tr>
<td valign="top" align="left">RMSE</td>
<td valign="top" align="center">3.35</td>
<td valign="top" align="center">3.32</td>
<td valign="top" align="center">3.17</td>
</tr>
<tr>
<td valign="top" align="left">R<sup>2</sup></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">MAPE</td>
<td valign="top" align="center">8.42%</td>
<td valign="top" align="center">8.25%</td>
<td valign="top" align="center">7.97%</td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>MAE, Mean Absolute Error (lower values indicate better fit); RMSE, Root Mean Squared Error (lower is better); R<sup>2</sup>, Coefficient of Determination (higher is better); MAPE, Mean Absolute Percentage Error; (lower is better).</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec sec-type="discussion" id="s4">
<title>4 Discussion</title>
<p>This study aimed to address the limitations of response-dependent short form development methods by proposing a novel item reduction approach using transformer-based semantic embeddings. The proposed method clustered items based on semantic similarity. The Hungarian algorithm was applied to evaluate the alignment between the semantic clusters and the original subfactors, yielding an accuracy of 96%, which reflects a high level of alignment and is considered a strong result for cluster evaluation. Furthermore, the average silhouette score for the selected items was 0.50, suggesting an acceptable level of cohesion within clusters and separation between them. It demonstrates that even without response data, semantic similarity can serve as a reliable basis for organizing psychological items, and suggests the potential for developing short forms that maintain the conceptual integrity of the original scale.</p>
<p>The results demonstrated that applying a transformer-based embedding technique, followed by K-means clustering, effectively grouped items in accordance with their underlying psychological dimensions. These findings suggest that semantic similarity among items was successfully captured through the embedding process, supporting the feasibility of using semantic representations.</p>
<p>Moreover, the short form derived using this method demonstrated high convergent validity with the original scale (correlations &#x0003E; 0.90) and acceptable internal consistency, with Cronbach&#x00027;s alpha coefficients exceeding 0.70 across all factors. Cronbach&#x00027;s alpha values exceeding 0.70 are generally considered acceptable for psychological scales (<xref ref-type="bibr" rid="B30">Taber, 2018</xref>).</p>
<p>Visual analysis of item correspondence further confirmed that the structural patterns observed in the full scale were preserved in the reduced form. These findings collectively suggest that the proposed method preserves the core psychometric properties of the original scale, indicating its potential applicability as an efficient and valid tool for psychological assessment.</p>
<p>In addition, when comparing the proposed transformer-based item reduction method with existing approaches&#x02014;namely, CTT, PCA and GA&#x02014;CTT/PCA yielded the highest internal consistency (Cronbach&#x00027;s &#x003B1;). This outcome is expected, as CTT selects items based on their correlations with total scores, which tends to inflate internal consistency (<xref ref-type="bibr" rid="B2">Cortina, 1993</xref>), And, PCA is also sensitive to inter-item correlations (<xref ref-type="bibr" rid="B14">Jolliffe and Cadima, 2016</xref>), which may contribute to the high internal consistency observed in the PCA-based short form.</p>
<p>In terms of convergent validity, however, all three methods yielded similarly strong results:.96 for the transformer-based method,.94 for CTT/PCA, and 0.97 for GA. When comparing predictive performance using key regression metrics (MAE, RMSE, R<sup>2</sup>, and MAPE), the transformer-based method demonstrated performance comparable to that of the GA- and CTT-based methods overall.</p>
<p>Overall, the item reduction method proposed in this study demonstrated competitive performance in terms of reliability, validity, and predictive accuracy when compared with exiting approaches. These findings suggest that transformer-based embeddings may serve as a valid and practical alternative for developing short forms of psychometric instruments.</p>
<p>The contributions of this study are threefold. First, the clustering of numerically embedded items based solely on their semantic content revealed that the factor structure could be recovered without access to response data. This indicates the method&#x00027;s potential for identifying latent dimensions based on item meaning, offering an alternative analytic approach for future exploratory or confirmatory factor analysis.</p>
<p>Second, the study introduces a novel application of transformer-based sentence embeddings&#x02014;specifically, the sentence-t5-xxl model&#x02014;in the development of short forms. This highlights the feasibility of using state-of-the-art natural language processing (NLP) techniques to inform item selection in psychological measurement.</p>
<p>Third, the proposed method makes use of semantic similarity between items to offer an alternative way to explore item structure before test administration. This approach could help improve time and cost efficiency in scale development and item refinement.</p>
<p>Despite its strengths, the study has several limitations. First, only one transformer model (sentence-t5-xxl) was employed, and the results may vary depending on the embedding model used. Future research should explore the impact of different embedding architectures on item selection and model performance.</p>
<p>Second, reverse-scored items in the original scale were rephrased into positively worded statements to enable semantic embedding and clustering. While this transformation was necessary for consistent vector representation, it may have altered the original semantic intent or psychometric properties of the items. Prior research (<xref ref-type="bibr" rid="B11">Hommel and Arslan, 2024</xref>) suggests that large language models can accurately preserve semantic structure even when negatively worded items are included without rewording. Future studies could therefore explore embedding the original item phrasing without rewording as an alternative. Additionally, to minimize potential researcher bias in the rewording process, objective methods such as automated paraphrasing using the generative capabilities of large language models should be considered.</p>
<p>Third, although the proposed method was compared with CTT, PCA and GA-based approaches, it was not benchmarked against other commonly used item reduction strategies such as Item Response Theory (IRT), factor analysis, or expert judgment. Comparative studies involving a broader range of reduction techniques will be essential to further assess the method&#x00027;s generalizability and relative strengths.</p>
<p>Fourth, while this study primarily focused on internal consistency and basic convergent validity in evaluating the reduced scale, comprehensive scale evaluation should also consider additional indicators, such as model fit, factorial validity, absence of correlated residuals, and especially criterion-related validity, which demonstrates whether the reduced form retains predictive utility. Future research should address these aspects to ensure the robustness of the proposed method.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s6">
<title>Author contributions</title>
<p>S-JJ: Writing &#x02013; original draft, Writing &#x02013; review &#x00026; editing. J-WS: Writing &#x02013; review &#x00026; editing.</p>
</sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research and/or publication of this article. The research received funding from the Brain Korea 21 Fourth Project of the Korea Research 348 Foundation (Jeonbuk National University, Psychology Department No. 4199990714213).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s8">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Zhu</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2024</year>). <article-title>A survey on evaluation of large language models</article-title>. <source>ACM Trans. Intell. Syst. Technol.</source> <volume>15</volume>:<fpage>1</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1145/3641289</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cortina</surname> <given-names>J. M.</given-names></name></person-group> (<year>1993</year>). <article-title>What is coefficient alpha? An examination of theory and applications</article-title>. <source>J. Appl. Psychol.</source> <volume>78</volume>, <fpage>98</fpage>&#x02013;<lpage>104</lpage>. <pub-id pub-id-type="doi">10.1037/0021-9010.78.1.98</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eisenbarth</surname> <given-names>H.</given-names></name> <name><surname>Lilienfeld</surname> <given-names>S. O.</given-names></name> <name><surname>Yarkoni</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>Using a genetic algorithm to abbreviate the psychopathic personality inventory&#x02013;revised (PPI-R)</article-title>. <source>Psychol. Assess.</source> <volume>27</volume>, <fpage>194</fpage>&#x02013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1037/pas0000032</pub-id><pub-id pub-id-type="pmid">25436663</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>D. E.</given-names></name> <name><surname>Korb</surname> <given-names>B.</given-names></name> <name><surname>Deb</surname> <given-names>K.</given-names></name></person-group> (<year>1989</year>). <article-title>Messy genetic algorithms: motivation, analysis, and first results</article-title>. <source>Complex Syst.</source> <volume>3</volume>, <fpage>493</fpage>&#x02013;<lpage>530</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goldberg</surname> <given-names>L. R.</given-names></name> <name><surname>Johnson</surname> <given-names>J. A.</given-names></name> <name><surname>Eber</surname> <given-names>H. W.</given-names></name> <name><surname>Hogan</surname> <given-names>R.</given-names></name> <name><surname>Ashton</surname> <given-names>M. C.</given-names></name> <name><surname>Cloninger</surname> <given-names>C. R.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>The international personality item pool and the future of public-domain personality measures</article-title>. <source>J. Res. Pers.</source> <volume>40</volume>, <fpage>84</fpage>&#x02013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1016/j.jrp.2005.08.007</pub-id><pub-id pub-id-type="pmid">30400846</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guenole</surname> <given-names>N.</given-names></name> <name><surname>D&#x00027;Urso</surname> <given-names>E. D.</given-names></name> <name><surname>Samo</surname> <given-names>A.</given-names></name> <name><surname>Sun</surname> <given-names>T.</given-names></name></person-group> (<year>2024</year>). <article-title>Pseudo factor analysis of language embedding similarity matrices: new ways to model latent constructs</article-title>. <source>OSF</source>. <pub-id pub-id-type="doi">10.31234/osf.io/vf3se</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hambleton</surname> <given-names>R. K.</given-names></name> <name><surname>Swaminathan</surname> <given-names>H.</given-names></name> <name><surname>Rogers</surname> <given-names>H. J.</given-names></name></person-group> (<year>1991</year>). <source>Fundamentals of Item Response Theory</source>. <publisher-loc>Newbury Park, CA</publisher-loc>: <publisher-name>Sage Publications</publisher-name>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hernandez</surname> <given-names>I.</given-names></name> <name><surname>Nie</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>The AI-IP: minimizing the guesswork of personality scale item development through artificial intelligence</article-title>. <source>Pers. Psychol.</source> <volume>76</volume>, <fpage>1011</fpage>&#x02013;<lpage>1035</lpage>. <pub-id pub-id-type="doi">10.1111/peps.12543</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herzog</surname> <given-names>A. R.</given-names></name> <name><surname>Bachman</surname> <given-names>J. G.</given-names></name></person-group> (<year>1981</year>). <article-title>Effects of questionnaire length on response quality</article-title>. <source>Public Opin. Q.</source> <volume>45</volume>, <fpage>549</fpage>&#x02013;<lpage>559</lpage>. <pub-id pub-id-type="doi">10.1086/268687</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Holland</surname> <given-names>J. H.</given-names></name></person-group> (<year>1992</year>). <source>Adaptation in Natural and Artificial Systems: An Introductory Analysis With Applications to Biology, Control, and Artificial Intelligence</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hommel</surname> <given-names>B. E.</given-names></name> <name><surname>Arslan</surname> <given-names>R. C.</given-names></name></person-group> (<year>2024</year>). <article-title>Language models accurately infer correlations between psychological items and scales from text alone</article-title>. <source>PsyArXiv</source>. <pub-id pub-id-type="doi">10.31234/osf.io/kjuce</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Howard</surname> <given-names>M. C.</given-names></name></person-group> (<year>2018</year>). <article-title>Scale pretesting</article-title>. <source>Pract. Assess. Res. Eval.</source> <volume>23</volume>:<fpage>5</fpage>. <pub-id pub-id-type="doi">10.7275/hwpz-jx61</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jolliffe</surname> <given-names>I. T.</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;Principal component analysis for special types of data,&#x0201D;</article-title> in <source>Principal Component Analysis</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>338</fpage>&#x02013;<lpage>372</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jolliffe</surname> <given-names>I. T.</given-names></name> <name><surname>Cadima</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Principal component analysis: a review and recent developments</article-title>. <source>Philos. Trans. R. Soc. A</source> <volume>374</volume>:<fpage>20150202</fpage>. <pub-id pub-id-type="doi">10.1098/rsta.2015.0202</pub-id><pub-id pub-id-type="pmid">26953178</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katoch</surname> <given-names>S.</given-names></name> <name><surname>Chauhan</surname> <given-names>S. S.</given-names></name> <name><surname>Kumar</surname> <given-names>V.</given-names></name></person-group> (<year>2021</year>). <article-title>A review on genetic algorithm: past, present, and future</article-title>. <source>Multimed. Tools Appl.</source> <volume>80</volume>, <fpage>8091</fpage>&#x02013;<lpage>8126</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-020-10139-6</pub-id><pub-id pub-id-type="pmid">33162782</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kline</surname> <given-names>R. B.</given-names></name></person-group> (<year>2004</year>). <source>Principles and Practice of Structural Equation Modeling</source>, <edition>2nd ed</edition>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Guilford Publications</publisher-name>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kriegel</surname> <given-names>H. P.</given-names></name> <name><surname>Kr&#x000F6;ger</surname> <given-names>P.</given-names></name> <name><surname>Schubert</surname> <given-names>E.</given-names></name> <name><surname>Zimek</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;A general framework for increasing the robustness of PCA-based correlation clustering algorithms,&#x0201D;</article-title> in <source>International Conference on Scientific and Statistical Database Management</source> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>418</fpage>&#x02013;<lpage>435</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>H. W.</given-names></name></person-group> (<year>1955</year>). <article-title>The Hungarian method for the assignment problem</article-title>. <source>Nav. Res. Logist. Q.</source> <volume>2</volume>, <fpage>83</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1002/nav.3800020109</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lloyd</surname> <given-names>S.</given-names></name></person-group> (<year>1982</year>). <article-title>Least squares quantization in PCM</article-title>. <source>IEEE Trans. Inf. Theory</source> <volume>28</volume>, <fpage>129</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1109/TIT.1982.1056489</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lord</surname> <given-names>F. M.</given-names></name> <name><surname>Novick</surname> <given-names>M. R.</given-names></name></person-group> (<year>2008</year>). <source>Statistical Theories of Mental Test Scores</source>. <publisher-loc>Charlotte, NC</publisher-loc>: <publisher-name>Information Age Publishing</publisher-name>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McInnes</surname> <given-names>L.</given-names></name> <name><surname>Healy</surname> <given-names>J.</given-names></name> <name><surname>Melville</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Umap: uniform manifold approximation and projection for dimension reduction</article-title>. <source>arXiv</source> [Preprint]. arXiv:1802.03426. <pub-id pub-id-type="doi">10.21105/joss.00861</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moret</surname> <given-names>L.</given-names></name> <name><surname>Nguyen</surname> <given-names>J. M.</given-names></name> <name><surname>Pillet</surname> <given-names>N.</given-names></name> <name><surname>Falissard</surname> <given-names>B.</given-names></name> <name><surname>Lombrail</surname> <given-names>P.</given-names></name> <name><surname>Gasquet</surname> <given-names>I.</given-names></name></person-group> (<year>2007</year>). <article-title>Improvement of psychometric properties of a scale measuring inpatient satisfaction with care: a better response rate and a reduction of the ceiling effect</article-title>. <source>BMC Health Serv. Res.</source> <volume>7</volume>:<fpage>197</fpage>. <pub-id pub-id-type="doi">10.1186/1472-6963-7-197</pub-id><pub-id pub-id-type="pmid">18053170</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ni</surname> <given-names>J.</given-names></name> <name><surname>&#x000C1;brego</surname> <given-names>G. H.</given-names></name> <name><surname>Constant</surname> <given-names>N.</given-names></name> <name><surname>Ma</surname> <given-names>J.</given-names></name> <name><surname>Hall</surname> <given-names>K. B.</given-names></name> <name><surname>Cer</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Sentence-T5: scalable sentence encoders from pre-trained text-to-text models</article-title>. <source>arXiv</source> [Preprint]. arXiv:2108.08877. <pub-id pub-id-type="doi">10.18653/v1/2022.findings-acl.146</pub-id><pub-id pub-id-type="pmid">36568019</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Porter</surname> <given-names>C.</given-names></name> <name><surname>Woo</surname> <given-names>S. E.</given-names></name> <name><surname>Tak</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Developing and validating short form protean and boundaryless career attitudes scales</article-title>. <source>J. Career Assess.</source> <volume>24</volume>, <fpage>162</fpage>&#x02013;<lpage>181</lpage>. <pub-id pub-id-type="doi">10.1177/1069072714565775</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>P. L.</given-names></name> <name><surname>Spirling</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Word embeddings: what works, what doesn&#x00027;t, and how to tell the difference for applied research</article-title>. <source>J. Polit.</source> <volume>84</volume>, <fpage>101</fpage>&#x02013;<lpage>115</lpage>. <pub-id pub-id-type="doi">10.1086/715162</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rousseeuw</surname> <given-names>P. J.</given-names></name></person-group> (<year>1987</year>). <article-title>Silhouettes: a graphical aid to the interpretation and validation of cluster analysis</article-title>. <source>J. Comput. Appl. Math.</source> <volume>20</volume>, <fpage>53</fpage>&#x02013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1016/0377-0427(87)90125-7</pub-id><pub-id pub-id-type="pmid">15760469</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samejima</surname> <given-names>F.</given-names></name></person-group> (<year>1969</year>). <article-title>Estimation of latent ability using a response pattern of graded scores</article-title>. <source>Psychometrika</source> <volume>34</volume>, <fpage>1</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1007/BF03372160</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sijtsma</surname> <given-names>K.</given-names></name></person-group> (<year>2009</year>). <article-title>On the use, the misuse, and the very limited usefulness of Cronbach&#x00027;s alpha</article-title>. <source>Psychometrika</source> <volume>74</volume>, <fpage>107</fpage>&#x02013;<lpage>120</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-008-9101-0</pub-id><pub-id pub-id-type="pmid">20037639</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>G. T.</given-names></name> <name><surname>McCarthy</surname> <given-names>D. M.</given-names></name> <name><surname>Anderson</surname> <given-names>K. G.</given-names></name></person-group> (<year>2000</year>). <article-title>On the sins of short-form development</article-title>. <source>Psychol. Assess.</source> <volume>12</volume>, <fpage>102</fpage>&#x02013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1037/1040-3590.12.1.102</pub-id><pub-id pub-id-type="pmid">10752369</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taber</surname> <given-names>K. S.</given-names></name></person-group> (<year>2018</year>). <article-title>The use of Cronbach&#x00027;s alpha when developing and reporting research instruments in science education</article-title>. <source>Res. Sci. Educ.</source> <volume>48</volume>, <fpage>1273</fpage>&#x02013;<lpage>1296</lpage>. <pub-id pub-id-type="doi">10.1007/s11165-016-9602-2</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yarkoni</surname> <given-names>T.</given-names></name></person-group> (<year>2010</year>). <article-title>Personality in 100,000 words: a large-scale analysis of personality and word use among bloggers</article-title>. <source>J. Res. Pers.</source> <volume>44</volume>, <fpage>363</fpage>&#x02013;<lpage>373</lpage>. <pub-id pub-id-type="doi">10.1016/j.jrp.2010.04.001</pub-id><pub-id pub-id-type="pmid">20563301</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>J.</given-names></name> <name><surname>Fu</surname> <given-names>G.</given-names></name> <name><surname>Anderson</surname> <given-names>K.</given-names></name> <name><surname>Chu</surname> <given-names>H.</given-names></name> <name><surname>Rakovski</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>A 12-lead ECG database to identify origins of idiopathic ventricular arrhythmia containing 334 patients</article-title>. <source>Sci. Data</source> <volume>7</volume>:<fpage>98</fpage>. <pub-id pub-id-type="doi">10.1038/s41597-020-00588-0</pub-id><pub-id pub-id-type="pmid">32251335</pub-id></citation></ref>
</ref-list>
</back>
</article>