<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Educ.</journal-id>
<journal-title>Frontiers in Education</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Educ.</abbrev-journal-title>
<issn pub-type="epub">2504-284X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/feduc.2025.1624827</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Education</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>HSNMF enables accurate and effective analysis for the college students&#x00027; psychological health education data and student life data</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Ma</surname> <given-names>Yuanyuan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/790832/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Lifang</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/3155192/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Computer Engineering, Hubei University of Arts and Science</institution>, <addr-line>Xiangyang</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Hubei Key Laboratory of Power System Design and Test for Electrical Vehicle, Hubei University of Arts and Science</institution>, <addr-line>Xiangyang</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Physics and Electronic Engineering, Hubei University of Arts and Science</institution>, <addr-line>Xiangyang</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jingpu Zhang, Henan University of Urban Construction, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Yingjun Ma, Xiamen University of Technology, China</p>
<p>Ping Yu, Shanxi Normal University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Yuanyuan Ma <email>chonghua_1983&#x00040;126.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>08</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>10</volume>
<elocation-id>1624827</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>05</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>07</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2025 Ma and Liu.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Ma and Liu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>College students face different levels of anxiety, depression, and other psychological problems due to various factors such as academic stress, excess workload, and family responsibilities. The state of mind plays a crucial role in shaping individuals&#x00027; daily behaviors and academic performance. To comprehensively analyze the psychological health status of college students and research domains related to psychological health education, it is urgently needed to develop effective tools and models.</p></sec>
<sec>
<title>Methods</title>
<p>In this study, we proposed a novel framework called hypergraph-induced semi-orthogonal nonnegative matrix factorization (HSNMF). By using this framework, we can effectively evaluate the college students&#x00027; psychological health levels.</p></sec>
<sec>
<title>Results</title>
<p>We implemented the proposed algorithm on two real datasets, and the results showed that the proposed algorithm outperformed other competing methods. The identified research domains provided insights into psychological health education. We also implemented a depression-level classification task on the student life dataset. The results showed that the low-dimensional latent variables learned from HSNMF contained rich semantic information, further improving the performance of traditional machine learning models. Clustering and regression analyses performed on the student life dataset showed that the depression status of students was significantly correlated with their performance in class and social life, as indicated by variables such as &#x0201C;Number of friends (<italic>p</italic>-value = 0.000598),&#x0201D; &#x0201C;Gender (<italic>p</italic>-value = 0.000034),&#x0201D; and &#x0201C;Taking notes in class (<italic>p</italic>-value = 0.03).&#x0201D;</p></sec>
<sec>
<title>Discussion</title>
<p>The significance of student psychological health study is discussed.</p></sec></abstract>
<kwd-group>
<kwd>psychological health education</kwd>
<kwd>matrix factorization</kwd>
<kwd>hypergraph learning</kwd>
<kwd>data visualization</kwd>
<kwd>depress status association analysis</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="0"/>
<equation-count count="19"/>
<ref-count count="32"/>
<page-count count="9"/>
<word-count count="5942"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Mental Health and Wellbeing in Education</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>The psychological wellbeing of college students is of great significance because the state of mind plays a crucial role in shaping individuals&#x00027; daily behaviors and academic performance, consequently affecting their learning outcomes and willingness to engage in educational activities (<xref ref-type="bibr" rid="B9">Jao et al., 2019</xref>). Accurately and timely assessing the psychological health status of students is crucial to ensure the smooth progress of their learning activities and serves as a basis for implementing intelligent psychological health education in colleges (<xref ref-type="bibr" rid="B22">Liu et al., 2020</xref>; <xref ref-type="bibr" rid="B15">Kontoangelos et al., 2020</xref>; <xref ref-type="bibr" rid="B23">Lu, 2022</xref>). The rapid accumulation of data on student education and student behavior provides us with an unprecedented opportunity to analyze the relationships between students&#x00027; psychological health levels and their life behaviors.</p>
<p>Yi et al., established an association between risk behaviors and psychological health and physical activity using two-step clustering and regression analysis and found that a specific behavior cluster was significantly correlated with psychological health and physical activity (<xref ref-type="bibr" rid="B31">Yi et al., 2020</xref>). Opoku Asare et al., utilized smartphone data to analyze the relationships between human behaviors and depression, and identified some behavior markers related to depression (<xref ref-type="bibr" rid="B26">Opoku et al., 2021</xref>). Wang et al., conducted a Student Life study and found significant correlations among the following variables: lifecycle and stress, conversation, and activity (<xref ref-type="bibr" rid="B29">Wang et al., 2014</xref>). These studies have provided some valuable insights into students&#x00027; psychological health problems; however, the clustering performance and extension of algorithms are poor, especially for high-dimensional sparse data analysis tasks.</p>
<p>Recently, non-negative matrix factorization (NMF)-based methods have attracted wide interests in data mining and visualization. Cai et al., proposed a graph regularized non-negative matrix factorization algorithm for data representation and clustering (<xref ref-type="bibr" rid="B3">Cai et al., 2010</xref>). Jiang et al., developed an NMF-based framework to analyze metagenomic data, and identified some canonical sample types (<xref ref-type="bibr" rid="B11">Jiang et al., 2012</xref>). Chavoshinejad et al., proposed an effective semi-supervised NMF algorithm to learn discriminative representations (<xref ref-type="bibr" rid="B5">Chavoshinejad et al., 2023</xref>). These studies have taken full advantages of NMF, and obtained the better part-based representation that can be used for data clustering and visualization. To the best of our knowledge, however, NMF-based methods are seldom used for the students&#x00027; psychological health education data analysis and student life data analysis. Compared to other methods, the advantages of NMF lies in 2 folds: (1) It provides better explanation for many real-world problems. Specifically, NMF factories a non-negative data matrix <inline-formula><mml:math id="M1"><mml:mi>X</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> into two low-rank factor matrices <inline-formula><mml:math id="M2"><mml:mi>W</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M3"><mml:mi>H</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. For the entries in the coefficient matrix <italic>H</italic>, it presents the probability of one sample belonging to a certain cluster. However, for other clustering methods, such as spectral clustering and singular value decomposition (SVD), the factorized low rank matrices may contain negative elements, which are difficult to be viewed as probabilities. (2) NMF cannot only cluster rows in a data matrix, but also simultaneously groups columns in the data matrix. Hence, it can naturally capture associations between samples (students) and features (variables), and is further used to analyze the behaviors of the students.</p>
<p>In this study, we proposed a novel unsupervised learning framework called hypergraph-induced semi-orthogonal non-negative matrix factorization (HSNMF) to analyze the students&#x00027; psychological health education data. HSNMF could evaluate their psychological health status. HSNMF is a versatile toolkit that enables data clustering and visualization and facilitates the analysis of the association between student depression levels and their daily behaviors. Unlike spectral clustering methods based on pairwise interaction relationships, HSNMF uses hypergraphs to encode the high-order interactions between more than two nodes. In addition, a semi-orthogonal constraint on the low-dimensional factor representation matrix ensures the uniqueness and interpretability of the solution. By implementing the HSNMF framework on two real students&#x00027; psychological datasets, we demonstrated the effectiveness of HSNMF in identifying topic domains and student depression levels. HSNMF achieved superior performance in clustering and captured clear and meaningful clustering structures based on the learned similarity matrix. <xref ref-type="fig" rid="F1">Figure 1</xref> shows an overview of the proposed HSNMF framework.</p>
<fig position="float" id="F1">
<label>Figure 1</label>
<caption><p>The flowchart of HSNMF. HSNMF is developed for analyzing students&#x00027; psychological data. <bold>(a)</bold> The original data matrix <italic>X</italic> and sample-sample similarity matrix <italic>S</italic> that were used as inputs of HSNMF. <bold>(b)</bold> Hypergraph construction based the initially sample clustering matrix obtained by louvain algorithm. <bold>(c)</bold> The outputs of HSNMF, including the low-dimensional factor representation matrices <italic>W</italic>, <italic>H</italic>, and the consensus sample-sample similarity matrix <italic>S</italic>. <bold>(d)</bold> Based on the outputs of HSNMF, downstream analysis tasks can be implemented, including clustering, data visualization, and association analysis between depression status and students&#x00027; behaviors.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-10-1624827-g0001.tif">
<alt-text>Three bar graphs labeled a, b, and c compare clustering metrics across four methods: NMF, SC, SNMF, and HSNMF. Graph a shows CHI scores, with HSNMF scoring highest. Graph b displays DBI scores, where NMF is highest. Graph c illustrates Silhouette scores, with HSNMF performing best.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s2">
<title>2 Materials and methods</title>
<sec>
<title>2.1 Datasets and data preprocessing</title>
<p>The college students&#x00027; psychological health education dataset downloaded from China national knowledge infrastructure (CNKI, <ext-link ext-link-type="uri" xlink:href="https://www.cnki.net">https://www.cnki.net</ext-link>), consisted of 865 articles. The terms &#x0201C;college students&#x00027; psychological&#x0201D; and &#x0201C;education&#x0201D; were used as retrieve articles based on the Chinese social sciences citation Index (CSSCI), the Chinese science citation database (CSCD) and the Chinese core journal criterion of PKU. The time range was set between 1994 and 2024. We excluded the articles that were irrelevant to the current research or not an original research article (e.g., reviews, newsletters, or conference reports). Finally, 676 articles with bibliographic data including titles, keywords, authors and publish data were retained, and then these data were imported into the Bibliographic items co-occurrence matrix builder (Bicomb) software (<xref ref-type="bibr" rid="B7">Guo et al., 2015</xref>; <xref ref-type="bibr" rid="B21">Li and Cheng, 2021</xref>). Via Bicomb we extracted keywords from these articles, and generated a term-document matrix and a term-term co-occurrence matrix that were used to conduct downstream analysis tasks.</p>
<p>The student life dataset downloaded from Kaggle site (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets">https://www.kaggle.com/datasets</ext-link>) comprised survey results from 100 computer science students. Demographic data, depression status, and performance variables, such as academic performance, taking notes in class, and presentation frequency, were collected. These variables have been measured before this study and are publicly accessibly (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets">https://www.kaggle.com/datasets</ext-link>).</p>
<p>For the college students&#x00027; psychological health education dataset, we implemented statistical analysis on keywords that appeared in each record, and obtained keyword rankings, in which high frequency indicated close attentions on the corresponding field. Keywords that occurred &#x0003C;2 times across all records were removed. For student life dataset, the Python toolkit package Scikit-learn was used to transform numerical variables and categorical variables.</p></sec>
<sec>
<title>2.2 Construction of hypergraph</title>
<p>Different from traditional graph where an edge can only connect to two nodes, in hypergraph an edge can connect more than two nodes. This mechanism of hypergraph effectively handles with information loss problem. For example, to group a set of articles into different topics, the common practice is to first construct a pairwise interaction network (simple graph) where two nodes are linked with an edge if there is at least one common author writes them, and then graph clustering method is implemented on this graph to obtain the final cluster assignments (<xref ref-type="bibr" rid="B32">Zhou et al., 2006</xref>). However, the graph constructing strategy above obviously ignore some useful information when the same author contributes more than two articles. Such unexpected lost information is important to clustering or knowledge findings.</p>
<p>A natural approach to address information loss issue is to use hypergraph to organize data relationships. <xref ref-type="fig" rid="F2">Figure 2</xref> gives an illustrative example to construct hypergraph. Given the weighted hypergraph <italic>G</italic> &#x0003D; (<italic>V, E, W</italic>), where <italic>V</italic> is the set of nodes and <italic>E</italic> denotes the of hyperedges. For each hyperedge <italic>e</italic>, we used <italic>w</italic>(<italic>e</italic>) to represent its weight. In this manuscript, the gaussian kernel function is used to compute the weights of hyperedges. The incidence matrix <italic>P</italic>&#x02208;<italic>R</italic><sup>|<italic>V</italic>| &#x000D7; |<italic>E</italic>|</sup> corresponding to <italic>G</italic> with entry <italic>p</italic>(<italic>v, e</italic>) is defined as follows:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M4"><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>v</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>e</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>v</mml:mi><mml:mo>&#x02209;</mml:mo><mml:mi>e</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where |<italic>V</italic>| and |<italic>E</italic>| represent the number of nodes and hyperedges, respectively. The degree of node <italic>v</italic> is defined as <inline-formula><mml:math id="M5"><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>E</mml:mi></mml:mrow></mml:munder><mml:mi>w</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. The degree of hyperedge <italic>e</italic> is defined as <inline-formula><mml:math id="M6"><mml:mi>&#x003B4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>V</mml:mi></mml:mrow></mml:munder><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Let <italic>D</italic><sub><italic>v</italic></sub> and <italic>D</italic><sub><italic>e</italic></sub> denote degree matrices of nodes and hyperedges, respectively. The hypergraph Laplacian matrix <italic>L</italic><sub><italic>hg</italic></sub> can be defined as follows:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>P</mml:mi><mml:mi>W</mml:mi><mml:msubsup><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Note that in this manuscript we used Louvain community detection algorithm (<xref ref-type="bibr" rid="B1">Blondel et al., 2008</xref>) instead of <italic>k-</italic>nearest neighbors (KNN) to generate hyperedges, i.e., each cluster represents a hyperedge. By using this strategy, noisy information or outliers are filtered out to some extent. The hypergraph captures the high-order interaction among nodes. The hypergraph regularization can be defined as follows:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>O</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>E</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02208;</mml:mo><mml:mi>e</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mfrac><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mi>H</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here, <italic>H</italic> denotes the low-dimensional representation matrix of nodes.</p>
<fig position="float" id="F2">
<label>Figure 2</label>
<caption><p>An illustrative example for hypergraph organization. <bold>(a)</bold> Hyperedges contains more than two nodes in hypergraph. <italic>v</italic><sub><italic>i</italic></sub> denotes <italic>i</italic>th nodes, <italic>e</italic><sub><italic>j</italic></sub> denotes <italic>j</italic>th hyperedge. <bold>(b)</bold> The incidence matrix of hypergraph. The entry (<italic>v</italic><sub><italic>i</italic></sub>, <italic>e</italic><sub><italic>j</italic></sub>) is set to be 1 when <italic>v</italic><sub><italic>i</italic></sub> belongs to <italic>e</italic><sub><italic>j</italic></sub>, and 0 otherwise.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-10-1624827-g0002.tif">
<alt-text>Diagram illustrating a hypergraph and corresponding incidence matrix. Part a shows a hypergraph with vertices (v_1) to (v_7) connected by hyperedges (e_1) to (e_6). Part b, the incidence matrix, displays connections between vertices and hyperedges, with entries of 1 indicating an edge connection.</alt-text>
</graphic>
</fig>
<p>Next, we will introduce the proposed HSNMF algorithm by integrating hypergraph into this objective.</p></sec>
<sec>
<title>2.3 HSNMF model</title>
<p>To identify the college students&#x00027; psychological health education research topics, we introduce hypergraph induced semi-orthogonal non-negative matrix factorization (HSNMF) model. Let <inline-formula><mml:math id="M9"><mml:mi>X</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the original data matrix, HSNMF aims to learn two low-dimensional representation matrices <inline-formula><mml:math id="M10"><mml:mi>W</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> (basis matrix) and <inline-formula><mml:math id="M11"><mml:mi>H</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> (coefficient matrix), and a feature similarity matrix <inline-formula><mml:math id="M12"><mml:mi>S</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, where <italic>p, n</italic> denote the numbers of features and samples, respectively. <italic>k</italic>&#x0226A;<italic>min</italic>(<italic>p, n</italic>) is the rank of factorized matrix. The objective function of HSNMF can be written as follows:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M13"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:munder><mml:mrow><mml:mi>min</mml:mi></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>H</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>S</mml:mi></mml:mrow></mml:munder><mml:mtext>&#x000A0;</mml:mtext><mml:mi>J</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>W</mml:mi><mml:mi>H</mml:mi></mml:mrow><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x003B1;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>W</mml:mi><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x003B3;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:msubsup><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>I</mml:mi></mml:mrow><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;s.t.</mml:mtext><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>H</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B1;</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B2;</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>&#x003B3;</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x02265;</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x000A0;</mml:mo><mml:mi>S</mml:mi><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mn>1</mml:mn></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mn>1</mml:mn></mml:mstyle><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>I</italic> is identity matrix, <bold>1</bold> is a column vector with all its elements to be 1s. <italic>S</italic> is the learned feature similarity matrix that can be used for clustering and data visualization. &#x003B1;, &#x003B2; and &#x003B3; are hyperparameters to be tuned. We will discuss how to select their values in the later section.</p>
<p>In the object of HSNMF, the first term, <inline-formula><mml:math id="M14"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>X</mml:mi><mml:mo>-</mml:mo><mml:mi>W</mml:mi><mml:mi>H</mml:mi><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, is standard non-negative matrix factorization (NMF) loss function for student psychological education or student life data. The second term, <inline-formula><mml:math id="M15"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>S</mml:mi><mml:mo>-</mml:mo><mml:mi>W</mml:mi><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, is a consensus graph factorization strategy which regularizes kernel <italic>WW</italic><sup><italic>T</italic></sup> toward a consensus graph <italic>S</italic>, and generating a meaningful factorization. Through iterative training, the first two terms in <xref ref-type="disp-formula" rid="E4">Equation 4</xref> can learned the low-dimensional representation for samples and features, however, the high-order relationships among nodes may be ignored. For example, the depression level of student may be correlated to the mixed effects of academic pressure, numerous work, and family responsibilities (<xref ref-type="bibr" rid="B26">Opoku et al., 2021</xref>; <xref ref-type="bibr" rid="B29">Wang et al., 2014</xref>). Hence, modeling high-order interactions among variables with hypergraph is important to mine the latent feature associations. Therefore, we include the third term, <inline-formula><mml:math id="M16"><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, in HSNMF model. The fourth term, <inline-formula><mml:math id="M17"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, encourages <italic>W</italic> to be column-orthogonal. One of the advantages is the uniqueness and interpretability of solution <italic>W</italic>. The term, <italic>S</italic><bold>1&#x0003D;1</bold>, is a constraint term on <italic>S</italic> that enforces each row of <italic>S</italic> to have summation close to 1.</p>
<p>Unlike other clustering methods based on spectral graph theory including spectral clustering and its variants which used the eigenvectors corresponding to large eigenvalues to conduct clustering (<xref ref-type="bibr" rid="B30">White and Smyth, 2005</xref>; <xref ref-type="bibr" rid="B25">Ng et al., 2001</xref>; <xref ref-type="bibr" rid="B18">Law et al., 2017</xref>), HSNMF adopts hypergraph Laplacian to explore the complicated interaction relationships among variables. The low-dimensional factor matrices obtained from HSNMF own stronger representation ability. In addition, orthogonal constraints on basis matrix <italic>W</italic> leads to better clustering solution and interpretability: the columns in <italic>W</italic> will tend to be sparse. The optimization algorithm of HSNMF is presented in the next subsection.</p></sec>
<sec>
<title>2.4 Optimization of HSNMF model</title>
<p>The optimal problem of objective function (<xref ref-type="disp-formula" rid="E4">Equation 4</xref>) can be divided into three sub-problems and can be solved alternately.</p>
<sec>
<title>2.4.1 Fixing <italic>W</italic> and <italic>S</italic>, updating <italic>H</italic></title>
<p>For <italic>H</italic>, the constrained optimization problem of <xref ref-type="disp-formula" rid="E4">Equation 4</xref> can be solved easily by multiplicative update rule as traditional NMF did (<xref ref-type="bibr" rid="B19">Lee and Seung, 2000</xref>, <xref ref-type="bibr" rid="B20">1999</xref>). Based on trace operation, <xref ref-type="disp-formula" rid="E4">Equation 4</xref> can be rewritten as follows:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M18"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:mi>H</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>H</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:mi>H</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x003B1;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>S</mml:mi></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>S</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi>W</mml:mi><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x003B3;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x003B3;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>W</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>I</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We only consider the terms related to <italic>H</italic>, and introduced Lagrange multiplier &#x003A6;<sup>(1)</sup> to solve the optimal problem. Taking the partial derivatives of <italic>L</italic> with respect to <italic>H</italic> gives:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>H</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>X</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:mi>H</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003A6;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Using KKT condition, we can obtain the following updating rule for <italic>H</italic>:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:mi>H</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></sec>
<sec>
<title>2.4.2 Fixing <italic>H</italic> and <italic>S</italic>, updating <italic>W</italic></title>
<p>Similarly, we can obtain the updating rule for <italic>W</italic>:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:msup><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:msup><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:msubsup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo></mml:mrow></mml:msubsup><mml:mi>W</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:mi>W</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mi>H</mml:mi><mml:msup><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>W</mml:mi><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:msubsup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msubsup><mml:mi>W</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B3;</mml:mi><mml:mi>W</mml:mi><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M23"><mml:msubsup><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>.</p></sec>
<sec>
<title>2.4.3 Fixing <italic>W</italic> and <italic>H</italic>, updating <italic>S</italic></title>
<disp-formula id="E9"><label>(9)</label><mml:math id="M24"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>W</mml:mi><mml:msup><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003D1;</mml:mi><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>S</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003D1;</mml:mi><mml:mi>S</mml:mi><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mn>1</mml:mn></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec></sec>
<sec>
<title>2.5 Parameter selection</title>
<p>In HSNMF, there are three parameters &#x003B1;, &#x003B2; and &#x003B3; that need to be determined. First, we used NNDSVD (<xref ref-type="bibr" rid="B2">Boutsidis and Gallopoulos, 2008</xref>) to solved the optimization problem <inline-formula><mml:math id="M25"><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>X</mml:mi><mml:mo>-</mml:mo><mml:mi>W</mml:mi><mml:mi>H</mml:mi><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> and obtain the initial solutions <inline-formula><mml:math id="M26"><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M27"><mml:mover accent="false"><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>. Second, we used Ochiai coefficients (<xref ref-type="bibr" rid="B28">Vancraeynest et al., 2024</xref>; <xref ref-type="bibr" rid="B13">Kalgotra et al., 2020</xref>) to set the initial value <inline-formula><mml:math id="M28"><mml:mover accent="false"><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> of <italic>S</italic>. Finally, &#x003B1;, &#x003B2; and &#x003B3; are set as:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M29"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003B1;</mml:mi><mml:mo>=</mml:mo><mml:mo>||</mml:mo><mml:mi>X</mml:mi><mml:mo>-</mml:mo><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mover accent="false"><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>||</mml:mo><mml:mover accent="false"><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:msup><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E11"><label>(11)</label><mml:math id="M30"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003B2;</mml:mi><mml:mo>=</mml:mo><mml:mo>||</mml:mo><mml:mi>X</mml:mi><mml:mo>-</mml:mo><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mover accent="false"><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mi>W</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E12"><label>(12)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003B3;</mml:mi><mml:mo>=</mml:mo><mml:mo>||</mml:mo><mml:mi>X</mml:mi><mml:mo>-</mml:mo><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mover accent="false"><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>||</mml:mo><mml:msup><mml:mrow><mml:mover accent="false"><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo><mml:mtext>&#x000A0;</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Noting that for the sake of fairness we adopted parameter selection rules defined in <xref ref-type="disp-formula" rid="E10">Equations 10</xref>&#x02013;<xref ref-type="disp-formula" rid="E12">12</xref> across all the experiments. The ablation experiments in the following subsection also demonstrated the effectiveness of the objection of HSNMF algorithm.</p></sec>
<sec>
<title>2.6 Evaluation metrics</title>
<p>Unsupervised clustering metric silhouette coefficient (<xref ref-type="bibr" rid="B14">Kaufman and Rousseeuw, 2009</xref>), Calinski-Harabasz index (CHI; <xref ref-type="bibr" rid="B12">Jos&#x000E9;-Garc&#x000ED;a and G&#x000F3;mez-Flores, 2023</xref>; <xref ref-type="bibr" rid="B4">Cali&#x00144;ski and Harabasz, 1974</xref>), and Davies-Bouldin index (DBI; <xref ref-type="bibr" rid="B12">Jos&#x000E9;-Garc&#x000ED;a and G&#x000F3;mez-Flores, 2023</xref>; <xref ref-type="bibr" rid="B6">Davies and Bouldin, 1979</xref>) are used to evaluate the performance of the clustering methods.</p>
<p>Let <italic>a</italic>(<italic>i</italic>) denote the average distance of data point <italic>i</italic> to all other data within the same cluster with <italic>i</italic>, and <italic>b</italic>(<italic>i</italic>) denote the average distance of <italic>i</italic> to all data points to the neighboring cluster, i.e., the smallest average distance to the cluster of <italic>i</italic>. The silhouette coefficient for data point <italic>i</italic> is defined as:</p>
<disp-formula id="E13"><label>(13)</label><mml:math id="M32"><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mi>b</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>b</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mrow><mml:mi>b</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x0003E;</mml:mo><mml:mi>b</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>High silhouette score indicates good clustering performance. The average values of silhouette scores of all the data points are reported.</p>
<p>Let <italic>n</italic>, <italic>k</italic> denote the number of data points and clusters, respectively, the CHI is defined as follows:</p>
<disp-formula id="E14"><label>(14)</label><mml:math id="M33"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>H</mml:mi><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>||</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mo>||</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>n</italic><sub><italic>i</italic></sub> is the number of data points belonging to <italic>i</italic>th cluster, <italic>C</italic><sub><italic>i</italic></sub> is the data set belonging to <italic>i</italic>th cluster, <italic>c</italic><sub><italic>i</italic></sub> is the centroid of <italic>i</italic>th cluster, <italic>c</italic> is the centroid of all data points.</p>
<p>The DBI is defined as follows:</p>
<disp-formula id="E15"><label>(15)</label><mml:math id="M34"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mi>B</mml:mi><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E16"><label>(16)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where |<italic>C</italic><sub><italic>i</italic></sub>| denotes the number of data points in <italic>C</italic><sub><italic>i</italic></sub>, <italic>d</italic>(<italic>x, c</italic><sub><italic>i</italic></sub>) denotes the distance between <italic>x</italic> and <italic>c</italic><sub><italic>i</italic></sub>.</p></sec></sec>
<sec id="s3">
<title>3 Results and discussion</title>
<sec>
<title>3.1 HSNMF achieves superior performance on the college students&#x00027; psychological health education dataset</title>
<p>In the college students&#x00027; psychology health education dataset, we used the term-document matrix and the term-term co-occurrence matrix as inputs for HSNMF, and obtained the low-dimensional representation matrices <italic>W</italic>, <italic>H</italic> and a similarity matrix <italic>S</italic>. Clustering and data visualization were implemented using <italic>S</italic>. We compared the HSNMF algorithm with other competing methods, including NMF (<xref ref-type="bibr" rid="B19">Lee and Seung, 2000</xref>, <xref ref-type="bibr" rid="B20">1999</xref>), spectral clustering (SC; <xref ref-type="bibr" rid="B30">White and Smyth, 2005</xref>; <xref ref-type="bibr" rid="B10">Jia et al., 2014</xref>), and symmetric non-negative matrix factorization (SNMF; <xref ref-type="bibr" rid="B16">Kuang et al., 2012</xref>; <xref ref-type="bibr" rid="B24">Ma et al., 2020</xref>), for the topic analysis of students&#x00027; psychological health research. For SC, we first constructed similarity matrix with the cosine function, and then implemented spectral clustering on the similarity matrix. For HSNMF, we first constructed a <italic>k</italic>-nearest neighbor (KNN) graph based on the learned similarity matrix <italic>S</italic>, and then implemented Louvain clustering on the KNN graph.</p>
<p>The clustering performance evaluated by CHI, DBI and silhouette scores are presented in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig position="float" id="F3">
<label>Figure 3</label>
<caption><p>Comparison of clustering performance in the college students&#x00027; psychological health education data. <bold>(a)</bold> Assessment of clustering performance in terms of CHI. <bold>(b)</bold> Assessment of clustering performance in terms of DBI. <bold>(c)</bold> Assessment of clustering performance in terms of the averaged silhouette scores. These three unsupervised metrics were computed based on the original term-document matrix and cluster results obtain from each method. For CHI and silhouette score metrics, high values indicate good clustering performance. For DBI, The smaller the values of the DBI, the better the performance of the algorithms.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-10-1624827-g0003.tif">
<alt-text>Panel a shows a heatmap titled &#x0201D;Visualization on W&#x0201D; with terms as rows and topics as columns. Colors range from blue to green, indicating values from -2 to 2. Panel b displays another heatmap titled &#x0201D;Visualization on S&#x0201D; with terms on both axes. Colors range from blue to yellow, representing a different value scale. Both panels include labeled color bars on the right side.</alt-text>
</graphic>
</fig>
<p>As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, HSNMF achieved the best performance on the college students&#x00027; psychological health education dataset in terms of CHI, DBI and silhouette scores. NMF achieves the second-best performance in terms of CHI. SC also performed well in terms of silhouette score. The results demonstrate that the proposed HSNMF algorithm is effective on students&#x00027; psychological health education dataset. One of the possible reasons is that the introducing hypergraph regularization term and the semi-orthogonal constraints on the low-dimensional representation <italic>W</italic> into the HSNMF objective function.</p></sec>
<sec>
<title>3.2 Ablation study</title>
<p>We further validate the effectiveness of HSNMF via ablation experiments. We set &#x003B1;, &#x003B2; and &#x003B3; to 0 in turn. When &#x003B1; &#x0003D; 0, the values of CHI and DBI were 2.7e-5 and 45.87, respectively. When &#x003B2; &#x0003D; 0, the values of CHI and DBI were 0.0075 and 17.51, respectively. When &#x003B3; &#x0003D; 0, the values of CHI and DBI were 0.0082 and 19.05, respectively, and the averaged silhouette score was 0.7206. The results demonstrate that the effectiveness of introducing the hypergraph regularization term, consensus graph factorization strategy, and semi-orthogonal constraints on the columns of <italic>W</italic> into the object of the HSNMF algorithm.</p></sec>
<sec>
<title>3.3 HSNMF facilitates visualization for topic terms of psychological health education fields</title>
<p>HSNMF can not only be used to cluster, but also be used to data visualization. Based on the learned the low-dimensional factor matrix <italic>W</italic>, the term similarity matrix <italic>S</italic>, and the clustering indices of all terms, we implemented visualization analysis, and the term clusters were represented in <xref ref-type="fig" rid="F4">Figures 4a</xref>, <xref ref-type="fig" rid="F4">b</xref>. We can see that the learned the low-dimensional factor matrix <italic>W</italic> and similarity matrix <italic>S</italic> had clear clustering structures. The terms in pink box in the figure consist of research topics related to psychological health education. As <xref ref-type="fig" rid="F4">Figure 4</xref> shown, <italic>W</italic> and <italic>S</italic> had consistent clustering structures in some topics.</p>
<fig position="float" id="F4">
<label>Figure 4</label>
<caption><p>Visualization of terms in the psychological health education research fields. <bold>(a)</bold> Visualization on <italic>W</italic>. <bold>(b)</bold> Visualization on <italic>S</italic>. The ranges in pink boxes indicate five main research topics in the psychological health education research fields.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-10-1624827-g0004.tif">
<alt-text>Diagram illustrating a process in four parts (a-d). Part a shows matrices labeled (X) and (S), representing samples and features. Part b illustrates a sample classification and hypergraph. Part c presents matrices of factors, with the note &#x0201D;HSNMF.&#x0201D; Part d displays a clustering heatmap, visualization matrix, and association analysis with factors like gender, number of friends, like presentation, academic performance, and depression status.</alt-text>
</graphic>
</fig>
<p>We conducted further investigation into the five topics and found that they primarily focused on the following several fields: 1) Psychological health problem, including psychological characteristics, psychological capital, and psychological stress. 2) Psychological health education and teaching, including psychological health education for college students, teaching reform, psychological counseling, art education, curriculum system, moral education, frustration education. 3) Strategy and service system, including growth path, student psychological health services, and education strategy.</p></sec>
<sec>
<title>3.4 HSNMF identified the latent associations between depression levels and student class behavior</title>
<p>We further validated the performance of HSNMF on student life dataset. <xref ref-type="fig" rid="F5">Figures 5a</xref>, <xref ref-type="fig" rid="F5">b</xref> show the visualization results for the original data matrix and the learned low-dimensional factor matrix <italic>W</italic> (the averaged silhouette score equals 0.9128). We can see that the individuals with different depression levels show the clearly clustering structure (<xref ref-type="fig" rid="F5">Figure 5b</xref>). Additional analysis was conducted on the factor matrices <italic>W</italic> and <italic>H</italic>, where we inspected the rows of <italic>H</italic> with the large entries, and identified several features related to student depression status. The experimental results showed that three kinds of behaviors including &#x0201C;Academic performance,&#x0201D; &#x0201C;Taking note in class&#x0201D; and &#x0201C;Number of friends&#x0201D; were related to student depression status (depression). &#x0201C;Gender,&#x0201D; &#x0201C;Taking note in class&#x0201D; and &#x0201C;Like presentation&#x0201D; are related to student depression status (sometimes depression).</p>
<fig position="float" id="F5">
<label>Figure 5</label>
<caption><p>Analysis on student life data. <bold>(a)</bold> Visualization on the original student life dataset. Row represents the individual, and column represents the student behavior. <bold>(b)</bold> Visualization on the low-dimensional factor matrix <italic>W</italic>. Column represents the latent depression levels of students. <bold>(c)</bold> Performance assessment of different methods in terms of accuracy and f1 score. <bold>(d)</bold> Macro-average ROC curves of different methods.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="feduc-10-1624827-g0005.tif">
<alt-text>Panels a and b display heatmaps representing individual performance and depression levels, respectively, with a color gradient from blue (low) to green (high). Panel c shows a bar chart assessing performance with accuracy and F1 scores for methods: HS-NMF combined with SVC, LR, LightGBM, XGBoost, and Random Forest. Panel d features an ROC curve comparison on student life data, highlighting areas under the curve for the same methods.</alt-text>
</graphic>
</fig>
<p>Next, we implemented an association analysis between depression levels and student class behaviors based on student life data. Logistic regression analysis (LogisticRegression function in Sklearn package) was implemented to measure the associations between student depression level and behavior variables. The regression coefficients for features &#x0201C;Number of friends,&#x0201D; &#x0201C;Gender&#x0201D; and &#x0201C;Taking note in class&#x0201D; are &#x02212;1.32 (<italic>p</italic>-value = 0.000598), &#x02212;0.97(<italic>p</italic>-value = 0.000034) and &#x02212;0.27(<italic>p</italic>-value = 0.03), respectively. The results indicate that students&#x00027; depression level is significantly correlated to their class performance and social communication strength (number of friends).</p>
<p>HSNMF can not only be used for sample clustering, but also be used for classification. We further validate the effectiveness of HSNMF via cross-validation where the original data are divided into two parts. The one is used to train model, the other is used to test. The results were presented in <xref ref-type="fig" rid="F5">Figures 5c</xref>, <xref ref-type="fig" rid="F5">d</xref>. We can see that &#x0201C;HSNMF &#x0002B; random forest&#x0201D; achieves the best performance compared to other competing methods in terms of accuracy, F1, and AUC metrics. Different from random forest, we used the low-dimensional factor matrix <italic>W</italic> obtained from HSNMF to train classification models, and obtain the better performance (random forest, accuracy 0.85). The results demonstrate that integrating HSNMF into the traditional machine learning algorithms can effectively improves the model&#x00027;s performance. One of the possible reasons is that the learned low-dimensional latent variable <italic>W</italic> contain more semantic information that can guide the training process for more effective classification tasks.</p>
<p>Although HSNMF achieves better performance in terms of accuracy, F1, and AUC metrics, it still needs to be further validated in more student life data. In this manuscript, we aimed to develop a novel algorithm to analyze the students&#x00027; behaivour, and explored the relationships between depression levels and student class behaviors. For other questions, such as population study and behavior variables measurement, they are beyond the scope of our research.</p></sec></sec>
<sec id="s4">
<title>4 Conclusion</title>
<p>The rapid accumulation of student education data and student behavior data provides us an unprecedented chance to analyze the relationships between students&#x00027; psychological health levels and their life behaviors. To this end, we proposed an effective analysis framework HSNMF, which utilizes hypergraph to learn the low-dimensional factor representation and sample-sample similarity matrix. One advantage of hypergraph lies that it can encode the high-order interaction relationships between objects, thus leading to better clustering qualities. Extensive experiments were implemented on two real datasets. The experimental results showed that the proposed HSNMF algorithm achieved the best performance compared to other competing methods. We also implemented depression level classification task on student life dataset by cross-validation experiments. The experimental results showed that integrating HSNMF into the traditional machine learning models can effectively improve their performance, which indicated that the low-dimensional latent variables learned from HSNMF may contain more semantic information. In terms of the versatileness of HSNMF, we can use it in various fields for different populations and purposes, such as informetric, knowledge graph, and so on.</p>
<p>Analysis on the college students&#x00027; psychology health education data identified several meaningful topic domains, including psychological health problem, psychological health education and teaching, and the psychological teaching strategies and service system. Clustering analysis and regression analysis on the student life dataset showed student depression status is significantly correlated to their performance in class and social life, such as &#x0201C;Number of friends (<italic>p</italic>-value = 0.000598),&#x0201D; &#x0201C;Gender (<italic>p</italic>-value = 0.000034)&#x0201D; and &#x0201C;Taking note in class (<italic>p</italic>-value = 0.03).&#x0201D;</p>
<p>The college students&#x00027; psychological health is one of the most important problems in current higher education. This study aimed to develop a novel framework to group students with various mental statuses into different clusters and further identify the latent associations between depression status and behavior variables based on the college student life data. The conclusions from association analysis help to shine a light on student life, and help teachers and parents to preliminarily assess the status of mental of the students in the college. So, some interventions can be timely adopted to improve the mental health of students.</p>
<p>The limitations of this study lie in: 1) for the college students&#x00027; psychological health education data, we only analyzed the CNKI data source, other source including Web of Science, Scopus, Engineering Index (EI) were not included in the study. So, extensive data analysis based on different data sources is necessary to validate the generalization performance of HSNMF. 2) For student depression status analysis, more complicated interactions relationships between variables may exist. The interplay among these variables may be beyond the interaction between two variables, high-order interaction relationships may be ubiquitous in student behavior analysis. In this study, we used hypergraph to encode the high-order interactions, however, the association analysis was still based on the correlation between two variables.</p>
<p>The future directions of HSNMF framework mainly focus on the following aspects. 1) Identification of depression risk factors from health surveys and biomedicine data (<xref ref-type="bibr" rid="B8">Jamali et al., 2024</xref>). 2) Extending HSNMF to multi-modality student psychological health data, including psychological health survey, biomedicine image data, metagenomics data (<xref ref-type="bibr" rid="B17">Lai et al., 2022</xref>), microbiome data (<xref ref-type="bibr" rid="B27">Rong et al., 2019</xref>), and so on.</p></sec>
</body>
<back>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="https://github.com/chonghua-1983/student_life_analysis">https://github.com/chonghua-1983/student_life_analysis</ext-link>.</p>
</sec>
<sec sec-type="author-contributions" id="s6">
<title>Author contributions</title>
<p>YM: Conceptualization, Methodology, Software, Writing &#x02013; review &#x00026; editing. LL: Data curation, Validation, Writing &#x02013; original draft.</p>
</sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research and/or publication of this article. This work was supported by Hubei Provincial Key Project of Educational Science Planning &#x0201C;Research on the Construction and Measurement of Indicator System, and Improvement Strategies of the Aesthetic Education Integration Campus&#x0201D; [24GA095] and Hubei Superior and Distinctive Discipline Group of &#x0201C;New Energy Vehicle and Smart Transportation.&#x0201D;</p>
</sec>
<ack><p>The authors thank LetPub (<ext-link ext-link-type="uri" xlink:href="https://www.letpub.com.cn">https://www.letpub.com.cn</ext-link>) for its linguistic assistance during the preparation of this manuscript.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="ai-statement" id="s8">
<title>Generative AI statement</title>
<p>The author(s) declare that no Gen AI was used in the creation of this manuscript.</p></sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Blondel</surname> <given-names>V. D.</given-names></name> <name><surname>Guillaume</surname> <given-names>J.-L.</given-names></name> <name><surname>Lambiotte</surname> <given-names>R.</given-names></name> <name><surname>Lefebvre</surname> <given-names>E.</given-names></name></person-group> (<year>2008</year>). <article-title>Fast unfolding of communities in large networks</article-title>. <source>J. Stat. Mech. Theory Exp.</source> <volume>2008</volume>:<fpage>10008</fpage>. <pub-id pub-id-type="doi">10.1088/1742-5468/2008/10/P10008</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boutsidis</surname> <given-names>C.</given-names></name> <name><surname>Gallopoulos</surname> <given-names>E. S. V. D.</given-names></name></person-group> (<year>2008</year>). <article-title>based initialization: a head start for nonnegative matrix factorization</article-title>. <source>Pattern recognition</source>. <volume>41</volume>, <fpage>1350</fpage>&#x02013;<lpage>1362</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2007.09.010</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cai</surname> <given-names>D.</given-names></name> <name><surname>He</surname> <given-names>X.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Huang</surname> <given-names>T. S.</given-names></name></person-group> (<year>2010</year>). <article-title>Graph regularized nonnegative matrix factorization for data representation</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>33</volume>, <fpage>1548</fpage>&#x02013;<lpage>1560</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2010.231</pub-id><pub-id pub-id-type="pmid">21173440</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cali&#x00144;ski</surname> <given-names>T.</given-names></name> <name><surname>Harabasz</surname> <given-names>J.</given-names></name></person-group> (<year>1974</year>). <article-title>A dendrite method for cluster analysis</article-title>. <source>Commun. Stat. Theory Methods</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1080/03610927408827101</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chavoshinejad</surname> <given-names>J.</given-names></name> <name><surname>Seyedi</surname> <given-names>S. A.</given-names></name> <name><surname>Tab</surname> <given-names>F. A.</given-names></name> <name><surname>Salahian</surname> <given-names>N.</given-names></name></person-group> (<year>2023</year>). <article-title>Self-supervised semi-supervised nonnegative matrix factorization for data clustering</article-title>. <source>Pattern Recognit.</source> <volume>137</volume>:<fpage>109282</fpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2022.109282</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Davies</surname> <given-names>D. L.</given-names></name> <name><surname>Bouldin</surname> <given-names>D. W.</given-names></name></person-group> (<year>1979</year>). <article-title>A cluster separation measure</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>1</volume>, <fpage>224</fpage>&#x02013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.1979.4766909</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>R.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Fan</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name></person-group> (<year>2015</year>). <article-title>Mapping knowledge domain of counseling and psychotherapy researches in China</article-title>. <source>Chin. Ment. Health J</source>. <volume>12</volume>, <fpage>510</fpage>&#x02013;<lpage>515</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jamali</surname> <given-names>A. A.</given-names></name> <name><surname>Berger</surname> <given-names>C.</given-names></name> <name><surname>Spiteri</surname> <given-names>R. J.</given-names></name></person-group> (<year>2024</year>). <article-title>Identification of depression predictors from standard health surveys using machine learning</article-title>. <source>Curr. Res. Behav. Sci.</source> <volume>2024</volume>:<fpage>100157</fpage>. <pub-id pub-id-type="doi">10.1016/j.crbeha.2024.100157</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jao</surname> <given-names>N. C.</given-names></name> <name><surname>Robinson</surname> <given-names>L. D.</given-names></name> <name><surname>Kelly</surname> <given-names>P. J.</given-names></name> <name><surname>Ciecierski</surname> <given-names>C. C.</given-names></name> <name><surname>Hitsman</surname> <given-names>B.</given-names></name></person-group> (<year>2019</year>). <article-title>Unhealthy behavior clustering and mental health status in United States college students</article-title>. <source>J. Am. Coll. Health</source> <volume>67</volume>, <fpage>790</fpage>&#x02013;<lpage>800</lpage>. <pub-id pub-id-type="doi">10.1080/07448481.2018.1515744</pub-id><pub-id pub-id-type="pmid">30485154</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>H.</given-names></name> <name><surname>Ding</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>X.</given-names></name> <name><surname>Nie</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>The latest research progress on spectral clustering</article-title>. <source>Neural Comput. Appl.</source> <volume>24</volume>, <fpage>1477</fpage>&#x02013;<lpage>1486</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-013-1439-2</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>X.</given-names></name> <name><surname>Weitz</surname> <given-names>J. S.</given-names></name> <name><surname>Dushoff</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>A non-negative matrix factorization framework for identifying modular patterns in metagenomic profile data</article-title>. <source>J. Math. Biol.</source> <volume>64</volume>, <fpage>697</fpage>&#x02013;<lpage>711</lpage>. <pub-id pub-id-type="doi">10.1007/s00285-011-0428-2</pub-id><pub-id pub-id-type="pmid">21630089</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jos&#x000E9;-Garc&#x000ED;a</surname> <given-names>A.</given-names></name> <name><surname>G&#x000F3;mez-Flores</surname> <given-names>W. C. V. I. K.</given-names></name></person-group> (<year>2023</year>). <article-title>A Matlab-based cluster validity index toolbox for automatic data clustering</article-title>. <source>SoftwareX</source> <volume>22</volume>:<fpage>101359</fpage>. <pub-id pub-id-type="doi">10.1016/j.softx.2023.101359</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kalgotra</surname> <given-names>P.</given-names></name> <name><surname>Sharda</surname> <given-names>R.</given-names></name> <name><surname>Luse</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Which similarity measure to use in network analysis: impact of sample size on phi correlation coefficient and ochiai index</article-title>. <source>Int. J. Inf. Manag.</source> <volume>55</volume>:<fpage>102229</fpage>. <pub-id pub-id-type="doi">10.1016/j.ijinfomgt.2020.102229</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kaufman</surname> <given-names>L.</given-names></name> <name><surname>Rousseeuw</surname> <given-names>P. J.</given-names></name></person-group> (<year>2009</year>). <source>Finding Groups in Data: An Introduction to Cluster Analysis</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x00026; Sons</publisher-name>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kontoangelos</surname> <given-names>K.</given-names></name> <name><surname>Economou</surname> <given-names>M.</given-names></name> <name><surname>Papageorgiou</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>Mental health effects of COVID-19 pandemia: a review of clinical and psychological traits</article-title>. <source>Psychiatry Investig.</source> <volume>17</volume>:<fpage>491</fpage>. <pub-id pub-id-type="doi">10.30773/pi.2020.0161</pub-id><pub-id pub-id-type="pmid">32570296</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kuang</surname> <given-names>D.</given-names></name> <name><surname>Ding</surname> <given-names>C.</given-names></name> <name><surname>Park</surname> <given-names>H.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Symmetric nonnegative matrix factorization for graph clustering,&#x0201D;</article-title> in <source>Proceedings of the 2012 SIAM International Conference on Data Mining (SIAM)</source> (<publisher-loc>Anaheim, CA</publisher-loc>: <publisher-name>Society for Industrial and Applied Mathematics (SIAM</publisher-name>)), <fpage>106</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1137/1.9781611972825.10</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lai</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>A.</given-names></name> <name><surname>Jiang</surname> <given-names>J.</given-names></name> <name><surname>Yuan</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name> <name><surname>Xi</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Metagenomic analysis reveals gut bacterial signatures for diagnosis and treatment outcome prediction in bipolar depression</article-title>. <source>Psychiatry Res.</source> <volume>307</volume>:<fpage>114326</fpage>. <pub-id pub-id-type="doi">10.1016/j.psychres.2021.114326</pub-id><pub-id pub-id-type="pmid">34896845</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Law</surname> <given-names>M. T.</given-names></name> <name><surname>Urtasun</surname> <given-names>R.</given-names></name> <name><surname>Zemel</surname> <given-names>R. S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Deep spectral clustering learning,&#x0201D;</article-title> in <source>International Conference on Machine Learning</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>1985</fpage>&#x02013;<lpage>1994</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>D.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>2000</year>). <article-title>Algorithms for non-negative matrix factorization</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>13</volume>, <fpage>556</fpage>&#x02013;<lpage>562</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>D. D.</given-names></name> <name><surname>Seung</surname> <given-names>H. S.</given-names></name></person-group> (<year>1999</year>). <article-title>Learning the parts of objects by non-negative matrix factorization</article-title>. <source>Nature</source> <volume>401</volume>, <fpage>788</fpage>&#x02013;<lpage>791</lpage>. <pub-id pub-id-type="doi">10.1038/44565</pub-id><pub-id pub-id-type="pmid">10548103</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Cheng</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Bibliometric analysis of researches of orem self-care model in China based on BICOMB</article-title>. <source>TMR Integr. Nurs.</source> <volume>2021</volume>:<fpage>5</fpage>. <pub-id pub-id-type="doi">10.53388/TMRIN20191214</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Xiang</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Hu</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Online mental health services in China during the COVID-19 outbreak</article-title>. <source>Lancet Psychiatry</source> <volume>7</volume>, <fpage>e17</fpage>&#x02013;<lpage>e18</lpage>. <pub-id pub-id-type="doi">10.1016/S2215-0366(20)30077-8</pub-id><pub-id pub-id-type="pmid">32085841</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>Z.</given-names></name></person-group> (<year>2022</year>). <article-title>Analysis model of college students&#x00027; mental health based on online community topic mining and emotion analysis in novel coronavirus epidemic situation</article-title>. <source>Front. Public Health</source> <volume>10</volume>:<fpage>1000313</fpage>. <pub-id pub-id-type="doi">10.3389/fpubh.2022.1000313</pub-id><pub-id pub-id-type="pmid">36187685</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Ma</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>MHSNMF: multi-view hessian regularization based symmetric nonnegative matrix factorization for microbiome data analysis</article-title>. <source>BMC Bioinformatics</source> <volume>21</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-03555-w</pub-id><pub-id pub-id-type="pmid">33203357</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ng</surname> <given-names>A.</given-names></name> <name><surname>Jordan</surname> <given-names>M.</given-names></name> <name><surname>Weiss</surname> <given-names>Y.</given-names></name></person-group> (<year>2001</year>). <article-title>On spectral clustering: analysis and an algorithm</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>14</volume>, <fpage>849</fpage>&#x02013;<lpage>856</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Opoku</surname> <given-names>A. K.</given-names></name> <name><surname>Terhorst</surname> <given-names>Y.</given-names></name> <name><surname>Vega</surname> <given-names>J.</given-names></name> <name><surname>Peltonen</surname> <given-names>E.</given-names></name> <name><surname>Lagerspetz</surname> <given-names>E.</given-names></name> <name><surname>Ferreira</surname> <given-names>D.</given-names></name></person-group> (<year>2021</year>). <article-title>Predicting depression from smartphone behavioral markers using machine learning methods, hyperparameter optimization, and feature importance analysis: exploratory study</article-title>. <source>JMIR Mhealth Uhealth</source> <volume>9</volume>:<fpage>e26540</fpage>. <pub-id pub-id-type="doi">10.2196/26540</pub-id><pub-id pub-id-type="pmid">34255713</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rong</surname> <given-names>H.</given-names></name> <name><surname>Xie</surname> <given-names>X-h.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Lai</surname> <given-names>W-t.</given-names></name> <name><surname>Wang</surname> <given-names>M-b.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Similarly in depression, nuances of gut microbiota: evidences from a shotgun metagenomics sequencing study on major depressive disorder versus bipolar disorder with current major depressive episode patients</article-title>. <source>J. Psychiatr. Res.</source> <volume>113</volume>, <fpage>90</fpage>&#x02013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1016/j.jpsychires.2019.03.017</pub-id><pub-id pub-id-type="pmid">30927646</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vancraeynest</surname> <given-names>B.</given-names></name> <name><surname>Pham</surname> <given-names>H-S.</given-names></name> <name><surname>Ali-Eldin</surname> <given-names>A.</given-names></name></person-group> (<year>2024</year>). <article-title>A new approach to computing the distances between research disciplines based on researcher collaborations and similarity measurement techniques</article-title>. <source>J. Informetr.</source> <volume>18</volume>:<fpage>101527</fpage>. <pub-id pub-id-type="doi">10.1016/j.joi.2024.101527</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>R.</given-names></name> <name><surname>Chen</surname> <given-names>F.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>T.</given-names></name> <name><surname>Harari</surname> <given-names>G.</given-names></name> <name><surname>Tignor</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>&#x0201C;StudentLife: assessing mental health, academic performance and behavioral trends of college students using smartphones,&#x0201D;</article-title> in <source>Proceedings of the 2014 ACM International Joint Conference on Pervasive and Ubiquitous Computing</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>3</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1145/2632048.2632054</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>White</surname> <given-names>S.</given-names></name> <name><surname>Smyth</surname> <given-names>P.</given-names></name></person-group> (<year>2005</year>). <article-title>&#x0201C;A spectral clustering approach to finding communities in graphs,&#x0201D;</article-title> in <source>Proceedings of the 2005 SIAM International Conference on Data Mining</source> (<publisher-loc>Newport Beach, CA</publisher-loc>: <publisher-name>SIAM</publisher-name>), <fpage>274</fpage>&#x02013;<lpage>285</lpage>. <pub-id pub-id-type="doi">10.1137/1.9781611972757.25</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yi</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Qiao</surname> <given-names>W.</given-names></name> <name><surname>Xie</surname> <given-names>X.</given-names></name> <name><surname>Yi</surname> <given-names>N.</given-names></name> <name><surname>Dong</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Clustering effects of health risk behavior on mental health and physical activity in Chinese adolescents</article-title>. <source>Health Qual. Life Outcomes</source> <volume>18</volume>, <fpage>1</fpage>&#x02013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1186/s12955-020-01468-z</pub-id><pub-id pub-id-type="pmid">32620107</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Sch&#x000F6;lkopf</surname> <given-names>B.</given-names></name></person-group> (<year>2006</year>). <article-title>Learning with hypergraphs: clustering, classification, and embedding</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>19</volume>, <fpage>1601</fpage>&#x02013;<lpage>1608</lpage>. <pub-id pub-id-type="doi">10.7551/mitpress/7503.003.0205</pub-id></citation>
</ref>
</ref-list>
</back>
</article>