<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="discussion">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2023.1243092</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Opinion</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Clustering and disease subtyping in Neuroscience, toward better methodological adaptations</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Poulakis</surname> <given-names>Konstantinos</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2347254/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Westman</surname> <given-names>Eric</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/71759/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Division of Clinical Geriatrics, Department of Neurobiology, Care Sciences and Society, Karolinska Institutet</institution>, <addr-line>Stockholm</addr-line>, <country>Sweden</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Neuroimaging, Centre for Neuroimaging Sciences, Institute of Psychiatry, Psychology and Neuroscience, King&#x00027;s College London</institution>, <addr-line>London</addr-line>, <country>United Kingdom</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Noemi Montobbio, University of Genoa, Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Michael Thrun, University of Marburg, Germany; Eduardo Castro, University of New Mexico, United States; Andrea Chincarini, National Institute of Nuclear Physics of Genoa, Italy</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Konstantinos Poulakis <email>konstantinos.poulakis&#x00040;ki.se</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>10</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1243092</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Poulakis and Westman.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Poulakis and Westman</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license></permissions> 
<kwd-group>
<kwd>clustering</kwd>
<kwd>unsupervised learning</kwd>
<kwd>Neuroscience</kwd>
<kwd>disease subtypes</kwd>
<kwd>Alzheimer&#x00027;s disease</kwd>
</kwd-group>
<counts>
<fig-count count="0"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="36"/>
<page-count count="4"/>
<word-count count="3233"/>
</counts>
</article-meta>
</front>
<body>
<p>The increasing interest in identifying disease biomarkers to understand psychiatric and neurological conditions has led to large patient registries and cohorts. Traditionally, clinically defined labels (e.g., disease vs. control group) were associated statistically with potential biomarkers to draw useful information about brain function related to a disease (supervised analysis) (Deo, <xref ref-type="bibr" rid="B5">2015</xref>). However, the observed biomarker variability and the presence of clinical disease subtypes have sparked interest in quantitatively exploring heterogeneity (Feczko et al., <xref ref-type="bibr" rid="B9">2019</xref>; Ferreira et al., <xref ref-type="bibr" rid="B10">2020</xref>). The unsupervised<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> exploration of a disease population (without any clinical labels) through a selected sample is a demanding task that differs from supervised analysis by definition (Habes et al., <xref ref-type="bibr" rid="B11">2020</xref>). However, in research the differences between the two are often overlooked. Therefore, we want to highlight the applications and challenges of clustering, where supervised analysis principles are sometimes misapplied. We also demonstrate how such practices can negatively impact clustering results.</p>
<p>Some common challenges in clustering methods include selecting relevant features to describe data heterogeneity, preprocessing to remove biases, choosing appropriate similarity measures to summarize critical information, selecting a suitable method for meaningful clustering, tuning clustering model parameters (such as cluster size) without ground truth, and validating clustering results (Halkidi et al., <xref ref-type="bibr" rid="B12">2001</xref>; Hennig et al., <xref ref-type="bibr" rid="B14">2015</xref>).</p>
<p>The most common clustering applications in medicine (Halkidi et al., <xref ref-type="bibr" rid="B12">2001</xref>):</p>
<list list-type="bullet">
<list-item><p>Data reduction (Hennig et al., <xref ref-type="bibr" rid="B14">2015</xref>). When dealing with large datasets, like genomics, proteomics, or medical imaging data, clustering can condense the information into representative vectors or filter out uninformative features.</p></list-item>
<list-item><p>Generate new hypotheses. Discovering specific disease subtypes can lead to the development of new hypotheses, altering existing theories.</p></list-item>
<list-item><p>Hypothesis testing (Thrun and Ultsch, <xref ref-type="bibr" rid="B29">2021</xref>). Clustering can be used for hypothesis testing. For example, it can assess whether clinical observations align with biological data in diseases with known subtypes without forcing the association between biological data and clinical labels (supervised approach).</p></list-item>
<list-item><p>Prediction in new patients (Wu et al., <xref ref-type="bibr" rid="B33">2019</xref>). Clustering can identify disease subtypes and scientific theories that investigators can use to create supervised classification models for grouping new patients. This new classification is valuable for personalized medicine and future patient treatment, among other applications.</p></list-item>
</list>
<p>When working with unsupervised methods, it&#x00027;s crucial to understand their limitations and nuances. Clustering encompasses a wide range of techniques which handle population structures and characteristics differently. Understanding the idiosyncrasies of a dataset is essential for applying clustering successfully. Questions about how clustering results generalize to the disease population, which are the optimal model parameters, and why results change with slight dataset modifications often emerge during study design, model optimization, interpretation, and peer review. One intriguing approach that combines automatic machine learning with expert knowledge from the field is the &#x00027;human-in-the-loop&#x00027; method (Holzinger, <xref ref-type="bibr" rid="B15">2016</xref>). This approach is particularly effective in neurological applications and can help address the abovementioned questions.</p>
<p>Regarding cluster size and type, we may know in advance whether there is excess variation in a disease population, some heterogeneous disease features, and even subtype proportions. This knowledge is vital in the model selection process so that we can sort out methods that are wrong methodological fits for the population of interest. For example, k-means, one of the most popular clustering methods, tends to produce convex-shaped clusters (it tends to equalize the spatial variance) that are spherical and often become similar in size (Celebi et al., <xref ref-type="bibr" rid="B4">2013</xref>). Therefore, if in a specific disease population, we are aware of rare disease subtypes that may also exist in our sample, we may want to avoid k-means. Instead, we should focus on clustering methods to identify outliers/outlier clusters (Campello et al., <xref ref-type="bibr" rid="B3">2015</xref>). Further, the more variables we use in a clustering method, the more the dimensionality of the dataset increases. A good practice is to use methods that either pretreat data to reduce the dimensionality and then apply regular clustering to them or select a method that can cope with high dimensional datasets (Babu et al., <xref ref-type="bibr" rid="B2">2011</xref>; Thrun, <xref ref-type="bibr" rid="B28">2021</xref>). While the gold standard in machine learning, some studies fail to utilize suitable models for high-dimensional data (Noh et al., <xref ref-type="bibr" rid="B21">2014</xref>; Hwang et al., <xref ref-type="bibr" rid="B16">2016</xref>; Jeon et al., <xref ref-type="bibr" rid="B17">2019</xref>; Levin et al., <xref ref-type="bibr" rid="B18">2021</xref>), limiting our ability to assess the success of clustering.</p>
<p>Further, all clustering methods cannot cope with all types of data (ordinal/nominal categorical, numerical) (Halkidi et al., <xref ref-type="bibr" rid="B12">2001</xref>). When we binarize continuous variables to utilize a clustering algorithm for binary data only, the reduction of information due to data transformation must be at least considered when interpreting the results (Zhang et al., <xref ref-type="bibr" rid="B36">2016</xref>). Some algorithms use mixed data types and should be preferred when mixed data distributions are present (Szepannek, <xref ref-type="bibr" rid="B26">2019</xref>). If not accounted for, data biases may render a clustering result misleading. For example, we may be interested in understanding the heterogeneity of a particular biological process during aging. Understanding and adjusting the data to consider the participants&#x00027; age variability results in clusters of participants that are not driven by age differences but by differences in the biological process under investigation if those exist (given that other biases are not present). However, due to complex data/aging relationships, these effects may persist even after statistical accounting for aging. Other sampling features that can drive clustering results are sex, disease stage, comorbidities, medication exposure, and geographical position. For example, it is known that the disease stage may contribute to the observed heterogeneity in Alzheimer&#x00027;s disease (AD) (Ferreira et al., <xref ref-type="bibr" rid="B10">2020</xref>), we have only recently started accounting for this or trying to assess its contribution (Young et al., <xref ref-type="bibr" rid="B35">2017</xref>; Vogel et al., <xref ref-type="bibr" rid="B32">2021</xref>; Yang et al., <xref ref-type="bibr" rid="B34">2021</xref>; Poulakis et al., <xref ref-type="bibr" rid="B24">2022</xref>) while in previous studies (Noh et al., <xref ref-type="bibr" rid="B21">2014</xref>; Dong et al., <xref ref-type="bibr" rid="B7">2016</xref>; Hwang et al., <xref ref-type="bibr" rid="B16">2016</xref>; Zhang et al., <xref ref-type="bibr" rid="B36">2016</xref>; Park et al., <xref ref-type="bibr" rid="B22">2017</xref>; Poulakis et al., <xref ref-type="bibr" rid="B23">2018</xref>; ten Kate et al., <xref ref-type="bibr" rid="B27">2018</xref>) we did not assess or account for this effect.</p>
<p>Clustering results must generalize well to the population, which makes validation a central topic. Traditionally, cross-validation (CV), bootstrapping, external data testing (training, validating, and testing), and careful sample selection have been some of the most popular approaches in supervised analysis. However, validation in clustering is not straightforward since no ground truth exists. The adaptation of training and testing a clustering model using independent datasets can sometimes mislead us. For example, three subtypes are present in a hypothetical disease population N (s1, s2, and s3). One is the most prevalent (s1) (typical presentation), the second subtype (s2) has half of the prevalence of the first one (n<sub>s2</sub> = <inline-formula><mml:math id="M1"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>n<sub>s1</sub>), and the third subtype has a low prevalence (one-tenth of the first subtype, n<sub>s3</sub> = <inline-formula><mml:math id="M2"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>n<sub>s1</sub>) (s3). The disease population N equals n<sub>1</sub> &#x0002B; n<sub>2</sub> &#x0002B; n<sub>3</sub>. A perfectly representative random sample of 100 patients from the disease population will include approximately 63 patients from s1, 31 from s2, and six from s3. A clustering model can then be trained on 70% (70 patients) and tested using 30% (30 patients). Suppose the data in the training set perfectly represent the population, a rare phenomenon, and clustering accurately identifies the subtypes. In that case, 44 patients will end up in Cluster 1, 22 in Cluster 2, and 4 in Cluster 3. The test set should have 19 patients in s1, 9 in s2, and 2 in s3. Clustering can then be applied to identify subtypes s1, s2, and s3. Since the actual data labels are unknown, which is what clustering should discover, the test set results will be compared to the training set. The problem arises with rare subtypes, such as the hypothetical s3 subtype (six patients in the sample, four in the training set, and two in the test set). Patients of such subtypes may end up in larger clusters when the overall dataset is split into small segments for the needs of the analysis. Unfortunately, the most interesting heterogeneous characteristics will enrich another cluster&#x00027;s greater information pool, especially in high-dimensional datasets. In the best-case scenario, those patients will be single outliers (if the algorithm can recognize outlier clusters) (Campello et al., <xref ref-type="bibr" rid="B3">2015</xref>). Understanding their features is pivotal for the assessment of heterogeneity in the disease.</p>
<p>To the best of our knowledge, cross-validation has been successfully combined with clustering in two studies to assess the consistency of observations within the same cluster and to determine the optimal model solution (Varol et al., <xref ref-type="bibr" rid="B31">2017</xref>; Yang et al., <xref ref-type="bibr" rid="B34">2021</xref>). On the other hand, leave 10% of patients out-CV (a semi-supervised application where a control group is contrasted to a disease group) to decide the optimal clustering (Dong et al., <xref ref-type="bibr" rid="B7">2016</xref>, <xref ref-type="bibr" rid="B8">2017</xref>), may reveal the dominant patterns in the dataset. An interesting question is whether clusters of low/very low prevalence can survive this process. In AD, genetic mutations account for &#x0003C;1% of all AD (<xref ref-type="bibr" rid="B1">2020</xref>) cases, while early-onset AD accounts for 4%&#x02212;6% (Mendez, <xref ref-type="bibr" rid="B20">2017</xref>). Another evaluation approach is to compare clustering agreement after application of the same algorithm in different cohorts. We do not suggest that these results are wrong, but they may be misleading if different clustering findings in different cohorts are interpreted as a methodological failure, while convergence of findings between cohorts is the aim (ten Kate et al., <xref ref-type="bibr" rid="B27">2018</xref>; Vogel et al., <xref ref-type="bibr" rid="B32">2021</xref>). Sometimes, it is a requirement that clustering should be repeated cohort-wise to prove model robustness (Poulakis et al., <xref ref-type="bibr" rid="B23">2018</xref>, <xref ref-type="bibr" rid="B24">2022</xref>). Instead of reducing data variability in clustering by splitting the available data into segments, we should acknowledge that cluster-cohort agreement-based evaluation criteria can potentially interrupt the discovery of rare data patterns. Another issue with the cohort-wise analysis is the potential sample imbalance between cohorts that may render one cohort solution less reliable than another. Of note, cohort-wise analysis is reasonable when cohorts have different feature sets or systematic differences (Marinescu et al., <xref ref-type="bibr" rid="B19">2019</xref>; Tijms et al., <xref ref-type="bibr" rid="B30">2020</xref>). Prior knowledge (subtype prevalence or number of subtypes) is essential when formulating a clustering experimental design (Halkidi et al., <xref ref-type="bibr" rid="B12">2001</xref>, <xref ref-type="bibr" rid="B13">2002</xref>). Another example, hypothetically, two separate clusters of patients may be formed because a clustering validation criterion gives marginally better scores instead of grouping the patients in one cluster. Field experts and not only clustering internal evaluation criteria should conclude whether differences between clusters are essential enough to suggest heterogeneity (Halkidi et al., <xref ref-type="bibr" rid="B13">2002</xref>; Dolnicar and Leisch, <xref ref-type="bibr" rid="B6">2010</xref>). It is also often observed that clustering algorithms optimally select two-cluster solutions. This finding may not provide any insight of the disease process when it only reveals biomarker severity differences of no clinical interest (Poulakis et al., <xref ref-type="bibr" rid="B25">2021</xref>; Yang et al., <xref ref-type="bibr" rid="B34">2021</xref>). Based on the above, we believe that as large datasets as possible should be used when training a clustering model. In contrast, datasets should not be divided for validation purposes if the focus is on revealing heterogeneity in a population.</p>
<p>Clustering is a valuable approach to understand heterogeneity in brain disorders and healthy aging. The machine learning community has invested a great deal of research in addressing the methodological issues discussed above. As with every statistical tool, these methods should be carefully applied, and understanding their properties and limitations is essential.</p>
<sec sec-type="author-contributions" id="s1">
<title>Author contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s2">
<title>Funding</title>
<p>We would like to thank the Swedish Foundation for Strategic Research (SSF), the Swedish Research Council (VR), the Center for Innovative Medicine (CIMED), the Strategic Research Programme in Neuroscience at Karolinska Institutet (StratNeuro), Swedish Brain Power, the regional agreement on medical training and clinical research (ALF) between Stockholm County Council and Karolinska Institutet, Hj&#x000E4;rnfonden, Alzheimerfonden, the &#x000C5;ke Wiberg Foundation, the King Gustaf V:s and Queen Victorias Foundation, and Birgitta och Sten Westerberg for additional financial support.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s3">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>For the needs of this text, unsupervised analysis refers to clustering only, association analysis is not covered.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><collab>AD</collab></person-group> (<year>2020</year>). <article-title>2020 Alzheimer&#x00027;s disease facts and figures</article-title>. <source>Alzheimers Dement.</source> <volume>16</volume>, <fpage>391</fpage>&#x02013;<lpage>460</lpage>. <pub-id pub-id-type="doi">10.1002/alz.12068</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Babu</surname> <given-names>B.</given-names></name> <name><surname>Subash</surname> <given-names>C. N.</given-names></name> <name><surname>Gopal</surname> <given-names>T. V.</given-names></name></person-group> (<year>2011</year>). <article-title>Clustering algorithms for high dimensional data &#x02013; a survey of issues and existing approaches</article-title>. <source>Spec. Issue Int. J. Comput. Sci. Inform.</source> <volume>2</volume>, <fpage>13</fpage>. <pub-id pub-id-type="doi">10.47893/IJCSI.2013.1108</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Campello</surname> <given-names>R. J. G. B.</given-names></name> <name><surname>Moulavi</surname> <given-names>D.</given-names></name> <name><surname>Zimek</surname> <given-names>A.</given-names></name> <name><surname>Sander</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Hierarchical density estimates for data clustering, visualization, and outlier detection</article-title>. <source>ACM Trans. Knowl. Discov. Data</source> <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1145/2733381</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Celebi</surname> <given-names>M. E.</given-names></name> <name><surname>Kingravi</surname> <given-names>H. A.</given-names></name> <name><surname>Vela</surname> <given-names>P. A. A.</given-names></name></person-group> (<year>2013</year>). <article-title>comparative study of efficient initialization methods for the k-means clustering algorithm</article-title>. <source>Expert Syst. Appl.</source> <volume>40</volume>, <fpage>200</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2012.07.021</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deo</surname> <given-names>R. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Machine learning in medicine</article-title>. <source>Circulation</source> <volume>132</volume>, <fpage>1920</fpage>&#x02013;<lpage>1930</lpage>. <pub-id pub-id-type="doi">10.1161/CIRCULATIONAHA.115.001593</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dolnicar</surname> <given-names>S.</given-names></name> <name><surname>Leisch</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>Evaluation of structure and reproducibility of cluster solutions using the bootstrap</article-title>. <source>Mark. Lett.</source> <volume>21</volume>, <fpage>83</fpage>&#x02013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.1007/s11002-009-9083-4</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>A.</given-names></name> <name><surname>Honnorat</surname> <given-names>N.</given-names></name> <name><surname>Gaonkar</surname> <given-names>B.</given-names></name> <name><surname>Davatzikos</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>CHIMERA: clustering of heterogeneous disease effects via distribution matching of imaging patterns</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>35</volume>, <fpage>612</fpage>&#x02013;<lpage>621</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2015.2487423</pub-id><pub-id pub-id-type="pmid">26452275</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>A.</given-names></name> <name><surname>Toledo</surname> <given-names>J. B.</given-names></name> <name><surname>Honnorat</surname> <given-names>N.</given-names></name> <name><surname>Doshi</surname> <given-names>J.</given-names></name> <name><surname>Varol</surname> <given-names>E.</given-names></name> <name><surname>Sotiras</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Heterogeneity of neuroanatomical patterns in prodromal Alzheimer&#x00027;s disease: links to cognition, progression and biomarkers</article-title>. <source>Brain</source> <volume>140</volume>, <fpage>735</fpage>&#x02013;<lpage>747</lpage>. <pub-id pub-id-type="doi">10.1093/brain/aww319</pub-id><pub-id pub-id-type="pmid">28003242</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feczko</surname> <given-names>E.</given-names></name> <name><surname>Miranda-Dominguez</surname> <given-names>O.</given-names></name> <name><surname>Marr</surname> <given-names>M.</given-names></name> <name><surname>Graham</surname> <given-names>A. M.</given-names></name> <name><surname>Nigg</surname> <given-names>J. T.</given-names></name> <name><surname>Fair</surname> <given-names>D. A.</given-names></name></person-group> (<year>2019</year>). <article-title>The heterogeneity problem: approaches to identify psychiatric subtypes</article-title>. <source>Trends Cogn. Sci.</source> <volume>23</volume>, <fpage>584</fpage>&#x02013;<lpage>601</lpage>. <pub-id pub-id-type="doi">10.1016/j.tics.2019.03.009</pub-id><pub-id pub-id-type="pmid">31153774</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferreira</surname> <given-names>D.</given-names></name> <name><surname>Nordberg</surname> <given-names>A.</given-names></name> <name><surname>Westman</surname> <given-names>E.</given-names></name></person-group> (<year>2020</year>). <article-title>Biological subtypes of Alzheimer disease</article-title>. <source>Neurology</source> <volume>94</volume>, <fpage>436</fpage>&#x02013;<lpage>448</lpage>. <pub-id pub-id-type="doi">10.1212/WNL.0000000000009058</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Habes</surname> <given-names>M.</given-names></name> <name><surname>Grothe</surname> <given-names>M. J.</given-names></name> <name><surname>Tunc</surname> <given-names>B.</given-names></name> <name><surname>McMillan</surname> <given-names>C.</given-names></name> <name><surname>Wolk</surname> <given-names>D. A.</given-names></name> <name><surname>Davatzikos</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>Disentangling heterogeneity in Alzheimer&#x00027;s disease and related dementias using data-driven methods</article-title>. <source>Biol. Psychiatry</source> <volume>88</volume>, <fpage>70</fpage>&#x02013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1016/j.biopsych.2020.01.016</pub-id><pub-id pub-id-type="pmid">32201044</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Halkidi</surname> <given-names>M.</given-names></name> <name><surname>Batistakis</surname> <given-names>Y.</given-names></name> <name><surname>Vazirgiannis</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>On clustering validation techniques</article-title>. <source>J. Intell. Inf. Syst.</source> <volume>17</volume>, <fpage>107</fpage>&#x02013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1023/A:1012801612483</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Halkidi</surname> <given-names>M.</given-names></name> <name><surname>Batistakis</surname> <given-names>Y.</given-names></name> <name><surname>Vazirgiannis</surname> <given-names>M.</given-names></name></person-group> (<year>2002</year>). <article-title>Clustering validity checking methods</article-title>. <source>ACM SIGMOD Rec.</source> <volume>31</volume>, <fpage>19</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1145/601858.601862</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hennig</surname> <given-names>C.</given-names></name> <name><surname>Meila</surname> <given-names>M.</given-names></name> <name><surname>Murtagh</surname> <given-names>F.</given-names></name> <name><surname>Rocci</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <source>Handbook of Cluster Analysis</source>. <publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>Chapman and Hall/CRC</publisher-name>. <pub-id pub-id-type="doi">10.1201/b19706</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holzinger</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Interactive machine learning for health informatics: when do we need the human-in-the-loop?</article-title> <source>Brain Inform.</source> <volume>3</volume>, <fpage>119</fpage>&#x02013;<lpage>131</lpage>. <pub-id pub-id-type="doi">10.1007/s40708-016-0042-6</pub-id><pub-id pub-id-type="pmid">27747607</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hwang</surname> <given-names>J.</given-names></name> <name><surname>Kim</surname> <given-names>C. M.</given-names></name> <name><surname>Jeon</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>J. M.</given-names></name> <name><surname>Hong</surname> <given-names>Y. J.</given-names></name> <name><surname>Roh</surname> <given-names>J. H.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Prediction of Alzheimer&#x00027;s disease pathophysiology based on cortical thickness patterns</article-title>. <source>Alzheimer&#x00027;s Dement. Diagnosis, Assess. Dis. Monit.</source> <volume>2</volume>, <fpage>58</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1016/j.dadm.2015.11.008</pub-id><pub-id pub-id-type="pmid">27239533</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jeon</surname> <given-names>S.</given-names></name> <name><surname>Kang</surname> <given-names>J. M.</given-names></name> <name><surname>Seo</surname> <given-names>S.</given-names></name> <name><surname>Jeong</surname> <given-names>H. J.</given-names></name> <name><surname>Funck</surname> <given-names>T.</given-names></name> <name><surname>Lee</surname> <given-names>S.-Y.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Topographical heterogeneity of Alzheimer&#x00027;s disease based on MR imaging, tau PET, and amyloid PET</article-title>. <source>Front. Aging Neurosci.</source> <volume>11</volume>, <fpage>211</fpage>. <pub-id pub-id-type="doi">10.3389/fnagi.2019.00211</pub-id><pub-id pub-id-type="pmid">31481888</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Levin</surname> <given-names>F.</given-names></name> <name><surname>Ferreira</surname> <given-names>D.</given-names></name> <name><surname>Lange</surname> <given-names>C.</given-names></name> <name><surname>Dyrba</surname> <given-names>M.</given-names></name> <name><surname>Westman</surname> <given-names>E.</given-names></name> <name><surname>Buchert</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Data-driven FDG-PET subtypes of Alzheimer&#x00027;s disease-related neurodegeneration</article-title>. <source>Alzheimers Res Ther</source>. <volume>13</volume>, <fpage>49</fpage>. <pub-id pub-id-type="doi">10.1186/s13195-021-00785-9</pub-id><pub-id pub-id-type="pmid">33608059</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marinescu</surname> <given-names>R. V.</given-names></name> <name><surname>Eshaghi</surname> <given-names>A.</given-names></name> <name><surname>Lorenzi</surname> <given-names>M.</given-names></name> <name><surname>Young</surname> <given-names>A. L.</given-names></name> <name><surname>Oxtoby</surname> <given-names>N. P.</given-names></name> <name><surname>Garbarino</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>DIVE: a spatiotemporal progression model of brain pathology in neurodegenerative disorders</article-title>. <source>Neuroimage</source> <volume>192</volume>, <fpage>166</fpage>&#x02013;<lpage>177</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2019.02.053</pub-id><pub-id pub-id-type="pmid">30844504</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mendez</surname> <given-names>M. F.</given-names></name></person-group> (<year>2017</year>). <article-title>Early-onset Alzheimer disease</article-title>. <source>Neurol. Clin.</source> <volume>35</volume>, <fpage>263</fpage>&#x02013;<lpage>281</lpage>. <pub-id pub-id-type="doi">10.1016/j.ncl.2017.01.005</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Noh</surname> <given-names>Y.</given-names></name> <name><surname>Jeon</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>J. M.</given-names></name> <name><surname>Seo</surname> <given-names>S. W.</given-names></name> <name><surname>Kim</surname> <given-names>G. H.</given-names></name> <name><surname>Cho</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Anatomical heterogeneity of Alzheimer disease based on cortical thickness on MRIs</article-title>. <source>Neurology</source> <volume>83</volume>, <fpage>1936</fpage>&#x02013;<lpage>1944</lpage>. <pub-id pub-id-type="doi">10.1212/WNL.0000000000001003</pub-id><pub-id pub-id-type="pmid">25344382</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>J.-Y.</given-names></name> <name><surname>Na</surname> <given-names>H. K.</given-names></name> <name><surname>Kim</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>H. J.</given-names></name> <name><surname>Seo</surname> <given-names>S. W.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Robust Identification of Alzheimer&#x00027;s disease subtypes based on cortical atrophy patterns</article-title>. <source>Sci. Rep.</source> <volume>7</volume>, <fpage>43270</fpage>. <pub-id pub-id-type="doi">10.1038/srep43270</pub-id><pub-id pub-id-type="pmid">28276464</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poulakis</surname> <given-names>K.</given-names></name> <name><surname>Pereira</surname> <given-names>J. B.</given-names></name> <name><surname>Mecocci</surname> <given-names>P.</given-names></name> <name><surname>Vellas</surname> <given-names>B.</given-names></name> <name><surname>Tsolak</surname> <given-names>M.</given-names></name> <name><surname>K&#x00142;oszewska</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Heterogeneous patterns of brain atrophy in Alzheimer&#x00027;s disease</article-title>. <source>Neurobiol. Aging</source> <volume>65</volume>, <fpage>98</fpage>&#x02013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1016/j.neurobiolaging.2018.01.009</pub-id><pub-id pub-id-type="pmid">29455029</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poulakis</surname> <given-names>K.</given-names></name> <name><surname>Pereira</surname> <given-names>J. B.</given-names></name> <name><surname>Muehlboeck</surname> <given-names>J.-S.</given-names></name> <name><surname>Wahlund</surname> <given-names>L.-O.</given-names></name> <name><surname>Smedby</surname> <given-names>&#x000D6;.</given-names></name> <name><surname>Volpe</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Multi-cohort and longitudinal Bayesian clustering study of stage and subtype in Alzheimer&#x00027;s disease</article-title>. <source>Nat. Commun.</source> <volume>13</volume>, <fpage>4566</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-022-32202-6</pub-id><pub-id pub-id-type="pmid">35931678</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poulakis</surname> <given-names>K.</given-names></name> <name><surname>Reid</surname> <given-names>R. I.</given-names></name> <name><surname>Przybelski</surname> <given-names>S. A.</given-names></name> <name><surname>Knopman</surname> <given-names>D. S.</given-names></name> <name><surname>Graff-Radford</surname> <given-names>J.</given-names></name> <name><surname>Lowe</surname> <given-names>V. J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Longitudinal deterioration of white-matter integrity: heterogeneity in the ageing population</article-title>. <source>Brain Commun.</source> <volume>3</volume>, <fpage>fcaa238</fpage>. <pub-id pub-id-type="doi">10.1093/braincomms/fcaa238</pub-id><pub-id pub-id-type="pmid">33615218</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Szepannek</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>clustMixType: user-friendly clustering of mixed-type data in R</article-title>. <source>R J.</source> <volume>10</volume>, <fpage>200</fpage>. <pub-id pub-id-type="doi">10.32614/RJ-2018-048</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>ten Kate</surname> <given-names>M.</given-names></name> <name><surname>Dicks</surname> <given-names>E.</given-names></name> <name><surname>Visser</surname> <given-names>P. J.</given-names></name> <name><surname>van der Flier</surname> <given-names>W. M.</given-names></name> <name><surname>Teunissen</surname> <given-names>C. E.</given-names></name> <name><surname>Barkhof</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Atrophy subtypes in prodromal Alzheimer&#x00027;s disease are associated with cognitive decline</article-title>. <source>Brain</source> <volume>141</volume>, <fpage>3443</fpage>&#x02013;<lpage>3456</lpage>. <pub-id pub-id-type="doi">10.1093/brain/awy264</pub-id><pub-id pub-id-type="pmid">30351346</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thrun</surname> <given-names>M. C.</given-names></name></person-group> (<year>2021</year>). <article-title>Distance-based clustering challenges for unbiased benchmarking studies</article-title>. <source>Sci. Rep.</source> <volume>11</volume>, <fpage>18988</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-021-98126-1</pub-id><pub-id pub-id-type="pmid">34556686</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thrun</surname> <given-names>M. C.</given-names></name> <name><surname>Ultsch</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Swarm intelligence for self-organized clustering</article-title>. <source>Artif. Intell.</source> <volume>290</volume>, <fpage>103237</fpage>. <pub-id pub-id-type="doi">10.1016/j.artint.2020.103237</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tijms</surname> <given-names>B. M.</given-names></name> <name><surname>Gobom</surname> <given-names>J.</given-names></name> <name><surname>Reus</surname> <given-names>L.</given-names></name> <name><surname>Jansen</surname> <given-names>I.</given-names></name> <name><surname>Hong</surname> <given-names>S.</given-names></name> <name><surname>Dobricic</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Pathophysiological subtypes of Alzheimer&#x00027;s disease based on cerebrospinal fluid proteomics</article-title>. <source>Brain</source> <volume>143</volume>, <fpage>3776</fpage>&#x02013;<lpage>3792</lpage>. <pub-id pub-id-type="doi">10.1093/brain/awaa325</pub-id><pub-id pub-id-type="pmid">33439986</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Varol</surname> <given-names>E.</given-names></name> <name><surname>Sotiras</surname> <given-names>A.</given-names></name> <name><surname>Davatzikos</surname> <given-names>C.</given-names></name></person-group> (<year>2017</year>). <article-title>HYDRA: revealing heterogeneity of imaging and genetic patterns through a multiple max-margin discriminative analysis framework</article-title>. <source>Neuroimage</source> <volume>145</volume>, <fpage>346</fpage>&#x02013;<lpage>364</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2016.02.041</pub-id><pub-id pub-id-type="pmid">26923371</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vogel</surname> <given-names>J. W.</given-names></name> <name><surname>Young</surname> <given-names>A. L.</given-names></name> <name><surname>Oxtoby</surname> <given-names>N. P.</given-names></name> <name><surname>Smith</surname> <given-names>R.</given-names></name> <name><surname>Ossenkoppele</surname> <given-names>R.</given-names></name> <name><surname>Strandberg</surname> <given-names>O. T.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Four distinct trajectories of tau deposition identified in Alzheimer&#x00027;s disease</article-title>. <source>Nat. Med.</source> <volume>27</volume>, <fpage>871</fpage>&#x02013;<lpage>881</lpage>. <pub-id pub-id-type="doi">10.1038/s41591-021-01309-6</pub-id><pub-id pub-id-type="pmid">33927414</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Bang</surname> <given-names>S.</given-names></name> <name><surname>Bleecker</surname> <given-names>E. R.</given-names></name> <name><surname>Castro</surname> <given-names>M.</given-names></name> <name><surname>Denlinger</surname> <given-names>L.</given-names></name> <name><surname>Erzurum</surname> <given-names>S. C.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Multiview cluster analysis identifies variable corticosteroid response phenotypes in severe asthma</article-title>. <source>Am. J. Respir. Crit. Care Med.</source> <volume>199</volume>, <fpage>1358</fpage>&#x02013;<lpage>1367</lpage>. <pub-id pub-id-type="doi">10.1164/rccm.201808-1543OC</pub-id><pub-id pub-id-type="pmid">30682261</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Nasrallah</surname> <given-names>I.</given-names></name> <name><surname>Shou</surname> <given-names>H.</given-names></name> <name><surname>Wen</surname> <given-names>J.</given-names></name> <name><surname>Doshi</surname> <given-names>J.</given-names></name> <name><surname>Habes</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Disentangling brain heterogeneity via semi-supervised deep-learning and MRI: dimensional representations of Alzheimer&#x00027;s disease</article-title>. <source>Alzheimers Dement.</source> 17<italic>:</italic> <pub-id pub-id-type="doi">10.1002/alz.052735</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Young</surname> <given-names>A. L.</given-names></name> <name><surname>Marinescu</surname> <given-names>R. V.</given-names></name> <name><surname>Oxtoby</surname> <given-names>N. P.</given-names></name> <name><surname>Bocchetta</surname> <given-names>M.</given-names></name> <name><surname>Yong</surname> <given-names>K.</given-names></name> <name><surname>Firth</surname> <given-names>N. C.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Uncovering the heterogeneity and temporal complexity of neurodegenerative diseases with subtype and stage inference</article-title>. <source>bioRxiv</source> [preprint]. <pub-id pub-id-type="doi">10.1101/236604</pub-id><pub-id pub-id-type="pmid">30323170</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Mormino</surname> <given-names>E. C.</given-names></name> <name><surname>Sun</surname> <given-names>N.</given-names></name> <name><surname>Sperling</surname> <given-names>R. A.</given-names></name> <name><surname>Sabuncu</surname> <given-names>M. R.</given-names></name> <name><surname>Yeo</surname> <given-names>B. T. T.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Bayesian model reveals latent atrophy factors with dissociable cognitive trajectories in Alzheimer&#x00027;s disease</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>113</volume>, <fpage>E6535</fpage>&#x02013;<lpage>E6544</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1611073113</pub-id><pub-id pub-id-type="pmid">27702899</pub-id></citation></ref>
</ref-list>
</back>
</article> 