<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Psychiatry</journal-id>
<journal-title>Frontiers in Psychiatry</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Psychiatry</abbrev-journal-title>
<issn pub-type="epub">1664-0640</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpsyt.2021.637022</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Psychiatry</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Identifying Subgroups of Patients With Autism by Gene Expression Profiles Using Machine Learning Algorithms</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Lin</surname> <given-names>Ping-I</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1063429/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Moni</surname> <given-names>Mohammad Ali</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/109205/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Gau</surname> <given-names>Susan Shur-Fen</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Eapen</surname> <given-names>Valsamma</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/54157/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Psychiatry, The University of New South Wales</institution>, <addr-line>Sydney, NSW</addr-line>, <country>Australia</country></aff>
<aff id="aff2"><sup>2</sup><institution>South Western Sydney Local Health District</institution>, <addr-line>Liverpool, NSW</addr-line>, <country>Australia</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Psychiatry, National Taiwan University Hospital and College of Medicine</institution>, <addr-line>Taipei</addr-line>, <country>Taiwan</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Claudio Toma, Severo Ochoa Molecular Biology Center (CSIC-UAM), Spain</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Richard Woodman, Flinders University, Australia; No&#x000E8;lia Fern&#x000E0;ndez-Castillo, Centre for Biomedical Network Research (CIBER), Spain</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Ping-I Lin <email>daniel.lin&#x00040;unsw.edu.au</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Behavioral and Psychiatric Genetics, a section of the journal Frontiers in Psychiatry</p></fn></author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>05</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>637022</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>12</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>04</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Lin, Moni, Gau and Eapen.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Lin, Moni, Gau and Eapen</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license></permissions>
<abstract><p><bold>Objectives:</bold> The identification of subgroups of autism spectrum disorder (ASD) may partially remedy the problems of clinical heterogeneity to facilitate the improvement of clinical management. The current study aims to use machine learning algorithms to analyze microarray data to identify clusters with relatively homogeneous clinical features.</p>
<p><bold>Methods:</bold> The whole-genome gene expression microarray data were used to predict communication quotient (SCQ) scores against all probes to select differential expression regions (DERs). Gene set enrichment analysis was performed for DERs with a fold-change &#x0003E;2 to identify hub pathways that play a role in the severity of social communication deficits inherent to ASD. We then used two machine learning methods, random forest classification (RF) and support vector machine (SVM), to identify two clusters using DERs. Finally, we evaluated how accurately the clusters predicted language impairment.</p>
<p><bold>Results:</bold> A total of 191 DERs were initially identified, and 54 of them with a fold-change &#x0003E;2 were selected for the pathway analysis. Cholesterol biosynthesis and metabolisms pathways appear to act as hubs that connect other trait-associated pathways to influence the severity of social communication deficits inherent to ASD. Both RF and SVM algorithms can yield a classification accuracy level &#x0003E;90% when all 191 DERs were analyzed. The ASD subtypes defined by the presence of language impairment, a strong indicator for prognosis, can be predicted by transcriptomic profiles associated with social communication deficits and cholesterol biosynthesis and metabolism.</p>
<p><bold>Conclusion:</bold> The results suggest that both RF and SVM are acceptable options for machine learning algorithms to identify AD subgroups characterized by clinical homogeneity related to prognosis.</p></abstract>
<kwd-group>
<kwd>autism spectrum disorder</kwd>
<kwd>genomics</kwd>
<kwd>social cognition</kwd>
<kwd>language</kwd>
<kwd>machine learning</kwd>
</kwd-group>
<counts>
<fig-count count="5"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="67"/>
<page-count count="10"/>
<word-count count="6555"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Clinical heterogeneity is a norm rather than an exception in autism spectrum disorder (ASD), a complex neurodevelopmental disorder characterized by social communication deficits and stereotyped behaviors. Heterogeneous clinical features pose great challenges for diagnostics for ASD, such that children who receive a diagnosis of ASD have a range of vastly different presentations, trajectories, and outcomes. Further, the diagnostic criteria for ASD have been continuously revised through different editions of the Diagnostic and Statistical Manual for Mental Disorders (DSM), particularly the substantial changes in the 5th edition (DSM 5) where the wide range of clinical presentations have been brought together under a single ASD diagnostic entity (<xref ref-type="bibr" rid="B1">1</xref>). The current diagnostic system lacks an evidence-based approach and we urgently require a scientific approach to understanding which interventions are likely to be the most effective for which child with ASD (<xref ref-type="bibr" rid="B2">2</xref>). Accumulating evidence has shown that no pharmaceutical treatments have thus far been conclusively found to substantially reduce core symptoms of ASD (<xref ref-type="bibr" rid="B3">3</xref>). This may be partially attributable to the fact that most clinical trials did not take clinical heterogeneity into account and hence treatment effects remain equivocal. Variable clinical presentations may reflect different biological pathways. The identification of biomarkers for etiological pathways may hence hold the key to unraveling mechanisms underlying the variation in clinical presentations (<xref ref-type="bibr" rid="B4">4</xref>), which in turn may pave the way for personalized medicine in ASD.</p>
<p>The goal of identifying biomarkers for clinical homogeneity is to tackle challenges arising from clinical heterogeneity for research on either etiologies or treatments of ASD. One of the most extensively studied biomarkers for ASD is genetic factors. There are two different strategies to evaluate genetic markers for clinical heterogeneity: bottom-up and top-down approaches. The bottom-up approach is to define a priori subgroups using phenotypic information under the premise that some genetic loci are more likely to contribute to susceptibility to disease in a certain subgroup(s). Therefore, stratifying the population by a clinical marker (e.g., age of onset) will allow investigators to detect genetic association effects that are larger in certain subgroups. The top-down approach, on the other hand, is based on the premise that certain genetic markers can be used to distinguish subgroups, each of which is characterized by relatively homogeneous phenotypic profiles underscored by similar biological pathways&#x02014;which imply similar therapeutic targets. Many of the earlier genome-wide linkage or association studies that aimed to unravel genetic underpinnings of clinical heterogeneity chose the second approach, which is to identify genetic markers associated with the phenotype defined by strict diagnostic criteria of ASD (<xref ref-type="bibr" rid="B5">5</xref>&#x02013;<xref ref-type="bibr" rid="B7">7</xref>). Using the data from the Autism Diagnostic Interview-Revised (ADI-R) (<xref ref-type="bibr" rid="B8">8</xref>), Autism Diagnostic Observation Schedule (ADOS) (<xref ref-type="bibr" rid="B9">9</xref>), Vineland Adaptive Behavior Scales (VABS) (<xref ref-type="bibr" rid="B10">10</xref>), head circumferences, and ages at assessment as classifying variables, Veatch et al. identified clinically similar subgroups of individuals with ASD and found that the genotypes were more similar within subgroups compared to the whole population&#x02014;the proportion of the total genetic variance contained in a subpopulation was 0.17 (<xref ref-type="bibr" rid="B11">11</xref>). However, this approach has not yielded highly replicable and clinically meaningful findings that can lead to conclusively validated etiological factors yet (<xref ref-type="bibr" rid="B12">12</xref>). Furthermore, another genome-wide association study of 2,576 families with ASD probands did not discover any genetic loci that exert a larger effect on the disease risk in subpopulations defined by the diagnosis, IQ, and symptom profiles; heritability estimates were also found to be similar in subpopulations to the whole population (<xref ref-type="bibr" rid="B13">13</xref>). Results from different groups show that an increased number of gene-truncating variants (highly pathogenic variants) may exert a considerable impact on IQ in ASD patients (<xref ref-type="bibr" rid="B14">14</xref>, <xref ref-type="bibr" rid="B15">15</xref>); and higher burden of this pool of variants in ASD patients correlates with lower IQ scores. These studies showed that genomic approaches are able to identify genetic loci exerting larger effect on disease risk or associated with clinical outcomes, although genetic loci must be considered in an additive manner.</p>
<p>The top-down approach often starts with a few selected genetic loci associated with the disease. Despite fruitful findings from genome-wide and candidate gene-based association studies, few genetic loci can be used to improve accuracy in diagnostics or optimize treatment effects of therapeutics for ASD. Nevertheless, several genetic markers are found to be useful for classifying patients with ASD into relatively homogeneous subgroups. For example, Bruining et al. reported prominently higher symptom homogeneity in both the ASD group with 22q11 deletions and ASD group with Klinefelter Syndrome (KS), compared to the heterogeneous ASD sample (<xref ref-type="bibr" rid="B16">16</xref>). Transcriptomic profiles have also been used to identify genetic markers to classify individuals with ASD. Hu and Lai used the gene expression data to identify a subset of the &#x0201C;classifier&#x0201D; genes, which resulted in an overall class prediction accuracy of nearly 82%, &#x0007E;90% sensitivity, and 75% specificity (<xref ref-type="bibr" rid="B17">17</xref>). These results seem to demonstrate the value of the top-down approach.</p>
<p>Determining subgroups of ASD is challenging mainly because of the complexity of biological factors and clinical heterogeneity inherent to ASD. To tackle these challenges, one of the solutions is to implement state-of-the-art statistical methods that can efficiently parse through high-dimensionality data, such as machine learning (ML) algorithms, to differentiate subgroups with meaningful etiological, diagnostic, or therapeutic implications (<xref ref-type="bibr" rid="B18">18</xref>). Previous evidence suggests that ML algorithms can be used to reduce the number of items from standardized ASD assessment tools to make the assessment more efficient (<xref ref-type="bibr" rid="B19">19</xref>) and predict clinical outcomes with ASD phenotypic clusters and genetic data of copy number variations (<xref ref-type="bibr" rid="B20">20</xref>). The ML algorithms appear to be useful to identify phenotypic clusters as ASD subgroups that can predict clinical outcomes (<xref ref-type="bibr" rid="B21">21</xref>). In the current study, we attempted to implement the ML algorithms in the context of the bottom-up approach, which is to identify clusters using genomic information, and then explore the relationship between the genomic clusters and clinical features of ASD.</p></sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<sec>
<title>Data Collection</title>
<p>The goal of the current study is to evaluate whether transcriptomic profiles correlated with clinical severity levels of ASD&#x02014;which were measured with social communication questionnaire (SCQ) (<xref ref-type="bibr" rid="B22">22</xref>), can classify patients into two subgroups defined on the basis of language (i.e., the subgroup with language impairment vs. the subgroup without language impairment). The language function is considered as a strong predictor for cognitive ability and adaptive skills in children with ASD (<xref ref-type="bibr" rid="B23">23</xref>), and its variation within ASD patients is influenced by genetic factors (<xref ref-type="bibr" rid="B24">24</xref>&#x02013;<xref ref-type="bibr" rid="B26">26</xref>). The presence of language impairment was defined as the total score (verbal) &#x0003E;10 in the section of Qualitative Abnormalities in Communication in Autism Diagnostic Interview-Revised (ADI-R) (<xref ref-type="bibr" rid="B27">27</xref>). A total of 31 children diagnosed with ASD were recruited in the current study. The clinical diagnoses were made by Gau, a board-certified child psychiatrist, and confirmed by the ADI-R interview with the parents. The Chinese version of the ADI-R been approved by the Western Psychological Services in May 2007 (<xref ref-type="bibr" rid="B28">28</xref>) mRNA was extracted from lymphoblastoid cell lines (LCL) of all participants. The microarray experiment was performed at the Core Laboratory of National Taiwan University Hospital in Taiwan, using the Affymetrix Human Genome U133 Plus 2.0 Array (Affymetrix Inc., Santa Clara, CA, USA). The experimental procedures followed the protocols provided by the manufacturer. The study was conducted with the ethical approval by the Institutional Research Board at National Taiwan University Hospital in Taiwan.</p></sec>
<sec>
<title>Statistical Methods</title>
<sec>
<title>Transcriptome-Wide Association Analysis</title>
<p>We evaluated the integrity of 28S and 18S rRNA by electrophoresis of 2 mg of total RNA in 1.2% agarose gel containing 2.2 M formaldehyde and in a running buffer containing 0.2 M of MOPS (pH 7.0), 20 mM of sodium acetate and 10 mM of EDTA (pH 8.0). The A260/A280 ratio was used to measure the quality of RNA. The ratio between 1.9 and 2.1 was considered good quality. The intensity files of all the subjects were input into the computer program GAP: Generalized Association Plots (<xref ref-type="bibr" rid="B29">29</xref>, <xref ref-type="bibr" rid="B30">30</xref>) for quality control using visualization and descriptive statistics. We used the Robust Multi-array Analysis (RMA) method to normalize the data (<xref ref-type="bibr" rid="B31">31</xref>). In order to filter out probe sets with low variations and to reduce the impact of multiple comparisons, we kept only the 1,000 probe sets with the largest standard deviations. We searched for differential expression regions (DERs) by prioritizing the gene expression levels associated with the clinical severity indicated by SCQ scores, we used the generalized linear model to screen for probes across the whole genome with mRNA levels associated with the SCQ scores with unadjusted <italic>p</italic> &#x0003C; 0.00001. All original intensity ratio data were transformed into logarithmic 2 values after being normalized. We controlled for the batch effect by adjusting for the batch as a binary covariate since there were two batches. These probes constitute the primary source of predictors to determine ASD subgroups.</p></sec>
<sec>
<title>Gene Ontology and Pathway Analysis</title>
<p>The DERs with a fold-change &#x0003E;2 were selected for the gene ontology and pathway analysis to evaluate the biological relevance and functional pathways of the significant genes. We have incorporated the KEGG (<xref ref-type="bibr" rid="B32">32</xref>), WikiPathways (<xref ref-type="bibr" rid="B33">33</xref>), BioCarta (<xref ref-type="bibr" rid="B34">34</xref>), and Reactome (<xref ref-type="bibr" rid="B35">35</xref>) pathway database for the cell signaling pathways. We have also considered the GO Biological Process (2018) database for gene ontological analysis (<xref ref-type="bibr" rid="B36">36</xref>). The GO terms and pathways enriched by the list of genes were identified using the hypergeometric analyses with an adjusted <italic>P</italic> &#x02264; 0.05 was considered as statistically significant.</p></sec>
<sec>
<title>Gene Over-Representation Analysis</title>
<p>Then we used the webtool at ConsensusPathDB (<ext-link ext-link-type="uri" xlink:href="http://cpdb.molgen.mpg.de/">http://cpdb.molgen.mpg.de/</ext-link>) to identify pathway-pathway interaction network (CPDB analysis) (<xref ref-type="bibr" rid="B37">37</xref>). The analysis criteria included: (1) one-next neighbors for the radius with <italic>p</italic> &#x0003C; 0.01, (2) pathway-based sets at least two overlapped genes and <italic>p</italic> &#x0003C; 0.01, and (3) gene ontology level 2 categories with <italic>p</italic> &#x0003C; 0.01. The results from the second approach helped visualize the possible &#x0201C;hub&#x0201D; pathway from the top 10 networks associated with the candidate genes.</p>
<p>We chose two machine learning (ML) algorithms to evaluate the clustering results: random forest classification and support vector machine algorithms. The presence of language impairment was considered as a dichotomous clinical outcome to determine classification errors. We chose the first ML algorithm proposed by Shi and Horvath (<xref ref-type="bibr" rid="B38">38</xref>). We used the Random Forest classification (RF) algorithm in an unsupervised mode to generate a proximity matrix. The gene expression data were analyzed using RF using two different approaches for comparison. The first approach is to reduce data dimensionality using principal component analysis to identify principal component (PC) scores for each subject. The top 10 PCs were selected to calculate the proximity matrix that provides a rough estimate of the distance between samples based on the proportion of times the samples end up in the same leaf node. The proximity matrix values were then converted to a dissimilarity matrix to classify the sample into two subgroups using partitioning around medoid (PAM) (<xref ref-type="bibr" rid="B39">39</xref>). The second approach is to use the information of all 191 probes with gene expression levels significantly associated with SCQ scores to generate the RF proximity matrix. Similarly, the RF proximity matrix was used to classify the sample into two subgroups using the PAM clustering analysis (<xref ref-type="bibr" rid="B39">39</xref>) to classify the patients into two clusters to determine the final cluster assignment. The RF-PAM clustering analysis could allow us to evaluate the classification error by calculating the frequency of patients with language impairment in the cluster, in which the majority of patients had no language impairment, and vice versa.</p>
<p>We further chose Support Vector Machine (SVM) as the second ML algorithm to classify the patients into two subgroups (<xref ref-type="bibr" rid="B40">40</xref>). To reduce data dimensionality, we implemented principal component analysis to identify the principal component (PC) scores for each subject. The data of PC scores were split in a 7:3 ratio&#x02014;in other words, 70% of the data was used for training the model and the remaining 30% was for testing the model. Estimating the C (Cost) parameter to classify the data was performed using SVM with the linear kernel function. The choice of kernel function was made based on the recommendation from a prior study that using microarray data to predict the diagnosis of colon cancer&#x02014;which concludes that linear kernel function leads to a lower prediction error than the RBF, quadratic, and polynomial kernel functions (<xref ref-type="bibr" rid="B41">41</xref>). The prediction accuracy and Kappa value estimated when the C value was held constant at 1. The Kappa value was calculated using the formula (<italic>p</italic><sub>o</sub> &#x02013; <italic>p</italic><sub>e</sub>)/(1-<italic>p</italic><sub>e</sub>), where <italic>p</italic><sub>o</sub> and <italic>p</italic><sub>e</sub> denote the observed agreement and expected agreement for classification, respectively. We further used the confusion matrix, which contains the number of correct and incorrect predictions summarized with count values and broken down by each class, to predict the prediction accuracy of the SVM model. The accuracy is calculated as (TP &#x0002B; TN)/(TP&#x0002B;TN&#x0002B;FP&#x0002B;FN), where TP and TN refer to true positives and true negatives, respectively; FP and FN refer to false positives and false negatives, respectively. These two measures (i.e., accuracy and Kappa value) were chose to evaluate the SVM performance as recommended by previous studies (<xref ref-type="bibr" rid="B42">42</xref>, <xref ref-type="bibr" rid="B43">43</xref>). The Kappa statistics could lead to a biased performance estimate in unbalanced situations (<xref ref-type="bibr" rid="B44">44</xref>), which is not the characteristic of the current sample. The SVM analysis was performed using the R package &#x0201C;<italic>caret</italic>&#x0201D; (<xref ref-type="bibr" rid="B45">45</xref>).</p></sec></sec></sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>The workflow of the current project is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The clinical features of the 31 subjects are summarized in <xref ref-type="table" rid="T1">Table 1</xref>. The group with language impairment and the group without language impairment has significant differences in clinical features associated with both social communication function and verbal IQ scores.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The workflow of the study scheme.</p></caption>
<graphic xlink:href="fpsyt-12-637022-g0001.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Clinical features of the patients in the current study.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>Language impairment (51.3%)</bold></th>
<th valign="top" align="center"><bold>No language impairment (48.7%)</bold></th>
<th valign="top" align="center"><bold>Relationship with language impairment<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Age</td>
<td valign="top" align="center">9.00 (<italic>SD</italic>: 2.52)</td>
<td valign="top" align="center">8.91 (<italic>SD</italic>: 3.99)</td>
<td valign="top" align="center"><italic>P</italic> &#x0003E; 0.05</td>
</tr>
<tr>
<td valign="top" align="left">ADIR-BV</td>
<td valign="top" align="center">17.83 (<italic>SD</italic>: 3.27)</td>
<td valign="top" align="center">8.55 (<italic>SD</italic>: 1.13)</td>
<td valign="top" align="center"><italic>P</italic> &#x0003C; 0.0001</td>
</tr>
<tr>
<td valign="top" align="left">ADIR-BN</td>
<td valign="top" align="center">8.92 (<italic>SD</italic>: 2.71)</td>
<td valign="top" align="center">3.64 (<italic>SD</italic>: 1.43)</td>
<td valign="top" align="center"><italic>P</italic> &#x0003C; 0.0001</td>
</tr>
<tr>
<td valign="top" align="left">SCQ</td>
<td valign="top" align="center">22.19 (<italic>SD</italic>: 4.84)</td>
<td valign="top" align="center">11.47 (<italic>SD</italic>: 4.84)</td>
<td valign="top" align="center"><italic>P</italic> &#x0003C; 0.0001</td>
</tr>
<tr>
<td valign="top" align="left">VIQ</td>
<td valign="top" align="center">82.08 (<italic>SD</italic>: 20.77)</td>
<td valign="top" align="center">111.91 (<italic>SD</italic>: 10.12)</td>
<td valign="top" align="center"><italic>P</italic> = 0.0003</td>
</tr>
<tr>
<td valign="top" align="left">PIQ</td>
<td valign="top" align="center">90.83 (<italic>SD</italic>: 15.74)</td>
<td valign="top" align="center">101.36 (<italic>SD</italic>: 15.34)</td>
<td valign="top" align="center"><italic>P</italic> &#x0003E; 0.05</td>
</tr>
<tr>
<td valign="top" align="left">SRS</td>
<td valign="top" align="center">89.61 (<italic>SD</italic>: 16.12)</td>
<td valign="top" align="center">79.55 (<italic>SD</italic>: 27.99)</td>
<td valign="top" align="center"><italic>P</italic> &#x0003E; 0.05</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>ADIR-BV, Autism Diagnostic Interview&#x02013;Revised, Qualitative Abnormalities in Communication, Total Verbal score. ADIR-BN, Autism Diagnostic Interview&#x02013;Revised, Qualitative Abnormalities in Communication, Total Non-Verbal score. SCQ, Social Communication Questionnaire score; VIQ, verbal IQ; PIQ, performance IQ; SRS, Social Responsiveness Scale score</italic>.</p>
<fn id="TN1"><label>&#x0002A;</label><p><italic>The student&#x00027;s t-test was performed to evaluate whether the the two subgroups classified by the presence of language impairment had different values in each continuous variable</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>The transcriptomic association study reveals 191 probes that were statistically significantly associated with SCQ scores with a <italic>p</italic> &#x0003C; 0.00001. The batch effect seemingly did not affect the association results (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure 1</xref>). We selected 54 of them with a fold-change &#x0003E;2 for the pathway analysis. Differentially expressed 54 genes with logarithmic fold changes and &#x02013;logarithmic 10 adjusted <italic>p</italic>-values are listed in <xref ref-type="fig" rid="F2">Figure 2</xref>. Only three pathways were found to be over-represented by these 54 genes with adjusted <italic>p</italic> &#x0003C; 0.05: cholesterol biosynthetic process (GO:0006695), secondary alcohol biosynthetic process (GO:1902653), and regulation of signal transduction by p53 class mediator (GO:1901796). The CPBD analysis shows that Sterol Regulatory Element-Binding Proteins (SREBP) signaling pathway is the pathway connected with 9 of the 10 pathways including cholesterol biosynthetic pathway, so it can be regarded as the &#x0201C;hub&#x0201D; associated with genetic network for ASD (<xref ref-type="fig" rid="F3">Figure 3</xref>). This pathway of SREBP focuses on the regulation of lipid metabolism by SREBP.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Differentially expressed 54 genes with fold changes and &#x02013;logarithmic 10 adjusted <italic>p</italic>-values. The red circle represents logarithmic fold change and the blue color circle represents &#x02013;logarithmic 10 adjusted <italic>p</italic>-value for each significant gene.</p></caption>
<graphic xlink:href="fpsyt-12-637022-g0002.tif"/>
</fig>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Gene network analysis. The relationship among pathways enriched with candidate genes with expression levels associated with SCQ scores is shown.</p></caption>
<graphic xlink:href="fpsyt-12-637022-g0003.tif"/>
</fig>
<p>The RF-PAM analysis identified two clusters (<xref ref-type="fig" rid="F4">Figure 4</xref>). The classification accuracy was 67.7% when the top 10 PCs were used to generate the proximity matrix, while the classification accuracy was 96.9% when all 191 probes were used to generate the proximity matrix. The SVM analysis based on the top 10 PC scores shows that the clustering results reached classification accuracy at 93.3% (95% CI 68.1&#x02013;99.8%) and no-information rate (i.e., the largest proportion of the observed classes) at 53.3% (<italic>p</italic> = 0.0011). Other parameters relevant to prediction performance include Kappa value = 0.86, sensitivity = 0.86, specificity = 1.00, and balanced accuracy = 0.93. The SVM analysis using the information of all probes with differential gene expressions associated with SCQ scores yielded a slightly higher classification accuracy than the SVM analysis based on the top 10 PC scores. The classification accuracy at 99.9% (95% CI 78.2&#x02013;100%) and no-information rate (i.e., the largest proportion of the observed classes) at 53.3% (<italic>p</italic> = 8.035 &#x000D7; 10<sup>&#x02212;5</sup>) were achieved when 191 probes were analyzed. This classification accuracy can be demonstrated in gene expression level distributions stratified by language impairment (<xref ref-type="supplementary-material" rid="SM2">Supplementary Figure 2</xref>). The SVM clustering results are shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. The results suggest that the first two principal components could identify support vectors that fell in the area with better prediction confidence (<xref ref-type="fig" rid="F5">Figure 5A</xref>), compared with the results predicted by individual probes (<xref ref-type="fig" rid="F5">Figure 5B</xref>). The predicting performance of the RF-PAM and SVM algorithms is listed in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>ASD subgroups identified using RF and PAM clustering algorithms. Dim1 and Dim2 correspond to principal components 1 and 2, respectively. <bold>(A,B)</bold> The results based on the top 10 principal components (PCs) and the 191 probes, respectively. We used the first two predictors to make the plots to demonstrate how different approaches classified the sample.</p></caption>
<graphic xlink:href="fpsyt-12-637022-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>SVM clustering results based on the top PCs. <bold>(A)</bold> Shows the color gradient that indicates how confidently a new point would be classified based on its features. PC1 and PC2 represent the first and second principal components, respectively. <bold>(B)</bold> Shows the color gradient that indicates how confidently a new point would be classified based on its features when predictors were based on all SCQ-associated probes. Probe 1 and probe 2 represent the first and second probes, respectively. The solid symbols indicate the support vectors and the hollow circles indicates other subjects. The circles and triangles represent the first and second subgroups, respectively.</p></caption>
<graphic xlink:href="fpsyt-12-637022-g0005.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Predicting performance of two machine learning algorithms.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Predictors</bold></th>
<th valign="top" align="center"><bold>Prediction accuracy</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">RF-PAM</td>
<td valign="top" align="center">191 probes</td>
<td valign="top" align="center">96.90%</td>
</tr>
<tr>
<td valign="top" align="left">RF-PAM</td>
<td valign="top" align="center">10 PC<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">67.70%</td>
</tr>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">191 probes</td>
<td valign="top" align="center">99.90%</td>
</tr>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">10 PC<xref ref-type="table-fn" rid="TN2"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">93.30%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN2"><label>&#x0002A;</label><p><italic>Principal component</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4">
<title>Discussions</title>
<p>We conducted a proof-of-concept study to demonstrate how transcriptomic data from a small sample could provide useful biomarkers to classify ASD subgroups. The selection of the predictors was based on DERs associated with SCQ scores, which indicate the variation in severity levels of social communication deficits, a hallmark clinical feature of ASD. The DER with strongest evidence for the association with social deficits in our sample is matched with the HEATR1 gene (HEAT Repeat Containing 1). The HEART1 gene is associated with schizophrenia (<xref ref-type="bibr" rid="B46">46</xref>). The HEATR1 gene abnormalities in the brain during the embryonic stage has been reported in zebrafish (<xref ref-type="bibr" rid="B47">47</xref>). The candidate genes that harbor these DERs suggest several genetic pathways that modulate the variation in social communication functions. Among these pathways, the pathway of cholesterol biosynthesis/metabolism and sterol regulatory element-binding proteins (SREBP) pathway&#x02014;cholesterol metabolism appear to act as hubs that connect other top SCQ-associated pathways. Particularly, the SREBP pathway shares most genes with other SCQ-associated pathways. These two pathways are related to lipid metabolism. Cholesterol synthesis and uptake are tightly modulated at the transcriptional level through negative feedback control, which is regulated by SREBPs (<xref ref-type="bibr" rid="B48">48</xref>). The relationship between lipid metabolism and brain functions has been well-documented. A growing body of evidence has indicated that cholesterol metabolism plays a key role in synaptic functions (<xref ref-type="bibr" rid="B49">49</xref>&#x02013;<xref ref-type="bibr" rid="B51">51</xref>). Dysregulated cholesterol metabolism has been extensively documented in ASD (<xref ref-type="bibr" rid="B51">51</xref>&#x02013;<xref ref-type="bibr" rid="B58">58</xref>). A recent study implemented a personalized medicine approach combining healthcare claims, electronic health records, familial whole-exome sequences, and neurodevelopmental gene expression patterns, and identified an ASD subtype characterized by dyslipidemia (<xref ref-type="bibr" rid="B59">59</xref>). There are certainly several other genetic pathways involved in molecular mechanisms underlying social communication deficits. Nevertheless, our results indicate that cholesterol synthesis/metabolism pathways act as hubs that connect most other biological pathways, which suggest that the genomic functional changes associated with lipid metabolism may moderate other genomic changes, such as the p53 signaling pathway, that regulate social communication functions.</p>
<p>Using the DERs as biomarkers, we clustered the sample into two subgroups using two different ML algorithms. Both the RF-PAM and SVM analyses yielded similar levels of classification accuracy when all 191 markers were utilized. However, compared to the analysis using the RF-PAM algorithm, the analysis using the SVM algorithm seemed to be more robust when we performed dimension reduction for all the 191 markers with the PCA method. The RF algorithm is applicable when there are more predictors than observations, relatively insensitive to the noise (e.g., a large number of irrelevant genes), and does not rely on excessive fine-tuning of parameters (<xref ref-type="bibr" rid="B60">60</xref>). RF algorithm is more robust to small sample size as the SVM algorithm (<xref ref-type="bibr" rid="B61">61</xref>, <xref ref-type="bibr" rid="B62">62</xref>). However, Brown et al. found that SVM outperforms other techniques that include Fisher&#x00027;s linear discriminant, Parzen window, and tow decision tree learners when using gene expression data to predict clinical outcomes (<xref ref-type="bibr" rid="B63">63</xref>). Additionally, Statnikov et al. conducted a comprehensive comparison of RF and SVM using microarray data for 22 diagnostic and prognostic datasets and concluded that SVM is superior to RF in terms of classification accuracy (<xref ref-type="bibr" rid="B64">64</xref>). Although the purpose of this study is not to comprehensively evaluate which ML algorithm outperforms the other ML algorithm, our results seem to lend some support to the robustness of the SVM algorithm. Nevertheless, the RF algorithm is at least as robust as the SVM algorithm when the dimension of input variables is not substantially reduced.</p>
<p>One of the major limitation of the current study is the small sample size. Nevertheless, some machine learning algorithm, such as SVM, can handle a small sample with a large number of features. Additionally, model overfitting may arise due to a lack of another independent sample for validation. Furthermore, unknown confounders may cause spurious associations between the phenotype and genomic markers. However, the goal of this proof-of-concept study is prediction of subtypes rather than the identification of etiologies. Therefore, confounders would not yield a substantial impact on prediction results (<xref ref-type="bibr" rid="B65">65</xref>).</p>
<p>The clinical and etiological heterogeneity in ASD has meant that there is considerable variability in treatment outcomes across different interventions and between individuals receiving the same intervention. Hence the traditional diagnostic and &#x0201C;one size fits all&#x0201D; approach to ASD intervention needs improvement. Further, we currently do not have a sufficient understanding of &#x0201C;what would work for whom,&#x0201D; thereby limiting opportunities for maximizing outcomes for children and their families with economic ramifications for broader society. In this context, ML algorithms have been found to be useful in predicting diagnostic accuracy in ASD with neuroimaging data (<xref ref-type="bibr" rid="B66">66</xref>). Further, one recent study used Gaussian Mixture Models and Hierarchical Agglomerative Clustering, which provide a statistical framework for learning latent cluster memberships to determine ASD subgroups with differentiated treatment responses (<xref ref-type="bibr" rid="B67">67</xref>). Our findings that using ML algorithms, children could be classified into two groups based on the presence of language impairment, offers promise for unraveling clinically meaningful subgroups in ASD. This, in turn, can be used for predicting likely responsiveness (and non-responsiveness) to specific treatment pathways. This &#x0201C;precision&#x0201D; approach to assessment and intervention will ensure that resources for appropriate intervention and supports are allocated in an evidence-based manner. This is critical as without timely recognition of the variability in the clinical presentation, neurocognitive level of functioning, and psychosocial circumstances coupled with individualized intervention, children and their families may miss key opportunities of brain plasticity available in the critical early years. ML techniques as utilized in this study offer a viable solution to address this by allowing matching interventions and supports that are tailored to the individual profile and needs.</p></sec>
<sec sec-type="data-availability-statement" id="s5">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found at: <ext-link ext-link-type="uri" xlink:href="https://figshare.com/articles/dataset/Autism_gene_expression_data/14251328">https://figshare.com/articles/dataset/Autism_gene_expression_data/14251328</ext-link>.</p></sec>
<sec id="s6">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by Research Ethics Committee of the National Taiwan University Hospital. Written informed consent to participate in this study was provided by the participants&#x00027; legal guardian/next of kin.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>P-IL and MM carried out the statistical analysis. P-IL and VE conceived of the study and drafted the manuscript. SG participated in the study design and coordination. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
</body>
<back>
<ack><p>The authors would like to thank the subjects who participated in this study and facility support at National Taiwan University Hospital.</p>
</ack>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpsyt.2021.637022/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpsyt.2021.637022/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Image_1.TIFF" id="SM1" mimetype="image/tiff" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 1</label>
<caption><p>The evaluation of potential batch effect due to the microarrays timing. <bold>(A)</bold> The kernel density distributions of gene expression levels of the two batches are shown. <bold>(B)</bold> Time 1 and time 2 indicate the association test results that adjusted for the time (i.e., batch) vs. the results without adjusting for the time.</p></caption></supplementary-material>
<supplementary-material xlink:href="Image_2.TIFF" id="SM2" mimetype="image/tiff" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Supplementary Figure 2</label>
<caption><p>Randomly selected four probes associated with SCQ scores stratified by the presence of language impairment. The red and blue curves represent the group without language impairment and the group with language impairment, respectively.</p></caption></supplementary-material></sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="book"><person-group person-group-type="author"><collab>American Psychiatric Association</collab></person-group> (<year>2013</year>). <source>Diagnostic and statistical manual of mental disorders (5th ed.)</source>. <publisher-loc>Arlington</publisher-loc>: <publisher-name>American Psychiatric Association</publisher-name>. p. <fpage>31</fpage>&#x02013;<lpage>2</lpage>. <pub-id pub-id-type="doi">10.1176/appi.books.9780890425596</pub-id></citation></ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eapen</surname> <given-names>V</given-names></name></person-group>. <article-title>Genetic basis of autism: is there a way forward?</article-title> <source>Curr Opin Psychiatry.</source> (<year>2011</year>) <volume>24</volume>:<fpage>226</fpage>&#x02013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1097/YCO.0b013e328345927e</pub-id><pub-id pub-id-type="pmid">21460645</pub-id></citation></ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bowers</surname> <given-names>K</given-names></name> <name><surname>Lin</surname> <given-names>P-I</given-names></name> <name><surname>Erickson</surname> <given-names>C</given-names></name></person-group>. <article-title>Pharmacogenomic medicine in autism: challenges and opportunities</article-title>. <source>Pediatr Drugs.</source> (<year>2015</year>) <volume>17</volume>:<fpage>115</fpage>&#x02013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1007/s40272-014-0106-0</pub-id><pub-id pub-id-type="pmid">25420674</pub-id></citation></ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McPartland</surname> <given-names>JC</given-names></name> <name><surname>Bernier</surname> <given-names>RA</given-names></name> <name><surname>Jeste</surname> <given-names>SS</given-names></name> <name><surname>Dawson</surname> <given-names>G</given-names></name> <name><surname>Nelson</surname> <given-names>CA</given-names></name> <name><surname>Chawarska</surname> <given-names>K</given-names></name> <etal/></person-group>. <article-title>The autism biomarkers consortium for clinical trials (ABC-CT): scientific context, study design, and progress toward biomarker qualification</article-title>. <source>Front Integr Neurosci.</source> (<year>2020</year>) <volume>14</volume>:<fpage>16</fpage>. <pub-id pub-id-type="doi">10.3389/fnint.2020.00016</pub-id><pub-id pub-id-type="pmid">32346363</pub-id></citation></ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anney</surname> <given-names>R</given-names></name> <name><surname>Klei</surname> <given-names>L</given-names></name> <name><surname>Pinto</surname> <given-names>D</given-names></name> <name><surname>Regan</surname> <given-names>R</given-names></name> <name><surname>Conroy</surname> <given-names>J</given-names></name> <name><surname>Magalhaes</surname> <given-names>TR</given-names></name> <etal/></person-group>. <article-title>A genome-wide scan for common alleles affecting risk for autism</article-title>. <source>Hum Mol Genet.</source> (<year>2010</year>) <volume>19</volume>:<fpage>4072</fpage>&#x02013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1093/hmg/ddq307</pub-id><pub-id pub-id-type="pmid">20663923</pub-id></citation></ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yonan</surname> <given-names>AL</given-names></name> <name><surname>Alarc&#x000F3;n</surname> <given-names>M</given-names></name> <name><surname>Cheng</surname> <given-names>R</given-names></name> <name><surname>Magnusson</surname> <given-names>PKE</given-names></name> <name><surname>Spence</surname> <given-names>SJ</given-names></name> <name><surname>Palmer</surname> <given-names>AA</given-names></name> <etal/></person-group>. <article-title>A genomewide screen of 345 families for autism-susceptibility loci</article-title>. <source>Am J Hum Genet.</source> (<year>2003</year>) <volume>73</volume>:<fpage>886</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1086/378778</pub-id><pub-id pub-id-type="pmid">13680528</pub-id></citation></ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J</given-names></name> <name><surname>Nyholt</surname> <given-names>DR</given-names></name> <name><surname>Magnussen</surname> <given-names>P</given-names></name> <name><surname>Parano</surname> <given-names>E</given-names></name> <name><surname>Pavone</surname> <given-names>P</given-names></name> <name><surname>Geschwind</surname> <given-names>D</given-names></name> <etal/></person-group>. <article-title>A genomewide screen for autism susceptibility loci</article-title>. <source>Am J Hum Genet.</source> (<year>2001</year>) <volume>69</volume>:<fpage>327</fpage>&#x02013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1086/321980</pub-id><pub-id pub-id-type="pmid">13680528</pub-id></citation></ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lord</surname> <given-names>C</given-names></name> <name><surname>Rutter</surname> <given-names>M</given-names></name> <name><surname>Le Couteur</surname> <given-names>A</given-names></name></person-group>. <article-title>Autism Diagnostic Interview-Revised: a revised version of a diagnostic interview for caregivers of individuals with possible pervasive developmental disorders</article-title>. <source>J Autism Dev Disord.</source> (<year>1994</year>) <volume>24</volume>:<fpage>659</fpage>&#x02013;<lpage>85</lpage>.<pub-id pub-id-type="pmid">7814313</pub-id></citation></ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gotham</surname> <given-names>K</given-names></name> <name><surname>Pickles</surname> <given-names>A</given-names></name> <name><surname>Lord</surname> <given-names>C</given-names></name></person-group>. <article-title>Standardizing ADOS scores for a measure of severity in autism spectrum disorders</article-title>. <source>J Autism Dev Disord.</source> (<year>2009</year>) <volume>39</volume>:<fpage>693</fpage>&#x02013;<lpage>705</lpage>. <pub-id pub-id-type="doi">10.1007/s10803-008-0674-3</pub-id><pub-id pub-id-type="pmid">19082876</pub-id></citation></ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Icabone</surname> <given-names>DG</given-names></name></person-group>. <article-title>Vineland adaptive behavior scales</article-title>. <source>Diagnostique.</source> (<year>1999</year>) <volume>24</volume>:<fpage>257</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1177/153450849902401-423</pub-id></citation></ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Veatch</surname> <given-names>OJ</given-names></name> <name><surname>Veenstra-Vanderweele</surname> <given-names>J</given-names></name> <name><surname>Potter</surname> <given-names>M</given-names></name> <name><surname>Pericak-Vance</surname> <given-names>MA</given-names></name> <name><surname>Haines</surname> <given-names>JL</given-names></name></person-group>. <article-title>Genetically meaningful phenotypic subgroups in autism spectrum disorders</article-title>. <source>Genes Brain Behav.</source> (<year>2014</year>) <volume>13</volume>:<fpage>276</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1111/gbb.12117</pub-id><pub-id pub-id-type="pmid">24373520</pub-id></citation></ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Anney</surname> <given-names>R</given-names></name> <name><surname>Klei</surname> <given-names>L</given-names></name> <name><surname>Pinto</surname> <given-names>D</given-names></name> <name><surname>Almeida</surname> <given-names>J</given-names></name> <name><surname>Bacchelli</surname> <given-names>E</given-names></name> <name><surname>Baird</surname> <given-names>G</given-names></name> <etal/></person-group>. <article-title>Individual common variants exert weak effects on the risk for autism spectrum disorders</article-title>. <source>Hum Mol Genet.</source> (<year>2012</year>) <volume>21</volume>:<fpage>4781</fpage>&#x02013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1093/hmg/dds301</pub-id><pub-id pub-id-type="pmid">22843504</pub-id></citation></ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chaste</surname> <given-names>P</given-names></name> <name><surname>Klei</surname> <given-names>L</given-names></name> <name><surname>Sanders</surname> <given-names>SJ</given-names></name> <name><surname>Hus</surname> <given-names>V</given-names></name> <name><surname>Murtha</surname> <given-names>MT</given-names></name> <name><surname>Lowe</surname> <given-names>JK</given-names></name> <etal/></person-group>. <article-title>A genome-wide association study of autism using the Simons simplex collection: does reducing phenotypic heterogeneity in autism increase genetic homogeneity?</article-title> <source>Biol Psychiatry.</source> (<year>2015</year>) <volume>77</volume>:<fpage>775</fpage>&#x02013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1016/j.biopsych.2014.09.017</pub-id><pub-id pub-id-type="pmid">25534755</pub-id></citation></ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Torrico</surname> <given-names>B</given-names></name> <name><surname>Shaw</surname> <given-names>AD</given-names></name> <name><surname>Mosca</surname> <given-names>R</given-names></name> <name><surname>Viv&#x000F3;-Luque</surname> <given-names>N</given-names></name> <name><surname>Herv&#x000E1;s</surname> <given-names>A</given-names></name> <name><surname>Fern&#x000E0;ndez-Castillo</surname> <given-names>N</given-names></name> <etal/></person-group>. <article-title>Truncating variant burden in high-functioning autism and pleiotropic effects of LRP1 across psychiatric phenotypes</article-title>. <source>J Psychiatry Neurosci.</source> (<year>2019</year>) <volume>44</volume>:<fpage>350</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1503/jpn.180184</pub-id><pub-id pub-id-type="pmid">31094488</pub-id></citation></ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chiang</surname> <given-names>AH</given-names></name> <name><surname>Chang</surname> <given-names>J</given-names></name> <name><surname>Wang</surname> <given-names>J</given-names></name> <name><surname>Vitkup</surname> <given-names>D</given-names></name></person-group>. <article-title>Exons as units of phenotypic impact for truncating mutations in autism</article-title>. <source>Mol Psychiatry</source> (<year>2020</year>) <volume>25</volume>:<fpage>1</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1038/s41380-020-00876-3</pub-id><pub-id pub-id-type="pmid">33110259</pub-id></citation></ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bruining</surname> <given-names>H</given-names></name> <name><surname>de Sonneville</surname> <given-names>L</given-names></name> <name><surname>Swaab</surname> <given-names>H</given-names></name> <name><surname>de Jonge</surname> <given-names>M</given-names></name> <name><surname>Kas</surname> <given-names>M</given-names></name> <name><surname>van Engeland</surname> <given-names>H</given-names></name> <etal/></person-group>. <article-title>Dissecting the clinical heterogeneity of autism spectrum disorders through defined genotypes</article-title>. <source>PLoS ONE.</source> (<year>2010</year>) <volume>5</volume>:<fpage>e10887</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0010887</pub-id><pub-id pub-id-type="pmid">20526357</pub-id></citation></ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>VW</given-names></name> <name><surname>Lai</surname> <given-names>Y</given-names></name></person-group>. <article-title>Developing a Predictive Gene Classifier for Autism Spectrum Disorders Based upon Differential Gene Expression Profiles of Phenotypic Subgroups</article-title>. <source>N Am J Med Sci (Boston)</source> (<year>2013</year>) <volume>6</volume>:<fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.7156/najms.2013.0603107</pub-id><pub-id pub-id-type="pmid">24363828</pub-id></citation></ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mottron</surname> <given-names>L</given-names></name> <name><surname>Bzdok</surname> <given-names>D</given-names></name></person-group>. <article-title>Autism spectrum heterogeneity: fact or artifact?</article-title> <source>Mol Psychiatry.</source> (<year>2020</year>) <volume>25</volume>:<fpage>3178</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1038/s41380-020-0748-y</pub-id><pub-id pub-id-type="pmid">32355335</pub-id></citation></ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000FC;pper</surname> <given-names>C</given-names></name> <name><surname>Stroth</surname> <given-names>S</given-names></name> <name><surname>Wolff</surname> <given-names>N</given-names></name> <name><surname>Hauck</surname> <given-names>F</given-names></name> <name><surname>Kliewer</surname> <given-names>N</given-names></name> <name><surname>Schad-Hansjosten</surname> <given-names>T</given-names></name> <etal/></person-group>. <article-title>Identifying predictive features of autism spectrum disorders in a clinical sample of adolescents and adults using machine learning</article-title>. <source>Sci Rep.</source> (<year>2020</year>) <volume>10</volume>:<fpage>4805</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-61607-w</pub-id><pub-id pub-id-type="pmid">32188882</pub-id></citation></ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Asif</surname> <given-names>M</given-names></name> <name><surname>Martiniano</surname> <given-names>HFMC</given-names></name> <name><surname>Marques</surname> <given-names>AR</given-names></name> <name><surname>Santos</surname> <given-names>JX</given-names></name> <name><surname>Vilela</surname> <given-names>J</given-names></name> <name><surname>Rasga</surname> <given-names>C</given-names></name> <etal/></person-group>. <article-title>Identification of biological mechanisms underlying a multidimensional ASD phenotype using machine learning</article-title>. <source>Transl Psychiatry.</source> (<year>2020</year>) <volume>10</volume>:<fpage>43</fpage>. <pub-id pub-id-type="doi">10.1038/s41398-020-0721-1</pub-id><pub-id pub-id-type="pmid">32066720</pub-id></citation></ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Akter</surname> <given-names>T</given-names></name> <name><surname>Shahriare Satu</surname> <given-names>M</given-names></name> <name><surname>Khan</surname> <given-names>MI</given-names></name> <name><surname>Ali</surname> <given-names>MH</given-names></name> <name><surname>Uddin</surname> <given-names>S</given-names></name> <name><surname>Lio</surname> <given-names>P</given-names></name> <etal/></person-group>. <article-title>Machine learning-based models for early stage detection of autism spectrum disorders</article-title>. <source>IEEE Access.</source> (<year>2019</year>) <volume>7</volume>:<fpage>166509</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2952609</pub-id></citation></ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schanding</surname> <given-names>Jr</given-names></name></person-group>. <article-title>GT, Nowell KP, Goin-Kochel RP. Utility of the social communication questionnaire-current and social responsiveness scale as teacher-report screening tools for autism spectrum disorders</article-title>. <source>J Autism Dev Disord.</source> (<year>2012</year>) <volume>42</volume>:<fpage>1705</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1007/s10803-011-1412-9</pub-id></citation></ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mayo</surname> <given-names>J</given-names></name> <name><surname>Chlebowski</surname> <given-names>C</given-names></name> <name><surname>Fein</surname> <given-names>DA</given-names></name> <name><surname>Eigsti</surname> <given-names>IM</given-names></name></person-group>. <article-title>Age of first words predicts cognitive ability and adaptive skills in children with ASD</article-title>. <source>J Autism Dev Disord.</source> (<year>2013</year>) <volume>43</volume>:<fpage>253</fpage>&#x02013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1007/s10803-012-1558-0</pub-id><pub-id pub-id-type="pmid">22673858</pub-id></citation></ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>PI</given-names></name> <name><surname>Kuo</surname> <given-names>PH</given-names></name> <name><surname>Chen</surname> <given-names>CH</given-names></name> <name><surname>Wu</surname> <given-names>JY</given-names></name> <name><surname>Gau</surname> <given-names>SSF</given-names></name> <name><surname>Wu</surname> <given-names>YY</given-names></name> <etal/></person-group>. <article-title>Runs of homozygosity associated with speech delay in autism in a taiwanese Han population: evidence for the recessive model</article-title>. <source>PLoS ONE.</source> (<year>2013</year>) <volume>8</volume>:<fpage>e72056</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0072056</pub-id><pub-id pub-id-type="pmid">23977206</pub-id></citation></ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>PI</given-names></name> <name><surname>Chien</surname> <given-names>YL</given-names></name> <name><surname>Wu</surname> <given-names>YY</given-names></name> <name><surname>Chen</surname> <given-names>CH</given-names></name> <name><surname>Gau</surname> <given-names>SSF</given-names></name> <name><surname>Huang</surname> <given-names>YS</given-names></name> <etal/></person-group>. <article-title>The WNT2 gene polymorphism associated with speech delay inherent to autism</article-title>. <source>Res Dev Disabil.</source> (<year>2012</year>) <volume>33</volume>:<fpage>1533</fpage>&#x02013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1016/j.ridd.2012.03.004</pub-id><pub-id pub-id-type="pmid">22522212</pub-id></citation></ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eicher</surname> <given-names>JD</given-names></name> <name><surname>Gruen</surname> <given-names>JR</given-names></name></person-group>. <article-title>Language impairment and dyslexia genes influence language skills in children with autism spectrum disorders</article-title>. <source>Autism Res.</source> (<year>2015</year>) <volume>8</volume>:<fpage>229</fpage>&#x02013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1002/aur.1436</pub-id><pub-id pub-id-type="pmid">25448322</pub-id></citation></ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lord</surname> <given-names>C</given-names></name> <name><surname>Rutter</surname> <given-names>M</given-names></name> <name><surname>Le Couteur</surname> <given-names>A</given-names></name></person-group>. <article-title>Autism Diagnostic Interview-Revised: a revised version of a diagnostic interview for caregivers of individuals with possible pervasive developmental disorders</article-title>. <source>J Autism Dev Disord.</source> (<year>1994</year>) <volume>24</volume>:<fpage>659</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1007/BF02172145</pub-id><pub-id pub-id-type="pmid">7814313</pub-id></citation></ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gau</surname> <given-names>SSF</given-names></name> <name><surname>Lee</surname> <given-names>CM</given-names></name> <name><surname>Lai</surname> <given-names>MC</given-names></name> <name><surname>Chiu</surname> <given-names>YN</given-names></name> <name><surname>Huang</surname> <given-names>YF</given-names></name> <name><surname>Kao</surname> <given-names>J Der</given-names></name> <etal/></person-group>. <article-title>Psychometric properties of the Chinese version of the social communication questionnaire</article-title>. <source>Res Autism Spectr Disord.</source> (<year>2011</year>) <volume>5</volume>:<fpage>809</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/j.rasd.2010.09.010</pub-id></citation></ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>CH</given-names></name></person-group>. <article-title>Generalized association plots: information visualization via iteratively generated correlation matrices</article-title>. <source>Stat Sin.</source> (<year>2002</year>) <volume>12</volume>:<fpage>7</fpage>&#x02013;<lpage>29</lpage>.</citation></ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>HM</given-names></name> <name><surname>Tien</surname> <given-names>YJ</given-names></name> <name><surname>Chen</surname> <given-names>CH</given-names></name></person-group>. <article-title>GAP: a graphical environment for matrix visualization and cluster analysis</article-title>. <source>Comput Stat Data Anal.</source> (<year>2010</year>) <volume>54</volume>:<fpage>767</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1016/j.csda.2008.09.029</pub-id></citation></ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Irizarry</surname> <given-names>RA</given-names></name> <name><surname>Bolstad</surname> <given-names>BM</given-names></name> <name><surname>Collin</surname> <given-names>F</given-names></name> <name><surname>Cope</surname> <given-names>LM</given-names></name> <name><surname>Hobbs</surname> <given-names>B</given-names></name> <name><surname>Speed</surname> <given-names>TP</given-names></name></person-group>. <article-title>Summaries of Affymetrix GeneChip probe level data</article-title>. <source>Nucleic Acids Res.</source> (<year>2003</year>) <volume>31</volume>:<fpage>e15</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gng015</pub-id><pub-id pub-id-type="pmid">12582260</pub-id></citation></ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kanehisa</surname> <given-names>M</given-names></name> <name><surname>Goto</surname> <given-names>S</given-names></name></person-group>. <article-title>KEGG: Kyoto encyclopedia of genes and genomes</article-title>. <source>Nucleic Acids Res.</source> (<year>2000</year>) <volume>28</volume>:<fpage>27</fpage>&#x02013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.27</pub-id><pub-id pub-id-type="pmid">25811933</pub-id></citation></ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Slenter</surname> <given-names>DN</given-names></name> <name><surname>Kutmon</surname> <given-names>M</given-names></name> <name><surname>Hanspers</surname> <given-names>K</given-names></name> <name><surname>Riutta</surname> <given-names>A</given-names></name> <name><surname>Windsor</surname> <given-names>J</given-names></name> <name><surname>Nunes</surname> <given-names>N</given-names></name> <etal/></person-group>. <article-title>WikiPathways: a multifaceted pathway database bridging metabolomics to other omics research</article-title>. <source>Nucleic Acids Res.</source> (<year>2018</year>) <volume>46</volume>:<fpage>D661</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkx1064</pub-id><pub-id pub-id-type="pmid">29136241</pub-id></citation></ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nishimura</surname> <given-names>D</given-names></name></person-group>. <article-title>BioCarta</article-title>. <source>Biotech Softw Internet Rep.</source> (<year>2001</year>) <fpage>117</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1089/152791601750294344</pub-id></citation></ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fabregat</surname> <given-names>A</given-names></name> <name><surname>Jupe</surname> <given-names>S</given-names></name> <name><surname>Matthews</surname> <given-names>L</given-names></name> <name><surname>Sidiropoulos</surname> <given-names>K</given-names></name> <name><surname>Gillespie</surname> <given-names>M</given-names></name> <name><surname>Garapati</surname> <given-names>P</given-names></name> <etal/></person-group>. <article-title>The reactome pathway knowledgebase</article-title>. <source>Nucleic Acids Res.</source> (<year>2018</year>) <volume>44</volume>:<fpage>D481</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkx1132</pub-id></citation></ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carbon</surname> <given-names>S</given-names></name> <name><surname>Douglass</surname> <given-names>E</given-names></name> <name><surname>Dunn</surname> <given-names>N</given-names></name> <name><surname>Good</surname> <given-names>B</given-names></name> <name><surname>Harris</surname> <given-names>NL</given-names></name> <name><surname>Lewis</surname> <given-names>SE</given-names></name> <etal/></person-group>. <article-title>The gene ontology resource: 20 years and still GOing strong</article-title>. <source>Nucleic Acids Res.</source> (<year>2019</year>) <volume>47</volume>:<fpage>D330</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1055</pub-id><pub-id pub-id-type="pmid">30395331</pub-id></citation></ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kamburov</surname> <given-names>A</given-names></name> <name><surname>Stelzl</surname> <given-names>U</given-names></name> <name><surname>Lehrach</surname> <given-names>H</given-names></name> <name><surname>Herwig</surname> <given-names>R</given-names></name></person-group>. <article-title>The ConsensusPathDB interaction database: 2013 Update</article-title>. <source>Nucleic Acids Res.</source> (<year>2013</year>) <volume>41</volume>:<fpage>D793</fpage>&#x02013;<lpage>800</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks1055</pub-id><pub-id pub-id-type="pmid">23143270</pub-id></citation></ref>
<ref id="B38">
<label>38.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>T</given-names></name> <name><surname>Horvath</surname> <given-names>S</given-names></name></person-group>. <article-title>Unsupervised learning with random forest predictors</article-title>. <source>J Comput Graph Stat.</source> (<year>2006</year>) <volume>15</volume>:<fpage>118</fpage>&#x02013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1198/106186006X94072</pub-id></citation></ref>
<ref id="B39">
<label>39.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kaufman</surname> <given-names>L</given-names></name> <name><surname>Rousseeuw</surname> <given-names>PJ</given-names></name></person-group>. <source>Partitioning Around Medoids (Program PAM), in Finding Groups in Data: An Introduction to Cluster Analysis</source>. <publisher-loc>Hoboken</publisher-loc>: <publisher-name>Wiley</publisher-name> (<year>2008</year>). <pub-id pub-id-type="doi">10.1002/9780470316801.ch2</pub-id></citation></ref>
<ref id="B40">
<label>40.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cortes</surname> <given-names>C</given-names></name> <name><surname>Vapnik</surname> <given-names>V</given-names></name></person-group>. <article-title>Support-vector networks</article-title>. <source>Mach Learn.</source> (<year>1995</year>) <volume>20</volume>:<fpage>273</fpage>&#x02013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1023/A:1022627411411</pub-id></citation></ref>
<ref id="B41">
<label>41.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Devi Arockia Vanitha</surname> <given-names>C</given-names></name> <name><surname>Devaraj</surname> <given-names>D</given-names></name> <name><surname>Venkatesulu</surname> <given-names>M</given-names></name></person-group>. <article-title>Gene expression data classification using Support Vector Machine and mutual information-based gene selection</article-title>. <source>Procedia Comput Sci.</source> (<year>2014</year>) <volume>47</volume>:<fpage>13</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2015.03.178</pub-id><pub-id pub-id-type="pmid">31209677</pub-id></citation></ref>
<ref id="B42">
<label>42.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soleymani</surname> <given-names>A</given-names></name> <name><surname>Pennekamp</surname> <given-names>F</given-names></name> <name><surname>Petchey</surname> <given-names>OL</given-names></name> <name><surname>Weibel</surname> <given-names>R</given-names></name></person-group>. <article-title>Developing and integrating advanced movement features improves automated classification of ciliate species</article-title>. <source>PLoS ONE.</source> (<year>2015</year>) <volume>11</volume>:<fpage>e0145345</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0145345</pub-id><pub-id pub-id-type="pmid">26824617</pub-id></citation></ref>
<ref id="B43">
<label>43.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verda</surname> <given-names>D</given-names></name> <name><surname>Parodi</surname> <given-names>S</given-names></name> <name><surname>Ferrari</surname> <given-names>E</given-names></name> <name><surname>Muselli</surname> <given-names>M</given-names></name></person-group>. <article-title>Analyzing gene expression data for pediatric and adult cancer diagnosis using logic learning machine and standard supervised methods</article-title>. <source>BMC Bioinformatics.</source> (<year>2019</year>) <volume>20</volume>:<fpage>390</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-019-2953-8</pub-id><pub-id pub-id-type="pmid">31757200</pub-id></citation></ref>
<ref id="B44">
<label>44.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Delgado</surname> <given-names>R</given-names></name> <name><surname>Tibau</surname> <given-names>XA</given-names></name></person-group>. <article-title>Why Cohen&#x00027;s Kappa should be avoided as performance measure in classification</article-title>. <source>PLoS ONE.</source> (<year>2019</year>) <volume>14</volume>:<fpage>e0222916</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0222916</pub-id><pub-id pub-id-type="pmid">31557204</pub-id></citation></ref>
<ref id="B45">
<label>45.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>M</given-names></name></person-group>. <article-title>caret Package</article-title>. <source>J Stat Softw.</source> (<year>2008</year>) <volume>28</volume>:<fpage>1</fpage>&#x02013;<lpage>26</lpage>.</citation></ref>
<ref id="B46">
<label>46.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roussos</surname> <given-names>P</given-names></name> <name><surname>Guennewig</surname> <given-names>B</given-names></name> <name><surname>Kaczorowski</surname> <given-names>DC</given-names></name> <name><surname>Barry</surname> <given-names>G</given-names></name> <name><surname>Brennand</surname> <given-names>KJ</given-names></name></person-group>. <article-title>Activity-dependent changes in gene expression in schizophrenia human-induced pluripotent stem cell neurons</article-title>. <source>JAMA Psychiatry.</source> (<year>2016</year>) <volume>73</volume>:<fpage>1180</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1001/jamapsychiatry.2016.2575</pub-id><pub-id pub-id-type="pmid">27732689</pub-id></citation></ref>
<ref id="B47">
<label>47.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Azuma</surname> <given-names>M</given-names></name> <name><surname>Toyama</surname> <given-names>R</given-names></name> <name><surname>Laver</surname> <given-names>E</given-names></name> <name><surname>Dawid</surname> <given-names>IB</given-names></name></person-group>. <article-title>Perturbation of rRNA synthesis in the bap28 mutation leads to apoptosis mediated by p53 in the zebrafish central nervous system</article-title>. <source>J Biol Chem.</source> (<year>2006</year>) <volume>281</volume>:<fpage>13309</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1074/jbc.M601892200</pub-id><pub-id pub-id-type="pmid">16531401</pub-id></citation></ref>
<ref id="B48">
<label>48.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sato</surname> <given-names>R</given-names></name></person-group>. <article-title>Sterol metabolism and SREBP activation</article-title>. <source>Arch Biochem Biophys.</source> (<year>2010</year>) <volume>501</volume>:<fpage>177</fpage>&#x02013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1016/j.abb.2010.06.004</pub-id></citation></ref>
<ref id="B49">
<label>49.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paul</surname> <given-names>SM</given-names></name> <name><surname>Doherty</surname> <given-names>JJ</given-names></name> <name><surname>Robichaud</surname> <given-names>AJ</given-names></name> <name><surname>Belfort</surname> <given-names>GM</given-names></name> <name><surname>Chow</surname> <given-names>BY</given-names></name> <name><surname>Hammond</surname> <given-names>RS</given-names></name> <etal/></person-group>. <article-title>The major brain cholesterol metabolite 24(S)-hydroxycholesterol is a potent allosteric modulator of N-Methyl-D-Aspartate receptors</article-title>. <source>J Neurosci.</source> (<year>2013</year>) <volume>33</volume>:<fpage>17290</fpage>&#x02013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2619-13.2013</pub-id><pub-id pub-id-type="pmid">24174662</pub-id></citation></ref>
<ref id="B50">
<label>50.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>H</given-names></name></person-group>. <article-title>Lipid rafts: a signaling platform linking cholesterol metabolism to synaptic deficits in autism spectrum disorders</article-title>. <source>Front Behav Neurosci.</source> (<year>2014</year>) <volume>8</volume>:<fpage>104</fpage>. <pub-id pub-id-type="doi">10.3389/fnbeh.2014.00104</pub-id><pub-id pub-id-type="pmid">24723866</pub-id></citation></ref>
<ref id="B51">
<label>51.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Petrov</surname> <given-names>AM</given-names></name> <name><surname>Kasimov</surname> <given-names>MR</given-names></name> <name><surname>Zefirov</surname> <given-names>AL</given-names></name></person-group>. <article-title>Cholesterol in the pathogenesis of alzheimer&#x00027;s, parkinson&#x00027;s diseases and autism: link to synaptic dysfunction</article-title>. <source>Acta Naturae.</source> (<year>2017</year>) <volume>9</volume>:<fpage>26</fpage>&#x02013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.32607/20758251-2017-9-1-26-37</pub-id><pub-id pub-id-type="pmid">28461971</pub-id></citation></ref>
<ref id="B52">
<label>52.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tamiji</surname> <given-names>J</given-names></name> <name><surname>Crawford</surname> <given-names>DA</given-names></name></person-group>. <article-title>The neurobiology of lipid metabolism in autism spectrum disorders</article-title>. <source>NeuroSignals.</source> (<year>2011</year>) <volume>18</volume>:<fpage>98</fpage>&#x02013;<lpage>112</lpage>. <pub-id pub-id-type="doi">10.1159/000323189</pub-id><pub-id pub-id-type="pmid">21346377</pub-id></citation></ref>
<ref id="B53">
<label>53.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gillberg</surname> <given-names>C</given-names></name> <name><surname>Fernell</surname> <given-names>E</given-names></name> <name><surname>Ko&#x0010D;ovsk&#x000E1;</surname> <given-names>E</given-names></name> <name><surname>Minnis</surname> <given-names>H</given-names></name> <name><surname>Bourgeron</surname> <given-names>T</given-names></name> <name><surname>Thompson</surname> <given-names>L</given-names></name> <etal/></person-group>. <article-title>The role of cholesterol metabolism and various steroid abnormalities in autism spectrum disorders: a hypothesis paper</article-title>. <source>Autism Res.</source> (<year>2017</year>) <volume>10</volume>:<fpage>1022</fpage>&#x02013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1002/aur.1777</pub-id><pub-id pub-id-type="pmid">28401679</pub-id></citation></ref>
<ref id="B54">
<label>54.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Richardson</surname> <given-names>AJ</given-names></name> <name><surname>Ross</surname> <given-names>MA</given-names></name></person-group>. <article-title>Fatty acid metabolism in neurodevelopmental disorder: a new perspective on associations between attention-deficit/hyperactivity disorder, dyslexia, dyspraxia and the autistic spectrum</article-title>. <source>Prostaglandins Leukot Essent Fat Acids.</source> (<year>2000</year>) <volume>63</volume>:<fpage>1</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1054/plef.2000.0184</pub-id><pub-id pub-id-type="pmid">10970706</pub-id></citation></ref>
<ref id="B55">
<label>55.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aneja</surname> <given-names>A</given-names></name> <name><surname>Tierney</surname> <given-names>E</given-names></name></person-group>. <article-title>Autism: the role of cholesterol in treatment</article-title>. <source>Int Rev Psychiatry.</source> (<year>2008</year>) <volume>20</volume>:<fpage>165</fpage>&#x02013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.1080/09540260801889062</pub-id></citation></ref>
<ref id="B56">
<label>56.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cartocci</surname> <given-names>V</given-names></name> <name><surname>Catallo</surname> <given-names>M</given-names></name> <name><surname>Tempestilli</surname> <given-names>M</given-names></name> <name><surname>Segatto</surname> <given-names>M</given-names></name> <name><surname>Pfrieger</surname> <given-names>FW</given-names></name> <name><surname>Bronzuoli</surname> <given-names>MR</given-names></name> <etal/></person-group>. <article-title>Altered brain cholesterol/isoprenoid metabolism in a rat model of autism spectrum disorders</article-title>. <source>Neuroscience.</source> (<year>2018</year>) <volume>372</volume>:<fpage>27</fpage>&#x02013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroscience.2017.12.053</pub-id><pub-id pub-id-type="pmid">29309878</pub-id></citation></ref>
<ref id="B57">
<label>57.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Esparham</surname> <given-names>AE</given-names></name> <name><surname>Smith</surname> <given-names>T</given-names></name> <name><surname>Belmont</surname> <given-names>JM</given-names></name> <name><surname>Haden</surname> <given-names>M</given-names></name> <name><surname>Wagner</surname> <given-names>LE</given-names></name> <name><surname>Evans</surname> <given-names>RG</given-names></name> <etal/></person-group>. <article-title>Nutritional and metabolic biomarkers in autism spectrum disorders: an exploratory study</article-title>. <source>Integr Med.</source> (<year>2015</year>) <volume>14</volume>:<fpage>40</fpage>&#x02013;<lpage>53</lpage>.<pub-id pub-id-type="pmid">26770138</pub-id></citation></ref>
<ref id="B58">
<label>58.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tierney</surname> <given-names>E</given-names></name> <name><surname>Bukelis</surname> <given-names>I</given-names></name> <name><surname>Thompson</surname> <given-names>RE</given-names></name> <name><surname>Ahmed</surname> <given-names>K</given-names></name> <name><surname>Aneja</surname> <given-names>A</given-names></name> <name><surname>Kratz</surname> <given-names>L</given-names></name> <etal/></person-group>. <article-title>Abnormalities of cholesterol metabolism in autism spectrum disorders</article-title>. <source>Am J Med Genet Part B Neuropsychiatr Genet.</source> (<year>2006</year>) <volume>141B</volume>:<fpage>666</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1002/ajmg.b.30368</pub-id><pub-id pub-id-type="pmid">28401679</pub-id></citation></ref>
<ref id="B59">
<label>59.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>Y</given-names></name> <name><surname>Eran</surname> <given-names>A</given-names></name> <name><surname>Palmer</surname> <given-names>N</given-names></name> <name><surname>Avillach</surname> <given-names>P</given-names></name> <name><surname>Levy-Moonshine</surname> <given-names>A</given-names></name> <name><surname>Szolovits</surname> <given-names>P</given-names></name> <etal/></person-group>. <article-title>A multidimensional precision medicine approach identifies an autism subtype characterized by dyslipidemia</article-title>. <source>Nat Med.</source> (<year>2020</year>) <volume>26</volume>:<fpage>1375</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1038/s41591-020-1007-0</pub-id><pub-id pub-id-type="pmid">32778826</pub-id></citation></ref>
<ref id="B60">
<label>60.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L</given-names></name></person-group>. <article-title>Random forests</article-title>. <source>Mach Learn.</source> (<year>2001</year>) <volume>45</volume>:<fpage>5</fpage>&#x02013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/A:1010933404324</pub-id></citation></ref>
<ref id="B61">
<label>61.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>D&#x000ED;az-Uriarte</surname> <given-names>R</given-names></name> <name><surname>Alvarez de Andr&#x000E9;s</surname> <given-names>S</given-names></name></person-group>. <article-title>Gene selection and classification of microarray data using random forest</article-title>. <source>BMC Bioinformatics.</source> (<year>2006</year>) <volume>7</volume>:<fpage>3</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-7-3</pub-id><pub-id pub-id-type="pmid">22125385</pub-id></citation></ref>
<ref id="B62">
<label>62.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>SY</given-names></name></person-group>. <article-title>Effects of sample size on robustness and prediction accuracy of a prognostic gene signature</article-title>. <source>BMC Bioinformatics.</source> (<year>2009</year>) <volume>10</volume>:<fpage>147</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-10-147</pub-id><pub-id pub-id-type="pmid">19445687</pub-id></citation></ref>
<ref id="B63">
<label>63.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brown</surname> <given-names>MPS</given-names></name> <name><surname>Grundy</surname> <given-names>WN</given-names></name> <name><surname>Lin</surname> <given-names>D</given-names></name> <name><surname>Cristianini</surname> <given-names>N</given-names></name> <name><surname>Sugnet</surname> <given-names>CW</given-names></name> <name><surname>Furey</surname> <given-names>TS</given-names></name> <etal/></person-group>. <article-title>Knowledge-based analysis of microarray gene expression data by using support vector machines</article-title>. <source>Proc Natl Acad Sci USA.</source> (<year>2000</year>) <volume>97</volume>:<fpage>262</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.97.1.262</pub-id><pub-id pub-id-type="pmid">10618406</pub-id></citation></ref>
<ref id="B64">
<label>64.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Statnikov</surname> <given-names>A</given-names></name> <name><surname>Wang</surname> <given-names>L</given-names></name> <name><surname>Aliferis</surname> <given-names>CF</given-names></name></person-group>. <article-title>A comprehensive comparison of random forests and support vector machines for microarray-based cancer classification</article-title>. <source>BMC Bioinformatics.</source> (<year>2008</year>) <volume>9</volume>:<fpage>319</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-9-319</pub-id><pub-id pub-id-type="pmid">18647401</pub-id></citation></ref>
<ref id="B65">
<label>65.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Diepen</surname> <given-names>M</given-names></name> <name><surname>Ramspek</surname> <given-names>CL</given-names></name> <name><surname>Jager</surname> <given-names>KJ</given-names></name> <name><surname>Zoccali</surname> <given-names>C</given-names></name> <name><surname>Dekker</surname> <given-names>FW</given-names></name></person-group>. <article-title>Prediction versus aetiology: common pitfalls and how to avoid them</article-title>. <source>Nephrol Dial Transplant.</source> (<year>2017</year>) <volume>32</volume>:<fpage>ii1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1093/ndt/gfw459</pub-id><pub-id pub-id-type="pmid">28339854</pub-id></citation></ref>
<ref id="B66">
<label>66.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moon</surname> <given-names>SJ</given-names></name> <name><surname>Hwang</surname> <given-names>J</given-names></name> <name><surname>Kana</surname> <given-names>R</given-names></name> <name><surname>Torous</surname> <given-names>J</given-names></name> <name><surname>Kim</surname> <given-names>JW</given-names></name></person-group>. <article-title>Accuracy of machine learning algorithms for the diagnosis of autism spectrum disorder: systematic review and meta-analysis of brain magnetic resonance imaging studies</article-title>. <source>J Med Internet Res</source>. (<year>2019</year>) <volume>6</volume>:<fpage>e14108</fpage>. <pub-id pub-id-type="doi">10.2196/14108</pub-id><pub-id pub-id-type="pmid">31562756</pub-id></citation></ref>
<ref id="B67">
<label>67.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stevens</surname> <given-names>E</given-names></name> <name><surname>Dixon</surname> <given-names>DR</given-names></name> <name><surname>Novack</surname> <given-names>MN</given-names></name> <name><surname>Granpeesheh</surname> <given-names>D</given-names></name> <name><surname>Smith</surname> <given-names>T</given-names></name> <name><surname>Linstead</surname> <given-names>E</given-names></name></person-group>. <article-title>Identification and analysis of behavioral phenotypes in autism spectrum disorder via unsupervised machine learning</article-title>. <source>Int J Med Inform.</source> (<year>2019</year>) <volume>129</volume>:<fpage>29</fpage>&#x02013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijmedinf.2019.05.006</pub-id><pub-id pub-id-type="pmid">31445269</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> The genetic data collection and analysis were supported by grants from the Ministry of Science and Technology (NSC 99-3112-B-002-036), Taiwan, and National Taiwan University Hospital (NCTRC201114), Taiwan, awarded to SG.</p>
</fn>
</fn-group>
</back>
</article>