<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2022.872848</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Estimating High-Order Brain Functional Networks in Bayesian View for Autism Spectrum Disorder Identification</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Jiang</surname> <given-names>Xiao</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x2020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1672323/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhou</surname> <given-names>Yueying</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x2020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/503417/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Yining</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Limei</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Qiao</surname> <given-names>Lishan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>De Leone</surname> <given-names>Renato</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x002A;</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Mathematics Science, Liaocheng University</institution>, <addr-line>Liaocheng</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Science and Technology, University of Camerino</institution>, <addr-line>Camerino</addr-line>, <country>Italy</country></aff>
<aff id="aff3"><sup>3</sup><institution>College of Computer Science and Technology, Nanjing University of Aeronautics</institution>, <addr-line>Nanjing</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>School of Computer Science and Technology, Shandong Jianzhu University</institution>, <addr-line>Jinan</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Zhengxia Wang, Hainan University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Benzheng Wei, Shandong University of Traditional Chinese Medicine, China; Linling Li, Shenzhen University, China; Jun Wang, Shanghai University, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Lishan Qiao, <email>qiaolishan@lcu.edu.cn</email></corresp>
<corresp id="c002">Renato De Leone, <email>renato.deleone@unicam.it</email></corresp>
<fn fn-type="equal" id="fn002"><p><sup>&#x2020;</sup>These authors have contributed equally to this work and share first authorship</p></fn>
<fn fn-type="other" id="fn004"><p>This article was submitted to Brain Imaging Methods, a section of the journal Frontiers in Neuroscience</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>872848</elocation-id>
<history>
<date date-type="received">
<day>10</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Jiang, Zhou, Zhang, Zhang, Qiao and De Leone.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Jiang, Zhou, Zhang, Zhang, Qiao and De Leone</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Brain functional network (BFN) has become an increasingly important tool to understand the inherent organization of the brain and explore informative biomarkers of neurological disorders. Pearson&#x2019;s correlation (PC) is the most widely accepted method for constructing BFNs and provides a basis for designing new BFN estimation schemes. Particularly, a recent study proposes to use two sequential PC operations, namely, correlation&#x2019;s correlation (CC), for constructing the high-order BFN. Despite its empirical effectiveness in identifying neurological disorders and detecting subtle changes of connections in different subject groups, CC is defined intuitively without a solid and sustainable theoretical foundation. For understanding CC more rigorously and providing a systematic BFN learning framework, in this paper, we reformulate it in the Bayesian view with a prior of matrix-variate normal distribution. As a result, we obtain a probabilistic explanation of CC. In addition, we develop a Bayesian high-order method (BHM) to automatically and simultaneously estimate the high- and low-order BFN based on the probabilistic framework. An efficient optimization algorithm is also proposed. Finally, we evaluate BHM in identifying subjects with autism spectrum disorder (ASD) from typical controls based on the estimated BFNs. Experimental results suggest that the automatically learned high- and low-order BFNs yield a superior performance over the artificially defined BFNs <italic>via</italic> conventional CC and PC.</p>
</abstract>
<kwd-group>
<kwd>brain functional network</kwd>
<kwd>high-order network</kwd>
<kwd>Pearson&#x2019;s correlation</kwd>
<kwd>Bayesian statistics</kwd>
<kwd>matrix-variate normal distribution</kwd>
<kwd>autism spectrum disorder</kwd>
</kwd-group>
<contract-num rid="cn001">61976110</contract-num>
<contract-num rid="cn001">62176112</contract-num>
<contract-num rid="cn001">11931008</contract-num>
<contract-num rid="cn002">ZR202102270451</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<contract-sponsor id="cn002">Natural Science Foundation of Shandong Province<named-content content-type="fundref-id">10.13039/501100007129</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="3"/>
<equation-count count="44"/>
<ref-count count="39"/>
<page-count count="13"/>
<word-count count="9288"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>Resting-state functional magnetic resonance imaging (rs-fMRI)-based brain functional network (BFN) analysis without a specific task, has shown a great potential to discover biomarkers for identifying neurological/mental disorders, such as autism spectrum disorder (ASD) (<xref ref-type="bibr" rid="B31">Wang et al., 2019</xref>), major depressive disorder (MDD) (<xref ref-type="bibr" rid="B22">Long et al., 2020</xref>), schizophrenia (<xref ref-type="bibr" rid="B1">Ariana and Cohen, 2013</xref>), Parkinson&#x2019;s disease (PD) (<xref ref-type="bibr" rid="B2">Baggio et al., 2014</xref>), Alzheimer&#x2019;s disease (AD) (<xref ref-type="bibr" rid="B15">Hahn et al., 2013</xref>), and its early stage, namely, mild cognitive impairment (MCI) (<xref ref-type="bibr" rid="B17">Jiang et al., 2019</xref>). However, the identification of brain disorders based on the BFN remains a critical challenge, since its great performance depends on multiple interactive factors including reasonable brain parcellation, well-parametrized network estimation, discriminative feature selection/extraction, and powerful classifier design (<xref ref-type="bibr" rid="B7">Dadi et al., 2019</xref>; <xref ref-type="bibr" rid="B28">Pervaiz et al., 2020</xref>). Instead of considering all these aspects that have been empirically evaluated in recent studies (<xref ref-type="bibr" rid="B7">Dadi et al., 2019</xref>; <xref ref-type="bibr" rid="B28">Pervaiz et al., 2020</xref>), in this paper, we mainly focus on the BFN estimation issue. Recently, more advanced studies (<xref ref-type="bibr" rid="B33">Yu et al., 2017</xref>; <xref ref-type="bibr" rid="B37">Zhang et al., 2017</xref>; <xref ref-type="bibr" rid="B23">Mahjoub et al., 2018</xref>) have proposed the brain functional connectivity representations for estimating BFN at different connectivity levels, including low-order, high-order, etc. Low-order methods are designed to characterize the synchronization of blood oxygen level dependent (BOLD) signals and are insufficient to characterize a high level of interaction. Recent literature (<xref ref-type="bibr" rid="B6">Chen et al., 2016</xref>; <xref ref-type="bibr" rid="B35">Zhang et al., 2016</xref>) presented high-order methods to measure the relationship between the BFN connectivity. This paper specifically aims to capture the brain connectivity that is supposed to exist in a higher-order form.</p>
<p>Owing to its non-invasiveness and easy reproducibility, rs-fMRI (<xref ref-type="bibr" rid="B32">Wang et al., 2020</xref>) has become a widely used technique to estimate BFN whose nodes correspond to spatial regions of interest (ROIs) and edges describe the relationship (e.g., similarity, correlation, synchronism, etc.) between the rs-fMRI signals associated with these ROIs. In the past decades, researchers have developed many BFN estimation methods, including Pearson&#x2019;s correlation (PC) (<xref ref-type="bibr" rid="B4">Biswal et al., 1995</xref>; <xref ref-type="bibr" rid="B10">Eguiluz et al., 2005</xref>), partial correlation (<xref ref-type="bibr" rid="B24">Marrelec et al., 2006</xref>), regularized full/partial correlation (<xref ref-type="bibr" rid="B11">Friedman et al., 2008</xref>; <xref ref-type="bibr" rid="B18">Jie et al., 2009</xref>; <xref ref-type="bibr" rid="B21">Li et al., 2017</xref>), structural equation modeling (<xref ref-type="bibr" rid="B25">Mclntosh and Gonzalez-Lima, 1994</xref>), and dynamic causal modeling (<xref ref-type="bibr" rid="B12">Friston et al., 2003</xref>), etc. According to a recent comparative study (<xref ref-type="bibr" rid="B30">Smith et al., 2013</xref>), the correlation-based approaches are &#x201C;quite successful&#x201D; for estimating informative BFNs. Particularly, PC is the fundamental and most widely used correlation-based method for BFN estimation. Despite its empirical effectiveness, PC only considers a pair of ROIs at a time, and thus suffers the confounding effect from other ROIs. Partial correlation can tackle this problem by regressing out the confounding variables. However, that may lead to an ill-posed estimation since the partial correlation is usually calculated by inverting a covariance matrix that may be singular. In practice, a regularizer is generally introduced into the partial correlation model, which not only deals with the ill-posed problem but also provides a natural way to introduce topological priors of the brain network into the estimation models. Specifically, L<sub>1</sub>-norm is commonly used to encode the sparsity prior of the BFN (<xref ref-type="bibr" rid="B20">Lee et al., 2011</xref>), a weighted version of the L<sub>1</sub>-norm to capture the hub structure (prior) (<xref ref-type="bibr" rid="B21">Li et al., 2017</xref>), the L<sub>2,1</sub>-norm to model group sparsity (or population prior) that imposes all the subjects share the same BFN topology (<xref ref-type="bibr" rid="B34">Zhang, 2010</xref>), and a combination of L<sub>1</sub>-norm with trace-norm to encode the modularity (prior) of the BFN (<xref ref-type="bibr" rid="B29">Qiao et al., 2016</xref>), to just name a few.</p>
<p>No matter which prior or regularizer is introduced, most of the correlation-based methods only estimate low-order BFNs whose edges are the full or partial correlation of the rs-fMRI time series. Beyond these traditional low-order correlations, researchers found that some forms of high-order correlations may contain useful feature information for BFN analysis and classification (<xref ref-type="bibr" rid="B6">Chen et al., 2016</xref>; <xref ref-type="bibr" rid="B35">Zhang et al., 2016</xref>; <xref ref-type="bibr" rid="B38">Zhou et al., 2018a</xref>,<xref ref-type="bibr" rid="B39">b</xref>). For example, <xref ref-type="bibr" rid="B6">Chen et al. (2016)</xref> defined the high-order correlation as the dependency between functional connectivity fluctuation, with clustered mean correlation time series as input. Different from characterizing a temporal correlation, <xref ref-type="bibr" rid="B35">Zhang et al. (2016)</xref> proposed to construct the high-order BFN to examine spatial properties of the functional connectivity network. Specifically, such a scheme is achieved by two sequential PC operations, where the first PC operation is used to construct a traditional low-order BFN, and the ensuing PC operation is conducted on the edge weights of the estimated BFN to generate the high-order BFN. Despite encoding the network information from different dimensions, the above methods are uniformly called correlation&#x2019;s correlation (CC) (<xref ref-type="bibr" rid="B37">Zhang et al., 2017</xref>) since they both involve two PC operations in the high-order BFN construction. However, the CC-based high-order BFNs are estimated intuitively and heuristically without the support of any strong theoretical basis.</p>
<p>Toward a better understanding of CC, in this paper, we reformulate it in the Bayesian framework with a prior that the low-order BFN follows the matrix-variate normal (MVN) distribution. As a result, we obtain a probabilistic explanation for CC and develop a new method that both learns low- and high-order BFNs from data based on the rigorous theoretical framework. In brief, we summarize the main contributions of this paper as follows.</p>
<list list-type="simple">
<list-item>
<label>1.</label>
<p>We reformulate PC from a statistical point of view. Based on this, a regularized statistical framework is derived by introducing Gaussian distribution to the error term, which provides a more flexible modeling idea.</p>
</list-item>
<list-item>
<label>2.</label>
<p>A mathematical model for a high-order learning method based on CC is developed by assuming the adjacency matrix of low-order brain networks follows an aprior normal distribution.</p>
</list-item>
<list-item>
<label>3.</label>
<p>Based on the probabilistic framework derived above, an automatic learning model, namely, BHM, is proposed. Compared with the traditional high-order network learning method (i.e., CC), the model simultaneously learns low-order and high-order brain networks. In the learning process, the direct information of the low-order network and the indirect information of the high-order network complement each other toward more reliable/discriminative brain networks.</p>
</list-item>
<list-item>
<label>4.</label>
<p>Finally, we empirically verify that the automatically learned BFNs outperform the artificially defined ones <italic>via</italic> CC and other baselines in the identification of ASD, even with a simple feature selection method and classifier.</p>
</list-item>
</list>
<p>For a consistent expression throughout the paper, we first describe the basic notations as follows. Scalars involving the variables, parameters, and constants are denoted by <italic>italic</italic> lowercase letters, e.g., <italic>x</italic>. Vectors are denoted by bold lowercase letters and the elements inside are stored in a column, e.g., <bold>x</bold> = (<italic>x</italic><sub>1</sub>, <italic>x</italic><sub>2</sub>, &#x22EF;, <italic>x</italic><sub><italic>n</italic></sub>)<sup><italic>T</italic></sup>. Matrices are denoted by bold uppercase letters such as <bold>X</bold>.</p>
<p>The rest of the paper is organized as follows. In Section &#x201C;Related Works,&#x201D; we review the related works including PC, sparse representation (SR), and CC. In Section &#x201C;High-Order Correlation Learning,&#x201D; we first introduce a theoretical framework for explaining CC and then develop a new framework for learning high-order BFN by reformatting CC in a view of the maximal posterior probability. In Section &#x201C;Experiments and Results,&#x201D; we conduct experiments to evaluate the discrimination of the automatically learned high-order BFNs. In Section &#x201C;Discussions,&#x201D; we discuss the main findings. Finally, the conclusion are reported in Section &#x201C;Conclusion.&#x201D;</p>
</sec>
<sec id="S2">
<title>Related Works</title>
<p>In this section, we review three related works: PC, SR, and CC. As discussed previously, PC and SR are used to construct the traditional <bold>low-order</bold> BFNs, while CC as a two-step sequential PC method is used to estimate the <bold>high-order</bold> BFNs.</p>
<sec id="S2.SS1">
<title>Pearson&#x2019;s Correlation</title>
<p>Suppose <bold>x</bold><sub><italic>i</italic></sub> is the multivariate random variable (random vector) associated with the <italic>i<sup>th</sup></italic> ROI. Then, the observed rs-fMRI signals<sup><xref ref-type="fn" rid="footnote1">1</xref></sup> <bold>x</bold><sub><italic>i</italic></sub> = (<italic>x</italic><sub>1<italic>i</italic></sub>, <italic>x</italic><sub>2<italic>i</italic></sub>, &#x22EF;, <italic>x</italic><sub><italic>ni</italic></sub>)<sup><italic>T</italic></sup>, <italic>i</italic> = 1, 2, &#x22EF;, <italic>p</italic> can be considered as a sampling of the multivariate random variable (or population) <bold>x</bold><sub><italic>i</italic></sub>, where <italic>p</italic> is the number of ROIs and <italic>n</italic> is the number of time points. Since our goal is to estimate the edge weights of the BFN, the simplest and empirically effective way is to calculate the sample PC coefficient <inline-formula><mml:math id="INEQ6"><mml:mrow><mml:mpadded width="+5pt"><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> of pair-wise ROIs, as follows.</p>
<disp-formula id="S2.E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msqrt><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="INEQ7"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">x</mml:mtext><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the mean vector corresponding to <bold>x</bold><sub><italic>i</italic></sub>. Under Gaussian assumption, Eq. 1 gives an asymptotically unbiased estimation for the population PC coefficient. Without loss of generality, we redefine <inline-formula><mml:math id="INEQ9"><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x225C;</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">x</mml:mtext><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">x</mml:mtext><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">x</mml:mtext><mml:mo>&#x00AF;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mrow></mml:mrow></mml:math></inline-formula>. Then, the estimator of population PC coefficient can be simplified as follows:</p>
<disp-formula id="S2.E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mpadded><mml:mi>o</mml:mi><mml:mpadded width="+3.3pt"><mml:mi>r</mml:mi></mml:mpadded><mml:mpadded width="+3.3pt"><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <bold>X</bold> = [<bold>x</bold><sub>1</sub>, <bold>x</bold><sub>2</sub>, &#x22EF;, <bold>x</bold><sub><italic>p</italic></sub>] is the rs-fMRI data matrix whose columns are the rs-fMRI time series associated with different ROIs. Therein, <inline-formula><mml:math id="INEQ11"><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> is the generalized estimator in matrix form.</p>
</sec>
<sec id="S2.SS2">
<title>Sparse Representation</title>
<p>Sparse representation is one of the commonly used methods for calculating partial correlation among ROIs. A regularization term encoding sparsity prior is introduced into the BFN</p>
<p>estimation model. Specifically, the mathematical model of SR is given by:</p>
<disp-formula id="S2.E3"><label>(3)</label><mml:math id="M3"><mml:mrow><mml:mrow><mml:munder><mml:mo lspace="7.5pt" movablelimits="false">min</mml:mo><mml:mi mathvariant="bold">W</mml:mi></mml:munder><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:munder><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mrow><mml:munder><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <bold>W</bold> is the edge weight matrix of BFN.</p>
<p>Similar to PC, we can rewrite SR in matrix form:</p>
<disp-formula id="S2.Ex1"><mml:math id="M4"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mi mathvariant="bold">W</mml:mi></mml:munder><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:mtext mathvariant="bold">X</mml:mtext><mml:mo>-</mml:mo><mml:mtext mathvariant="bold">XW</mml:mtext></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:msub><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S2.E4"><label>(4)</label><mml:math id="M5"><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>.</mml:mo><mml:mi mathvariant="normal">t</mml:mi><mml:mo rspace="7.5pt">.</mml:mo><mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo rspace="7.5pt">,</mml:mo><mml:mrow><mml:mrow><mml:mo>&#x2200;</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>i</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where the constraint <italic>w</italic><sub><italic>ii</italic></sub> = 0 plays a role in removing <bold>x</bold><sub><italic>i</italic></sub> from <bold>X</bold> to avoid the trivial solution.</p>
</sec>
<sec id="S2.SS3">
<title>Correlation&#x2019;s Correlation</title>
<p>Despite its popularity and effectiveness, the traditional PC can only construct the low-order BFN. That is, the connection between two ROIs is determined by the correlation of the corresponding rs-fMRI time series. However, in practice, a connection can be described in <italic>both</italic> low-order <italic>and</italic> high-order views. For example, we can directly define a connection between ROI <italic>i</italic> and <italic>j</italic> if there is a relationship. Besides, if ROIs <italic>i</italic> and <italic>j</italic> are connected to the same brain region, we can infer with a great possibility that there is a connection between ROI <italic>i</italic> and <italic>j</italic>. The former corresponds to the correlation in the traditional low-order view, while the latter can be considered as CC in a high-order perspective. This results can be achieved by a two-step procedure. First, the low-order BFN is estimated <italic>via</italic> PC. According to the formula in Eqs 1 or 2, the adjacency matrix <inline-formula><mml:math id="INEQ16"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>p</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> of the PC-based BFN can be calculated as follows.</p>
<disp-formula id="S2.E5"><label>(5)</label><mml:math id="M6"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mtext mathvariant="bold">w</mml:mtext><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">w</mml:mtext><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="INEQ17"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">w</mml:mtext><mml:mo>^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="INEQ18"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">w</mml:mtext><mml:mo>^</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> are the <italic>i<sup>th</sup></italic> and <italic>j<sup>th</sup></italic> columns of <inline-formula><mml:math id="INEQ19"><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, respectively. For simplicity, in Eq. 5, <inline-formula><mml:math id="INEQ20"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">w</mml:mtext><mml:mo>^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="INEQ21"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">w</mml:mtext><mml:mo>^</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula> has been centralized and normalized as the case in Eq. 2. As a result, the CC-based high-order BFN is defined as follows,</p>
<disp-formula id="S2.E6"><label>(6)</label><mml:math id="M7"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mover accent="true"><mml:mtext mathvariant="bold">H</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mpadded width="+3.3pt"><mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>h</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>p</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover><mml:mi>T</mml:mi></mml:msup><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:math></disp-formula>
</sec>
</sec>
<sec id="S3">
<title>High-Order Correlation Learning</title>
<p>As described previously, CC constructs the high-order BFN based on two sequential correlation operations. Despite its empirical effectiveness in identifying neuro-disorders (<xref ref-type="bibr" rid="B6">Chen et al., 2016</xref>; <xref ref-type="bibr" rid="B35">Zhang et al., 2016</xref>), CC is a measure defined intuitively without a clear mathematical/probabilistic explanation. Therefore, in this section, we will construct a more rigorous mathematical model for CC, in order to provide a better understanding of the CC-based high-order BFN.</p>
<p>Since CC is based on the PC variant, in the following Section &#x201C;Pearson&#x2019;s Correlation-Based Brain Functional Network Learning Framework in Bayesian View,&#x201D; we first reformulate PC into a more flexible BFN estimation framework. Then, based on the framework, we establish a theoretical model for CC in Section &#x201C;Learning High-Order Brain Functional Network With a Matrix-Normal Penalty.&#x201D; Finally, we design an algorithm for learning the high-order BFN based on the theoretical model of CC in Section &#x201C;Algorithm&#x201D;.</p>
<sec id="S3.SS1">
<title>Pearson&#x2019;s Correlation-Based Brain Functional Network Learning Framework in Bayesian View</title>
<p>As we know, PC is a measure of the linear correlation between pair-wise rs-fMRI time series associated with the ROIs. In other words, a time series <bold>x</bold><sub><italic>i</italic></sub> = (<italic>x</italic><sub>1<italic>i</italic></sub>, <italic>x</italic><sub>2<italic>i</italic></sub>, &#x22EF;, <italic>x</italic><sub><italic>ni</italic></sub>)<sup><italic>T</italic></sup> can be linearly represented by other time series <bold>x</bold><sub><italic>j</italic></sub> = (<italic>x</italic><sub>1<italic>j</italic></sub>, <italic>x</italic><sub>2<italic>j</italic></sub>, &#x22EF;, <italic>x</italic><sub><italic>nj</italic></sub>)<sup><italic>T</italic></sup> as</p>
<disp-formula id="S3.E7"><label>(7)</label><mml:math id="M8"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B5;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>a</italic><sub><italic>ij</italic></sub> is the representative coefficient, and <italic>&#x03B5;</italic><sub><italic>i</italic></sub> = (&#x03B5;<sub>1<italic>i</italic></sub>, &#x03B5;<sub>2<italic>i</italic></sub>, &#x22EF;, &#x03B5;<sub><italic>ni</italic></sub>)<sup><italic>T</italic></sup> is the random error vector. That is, for each variable, we have:</p>
<disp-formula id="S3.E8"><label>(8)</label><mml:math id="M9"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mpadded width="+5pt"><mml:msub><mml:mi mathvariant="normal">&#x03B5;</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mrow><mml:mo>(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>;</mml:mo><mml:mo>;</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Generally, we assume that the random variable &#x03B5;<sub><italic>ki</italic></sub> follows a normal distribution with mathematical expectation 0, i.e., &#x03B5;<sub><italic>ki</italic></sub>&#x223C;&#x1D4A9;(0, &#x03C3;<sup>2</sup>). Therefore, given <italic>x</italic><sub><italic>kj</italic></sub> and <italic>a</italic><sub><italic>ij</italic></sub> are constants, <italic>x</italic><sub><italic>ki</italic></sub> follows the normal distribution <italic>x</italic><sub><italic>ki</italic></sub>&#x223C;&#x1D4A9;(<italic>a</italic><sub><italic>ij</italic></sub><italic>x</italic><sub><italic>kj</italic></sub>, &#x03C3;<sup>2</sup>). Then, The following formula can be obtained by the maximum likelihood estimation of <italic>x</italic><sub><italic>ki</italic></sub> (see <xref ref-type="app" rid="S10.SS1">Appendix A</xref> for details).</p>
<disp-formula id="S3.E9"><label>(9)</label><mml:math id="M10"><mml:mrow><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mo>-</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mi>&#x03C3;</mml:mi></mml:mrow></mml:msqrt><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext>x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext>x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>Note that Eq. 9 can be equivalently written as the following least-squares problem:</p>
<disp-formula id="S3.E10"><label>(10)</label><mml:math id="M11"><mml:mrow><mml:munder><mml:mtext>min</mml:mtext><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:math></disp-formula>
<p>The optimal solution to Problem (10) is given by <inline-formula><mml:math id="INEQ28"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mover accent="true"><mml:mi>a</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, considering that all of time series <bold>x</bold><sub><italic>i</italic></sub>, <italic>i</italic> = 1, 2, &#x22EF;, <italic>p</italic> have been normalized by <inline-formula><mml:math id="INEQ30"><mml:mfrac><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo fence="true" maxsize="142%" minsize="142%">||</mml:mo><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo fence="true" maxsize="142%" minsize="142%">||</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula>. This means that the solution of Eqs 9, 10 is the same as PC shown in Eq. 2. Therefore, in what follows, we only use <italic>w</italic><sub><italic>ij</italic></sub> instead of <italic>a</italic><sub><italic>ij</italic></sub> for the consistency of mathematical notations.</p>
<p>To provide a more flexible framework for BFN estimation, we further generalize PC in Bayesian view by introducing a prior distribution on <italic>w</italic><sub><italic>ij</italic></sub>. Although various distributions can be used as the prior, here we first consider the standard normal distribution, i.e., <italic>w</italic><sub><italic>ij</italic></sub>&#x223C;&#x1D4A9;(0,1), since it provides a basis for understanding more complex cases. However, in practice, the entries <italic>w</italic><sub><italic>ij</italic></sub> in <bold>W</bold> may not be apriori independent of each other, but exist a relationship. Due to <italic>w</italic><sub><italic>ij</italic></sub>&#x223C;&#x1D4A9;(0,1), the edge weight <italic>w</italic><sub><italic>ij</italic></sub> of BFN has the following prior probabilistic density:</p>
<disp-formula id="S3.E11"><label>(11)</label><mml:math id="M12"><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt></mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mmultiscripts><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Next, a maximal posterior estimation of <italic>w</italic><sub><italic>ij</italic></sub> (see <xref ref-type="app" rid="S10.SS2">Appendix B</xref> for details) can be obtained as follows,</p>
<disp-formula id="S3.E12"><label>(12)</label><mml:math id="M13"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mo>-</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>n</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mtext>log</mml:mtext></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mmultiscripts><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>The above problem is equivalent to the regularized least-squares problem:</p>
<disp-formula id="S3.E13"><label>(13)</label><mml:math id="M14"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext>x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext>x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where &#x03BB; = &#x03C3;<sup>2</sup> in the case of standard normal distribution. In practice, <inline-formula><mml:math id="INEQ35"><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mo>&#x225C;</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>/</mml:mo><mml:mmultiscripts><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>0</mml:mn><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts></mml:mrow></mml:mrow></mml:math></inline-formula> is a hyper-parameter that controls the balance between the two terms in Eq. 13, where <inline-formula><mml:math id="INEQ36"><mml:mmultiscripts><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>0</mml:mn><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts></mml:math></inline-formula> corresponds to the variance of normal distribution of the edge weight <italic>w</italic><sub><italic>ij</italic></sub>. Setting the gradient of the objective function to zero, we obtain the optimal solution of Eq. 13 as follows:</p>
<disp-formula id="S3.E14"><label>(14)</label><mml:math id="M15"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">+</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mn>1</mml:mn></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>We find that it is a shrinkage of the original estimation of PC which helps remove the weak connections in BFN.</p>
<p>Note that Eq. 13 only considers finding one edge weight at a time. Without loss of generality, with the assumption that the variables <italic>w</italic><sub><italic>ij</italic></sub> in <bold>W</bold> are independent, we can rewrite Eq. 13 in the following matrix form:</p>
<disp-formula id="S3.E15"><label>(15)</label><mml:math id="M16"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mi mathvariant="bold">W</mml:mi></mml:munder><mml:msubsup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mtext mathvariant="bold">WW</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="INEQ38"><mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mtext mathvariant="bold">WW</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msub><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mrow></mml:math></inline-formula> is the trace operator of <bold>WW</bold><sup><italic>T</italic></sup>. As a result, we achieve the estimation of BFN in a batching way as follows,</p>
<disp-formula id="S3.E16"><label>(16)</label><mml:math id="M17"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mn>1</mml:mn></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>This formula is essentially a generalization of Eq. 2 with a shrinkage factor (1+&#x03BB;)<sup>&#x2212;1</sup>. When &#x03BB;+0, Eq. 16 reduces to the traditional PC.</p>
</sec>
<sec id="S3.SS2">
<title>Learning High-Order Brain Functional Network With a Matrix-Normal Penalty</title>
<p>In Section &#x201C;Pearson&#x2019;s Correlation-Based Brain Functional Network Learning Framework in Bayesian View,&#x201D; we reformulate PC and then generalize it in Bayesian view by introducing a standard normal prior <italic>w</italic><sub><italic>ij</italic></sub>&#x223C;&#x1D4A9;(0,1) for each pair of ROIs (<italic>i</italic>, <italic>j</italic>), <italic>i</italic>, <italic>j</italic> = 1, 2, &#x22EF;, <italic>p</italic>. However, in practice, the entries <italic>w</italic><sub><italic>ij</italic></sub> in <bold>W</bold> may not be apriori independent of each other, but exist a relationship. Even so, Section &#x201C;Pearson&#x2019;s Correlation-Based Brain Functional Network Learning Framework in Bayesian View&#x201D; provides a flexible probabilistic framework to develop new brain network estimation methods. Inspired by this point, we lay down theoretical support for CC from the Bayesian perspective. More importantly, instead of assuming that <italic>w</italic><sub><italic>ij</italic></sub> in <bold>W</bold> are independent, we propose a Bayesian high-order model (BHM) for BFN estimation by introducing the prior of matrix-variate normal distribution to the low-order BFN <bold>W</bold>. BHM learns the high-order relationship from the data automatically, rather than manually define as the case in CC. Interestingly, such a scheme can simultaneously learn low-order and high-order BFNs by considering the spatial structure of network connections. To distinguish the low-order and high-order correlations, an illustration is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The connections among <italic>w</italic><sub><italic>ij</italic></sub> can be considered as a high-order correlation <italic>h</italic><sub><italic>ij,kl</italic></sub>, while <italic>w</italic><sub><italic>ij</italic></sub> denotes the traditional low-order linear correlation between ROIs.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The diagram of low- and high-order connections.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g001.tif"/>
</fig>
<sec id="S3.SS2.SSS1">
<title>Model</title>
<p>Since we can vectorize the low-order edge weight matrix <bold>W</bold> into a <italic>p</italic><sup>2</sup> = 1 vector, the low- and high- order correlation can be modeled by a multivariate normal distribution, <italic>vec</italic>(<bold>W</bold>&#x223C;)&#x1D4A9;(<bold>O</bold>, <bold>&#x03A9;</bold>), where <bold>W</bold> encodes the low-order relationship and <bold>&#x03A9;</bold> &#x2208; <italic>R</italic><sup><italic>p</italic><sup>2</sup> = <italic>p</italic><sup>2</sup></sup> is the covariance matrix for modeling the relationship between the entries in <bold>W</bold>. Despite the theoretical feasibility for encoding the high-order relationship, <bold>W</bold>-vectorization ignores its spatial structure as a matrix. Even worse, the estimation of <bold>&#x03A9;</bold> is extremely challenging due to its high dimension. Such a scale <italic>not only</italic> goes beyond the storage ability of the general memory (Specifically, in our experiment, <italic>p</italic> is 160, which takes up about 4.9 GB storage), <italic>but also</italic> may lead to the overfitting problem. Therefore, we further assume that the covariance matrix <bold>&#x03A9;</bold> has the Kronecker product decomposition (<xref ref-type="bibr" rid="B14">Gupta and Nagar, 2000</xref>), i.e., <bold>&#x03A9;</bold> = <bold>&#x03A9;</bold><sub>1</sub>&#x2297;<bold>&#x03A9;</bold><sub>2</sub>, where <bold>&#x03A9;</bold><sub>1</sub> and <bold>&#x03A9;</bold><sub>2</sub> denote the <bold>row</bold> and <bold>column</bold> covariance matrices, respectively. That is, <bold>W</bold> follows the distribution <bold>W</bold>&#x223C;&#x2133;&#x1D4A9;(<bold>O</bold>, <bold>&#x03A9;</bold><sub>1</sub>&#x2297;<bold>&#x03A9;</bold><sub>2</sub>). As described earlier, in this paper, we mainly focus on correlation-based methods that generally result in the symmetric BFN. Therefore, the row and column covariance matrices of <bold>W</bold> are the same, i.e., <bold>&#x03A9;</bold><sub>1</sub> = <bold>&#x03A9;</bold><sub>2</sub>, and without loss of generality, we define <bold>&#x03A9;</bold>&#x225C;<bold>&#x03A9;</bold><sub>1</sub> = <bold>&#x03A9;</bold><sub>2</sub>. As a result, the matrix-variate normal distribution (<xref ref-type="bibr" rid="B14">Gupta and Nagar, 2000</xref>) of the low-order BFN <bold>W</bold> has a probability density:</p>
<disp-formula id="S3.E17"><label>(17)</label><mml:math id="M18"><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold">W</mml:mi><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold">W</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Similar to the formulation in Eqs 11&#x2013;16, we take Eq. 17 as a prior distribution of low-order network <bold>W</bold>. Then, we can formulate the posterior probability of <bold>W</bold> based on the Bayesian rule (see <xref ref-type="app" rid="S10.SS3">Appendix C</xref> for details). By maximizing the posterior probability, the low- and high-order BFN mutual learning model can be obtained as follows:</p>
<disp-formula id="S3.E18"><label>(18)</label><mml:math id="M19"><mml:mtable><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mi>W</mml:mi><mml:mo>-</mml:mo><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mo>|</mml:mo><mml:msubsup><mml:mo>|</mml:mo><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mrow><mml:mo maxsize="210%" minsize="210%">[</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mspace width="15em"/><mml:mrow><mml:mo>+</mml:mo><mml:mi>p</mml:mi><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo maxsize="210%" minsize="210%">]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <bold>W</bold> is the Bayesian low-order BFN, <bold>&#x03A9;</bold> corresponds to the Bayesian high-order BFN, <italic>p</italic> is the number of ROIs, and &#x03BB; is a hyper-parameter that controls the balance between the two terms in the objective function.</p>
</sec>
<sec id="S3.SS2.SSS2">
<title>Algorithm</title>
<p>The alternating optimization (AO) scheme (<xref ref-type="bibr" rid="B3">Bezdek and Hathaway, 2002</xref>) is employed to solve Problem (18). More specifically, we first initialize the low-order BFN using the PC estimator, i.e., <inline-formula><mml:math id="INEQ70"><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo rspace="5.8pt">=</mml:mo><mml:mpadded width="+3.3pt"><mml:mover accent="true"><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow></mml:math></inline-formula>, and then alternatively optimize <bold>W</bold> and &#x03A9;.</p>
<p><bold><underline>Step 1</underline></bold> Fix <bold>W</bold> and solve <bold>&#x03A9;</bold>. The optimization problem is</p>
<disp-formula id="S3.E19"><label>(19)</label><mml:math id="M20"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:munder><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo rspace="5.3pt">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>which can be solved by the following iterative formula (<xref ref-type="bibr" rid="B9">Dutilleul, 1999</xref>; <xref ref-type="bibr" rid="B36">Zhang and Schneider, 2010</xref>)</p>
<disp-formula id="S3.E20"><label>(20)</label><mml:math id="M21"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi mathvariant="bold">&#x03A9;</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Note that, by initializing <bold>&#x03A9;</bold> = <bold>I</bold> in Eq. 20, at the first iteration, we obtain <bold>&#x03A9;</bold> = <bold>W</bold><sup>T</sup><bold>W</bold>, which reduces to the traditional CC, as <inline-formula><mml:math id="INEQ77"><mml:mover accent="true"><mml:mtext mathvariant="bold">H</mml:mtext><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> defined in Eq. 6. In other words, the traditional CC is only a rough estimation of the theoretical value at the first iteration. We can continue the iteration toward a more accurate estimation of <bold>&#x03A9;</bold>. In fact, with the estimated <bold>&#x03A9;</bold>, we can further update <bold>W</bold> according to the AO scheme. In practice, we generally add a small quantity &#x03B4;<bold>I</bold> to Eq. 20 for a more stable numerical solution where &#x03B4; is a small positive constant.</p>
<p><bold><underline>Step 2</underline></bold> Fix <bold>&#x03A9;</bold> and solve <bold>W</bold>. The optimization problem is</p>
<disp-formula id="S3.E21"><label>(21)</label><mml:math id="M22"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mi mathvariant="bold">W</mml:mi></mml:munder><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:mfrac><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>With the fixed <bold>&#x03A9;</bold>, the gradient of Eq. 21 with respect to <bold>W</bold> is</p>
<disp-formula id="S3.E22"><label>(22)</label><mml:math id="M23"><mml:mrow><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Setting the gradient equal to zero, we obtain:</p>
<disp-formula id="S3.E23"><label>(23)</label><mml:math id="M24"><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>We summarize the algorithm for solving Problem (18) in <xref ref-type="other" rid="Box1">Algorithm 1</xref>.</p>
<boxed-text id="Box1" position="float">
<title>Algorithm 1: Estimating BFN with BHM model.</title>
<graphic xlink:href="fnins-16-872848-t00a1.jpg"/>
</boxed-text>
</sec>
</sec>
</sec>
<sec id="S4" sec-type="results">
<title>Experiments and Results</title>
<sec id="S4.SS1">
<title>Data Acquisitions and Processing</title>
<p>To evaluate the effectiveness of the proposed BHM, we conduct experiments on Autism Brain Imaging Data Exchange (ABIDE) database. The objective is to identify subjects with ASD from typical controls (TCs). Considering the heterogeneity of multi-site data, we only use data from the <italic>NYU</italic> site in our study. The dataset includes 184 subjects (79 ASD patients and 105 TCs). The detailed scan procedures and protocols are described on the ABIDE website.<sup><xref ref-type="fn" rid="footnote2">2</xref></sup> The demographic information of all participants is summarized and displayed in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Demographic information of the used dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td/>
<td valign="top" align="center">ASD (<italic>n</italic> = 79)</td>
<td valign="top" align="center">TC (<italic>n</italic> = 105)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Gender (M/F)</td>
<td valign="top" align="center">68/11</td>
<td valign="top" align="center">79/26</td>
</tr>
<tr>
<td valign="top" align="left">Age (year SD)</td>
<td valign="top" align="center">14.51 &#x00B1; 6.23</td>
<td valign="top" align="center">15.80 &#x00B1; 3.23</td>
</tr>
<tr>
<td valign="top" align="left">FIQ (mean SD)</td>
<td valign="top" align="center">107.91 &#x00B1; 16.62</td>
<td valign="top" align="center">113.15 &#x00B1; 13.12</td>
</tr>
<tr>
<td valign="top" align="left">ADOS (mean SD)</td>
<td valign="top" align="center">11.3 &#x00B1; 4.08</td>
<td valign="top" align="center">&#x2013;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>ASD, autism spectrum disorders; TC, typical control; FIQ, full intelligence quotient; ADOS, autism diagnostic observation schedule.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>All rs-fMRI images were acquired using a standard echo-planar imaging sequence on a clinical routine 3T Siemens Allegra scanner. During the 6-min rs-fMRI scanning procedure, most subjects were required to relax with their eyes focusing on a white fixation cross in the middle of the black background screen projected on a screen. A few participants close their eyes. The functional scan parameters include the flip angle = 90&#x00B0;, 33 slices, TR/TE = 2000/15 ms with 180 volumes, FOV = 240 mm and voxel size = 3 &#x00D7;3 &#x00D7;4 mm<sup>3</sup>. The rs-fMRI data were preprocessed by DPARSF<sup><xref ref-type="fn" rid="footnote3">3</xref></sup> software. Specifically, to avoid the interference of early signal instability, the first 5 rs-fMRI volumes of each subject were discarded. The remaining volumes were calibrated as follows: (1) Slice timing correction and head motion correction; (2) Regression of nuisance signals (ventricle, white matter) and head-motion with Friston 24-parameter model (<xref ref-type="bibr" rid="B13">Friston et al., 1996</xref>); (3) Normalization and register to MNI space with resolution of 3 3 3 mm<sup>3</sup>; (4) Segmentation using DATTEL; (5) Spatial smoothing by a kernel of 6 mm. After that, since our focus is functional connectivity, the rs-fMRI time series signals were partitioned into 160 ROIs, according to the functional atlas Dosenbach 160 (<xref ref-type="bibr" rid="B8">Dosenbach et al., 2010</xref>). Finally, the mean time series of the ROI were put into a data matrix <bold>X</bold> &#x2208; <italic>R</italic><sup>175=160</sup>, which will be used for the subsequent BFN estimation.</p>
</sec>
<sec id="S4.SS2">
<title>Brain Functional Network Construction, Feature Selection, and Classification</title>
<p>With the preprocessed rs-fMRI data, we estimate the low- and high-order BFNs using the proposed method, i.e., Bayesian low-order Network <bold>W</bold> and Bayesian high-order Network <bold>&#x03A9;</bold>, respectively. For comparison, we also choose PC, SR, and traditional CC as baseline methods to construct BFNs.</p>
<p>Once the BFNs are constructed, the next step is feature selection and classification. In our study, we directly use edge weights of the estimated BFN as features for ASD identification. Despite its simplicity (without complex feature design), such a scheme easily causes the curse of dimensionality due to limited sample size. As described previously, the number of ROIs is 160 and thus the estimated feature edges are 160=(160&#x2212;1)/2=12720, which is far greater than the sample size (i.e., the number of subjects 184). To alleviate the problem of small sample size, we adopt a two-sample <italic>t</italic>-test with an empirically fixed <italic>p</italic> values to select features before ASD classification. In our experiments, we evaluate five candidate parametric values of <italic>p</italic>, that is [0.001,0.005,0.01,0.05,0.1]. The specific parameter analysis results are given in Section &#x201C;Sensitivity to Network Modeling Parameters.&#x201D;</p>
<p>To perform the following classification task, we use a linear support vector machine (SVM) (<xref ref-type="bibr" rid="B5">Chang and Lin, 2011</xref>) with default <italic>C</italic> = 1 as the classifier. To evaluate the model, we adopt leave-one-out cross-validation (LOOCV) in our experiments due to the limited data samples. Specifically, a LOOCV works in each run and only one sample is used to test while the rest are used to train a classifier. The final performance is obtained by the averaged results of all the runs. Note that the model parameters are involved in certain methods, including SR and BHM. Therefore, we additionally adopt an inner LOOCV procedure on the training data to obtain the optimal parametric value. Specifically, for SR, the regularization parameter &#x03BB; is set to [2<sup>&#x2212;2</sup>,&#x2005;2<sup>&#x2212;1</sup>,&#x2005;2<sup>0</sup>,&#x2005;2<sup>1</sup>,&#x2005;2<sup>2</sup>]. For the proposed BHM, the regularization parameter &#x03BB; is set to [0.0001,&#x2005;0.001,&#x2005;0.01,&#x2005;0.1,&#x2005;1]. To be consistent with the number of parameters in other methods, the coefficient &#x03B4; of the perturbation involved in Eq. 20 is set to 0.1 empirically.</p>
</sec>
<sec id="S4.SS3">
<title>Classification Results</title>
<p>To evaluate the classification results of different methods, we use accuracy (ACC), sensitivity (SEN), specificity (SPE) as performance metrics. The definition of these quantities are reported in <xref ref-type="table" rid="T2">Table 2</xref>. Note that, in this work, we treat ASD patients as the positive class while the NCs as the negative class.</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>Different performance metrics.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Performance metrics</td>
<td valign="top" align="center">Abbreviations</td>
<td valign="top" align="center">Definitions</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Accuracy</td>
<td valign="top" align="center">ACC</td>
<td valign="top" align="center"><inline-formula><mml:math id="INEQ117"><mml:mfrac><mml:mrow><mml:mpadded width="+3.3pt"><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mpadded width="+3.3pt"><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mpadded width="+3.3pt"><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:mtext>FN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Sensitivity</td>
<td valign="top" align="center">SEN</td>
<td valign="top" align="center"><inline-formula><mml:math id="INEQ118"><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:mtext>FN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>
<tr>
<td valign="top" align="left">Specificity</td>
<td valign="top" align="center">SPE</td>
<td valign="top" align="center"><inline-formula><mml:math id="INEQ119"><mml:mfrac><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow><mml:mrow><mml:mpadded width="+3.3pt"><mml:mrow><mml:mtext>TN</mml:mtext></mml:mrow></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>TP, TN, FP, and FN indicate true positive, true negative, false positive, and false negative, respectively.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>In <xref ref-type="table" rid="T3">Table 3</xref>, we report the ASD classification results of five methods. It can be observed that the Bayesian low-order network (BHM-<bold>W</bold>) and Bayesian high-order network (BHM-<bold>&#x03A9;</bold>) constructed by the proposed BHM perform better than the BFNs constructed by the traditional PC and CC, respectively. Moreover, BHM-<bold>&#x03A9;</bold> achieves the best performance. Besides, the high-order BFNs (traditional CC and BHM-<bold>&#x03A9;</bold>) are associated with better recognition performance when they are compared with the baseline methods PC, SR, and BHM-<bold>W</bold>. This means that the high-order network structure can provide more helpful information for BFN analysis to some extent. Furthermore, for two corresponding low-order methods, the performance of the traditional PC and BHM-<bold>W</bold> are approximately similar whereas BHM-<bold>W</bold> has slightly better accuracy than the traditional PC. This may benefit from the guidance information provided by the Bayesian high-order network <bold>&#x03A9;</bold> in the optimization process.</p>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>The classification results based on five different methods for ASD identification.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Methods</td>
<td valign="top" align="center">ACC</td>
<td valign="top" align="center">SEN</td>
<td valign="top" align="center">SPE</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">PC</td>
<td valign="top" align="center">0.6359</td>
<td valign="top" align="center">0.6222</td>
<td valign="top" align="center">0.6383</td>
</tr>
<tr>
<td valign="top" align="left">SR</td>
<td valign="top" align="center">0.6033</td>
<td valign="top" align="center">0.2658</td>
<td valign="top" align="center">0.8571</td>
</tr>
<tr>
<td valign="top" align="left">CC</td>
<td valign="top" align="center">0.6630</td>
<td valign="top" align="center">0.5570</td>
<td valign="top" align="center">0.7429</td>
</tr>
<tr>
<td valign="top" align="left">BHM-<bold>W</bold></td>
<td valign="top" align="center"><bold>0.6576</bold></td>
<td valign="top" align="center"><bold>0.5316</bold></td>
<td valign="top" align="center"><bold>0.7143</bold></td>
</tr>
<tr>
<td valign="top" align="left">BHM-<bold>&#x03A9;</bold></td>
<td valign="top" align="center"><bold>0.7283</bold></td>
<td valign="top" align="center"><bold>0.6329</bold></td>
<td valign="top" align="center"><bold>0.8000</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="S5" sec-type="discussion">
<title>Discussion</title>
<sec id="S5.SS1">
<title>Brain Functional Network Visualization</title>
<p>To evaluate the BFNs estimated by different methods, we randomly select a subject and visualize the BFNs constructed by PC, SR, CC, BHM, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. Specifically, the different colors of <xref ref-type="fig" rid="F2">Figure 2</xref> indicate different weights of the edge weights matrix (i.e., the BFN), ranging from &#x2212;1 to 1. It is observed that: (1) Compared with SR, the BFNs estimated by the correlation-based methods (i.e., PC, CC, BHM) are denser since the sparsity prior is introduced into SR. (2) There are fewer areas of cold colors in the PC network heatmap, implying that the edges with the negative weights are less. (3) Compared with PC, CC&#x2019;s network heatmap shows a sharper distinction between the areas of warm and cold colors, indicating that the positive edge weights of CC&#x2019;s BFN are larger and the negative edge weights are smaller. (4) The BHM-<bold>W</bold> estimated from the Bayesian perspective has a greater distinction between positive and negative edge weights than PC. (5) The BHM-<bold>&#x03A9;</bold> as a Bayesian version of the traditional CC tends to produce a greater distinction between positive and negative edge weights than CC. Combined with the fact that the high classification accuracy of the BHM method, we can infer that the negative edge weights of BFN also have important information for classification. (6) The BFNs based on the correlation methods show a degree of consistency., as shown in the black box in <xref ref-type="fig" rid="F2">Figure 2</xref>. Similar structures appear in the four BFNs estimated by PC, CC, BHM, which can provide certain support for the reliability of the three correlation methods. In addition, we cluster the brain regions using spectral clustering (<xref ref-type="bibr" rid="B26">Ng et al., 2001</xref>) and visualize the clustered adjacency matrices of these 5 methods in <xref ref-type="fig" rid="F3">Figure 3</xref>. It is observed that the BFN estimated by BHM-<bold>&#x03A9;</bold> shows a more significant modular structure.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>The BFN adjacency matrices of different methods. The patches marked by the black box are the consistent part of the network constructed by different methods.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Five clustered edge weight matrices of the same subject estimated by different methods.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g003.tif"/>
</fig>
</sec>
<sec id="S5.SS2">
<title>Sensitivity to Network Modeling Parameters</title>
<p>As stated earlier, some methods including SR and BHM involve optional model parameters. Different parameter values may have a significant impact on the results. Therefore, we calculate the accuracy of different methods under different parameter values, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. It is worth noting that, traditional PC and CC models do not involve optional parameters. However, for comparison, we fix their values with the final classification accuracy in <xref ref-type="fig" rid="F4">Figure 4</xref> for visualization. We can observe that BHM-<bold>&#x03A9;</bold> and BHM-<bold>W</bold> are quite sensitive to the parameters. When &#x03BB; in BHM is set to a large real number, the accuracy of BHM-<bold>&#x03A9;</bold> decreases, which may be because the large value of &#x03BB;, the algorithm has difficulties to converge. Besides, the accuracy of BHM-<bold>W</bold> increases as &#x03BB; increases. For this, we empirically tested a larger lambda range [2<sup>&#x2212;5</sup>,&#x2005;2<sup>&#x2212;4</sup>, &#x22EF;,&#x2005;2<sup>0</sup>, &#x22EF;,&#x2005;2<sup>4</sup>,&#x2005;2<sup>5</sup>] and find that as &#x03BB; continues to increase, the accuracy decreases, which is consistent with the performance of BHM-<bold>&#x03A9;</bold>. Moreover, SR is not sensitive to different parameter values, but its accuracy performs average in general.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Classification accuracy of ASD identification based on 5 BFNs estimated by PC, CC, SR, and BHM with 5 different parametric values. Although PC and CC have no optional parameters, in order to facilitate comparison, we visualize the accuracy of PC and CC in the left chart.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g004.tif"/>
</fig>
<p>Considering that different <italic>p</italic>-values significantly influence the results, we show the classification accuracies of 5 methods under different <italic>p</italic>-values in <xref ref-type="fig" rid="F5">Figure 5</xref>. Note that all 5 methods are sensitive to different <italic>p</italic>-values. We selected the optimal parameter value for feature selection, so that different methods can get the best classification performance.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Classification accuracy of ASD identification based on 5 BFNs estimated by PC, CC, SR, and BHM for 5 different <italic>p</italic>-values.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g005.tif"/>
</fig>
</sec>
<sec id="S5.SS3">
<title>Top Discriminative Features</title>
<p>In this work, for the ASD classification task, we use the edge weights of the estimated BFN as features. With the empirically optimal parameter, we construct the BFNs using the proposed BHM, then apply a two-sample <italic>t</italic>-test to rearrange the features according to the <italic>p</italic>-values. Particularly, we choose the BHM-<bold>&#x03A9;</bold> since it outperforms the BFNs estimated by the other methods. As a result, we obtain the discriminative edge connections with a threshold value <italic>p</italic> &#x003C; 0.001 as shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. Here, the thickness of each arc represents the discriminative power that is inversely proportional to the corresponding <italic>p</italic>-value. The colors of each arc are assigned randomly for better visualization.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>The most discriminative edge features of the BHM-<bold>&#x03A9;</bold> involved in the ASD classification task by using a <italic>t</italic>-test (<italic>p</italic> &#x003C; 0.001). This figure is created by the circularGraph tool, which is designed by Paul Kassebaum and can be downloaded from <ext-link ext-link-type="uri" xlink:href="http://www.mathworks.com/matlabcentral/fileexchange/48576-circulargraph">http://www.mathworks.com/matlabcentral/fileexchange/48576-circulargraph</ext-link>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g006.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F6">Figure 6</xref>, the top discriminative features and the corresponding brain regions, that may contribute to ASD identification include occipital lobe, post-cingulate, dorsal frontal cortex, inferior parietal lobule, precuneus, anterior prefrontal cortex, lateral cerebellum, temporal lobe, fusiform gyrus, mid insula, etc. in order of discriminant ability. The findings are consistent with previous studies (<xref ref-type="bibr" rid="B27">Nickl-Jockschat et al., 2012</xref>; <xref ref-type="bibr" rid="B16">Hashem et al., 2020</xref>; <xref ref-type="bibr" rid="B19">Lau et al., 2020</xref>). We visualize the ROIs using the Brainmesh of Ch2 with Cerebellum in <xref ref-type="fig" rid="F7">Figure 7</xref>, where the size of node spheres depends on the original value in the node file provided by the Dosenbach 160 template.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>The full view of most relevant ROI associated with the ASD classification task based on BHM-&#x03A9;. This visualization is created using the BrainNet Viewer (<ext-link ext-link-type="uri" xlink:href="https://www.nitrc.org/projects/bnv/">https://www.nitrc.org/projects/bnv/</ext-link>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-872848-g007.tif"/>
</fig>
</sec>
<sec id="S5.SS4">
<title>Other Distribution Priors</title>
<p>As described before, we first give an equivalent probability explanation for PC by introducing a normal distribution for the rs-fMRI signal values. Then we reformulate PC with Bayesian rule, thus getting two perspectives of PC. This provides a platform for generalizing PC to CC by assuming that the edge weight matrix <bold>W</bold> follows the MVN distribution prior. As a result, we derive a probabilistic explanation of CC and develop a high-order BFN estimation framework that allows the introduction of different priors (or regularizers).</p>
<p>Besides the introduced normal distribution prior on <bold>W</bold> for BFN estimation, we can also introduce other priors on <bold>W</bold>. For example, considering Laplacian distribution prior for <italic>w</italic><sub><italic>ij</italic></sub>, e.g., <inline-formula><mml:math id="INEQ144"><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow></mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></inline-formula> where &#x03B2; is a scale parameter. In this way, the regularized least square problem is</p>
<disp-formula id="S5.E24"><label>(24)</label><mml:math id="M25"><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mrow><mml:munder><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>We can find that the Laplacian distribution generates sparse BFN due to the regularizer <italic>w</italic><sub><italic>ij</italic></sub>. Besides, we can get the optimal solution by the soft thresholding, as follows:</p>
<disp-formula id="S5.E25"><label>(25)</label><mml:math id="M26"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>w</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnspacing="5pt" displaystyle="true" rowspacing="0pt"><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mrow><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign="right"><mml:mrow><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="right"><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign="right"><mml:mrow><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>&lt;</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mi/></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Although different prior distributions can be tried to introduce the proposed probabilistic framework, we do not formulate their models in detail since this paper focuses on the formulation of CC.</p>
</sec>
</sec>
<sec id="S6" sec-type="conclusion">
<title>Conclusion</title>
<p>In this paper, we propose a probabilistic high-order BFN learning framework with a matrix normal penalty for ASD identification. As pointed out previously, CC is intuitively defined based on two sequential PC operations and falls short of a rigorous mathematical basis. To address this issue, we first reformulate PC with Bayesian rule and then generalize PC to CC by assuming that the edge weight matrix follows a matrix-variate normal distribution prior. This work lays the theoretical foundation for CC methods, leading to a better understanding of high-order BFN learning. In this base, we develop a Bayesian High-order Model to simultaneously estimate the high- and low-order BFN. To efficiently solve the proposed objective function, an alternating optimization algorithm is proposed. Extensive experiments on the NYU site of ABIDE dataset demonstrate the effectiveness of the proposed method, in comparison to the baseline methods. Especially for the BHM-<bold>&#x03A9;</bold>, it achieves the best performance. Note that we only construct high-order BFN based on PC. In principle, any correlation-based BFN estimation method (e.g., SR) can be embedded in the proposed probabilistic framework. In the future, we plan to validate the scheme on the other correlation-based BFN models.</p>
</sec>
<sec id="S7" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="S8">
<title>Author Contributions</title>
<p>LQ proposed the idea of Bayesian high-order model. RD steered the structure of the manuscript and polished the language. XJ validated the models&#x2019; derivation, coded and performed the experiments, and planned and wrote the manuscript. YuZ set the procedures of ASD identification experiments. YiZ and LZ designed a core derivation of the models. All authors developed the estimation algorithm and contributed to the preparation of the manuscript.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="pudiscl1" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="S9" sec-type="funding-information">
<title>Funding</title>
<p>This work was partly supported by the National Natural Science Foundation of China (Nos. 61976110, 62176112, and 11931008), Natural Science Foundation of Shandong Province (No. ZR202102270451), and The Open Project of Liaocheng University Animal Husbandry Discipline (No. 319312101-01).</p>
</sec>
<app-group>
<app id="A1">
<title>Appendix</title>
<p>This Appendix consists of three appendices: Appendix A gives a detailed explanation of Eq. 9, Appendix B presents a detail for Eq. 12 and Appendix C gives a detailed derivation of Eq. 18. To keep the process of derivation smooth, we write the formulas that appeared above with the original number in the appendix, while the new formulas in the process of derivation was renumbered.</p>
<sec id="S10.SS1">
<title>Appendix A</title>
<p>Given <italic>x</italic><sub><italic>kj</italic></sub> and <italic>a</italic><sub><italic>ij</italic></sub> are constants, <italic>x</italic><sub><italic>ki</italic></sub> follows the normal distribution <italic>x</italic><sub><italic>ki</italic></sub>&#x223C;&#x1D4A9;(<italic>a</italic><sub><italic>ij</italic></sub><italic>x</italic><sub><italic>kj</italic></sub>, &#x03C3;<sup>2</sup>). The conditional distribution can be written as</p>
<disp-formula id="S10.E26"><label>(A1&#x2032;)</label><mml:math id="M27"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow></mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>Assuming that the variables <italic>x</italic><sub><italic>ki</italic></sub> of rs-fMRI time series <bold>x</bold><sub><italic>i</italic></sub>, <italic>i</italic> = 1, &#x22EF;, <italic>p</italic> are independent identically distributed, the likelihood function can be written as follows:</p>
<disp-formula id="S10.E27"><label>(A2&#x2032;)</label><mml:math id="M28"><mml:mtable><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x220F;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mspace width="7.3em"/><mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>To avoid overflow caused by multiplying operations in Eq. A2&#x2032;, we use the log-likelihood function and further maximize it:</p>
<disp-formula id="S10.E28"><label>(A3&#x2032;)</label><mml:math id="M29"><mml:mrow><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mtext>log</mml:mtext><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo rspace="7.5pt">,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mo>-</mml:mo><mml:mi>n</mml:mi><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo>)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
</sec>
<sec id="S10.SS2">
<title>Appendix B</title>
<p>The edge weight <italic>w</italic><sub><italic>ij</italic></sub> of BFN has the following prior probabilistic density:</p>
<disp-formula id="S10.E29"><label>(A11)</label><mml:math id="M30"><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt></mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mmultiscripts><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>According to Bayesian rule, the posterior distribution of <italic>w</italic><sub><italic>ij</italic></sub> is proportional to <italic>P</italic>(<bold>x</bold><sub><italic>i</italic></sub>|<italic>w</italic><sub><italic>ij</italic></sub>, <bold>x</bold><sub><italic>j</italic></sub>)<italic>P</italic>(<italic>w</italic><sub><italic>ij</italic></sub>):</p>
<disp-formula id="S10.E30"><label>(A4&#x2032;)</label><mml:math id="M31"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x221D;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>P</italic>(<bold>x</bold><sub><italic>i</italic></sub>|<italic>w</italic><sub><italic>ij</italic></sub>, <bold>x</bold><sub><italic>j</italic></sub>) is the likelihood of <italic>w</italic><sub><italic>ij</italic></sub> for <bold>x</bold><sub><italic>i</italic></sub>. Based on Eqs A2&#x2032;, 11</p>
<disp-formula id="S10.E31"><label>(A5&#x2032;)</label><mml:math id="M32"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="10.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>n</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mmultiscripts><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>Taking the logarithm on Eq. A4&#x2032;, we obtain</p>
<disp-formula id="S10.E32"><label>(A6&#x2032;)</label><mml:math id="M33"><mml:mrow><mml:mtext>log</mml:mtext><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo rspace="7.5pt">,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x221D;</mml:mo><mml:mo>-</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>n</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mtext>log</mml:mtext><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac><mml:mo>-</mml:mo><mml:mfrac><mml:mmultiscripts><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>Next, the log-posterior probability is maximized (i.e., maximal posterior estimation) as follows,</p>
<disp-formula id="S10.E33"><label>(A12)</label><mml:math id="M34"><mml:mrow><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:munder><mml:mo>-</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>n</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mtext>log</mml:mtext><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mpadded width="+3.3pt"><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mmultiscripts><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:none/><mml:mn>2</mml:mn></mml:mmultiscripts></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
</sec>
<sec id="S10.SS3">
<title>Appendix C</title>
<p>As shown in Section &#x201C;Learning High-Order Brain Functional Network With a Matrix-Normal Penalty,&#x201D; the proposed model is formulated as</p>
<disp-formula id="S10.E34"><label>(A18)</label><mml:math id="M35"><mml:mrow><mml:mrow><mml:mi>J</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mrow><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>As described before, we assume that any time series <bold>x</bold><sub><italic>i</italic></sub> from the data set <bold>X</bold> = [<bold>x</bold><sub>1</sub>, <bold>x</bold><sub>2</sub>, &#x22EF;, <bold>x</bold><sub><italic>p</italic></sub>]<sup><italic>T</italic></sup> follows the conditional distribution as:</p>
<disp-formula id="S10.E35"><label>(A2&#x2032;)</label><mml:math id="M36"><mml:mtable><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo rspace="7.5pt">,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x220F;</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo rspace="7.5pt">,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mspace width="7.7em"/><mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mi mathvariant="normal">&#x03C3;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Using the maximum likelihood estimation for <italic>a</italic><sub><italic>ij</italic></sub>, we get <inline-formula><mml:math id="INEQ153"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msubsup><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. Therefore, we rewrite Eq. A2&#x2032; as follows.</p>
<disp-formula id="S10.E36"><label>(A7&#x2032;)</label><mml:math id="M37"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo rspace="7.5pt">,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:msup><mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>We can further convert Eq. A7&#x2032; to the matrix form:</p>
<disp-formula id="S10.E37"><label>(A8&#x2032;)</label><mml:math id="M38"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mtext mathvariant="bold">X</mml:mtext><mml:mo stretchy="false">|</mml:mo><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo rspace="5.8pt" stretchy="false">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x220F;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub><mml:mo rspace="7.5pt">,</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msqrt><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow></mml:msqrt><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mi>n</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="false"><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>Furthermore, as mentioned in Section &#x201C;Learning High-Order Brain Functional Network With a Matrix-Normal Penalty,&#x201D; we assume that the low-order correlation matrix <bold>W</bold> follows the MVN distribution with the probabilistic density</p>
<disp-formula id="S10.E38"><label>(A9&#x2032;)</label><mml:math id="M39"><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold">W</mml:mi><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold">W</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Note that the row\column variance matrix <bold>&#x03A9;</bold> is considered as the high-order correlation matrix between <bold>x</bold><sub><italic>i</italic></sub> since it models the relationships between <italic>w</italic><sub><italic>ij</italic></sub>.</p>
<p>Based on the Bayesian rule, the posterior probability of <bold>W</bold> is given by</p>
<disp-formula id="S10.E39"><label>(A10&#x2032;)</label><mml:math id="M40"><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x221D;</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">X</mml:mtext><mml:mo lspace="2.5pt" rspace="2.5pt" stretchy="false">|</mml:mo><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold">W</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>From Eqs A8&#x2032;, A9&#x2032;, we obtain</p>
<disp-formula id="S10.E40"><label>(A11&#x2032;)</label><mml:math id="M41"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">X</mml:mtext><mml:mo>|</mml:mo><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x03C0;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:mo>-</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>n</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:msup><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mi>n</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:msup><mml:mo>|</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mrow><mml:mtext>tr</mml:mtext></mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold">W</mml:mi><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold">W</mml:mi><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="false"><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold">x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true" maxsize="200%" minsize="200%">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>Taking the logarithm of the above likelihood function and then maximizing it, we get</p>
<disp-formula id="S10.E41"><label>(A12&#x2032;)</label><mml:math id="M42"><mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">max</mml:mo><mml:mrow><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mo>-</mml:mo><mml:mpadded width="+3.3pt"><mml:mfrac><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:msubsup><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>log</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>The above is equivalent to the optimization problem:</p>
<disp-formula id="S10.E42"><label>(A13&#x2032;)</label><mml:math id="M43"><mml:mrow><mml:mrow><mml:munder><mml:mtext>min</mml:mtext><mml:mrow><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi></mml:mrow></mml:munder><mml:mpadded width="+3.3pt"><mml:mfrac><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:msubsup><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi mathvariant="normal">&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:mo rspace="5.8pt">}</mml:mo></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>log</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Consider that the solution of the former is <inline-formula><mml:math id="INEQ158"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mmultiscripts><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi><mml:none/><mml:none/><mml:mi>T</mml:mi></mml:mmultiscripts><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math></inline-formula>. Note that <inline-formula><mml:math id="INEQ159"><mml:mrow><mml:msub><mml:mo>min</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:msub><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>j</mml:mi></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:msubsup><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mtext mathvariant="bold">x</mml:mtext><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:math></inline-formula> can be rewritten in matrix form as <inline-formula><mml:math id="INEQ160"><mml:mrow><mml:msub><mml:mo>min</mml:mo><mml:mi mathvariant="bold">W</mml:mi></mml:msub><mml:msubsup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>-</mml:mo><mml:mrow><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext></mml:mrow></mml:mrow><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>. Therefore, problem (A13&#x2032;) can be transformed as</p>
<disp-formula id="S10.E43"><label>(A14&#x2032;)</label><mml:math id="M44"><mml:mrow><mml:mpadded width="+5pt"><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mrow><mml:mi mathvariant="bold">W</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi></mml:mrow></mml:munder></mml:mpadded><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:mtext mathvariant="bold">W</mml:mtext><mml:mo>-</mml:mo><mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mtext mathvariant="bold">X</mml:mtext><mml:mo>|</mml:mo><mml:mpadded width="+3.3pt"><mml:msubsup><mml:mo>|</mml:mo><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:msup><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mtext mathvariant="bold">W</mml:mtext><mml:mi>T</mml:mi></mml:msup><mml:mo rspace="5.8pt">)</mml:mo></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mi>p</mml:mi><mml:mtext>log</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mo>|</mml:mo><mml:mi mathvariant="bold">&#x03A9;</mml:mi><mml:mo>|</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
</sec>
</app>
</app-group>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ariana</surname> <given-names>A.</given-names></name> <name><surname>Cohen</surname> <given-names>M. S.</given-names></name></person-group> (<year>2013</year>). <article-title>Decreased small-world functional network connectivity and clustering across resting state networks in Schizophrenia: an fMRI classification tutorial.</article-title> <source><italic>Front. Human Neurosci.</italic></source> <volume>7</volume>:<issue>520</issue>. <pub-id pub-id-type="doi">10.3389/fnhum.2013.00520</pub-id> <pub-id pub-id-type="pmid">24032010</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baggio</surname> <given-names>H. C.</given-names></name> <name><surname>Sala-Llonch</surname> <given-names>R.</given-names></name> <name><surname>Segura</surname> <given-names>B.</given-names></name> <name><surname>Marti</surname> <given-names>M.-J.</given-names></name> <name><surname>Valldeoriola</surname> <given-names>F.</given-names></name> <name><surname>Compta</surname> <given-names>Y.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Functional brain networks and cognitive deficits in Parkinson&#x2019;s disease.</article-title> <source><italic>Human Brain Mapp.</italic></source> <volume>35</volume> <fpage>4620</fpage>&#x2013;<lpage>4634</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.22499</pub-id> <pub-id pub-id-type="pmid">24639411</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bezdek</surname> <given-names>J. C.</given-names></name> <name><surname>Hathaway</surname> <given-names>R. J.</given-names></name></person-group> (<year>2002</year>). <article-title>Some notes on alternating optimization.</article-title> <source><italic>Lect. Notes Comp. Sci.</italic></source> <volume>2275</volume> <fpage>187</fpage>&#x2013;<lpage>195</lpage>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Biswal</surname> <given-names>B.</given-names></name> <name><surname>Yetkin</surname> <given-names>F. Z.</given-names></name> <name><surname>Haughton</surname> <given-names>V. M.</given-names></name> <name><surname>Hyde</surname> <given-names>J. S.</given-names></name></person-group> (<year>1995</year>). <article-title>Functional connectivity in the motor cortex of resting human brain using echo-planar mri.</article-title> <source><italic>Magn. Reson. Med.</italic></source> <volume>34</volume> <fpage>537</fpage>&#x2013;<lpage>541</lpage>. <pub-id pub-id-type="doi">10.1002/mrm.1910340409</pub-id> <pub-id pub-id-type="pmid">8524021</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>C.-C.</given-names></name> <name><surname>Lin</surname> <given-names>C.-J.</given-names></name></person-group> (<year>2011</year>). <article-title>LIBSVM: a library for support vector machines.</article-title> <source><italic>ACM Trans. Intell. Syst. Technol.</italic></source> <volume>2</volume> <fpage>1</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1145/1961189.1961199</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Gao</surname> <given-names>Y.</given-names></name> <name><surname>Wee</surname> <given-names>C.-Y.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>High-order resting-state functional connectivity network for MCI classification.</article-title> <source><italic>Human Brain Mapp.</italic></source> <volume>37</volume> <fpage>3282</fpage>&#x2013;<lpage>3296</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.23240</pub-id> <pub-id pub-id-type="pmid">27144538</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dadi</surname> <given-names>K.</given-names></name> <name><surname>Rahim</surname> <given-names>M.</given-names></name> <name><surname>Abraham</surname> <given-names>A.</given-names></name> <name><surname>Chyzhyk</surname> <given-names>D.</given-names></name> <name><surname>Milham</surname> <given-names>M.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name><etal/></person-group> (<year>2019</year>). <article-title>Benchmarking functional connectome-based predictive models for resting-state fMRI.</article-title> <source><italic>NeuroImage</italic></source> <volume>192</volume> <fpage>115</fpage>&#x2013;<lpage>134</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2019.02.062</pub-id> <pub-id pub-id-type="pmid">30836146</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dosenbach</surname> <given-names>N. U. F.</given-names></name> <name><surname>Nardos</surname> <given-names>B.</given-names></name> <name><surname>Cohen</surname> <given-names>A. L.</given-names></name> <name><surname>Fair</surname> <given-names>D. A.</given-names></name> <name><surname>Power</surname> <given-names>J. D.</given-names></name> <name><surname>Church</surname> <given-names>J. A.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>Prediction of individual brain maturity using fMRI.</article-title> <source><italic>Science</italic></source> <volume>329</volume> <fpage>1358</fpage>&#x2013;<lpage>1361</lpage>. <pub-id pub-id-type="doi">10.1126/science.1194144</pub-id> <pub-id pub-id-type="pmid">20829489</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dutilleul</surname> <given-names>P.</given-names></name></person-group> (<year>1999</year>). <article-title>The MLE algorithm for the matrix normal distribution.</article-title> <source><italic>J. Stat. Comp. Simul.</italic></source> <volume>64</volume> <fpage>105</fpage>&#x2013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1080/00949659908811970</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eguiluz</surname> <given-names>V. M.</given-names></name> <name><surname>Chialvo</surname> <given-names>D. R.</given-names></name> <name><surname>Cecchi</surname> <given-names>G. A.</given-names></name> <name><surname>Baliki</surname> <given-names>M.</given-names></name> <name><surname>Apkarian</surname> <given-names>A. V.</given-names></name></person-group> (<year>2005</year>). <article-title>Scale-free brain functional networks.</article-title> <source><italic>Phys. Rev. Lett.</italic></source> <volume>94</volume>:<issue>18102</issue>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.94.018102</pub-id> <pub-id pub-id-type="pmid">15698136</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friedman</surname> <given-names>J. H.</given-names></name> <name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Sparse inverse covariance estimation with the graphical lasso.</article-title> <source><italic>Biostatistics</italic></source> <volume>9</volume> <fpage>432</fpage>&#x2013;<lpage>441</lpage>. <pub-id pub-id-type="doi">10.1093/biostatistics/kxm045</pub-id> <pub-id pub-id-type="pmid">18079126</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Harrison</surname> <given-names>L.</given-names></name> <name><surname>Penny</surname> <given-names>W.</given-names></name></person-group> (<year>2003</year>). <article-title>Dynamic causal modelling.</article-title> <source><italic>NeuroImage</italic></source> <volume>19</volume> <fpage>1273</fpage>&#x2013;<lpage>1302</lpage>. <pub-id pub-id-type="doi">10.1016/s1053-8119(03)00202-7</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Friston</surname> <given-names>K. J.</given-names></name> <name><surname>Williams</surname> <given-names>S.</given-names></name> <name><surname>Howard</surname> <given-names>R.</given-names></name> <name><surname>Frackowiak</surname> <given-names>R. S.</given-names></name> <name><surname>Turner</surname> <given-names>R.</given-names></name></person-group> (<year>1996</year>). <article-title>Movement-related effects in fMRI time-series.</article-title> <source><italic>Magn. Reson. Med.</italic></source> <volume>35</volume> <fpage>346</fpage>&#x2013;<lpage>355</lpage>. <pub-id pub-id-type="doi">10.1002/mrm.1910350312</pub-id> <pub-id pub-id-type="pmid">8699946</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gupta</surname> <given-names>A. K.</given-names></name> <name><surname>Nagar</surname> <given-names>D. K.</given-names></name></person-group> (<year>2000</year>). <source><italic>Matrix Variate Distributions.</italic></source> <publisher-loc>London</publisher-loc>: <publisher-name>Chapman and Hall</publisher-name>.</citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hahn</surname> <given-names>K.</given-names></name> <name><surname>Myers</surname> <given-names>N.</given-names></name> <name><surname>Prigarin</surname> <given-names>S.</given-names></name> <name><surname>Rodenacker</surname> <given-names>K.</given-names></name> <name><surname>Kurz</surname> <given-names>A.</given-names></name> <name><surname>F&#x00F6;rstl</surname> <given-names>H.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Selectively and progressively disrupted structural connectivity of functional brain networks in Alzheimer&#x2019;s disease &#x2014; Revealed by a novel framework to analyze edge distributions of networks detecting disruptions with strong statistical evidence.</article-title> <source><italic>NeuroImage</italic></source> <volume>81</volume> <fpage>96</fpage>&#x2013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.05.011</pub-id> <pub-id pub-id-type="pmid">23668966</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hashem</surname> <given-names>S.</given-names></name> <name><surname>Nisar</surname> <given-names>S.</given-names></name> <name><surname>Bhat</surname> <given-names>A. A.</given-names></name> <name><surname>Yadav</surname> <given-names>S. K.</given-names></name> <name><surname>Azeem</surname> <given-names>M. W.</given-names></name> <name><surname>Bagga</surname> <given-names>P.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Genetics of structural and functional brain changes in autism spectrum disorder.</article-title> <source><italic>Transl. Psychiatry</italic></source> <volume>10</volume> <fpage>229</fpage>&#x2013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1038/s41398-020-00921-3</pub-id> <pub-id pub-id-type="pmid">32661244</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Qiao</surname> <given-names>L.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Estimating functional connectivity networks via low-rank tensor approximation with applications to MCI identification.</article-title> <source><italic>IEEE Trans. Bio-Med. Eng.</italic></source> <volume>67</volume> <fpage>1912</fpage>&#x2013;<lpage>1920</lpage>. <pub-id pub-id-type="doi">10.1109/TBME.2019.2950712</pub-id> <pub-id pub-id-type="pmid">31675312</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jie</surname> <given-names>P.</given-names></name> <name><surname>Pei</surname> <given-names>W.</given-names></name> <name><surname>Zhou</surname> <given-names>N.</given-names></name> <name><surname>Ji</surname> <given-names>Z.</given-names></name></person-group> (<year>2009</year>). <article-title>Partial correlation estimation by joint sparse regression models.</article-title> <source><italic>J. Am. Stat. Assoc.</italic></source> <volume>104</volume> <fpage>735</fpage>&#x2013;<lpage>746</lpage>. <pub-id pub-id-type="doi">10.1198/jasa.2009.0126</pub-id> <pub-id pub-id-type="pmid">19881892</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lau</surname> <given-names>W. K. W.</given-names></name> <name><surname>Leung</surname> <given-names>M.-K.</given-names></name> <name><surname>Zhang</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Hypofunctional connectivity between the posterior cingulate cortex and ventromedial prefrontal cortex in autism: Evidence from coordinate-based imaging meta-analysis.</article-title> <source><italic>Prog. Neuro-Psychopharm. Biol. Psychiatry</italic></source> <volume>103</volume>:<issue>109986</issue>. <pub-id pub-id-type="doi">10.1016/j.pnpbp.2020.109986</pub-id> <pub-id pub-id-type="pmid">32473190</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Lee</surname> <given-names>D. S.</given-names></name> <name><surname>Kang</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>B.</given-names></name> <name><surname>Chung</surname> <given-names>M. K.</given-names></name></person-group> (<year>2011</year>). <article-title>Sparse brain network recovery under compressed sensing.</article-title> <source><italic>IEEE Trans. Med. Imag.</italic></source> <volume>30</volume> <fpage>1154</fpage>&#x2013;<lpage>1165</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2011.2140380</pub-id> <pub-id pub-id-type="pmid">21478072</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Qiao</surname> <given-names>L.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Remodeling pearson&#x2019;s correlation for functional brain network eetimation and autism spectrum disorder identification.</article-title> <source><italic>Front. Neuroinform.</italic></source> <volume>11</volume>:<issue>55</issue>. <pub-id pub-id-type="doi">10.3389/fninf.2017.00055</pub-id> <pub-id pub-id-type="pmid">28912708</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>H.</given-names></name> <name><surname>Yan</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Castellanos</surname> <given-names>F. X.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Altered resting-state dynamic functional brain networks in major depressive disorder: findings from the REST-meta-MDD consortium.</article-title> <source><italic>NeuroImage. Clin.</italic></source> <volume>26</volume>:<issue>102163</issue>. <pub-id pub-id-type="doi">10.1016/j.nicl.2020.102163</pub-id> <pub-id pub-id-type="pmid">31953148</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mahjoub</surname> <given-names>I.</given-names></name> <name><surname>Mahjoub</surname> <given-names>M. A.</given-names></name> <name><surname>Rekik</surname> <given-names>I.</given-names></name> <name><surname>Weiner</surname> <given-names>M.</given-names></name> <name><surname>Aisen</surname> <given-names>P.</given-names></name> <name><surname>Petersen</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2018</year>). <article-title>Brain multiplexes reveal morphological connectional biomarkers fingerprinting late brain dementia states.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>8</volume>:<issue>4103</issue>. <pub-id pub-id-type="doi">10.1038/s41598-018-21568-7</pub-id> <pub-id pub-id-type="pmid">29515158</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marrelec</surname> <given-names>G.</given-names></name> <name><surname>Krainik</surname> <given-names>A.</given-names></name> <name><surname>Duffau</surname> <given-names>H.</given-names></name> <name><surname>P&#x00E9;l&#x00E9;grini-Issac</surname> <given-names>M.</given-names></name> <name><surname>Leh&#x00E9;ricy</surname> <given-names>S.</given-names></name> <name><surname>Doyon</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2006</year>). <article-title>Partial correlation for functional brain interactivity investigation in functional MRI.</article-title> <source><italic>NeuroImage</italic></source> <volume>32</volume> <fpage>228</fpage>&#x2013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2005.12.057</pub-id> <pub-id pub-id-type="pmid">16777436</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mclntosh</surname> <given-names>A. R.</given-names></name> <name><surname>Gonzalez-Lima</surname> <given-names>F.</given-names></name></person-group> (<year>1994</year>). <article-title>Structural equation modeling and its application to network analysis in functional brain imaging.</article-title> <source><italic>Human Brain Mapp.</italic></source> <volume>2</volume> <fpage>2</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.460020104</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ng</surname> <given-names>A. Y.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name> <name><surname>Weiss</surname> <given-names>Y.</given-names></name></person-group> (<year>2001</year>). <article-title>On spectral clustering: analysis and an algorithm.</article-title> <source><italic>Adv. Neural Inform. Proc. Syst.</italic></source> <volume>14</volume> <fpage>849</fpage>&#x2014;-<lpage>856</lpage>.</citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nickl-Jockschat</surname> <given-names>T.</given-names></name> <name><surname>Habel</surname> <given-names>U.</given-names></name> <name><surname>Michel</surname> <given-names>T. M.</given-names></name> <name><surname>Manning</surname> <given-names>J.</given-names></name> <name><surname>Laird</surname> <given-names>A. R.</given-names></name> <name><surname>Fox</surname> <given-names>P. T.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Brain structure anomalies in autism spectrum disorder&#x2013;a meta-analysis of VBM studies using anatomic likelihood estimation.</article-title> <source><italic>Human Brain Mapp.</italic></source> <volume>33</volume> <fpage>1470</fpage>&#x2013;<lpage>1489</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.21299</pub-id> <pub-id pub-id-type="pmid">21692142</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pervaiz</surname> <given-names>U.</given-names></name> <name><surname>Vidaurre</surname> <given-names>D.</given-names></name> <name><surname>Woolrich</surname> <given-names>M. W.</given-names></name> <name><surname>Smith</surname> <given-names>S. M.</given-names></name></person-group> (<year>2020</year>). <article-title>Optimising network modelling methods for fMRI.</article-title> <source><italic>NeuroImage</italic></source> <volume>211</volume>:<issue>116604</issue>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2020.116604</pub-id> <pub-id pub-id-type="pmid">32062083</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiao</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name> <name><surname>Teng</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Estimating functional brain networks by incorporating a modularity prior.</article-title> <source><italic>NeuroImage</italic></source> <volume>141</volume> <fpage>399</fpage>&#x2013;<lpage>407</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2016.07.058</pub-id> <pub-id pub-id-type="pmid">27485752</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>S. M.</given-names></name> <name><surname>Vidaurre</surname> <given-names>D.</given-names></name> <name><surname>Beckmann</surname> <given-names>C. F.</given-names></name> <name><surname>Glasser</surname> <given-names>M. F.</given-names></name> <name><surname>Jenkinson</surname> <given-names>M.</given-names></name> <name><surname>Miller</surname> <given-names>K. L.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Functional connectomics from resting-state fMRI.</article-title> <source><italic>Trends Cogn. Sci.</italic></source> <volume>17</volume> <fpage>666</fpage>&#x2013;<lpage>682</lpage>.</citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Sparse multiview task-centralized ensemble learning for ASD diagnosis based on age- and sex-related functional connectivity patterns.</article-title> <source><italic>IEEE Trans. Cybernet.</italic></source> <volume>49</volume> <fpage>3141</fpage>&#x2013;<lpage>3154</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2018.2839693</pub-id> <pub-id pub-id-type="pmid">29994137</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Multi-class ASD classification based on functional connectivity and functional correlation tensor via multi-source domain adaptation and multi-view sparse representation.</article-title> <source><italic>IEEE Trans. Med. Imag.</italic></source> <volume>39</volume> <fpage>3137</fpage>&#x2013;<lpage>3147</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2020.2987817</pub-id> <pub-id pub-id-type="pmid">32305905</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>An</surname> <given-names>L.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Wei</surname> <given-names>Z.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Connectivity strength-weighted sparse group representation-based brain network construction for MCI classification.</article-title> <source><italic>Hum. Brain Mapp.</italic></source> <volume>38</volume> <fpage>2370</fpage>&#x2013;<lpage>2383</lpage>. <pub-id pub-id-type="doi">10.1002/hbm.23524</pub-id> <pub-id pub-id-type="pmid">28150897</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H. T.</given-names></name></person-group> (<year>2010</year>). <article-title>The benefit of group sparsity.</article-title> <source><italic>Ann. Stat.</italic></source> <volume>38</volume> <fpage>1978</fpage>&#x2013;<lpage>2004</lpage>.</citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Shi</surname> <given-names>F.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name> <name><surname>Giannakopoulos</surname> <given-names>P.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Topographical information-based high-order functional connectivity and its application in abnormality detection for Mild Cognitive Impairment.</article-title> <source><italic>J. Alzheimer&#x2019;s Dis.</italic></source> <volume>54</volume> <fpage>1095</fpage>&#x2013;<lpage>1112</lpage>. <pub-id pub-id-type="doi">10.3233/JAD-160092</pub-id> <pub-id pub-id-type="pmid">27567817</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Learning multiple tasks with a sparse matrix-normal penalty.</article-title> <source><italic>Adv. Neural Inform. Proc. Syst.</italic></source> <volume>23</volume> <fpage>2550</fpage>&#x2013;<lpage>2558</lpage>.</citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Lee</surname> <given-names>S. W.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Hybrid high-order functional connectivity networks using resting-state functional MRI for Mild Cognitive Impairment diagnosis.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>7</volume>:<issue>6530</issue>. <pub-id pub-id-type="doi">10.1038/s41598-017-06509-0</pub-id> <pub-id pub-id-type="pmid">28747782</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Qiao</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2018a</year>). <article-title>Simultaneous estimation of low- and high-order functional connectivity for identifying Mild Cognitive Impairment.</article-title> <source><italic>Front. Neuroinform.</italic></source> <volume>12</volume>:<issue>3</issue>. <pub-id pub-id-type="doi">10.3389/fninf.2018.00003</pub-id> <pub-id pub-id-type="pmid">29467643</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Teng</surname> <given-names>S.</given-names></name> <name><surname>Qiao</surname> <given-names>L.</given-names></name> <name><surname>Shen</surname> <given-names>D.</given-names></name></person-group> (<year>2018b</year>). <article-title>Improving sparsity and modularity of high-order functional connectivity networks for MCI and ASD identification.</article-title> <source><italic>Front. Neurosci.</italic></source> <volume>12</volume>:<issue>959</issue>. <pub-id pub-id-type="doi">10.3389/fnins.2018.00959</pub-id> <pub-id pub-id-type="pmid">30618582</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="footnote1">
<label>1</label>
<p>Note that the rs-fMRI signals have been preprocessed as described in Section &#x201C;Data Acquisitions and Processing.&#x201D;</p></fn>
<fn id="footnote2">
<label>2</label>
<p><ext-link ext-link-type="uri" xlink:href="http://fcon_1000.projects.nitrc.org/indi/abide/">http://fcon_1000.projects.nitrc.org/indi/abide/</ext-link></p></fn>
<fn id="footnote3">
<label>3</label>
<p><ext-link ext-link-type="uri" xlink:href="http://rfmri.org/dpabi">http://rfmri.org/dpabi</ext-link></p></fn>
</fn-group>
</back>
</article>
