<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2022.839044</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>CRIA: An Interactive Gene Selection Algorithm for Cancers Prediction Based on Copy Number Variations</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Wu</surname> <given-names>Qiang</given-names></name>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Li</surname> <given-names>Dongxi</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1604735/overview"/>
</contrib>
</contrib-group>
<aff><institution>College of Data Science, Taiyuan University of Technology</institution>, <addr-line>Taiyuan</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Weihua Pan, Agricultural Genomics Institute at Shenzhen (CAAS), China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Lin Wang, Tianjin University of Science and Technology, China; Jiazhou Chen, South China University of Technology, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Dongxi Li <email>dxli0426&#x00040;126.com</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Plant Bioinformatics, a section of the journal Frontiers in Plant Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>03</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>839044</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>01</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Wu and Li.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Wu and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Genomic copy number variations (CNVs) are among the most important structural variations of genes found to be related to the risk of individual cancer and therefore they can be utilized to provide a clue to the research on the formation and progression of cancer. In this paper, an improved computational gene selection algorithm called CRIA (correlation-redundancy and interaction analysis based on gene selection algorithm) is introduced to screen genes that are closely related to cancer from the whole genome based on the value of gene CNVs. The CRIA algorithm mainly consists of two parts. Firstly, the main effect feature is selected out from the original feature set that has the largest correlation with the class label. Secondly, after the analysis involving correlation, redundancy and interaction for each feature in the candidate feature set, we choose the feature that maximizes the value of the custom selection criterion and add it into the selected feature set and then remove it from the candidate feature set in each selection round. Based on the real datasets, CRIA selects the top 200 genes to predict the type of cancer. The experiments&#x00027; results of our research show that, compared with the state-of-the-art related methods, the CRIA algorithm can extract the key features of CNVs and a better classification performance can be achieved based on them. In addition, the interpretable genes highly related to cancer can be known, which may provide new clues at the genetic level for the treatment of the cancer.</p></abstract>
<kwd-group>
<kwd>gene selection</kwd>
<kwd>correlation-redundancy analysis</kwd>
<kwd>interaction analysis</kwd>
<kwd>copula entropy</kwd>
<kwd>copy number variations (CNVs)</kwd>
<kwd>cancers prediction</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="8"/>
<equation-count count="35"/>
<ref-count count="48"/>
<page-count count="15"/>
<word-count count="9322"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>The occurrence of many diseases is associated with genome structural variations. Human genome variations include single nucleotide polymorphisms (SNPs), copy number variations (CNVs), etc. The copy number variations refer to the amplification, deletion, and more complex mutations in the genome of DNA fragments longer than 1 kb in length (Redon et al., <xref ref-type="bibr" rid="B36">2006</xref>). SNPs account for 0.5% of the human genome, and nearly 12% of the human genome often undergoes copy number variations (Redon et al., <xref ref-type="bibr" rid="B36">2006</xref>). Copy number variations have become an important genomic variation, and their role in the pathogenesis of complex human diseases is still being revealed.</p>
<p>The close relationship between CNVs and diseases has been widely recognized. Numerous studies have demonstrated that not a few human diseases involved copy number variations that could change the diploid status of particular locus of the genome (Zhang et al., <xref ref-type="bibr" rid="B47">2016</xref>). The Flierl research team found that the higher vulnerability of Parkinson&#x00027;s disease and stress sensitivity of neuronal precursor cells carry an &#x003B1;-synuclein gene triplication (Flierl et al., <xref ref-type="bibr" rid="B15">2014</xref>). Grangeon et al. (<xref ref-type="bibr" rid="B22">2021</xref>) discovered that early-onset cerebral amyloid angiopathy and Alzheimer Disease (AD) were related to an amyloid precursor protein (App) gene triple amplification. Breunis et al. (<xref ref-type="bibr" rid="B4">2008</xref>) reported that the copy number variations of FCGR2C gene promoted idiopathic thrombocytopenic purpura. Zheng et al. (<xref ref-type="bibr" rid="B48">2017</xref>) found that the low copy number of FCGR3B was associated with lupus nephritis in a Chinese population. And Pandey et al. (<xref ref-type="bibr" rid="B34">2015</xref>) revealed that there was both direct and indirect evidence suggesting abnormalities of glycogen synthase kinase (GSK)-3&#x003B2; and &#x003B2;-catenin in the pathophysiology of bipolar illness and possibly schizophrenia (SZ). Moreover, several neuro-developmental relevant genes, such as A2BP1, IMMP2, and AUTS2, were reported with mutational CNVs (Elia et al., <xref ref-type="bibr" rid="B12">2010</xref>). In 2006, a research team composed of researchers from the United Kingdom, Japan, the United States, Canada and other countries studied 270 individuals in 4 groups of the HapMap project, and constructed the first-generation copy number variations map of the human genome, and obtained 144 CNVs region (about 12% of the size of the human genome). Among them, 285 CNVs regions were related to the occurrence of known diseases (Redon et al., <xref ref-type="bibr" rid="B36">2006</xref>). Compared with SNPs, CNVs regions contained more DNA sequences, disease sites and functional elements, which could provide more clues for disease research. The publication of this map has become an important tool for studying the complex structural variations of the human genome and human diseases.</p>
<p>Cancer is a kind of diseases which involves uncontrolled abnormal cell growth and can spread to other tissues (Du and Elemento, <xref ref-type="bibr" rid="B11">2015</xref>). The formation and development of cancer are also associated with copy number variations (Frank et al., <xref ref-type="bibr" rid="B17">2007</xref>). Van Bockstal et al. (<xref ref-type="bibr" rid="B42">2020</xref>) discovered that HER2 gene amplification had a relationship with a bad result in invasive breast cancer and the amplification of heterogeneous HER2 had been described in 5&#x02013;41% of breast cancer. The experimental results of Buchynska et al. (<xref ref-type="bibr" rid="B5">2019</xref>) shown that assessment of copy number variations of HER-2/neu, c-MYC and CCNE1 genes revealed their amplification in the tumors of 18.8, 25.0 and 14.3% of endometrial cancer patients, respectively. Heo et al. (<xref ref-type="bibr" rid="B25">2020</xref>) pointed out that CNVs were related to the mechanism of lung cancer development through a comparative experiment. Moreover, Tian et al. (<xref ref-type="bibr" rid="B41">2020</xref>) found that CNVs of CYLD, USP9X and USP11 were significantly associated with the risk of colorectal cancer. A latest global cancer burden data released by the International Agency for Research on Cancer(IARC) of the WHO showed that the number of patients with new cancer and cancer deaths in China ranked first around the world with 4.57 million patients with new cancer and 3 million cancer deaths, accounting for 23.7 and 30%, respectively. It is of great significance to investigate cancer causes and its treatment. Because the gene expression patterns in cancer tumor have high specificity (Liang et al., <xref ref-type="bibr" rid="B30">2020</xref>), studying the relationship between these genetic information and cancer can provide a new idea for investigating the causes of cancer and help in early cancer diagnosis.</p>
<p>However, few studies have utilized machine learning (ML) or deep learning (DL) methods to use copy number variations data for the prediction of various cancer types. Zhang et al. (<xref ref-type="bibr" rid="B47">2016</xref>) used the mRMR and IFS methods to select 19 features from the 24,174 gene features of the copy number variations data set, which contained a total of 3,480 samples of 6 cancer types. They applied the Dagging algorithm with ten-fold cross-validation to classify cancer. But the accuracy of final result only reached 75%. Liang et al. (<xref ref-type="bibr" rid="B30">2020</xref>) used CNA_origin for cancer classification on the same data set. CNA_origin was an intelligent combined deep learning network, which was composed of two parts&#x02014;a stacked autoencoder and a one-dimensional convolutional neural network with multiscale convolutional kernels. CNA_origin eventually had an overall accuracy of 83.81% on ten-fold cross-validation. But it could not identify which gene features were more important and more closely associated with cancer classification.</p>
<p>Here, we present an improved novel computational algorithm named CRIA, which can successfully classify cancer based on the information of gene CNVs levels from the same dataset. CRIA can not only effectively perform dimensionality reduction operation on high-dimensional gene CNVs data, which can improve the efficiency of the experiment, but also selects specific gene features closely related to cancer, making it clear which genes are more important in cancer classification. And the final results had higher classification accuracy than the state-of-the-art methods.</p>
<p>The rest sections of this paper are structured as follows: Section Background describes the theoretical background and related work. Section The Proposed Method-CRIA introduces the collection of CNVs dataset, the implementation details and performance of the proposed algorithm. Section Results and Discussions demonstrates the experimental results on CNVs dataset and the performance comparison with the recent methods. In section Conclusions, we summarize the conclusions and point out our future work.</p></sec>
<sec id="s2">
<title>Background</title>
<p>In section Information Theory, we introduce some basic information theory knowledge, which is the core of our proposed algorithm. Before proposing our algorithm, we summarize some related work on gene selection methods and point out their drawbacks in section Related Work.</p>
<sec>
<title>Information Theory</title>
<p>As early as 1948, Shannon&#x00027;s information theory had been proposed (Shannon, <xref ref-type="bibr" rid="B38">2001</xref>), providing an effective method for measuring random variables&#x00027; information. The entropy can be understood as a measure of the uncertainty of a random variable (Cover and Thomas, <xref ref-type="bibr" rid="B10">1991</xref>). The greater the entropy of a random variable, the greater its uncertainty. If <italic>X</italic> &#x0003D; {<italic>x</italic><sub>1</sub>, <italic>x</italic><sub>2</sub>, &#x02026;, <italic>x</italic><sub><italic>l</italic></sub>} is a discrete random variable, its probability distribution is <italic>p</italic>(<italic>x</italic>) &#x0003D; <italic>P</italic>(<italic>X</italic> &#x0003D; <italic>x</italic>), <italic>x</italic> &#x02208; <italic>X</italic>. The entropy of <italic>X</italic> is defined as:
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>p</italic>(<italic>x</italic><sub><italic>i</italic></sub>) is the probability of <italic>x</italic><sub><italic>i</italic></sub>. Here the base of log is 2 and specified that 0 log 0 &#x0003D; 0.</p>
<p>If <italic>Y</italic> &#x0003D; {<italic>y</italic><sub>1</sub>, <italic>y</italic><sub>2</sub>, &#x02026;, <italic>y</italic><sub><italic>m</italic></sub>} is a discrete random variable, <italic>p</italic>(<italic>x</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>j</italic></sub>) is the joint probability of <italic>X</italic> and <italic>Y</italic>. Then, their joint entropy is defined as:
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
If the random variable <italic>X</italic> is in a given situation, the uncertainty measure of the variable <italic>Y</italic> can be defined by conditional entropy as follows:
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>p</italic>(<italic>y</italic><sub><italic>j</italic></sub>|<italic>x</italic><sub><italic>i</italic></sub>) is the conditional probability of <italic>Y</italic> under the condition of <italic>X</italic>.</p>
<p><bold>Definition 1:</bold> Mutual information (MI) (Cover and Thomas, <xref ref-type="bibr" rid="B10">1991</xref>) is a measure of useful information in information theory. It can be regarded as the amount of information shared by two random variables. MI can be defined as:
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><bold>Definition 2:</bold> Conditional mutual information (CMI) (Cover and Thomas, <xref ref-type="bibr" rid="B10">1991</xref>) can be defined as the amount of information that shared by variables <italic>X</italic> and <italic>Y</italic>, if a discrete random variable <italic>Z</italic> &#x0003D; {<italic>z</italic><sub>1</sub>, <italic>z</italic><sub>2</sub>, &#x02026;, <italic>z</italic><sub><italic>n</italic></sub>} is known.
<disp-formula id="E6"><label>(5)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><bold>Definition 3:</bold> Joint mutual information (JMI) (Cover and Thomas, <xref ref-type="bibr" rid="B10">1991</xref>) measures the amount of information shared by a joint random variable (<italic>X</italic><sub>1</sub>, <italic>X</italic><sub>2</sub>, &#x022EF;<italic>X</italic><sub><italic>q</italic></sub>) and Y and it can be defined as:
<disp-formula id="E7"><label>(6)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mo>&#x022EF;</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>y</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>Y</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><bold>Definition 4:</bold> Interaction gain (IG) had been introduced by Jakulin (<xref ref-type="bibr" rid="B26">2003</xref>), Jakulin and Bratko (<xref ref-type="bibr" rid="B27">2004</xref>) to measure the amount of information shared by three random variables at the same time. Mutual information can be regarded as a two-way interaction gain. IG is defined as follows:
<disp-formula id="E8"><label>(7)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo>;</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Y</mml:mi><mml:mo>;</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mo>;</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi><mml:mo>;</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>;</mml:mo><mml:mi>Z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></sec>
<sec>
<title>Related Work</title>
<p>The irrelevant features and redundant features existed in high-dimensional data will damage the performance of the learning algorithm and reduce the efficiency of the learning algorithm. Therefore, the dimensionality reduction of features is one of the most common methods of data preprocessing (Orsenigo and Vercellis, <xref ref-type="bibr" rid="B33">2013</xref>) and its purpose is to reduce the training time of the algorithm and improve the accuracy of final results (Bennasar et al., <xref ref-type="bibr" rid="B2">2015</xref>). In recent years, the research of gene selection methods based on mutual information has received wide attention from scholars. Best individual gene selection (BIF) (Chandrashekar and Sahin, <xref ref-type="bibr" rid="B7">2014</xref>) is the simplest and fastest filtering gene selection algorithm, especially suitable for high-dimensional data.</p>
<p>Battiti utilized the mutual information (MI) between features and class labels [<italic>I</italic>(<italic>f</italic><sub><italic>i</italic></sub>; <italic>c</italic>)] to measure the relevance and the mutual information between features [<italic>I</italic>(<italic>f</italic><sub><italic>i</italic></sub>; <italic>f</italic><sub><italic>s</italic></sub>)] to measure the redundancy (Battiti, <xref ref-type="bibr" rid="B1">1994</xref>). He proposed the Mutual Information Gene selection (MIFS) criterion and it is defined as:
<disp-formula id="E9"><label>(8)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where F is the original feature set, &#x003A9;<sub><italic>S</italic></sub> is the selected feature subset, F &#x02212; &#x003A9;<sub><italic>S</italic></sub> is the candidate feature subset and <italic>c</italic> is the class label. &#x003B2; is a configurable parameter to determine the trade&#x02013;off between relevance and redundancy. However, &#x003B2; is set experimentally, which results in an unstable performance.</p>
<p>Peng et al. (<xref ref-type="bibr" rid="B35">2005</xref>) proposed the Minimum-Redundancy Maximum-Relevance (MRMR) criterion and its evaluation function is defined as:
<disp-formula id="E10"><label>(9)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where |<italic>n</italic><sub><italic>s</italic></sub>| is the number of selected features.</p>
<p>Similarly, other gene selection methods that consider relevance between features and the class label and redundancy between features are concluded, such as Normalized Mutual Information Gene selection (NMIFS) and Conditional Mutual Information (CMI), and they were proposed by Est&#x000E9;vez et al. (<xref ref-type="bibr" rid="B13">2009</xref>) and Liang et al. (<xref ref-type="bibr" rid="B29">2019</xref>) respectively. Their evaluation function are defined as follows:
<disp-formula id="E11"><label>(10)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo class="qopname">min</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E13"><label>(11)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>M</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>H</italic>(<italic>f</italic><sub><italic>i</italic></sub>) is the information entropy and <italic>H</italic>(<italic>f</italic><sub><italic>i</italic></sub>|<italic>c</italic>) is the conditional entropy.</p>
<p>Many gene selection algorithms based on information theory tend to use mutual information as a measure of relevance, which will bring a disadvantage that mutual information tends to select features with more discrete values (Foithong et al., <xref ref-type="bibr" rid="B16">2012</xref>). Thus, the symmetrical uncertainty (Witten and Frank, <xref ref-type="bibr" rid="B44">2002</xref>) (a normalized form of mutual information, <italic>SU</italic>) is adopted to solve this problem. The symmetrical uncertainty can be described as:
<disp-formula id="E15"><label>(12)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
The <italic>SU</italic> can redress the bias of mutual information as much as possible and scale its values to [0,1] by penalizing inputs with large entropies. It will make the performance of gene selection better. Same as MI, for any two features <italic>f</italic><sub><italic>i</italic>1</sub> and <italic>f</italic><sub><italic>i</italic>2</sub>, if <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic>1</sub>; <italic>c</italic>) &#x0003E; <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic>2</sub>; <italic>c</italic>), due to more information can be provided by the former, <italic>f</italic><sub><italic>i</italic>1</sub> and <italic>c</italic> are more relevant. If <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic>1</sub>; <italic>f</italic><sub><italic>s</italic></sub>) &#x0003E; <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic>2</sub>; <italic>f</italic><sub><italic>s</italic></sub>), owing to the information shared by <italic>f</italic><sub><italic>i</italic>1</sub> and <italic>f</italic><sub><italic>s</italic></sub> being more and providing less information, <italic>f</italic><sub><italic>i</italic>1</sub> and <italic>f</italic><sub><italic>s</italic></sub> have greater redundancy.</p>
<p>Additionally, these gene selection algorithms mentioned above fail to take the feature interaction into consideration. After relevance and redundancy analysis, one feature deemed useless may interact with other features to provide more useful information. Especially in complicated biology systems, molecules interacting with each other, they work together to express physiological and pathological changes. If we only consider relevance and redundancy but ignore the feature interaction in data analysis, we may miss some useful features and affect the analysis results (Chen et al., <xref ref-type="bibr" rid="B8">2015</xref>).</p>
<p>Sun et al. (<xref ref-type="bibr" rid="B40">2013</xref>), Zeng et al. (<xref ref-type="bibr" rid="B46">2015</xref>), and Gu et al. (<xref ref-type="bibr" rid="B23">2020</xref>), respectively proposed a gene selection method using dynamic feature weights: Dynamic Weighting-based Gene selection algorithm (DWFS), Interaction Weight based Gene selection algorithm (IWFS) and Redundancy Analysis and Interaction Weight-based gene selection algorithm (RAIW). All of them employ the symmetric uncertainty to measure the relevance between features and the class label, and exploit the three-dimensional interaction information (mentioned at <bold>Information Theory Definition 4</bold>) to measure the interaction between two features and the class label. The evaluation functions are defined as follow:
<disp-formula id="E16"><label>(13)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E17"><label>(14)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E18"><label>(15)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>A</mml:mi><mml:mi>I</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>A</mml:mi><mml:mi>I</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>&#x003F5;</mml:mi><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>w</italic>(<italic>f</italic><sub><italic>i</italic></sub>) is the weight of each feature and its initial value is set to 1, &#x003B1; is a redundancy coefficient and the value is relevant to the number of dataset&#x00027;s features, <italic>f</italic><sub><italic>s</italic></sub> is one of features in the selected feature subset. In each round, the feature weight <italic>w</italic>(<italic>f</italic><sub><italic>i</italic></sub>) is updated by their interaction weight factors.
<disp-formula id="E20"><label>(16)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E21"><label>(17)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>I</mml:mi><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mi>W</mml:mi><mml:mi>F</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E23"><label>(18)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>A</mml:mi><mml:mi>I</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>A</mml:mi><mml:mi>I</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mtext>&#x000A0;</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>I</mml:mi><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>A</mml:mi><mml:mi>I</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M25"><mml:mi>w</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the feature weight of the previous round, <italic>I</italic>(<italic>f</italic><sub><italic>i</italic></sub>; <italic>c</italic>|<italic>f</italic><sub><italic>s</italic></sub>) is the conditional mutual information of <italic>f</italic><sub><italic>i</italic></sub> and <italic>c</italic> when <italic>f</italic><sub><italic>s</italic></sub> is given. <italic>I</italic>(<italic>f</italic><sub><italic>i</italic></sub>; <italic>f</italic><sub><italic>s</italic></sub>; <italic>c</italic>) is three-dimensional interaction information. However, we can find that although DWFS and IWFS take into account relevance and interaction, they ignore the redundancy between features. Correlation, redundancy and interaction are all taken into account by RAIW, but there is a no reasonable value for &#x003B1; in a specific dataset.</p>
<p>Furthermore, some other gene selection methods about three-way mutual information are listed and their evaluation function are defined as follows, such as Composition of Feature Relevance (CFR) (Gao et al., <xref ref-type="bibr" rid="B20">2018a</xref>), Joint Mutual Information Maximization (JMIM) (Bennasar et al., <xref ref-type="bibr" rid="B2">2015</xref>), Dynamic Change of Selected Feature with the class (DCSF) (Gao et al., <xref ref-type="bibr" rid="B19">2018b</xref>) and Max-Relevance and Max-Independence (MRI) (Wang et al., <xref ref-type="bibr" rid="B43">2017</xref>).</p>
<disp-formula id="E26"><label>(19)</label><mml:math id="M27"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>F</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo><mml:mtext>&#x02009;&#x02009;</mml:mtext><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula>
<disp-formula id="E41"><label>(20)</label><mml:math id="M52"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>J</mml:mi><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>max</mml:mi><mml:mo stretchy='false'>[</mml:mo><mml:munder><mml:mrow><mml:mi>min</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mo stretchy='false'>(</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>]</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x02009;&#x02009;&#x02009;</mml:mtext><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula>
<disp-formula id="E42"><label>(21)</label><mml:math id="M53"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>C</mml:mi><mml:mi>S</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo><mml:mtext>&#x02009;&#x02009;&#x02009;</mml:mtext><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E43"><label>(22)</label><mml:math id="M54"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mrow></mml:munder><mml:mrow><mml:mi>I</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo><mml:mtext>&#x02009;&#x02009;&#x02009;</mml:mtext><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mo>&#x003A9;</mml:mo><mml:mi>S</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>I</italic>(<italic>f</italic><sub><italic>i</italic></sub>, <italic>f</italic><sub><italic>s</italic></sub>; <italic>c</italic>) is the joint mutual information of <italic>f</italic><sub><italic>i</italic></sub>, <italic>f</italic><sub><italic>s</italic></sub> and <italic>c</italic>. <italic>I</italic>(<italic>f</italic><sub><italic>s</italic></sub>; <italic>c</italic>|<italic>f</italic><sub><italic>i</italic></sub>) is the conditional mutual information of <italic>f</italic><sub><italic>s</italic></sub> and <italic>c</italic> when <italic>f</italic><sub><italic>i</italic></sub> is given. However, these algorithms only take into account three-way mutual information among features and the class label, and none of them considers relevance, redundancy and three-dimensional mutual information between features at the same time, which will affect the performance of these algorithms.</p></sec></sec>
<sec id="s3">
<title>The Proposed Method-CRIA</title>
<p>In section CNVs Dataset, we firstly introduce the collection of datasets and the process of data processing specifically. Subsequently, we redress other methods&#x00027; shortcomings and propose an improved gene selection algorithm called CRIA in section The Proposed Algorithm and give it a specific implementation in section Algorithm Implementation. Finally, in section Verify the Performance of CRIA, we verify the performance of CRIA by comparing the experimental results of CRIA and other 8 algorithms on 5 datasets.</p>
<sec>
<title>CNVs Dataset</title>
<p>The datasets of copy number variations in different cancer types used in this paper comes from the cBioPortal for Cancer Genomics (<ext-link ext-link-type="uri" xlink:href="http://cbio.mskcc.org/cancergenomics/pancan_tcga/">http://cbio.mskcc.org/cancergenomics/pancan_tcga/</ext-link>, Release 2/4/2013) (Cerami et al., <xref ref-type="bibr" rid="B6">2012</xref>; Ciriello et al., <xref ref-type="bibr" rid="B9">2013</xref>; Gao et al., <xref ref-type="bibr" rid="B18">2013</xref>). The copy number values in the dataset are generated by Affymetrix SNP 6.0 arrays for the set of samples in the cancer genome atlas (TCGA) study (Liang et al., <xref ref-type="bibr" rid="B30">2020</xref>). The preprocessing analysis of the dataset is performed with GISTIC (Beroukhim et al., <xref ref-type="bibr" rid="B3">2007</xref>). There are 11 cancer types in the cBioPortal database with the largest sample number was 847 and the smallest sample was 135. In order to avoid affecting the experimental results due to the large difference in the number of samples of cancer types, we only select six cancer types with more than 400 samples as our experimental data. The details of six cancer types are listed in <xref ref-type="table" rid="T1">Table 1</xref>, and totally there are 3480 samples in our experimental dataset.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>The number of samples for each cancer type in this dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Class label</bold></th>
<th valign="top" align="left"><bold>Histology</bold></th>
<th valign="top" align="center"><bold>Samples</bold></th>
<th valign="top" align="center"><bold>Percentage</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="left">UCEC (Uterine corpus endometrial carcinoma)</td>
<td valign="top" align="center">443</td>
<td valign="top" align="center">12.73%</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="left">KIRC (Kidney renal clear cell carcinoma)</td>
<td valign="top" align="center">490</td>
<td valign="top" align="center">14.08%</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">OV (Ovarian serous cystadenocarcinoma)</td>
<td valign="top" align="center">562</td>
<td valign="top" align="center">16.15%</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="left">GBM (Glioblastoma multiforme)</td>
<td valign="top" align="center">563</td>
<td valign="top" align="center">16.18%</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">COAD/READ (Colon adenocarcinoma/Rect-um adenocarcinoma)</td>
<td valign="top" align="center">575</td>
<td valign="top" align="center">16.52%</td>
</tr>
<tr>
<td valign="top" align="left">6</td>
<td valign="top" align="left">BRCA (Breast invasive carcinoma)</td>
<td valign="top" align="center">847</td>
<td valign="top" align="center">24.34%</td>
</tr>
<tr>
<td valign="top" align="left">Total</td>
<td/>
<td valign="top" align="center">3,480</td>
<td valign="top" align="center">100%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In this dataset, each sample consists of labels for 24174 genetic cytobands. The CNV spectrum is divided into five regions/labels by setting four thresholds in cancer algorithm (Mermel et al., <xref ref-type="bibr" rid="B32">2011</xref>). Then, the CNV values are discretized into 5 different values&#x02014;&#x0201C;-2,&#x0201D; &#x0201C;-1,&#x0201D; &#x0201C;0,&#x0201D; &#x0201C;1,&#x0201D; &#x0201C;2,&#x0201D; where &#x0201C;-2&#x0201D; denotes the deletion of both copies (possibly homozygous deletion), &#x0201C;&#x02212;1&#x0201D; means the deletion of one copy (possibly heterozygous deletion), &#x0201C;0&#x0201D; corresponds to exactly two copies, i.e., no gain/loss (diploid), &#x0201C;1&#x0201D; denotes a low-level copy number gain and &#x0201C;2&#x0201D; means a high-level copy number amplification (Ciriello et al., <xref ref-type="bibr" rid="B9">2013</xref>).</p>
<p>The CNVs values are preprocessed to the range of [&#x02212;1,1] with Equation (23).
<disp-formula id="E27"><label>(23)</label><mml:math id="M28"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:msup><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>val</italic> is the value of gene copy number variations of each sample, |<italic>val</italic>|<sub>max</sub> is the maximum absolute value of gene CNVs among samples and <italic>val</italic>&#x02032; is the recalculated value.</p></sec>
<sec>
<title>The Proposed Algorithm</title>
<p>In section Related Work, we analyze the 11 gene selection methods and point out their shortcomings. In view of the defects of these algorithms, we propose an improved gene selection algorithm to redress their shortcomings: Correlation-Redundancy and Interaction Analysis based gene selection algorithm (CRIA). This method uses the symmetric uncertainty (<italic>SU</italic>) to measure the correlation between features and the class label and the redundancy among features. In addition, copula entropy is introduced to measure the feature interaction information. Different from the three-way interaction of DWFS, IWFS and RAIW, the proposed algorithm considers the interaction between the candidate feature and the entire set of selected features, instead of being limited to the three-dimensional interaction.</p>
<p>As we know, Shannon&#x00027;s definition of mutual information aims at a pair of random variables, and it measures the correlation between two random variables. Therefore, naturally, many researchers have tried to study how to extend the definition of mutual information from two variables to multivariate situations. In 2011, Ma and Sun published a paper (Ma and Sun, <xref ref-type="bibr" rid="B31">2011</xref>), which contributed to the entropy of information theory. They defined a new concept of entropy in that paper, called Copula Entropy. Copula Entropy is defined on a set of random variables and conformed to symmetry. Therefore, it is a multivariate extension of mutual information, which can be utilized to measure the full-order, non-linear correlation among random variables. They proved the equivalence between copula entropy and the concept of mutual information, which was, mutual information was equal to negative copula entropy (Ma and Sun, <xref ref-type="bibr" rid="B31">2011</xref>).</p>
<p>The copula entropy of <inline-formula><mml:math id="M29"><mml:mover class="overset"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is defined as:
<disp-formula id="E28"><label>(24)</label><mml:math id="M30"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:mrow><mml:mo>&#x0222B;</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mover class="overset"><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow></mml:mstyle></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula><mml:math id="M31"><mml:mover class="overset"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:math></inline-formula> are random variables with marginal functions <inline-formula><mml:math id="M32"><mml:mover class="overset"><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and copula density <inline-formula><mml:math id="M33"><mml:mi>c</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p>
<p>Thus, we can use interaction factor <italic>IF</italic><sub><italic>CRIA</italic></sub>, which is defined in Equation (25) to measure the interaction between the candidate feature and the selected feature subset. The meaning of <italic>IF</italic><sub><italic>CRIA</italic></sub> is that, after adding a random candidate feature <italic>f</italic><sub><italic>i</italic></sub> into the selected feature subset &#x003A9;<sub><italic>S</italic></sub>, the amount of interaction information increased relative to the original selected feature subset. So, the bigger value of <italic>IF</italic><sub><italic>CRIA</italic></sub>, the bigger value of interaction between <italic>f</italic><sub><italic>i</italic></sub> and &#x003A9;<sub><italic>S</italic></sub>. In each round of calculation, we are supposed to choose the variable that maximizes the <italic>IF</italic><sub><italic>CRIA</italic></sub> value.
<disp-formula id="E29"><label>(25)</label><mml:math id="M34"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
Where <italic>c</italic> is the target class label.</p>
<p>Integrating the correlation between the features and the class label and the redundancy between features that we improved, and the interaction factor we proposed, we can define the evaluation criterion of a candidate feature as follows:
<disp-formula id="E30"><label>(26)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>I</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
For the equation (26), we can see that the proposed algorithm can take into account the relevance between the candidate feature and the class label, redundancy and multi-dimensional interaction among the candidate feature and the selected features at the same time. The formula <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic></sub>, <italic>c</italic>) can denote the relevance and <inline-formula><mml:math id="M36"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> calculates the redundancy. Also, the formula <inline-formula><mml:math id="M37"><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> denotes the interaction among the features.</p>
<p>According to the definition of copula entropy and Equation (24), there is a theorem.</p>
<p><bold>Theorem 1:</bold> The mutual information of random variables is equivalent to their negative copula entropy (Ma and Sun, <xref ref-type="bibr" rid="B31">2011</xref>):
<disp-formula id="E31"><label>(27)</label><mml:math id="M38"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover class="overset"><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>According to Theorem 1, the value of copula entropy can be calculated by the MI of multivariates. The definition of mutual information extended from two variables to multivariate is described as follows:
<disp-formula id="E32"><label>(28)</label><mml:math id="M39"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x0222C;</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mi>&#x0222C;</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mo>&#x0222B;</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>&#x0222B;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo class="qopname">&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo class="qopname">&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo class="qopname">&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>X</italic><sub><italic>m</italic></sub> &#x0003D; {<italic>x</italic><sub>1</sub>, <italic>x</italic><sub>2</sub>, &#x02026;, <italic>x</italic><sub><italic>m</italic>&#x02212;1</sub>, <italic>x</italic><sub><italic>m</italic></sub>} &#x0003D; {<italic>X</italic><sub><italic>m</italic>&#x02212;1</sub>, <italic>x</italic><sub><italic>m</italic></sub>}.</p>
<p>According to the Equation (28), we have:
<disp-formula id="E33"><label>(29)</label><mml:math id="M40"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
Therefore,
<disp-formula id="E34"><label>(30)</label><mml:math id="M41"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
According to Equation (26), (27) and (30), we have:
<disp-formula id="E35"><label>(31)</label><mml:math id="M42"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>I</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
Let &#x003A9;<sub><italic>S</italic></sub> &#x0003D; {<italic>f</italic><sub>1</sub>, <italic>f</italic><sub>2</sub>, &#x02026;, <italic>f</italic><sub><italic>m</italic></sub>}, Since,
<disp-formula id="E36"><label>(32)</label><mml:math id="M43"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
Therefore,
<disp-formula id="E37"><label>(33)</label><mml:math id="M44"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>F</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo stretchy="false">}</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
The general flow chart of the proposed algorithm is presented in <xref ref-type="fig" rid="F1">Figure 1</xref> we can see that an original feature set F is first given, from which we select the main effect feature that maximizes the value of (12). Then the main feature is put into the selected feature subset &#x003A9;<sub><italic>S</italic></sub>. For each feature in the candidate feature set, after conducting correlation and redundancy analysis, we are next supposed to use (25) to perform interaction analysis on it. Choose the feature that maximizes the value of (33), which then is put into the selected feature set. If the number of the selected features meets the threshold condition, the above steps will be executed again, otherwise the program ends directly.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>General flow chart of the proposed algorithm.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-839044-g0001.tif"/>
</fig></sec>
<sec>
<title>Algorithm Implementation</title>
<p>We propose a gene selection method based on correlation-redundancy and interaction analysis. The pseudo code of CRIA algorithm is described as follows.</p>
<p>Here, for the CNVs dataset, we set the value of the threshold <italic>M</italic> to be 200 to reduce the calculation time and avoid curse of dimensionality. In addition, we need to control the number of selected features to be same as the method proposed by Zhang et al. (<xref ref-type="bibr" rid="B47">2016</xref>).</p>
<p>The CRIA algorithm consists of two stages:</p>
<p>Stage 1 (lines 1&#x02013;7): In this part, the selected feature subset &#x003A9;<sub><italic>S</italic></sub> and the original feature set F are first initialized. For each feature in the original feature set <italic>f</italic><sub><italic>i</italic></sub>, the symmetrical uncertainty <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic></sub>; <italic>c</italic>) between <italic>f</italic><sub><italic>i</italic></sub> and class label <italic>c</italic> is calculated. The feature whose value of symmetrical uncertainty with class label is the maximum is selected out and added into the selected features subset &#x003A9;<sub><italic>S</italic></sub>, which we name &#x0201C;the main effect feature.&#x0201D;</p>
<p>Stage 2 (lines 8-18): The second stage mainly calculates the correlation measure <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic></sub>; <italic>c</italic>) and the redundancy measure <inline-formula><mml:math id="M48"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Then the interaction value <italic>IF</italic><sub><italic>CRIA</italic></sub> between &#x003A9;<sub><italic>S</italic></sub>, <italic>f</italic><sub><italic>i</italic></sub> and <italic>c</italic> is updated. <italic>J</italic><sub><italic>CRIA</italic></sub>(<italic>f</italic><sub><italic>i</italic></sub>) is calculated and the feature with the maximum value is added into the selected feature subset &#x003A9;<sub><italic>S</italic></sub>. This procedure terminates until the number of selected features is no less than predefined threshold <italic>M</italic>.</p>
<p>According to <xref ref-type="table" rid="T9">Algorithm 1</xref>, when the size of the feature subset reaches the set threshold <italic>M</italic>, the procedure will be terminated. The value of the threshold setting should be determined by different datasets. A small <italic>M</italic> can reduce the amount of calculation but may also lose many effective features that are useful; a large <italic>M</italic> will increase the amount of calculation but may improve the accuracy of final result (Foithong et al., <xref ref-type="bibr" rid="B16">2012</xref>). Actually, when the threshold exceeds a certain value, the accuracy of the final result will not only not increase, but may decrease, and it will bring computational complexity. The selected features are ranked according to the value of the evaluation function <italic>J</italic><sub><italic>CRIA</italic></sub>(<italic>f</italic><sub><italic>i</italic></sub>) from largest to smallest.</p>
<table-wrap position="float" id="T9">
<label>Algorithm 1</label>
<caption><p>CRIA: correlation-redundancy and interaction analysis based gene selection algorithm.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td valign="top" align="left"><bold>Input</bold> <italic>N</italic>: the number of original features, <italic>M</italic>: the number of features to be selected, <italic>n</italic><sub><italic>s</italic></sub>: the<break/> number of selected features.</td></tr>
<tr>
<td valign="top" align="left"><bold>Output</bold>: the selected feature subset (&#x003A9;<sub><italic>S</italic></sub> &#x02286; <italic>F</italic>).</td></tr>
<tr>
<td valign="top" align="left">1 First initializes &#x003A9;<sub><italic>S</italic></sub> &#x0003D; &#x02205;, <italic>F</italic> &#x0003D; {<italic>f</italic><sub>1</sub>, <italic>f</italic><sub>2</sub>, &#x02026;, <italic>f</italic><sub><italic>N</italic></sub>};</td></tr>
<tr>
<td valign="top" align="left">2 for each <italic>f</italic><sub><italic>i</italic></sub> &#x02208; <italic>F</italic> do:</td></tr>
<tr>
<td valign="top" align="left">3 calculate <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic></sub>, <italic>c</italic>);</td></tr>
<tr>
<td valign="top" align="left">4 end for</td></tr>
<tr>
<td valign="top" align="left">5 select the feature <italic>f</italic><sub><italic>i</italic> max</sub> &#x02208; &#x1D53D; with the largest value of <italic>SU</italic>(<italic>f</italic><sub><italic>i</italic></sub>, <italic>c</italic>);</td></tr>
<tr>
<td valign="top" align="left">6 &#x003A9;<sub><italic>S</italic></sub> &#x0003D; &#x003A9;<sub><italic>S</italic></sub> &#x0222A; {<italic>f</italic><sub><italic>i</italic> max</sub>};</td></tr>
<tr>
<td valign="top" align="left">7 <italic>F</italic> &#x0003D; <italic>F</italic> &#x02212; {<italic>f</italic><sub><italic>i</italic> max</sub>};</td></tr>
<tr>
<td valign="top" align="left">8 while <italic>n</italic><sub><italic>s</italic></sub> &#x02264; <italic>M</italic> do:</td></tr>
<tr>
<td valign="top" align="left">9 for <italic>f</italic><sub><italic>i</italic></sub> &#x02208; &#x1D53D; do:</td></tr>
<tr>
<td valign="top" align="left">10 calculate <inline-formula><mml:math id="M45"><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>;</td></tr>
<tr>
<td valign="top" align="left">11 calculate <inline-formula><mml:math id="M46"><mml:mi>I</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula>;</td></tr>
<tr>
<td valign="top" align="left">12 calculate <inline-formula><mml:math id="M47"><mml:msub><mml:mrow><mml:mi>J</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x003A9;</mml:mo></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>S</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mi>I</mml:mi><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>I</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and append it into a list;</td></tr>
<tr>
<td valign="top" align="left">13 end for</td></tr>
<tr>
<td valign="top" align="left">14 select the feature <italic>f</italic><sub><italic>k</italic> max</sub> &#x02208; &#x1D53D; with the largest value of <italic>J</italic><sub><italic>CRIA</italic></sub>(<italic>f</italic><sub><italic>i</italic></sub>) from list;</td></tr>
<tr>
<td valign="top" align="left">15 &#x003A9;<sub><italic>S</italic></sub> &#x0003D; &#x003A9;<sub><italic>S</italic></sub> &#x0222A; {<italic>f</italic><sub><italic>k</italic> max</sub>};</td></tr>
<tr>
<td valign="top" align="left">16 <italic>F</italic> &#x0003D; <italic>F</italic> &#x02212; {<italic>f</italic><sub><italic>k</italic> max</sub>};</td></tr>
<tr>
<td valign="top" align="left">17 end while</td></tr>
<tr>
<td valign="top" align="left">18 output &#x003A9;<sub><italic>S</italic></sub>.</td></tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Verify the Performance of CRIA</title>
<p>Eight gene selection algorithms&#x02014;JMIM (Bennasar et al., <xref ref-type="bibr" rid="B2">2015</xref>), mRMR (Peng et al., <xref ref-type="bibr" rid="B35">2005</xref>), DWFS (Sun et al., <xref ref-type="bibr" rid="B40">2013</xref>), IWFS (Zeng et al., <xref ref-type="bibr" rid="B46">2015</xref>), RAIW (Gu et al., <xref ref-type="bibr" rid="B23">2020</xref>), CFR (Gao et al., <xref ref-type="bibr" rid="B20">2018a</xref>), DCSF (Gao et al., <xref ref-type="bibr" rid="B19">2018b</xref>), and MRI (Wang et al., <xref ref-type="bibr" rid="B43">2017</xref>) are used to compare with CRIA to examine the performance of our proposed method.</p>
<p>The datasets used in validation experiment come from Arizona State University (ASU) datasets (Li et al., <xref ref-type="bibr" rid="B28">2017</xref>), which include four biological data and one other type of data (digit recognition). They are all high-dimensional data. The smallest feature number is 2000 and the largest feature number is 9182 among them. The specific details of these datasets are shown in <xref ref-type="table" rid="T2">Table 2</xref>. We only use minimum description length method (Fayyad and Irani, <xref ref-type="bibr" rid="B14">1993</xref>) for gene selection and utilize it to convert these numerical features.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Datasets for comparison between CRIA algorithm and other algorithms.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Datasets type</bold></th>
<th valign="top" align="center"><bold>No</bold>.</th>
<th valign="top" align="left"><bold>Datasets</bold></th>
<th valign="top" align="center"><bold>Samples</bold></th>
<th valign="top" align="center"><bold>Features</bold></th>
<th valign="top" align="center"><bold>Classes</bold></th>
<th valign="top" align="left"><bold>Types</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Biological data</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left">leukemia</td>
<td valign="top" align="center">72</td>
<td valign="top" align="center">7,070</td>
<td valign="top" align="center">2</td>
<td valign="top" align="left">Discrete</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">2</td>
<td valign="top" align="left">Carcinoma</td>
<td valign="top" align="center">174</td>
<td valign="top" align="center">9,182</td>
<td valign="top" align="center">11</td>
<td valign="top" align="left">Continuous</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">3</td>
<td valign="top" align="left">colon</td>
<td valign="top" align="center">62</td>
<td valign="top" align="center">2,000</td>
<td valign="top" align="center">2</td>
<td valign="top" align="left">Discrete</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">4</td>
<td valign="top" align="left">TOX_171</td>
<td valign="top" align="center">171</td>
<td valign="top" align="center">5,748</td>
<td valign="top" align="center">4</td>
<td valign="top" align="left">Continuous</td>
</tr>
<tr>
<td valign="top" align="left">Digit recognition</td>
<td valign="top" align="center">5</td>
<td valign="top" align="left">Gisette</td>
<td valign="top" align="center">7,000</td>
<td valign="top" align="center">5,000</td>
<td valign="top" align="center">2</td>
<td valign="top" align="left">Continuous</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The number of features <italic>N</italic> used in the experiment is reduced to 50 and three classifiers&#x02014;IB1, J48 and Na&#x000EF;ve Bayes are exploited. The parameters of the classifiers are set to the default parameters of Waikato Environment for Knowledge Analysis (WEKA) (Hall et al., <xref ref-type="bibr" rid="B24">2009</xref>). We use 10 times of ten-fold cross-validation to avoid the influence of randomness on experimental results. Then mean value and Standard Deviation (STD) are taken as the comparison indices of performance of each algorithm and STD is defined as follows:
<disp-formula id="E38"><label>(34)</label><mml:math id="M49"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>T</mml:mi><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>n</italic><sub><italic>run</italic></sub> is the number of times of our experiments, here we set <italic>n</italic><sub><italic>run</italic></sub> &#x0003D; 10, <italic>ACC</italic> is the classification accuracy, <italic>u</italic> represents the average value of <italic>ACC</italic>, and <italic>N</italic> denotes the number of samples. The bigger <italic>ACC</italic>, the better performance, and the smaller <italic>STD</italic>, the higher stability.</p>
<p>The comparison results between the proposed algorithm and other gene selection algorithms are shown in <xref ref-type="table" rid="T3">Tables 3</xref>&#x02013;<xref ref-type="table" rid="T5">5</xref>. As shown in <xref ref-type="table" rid="T3">Table 3</xref>, for the five data sets in the experiment, we can see that the classification results of CRIA in four data sets are better than other eight algorithms, which ranking first, except TOX_171, ranking fourth. Compared with other algorithms, the average accuracy of CRIA is increased by 2.42&#x02013;5.08%. In <xref ref-type="table" rid="T4">Table 4</xref>, CRIA also outperforms the other 8 algorithms on four data sets except TOX_171, on which the experimental results of CRIA ranking third. The biggest improved rate of the proposed algorithm is 10.05% and the smallest one is 3.73%. From <xref ref-type="table" rid="T5">Table 5</xref>, we can find that the results of CRIA on the three data sets are superior to other algorithms, ranking first. However, on the datasets of Carcinoma and TOX_171, compared with the maximum values, the experimental accuracies of CRIA are slightly decreased by 0.50 and 1.71%, ranking second and third respectively. From the perspective of average accuracy, CRIA&#x00027;s result is better than other algorithms, and it is improved by 2.11&#x02013;8.67%.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Comparision (mean &#x000B1; std.dev.) of performance between CRIA and other 8 algorithms with J48 classifier.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Datasets</bold></th>
<th valign="top" align="center"><bold>CRIA</bold><break/> <bold>(proposed)</bold></th>
<th valign="top" align="center"><bold>RAIW</bold></th>
<th valign="top" align="center"><bold>mRMR</bold></th>
<th valign="top" align="center"><bold>DWFS</bold></th>
<th valign="top" align="center"><bold>IWFS</bold></th>
<th valign="top" align="center"><bold>JMIM</bold></th>
<th valign="top" align="center"><bold>MRI</bold></th>
<th valign="top" align="center"><bold>CFR</bold></th>
<th valign="top" align="center"><bold>DCFS</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">leukemia</td>
<td valign="top" align="center"><bold>95.00 &#x000B1; 1.17</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">93.08 &#x000B1; 0.67<break/> (4)</td>
<td valign="top" align="center">93.09 &#x000B1; 0.64<break/> (3)</td>
<td valign="top" align="center">92.51 &#x000B1; 1.14<break/> (7)</td>
<td valign="top" align="center">92.05 &#x000B1; 0.73<break/> (8)</td>
<td valign="top" align="center">93.36 &#x000B1; 0.59<break/> (2)</td>
<td valign="top" align="center">92.70 &#x000B1; 0.98<break/> (6)</td>
<td valign="top" align="center">93.07 &#x000B1; 1.10<break/> (5)</td>
<td valign="top" align="center">91.50 &#x000B1; 1.44<break/> (9)</td>
</tr>
<tr>
<td valign="top" align="left">Carcinoma</td>
<td valign="top" align="center"><bold>75.17 &#x000B1; 1.67</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">71.73 &#x000B1; 1.80<break/> (3)</td>
<td valign="top" align="center">71.78 &#x000B1; 2.26<break/> (2)</td>
<td valign="top" align="center">69.79 &#x000B1; 1.58<break/> (4)</td>
<td valign="top" align="center">64.47 &#x000B1; 1.82<break/> (9)</td>
<td valign="top" align="center">64,84 &#x000B1; 1.68<break/> (5)</td>
<td valign="top" align="center">68.65 &#x000B1; 2.20<break/> (6)</td>
<td valign="top" align="center">68.01 &#x000B1; 1.97<break/> (7)</td>
<td valign="top" align="center">65.53 &#x000B1; 1.89<break/> (8)</td>
</tr>
<tr>
<td valign="top" align="left">colon</td>
<td valign="top" align="center"><bold>79.68 &#x000B1; 2.54</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">76.55 &#x000B1; 3.86<break/> (5)</td>
<td valign="top" align="center">77.14 &#x000B1; 3.31<break/> (4)</td>
<td valign="top" align="center">74.32 &#x000B1; 2.48<break/> (8)</td>
<td valign="top" align="center">76.32 &#x000B1; 2.07<break/> (6)</td>
<td valign="top" align="center">77.28 &#x000B1; 3.77<break/> (3)</td>
<td valign="top" align="center">77.30 &#x000B1; 4.73<break/> (2)</td>
<td valign="top" align="center">74.71 &#x000B1; 3.96<break/> (7)</td>
<td valign="top" align="center">73.31 &#x000B1; 3.77<break/> (9)</td>
</tr>
<tr>
<td valign="top" align="left">TOX_171</td>
<td valign="top" align="center">62.01 &#x000B1; 1.61<break/> (4)</td>
<td valign="top" align="center">62.21 &#x000B1; 2.04<break/> (2)</td>
<td valign="top" align="center">56.61 &#x000B1; 2.21<break/> (8)</td>
<td valign="top" align="center">60.94 &#x000B1; 3.08<break/> (5)</td>
<td valign="top" align="center">62.09 &#x000B1; 2.14<break/> (3)</td>
<td valign="top" align="center">57.49 &#x000B1; 3.01<break/> (9)</td>
<td valign="top" align="center">59.78 &#x000B1; 2.02<break/> (7)</td>
<td valign="top" align="center">60.18 &#x000B1; 1.91<break/> (6)</td>
<td valign="top" align="center"><bold>62.26 &#x000B1; 2.71</bold><break/> <bold>(1)</bold></td>
</tr>
<tr>
<td valign="top" align="left">gisette</td>
<td valign="top" align="center"><bold>93.71 &#x000B1; 0.17</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">92.40 &#x000B1; 0.08<break/> (6)</td>
<td valign="top" align="center">92.02 &#x000B1; 0.08<break/> (8)</td>
<td valign="top" align="center">92.66 &#x000B1; 0.10<break/> (5)</td>
<td valign="top" align="center">92.05 &#x000B1; 0.12<break/> (7)</td>
<td valign="top" align="center">91.18 &#x000B1; 0.12<break/> (9)</td>
<td valign="top" align="center">92.84 &#x000B1; 0.08<break/> (3)</td>
<td valign="top" align="center">92.82 &#x000B1; 0.07<break/> (4)</td>
<td valign="top" align="center">93.35 &#x000B1; 0.08<break/> (2)</td>
</tr>
<tr>
<td valign="top" align="left">Avg.acc</td>
<td valign="top" align="center">81.11</td>
<td valign="top" align="center">79.19</td>
<td valign="top" align="center">78.53</td>
<td valign="top" align="center">78.04</td>
<td valign="top" align="center">77.40</td>
<td valign="top" align="center">77.63</td>
<td valign="top" align="center">78.25</td>
<td valign="top" align="center">77.76</td>
<td valign="top" align="center">77.19</td>
</tr>
<tr>
<td valign="top" align="left">Avg.rank</td>
<td valign="top" align="center">1.60</td>
<td valign="top" align="center">4.00</td>
<td valign="top" align="center">5.00</td>
<td valign="top" align="center">5.80</td>
<td valign="top" align="center">6.60</td>
<td valign="top" align="center">5.60</td>
<td valign="top" align="center">4.80</td>
<td valign="top" align="center">5.80</td>
<td valign="top" align="center">5.80</td>
</tr>
<tr>
<td valign="top" align="left">Improved<break/> rate</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">2.42%</td>
<td valign="top" align="center">3.29%</td>
<td valign="top" align="center">3.93%</td>
<td valign="top" align="center">4.79%</td>
<td valign="top" align="center">4.48%</td>
<td valign="top" align="center">3.65%</td>
<td valign="top" align="center">4.31%</td>
<td valign="top" align="center">5.08%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The meaning of the bold values represent the best performance achieved on a certain dataset for the nine methods</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Comparision (mean &#x000B1; std.dev.) of performance between CRIA and other 8 algorithms with IB1 classifier.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Datasets</bold></th>
<th valign="top" align="center"><bold>CRIA</bold><break/> <bold>(proposed)</bold></th>
<th valign="top" align="center"><bold>RAIW</bold></th>
<th valign="top" align="center"><bold>mRMR</bold></th>
<th valign="top" align="center"><bold>DWFS</bold></th>
<th valign="top" align="center"><bold>IWFS</bold></th>
<th valign="top" align="center"><bold>JMIM</bold></th>
<th valign="top" align="center"><bold>MRI</bold></th>
<th valign="top" align="center"><bold>CFR</bold></th>
<th valign="top" align="center"><bold>DCFS</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">leukemia</td>
<td valign="top" align="center"><bold>99.44 &#x000B1; 0.97</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">97.03 &#x000B1; 0.87<break/> (2)</td>
<td valign="top" align="center">96.16 &#x000B1; 0.43<break/> (4)</td>
<td valign="top" align="center">95.56 &#x000B1; 0.90<break/> (7)</td>
<td valign="top" align="center">88.75 &#x000B1; 1.88<break/> (9)</td>
<td valign="top" align="center">96.61 &#x000B1; 0.78<break/> (3)</td>
<td valign="top" align="center">96.13 &#x000B1; 0.90<break/> (5)</td>
<td valign="top" align="center">95.73 &#x000B1; 0.95<break/> (6)</td>
<td valign="top" align="center">94.22 &#x000B1; 0.80<break/> (8)</td>
</tr>
<tr>
<td valign="top" align="left">Carcinoma</td>
<td valign="top" align="center"><bold>86.84 &#x000B1; 0.50</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">82.45 &#x000B1; 1.33<break/> (2)</td>
<td valign="top" align="center">81.39 &#x000B1; 1.02<break/> (4.5)</td>
<td valign="top" align="center">82.32 &#x000B1; 1.07<break/> (3)</td>
<td valign="top" align="center">76.88 &#x000B1; 1.97<break/> (9)</td>
<td valign="top" align="center">81.10 &#x000B1; 1.06<break/> (7)</td>
<td valign="top" align="center">81.39 &#x000B1; 1.16<break/> (4.5)</td>
<td valign="top" align="center">81.35 &#x000B1; 1.08<break/> (6)</td>
<td valign="top" align="center">80.96 &#x000B1; 1.18<break/> (8)</td>
</tr>
<tr>
<td valign="top" align="left">colon</td>
<td valign="top" align="center"><bold>86.77 &#x000B1; 1.27</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">78.60 &#x000B1; 2.10<break/> (2)</td>
<td valign="top" align="center">78.22 &#x000B1; 1.54<break/> (3)</td>
<td valign="top" align="center">75.94 &#x000B1; 2.32<break/> (5)</td>
<td valign="top" align="center">70.87 &#x000B1; 1.97<break/> (9)</td>
<td valign="top" align="center">76.69 &#x000B1; 2.35<break/> (4)</td>
<td valign="top" align="center">71.77 &#x000B1; 3.41<break/> (8)</td>
<td valign="top" align="center">72.24 &#x000B1; 2.17<break/> (7)</td>
<td valign="top" align="center">74.87 &#x000B1; 1.80<break/> (6)</td>
</tr>
<tr>
<td valign="top" align="left">TOX_171</td>
<td valign="top" align="center">84.56 &#x000B1; 0.52<break/> (3)</td>
<td valign="top" align="center">85.13 &#x000B1; 1.26<break/> (2)</td>
<td valign="top" align="center">78.14 &#x000B1; 1.36<break/> (8)</td>
<td valign="top" align="center"><bold>85.19 &#x000B1; 1.30</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">82.59 &#x000B1; 1.78<break/> (5)</td>
<td valign="top" align="center">76.68 &#x000B1; 1.68<break/> (9)</td>
<td valign="top" align="center">81.69 &#x000B1; 1.64<break/> (7)</td>
<td valign="top" align="center">82.05 &#x000B1; 1.28<break/> (6)</td>
<td valign="top" align="center">84.04 &#x000B1; 1.48<break/> (4)</td>
</tr>
<tr>
<td valign="top" align="left">gisette</td>
<td valign="top" align="center"><bold>93.75 &#x000B1; 0.14</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">91.88 &#x000B1; 0.08<break/> (6)</td>
<td valign="top" align="center">91.26 &#x000B1; 0.09<break/> (7)</td>
<td valign="top" align="center">92.26 &#x000B1; 0.07<break/> (5)</td>
<td valign="top" align="center">91.05 &#x000B1; 0.15<break/> (8)</td>
<td valign="top" align="center">90.20 &#x000B1; 0.10<break/> (9)</td>
<td valign="top" align="center">92.70 &#x000B1; 0.06<break/> (3)</td>
<td valign="top" align="center">92.58 &#x000B1; 0.05<break/> (4)</td>
<td valign="top" align="center">93.13 &#x000B1; 0.10<break/> (2)</td>
</tr>
<tr>
<td valign="top" align="left">Avg.acc</td>
<td valign="top" align="center">90.27</td>
<td valign="top" align="center">87.02</td>
<td valign="top" align="center">85.03</td>
<td valign="top" align="center">86.25</td>
<td valign="top" align="center">82.03</td>
<td valign="top" align="center">84.26</td>
<td valign="top" align="center">84.74</td>
<td valign="top" align="center">84.79</td>
<td valign="top" align="center">85.44</td>
</tr>
<tr>
<td valign="top" align="left">Avg.rank</td>
<td valign="top" align="center">1.40</td>
<td valign="top" align="center">2.80</td>
<td valign="top" align="center">5.30</td>
<td valign="top" align="center">4.20</td>
<td valign="top" align="center">8.00</td>
<td valign="top" align="center">6.40</td>
<td valign="top" align="center">5.50</td>
<td valign="top" align="center">5.80</td>
<td valign="top" align="center">5.60</td>
</tr>
<tr>
<td valign="top" align="left">Improved<break/> rate</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">3.73%</td>
<td valign="top" align="center">6.16%</td>
<td valign="top" align="center">4.66%</td>
<td valign="top" align="center">10.05%</td>
<td valign="top" align="center">7.13%</td>
<td valign="top" align="center">6.53%</td>
<td valign="top" align="center">6.46%</td>
<td valign="top" align="center">5.65%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The meaning of the bold values represent the best performance achieved on a certain dataset for the nine methods</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Comparision (mean &#x000B1; std.dev.) of performance between CRIA and other 8 algorithms with Na&#x000EF;ve Bayes classifier.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Datasets</bold></th>
<th valign="top" align="center"><bold>CRIA</bold><break/> <bold>(proposed)</bold></th>
<th valign="top" align="center"><bold>RAIW</bold></th>
<th valign="top" align="center"><bold>mRMR</bold></th>
<th valign="top" align="center"><bold>DWFS</bold></th>
<th valign="top" align="center"><bold>IWFS</bold></th>
<th valign="top" align="center"><bold>JMIM</bold></th>
<th valign="top" align="center"><bold>MRI</bold></th>
<th valign="top" align="center"><bold>CFR</bold></th>
<th valign="top" align="center"><bold>DCFS</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">leukemia</td>
<td valign="top" align="center"><bold>99.58 &#x000B1; 0.67</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">97.44 &#x000B1; 0.66<break/> (2)</td>
<td valign="top" align="center">96.27 &#x000B1; 0.30<break/> (7)</td>
<td valign="top" align="center">97.03 &#x000B1; 0.70<break/> (4)</td>
<td valign="top" align="center">95.48 &#x000B1; 1.96<break/> (9)</td>
<td valign="top" align="center">96.18 &#x000B1; 0.39<break/> (8)</td>
<td valign="top" align="center">97.15 &#x000B1; 0.80<break/> (3)</td>
<td valign="top" align="center">96.79 &#x000B1; 0.71<break/> (5)</td>
<td valign="top" align="center">96.70 &#x000B1; 0.58<break/> (6)</td>
</tr>
<tr>
<td valign="top" align="left">Carcinoma</td>
<td valign="top" align="center">81.61 &#x000B1; 1.11<break/> (2)</td>
<td valign="top" align="center"><bold>82.02 &#x000B1; 0.83</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">80.23 &#x000B1; 1.33<break/> (7)</td>
<td valign="top" align="center">81.58 &#x000B1; 0.95<break/> (3)</td>
<td valign="top" align="center">76.38 &#x000B1; 1.85<break/> (9)</td>
<td valign="top" align="center">80.19 &#x000B1; 1.28<break/> (8)</td>
<td valign="top" align="center">80.41 &#x000B1; 0.86<break/> (5)</td>
<td valign="top" align="center">80.34 &#x000B1; 0.87<break/> (6)</td>
<td valign="top" align="center">80.46 &#x000B1; 1.09<break/> (4)</td>
</tr>
<tr>
<td valign="top" align="left">colon</td>
<td valign="top" align="center"><bold>88.71 &#x000B1; 0.00</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">82.97 &#x000B1; 1.34<break/> (2)</td>
<td valign="top" align="center">82.70 &#x000B1; 1.20<break/> (3)</td>
<td valign="top" align="center">80.72 &#x000B1; 1.47<break/> (6)</td>
<td valign="top" align="center">74.39 &#x000B1; 4.37<break/> (9)</td>
<td valign="top" align="center">81.66 &#x000B1; 1.45<break/> (5)</td>
<td valign="top" align="center">79.84 &#x000B1; 2.40<break/> (7)</td>
<td valign="top" align="center">78.95 &#x000B1; 1.72<break/> (8)</td>
<td valign="top" align="center">82.32 &#x000B1; 2.08<break/> (4)</td>
</tr>
<tr>
<td valign="top" align="left">TOX_171</td>
<td valign="top" align="center">69.53 &#x000B1; 0.70<break/> (3)</td>
<td valign="top" align="center"><bold>70.74 &#x000B1; 1.07</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">63.73 &#x000B1; 1.61<break/> (8)</td>
<td valign="top" align="center">68.68 &#x000B1; 1.26<break/> (4)</td>
<td valign="top" align="center">65.64 &#x000B1; 1.63<break/> (7)</td>
<td valign="top" align="center">60.41 &#x000B1; 2.35<break/> (9)</td>
<td valign="top" align="center">66.96 &#x000B1; 1.56<break/> (6)</td>
<td valign="top" align="center">67.04 &#x000B1; 1.74<break/> (5)</td>
<td valign="top" align="center">70.28 &#x000B1; 1.60<break/> (2)</td>
</tr>
<tr>
<td valign="top" align="left">gisette</td>
<td valign="top" align="center"><bold>93.16 &#x000B1; 0.05</bold><break/> <bold>(1)</bold></td>
<td valign="top" align="center">90.46 &#x000B1; 0.13<break/> (2)</td>
<td valign="top" align="center">88.26 &#x000B1; 0.02<break/> (4)</td>
<td valign="top" align="center">87.69 &#x000B1; 0.08<break/> (5)</td>
<td valign="top" align="center">86.23 &#x000B1; 0.24<break/> (8)</td>
<td valign="top" align="center">86.01 &#x000B1; 0.05<break/> (9)</td>
<td valign="top" align="center">87.60 &#x000B1; 0.04<break/> (6)</td>
<td valign="top" align="center">87.48 &#x000B1; 0.03<break/> (7)</td>
<td valign="top" align="center">89.46 &#x000B1; 0.05<break/> (3)</td>
</tr>
<tr>
<td valign="top" align="left">Avg.acc</td>
<td valign="top" align="center">86.52</td>
<td valign="top" align="center">84.73</td>
<td valign="top" align="center">82.24</td>
<td valign="top" align="center">83.14</td>
<td valign="top" align="center">79.62</td>
<td valign="top" align="center">80.89</td>
<td valign="top" align="center">82.39</td>
<td valign="top" align="center">82.12</td>
<td valign="top" align="center">83.84</td>
</tr>
<tr>
<td valign="top" align="left">Avg.rank</td>
<td valign="top" align="center">1.60</td>
<td valign="top" align="center">1.80</td>
<td valign="top" align="center">5.80</td>
<td valign="top" align="center">4.40</td>
<td valign="top" align="center">8.40</td>
<td valign="top" align="center">7.80</td>
<td valign="top" align="center">5.40</td>
<td valign="top" align="center">6.20</td>
<td valign="top" align="center">3.80</td>
</tr>
<tr>
<td valign="top" align="left">Improved<break/> rate</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">2.11%</td>
<td valign="top" align="center">5.20%</td>
<td valign="top" align="center">4.07%</td>
<td valign="top" align="center">8.67%</td>
<td valign="top" align="center">6.96%</td>
<td valign="top" align="center">5.01%</td>
<td valign="top" align="center">5.36%</td>
<td valign="top" align="center">3.20%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The meaning of the bold values represent the best performance achieved on a certain dataset for the nine methods</italic>.</p>
</table-wrap-foot>
</table-wrap></sec></sec>
<sec id="s4">
<title>Results and Discussions</title>
<sec>
<title>Evaluation Metrics of Experimental Results</title>
<p>Four evaluation metrics&#x02014;precision, recall, accuracy and F1-score are utilized to evaluate the performance of the corresponding method and values of these criteria are defined as equation (35).
<disp-formula id="E39"><label>(35)</label><mml:math id="M50"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <italic>T</italic><sub><italic>P</italic></sub>, <italic>T</italic><sub><italic>N</italic></sub>, F<sub><italic>P</italic></sub>, and <italic>F</italic><sub><italic>N</italic></sub> denotes the numbers of true positives, true negatives, false positives, and false negatives respectively.</p></sec>
<sec>
<title>The CRIA and IFS Results</title>
<p>As mentioned in section Evaluation Metrics of Experimental Results, each sample is represented by 24,174 features, each of which indicates the expression level of genes. The 24174 feature genes are sorted by CRIA value in descending order. However, we only select the top 200 features in this work for the consideration of computational time and curse of dimensionality. The top 15 key feature genes chosen by CRIA defined by equation (26) are listed in <xref ref-type="table" rid="T6">Table 6</xref>.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>The top 15 feature genes chosen by CRIA defined as equation (26).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Ranked order</bold></th>
<th valign="top" align="left"><bold>Official name</bold></th>
<th valign="top" align="left"><bold>Official full gene name</bold></th>
<th valign="top" align="left"><bold>Category</bold></th>
<th valign="top" align="center"><bold>CRIA value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="left">RPS15</td>
<td valign="top" align="left">ribosomal protein S15</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.168</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="left">TBC1D5</td>
<td valign="top" align="left">TBC1 Domain Family Member 5</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.089</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">CUL2</td>
<td valign="top" align="left">Cullin 2</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.093</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="left">SMPD3</td>
<td valign="top" align="left">Sphingomyelin Phosphodiesterase 3</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.089</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">CTAGE10P</td>
<td valign="top" align="left">CTAGE Family Member 10, Pseudogene</td>
<td valign="top" align="left">Pseudogene</td>
<td valign="top" align="center">0.071</td>
</tr>
<tr>
<td valign="top" align="left">6</td>
<td valign="top" align="left">C1orf98</td>
<td valign="top" align="left">Chromosome 1 Open Reading Frame 98</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.043</td>
</tr>
<tr>
<td valign="top" align="left">7</td>
<td valign="top" align="left">ZNF281</td>
<td valign="top" align="left">Zinc Finger Protein 281</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.061</td>
</tr>
<tr>
<td valign="top" align="left">8</td>
<td valign="top" align="left">CDKN2A</td>
<td valign="top" align="left">Cyclin Dependent Kinase Inhibitor 2A</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.161</td>
</tr>
<tr>
<td valign="top" align="left">9</td>
<td valign="top" align="left">EGFR</td>
<td valign="top" align="left">Epidermal Growth Factor Receptor</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.121</td>
</tr>
<tr>
<td valign="top" align="left">10</td>
<td valign="top" align="left">TMEM98</td>
<td valign="top" align="left">Transmembrane Protein 98</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.103</td>
</tr>
<tr>
<td valign="top" align="left">11</td>
<td valign="top" align="left">CTBP2</td>
<td valign="top" align="left">C-Terminal Binding Protein 2</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.083</td>
</tr>
<tr>
<td valign="top" align="left">12</td>
<td valign="top" align="left">SEMA6A</td>
<td valign="top" align="left">Semaphorin 6A</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.081</td>
</tr>
<tr>
<td valign="top" align="left">13</td>
<td valign="top" align="left">MIR1208</td>
<td valign="top" align="left">MicroRNA 1208</td>
<td valign="top" align="left">RNA Gene</td>
<td valign="top" align="center">0.077</td>
</tr>
<tr>
<td valign="top" align="left">14</td>
<td valign="top" align="left">RBFOX1</td>
<td valign="top" align="left">RNA Binding Fox-1 Homolog 1</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.069</td>
</tr>
<tr>
<td valign="top" align="left">15</td>
<td valign="top" align="left">CDC25A</td>
<td valign="top" align="left">Cell Division Cycle 25A</td>
<td valign="top" align="left">Protein Coding</td>
<td valign="top" align="center">0.066</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We use the Incremental Gene selection (IFS) (Yang et al., <xref ref-type="bibr" rid="B45">2019</xref>) to determine the optimal feature set. The first 200 features are added one by one to a feature subset in order. Each time a feature is added, a classifier is trained and examined. So, 200 classifiers are constructed. We use the criteria of accuracy to evaluate the performance of all the 200 classifiers and then we choose the classifier with the highest accuracy as the final one. The corresponding feature subset that the final classifier used is deemed to be the optimal feature set.</p>
<p>In this paper, three commonly used classifiers are adopted to verify the generalization performance of the proposed gene selection method on different classifiers. ten-fold cross-validation is used to evaluate our algorithm with the selected features. The complete data set is randomly split into 10 parts of approximately equal size. The three classifiers are trained 10 times; nine of the 10 subsets are used as the training datasets, and the remaining one is the test dataset. The average values of accuracy for each classifier are calculated and the IFS results are shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. Here, we name our methods as CRIA_CatBoost, CRIA_SVM and CRIA_LightGBM. From <xref ref-type="fig" rid="F2">Figure 2</xref>, it can be seen that the highest accuracy of 86.90% for CRIA_CatBoost method followed by 86.41% for CRIA_LightGBM and 85.98% for CRIA_SVM method, with only using the CNVs of 131 genes, 138 genes and 122 genes respectively.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Classification accuracies of three classifiers (CatBoost, LightGBM and SVM) with different numbers of features during the IFS procedure. The top 200 feature genes are selected by CRIA method.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-839044-g0002.tif"/>
</fig></sec>
<sec>
<title>The Proposed Algorithm Performance</title>
<p>For the different classifiers used in this work, after determining the optimal numbers of features according to the CRIA and IFS results, the classification performance can be further analyzed. The average values of three metrics-precision, recall and F1-score defined in Equation (35) on 10 test datasets are listed in <xref ref-type="table" rid="T7">Table 7</xref>.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Average performance of precision, recall and F1-score on 10 test datasets with three classifiers <italic>via</italic> ten-fold cross-validation (%).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left" colspan="2"><bold>Metrics</bold></th>
<th valign="top" align="center"><bold>UCEC</bold></th>
<th valign="top" align="center"><bold>KIRC</bold></th>
<th valign="top" align="center"><bold>OV</bold></th>
<th valign="top" align="center"><bold>GBM</bold></th>
<th valign="top" align="center"><bold>COAD/</bold><break/> <bold>READ</bold></th>
<th valign="top" align="center"><bold>BRCA</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Precision</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center">74.31</td>
<td valign="top" align="center">93.74</td>
<td valign="top" align="center">84.59</td>
<td valign="top" align="center">94.63</td>
<td valign="top" align="center">89.67</td>
<td valign="top" align="center">84.48</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center">70.47</td>
<td valign="top" align="center">90.33</td>
<td valign="top" align="center">85.40</td>
<td valign="top" align="center">95.64</td>
<td valign="top" align="center">88.54</td>
<td valign="top" align="center">84.76</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center">93.37</td>
<td valign="top" align="center">82.84</td>
<td valign="top" align="center">95.72</td>
<td valign="top" align="center">90.23</td>
<td valign="top" align="center">84.71</td>
</tr>
<tr>
<td valign="top" align="left">Recall</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center">73.14</td>
<td valign="top" align="center">91.63</td>
<td valign="top" align="center">87.90</td>
<td valign="top" align="center">90.76</td>
<td valign="top" align="center">86.09</td>
<td valign="top" align="center">88.67</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center">73.81</td>
<td valign="top" align="center">89.59</td>
<td valign="top" align="center">88.43</td>
<td valign="top" align="center">89.70</td>
<td valign="top" align="center">83.30</td>
<td valign="top" align="center">87.96</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center">72.91</td>
<td valign="top" align="center">92.04</td>
<td valign="top" align="center">89.32</td>
<td valign="top" align="center">91.30</td>
<td valign="top" align="center">83.48</td>
<td valign="top" align="center">87.01</td>
</tr>
<tr>
<td valign="top" align="left">F1-score</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center">73.72</td>
<td valign="top" align="center">92.67</td>
<td valign="top" align="center">86.21</td>
<td valign="top" align="center">92.65</td>
<td valign="top" align="center">87.84</td>
<td valign="top" align="center">86.52</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center">72.13</td>
<td valign="top" align="center">89.96</td>
<td valign="top" align="center">86.89</td>
<td valign="top" align="center">92.57</td>
<td valign="top" align="center">85.84</td>
<td valign="top" align="center">86.33</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center">72.18</td>
<td valign="top" align="center">92.70</td>
<td valign="top" align="center">85.96</td>
<td valign="top" align="center">93.46</td>
<td valign="top" align="center">86.72</td>
<td valign="top" align="center">85.84</td>
</tr>
</tbody>
</table>
</table-wrap></sec>
<sec>
<title>Performance Comparison With Other Methods</title>
<p>After selecting important features, we use three common classifiers&#x02014;CatBoost, SVM and LightGBM to predict cancer samples. The performance of our methods are compared with other two classification methods published before whose experimental dataset is the same as ours. Liang et al. (<xref ref-type="bibr" rid="B30">2020</xref>) used a method called CNA_origin which was composed of a stacked autoencoder and an one-dimensional convolutional neural network. The 24,174 gene features were extracted to 100 genes by the autoencoder, and then these 100 gene features were put into the 1D CNN for classification (Liang et al., <xref ref-type="bibr" rid="B30">2020</xref>). A computationally method for cancer types classification proposed by Zhang et al. (<xref ref-type="bibr" rid="B47">2016</xref>) was named as mRMR_Dagging here because there was no specific method name given by authors. It first used mRMR and IFS to select 19 of the 24,174 genes as classification features, and then used the Dagging algorithm to give the final results.</p>
<p>In <xref ref-type="table" rid="T8">Table 8</xref>, it can be seen that if the results of our methods are superior to CNA_origin and mRMR_Dagging, they are marked in bold. Similarly, if the largest of CNA_origin and mRMR_Dagging results is better than our method, it is also marked in bold. <xref ref-type="table" rid="T8">Table 8</xref> demonstrated that the performance of our methods is superior to CNA_origin and mRMR_Dagging for UCEC, KIRC, GBM, and COADREAD. For UCEC, the recall and F1-score of our methods (CRIA_Cat- Boost, CRIA_SVM and CRIA_LightGBM) are all superior to CNA_origin and mRMR_Dagging. The best precision of our methods is 0.12 percentage points higher than mRMR_Dagging. SVM and LightGBM are slightly worse than mRMR_Dagging with reductions of 5.28 and 3.82% in precision respectively. For KIRC, the precision and F1-score are all superior to CNA_origin and mRMR_Dagging except the F1-score of SVM, which performs slightly worse than the CNA_origin with reductions of 2.61%. Compared with the best, CNA_origin, the recall of our methods are decreased by 4.77% for CatBoost, 7.15% for SVM and 4.30% for LightGBM. For OV, compared with CNA_origin, the recall of our methods is at least increased by 1.36%. The precision and F1-score are slightly worse than CNA_origin, with reductions at most of 8.40, and 2.37%, respectively. For GBM and COADREAD, our methods are better than CNA_origin and mRMR_Dagging on all evaluation indicators. Compared with the best of the other two algorithms, the worst precision of our methods is increased by 1.64 and 8.53%, respectively, the worst recall is increased by 4.39 and 12.86%, respectively, and the worst F1-score is increased by 4.58 and 10.76%, respectively. For BRCA, the worst among our methods performs slightly worse than the best CNA_origin algorithm, with reductions of 3.57% in precision, 6.09% in recall and 4.66% in F1-score respectively.</p>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>Performance comparison of the proposed algorithm predictions with those of other methods (%).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Cancer</bold></th>
<th valign="top" align="left"><bold>Predictor</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>F1-score</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">UCEC</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center"><bold>74.31</bold></td>
<td valign="top" align="center"><bold>73.14</bold></td>
<td valign="top" align="center"><bold>73.72</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center">70.47</td>
<td valign="top" align="center"><bold>73.81</bold></td>
<td valign="top" align="center"><bold>72.13</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center">71.46</td>
<td valign="top" align="center"><bold>72.91</bold></td>
<td valign="top" align="center"><bold>72.18</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CNA_origin</td>
<td valign="top" align="center">67.92</td>
<td valign="top" align="center">72.00</td>
<td valign="top" align="center">69.90</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">mRMR_Dagging</td>
<td valign="top" align="center">74.19</td>
<td valign="top" align="center">46.93</td>
<td valign="top" align="center">57.50</td>
</tr>
<tr>
<td valign="top" align="left">KIRC</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center"><bold>93.74</bold></td>
<td valign="top" align="center">91.63</td>
<td valign="top" align="center"><bold>92.67</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center"><bold>90.33</bold></td>
<td valign="top" align="center">89.59</td>
<td valign="top" align="center">89.96</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center"><bold>93.37</bold></td>
<td valign="top" align="center">92.04</td>
<td valign="top" align="center"><bold>92.70</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CNA_origin</td>
<td valign="top" align="center">88.89</td>
<td valign="top" align="center"><bold>96.00</bold></td>
<td valign="top" align="center">92.31</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">mRMR_Dagging</td>
<td valign="top" align="center">80.85</td>
<td valign="top" align="center">92.68</td>
<td valign="top" align="center">86.36</td>
</tr>
<tr>
<td valign="top" align="left">OV</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center">84.59</td>
<td valign="top" align="center"><bold>87.90</bold></td>
<td valign="top" align="center">86.21</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center">85.40</td>
<td valign="top" align="center"><bold>88.43</bold></td>
<td valign="top" align="center">86.89</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center">82.84</td>
<td valign="top" align="center"><bold>89.32</bold></td>
<td valign="top" align="center">85.96</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CNA_origin</td>
<td valign="top" align="center"><bold>89.80</bold></td>
<td valign="top" align="center">86.72</td>
<td valign="top" align="center"><bold>88.00</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">mRMR_Dagging</td>
<td valign="top" align="center">84.61</td>
<td valign="top" align="center">75.86</td>
<td valign="top" align="center">80.00</td>
</tr>
<tr>
<td valign="top" align="left">GBM</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center"><bold>94.63</bold></td>
<td valign="top" align="center"><bold>90.76</bold></td>
<td valign="top" align="center"><bold>92.65</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center"><bold>95.64</bold></td>
<td valign="top" align="center"><bold>89.70</bold></td>
<td valign="top" align="center"><bold>92.57</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center"><bold>95.72</bold></td>
<td valign="top" align="center"><bold>91.30</bold></td>
<td valign="top" align="center"><bold>93.46</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CNA_origin</td>
<td valign="top" align="center">93.10</td>
<td valign="top" align="center">84.38</td>
<td valign="top" align="center">88.52</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">mRMR_Dagging</td>
<td valign="top" align="center">88.70</td>
<td valign="top" align="center">85.93</td>
<td valign="top" align="center">87.30</td>
</tr>
<tr>
<td valign="top" align="left">COADREAD</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center"><bold>89.67</bold></td>
<td valign="top" align="center"><bold>86.09</bold></td>
<td valign="top" align="center"><bold>87.84</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center"><bold>88.54</bold></td>
<td valign="top" align="center"><bold>83.30</bold></td>
<td valign="top" align="center"><bold>85.84</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center"><bold>90.23</bold></td>
<td valign="top" align="center"><bold>83.48</bold></td>
<td valign="top" align="center"><bold>86.72</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CNA_origin</td>
<td valign="top" align="center">81.58</td>
<td valign="top" align="center">73.81</td>
<td valign="top" align="center">77.50</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">mRMR_Dagging</td>
<td valign="top" align="center">60.00</td>
<td valign="top" align="center">73.46</td>
<td valign="top" align="center">66.05</td>
</tr>
<tr>
<td valign="top" align="left">BRCA</td>
<td valign="top" align="left">CRIA_CatBoost</td>
<td valign="top" align="center">84.48</td>
<td valign="top" align="center">88.67</td>
<td valign="top" align="center">86.52</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_SVM</td>
<td valign="top" align="center">84.76</td>
<td valign="top" align="center">87.96</td>
<td valign="top" align="center">86.33</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CRIA_LightGBM</td>
<td valign="top" align="center">84.71</td>
<td valign="top" align="center">87.01</td>
<td valign="top" align="center">85.84</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">CNA_origin</td>
<td valign="top" align="center"><bold>87.50</bold></td>
<td valign="top" align="center"><bold>92.31</bold></td>
<td valign="top" align="center"><bold>89.84</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">mRMR_Dagging</td>
<td valign="top" align="center">79.16</td>
<td valign="top" align="center">87.35</td>
<td valign="top" align="center">83.06</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, the macro-average results of four evaluation metrics: accuracy, precision, recall and F1-score are used to assess our methods and the other two algorithms on the datasets of six types of cancers. The results can be seen in <xref ref-type="fig" rid="F3">Figure 3</xref>. For accuracy, our methods have mean values of 86.90% for CatBoost, 86.41% for LightGBM and 85.98% for SVM respectively, which are increased by 3.69, 3.10, and 2.59% compared with CNA_origin. For precision, the average values of our methods are 86.61, 86.39, and 85.86%, which are increased by 3.49, 3.23, and 2.59%, respectively compared with the best among CNA_origin and mRMR_Dagging. For recall, our methods&#x00027; mean values are 86.37, 86.01, and 85.47%, which are 2.92, 2.56 and 2.02 percentage points higher than CNA_origin, respectively. For F1-score, compared with our methods, whose average values are 86.60, 86.14, and 85.62%, CNA_origin is decreased by 3.71, 3.19, and 2.60%, respectively.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Performance comparison of 4 evaluation metrics: accuracy, precision, recall and F1-score among our methods and the other two algorithms (CNA_origin and mRMR_Dagging).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-839044-g0003.tif"/>
</fig></sec>
<sec>
<title>Further Discussion</title>
<p>In order to study the relationship between the classes, we also summarize the confusion matrices in <xref ref-type="fig" rid="F4">Figure 4</xref> for class predictions using our methods. From <xref ref-type="fig" rid="F4">Figure 4</xref>, we can find that there existed a high error rate when predicting the samples of UCEC. Regardless of whether it is CRIA_CatBoost, CRIA_SVM or CRIA_LightGBM, more than 10% of the UCEC samples are incorrectly predicted as OV and BRCA. In <xref ref-type="fig" rid="F4">Figure 4A</xref>, 14.00% of UCEC samples are predicted as OV, while 11.06% of UCECsamples are predicted to be BRCA. In <xref ref-type="fig" rid="F4">Figure 4B</xref>, 14.45 and 11.74% of UCEC samples are predicted as OV and BRCA respectively. In <xref ref-type="fig" rid="F4">Figure 4C</xref>, 13.32 and 12.42% of UCEC samples are predicted as OV and BRCA respectively. The reasons may be that UCEC, OV and BRCA are hormone-dependent tumors and they relate closely in tumorigenesis. The 16 and 27 risk regions were identified by an independent genome-wide association study (GWAS) on endometrial cancer and ovarian cancer, respectively (Glubb et al., <xref ref-type="bibr" rid="B21">2020</xref>). Studies have shown that mutations in breast cancer susceptibility genes (BRCA1, BRCA2) have a relationship in hereditary ovarian cancer. Mutations at either end of the BRCA1 gene increase a person&#x00027;s risk of breast cancer, and its probability is higher than ovarian cancer. However, mutations in the middle of the BRCA1 gene put a person at a higher risk of ovarian cancer than breast cancer (Shi et al., <xref ref-type="bibr" rid="B39">2017</xref>). In addition, there is also a study indicated that UCEC, OV and BRCA all have a relationship with the changes in estrogen and estrogen receptors (Rodriguez et al., <xref ref-type="bibr" rid="B37">2019</xref>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Confusion matrices on the test groups: <bold>(A)</bold> CRIA_CatBoost; <bold>(B)</bold> CRIA_SVM; <bold>(C)</bold> CRIA_LightGBM.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-839044-g0004.tif"/>
</fig></sec></sec>
<sec sec-type="conclusions" id="s5">
<title>Conclusions</title>
<p>In this paper, we introduce a gene selection algorithm&#x02014;CRIA. We firstly apply this algorithm to 5 datasets and verify the effective performance of CRIA through comparison with other eight gene selection algorithms. The proposed algorithm can select features which are closely related to the class label. Then, we use this algorithm to select 200 genes that have a close relationship with cancer types from 24,174 genes features based on the value of copy number variations in the samples, and then combine three common classifiers&#x02014;CatBoost, SVM and LightGBM to predict the type of cancer. Our experimental results show that our methods have higher accuracies than the state-of-the-art methods for solving this problem. Our research has a certain degree of interpretability for cancer-related researches at the genetic level. As we all know, cancer is closely related to gene structural variations and the appearance of cancer is often accompanied by abnormalities in the deoxyribonucleic acid (DNA) sequence. Because CNVs is one of the most crucial structural variations of genes, studying the relationship between cancers and CNVs is of great significance. Many studies have tried to utilize the genetic information of cancers to predict cancer type, which can provide significant guidance for patient care and cancer therapy in promptly.</p>
<p>The future direction of this work can continue to develop from two aspects. First of all, because we only use the datasets of six cancer types and the total number of samples is only 3,480 in this paper, by collecting data sets of other cancer types and optimizing the proposed algorithm, we can continue to conduct further research in the field of cancer classification based on copy number variations. Moreover, integrating non-CNVs features for the samples can be taken into consideration. In addition to using CNVs for cancer prediction, we can also apply other genetic information for cancer prediction, or combine several biomarkers to reduce the error rate of classification as much as possible.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>QW conducted the experiments and wrote the manuscript. DL conceived and provided the main direction of the manuscript and guided the writing and modification of this manuscript. Both authors read and approved the manuscript.</p></sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This work was supported by the National Natural Science Foundation of China (Grant No. 11571009) and Applied Basic Research Programs of Shanxi Province (Grant No. 201901D111086).</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Battiti</surname> <given-names>R.</given-names></name></person-group> (<year>1994</year>). <article-title>Using mutual information for selecting features in supervised neural net learning</article-title>. <source>IEEE Trans. Neural Netw</source>. <volume>5</volume>, <fpage>537</fpage>&#x02013;<lpage>550</lpage>. <pub-id pub-id-type="doi">10.1109/72.298224</pub-id><pub-id pub-id-type="pmid">18267827</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bennasar</surname> <given-names>M.</given-names></name> <name><surname>Hicks</surname> <given-names>Y.</given-names></name> <name><surname>Setchi</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Gene selection using Joint Mutual Information Maximisation</article-title>. <source>Expert Syst. Appl.</source> <volume>42</volume>, <fpage>8520</fpage>&#x02013;<lpage>8532</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2015.07.007</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beroukhim</surname> <given-names>R.</given-names></name> <name><surname>Getz</surname> <given-names>G.</given-names></name> <name><surname>Nghiemphu</surname> <given-names>L.</given-names></name> <name><surname>Barretina</surname> <given-names>J.</given-names></name> <name><surname>Hsueh</surname> <given-names>T.</given-names></name> <name><surname>Linhart</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>Assessing the significance of chromosomal aberrations in cancer: Methodology and app- lication to glioma</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>104</volume>, <fpage>20007</fpage>&#x02013;<lpage>20012</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0710052104</pub-id><pub-id pub-id-type="pmid">18077431</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breunis</surname> <given-names>W. B.</given-names></name> <name><surname>van Mirre</surname> <given-names>E.</given-names></name> <name><surname>Bruin</surname> <given-names>M.</given-names></name> <name><surname>Geissler</surname> <given-names>J.</given-names></name> <name><surname>de Boer</surname> <given-names>M.</given-names></name> <name><surname>Peters</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>Copy number variation of the activating FCGR2C gene predisposes to idiopathic thrombocytopenic purpura</article-title>. <source>Blood</source> <volume>111</volume>, <fpage>1029</fpage>&#x02013;<lpage>1038</lpage>. <pub-id pub-id-type="doi">10.1182/blood-2007-03-079913</pub-id><pub-id pub-id-type="pmid">17827395</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buchynska</surname> <given-names>L. G.</given-names></name> <name><surname>Brieieva</surname> <given-names>O. V.</given-names></name> <name><surname>Iurchenko</surname> <given-names>N. P.</given-names></name></person-group> (<year>2019</year>). <article-title>Assessment of HER-2/neu, &#x003B1;-MYC and CCN- E1 gene copy number variations and protein expression in endometrial carcinomas</article-title>. <source>Exp. Oncol</source>. <fpage>41</fpage>. <pub-id pub-id-type="doi">10.32471/exp-oncology.2312-8852.vol-41-no-2.12973</pub-id><pub-id pub-id-type="pmid">31262163</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cerami</surname> <given-names>E.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Dogrusoz</surname> <given-names>U.</given-names></name> <name><surname>Gross</surname> <given-names>B. E.</given-names></name> <name><surname>Sumer</surname> <given-names>S. O.</given-names></name> <name><surname>Aksoy</surname> <given-names>B. A.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>The cBio cancer genomics portal: an open platform for exploring multidimensional cancer geno- mics data: figure 1</article-title>. <source>Cancer Discov.</source> <volume>2</volume>, <fpage>401</fpage>&#x02013;<lpage>404</lpage>. <pub-id pub-id-type="doi">10.1158/2159-8290.CD-12-0095</pub-id><pub-id pub-id-type="pmid">22588877</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chandrashekar</surname> <given-names>G.</given-names></name> <name><surname>Sahin</surname> <given-names>F.</given-names></name></person-group> (<year>2014</year>). <article-title>A survey on gene selection methods</article-title>. <source>Comput. Electr. Eng</source>. <volume>40</volume>, <fpage>16</fpage>&#x02013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1016/j.compeleceng.2013.11.024</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Wu</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>Z.</given-names></name> <name><surname>Ran</surname> <given-names>B.</given-names></name> <name><surname>Zhong</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Gene selection with redundancy-complementariness dispersion</article-title>. <source>Knowl. Based Syst.</source> <volume>89</volume>, <fpage>203</fpage>&#x02013;<lpage>217</lpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2015.07.004</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ciriello</surname> <given-names>G.</given-names></name> <name><surname>Miller</surname> <given-names>M. L.</given-names></name> <name><surname>Aksoy</surname> <given-names>B. A.</given-names></name> <name><surname>Senbabaoglu</surname> <given-names>Y.</given-names></name> <name><surname>Schultz</surname> <given-names>N.</given-names></name> <name><surname>Sander</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>Emerging landscape of oncogenic signatures across human cancers</article-title>. <source>Nat. Genet.</source> <volume>45</volume>, <fpage>1127</fpage>&#x02013;<lpage>1133</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2762</pub-id><pub-id pub-id-type="pmid">24071851</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cover</surname> <given-names>T. M.</given-names></name> <name><surname>Thomas</surname> <given-names>J. A.</given-names></name></person-group> (<year>1991</year>). <source>Elements of Information Theory</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>John Wiley and Sons</publisher-name>.</citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>W.</given-names></name> <name><surname>Elemento</surname> <given-names>O.</given-names></name></person-group> (<year>2015</year>). <article-title>Cancer systems biology: embracing complexity to develop better anticancer therapeutic strategies</article-title>. <source>Oncogene</source> <volume>34</volume>, <fpage>3215</fpage>&#x02013;<lpage>3225</lpage>. <pub-id pub-id-type="doi">10.1038/onc.2014.291</pub-id><pub-id pub-id-type="pmid">25220419</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elia</surname> <given-names>J.</given-names></name> <name><surname>Gai</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>H. M.</given-names></name> <name><surname>Perin</surname> <given-names>J. C.</given-names></name> <name><surname>Geiger</surname> <given-names>E.</given-names></name> <name><surname>Glessner</surname> <given-names>J. T.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Rare structural variants found in attention-deficit hyperactivity disorder are preferentially associated with neurodevelopmen- tal genes</article-title>. <source>Mol. Psychiatry</source> <volume>15</volume>, <fpage>637</fpage>&#x02013;<lpage>646</lpage>. <pub-id pub-id-type="doi">10.1038/mp.2009.57</pub-id><pub-id pub-id-type="pmid">19546859</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Est&#x000E9;vez</surname> <given-names>P. A.</given-names></name> <name><surname>Tesmer</surname> <given-names>M.</given-names></name> <name><surname>Perez</surname> <given-names>C. A.</given-names></name> <name><surname>Zurada</surname> <given-names>J. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Normalized Mutual Information Gene selection</article-title>. <source>IEEE Trans. Neural Netw</source>.<volume>20</volume>, <fpage>189</fpage>&#x02013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.1109/TNN.2008.2005601</pub-id><pub-id pub-id-type="pmid">19150792</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Fayyad</surname> <given-names>U. M.</given-names></name> <name><surname>Irani</surname> <given-names>K. B.</given-names></name></person-group> (<year>1993</year>). <article-title>Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning</article-title>, in <source>Pro-ceedings of International Joint Conference on Artificial Intel- ligence, pp</source> <fpage>1022</fpage>&#x02013;<lpage>1027</lpage></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Flierl</surname> <given-names>A.</given-names></name> <name><surname>Oliveira Lu&#x000ED;s</surname> <given-names>M. A.</given-names></name> <name><surname>Falomir-Lockhart Lisandro</surname> <given-names>J.</given-names></name> <name><surname>Mak Sally</surname> <given-names>K.</given-names></name> <name><surname>Hesley</surname> <given-names>J.</given-names></name> <name><surname>Soldner</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Higher vulnerability and stress sensitivity of neuronal precursor cells carrying an alpha-synuclein gene triplication</article-title>. <source>PLoS ONE</source> <volume>9</volume>, <fpage>e112413</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0112413</pub-id><pub-id pub-id-type="pmid">25390032</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Foithong</surname> <given-names>S.</given-names></name> <name><surname>Pinngern</surname> <given-names>O.</given-names></name> <name><surname>Attachoo</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <article-title>Feature subset selection wrapper based on mutual information and rough sets</article-title>. <source>Expert Syst. Appl.</source> <volume>39</volume>, <fpage>574</fpage>&#x02013;<lpage>584</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2011.07.048</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Frank</surname> <given-names>B.</given-names></name> <name><surname>Bermejo</surname> <given-names>J. L.</given-names></name> <name><surname>Hemminki</surname> <given-names>K.</given-names></name> <name><surname>Sutter</surname> <given-names>C.</given-names></name> <name><surname>Wappenschmidt</surname> <given-names>B.</given-names></name> <name><surname>Meindl</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>Copy number variant in the candidate tumor suppressor gene MTUS1 and familial breast cancer risk</article-title>. <source>Carcinogenesis</source> <volume>28</volume>, <fpage>1442</fpage>&#x02013;<lpage>1445</lpage>. <pub-id pub-id-type="doi">10.1093/carcin/bgm033</pub-id><pub-id pub-id-type="pmid">17301065</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Aksoy</surname> <given-names>B. A.</given-names></name> <name><surname>Dogrusoz</surname> <given-names>U.</given-names></name> <name><surname>Dresdner</surname> <given-names>G.</given-names></name> <name><surname>Gross</surname> <given-names>B.</given-names></name> <name><surname>Sumer</surname> <given-names>S. O.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Integrative analysis of complex cancer genomics and clinical profiles using the Cbioportal</article-title>. <source>Sci. Signaling</source> <volume>6</volume>, <fpage>pl1</fpage>&#x02013;<lpage>pl1</lpage>. <pub-id pub-id-type="doi">10.1126/scisignal.2004088</pub-id><pub-id pub-id-type="pmid">23550210</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>W.</given-names></name> <name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name></person-group> (<year>2018b</year>). <article-title>Class-specific mutual information variation for gene selection</article-title>. <source>Pattern Recogn</source>. <volume>79</volume>, <fpage>328</fpage>&#x02013;<lpage>339</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2018.02.020</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>W.</given-names></name> <name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name> <name><surname>He</surname> <given-names>J.</given-names></name></person-group> (<year>2018a</year>). <article-title>Gene selection considering the composition of feature relevancy</article-title>. <source>Pattern Recogn. Lett.</source> <volume>112</volume>, <fpage>70</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2018.06.005</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glubb</surname> <given-names>D. M.</given-names></name> <name><surname>Thompson</surname> <given-names>D. J.</given-names></name> <name><surname>Aben</surname> <given-names>K. K.</given-names></name> <name><surname>Alsulimani</surname> <given-names>A.</given-names></name> <name><surname>Amant</surname> <given-names>F.</given-names></name> <name><surname>Annibali</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Cross-cancer genome-wide association study of endometrial cancer and epithelial ovarian cancer identifies genetic risk regions associated with risk of both cancers</article-title>. <source>Cancer Epidemiol. Biomarkers Prev.</source> <volume>30</volume>, <fpage>217</fpage>&#x02013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1158/1055-9965.EPI-20-0739</pub-id><pub-id pub-id-type="pmid">33144283</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grangeon</surname> <given-names>L.</given-names></name> <name><surname>Cassinari</surname> <given-names>K.</given-names></name> <name><surname>Rousseau</surname> <given-names>S.</given-names></name> <name><surname>Croisile</surname> <given-names>B.</given-names></name> <name><surname>Formaglio</surname> <given-names>M.</given-names></name> <name><surname>Moreaud</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Early-onset cerebral amyloid angiopathy and alzheimer disease related to an app locus triplication</article-title>. <source>Neurol. Genet</source>. <volume>7</volume>, <fpage>e609</fpage>&#x02013;<lpage>e609</lpage>. <pub-id pub-id-type="doi">10.1212/NXG.0000000000000609</pub-id><pub-id pub-id-type="pmid">34532568</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>X.</given-names></name> <name><surname>Guo</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Xiao</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>A gene selection algorithm based on redundancy analysis and interaction weight</article-title>. <source>Appl. Intell.</source> <volume>51</volume>, <fpage>2672</fpage>&#x02013;<lpage>2686</lpage>. <pub-id pub-id-type="doi">10.1007/s10489-020-01936-5</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hall</surname> <given-names>M.</given-names></name> <name><surname>Frank</surname> <given-names>E.</given-names></name> <name><surname>Holmes</surname> <given-names>G.</given-names></name> <name><surname>Pfahringer</surname> <given-names>B.</given-names></name> <name><surname>Reutemann</surname> <given-names>P.</given-names></name> <name><surname>Witten</surname> <given-names>I. H.</given-names></name></person-group> (<year>2009</year>). <article-title>The WEKA data mining software: an update</article-title>. <source>SIGKDD Explor Newsl</source>. <volume>11</volume>, <fpage>10</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1145/1656274.1656278</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heo</surname> <given-names>Y.</given-names></name> <name><surname>Heo</surname> <given-names>J.</given-names></name> <name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>W. J.</given-names></name> <name><surname>Cheong</surname> <given-names>H. S.</given-names></name> <name><surname>Hong</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Difference of copy number variation in blood of patients with lung cancer</article-title>. <source>Int. J. Biol. Markers</source> <volume>36</volume>, <fpage>3</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1177/1724600820980739</pub-id><pub-id pub-id-type="pmid">33307925</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jakulin</surname> <given-names>A.</given-names></name></person-group> (<year>2003</year>). <source>Attribute Interactions in Machine Learning (Master thesis). Computer and Information Science, University of Ljubljana.</source></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jakulin</surname> <given-names>A.</given-names></name> <name><surname>Bratko</surname> <given-names>I.</given-names></name></person-group> (<year>2004</year>). <article-title>Testing the significance of attribute interactions</article-title>, in <source>Proceedings of the Twenty-first international conference on Machine learning - ICML&#x00027;04</source>. <publisher-loc>Banff, AL</publisher-loc>: <publisher-name>ACM Press</publisher-name>. pp. <fpage>409</fpage>&#x02013;<lpage>416</lpage>.</citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Cheng</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Morstatter</surname> <given-names>F.</given-names></name> <name><surname>Trevino</surname> <given-names>R. P.</given-names></name> <name><surname>Tang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Gene selection: a data perspective</article-title>. <source>ACM Comput. Surv</source>. <volume>50</volume>, <fpage>1</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1145/3136625</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>J.</given-names></name> <name><surname>Hou</surname> <given-names>L.</given-names></name> <name><surname>Luan</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). <article-title>Gene selection with conditional mutual information considering feature interaction</article-title>. <source>Symmetry</source> <volume>11</volume>, <fpage>858</fpage>. <pub-id pub-id-type="doi">10.3390/sym11070858</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Dai</surname> <given-names>C.</given-names></name> <name><surname>Shao</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A deep learning framework to predict tumor tissue-of-origin based on copy number alteration</article-title>. <source>Front. Bioeng. Biotech</source>. <volume>8</volume>, <fpage>701</fpage>. <pub-id pub-id-type="doi">10.3389/fbioe.2020.00701</pub-id><pub-id pub-id-type="pmid">32850687</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>Z.</given-names></name></person-group> (<year>2011</year>). <article-title>Mutual information is copula entropy</article-title>. <source>Tsinghua Sci. Technol</source>. <volume>16</volume>, <fpage>51</fpage>&#x02013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1016/S1007-0214(11)70008-6</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mermel</surname> <given-names>C. H.</given-names></name> <name><surname>Schumacher</surname> <given-names>S. E.</given-names></name> <name><surname>Hill</surname> <given-names>B.</given-names></name> <name><surname>Meyerson</surname> <given-names>M. L.</given-names></name> <name><surname>Beroukhim</surname> <given-names>R.</given-names></name> <name><surname>Getz</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <article-title>GISTIC2.0 facilitates sensitive and confident localization of the targets of focal somatic copy-number al teration in human cancers</article-title>. <source>Genome Biol.</source> <volume>12</volume>, <fpage>4</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2011-12-4-r41</pub-id><pub-id pub-id-type="pmid">21527027</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orsenigo</surname> <given-names>C.</given-names></name> <name><surname>Vercellis</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>A comparative study of non-linear manifold learning methods for cancer microarray data classification</article-title>. <source>Expert Syst. Appl.</source> <volume>40</volume>, <fpage>2189</fpage>&#x02013;<lpage>2197</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2012.10.044</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pandey</surname> <given-names>G. N.</given-names></name> <name><surname>Rizavi</surname> <given-names>H. S.</given-names></name> <name><surname>Tripathi</surname> <given-names>M.</given-names></name> <name><surname>Ren</surname> <given-names>X.</given-names></name></person-group> (<year>2015</year>). <article-title>Region-specific dysregulation of glycogen synthase kinase-3&#x003B2; and &#x003B2;-catenin in the postmortem brains of subjects with bipolar disorder and schizophrenia</article-title>. <source>Bipolar Disord</source>. <volume>17</volume>, <fpage>160</fpage>&#x02013;<lpage>171</lpage>. <pub-id pub-id-type="doi">10.1111/bdi.12228</pub-id><pub-id pub-id-type="pmid">25041379</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>H.</given-names></name> <name><surname>Long</surname> <given-names>F.</given-names></name> <name><surname>Ding</surname> <given-names>C.</given-names></name></person-group> (<year>2005</year>). <article-title>Gene selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy</article-title>. <source>IEEE T. Pattern Anal</source>. <volume>27</volume>, <fpage>1226</fpage>&#x02013;<lpage>1238</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2005.159</pub-id><pub-id pub-id-type="pmid">16119262</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Redon</surname> <given-names>R.</given-names></name> <name><surname>Ishikawa</surname> <given-names>S.</given-names></name> <name><surname>Fitch</surname> <given-names>K. R.</given-names></name> <name><surname>Feuk</surname> <given-names>L.</given-names></name> <name><surname>Perry</surname> <given-names>G. H.</given-names></name> <name><surname>Andrews</surname> <given-names>T. D.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>Global variation in copy number in the human genome</article-title>. <source>Nature</source> <volume>444</volume>, <fpage>444</fpage>&#x02013;<lpage>454</lpage>. <pub-id pub-id-type="doi">10.1038/nature05329</pub-id><pub-id pub-id-type="pmid">17122850</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rodriguez</surname> <given-names>A. C.</given-names></name> <name><surname>Blanchard</surname> <given-names>Z.</given-names></name> <name><surname>Maurer</surname> <given-names>K. A.</given-names></name> <name><surname>Gertz</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Estrogen signaling in endometrial cancer: a key oncogenic pathway with several open questions</article-title>. <source>HORM. CANC</source>. <volume>10</volume>, <fpage>51</fpage>&#x02013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1007/s12672-019-0358-9</pub-id><pub-id pub-id-type="pmid">30712080</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shannon</surname> <given-names>C. E.</given-names></name></person-group> (<year>2001</year>). <article-title>A mathematical theory of communication</article-title>. <source>SIGMOBILE Mob. Comput. Commun. Rev</source>. <volume>5</volume>, <fpage>3</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1145/584091.584093</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Xie</surname> <given-names>C.</given-names></name> <name><surname>Yin</surname> <given-names>S.</given-names></name> <name><surname>Shi</surname> <given-names>D.</given-names></name> <name><surname>Wei</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title><italic>BRCA1</italic> and <italic>BRCA2</italic> mutations in ovarian cancer patients from China: ethnic-related mutations in <italic>BRCA1</italic> associated with an increased risk of ovarian cancer: <italic>BRCA1/2</italic> mutation in Chinese ovarian cancer</article-title>. <source>Int. J. Cancer</source> <volume>140</volume>, <fpage>2051</fpage>&#x02013;<lpage>2059</lpage>. <pub-id pub-id-type="doi">10.1002/ijc.30633</pub-id><pub-id pub-id-type="pmid">28176296</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name></person-group> (<year>2013</year>). <article-title>Gene selection using dynamic weights for classification</article-title>. <source>Knowl. Based Syst.</source> <volume>37</volume>, <fpage>541</fpage>&#x02013;<lpage>549</lpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2012.10.001</pub-id><pub-id pub-id-type="pmid">12424115</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>T.</given-names></name> <name><surname>Bi</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Copy number variation of ubiquitin- specific proteases genes in blood leukocytes and colorectal cancer</article-title>. <source>Cancer Biol. Ther</source>. <volume>21</volume>, <fpage>637</fpage>&#x02013;<lpage>646</lpage>. <pub-id pub-id-type="doi">10.1080/15384047.2020.1750860</pub-id><pub-id pub-id-type="pmid">32364424</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Bockstal</surname> <given-names>M. R.</given-names></name> <name><surname>Agahozo</surname> <given-names>M. C.</given-names></name> <name><surname>van Marion</surname> <given-names>R.</given-names></name> <name><surname>Atmodimedjo</surname> <given-names>P. N.</given-names></name> <name><surname>Sleddens</surname> <given-names>H. F. B. M.</given-names></name> <name><surname>Dinjens</surname> <given-names>W. N. M.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Somatic mutations and copy number variations in breast cancers with heterogeneous <italic>HER2</italic> amplification</article-title>. <source>Mol. Oncol</source>. <volume>14</volume>, <fpage>671</fpage>&#x02013;<lpage>685</lpage>. <pub-id pub-id-type="doi">10.1002/1878-0261.12650</pub-id><pub-id pub-id-type="pmid">32058674</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Wei</surname> <given-names>J. M.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>S. Q.</given-names></name></person-group> (<year>2017</year>). <article-title>Gene selection by Maximizing Independent Classification Information</article-title>. <source>IEEE Trans. Knowl. Data Eng.</source> <volume>29</volume>, <fpage>828</fpage>&#x02013;<lpage>841</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2017.2650906</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Witten</surname> <given-names>I. H.</given-names></name> <name><surname>Frank</surname> <given-names>E.</given-names></name></person-group> (<year>2002</year>). <article-title>Data mining: practical machine learning tools and techniques with Java implementations</article-title>. <source>SIGMOD Rec</source>. <volume>31</volume>, <fpage>76</fpage>&#x02013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1145/507338.507355</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>S.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Discernible neighborhood counting based incremental gene selection for heterogeneous data</article-title>. <source>Int. J. Mach. Learn. Cybern</source>. <volume>11</volume>, <fpage>1115</fpage>&#x02013;<lpage>1127</lpage>. <pub-id pub-id-type="doi">10.1007/s13042-019-00997-4</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>R.</given-names></name> <name><surname>Yin</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>A novel gene selection method considering feature interaction</article-title>. <source>Pattern Recogn</source>. <volume>48</volume>, <fpage>2656</fpage>&#x02013;<lpage>2666</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2015.02.025</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>N.</given-names></name> <name><surname>Wang</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>P.</given-names></name> <name><surname>Huang</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Classification of cancers based on copy number variation landscapes</article-title>. <source>Biochim. Biophys. Acta, Gen. Subj</source>. <volume>1860</volume>, <fpage>2750</fpage>&#x02013;<lpage>2755</lpage>. <pub-id pub-id-type="doi">10.1016/j.bbagen.2016.06.003</pub-id><pub-id pub-id-type="pmid">27266344</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>Z.</given-names></name> <name><surname>Yu</surname> <given-names>R.</given-names></name> <name><surname>Gao</surname> <given-names>C.</given-names></name> <name><surname>Jian</surname> <given-names>X.</given-names></name> <name><surname>Quan</surname> <given-names>S.</given-names></name> <name><surname>Xing</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Low copy number of FCGR3B is associated with lupus nephritis in a Chinese population</article-title>. <source>Exp. Ther. Med</source>. <volume>14</volume>, <fpage>4497</fpage>&#x02013;<lpage>4502</lpage>. <pub-id pub-id-type="doi">10.3892/etm.2017.5069</pub-id><pub-id pub-id-type="pmid">29104657</pub-id></citation></ref>
</ref-list>
</back>
</article>