<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">864564</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.864564</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>SGII: Systematic Identification of Essential lncRNAs in Mouse and Human Genome With lncRNA-Protein-Protein Heterogeneous Interaction Network</article-title>
<alt-title alt-title-type="left-running-head">Xin et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Essential lncRNA Identification</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Xin</surname>
<given-names>Xiao-Hong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1656770/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Ying-Ying</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Gao</surname>
<given-names>Chu-Qiao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1499608/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Min</surname>
<given-names>Hui</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1692080/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wang</surname>
<given-names>Likun</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1665208/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Du</surname>
<given-names>Pu-Feng</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/778584/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>College of Intelligence and Computing</institution>, <institution>Tianjin University</institution>, <addr-line>Tianjin</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Institute of Systems Biomedicine, Department of Pathology, School of Basic Medical Sciences, Beijing Key Laboratory of Tumor Systems Biology, Peking-Tsinghua Center of Life Sciences, Peking University Health Science Center</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/552766/overview">Tao Huang</ext-link>, Shanghai Institute of Nutrition and Health (CAS), China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/861003/overview">Xiucai Ye</ext-link>, University of Tsukuba, Japan</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/459395/overview">Qi Zhao</ext-link>, University of Science and Technology Liaoning, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Likun Wang, <email>wanglk@bjmu.edu.cn</email>; Pu-Feng Du, <email>pdu@tju.edu.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>03</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>864564</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Xin, Zhang, Gao, Min, Wang and Du.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Xin, Zhang, Gao, Min, Wang and Du</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Long noncoding RNAs (lncRNAs) play important roles in a variety of biological processes. Knocking out or knocking down some lncRNA genes can lead to death or infertility. These lncRNAs are called essential lncRNAs. Identifying the essential lncRNA is of importance for complex disease diagnosis and treatments. However, experimental methods for identifying essential lncRNAs are always costly and time consuming. Therefore, computational methods can be considered as an alternative approach. We propose a method to identify essential lncRNAs by combining network centrality measures and lncRNA sequence information. By constructing a lncRNA-protein-protein interaction network, we measure the essentiality of lncRNAs from their role in the network and their sequence together. We name our method as the systematic gene importance index (SGII). As far as we can tell, this is the first attempt to identify essential lncRNAs by combining sequence and network information together. The results of our method indicated that essential lncRNAs have similar roles in the LPPI network as the essential coding genes in the PPI network. Another encouraging observation is that the network information can significantly boost the predictive performance of sequence-based method. All source code and dataset of SGII have been deposited in a GitHub repository (<ext-link ext-link-type="uri" xlink:href="https://github.com/ninglolo/SGII">https://github.com/ninglolo/SGII</ext-link>).</p>
</abstract>
<kwd-group>
<kwd>essential lncRNA</kwd>
<kwd>lncRNA-protein interaction network</kwd>
<kwd>protein-protein interaction network</kwd>
<kwd>network centrality</kwd>
<kwd>systematic method</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">National Key Research and Development Program of China<named-content content-type="fundref-id">10.13039/501100012166</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>Long noncoding RNAs (lncRNAs) refer to non-coding RNAs with a length over 200&#xa0;nt. LncRNAs play a major role in epigenetic control, cell differentiation, autophagy, apoptosis, and embryonic development (<xref ref-type="bibr" rid="B23">Mercer et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B27">Rinn and Chang, 2012</xref>; <xref ref-type="bibr" rid="B3">Chen, 2016</xref>). Many cellular processes are regulated by lncRNAs. For examples, RNA splicing, translation, and signal transductions are related to lncRNA regulations (<xref ref-type="bibr" rid="B14">Khalil and Rinn, 2011</xref>; <xref ref-type="bibr" rid="B4">Da Sacco et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B46">Zhu et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B10">Hu et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B39">Zhang et&#x20;al., 2018</xref>, <xref ref-type="bibr" rid="B41">2021</xref>; <xref ref-type="bibr" rid="B19">Li et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B26">Pyfrom et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B43">Zhao et&#x20;al., 2020</xref>). In addition, lncRNAs are related to a variety of complex diseases, including cancers, nervous system diseases, and cardiovascular diseases (<xref ref-type="bibr" rid="B5">Fenoglio et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B31">Uchida and Dimmeler, 2015</xref>; <xref ref-type="bibr" rid="B30">Schmitt and Chang, 2016</xref>).</p>
<p>Knocking out or knocking down some lncRNA genes can lead to death or infertility. These lncRNAs are called essential lncRNAs, which are of vital importance for survival and development. Identification of essential lncRNAs provides insight into the minimum requirements of normal cell functioning and normal organism development. Experimental methods have been applied to identify essential lncRNAs. Li <italic>et&#x20;al.</italic> established the single lncRNA knockout mouse model, as well as the multiple lncRNA knockout mouse model (<xref ref-type="bibr" rid="B17">Li and Chang, 2014</xref>). By large-scale phenotypic analysis, they found that knocking out lncRNAs, such as <italic>Fendrr</italic>, <italic>Peril</italic>, and <italic>Mdgt</italic>, showed perinatal and postpartum lethality (<xref ref-type="bibr" rid="B17">Li and Chang, 2014</xref>). Watanabe <italic>et&#x20;al.</italic> found that <italic>Dnm3os</italic> has an essential role in the normal growth and bone development of mice (<xref ref-type="bibr" rid="B34">Watanabe et&#x20;al., 2008</xref>). Zhou <italic>et&#x20;al.</italic> proposed that <italic>Meg3</italic> deletion in female rats can result in skeletal muscle defect and perinatal death (<xref ref-type="bibr" rid="B45">Zhou et&#x20;al., 2012</xref>). These studies provide helpful insights for identifying essential lncRNAs. However, experimental methods for identifying essential lncRNA genes are not always feasible due to many factors, which may also produce misleading results (<xref ref-type="bibr" rid="B11">Jathar et&#x20;al., 2017</xref>). Therefore, computational approaches are considered as alternative&#x20;ways.</p>
<p>Computing essentiality of a coding gene has been widely studied. Most of the existing methods define the essentiality measures based on the topological importance of a protein in protein-protein interaction networks. Various types of centralities have been introduced in this regard. For example, Jeong <italic>et&#x20;al.</italic> found that hub nodes with high connections in the protein-protein interaction (PPI) network are often indispensable, which allows them to use the degree centrality (DC) to identify essential proteins (<xref ref-type="bibr" rid="B12">Jeong et&#x20;al., 2001</xref>). Joey <italic>et&#x20;al.</italic> introduced the betweenness centrality (BC) to measure the essentiality of proteins, as they found that PPI network is modularized (<xref ref-type="bibr" rid="B13">Joy et&#x20;al., 2005</xref>). Wang <italic>et&#x20;al.</italic> used eigenvector centrality (EC) to predict essential proteins, which measures the importance of nodes by calculating the connection with high index nodes in the network (<xref ref-type="bibr" rid="B33">Wang et&#x20;al., 2013</xref>). Wuchty <italic>et&#x20;al.</italic> found that closeness centrality (CC) measure using local information is useful in predicting essential proteins (<xref ref-type="bibr" rid="B35">Wuchty and Stadler, 2003</xref>). Many more methods have tried to incorporate different types of information in predicting essential proteins (<xref ref-type="bibr" rid="B18">Li et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B32">Wang et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B44">Zhong et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B2">Campos et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B40">Zhang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B20">Liu et&#x20;al., 2021</xref>). However, these centrality measures are not always working, due to incomplete protein-protein networks and frequent false-positives in the high-throughput experiments for identifying protein-protein interactions. Therefore, sequence-based methods were also considered. Zeng <italic>et&#x20;al.</italic> defined the Gene Importance Calculator (GIC) score using only genomic sequence information (<xref ref-type="bibr" rid="B37">Zeng et&#x20;al., 2018</xref>). The GIC score was derived from a logistic regression model. It can score not only coding genes but also non-coding&#x20;genes.</p>
<p>As far as we can tell, the GIC score is the only available essentiality measure that can be applied on non-coding genes, including lncRNA genes (<xref ref-type="bibr" rid="B37">Zeng et&#x20;al., 2018</xref>). However, the design of the GIC score ignored all information that is buried in the lncRNA-protein interactions (LPI). We believe that the LPI information has a similar role in identifying essential lncRNAs to that of PPI in identifying essential coding&#x20;genes.</p>
<p>With the development of high-throughput experimental technologies, many databases have been established for non-coding genes and their interactions. The NPInter database provides a comprehensive archive of molecular interactions involving noncoding RNAs(<xref ref-type="bibr" rid="B8">Hao et&#x20;al., 2016</xref>). NONCODE database is an integrated knowledge database dedicated to non-coding RNAs and their annotations (<xref ref-type="bibr" rid="B42">Zhao et&#x20;al., 2016</xref>). However, essential gene databases, like the DEG database, focus more on recording essential coding genes (<xref ref-type="bibr" rid="B38">Zhang et&#x20;al., 2004</xref>). The essential non-coding genes are rarely recorded, particularly for complex organisms, like human and mouse. This is a primary challenge in developing a systematic method for measuring essentiality of non-coding&#x20;genes.</p>
<p>By curating data from various literatures, as well as public databases, we established a dataset as the basis for developing a computational method to measure non-coding gene essentiality. In this work, we proposed the systematic gene importance index (SGII) by combining various centralities on the lncRNA-protein-protein heterogeneous network and sequence-based essentiality scores. By comparing our measure to both network-based methods and sequence-based method, we found that network information can boost the sequence-based method significantly.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>Materials and Methods</title>
<sec id="s2-1">
<title>Dataset Curation</title>
<p>We downloaded human and mouse lncRNA-protein interactions from the NPInter database v4.0 (<xref ref-type="bibr" rid="B8">Hao et&#x20;al., 2016</xref>). Self-interactions and duplicates were removed. The mouse lncRNA-protein interaction network involves 33255 lncRNAs, 182 proteins, and 102051 interactions. The human lncRNA-protein interaction network contains 41589 lncRNAs, 3237 proteins, and 394895 interactions. We downloaded human and mouse protein-protein interaction data from BioGrid database version 4.4 (<xref ref-type="bibr" rid="B24">Oughtred et&#x20;al., 2021</xref>). The mouse protein-protein interaction network includes 9744 proteins and 52342 interactions. The human protein-protein interaction network includes 19106 proteins and 644235 interactions.</p>
<p>We combine the lncRNA-protein interactions and protein-protein interactions by matching the name of the proteins in both datasets, producing a heterogeneous network with two types of interactions. The mouse network was composed by 9845 proteins and 33255 lncRNAs with 102051&#x20;lncRNA-protein interactions and 52342&#x20;protein-protein interactions. The human network was composed by 19553 proteins and 41589 lncRNAs with 394895&#x20;lncRNA-protein interactions and 644235&#x20;protein-protein interactions. The sequences of all lncRNAs in both human and mouse interaction networks were obtained from the NONCODE database version 5 (<xref ref-type="bibr" rid="B42">Zhao et&#x20;al., 2016</xref>).</p>
<p>According to literatures (<xref ref-type="bibr" rid="B25">Penny et&#x20;al., 1996</xref>; <xref ref-type="bibr" rid="B22">Marahrens et&#x20;al., 1997</xref>; <xref ref-type="bibr" rid="B16">Lee, 2000</xref>; <xref ref-type="bibr" rid="B28">Sado et&#x20;al., 2001</xref>; <xref ref-type="bibr" rid="B7">Grote et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B15">Klattenhoff et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B29">Sauvageau et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B36">Yildirim et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B37">Zeng et&#x20;al., 2018</xref>), eight mouse lncRNAs, including <italic>Xist</italic>, <italic>Gas</italic>5, <italic>Meg</italic>3, <italic>Tsix</italic>, <italic>Gt</italic> (ROSA) 26<italic>Sor</italic>, <italic>Dnm</italic>3<italic>os</italic>, <italic>Fendrr</italic>, and <italic>Braveheart</italic>, were identified as essential lncRNAs. The remaining 33247 lncRNAs in the mouse network were marked with unknown status. For human lncRNAs, we curated a set of lncRNAs that are reported to be essential in various conditions from literatures (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). This set contains 63 lncRNAs. The names of these lncRNAs and the conditions that they are reported to be essential, are listed in <xref ref-type="sec" rid="s10">Supplementary Table S1</xref>, along with literatures of the original reports. In addition, 11 mouse lncRNAs, which are homologous of human essential lncRNAs, were also collected for validation purpose, as homologous usually have similar essentiality (<xref ref-type="bibr" rid="B6">Georgi et&#x20;al., 2013</xref>).</p>
</sec>
<sec id="s2-2">
<title>Gene Importance Calculator</title>
<p>Gene Importance Calculator (GIC) (<xref ref-type="bibr" rid="B37">Zeng et&#x20;al., 2018</xref>) is a useful essentiality indicator for both protein-coding genes and noncoding genes. It is based solely on sequence information. The GIC score (<italic>g</italic>) is defined as follows:<disp-formula id="e1">
<mml:math id="m1">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>&#x3b8;</italic>(<italic>p</italic>) is derived from a logistic regression model. <italic>&#x3b8;</italic>(<italic>p</italic>) can be defined as<disp-formula id="e2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfrac>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mi>L</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>L</mml:mi>
</mml:mfrac>
<mml:mi>e</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>5</mml:mn>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>&#x3b1;</italic>
<sub>1</sub>, <italic>&#x3b1;</italic>
<sub>2</sub>, &#x2026;, <italic>&#x3b1;</italic>
<sub>5</sub>, <italic>&#x3b2;</italic>
<sub>0</sub>, <italic>&#x3b2;</italic>
<sub>1</sub> and <italic>&#x3b2;</italic>
<sub>2</sub> are regression coefficients, <italic>L</italic> the length of RNA sequence, <italic>e</italic> the minimum free energy of RNA secondary structure, <italic>p</italic> the conditional probability that a gene is essential, and <italic>f</italic>
<sub>
<italic>i</italic>
</sub> the occurrence frequency of a triplet in the sequence. The five types of triplets, which are considered in the GIC, are CGA, GCG, TCG, ACG and TCA (<xref ref-type="bibr" rid="B37">Zeng et&#x20;al., 2018</xref>).</p>
<p>When calculating the GIC score, we need to use the external program RNAfold (<xref ref-type="bibr" rid="B21">Lorenz et&#x20;al., 2011</xref>), which requires a sequence length less than 20000&#xa0;nt. Therefore, only 24450 mouse lncRNAs and 29481 human lncRNAs can be calculated for GIC. All other lncRNAs have lengths too long for the RNAfold to&#x20;work.</p>
</sec>
<sec id="s2-3">
<title>Network Centralities</title>
<p>We formulate the heterogeneous graph as <bold>
<italic>G</italic>
</bold> &#x3d; (<bold>
<italic>V</italic>
</bold>, <bold>
<italic>E</italic>
</bold>), where <bold>
<italic>V</italic>
</bold> is the set of all nodes, including lncRNAs and proteins, and <bold>
<italic>E</italic>
</bold> the set of all interactions, including lncRNA-protein and protein-protein interactions. Without losing generality, we note the number of all nodes as <italic>n</italic>. The network can be represented as an adjacency matrix <bold>A</bold>&#x2208;{0.1}<sup>
<italic>n</italic>&#xd7;<italic>n</italic>
</sup>. The element on the <italic>i</italic>th row and the <italic>j</italic>th column of <bold>A</bold> can be denoted as <italic>a</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub>. If <italic>a</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> &#x3d; 1, the <italic>i</italic>th node and the <italic>j</italic>th node have interactions between them. If <italic>a</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> &#x3d; 0, there is no interaction between the <italic>i</italic>th node and the <italic>j</italic>th node. Given <italic>a</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub>, we can define four different centrality measures, including degree centrality (DC), betweenness centrality (BC), closeness centrality (CC), and eigenvector centrality (EC) for each node in the network.</p>
<p>The degree centrality of the <italic>i</italic>th node can be defined as follows:<disp-formula id="e3">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>The betweenness centrality of the <italic>i</italic>th node can be defined as follows:<disp-formula id="e4">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2260;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x2260;</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold">V</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>&#x3c3;</italic>
<sub>
<italic>u</italic>,<italic>v</italic>
</sub> is the number of shortest paths between the <italic>u</italic>th node and the <italic>v</italic>th node, and <italic>&#x3c3;</italic>
<sub>
<italic>u</italic>,<italic>v</italic>
</sub>(<italic>i</italic>) the number of shortest paths between the <italic>u</italic>th node and the <italic>v</italic>th node that pass the <italic>i</italic>th&#x20;node.</p>
<p>The closeness centrality of the <italic>i</italic>th node is defined as follows:<disp-formula id="e5">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold">R</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold">R</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>where <bold>
<italic>R</italic>
</bold>(<italic>i</italic>) is the set of nodes that can reach the <italic>i</italic>th node, <italic>d</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> the length of the shortest path between the <italic>i</italic>th node and the <italic>j</italic>th node, and &#x7c;.&#x7c; cardinal operator of a&#x20;set.</p>
<p>The eigenvector centrality of the <italic>i</italic>th node is defined as follows:<disp-formula id="e6">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <italic>x</italic>
<sub>max</sub>(<italic>i</italic>) is the <italic>i</italic>th dimension of the normalized eigen vector <bold>x</bold> that corresponds to the largest eigen value of adjacency matrix <bold>A</bold>. Let <italic>&#x3bb;</italic>
<sub>max</sub> be the largest eigen value of <bold>A</bold>, the following relationships are satisfied in finding <bold>x</bold>:<disp-formula id="e7">
<mml:math id="m7">
<mml:mrow>
<mml:mi mathvariant="bold">Ax</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mi mathvariant="normal">,</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi mathvariant="normal">and</mml:mi>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m8">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi mathvariant="bold">x</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where &#x7c;&#x7c;.&#x7c;&#x7c; is the vector norm operator.</p>
</sec>
<sec id="s2-4">
<title>Systematic Gene Importance Index</title>
<p>Our network model contains two types of nodes, lncRNAs, and proteins. It also involves two types of interactions, the lncRNA-protein interactions and protein-protein interactions. Essentially, it is a lncRNA-protein-protein interaction (LPPI) heterogeneous network. <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> illustrates a part of the LPPI network for human and mouse respectively.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>A part of the LPPI network. <bold>(A)</bold> Human dataset. The network contains the lncRNA <italic>CRNDE</italic> and 14 interacting proteins; <bold>(B)</bold> Mouse dataset. The network contains the lncRNA <italic>Xist</italic> and 14 interacting proteins.</p>
</caption>
<graphic xlink:href="fgene-13-864564-g001.tif"/>
</fig>
<p>We propose the Systematic Gene Importance Index (SGII) as a comprehensive measure of gene essentiality, particularly for non-coding genes. SGII is a combination of the sequence-based GIC score and centrality measures, which have been elaborated as&#x20;above.</p>
<p>For the <italic>i</italic>th node in the LPPI network, we compute its BC, CC, DC and EC, which can be noted as <italic>b</italic>
<sub>
<italic>i</italic>
</sub>, <italic>c</italic>
<sub>
<italic>i</italic>
</sub>, <italic>d</italic>
<sub>
<italic>i</italic>
</sub> and <italic>e</italic>
<sub>
<italic>i</italic>
</sub>, respectively. Its GIC score is noted as <italic>g</italic>
<sub>
<italic>i</italic>
</sub>. We sort all nodes according to their BC, CC, DC, EC and GIC in a descending order, respectively. The rank of the <italic>i</italic>th node after sorting according to BC, CC, DC, EC and GIC can be noted as <italic>r</italic>
<sub>
<italic>b</italic>
</sub>(<italic>i</italic>), <italic>r</italic>
<sub>
<italic>c</italic>
</sub>(<italic>i</italic>), <italic>r</italic>
<sub>
<italic>d</italic>
</sub>(<italic>i</italic>), <italic>r</italic>
<sub>
<italic>e</italic>
</sub>(<italic>i</italic>) and <italic>r</italic>
<sub>
<italic>g</italic>
</sub>(<italic>i</italic>), respectively.</p>
<p>Let <italic>s</italic>
<sub>
<italic>i</italic>
</sub> be the degree of the <italic>i</italic>th node, which can be computed as follows:<disp-formula id="e9">
<mml:math id="m9">
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>Given a threshold <italic>z</italic>, if <italic>s</italic>
<sub>
<italic>i</italic>
</sub> &#x2265; <italic>z</italic>, the centrality measures will determine the essentiality of a gene directly. For convenience, we define the centrality-based essentiality indicator function for the <italic>i</italic>th node according to BC, CC, DC, and EC respectively as follows:<disp-formula id="e10">
<mml:math id="m10">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>b</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>b</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
<disp-formula id="e11">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<disp-formula id="e12">
<mml:math id="m12">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mtext>and</mml:mtext>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>
<disp-formula id="e13">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>where <italic>k</italic> is a rank threshold parameter. The <italic>i</italic>th node is identified as essential when<disp-formula id="e14">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>b</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>is satisfied.</p>
<p>If <italic>s</italic>
<sub>
<italic>i</italic>
</sub> &#x3c; <italic>z</italic>, we rely on the GIC score to determine the essentiality of a gene. Similarly, we can define the indicator function for GIC ranking, as follows:<disp-formula id="e15">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>where <italic>t</italic> is another rank threshold parameter. The <italic>i</italic>th node is essential if<disp-formula id="e16">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>is satisfied.</p>
<p>The whole flowchart of SGII is illustrated in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The flowchart of SGII. The method SGII consists of two parts. For lncRNAs whose degree is greater than or equal to <italic>z</italic>, four types of centralities were used to determine whether they were essential lncRNAs. For lncRNAs whose degree is less than <italic>z</italic>, GIC was&#x20;used.</p>
</caption>
<graphic xlink:href="fgene-13-864564-g002.tif"/>
</fig>
</sec>
<sec id="s2-5">
<title>Performance Evaluation</title>
<p>In evaluating SGII, we use three statistics to describe its predictive performance. These statistics include sensitivity (<italic>s</italic>), false positive rate (<italic>r</italic>), and Fisher&#x2019;s exact test score (<italic>f</italic>), which are defined as follows:<disp-formula id="e17">
<mml:math id="m17">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
<disp-formula id="e18">
<mml:math id="m18">
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mtext>and</mml:mtext>
</mml:mrow>
</mml:math>
<label>(18)</label>
</disp-formula>
<disp-formula id="e19">
<mml:math id="m19">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>log</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:math>
<label>(19)</label>
</disp-formula>where <italic>n</italic>
<sub>
<italic>t</italic>
</sub> is the number of known essential lncRNAs that are identified as essential lncRNAs, <italic>n</italic>
<sub>&#x2b;</sub> the total number known essential lncRNAs, <italic>n</italic>
<sub>
<italic>f</italic>
</sub> the number of lncRNAs with unknown essentiality that are identified as essential, <italic>n</italic>
<sub>-</sub> the total number of lncRNAs with unknown essentiality and <italic>p</italic> the <italic>p</italic>-value of Fisher&#x2019;s exact test. Since SGII is a direct scoring method with manually configurable cutoff values, no training procedure is involved in the whole process. This is different to machine learning based methods. We cannot treat the above sensitivity and false positive rate as comparable to those in evaluating machine learning methods, as the knowledge of essential lncRNAs is too limited to perform any kind of cross-validations. This is also why we introduced the Fisher&#x2019;s exact test to further quantifying the quality of our results. It will measure how likely a result in whole is random or not. The bigger <italic>f</italic> value is, the results are less likely to be random.</p>
</sec>
<sec id="s2-6">
<title>Parameter Calibration</title>
<p>There are eight parameters in the GIC, which represent all the coefficients in the model built by GIC method. We took all the parameter values from literature (<xref ref-type="bibr" rid="B37">Zeng et&#x20;al., 2018</xref>). The values for the mouse model are <italic>&#x3b2;</italic>
<sub>0</sub> &#x3d; 0.1625, <italic>&#x3b2;</italic>
<sub>1</sub> &#x3d; 2.638 &#xd7; 10<sup>&#x2013;4</sup>, <italic>&#x3b2;</italic>
<sub>2</sub> &#x3d; 2.194, <italic>&#x3b1;</italic>
<sub>1</sub> &#x3d; 19.88 (for CGA), <italic>&#x3b1;</italic>
<sub>2</sub> &#x3d; 37.59 (for GCG), <italic>&#x3b1;</italic>
<sub>3</sub> &#x3d; 50.37 (for TCG), <italic>&#x3b1;</italic>
<sub>4</sub> &#x3d; 35.44 (for ACG), and <italic>&#x3b1;</italic>
<sub>5</sub> &#x3d; -64.66 (for TCA). The values for human model are <italic>&#x3b2;</italic>
<sub>0</sub> &#x3d; 0.7417, <italic>&#x3b2;</italic>
<sub>1</sub> &#x3d; 2.612 &#xd7; 10<sup>&#x2013;4</sup>, <italic>&#x3b2;</italic>
<sub>2</sub> &#x3d; 4.295, <italic>&#x3b1;</italic>
<sub>1</sub> &#x3d; 48.66, <italic>&#x3b1;</italic>
<sub>2</sub> &#x3d; 15.64, <italic>&#x3b1;</italic>
<sub>3</sub> &#x3d; 76.23, <italic>&#x3b1;</italic>
<sub>4</sub> &#x3d; -1.113, and <italic>&#x3b1;</italic>
<sub>5</sub> &#x3d; -60.29.</p>
<p>Three parameters are introduced in combining centralities and GIC, which are noted as <italic>z</italic>, <italic>k</italic> and <italic>t</italic>. We first perform a grid search of <italic>k</italic> and <italic>t</italic> with a given value of <italic>z</italic>. The pairs of <italic>k</italic> and <italic>t</italic>, which maximize the score <italic>f</italic>, are recorded for every different <italic>z</italic>. These values are further sorted to find the best <italic>z</italic>, <italic>k</italic> and <italic>t</italic> combination. When performing the grid search on the mouse dataset, <italic>k</italic>&#x20;&#x3d; 1, 3, 5, 7, 9, and <italic>t</italic>&#x20;&#x3d; 1, 3, 5, 7, 9. When performing the grid search on the human dataset, <italic>k</italic>&#x20;&#x3d; 5, 10, 15, 20, 25 and <italic>t</italic>&#x20;&#x3d; 5, 10, 15, 20, 25. For both datasets, <italic>z</italic>&#x20;&#x3d; 5, 10, 15, 20. Finally, we set <italic>z</italic>&#x20;&#x3d; 15, <italic>k</italic>&#x20;&#x3d; 5, <italic>t</italic>&#x20;&#x3d; 9 for mouse dataset, and <italic>z</italic>&#x20;&#x3d; 5, <italic>k</italic>&#x20;&#x3d; 20, <italic>t</italic>&#x20;&#x3d; 5 for human dataset. All results for different parameters are provided in supplementary materials, as <xref ref-type="sec" rid="s10">Supplementary Table&#x20;S2</xref>.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>Results and Discussions</title>
<sec id="s3-1">
<title>Characters of the lncRNA-Protein-Protein Heterogeneous Network</title>
<p>We first explore the basic statistical characters of the LPPI network. We plot the degree distribution of the mouse and human network respectively in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>. It is intuitively that the distribution of the degree follows the common power law distribution, which is similar to the PPI networks (<xref ref-type="bibr" rid="B12">Jeong et&#x20;al., 2001</xref>). Since in the PPI network, essential proteins are usually rare and with high degrees, we assume that in our LPPI network, the essential lncRNAs have similar properties.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The degree distribution of lncRNAs in mouse network and human network, respectively. <bold>(A)</bold> The degree distribution of lncRNAs in mouse network; <bold>(B)</bold> The degree distribution of lncRNAs in human network. No protein-protein interaction is counted in producing these distributions.</p>
</caption>
<graphic xlink:href="fgene-13-864564-g003.tif"/>
</fig>
<p>As we have mentioned in the method section, several lncRNAs with a length too long to calculate its secondary structure were not counted in our analysis. It becomes a question whether these lncRNAs have preferences to large or small amounts of interactions. We plot the degree distribution with and without those over-length lncRNAs for mouse and human datasets, respectively, in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>. It is hard to find differences on the degree distributions. We therefore believe that, for a lncRNA, its length alone is not a major contributing factor to its interactions in the LPPI network. This also implied that the essentiality, which we believe to be associated with local network structure, has no direct relationship with the length of the lncRNA. These over-length lncRNAs were kept in the network as dummy nodes, which means we did not compute their essentiality at all, regardless of whether they have a degree over the threshold or&#x20;not.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The degree distribution of lncRNAs in mouse network and human network with and without over-length lncRNAs. <bold>(A)</bold> The degree distribution of lncRNAs in mouse network with over-length lncRNAs; <bold>(B)</bold> The degree distribution of lncRNAs in mouse network without over-length lncRNAs; <bold>(C)</bold> The degree distribution of lncRNAs in human network with over-length lncRNAs; <bold>(D)</bold> The degree distribution of lncRNAs in human network without over-length lncRNAs.</p>
</caption>
<graphic xlink:href="fgene-13-864564-g004.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>Integrating Centrality Measures and the GIC Score</title>
<p>
<xref ref-type="fig" rid="F5">Figure&#x20;5</xref> gives scatter plots of GIC pairing with each of the four types of centralities on human and mouse datasets, respectively. For the mouse dataset, the red dots, which represent essential lncRNAs, tend to appear in the top-right part of the plots, while the blue dots, which denote all other lncRNAs, spread much wider. Although the red dots are relatively rare, but their top-right preference is still observable. For human dataset, this preference is not intuitively obvious.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The scatter plots of GIC pairing with each of the four types of centralities on mouse dataset and human dataset, respectively. <bold>(A)</bold> BC pairing with GIC on mouse dataset; <bold>(B)</bold> BC pairing with GIC on human dataset; <bold>(C)</bold> CC pairing with GIC on mouse dataset; <bold>(D)</bold> CC pairing with GIC on human dataset; <bold>(E)</bold> DC pairing with GIC on mouse dataset; <bold>(F)</bold> DC pairing with GIC on human dataset; <bold>(G)</bold> EC pairing with GIC on mouse dataset; <bold>(H)</bold> EC pairing with GIC on human dataset. Red dots represent known essential lncRNAs, while blue dots represented all others. When drawing panel (A), BC scores of mouse <italic>HOTAIR</italic> and <italic>Xist</italic> are too high to be plotted in the scope. Their (BC,GIC) values are (0.01.0.39) and (0.01.0.94). When drawing panel (B), <italic>NEAT1</italic>, <italic>MALAT1</italic>, <italic>U1</italic> are too distant to other dots, so they cannot be reasonably plotted in the scope. Their (BC,GIC) values are (0.03.0.40), (0.01.0.43) and (0.005.0.54).</p>
</caption>
<graphic xlink:href="fgene-13-864564-g005.tif"/>
</fig>
<p>This allows us to carry out further quantitative analysis on combining the centrality measures and the GIC scores. A primary challenge is that the number of known essential lncRNAs is too small for a machine learning algorithm to train on. In addition, some essential lncRNAs are only involved in a very limited number of interactions. For example, the <italic>Braveheart</italic> (Bvht) lncRNA, which is essential, has only one interaction record in the database. We think this may be due to the incomprehensive knowledge of the lncRNA-protein interaction network. As the estimation of centrality measures highly rely on the interaction enrichment of a node in the network, when dealing with a lncRNA with limited number of interactions, we turn to rely on the GIC&#x20;score.</p>
<p>With the settings in the method section, we combined four types of centrality measures and the GIC scores. On the mouse dataset, we identified 2284 essential lncRNAs from altogether 24450 lncRNAs. Among the 2284 lncRNAs, eight lncRNAs are known to be essential, accounting for 100% of all known essential lncRNAs, resulting a <italic>p</italic>-value &#x3d; 5.73 &#xd7; 10<sup>&#x2013;9</sup> (Fisher&#x2019;s exact test). On the human dataset, we identified 5063 essential lncRNAs, from altogether 29481 lncRNAs, Among the 5063 essential lncRNAs, 41 lncRNAs are reported to be essential in various conditions in literatures, accounting for 65% of all curated essential lncRNAs (<italic>p</italic>-value &#x3d; 3.59 &#xd7; 10<sup>&#x2013;17</sup>, Fisher&#x2019;s exact test). This result clearly indicates that our method is effective to identify essential lncRNAs.</p>
</sec>
<sec id="s3-3">
<title>Systematic Comparison Between Different Configurations of SGII</title>
<p>As SGII is the first attempt to combine the network information and sequence information to identify essential lncRNAs, we explore which kind of centrality measure is more capable to identify essential lncRNAs along with the GIC scores. We first plot the distribution density of different centralities and the GIC scores on mouse and human datasets respectively. As in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>, BC and DC centrality measures along with the GIC scores appear to have much better separation than the CC and EC measures on the mouse dataset, while on the human dataset, only BC and DC present an intuitive separation.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The distribution density of different centralities and the GIC scores on mouse dataset and human dataset respectively. <bold>(A)</bold> The distribution density of BC on mouse dataset; <bold>(B)</bold> The distribution density of BC on human dataset; <bold>(C)</bold> The distribution density of CC on mouse dataset; <bold>(D)</bold> The distribution density of CC on human dataset; <bold>(E)</bold> The distribution density of DC on mouse dataset; <bold>(F)</bold> The distribution density of DC on human dataset; <bold>(G)</bold> The distribution density of EC on mouse dataset; <bold>(H)</bold> The distribution density of EC on human dataset; <bold>(I)</bold> The distribution density of GIC on mouse dataset; <bold>(J)</bold> The distribution density of GIC on human dataset. The red bars represent known essential lncRNAs, while the blue bars for all others. The vertical axis for the red bars are on the right side of the panel, while blue on left. When drawing panel (A), BC scores of mouse <italic>HOTAIR</italic> and <italic>Xist</italic> are too high to be plotted in the scope. Their BC values are 0.01 and 0.01. When drawing panel (B), BC scores of human <italic>NEAT1</italic>, <italic>MALAT1</italic>, <italic>U1</italic> are too far to be drawn in the scope. Their BC values are 0.03, 0.01 and 0.005.</p>
</caption>
<graphic xlink:href="fgene-13-864564-g006.tif"/>
</fig>
<p>However, considering the large differences on axis scale for essential lncRNAs and all lncRNAs, these intuitive observations may be misleading. Therefore, we performed a quantitative comparison using eight different conditions, GIC alone, GIC combined with each one of four types of centralities, GIC combined with BC and DC, GIC combined with CC and EC, and GIC combined with all four types of centralities. The parameters of all comparison are optimized as in method section (<xref ref-type="table" rid="T1">Table&#x20;1</xref>).</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Comparison for different configurations of SGII on mouse and human datasets.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Methods</th>
<th align="center">Dataset</th>
<th align="center">Sen<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref> (%)</th>
<th align="center">FPR<xref ref-type="table-fn" rid="Tfn2">
<sup>b</sup>
</xref> (%)</th>
<th align="center">Fisher&#x2019;s exact test score<xref ref-type="table-fn" rid="Tfn3">
<sup>c</sup>
</xref>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">GIC<xref ref-type="table-fn" rid="Tfn4">
<sup>d</sup>
</xref>
</td>
<td align="center">Mouse</td>
<td align="char" char=".">75.00</td>
<td align="char" char=".">8.98</td>
<td align="char" char=".">4.90</td>
</tr>
<tr>
<td align="left">BC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">9.53</td>
<td align="char" char=".">8.16</td>
</tr>
<tr>
<td align="left">CC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">9.48</td>
<td align="char" char=".">8.18</td>
</tr>
<tr>
<td align="left">DC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">9.55</td>
<td align="char" char=".">8.16</td>
</tr>
<tr>
<td align="left">EC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">9.32</td>
<td align="char" char=".">8.24</td>
</tr>
<tr>
<td align="left">BC &#x2b; DC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">87.50</td>
<td align="char" char=".">4.54</td>
<td align="char" char=".">8.50</td>
</tr>
<tr>
<td align="left">CC &#x2b; EC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">9.31</td>
<td align="char" char=".">8.24</td>
</tr>
<tr>
<td align="left">BC &#x2b; CC &#x2b; DC &#x2b; EC &#x2b; GIC</td>
<td align="center">Mouse</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">9.31</td>
<td align="char" char=".">8.24</td>
</tr>
<tr>
<td align="left">GIC</td>
<td align="center">Human</td>
<td align="char" char=".">26.98</td>
<td align="char" char=".">14.97</td>
<td align="char" char=".">1.91</td>
</tr>
<tr>
<td align="left">BC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">66.67</td>
<td align="char" char=".">12.49</td>
<td align="char" char=".">22.61</td>
</tr>
<tr>
<td align="left">CC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">63.49</td>
<td align="char" char=".">16.37</td>
<td align="char" char=".">16.15</td>
</tr>
<tr>
<td align="left">DC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">71.43</td>
<td align="char" char=".">20.51</td>
<td align="char" char=".">17.25</td>
</tr>
<tr>
<td align="left">EC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">66.67</td>
<td align="char" char=".">19.94</td>
<td align="char" char=".">14.90</td>
</tr>
<tr>
<td align="left">BC &#x2b; DC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">71.43</td>
<td align="char" char=".">18.33</td>
<td align="char" char=".">19.23</td>
</tr>
<tr>
<td align="left">CC &#x2b; EC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">65.08</td>
<td align="char" char=".">18.19</td>
<td align="char" char=".">15.45</td>
</tr>
<tr>
<td align="left">BC &#x2b; CC &#x2b; DC &#x2b; EC &#x2b; GIC</td>
<td align="center">Human</td>
<td align="char" char=".">65.08</td>
<td align="char" char=".">17.07</td>
<td align="char" char=".">16.45</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="Tfn1">
<label>a</label>
<p>Sen stands for Sensitivity, as Eq.&#x20;17.</p>
</fn>
<fn id="Tfn2">
<label>b</label>
<p>FPR, stands for False Positive Rate, as Eq.&#x20;18.</p>
</fn>
<fn id="Tfn3">
<label>c</label>
<p>Fisher&#x2019;s Exact Test Score is defined in Eq.&#x20;19.</p>
</fn>
<fn id="Tfn4">
<label>d</label>
<p>When GIC, was used alone, it is applied on all lncRNAs.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The first observation on <xref ref-type="table" rid="T1">Table&#x20;1</xref> is that the best combination of centrality measure and the GIC is not the combination of all four types of centralities. For the mouse dataset, the BC &#x2b; DC &#x2b; GIC method has the best significance level and lowest FPR value. For the human dataset, the BC &#x2b; GIC method reaches the highest significance level. A second to the best significance level is obtained again by BC &#x2b; DC &#x2b; GIC method, with the highest sensitivity value. Therefore, we think that the BC &#x2b; DC &#x2b; GIC may be a better way to identify essential lncRNAs than the current configuration of SGII. This consists with the impression from <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>. However, due to the limited number of available data and current results, it is possible that this observation does not reflect a comprehensive scene of identifying essential lncRNAs. Therefore, we keep the configuration of SGII to combine all four kinds of centralities and the GIC score, for an unbiased way of identifying essential lncRNAs.</p>
</sec>
<sec id="s3-4">
<title>Comparative Analysis Between Human and Mouse Essential lncRNAs</title>
<p>At a closer look to <xref ref-type="table" rid="T1">Table&#x20;1</xref>, it appears that the Fisher&#x2019;s exact test reports much more significant results on both datasets when GIC is combined with centralities, which proves that integration of centrality measures and GIC is effective. Another observation is that SGII gives under-expected sensitivity values on the human dataset. However, the significance levels on the human dataset are generally way higher than that of the mouse dataset. This may be the results of two differences between the mouse and the human datasets. First, the human dataset is collected from literatures of lncRNAs in various conditions, including tumor cell line experiments. Essential lncRNAs, which are identified by one type of cell line experiments, may be different to those from the original essential gene definitions. As direct essential gene experiments on human are not feasible, the quality of the dataset is not comparable to the mouse dataset. This also applies to the coding gene data (<xref ref-type="bibr" rid="B1">Austin et&#x20;al., 2004</xref>). Secondly, the number of essential lncRNAs in the human dataset is roughly eight times of that of mouse dataset. Since the computation process of the significance level is affected by the raw counts, it is anticipated that systematic differences on significance levels&#x20;exist.</p>
<p>To further confirm the above explanations, we performed the following analysis. We find homologous genes of human essential lncRNAs in mouse. According to the studies in coding genes, these genes are likely to also produce essential lncRNAs(<xref ref-type="bibr" rid="B6">Georgi et&#x20;al., 2013</xref>). Altogether 11 homologous genes in mouse were identified as lncRNA genes in the mouse LPPI network. We used SGII to test if we can identify these homolog essential lncRNAs (<xref ref-type="table" rid="T2">Table&#x20;2</xref>).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Performance analysis on mouse homologs to human essential lncRNAs.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Methods</th>
<th align="center">Sen<xref ref-type="table-fn" rid="Tfn5">
<sup>a</sup>
</xref> (%)</th>
<th align="center">FPR<xref ref-type="table-fn" rid="Tfn6">
<sup>b</sup>
</xref> (%)</th>
<th align="center">Fisher&#x2019;s exact test score<xref ref-type="table-fn" rid="Tfn7">
<sup>c</sup>
</xref>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">GIC<xref ref-type="table-fn" rid="Tfn8">
<sup>d</sup>
</xref>
</td>
<td align="char" char=".">45.45</td>
<td align="char" char=".">8.98</td>
<td align="char" char=".">2.77</td>
</tr>
<tr>
<td align="left">BC &#x2b; GIC</td>
<td align="char" char=".">72.73</td>
<td align="char" char=".">1.79</td>
<td align="char" char=".">11.75</td>
</tr>
<tr>
<td align="left">CC &#x2b; GIC</td>
<td align="char" char=".">45.45</td>
<td align="char" char=".">1.22</td>
<td align="char" char=".">6.90</td>
</tr>
<tr>
<td align="left">DC &#x2b; GIC</td>
<td align="char" char=".">54.55</td>
<td align="char" char=".">1.25</td>
<td align="char" char=".">8.75</td>
</tr>
<tr>
<td align="left">EC &#x2b; GIC</td>
<td align="char" char=".">45.45</td>
<td align="char" char=".">1.17</td>
<td align="char" char=".">7.00</td>
</tr>
<tr>
<td align="left">BC &#x2b; DC &#x2b; GIC</td>
<td align="char" char=".">81.82</td>
<td align="char" char=".">5.89</td>
<td align="char" char=".">9.36</td>
</tr>
<tr>
<td align="left">CC &#x2b; EC &#x2b; GIC</td>
<td align="char" char=".">45.45</td>
<td align="char" char=".">1.17</td>
<td align="char" char=".">7.00</td>
</tr>
<tr>
<td align="left">BC &#x2b; CC &#x2b; DC &#x2b; EC &#x2b; GIC</td>
<td align="char" char=".">45.45</td>
<td align="char" char=".">1.17</td>
<td align="char" char=".">7.00</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="Tfn5">
<label>a</label>
<p>Sen stands for Sensitivity, as Eq.&#x20;17.</p>
</fn>
<fn id="Tfn6">
<label>b</label>
<p>FPR, stands for False Positive Rate, as Eq.&#x20;18.</p>
</fn>
<fn id="Tfn7">
<label>c</label>
<p>Fisher&#x2019;s Exact Test Score is defined in Eq.&#x20;19.</p>
</fn>
<fn id="Tfn8">
<label>d</label>
<p>When GIC, was used alone, it is applied on all lncRNAs.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Obviously, sensitivity is dropping in comparison to the mouse essential lncRNAs. However, it should be noted that the FPR is also dropping, which indicates much less false positives. The significance levels remain almost the same as the mouse essential lncRNAs. Again, the BC &#x2b; GIC method obtained the best significance level, while the BC &#x2b; DC &#x2b; GIC method obtained a second to the best significance level with the highest sensitivity. This result confirmed that the significance level difference between human and mouse dataset is largely caused by the raw counts of the dataset. It also suggests that the BC &#x2b; GIC or BC &#x2b; DC &#x2b; GIC method may be a better choice than combining all types of centralities and the GIC&#x20;score.</p>
<p>The importance of BC can be understood intuitively. If we think the cellular system as a system composed of molecules. The interactions between molecules transfer information. A high BC value indicated that the node is critical as an information hub in many shortest paths between other nodes. Therefore, dropping such nodes will easily break many information channels simultaneously, which will eventually destroy the whole system. That makes it an essential node in the network.</p>
<p>For the DC measure, the intrinsic mechanism is similar. The DC measure is directly associated to the degree of a node. If a node with many edges is dropped, it is more likely that the whole network collapses. This consists with the observations in coding genes. In addition, although some other kinds of centralities, like the NC (new centrality) (<xref ref-type="bibr" rid="B32">Wang et&#x20;al., 2012</xref>), can identify essential coding genes better, it does not work well in non-coding genes. This is an expected result. For NC to work in the LPPI network, it requires that dense interactions exist among the proteins that interacting the same lncRNAs. However, we did not observe this phenomenon in our dataset. The NC is difficult to be estimated for many lncRNAs, due to lacking such kind of interactions.</p>
</sec>
<sec id="s3-5">
<title>Functional Analysis of Essential lncRNA in the Mouse Genome</title>
<p>We took the essential lncRNA gene in mouse genome for functional analysis. For every lncRNA that was predicted as essential in mouse genome, we first map this lncRNA to the Ensembl database (<xref ref-type="bibr" rid="B9">Howe et&#x20;al., 2021</xref>) using either gene name or sequence information. The mapped genes are then uploaded to the Gene Ontology online system for functional enrichment analysis. The top three enrichment of functions are &#x201c;nucleic acid binding&#x201d; (GO:0003676), &#x201c;heterocyclic compound binding&#x201d; (GO:1901363) and &#x201c;organic cyclic compound binding&#x201d; (GO:0097159). As we have mentioned, this is expected for lncRNAs. They realize their functions through bindings with other molecules.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s4">
<title>Conclusion</title>
<p>SGII is the first attempt to combine lncRNA-protein interactions and lncRNA sequence information for identifying essential non-coding RNAs. Since the study on collecting and identifying essential coding genes has been performed for over a decade, it is time to step forward to the essentiality of non-coding genes, as non-coding genes are much more common than coding genes in mouse and human genomes. Due to the limited number of known essential lncRNAs, SGII does not use conventional machine learning algorithms, but applies simple scoring schemes and statistical tests. By combining BC, CC, DC, EC and GIC scores, SGII achieved a better performance than using only sequence information. Since the knowledge for constructing LPPI network may be incomprehensive, we applied the centrality measures only on those lncRNAs with enough interactions. For those lncRNAs with limited number of interactions, we turned to rely on its sequence to score the essentiality.</p>
<p>The results support our assumption that essential lncRNAs have similar roles as essential coding genes in the LPPI network. Particularly, we found that BC appears to be more important than other kinds of centrality measures. Due to the limited number of known essential lncRNAs, it is not feasible to explore further optimization of different weight on different centralities. When more essential lncRNAs are reported and recorded, we believe that modern machine learning algorithms will provide deeper insights in identifying essential non-coding genes. As a summary, we listed the prediction results of SGII on mouse and human datasets in <xref ref-type="sec" rid="s10">Supplementary Table S3</xref> in supplementary materials, which may be useful for life science studies. A more comprehensive collection of essential lncRNAs is being curated. We plan to establish a database that is dedicated in recording essential lncRNA information in future.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s10">Supplementary Material</xref>, further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>X-HX collected the data, implemented the algorithm, perform the experiments, analyzed the results, and partially wrote the manuscript; Y-YZ helped in designing the algorithm and analyzed the results; C-QG analyzed the results and partially wrote the manuscript; HM partially analyzed the results; LW and P-FD directed the whole study, conceptualize the algorithm, supervised the experiments, analyzed the results, and wrote the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work was supported by National Natural Science Foundation of China (NSFC 61872268) and National Key R and D Program of China (2018YFC0910405).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.864564/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2022.864564/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Table2.XLSX" id="SM1" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table3.XLSX" id="SM2" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Table1.XLSX" id="SM3" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Austin</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Battey</surname>
<given-names>J.&#x20;F.</given-names>
</name>
<name>
<surname>Bradley</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bucan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Capecchi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Collins</surname>
<given-names>F. S.</given-names>
</name>
<etal/>
</person-group> (<year>2004</year>). <article-title>The Knockout Mouse Project</article-title>. <source>Nat. Genet.</source> <volume>36</volume>, <fpage>921</fpage>&#x2013;<lpage>924</lpage>. <pub-id pub-id-type="doi">10.1038/ng0904-921</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Campos</surname>
<given-names>T. L.</given-names>
</name>
<name>
<surname>Korhonen</surname>
<given-names>P. K.</given-names>
</name>
<name>
<surname>Gasser</surname>
<given-names>R. B.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>N. D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>An Evaluation of Machine Learning Approaches for the Prediction of Essential Genes in Eukaryotes Using Protein Sequence-Derived Features</article-title>. <source>Comput. Struct. Biotechnol. J.</source> <volume>17</volume>, <fpage>785</fpage>&#x2013;<lpage>796</lpage>. <pub-id pub-id-type="doi">10.1016/j.csbj.2019.05.008</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L.-L.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Linking Long Noncoding RNA Localization and Function</article-title>. <source>Trends Biochem. Sci.</source> <volume>41</volume>, <fpage>761</fpage>&#x2013;<lpage>772</lpage>. <pub-id pub-id-type="doi">10.1016/j.tibs.2016.07.003</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Da Sacco</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Baldassarre</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Masotti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Bioinformatics Tools and Novel Challenges in Long Non-coding RNAs (lncRNAs) Functional Analysis</article-title>. <source>Ijms</source> <volume>13</volume>, <fpage>97</fpage>&#x2013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.3390/ijms13010097</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fenoglio</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ridolfi</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Galimberti</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Scarpini</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>An Emerging Role for Long Non-coding RNA Dysregulation in Neurological Disorders</article-title>. <source>Ijms</source> <volume>14</volume>, <fpage>20427</fpage>&#x2013;<lpage>20442</lpage>. <pub-id pub-id-type="doi">10.3390/ijms141020427</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Georgi</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Voight</surname>
<given-names>B. F.</given-names>
</name>
<name>
<surname>Bu&#x107;an</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>From Mouse to Human: Evolutionary Genomics Analysis of Human Orthologs of Essential Genes</article-title>. <source>Plos Genet.</source> <volume>9</volume>, <fpage>e1003484</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1003484</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grote</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Wittler</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Hendrix</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Koch</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>W&#xe4;hrisch</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Beisaw</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>The Tissue-specific lncRNA Fendrr Is an Essential Regulator of Heart and Body wall Development in the Mouse</article-title>. <source>Dev. Cel.</source> <volume>24</volume>, <fpage>206</fpage>&#x2013;<lpage>214</lpage>. <pub-id pub-id-type="doi">10.1016/j.devcel.2012.12.012</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>NPInter v3.0: an Upgraded Database of Noncoding RNA-Associated Interactions</article-title>. <source>Database</source> <volume>2016</volume>, <fpage>baw057</fpage>. <pub-id pub-id-type="doi">10.1093/database/baw057</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Howe</surname>
<given-names>K. L.</given-names>
</name>
<name>
<surname>Achuthan</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Allen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Allen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Alvarez-Jarreta</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Amode</surname>
<given-names>M. R.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Ensembl 2021</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D884</fpage>&#x2013;<lpage>D891</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa942</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ai</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>LPI-ETSLP: lncRNA-Protein Interaction Prediction Using Eigenvalue Transformation-Based Semi-supervised Link Prediction</article-title>. <source>Mol. Biosyst.</source> <volume>13</volume>, <fpage>1781</fpage>&#x2013;<lpage>1787</lpage>. <pub-id pub-id-type="doi">10.1039/c7mb00290d</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jathar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Srivastava</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tripathi</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Technological Developments in lncRNA Biology</article-title>. <source>Adv. Exp. Med. Biol.</source> <volume>1008</volume>, <fpage>283</fpage>&#x2013;<lpage>323</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-10-5203-3_10</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jeong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mason</surname>
<given-names>S. P.</given-names>
</name>
<name>
<surname>Barab&#xe1;si</surname>
<given-names>A.-L.</given-names>
</name>
<name>
<surname>Oltvai</surname>
<given-names>Z. N.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Lethality and Centrality in Protein Networks</article-title>. <source>Nature</source> <volume>411</volume>, <fpage>41</fpage>&#x2013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1038/35075138</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Joy</surname>
<given-names>M. P.</given-names>
</name>
<name>
<surname>Brock</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ingber</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>High-betweenness Proteins in the Yeast Protein Interaction Network</article-title>. <source>J.&#x20;Biomed. Biotechnol.</source> <volume>2005</volume>, <fpage>96</fpage>&#x2013;<lpage>103</lpage>. <pub-id pub-id-type="doi">10.1155/JBB.2005.96</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Khalil</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Rinn</surname>
<given-names>J.&#x20;L.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>RNA-protein Interactions in Human Health and Disease</article-title>. <source>Semin. Cel Dev. Biol.</source> <volume>22</volume>, <fpage>359</fpage>&#x2013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1016/j.semcdb.2011.02.016</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Klattenhoff</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Scheuermann</surname>
<given-names>J.&#x20;C.</given-names>
</name>
<name>
<surname>Surface</surname>
<given-names>L. E.</given-names>
</name>
<name>
<surname>Bradley</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>Fields</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Steinhauser</surname>
<given-names>M. L.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Braveheart, a Long Noncoding RNA Required for Cardiovascular Lineage Commitment</article-title>. <source>Cell</source> <volume>152</volume>, <fpage>570</fpage>&#x2013;<lpage>583</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2013.01.003</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>J.&#x20;T.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Disruption of Imprinted X Inactivation by Parent-Of-Origin Effects at Tsix</article-title>. <source>Cell</source> <volume>103</volume>, <fpage>17</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1016/s0092-8674(00)00101-x</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H. Y.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Physiological Roles of Long Noncoding RNAs: Insight from Knockout Mice</article-title>. <source>Trends Cel Biol.</source> <volume>24</volume>, <fpage>594</fpage>&#x2013;<lpage>602</lpage>. <pub-id pub-id-type="doi">10.1016/j.tcb.2014.06.003</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.-x.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>A New Essential Protein Discovery Method Based on the Integration of Protein-Protein Interaction and Gene Expression Data</article-title>. <source>BMC Syst. Biol.</source> <volume>6</volume>, <fpage>15</fpage>. <pub-id pub-id-type="doi">10.1186/1752-0509-6-15</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Egranov</surname>
<given-names>S. D.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Molecular Mechanisms of Long Noncoding RNAs&#x2010;mediated Cancer Metastasis</article-title>. <source>Genes Chromosomes Cancer</source> <volume>58</volume>, <fpage>200</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1002/gcc.22691</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Gan</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Inferring Gene Regulatory Networks Using the Improved Markov Blanket Discovery Algorithm</article-title>. <source>Interdiscip. Sci. Comput. Life Sci.</source> <pub-id pub-id-type="doi">10.1007/s12539-021-00478-9</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lorenz</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bernhart</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>H&#xf6;ner zu Siederdissen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tafer</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Flamm</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Stadler</surname>
<given-names>P. F.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>ViennaRNA Package 2.0</article-title>. <source>Algorithms Mol. Biol.</source> <volume>6</volume>, <fpage>26</fpage>. <pub-id pub-id-type="doi">10.1186/1748-7188-6-26</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marahrens</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Panning</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Dausman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Strauss</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Jaenisch</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Xist-deficient Mice Are Defective in Dosage Compensation but Not Spermatogenesis</article-title>. <source>Genes Dev.</source> <volume>11</volume>, <fpage>156</fpage>&#x2013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1101/gad.11.2.156</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mercer</surname>
<given-names>T. R.</given-names>
</name>
<name>
<surname>Dinger</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Mattick</surname>
<given-names>J.&#x20;S.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Long Non-coding RNAs: Insights into Functions</article-title>. <source>Nat. Rev. Genet.</source> <volume>10</volume>, <fpage>155</fpage>&#x2013;<lpage>159</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2521</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oughtred</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rust</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Breitkreutz</surname>
<given-names>B. J.</given-names>
</name>
<name>
<surname>Stark</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Willems</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>TheBioGRIDdatabase: A Comprehensive Biomedical Resource of Curated Protein, Genetic, and Chemical Interactions</article-title>. <source>Protein Sci.</source> <volume>30</volume>, <fpage>187</fpage>&#x2013;<lpage>200</lpage>. <pub-id pub-id-type="doi">10.1002/pro.3978</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Penny</surname>
<given-names>G. D.</given-names>
</name>
<name>
<surname>Kay</surname>
<given-names>G. F.</given-names>
</name>
<name>
<surname>Sheardown</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Rastan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Brockdorff</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Requirement for Xist in X Chromosome Inactivation</article-title>. <source>Nature</source> <volume>379</volume>, <fpage>131</fpage>&#x2013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1038/379131a0</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pyfrom</surname>
<given-names>S. C.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Payton</surname>
<given-names>J.&#x20;E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>PLAIDOH: a Novel Method for Functional Prediction of Long Non-coding RNAs Identifies Cancer-specific LncRNA Activities</article-title>. <source>BMC Genomics</source> <volume>20</volume>, <fpage>137</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-019-5497-4</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rinn</surname>
<given-names>J.&#x20;L.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H. Y.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Genome Regulation by Long Noncoding RNAs</article-title>. <source>Annu. Rev. Biochem.</source> <volume>81</volume>, <fpage>145</fpage>&#x2013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-biochem-051410-092902</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sado</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Sasaki</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Regulation of Imprinted X-Chromosome Inactivation in Mice by Tsix</article-title>. <source>Development</source> <volume>128</volume>, <fpage>1275</fpage>&#x2013;<lpage>1286</lpage>. <pub-id pub-id-type="doi">10.1242/dev.128.8.1275</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sauvageau</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Goff</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Lodato</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bonev</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Groff</surname>
<given-names>A. F.</given-names>
</name>
<name>
<surname>Gerhardinger</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Multiple Knockout Mouse Models Reveal lincRNAs Are Required for Life and Brain Development</article-title>. <source>Elife</source> <volume>2</volume>, <fpage>e01749</fpage>. <pub-id pub-id-type="doi">10.7554/eLife.01749</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schmitt</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H. Y.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Long Noncoding RNAs in Cancer Pathways</article-title>. <source>Cancer Cell</source> <volume>29</volume>, <fpage>452</fpage>&#x2013;<lpage>463</lpage>. <pub-id pub-id-type="doi">10.1016/j.ccell.2016.03.010</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Uchida</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dimmeler</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Long Noncoding RNAs in Cardiovascular Diseases</article-title>. <source>Circ. Res.</source> <volume>116</volume>, <fpage>737</fpage>&#x2013;<lpage>750</lpage>. <pub-id pub-id-type="doi">10.1161/CIRCRESAHA.116.302521</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Min Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Huan Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yi Pan</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Identification of Essential Proteins Based on Edge Clustering Coefficient</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinf.</source> <volume>9</volume>, <fpage>1070</fpage>&#x2013;<lpage>1080</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2011.147</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>F.-X.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Computational Approaches to Predicting Essential Proteins: a Survey</article-title>. <source>Proteomices. Clin. Appl.</source> <volume>7</volume>, <fpage>181</fpage>&#x2013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1002/prca.201200068</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Watanabe</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sato</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Amano</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kawamura</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kawamura</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kawaguchi</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2008</year>). <article-title>Dnm3os, a Non-coding RNA, Is Required for normal Growth and Skeletal Development in Mice</article-title>. <source>Dev. Dyn.</source> <volume>237</volume>, <fpage>3738</fpage>&#x2013;<lpage>3748</lpage>. <pub-id pub-id-type="doi">10.1002/dvdy.21787</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wuchty</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Stadler</surname>
<given-names>P. F.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Centers of Complex Networks</article-title>. <source>J.&#x20;Theor. Biol.</source> <volume>223</volume>, <fpage>45</fpage>&#x2013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-5193(03)00071-7</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yildirim</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Kirby</surname>
<given-names>J.&#x20;E.</given-names>
</name>
<name>
<surname>Brown</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Mercier</surname>
<given-names>F. E.</given-names>
</name>
<name>
<surname>Sadreyev</surname>
<given-names>R. I.</given-names>
</name>
<name>
<surname>Scadden</surname>
<given-names>D. T.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Xist RNA Is a Potent Suppressor of Hematologic Cancer in Mice</article-title>. <source>Cell</source> <volume>152</volume>, <fpage>727</fpage>&#x2013;<lpage>742</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2013.01.034</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Defining Essentiality Score of Protein-Coding Genes and Long Noncoding RNAs</article-title>. <source>Front. Genet.</source> <volume>9</volume>, <fpage>380</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2018.00380</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ou</surname>
<given-names>H.-Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.-T.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>DEG: a Database of Essential Genes</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume>, <fpage>271D</fpage>&#x2013;<lpage>272D</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh024</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yue</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>SFPEL-LPI: Sequence-Based Feature Projection Ensemble Learning for Predicting LncRNA-Protein Interactions</article-title>. <source>Plos Comput. Biol.</source> <volume>14</volume>, <fpage>e1006616</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1006616</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Novel Method to Predict Essential Proteins Based on Tensor and HITS Algorithm</article-title>. <source>Hum. Genomics</source> <volume>14</volume>, <fpage>14</fpage>. <pub-id pub-id-type="doi">10.1186/s40246-020-00263-7</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Using Network Distance Analysis to Predict lncRNA-miRNA Interactions</article-title>. <source>Interdiscip. Sci. Comput. Life Sci.</source> <volume>13</volume>, <fpage>535</fpage>&#x2013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1007/s12539-021-00458-z</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Hao</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>NONCODE 2016: an Informative and Valuable Data Source of Long Non-coding RNAs</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>D203</fpage>&#x2013;<lpage>D208</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv1252</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Teng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Yap</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Challenges and Strategies in Ascribing Functions to Long Noncoding RNAs</article-title>. <source>Cancers</source> <volume>12</volume>, <fpage>1458</fpage>. <pub-id pub-id-type="doi">10.3390/cancers12061458</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Prediction of Essential Proteins Based on Gene Expression Programming</article-title>. <source>BMC Genomics</source> <volume>14 Suppl 4</volume>, <fpage>S7</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-14-S4-S7</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Klibanski</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>MEG3 Noncoding RNA: a Tumor Suppressor</article-title>. <source>J.&#x20;Mol. Endocrinol.</source> <volume>48</volume>, <fpage>R45</fpage>&#x2013;<lpage>R53</lpage>. <pub-id pub-id-type="doi">10.1530/JME-12-0008</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Function of lncRNAs and Approaches to lncRNA-Protein Interactions</article-title>. <source>Sci. China Life Sci.</source> <volume>56</volume>, <fpage>876</fpage>&#x2013;<lpage>885</lpage>. <pub-id pub-id-type="doi">10.1007/s11427-013-4553-6</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>