<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2022.1091698</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Analysis on the interactions between the first introns and other introns in mitochondrial ribosomal protein genes</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Ruifang</given-names>
</name>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1155379/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Song</surname>
<given-names>Xinwei</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2116145/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Gao</surname>
<given-names>Shan</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2115785/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Peng</surname>
<given-names>Shiya</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2115806/overview"/>
</contrib>
</contrib-group>
<aff><institution>College of Physics and Electronic Information, Inner Mongolia Normal University</institution>, <addr-line>Hohhot</addr-line>, <country>China</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Yongchun Zuo, Inner Mongolia University, China</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Hao Wu, Shandong University, China; Shiyuan Wang, Harbin Medical University, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Ruifang Li, <email>liruifang@imnu.edu.cn</email></corresp>
<fn id="fn0003" fn-type="other"><p>This article was submitted to Evolutionary and Genomic Microbiology, a section of the journal Frontiers in Microbiology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>08</day>
<month>12</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>1091698</elocation-id>
<history>
<date date-type="received">
<day>07</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>11</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Li, Song, Gao and Peng.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Li, Song, Gao and Peng</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>It is realized that the first intron plays a key role in regulating gene expression, and the interactions between the first introns and other introns must be related to the regulation of gene expression. In this paper, the sequences of mitochondrial ribosomal protein genes were selected as the samples, based on the Smith-Waterman method, the optimal matched segments between the first intron and the reverse complementary sequences of other introns of each gene were obtained, and the characteristics of the optimal matched segments were analyzed. The results showed that the lengths and the ranges of length distributions of the optimal matched segments are increased along with the evolution of eukaryotes. For the distributions of the optimal matched segments with different GC contents, the peak values are decreased along with the evolution of eukaryotes, but the corresponding GC content of the peak values are increased along with the evolution of eukaryotes, it means most introns of higher organisms interact with each other though weak bonds binding. By comparing the lengths and matching rates of optimal matched segments with those of siRNA and miRNA, it is found that some optimal matched segments may be related to non-coding RNA with special biological functions, just like siRNA and miRNA, they may play an important role in the process of gene expression and regulation. For the relative position of the optimal matched segments, the peaks of relative position distributions of optimal matched segments are increased during the evolution of eukaryotes, and the positions of the first two peaks exhibit significant conservatism.</p>
</abstract>
<kwd-group>
<kwd>local matched alignment</kwd>
<kwd>first intron</kwd>
<kwd>optimal matched segments</kwd>
<kwd>mitochondrial ribosomal protein genes</kwd>
<kwd>interaction</kwd>
</kwd-group>
<contract-num rid="cn1">2019MS03042</contract-num>
<contract-num rid="cn2">31860304</contract-num>
<contract-sponsor id="cn1">Natural Science Foundation of Inner Mongolia<named-content content-type="fundref-id">10.13039/501100004763</named-content>
</contract-sponsor>
<contract-sponsor id="cn2">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
<counts>
<fig-count count="2"/>
<table-count count="2"/>
<equation-count count="7"/>
<ref-count count="22"/>
<page-count count="8"/>
<word-count count="5366"/>
</counts>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<title>Introduction</title>
<p>An intron sequence is regarded as a kind of non-coding sequence of interrupted gene, and its functions are being discovered. A large number of studies have shown that intron can regulate gene expression as a kind of regulatory element (<xref ref-type="bibr" rid="ref12">Palmiter et al., 1991</xref>; <xref ref-type="bibr" rid="ref7">Li et al., 2015</xref>; <xref ref-type="bibr" rid="ref1">Abou et al., 2020</xref>), for example, the heterologous introns can enhance expression of transgenes in mice (<xref ref-type="bibr" rid="ref12">Palmiter et al., 1991</xref>). In recent years, it has also been found that some introns can influence many stages of mRNA metabolism, including initial transcription of a gene, editing of pre-mRNA, and nuclear export, translation and decay of the mRNA (<xref ref-type="bibr" rid="ref2">Bo et al., 2019</xref>). Furthermore, introns contain kinds of non-coding RNA such as microRNA and snoRNA (<xref ref-type="bibr" rid="ref10">Mattick and Gagen, 2001</xref>), they also participate in the functional activities of a variety of non-coding RNA (<xref ref-type="bibr" rid="ref1">Abou et al., 2020</xref>). And it has shown that GC-AG introns are mainly associated with lncRNAs and are preferentially located in the first intron (<xref ref-type="bibr" rid="ref1">Abou et al., 2020</xref>), additionally, many studies have shown that introns are closely related to various diseases (<xref ref-type="bibr" rid="ref14">Sowalsky et al., 2015</xref>; <xref ref-type="bibr" rid="ref9">Malekkou et al., 2020</xref>; <xref ref-type="bibr" rid="ref11">Ong and Adusumalli, 2020</xref>), for example, a novel mutation deep within intron 7 of the GBA gene can cause Gaucher disease (<xref ref-type="bibr" rid="ref9">Malekkou et al., 2020</xref>), and increased intron retention is associated with Alzheimer&#x2019;s disease (<xref ref-type="bibr" rid="ref11">Ong and Adusumalli, 2020</xref>), and metastatic castration-resistant prostate cancer is related to some non-coding RNA (<xref ref-type="bibr" rid="ref14">Sowalsky et al., 2015</xref>).</p>
<p>The most basic and important interaction among bases is base matching, for example, the formation of a correct codon-anticodon pair is essential to ensure efficiency and fidelity during translation, and circRNA formed by exon cyclization or intron cyclization contains long flanking introns with complementary repeats (<xref ref-type="bibr" rid="ref5">Han et al., 2022</xref>). Besides, many studies indicated that intron complementary matching fragments are not only the cause of circular RNA, but also the potential factors for the complexity and diversity of gene expression at the transcriptional/post transcriptional level (<xref ref-type="bibr" rid="ref22">Zhang et al., 2013</xref>, <xref ref-type="bibr" rid="ref21">2014</xref>; <xref ref-type="bibr" rid="ref6">Jiao et al., 2021</xref>). Therefore, it is particularly important to study the circular matching problem of introns.</p>
<p>The first introns have gained increasing attentions in recent years because of their unique features that are located in close proximity to the transcription, and the distinct deposition of epigenetic marks and nucleosome density on the first intronic DNA sequence (<xref ref-type="bibr" rid="ref4">Fu et al., 2022</xref>; <xref ref-type="bibr" rid="ref13">Singh et al., 2022</xref>; <xref ref-type="bibr" rid="ref15">Spijker et al., 2022</xref>; <xref ref-type="bibr" rid="ref17">Vosseberg et al., 2022</xref>), and it is realized that the first introns play a key role in several mechanisms regulating gene expression. We determined that the matching features between the first introns and the corresponding reverse complementary sequences of other introns must provide many useful information.</p>
<p>The genome consists of an extremely complex network of interactions among functional elements, and its functions are achieved primarily through these interactions. We have known that a complete match between siRNA and targeted genes can lead to targeted genes silencing, and high but incomplete matching between miRNA and targeted genes can suppress gene expression. It means base matching is an important way for non-coding RNA to interact with targeted genes, intron as a kind of non-coding DNA is rich in eukaryote genomes, introns must interact with each other, and the interactions can be embodied by the modes of base matching. Based on this, in this work, the mitochondrial ribosomal protein gene sequences were selected as samples, the characteristics of the optimal matched segments between the first intron and the corresponding reverse complementary sequences of other introns were analyzed, and the variations of the characteristics along with the evolution of eukaryotes were studied.</p>
</sec>
<sec id="sec2" sec-type="materials|methods">
<title>Materials and methods</title>
<sec id="sec3">
<title>Datasets</title>
<p>All the sequences of mitochondrial ribosomal protein genes in the Ribosomal Protein Gene Database (RPG) were selected as our samples, they were from <italic>Homo sapiens, Mus musculus, Fugu rubripes, Drosophila melanogaster</italic> and <italic>Caenorhabditis elegans</italic>. Considering that the mitochondrial ribosomal protein gene has many advantages in biological research as a kind of housekeeping gene, they are involved in the key process of all protein translation and have very good evolutionary conservatism (<xref ref-type="bibr" rid="ref18">Yoshihama et al., 2002</xref>), thus forming a family of conservative genes. They exist widely in all eukaryotes, and their intron lengths and amounts have little difference in all eukaryotes. We believe that more reliable and functional interactions among introns can be obtained by selecting these conserved genes. Information about the protein genes is given in <xref rid="tab1" ref-type="table">Table 1</xref>.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Mitochondrial ribosomal protein genes.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Species</th>
<th align="center" valign="top">The amount of genes</th>
<th align="center" valign="top">The amount of introns</th>
<th align="center" valign="top">The amount of first introns</th>
</tr>
</thead>
<tbody>
<tr>
<td align="char" valign="top" char="."><italic>Homo sapiens</italic></td>
<td align="center" valign="top">114</td>
<td align="center" valign="top">512</td>
<td align="center" valign="top">114</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Mus musculus</italic></td>
<td align="center" valign="top">79</td>
<td align="center" valign="top">351</td>
<td align="center" valign="top">79</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Fugu rubripes</italic></td>
<td align="center" valign="top">69</td>
<td align="center" valign="top">266</td>
<td align="center" valign="top">64</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Drosophila melanogaster</italic></td>
<td align="center" valign="top">75</td>
<td align="center" valign="top">118</td>
<td align="center" valign="top">66</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Caenorhabditis elegans</italic></td>
<td align="center" valign="top">74</td>
<td align="center" valign="top">251</td>
<td align="center" valign="top">71</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="sec4">
<title>Matching method</title>
<p>The intron sequences were obtained from the above gene sequences, then, they were transformed into their reverse complementary sequences except the first introns. Next, similar alignments were done by the local similarity matching software called Smith-Waterman.<xref rid="fn0004" ref-type="fn"><sup>1</sup></xref> We adopt Ednafull matrix to similarity matching, and parameters chosen as follows, each Gap penalty is 50.0, in the gap each extend penalty is 5.0, thus, we got the optimal similar segment between the first intron and corresponding reverse complementary sequences of other introns in each gene sequence.</p>
</sec>
<sec id="sec5">
<title>The optimal matching frequency</title>
<p>We calculated the length, GC content, matching rate of each optimal matched segment considering that they must provide the basic characteristics of the optimal matched segments, then divided the optimal matching segments into several groups, respectively, according to their lengths, GC contents or matching rates. And then calculated the frequencies of the optimal matched segments with different ranges of lengths, GC contents and matching rates, marked with <italic>F</italic><sub><italic>L</italic>m</sub>, <italic>F</italic><sub>GCm</sub>, and <italic>F</italic><sub>mat</sub> respectively, they were defined as follows,</p>
<disp-formula id="EQ1">
<label>(1)</label>
<mml:math id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mi>N</mml:mi>
<mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi mathvariant="normal">m</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="EQ2">
<label>(2)</label>
<mml:math id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">GCm</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">GCm</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">GC</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">GCm</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="EQ3">
<label>(3)</label>
<mml:math id="M3">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi mathvariant="normal">mat</mml:mi>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">mat</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">mat</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">mat</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where, <italic>F</italic><sub><italic>L</italic>m<italic>i</italic></sub> is the frequency of the optimal matched segments whose length are within the <italic>i</italic>th group, <italic>N</italic><sub><italic>L</italic>m<italic>i</italic></sub> is the amount of the optimal matched segments in the <italic>i</italic>th group, and <italic>n</italic><sub><italic>L</italic></sub> is the amount of the groups divided according to their lengths. <italic>F</italic><sub>GCm<italic>j</italic></sub> is the frequency of the optimal matched segments whose GC contents are within the <italic>j</italic>th group, <italic>N</italic><sub>GCm<italic>j</italic></sub> is the amount of the optimal matched segments in the <italic>j</italic>th group, and <italic>n</italic><sub>GC</sub> is the amount of the groups divided according to their GC contents. <italic>F</italic><sub>mat<italic>k</italic></sub> is the frequency of the optimal matched segments whose matching rate are within the <italic>k</italic>th group, <italic>N</italic><sub>mat<italic>k</italic></sub> is the amount of the optimal matched segments in the <italic>k</italic>th group, and <italic>n</italic><sub>mat</sub> is the amount of the groups divided according to their matching rates.</p>
<p>The lengths of the first introns in different gene sequences are different, we standardized the first introns as sequences with 100&#x2009;bp length in order to conveniently compare the relative position distributions of the optimal matched segments. The method of length standardization as follows (<xref ref-type="bibr" rid="ref19">Zhang et al., 2016a</xref>),</p>
<disp-formula id="EQ4">
<label>(4)</label>
<mml:math id="M4">
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfenced close="" open="{">
<mml:mrow>
<mml:mtable equalrows="true" equalcolumns="true">
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mfenced close="]" open="[">
<mml:mrow>
<mml:mn>100</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mn>100</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi mathvariant="normal">
</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi mathvariant="normal"> </mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mfenced close="]" open="[">
<mml:mrow>
<mml:mn>100</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mi>L</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mn>100</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi mathvariant="normal">&#x2004;</mml:mi>
<mml:mspace width="0.25em"/>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi mathvariant="normal"> </mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x2010;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where, <italic>L<sub>i</sub></italic> is the length of the <italic>i</italic>th first intron, <italic>N<sub>ij</sub></italic> is the <italic>j</italic>th base site of the <italic>i</italic>th first intron, and <italic>n<sub>ij</sub></italic> is its relative position corresponding to the <italic>i</italic>th standardized first intron. In this way, the first introns are all transformed into 100&#x2009;bp long sequences.</p>
<p>According to the base site of each optimal matching sequence on the first intron, each base site of the first intron is scored, if in the optimal matching region, base site is scored 1, but if not, it is scored 0, and the definition of matching score as follows (<xref ref-type="bibr" rid="ref19">Zhang et al., 2016a</xref>),</p>
<disp-formula id="EQ5">
<label>(5)</label>
<mml:math id="M5">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mo>&#x2329;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.25em"/>
<mml:mspace width="0.25em"/>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mspace width="0.25em"/>
<mml:mspace width="0.25em"/>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo>&#x232A;</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>b</mml:mi>
<mml:mspace width="0.25em"/>
<mml:mspace width="0.25em"/>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where, <italic>f<sub>ij</sub></italic> is the score of the <italic>j</italic>th base site on the standardized <italic>i</italic>th first intron, <italic>n<sub>ia</sub></italic> and <italic>n<sub>ib</sub></italic> are the initiation base relative site and the termination base relative site of the optimal matched segments on the standardized <italic>i</italic>th first intron. Thus, for each optimal matching sequence, the first intron is transformed into a sequence consisted of 0 and 1. if there are <italic>m</italic> optimal matching sequences in a gene, we can obtain <italic>m</italic> sequences consisted of 0 and 1. On this basis, we divided the 100 sites of each number sequences into 10 regions on average, the relative position frequency of the optimal matched segments on each site and in each region are defined as follows,</p>
<disp-formula id="EQ6">
<label>(6)</label>
<mml:math id="M6">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="EQ7">
<label>(7)</label>
<mml:math id="M7">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where, <italic>F</italic><sub>r<italic>j</italic></sub> is the relative position frequency of the optimal matched segments on the <italic>j</italic>th base site of the standardized first intron, <italic>F</italic><sub>r<italic>k</italic></sub> is the relative position frequency of the optimal matched segments in the <italic>k</italic>th region, <italic>f<sub>ij</sub></italic> is the score of <italic>j</italic>th base site on the standardized first intron, <italic>p<sub>ka</sub></italic> and <italic>p<sub>kb</sub></italic> are the initiation base site and the termination base site of the <italic>k</italic>th region, <italic>N<sub>ia</sub></italic> and <italic>N<sub>ib</sub></italic> are the initiation base site and the termination base site of the optimal matched segments, and <italic>m</italic> is the total number of the optimal matched segments in the gene.</p>
</sec>
</sec>
<sec id="sec6" sec-type="results">
<title>Results</title>
<p>The optimal matched segments between the first introns and the reverse complementary sequences of other introns in each mitochondrial ribonucleo protein gene of five species were counted, and the dataset of the optimal matched segments was established. Then, the lengths of the optimal matched segments of each species were counted, and the frequencies of the optimal matched segments with different ranges of lengths were calculated by <xref ref-type="disp-formula" rid="EQ1">formula (1)</xref>. The GC contents of the optimal matched segments of each species were counted, and the frequencies of the optimal matched segments with different ranges of GC contents were calculated by <xref ref-type="disp-formula" rid="EQ2">formula (2)</xref>. The matching rates of the optimal matched segments of each species were calculated, and the frequencies of the optimal matched segments with different ranges of matching rates were calculated by <xref ref-type="disp-formula" rid="EQ3">formula (3)</xref>. The first introns were standardized to sequences with 100&#x2009;bp length, and their relative base positions were calculated according to <xref ref-type="disp-formula" rid="EQ4">formula (4)</xref>, and then, according to the base sites of each optimal matching sequence on the first intron, each base site of the first intron is scored according to <xref ref-type="disp-formula" rid="EQ5">formula (5)</xref>, based on this, the relative position frequency were calculated according to <xref ref-type="disp-formula" rid="EQ6">formula (6)</xref> and <xref ref-type="disp-formula" rid="EQ7">(7)</xref>. On this basis, the characteristics of the optimal matched segments of five species were analyzed. The results are presented in <xref rid="fig1" ref-type="fig">Figure 1</xref>.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>The distributions of the optimal matching frequencies. <bold>(A)</bold> The length (<italic>L<sub>m</sub></italic>) distributions of the optimal matched segments. The values on the <italic>x</italic>-axis are the length ranges of groups divided according to lengths of the optimal matched segments, and those on the <italic>y</italic>-axis are the length frequency of the optimal matched segments. <bold>(B)</bold> The GC content (<italic>C</italic><sub>GCm</sub>) distributions of the optimal matched segments. The values on the <italic>x</italic>-axis are GC content ranges of groups divided according to GC contents of the optimal matched segments, and those on the <italic>y</italic>-axis are the GC content frequencies of the optimal matched segments. <bold>(C)</bold> The matching rate (<italic>f</italic><sub>mat</sub>) distributions of the optimal matched segment of five species. The values on the <italic>x</italic>-axis are matching rate ranges of groups divided according to matching rates of the optimal matched segments, and those on the <italic>y</italic>-axis are the matching rate frequencies of the optimal matched segments. <bold>(D&#x2013;H)</bold> The relative position distributions of the optimal matched segments on the first intron sequences. The values on the <italic>x</italic>-axis are relative position ranges of groups divided according to relative positions of the optimal matched segments, and those on the <italic>y</italic>-axis are the relative position frequencies of the optimal matched segments. <bold>(D)</bold> <italic>Homo sapiens</italic>, <bold>(E)</bold> <italic>Mus musculus</italic>, <bold>(F)</bold> <italic>Fugu rubripes</italic>, <bold>(G)</bold> <italic>Drosophila melanogaster</italic>, <bold>(H)</bold> <italic>Caenorhabditis elegans</italic>. <bold>(I)</bold> The length (<italic>L<sub>m</sub></italic>) distributions of the optimal matched segments of 5 species. The calculations of the optimal matched segment lengths of 5 species are combined into one, the optimal matching segments were divided into several groups according to their lengths, the frequency of the optimal matched segments in each group was calculated, the values on the <italic>x</italic>-axis are the length ranges of groups, and those on the <italic>y</italic>-axis are the length frequency of the optimal matched segments.</p>
</caption>
<graphic xlink:href="fmicb-13-1091698-g001.tif"/>
</fig>
<sec id="sec7">
<title>The length distributions of the optimal matched segments</title>
<p>The lengths of the optimal matched segments of five species are mainly concentrated at 10&#x2013;50&#x2009;bp, while some optimal matched segments of <italic>Homo sapiens, Mus musculus</italic> and <italic>Fugu rubripes</italic> are up to 100&#x2009;bp in length, and the ratio of the optimal matched segments of <italic>Homo sapiens</italic> concentrating at 90&#x2013;100&#x2009;bp is up to 24.6 percent. In addition, the length distribution of the optimal matched segments of <italic>Homo sapiens</italic> is similar to that of <italic>Mus musculus</italic>. The results showed that the optimal matched segments of high eukaryotes have a longer length and a wider length distribution than that of the low eukaryotes, and it means the length and the ranges of length distribution of the optimal matched segments are increased along with the evolution of eukaryotes.</p>
</sec>
<sec id="sec8">
<title>The GC content distributions of the optimal matched segments</title>
<p>The distributions of GC content of the optimal matched segments of five species ranged from 0 to 0.9. And comparing the results of the five species, it is found that the peak values of <italic>F</italic><sub>GCm</sub> are decreased along with the evolution of eukaryotes, but the corresponding GC content of the peak values are increased along with the evolution of eukaryotes.</p>
</sec>
<sec id="sec9">
<title>The matching rate distributions of the optimal matched segments</title>
<p>As seen from the matching rate distributions of the optimal matched segments, most of the matching rates are distributed between 60% and 80%. Interestingly, studies showed that the matching rates between miRNA and target mRNA are distributed between 65% and 95% (<xref ref-type="bibr" rid="ref3">Cui et al., 2010</xref>). Comparing the matching rate ranges of the optimal matched segments with that of siRNA or miRNA with target mRNA (<xref ref-type="bibr" rid="ref16">Volpe et al., 2002</xref>; <xref ref-type="bibr" rid="ref8">Lim et al., 2005</xref>; <xref ref-type="bibr" rid="ref3">Cui et al., 2010</xref>; <xref ref-type="bibr" rid="ref20">Zhang et al., 2016b</xref>), it is found that there is a high similarity between the optimal matched segments with the most probable matching rates and siRNA or miRNA, this also suggests these optimal matched segments may be related to some non-coding RNAs with special biological functions, just like siRNA and miRNA, they may play an important role in the process of gene expression and regulation.</p>
</sec>
<sec id="sec10">
<title>The relative position distributions of the optimal matched segments in the first introns</title>
<p>It can be seen from <xref rid="fig1" ref-type="fig">Figure 1</xref> that the relative positions of the optimal matched segments vary with the different species. And the peaks with <italic>F</italic>r values bigger than 10% are analyzed, the results showed that for <italic>Homo sapiens</italic>, there are 3 peaks with Fr values bigger than 10%, which are at 20&#x2013;30&#x2009;bp, 50&#x2013;60&#x2009;bp and 90&#x2013;100&#x2009;bp, and there is two peaks with Fr values bigger than 10%, which are at 20&#x2013;30&#x2009;bp and 50&#x2013;60&#x2009;bp for <italic>Mus musculus</italic>, 30&#x2013;40&#x2009;bp and 60&#x2013;70&#x2009;bp for <italic>Fugu rubripes</italic> and <italic>Drosophila melanogaster</italic>, but for <italic>Caenorhabditis elegans</italic>, there is only one peak with Fr value bigger than 10%, which is at 70&#x2013;80&#x2009;bp. This also indicates that the relative position frequencies of the optimal matched segments are distinctively differences among the five different species, and the peaks of relative position distributions of optimal matched segments are increased along with the evolution of eukaryotes, but the positions of the first two peaks exhibit significant conservatism.</p>
<p>In order to further confirm the conservatism of the relative position distributions of optimal matched segments, <xref rid="fig2" ref-type="fig">Figure 2</xref> was made according to the calculations by <xref ref-type="disp-formula" rid="EQ6">formula (6)</xref>, which express the relative position frequencies with the base sites of the standardized first intron.</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p><bold>(A&#x2013;E)</bold> The distribution of the relative position frequencies with the base sites of the standardized first intron.</p>
</caption>
<graphic xlink:href="fmicb-13-1091698-g002.tif"/>
</fig>
<p>As we can see from <xref rid="fig2" ref-type="fig">Figure 2</xref>, the distributions do not accord with normal distribution, so, for the relative position frequency of the optimal matched segments on each site of the standardized first intron, the test for differences between any two species were performed by non-parametric test with level of significance 0.05 using R software, the results are presented in <xref rid="tab2" ref-type="table">Table 2</xref>.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Results of the test for differences of relative position frequency of the optimal matched segments on each site of the standardized first intron.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="char" valign="top" char="&#x00D7;">Species</th>
<th align="char" valign="top" char="&#x00D7;">Value of <italic>p</italic></th>
<th align="char" valign="top" char="&#x00D7;">Species</th>
<th align="char" valign="top" char="&#x00D7;">Value of <italic>p</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td align="char" valign="top" char="."><italic>Homo sapiens-Mus musculus</italic></td>
<td align="char" valign="top" char=".">0.54</td>
<td align="char" valign="top" char="&#x00B1;"><italic>Mus musculus-Drosophila melanogaster</italic></td>
<td align="char" valign="top" char=".">0.30</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Homo sapiens-Fugu rubripes</italic></td>
<td align="char" valign="top" char=".">0.40</td>
<td align="char" valign="top" char="&#x00B1;"><italic>Mus musculus-Caenorhabditis elegans</italic></td>
<td align="char" valign="top" char=".">0.33</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Homo sapiens-Drosophila melanogaster</italic></td>
<td align="char" valign="top" char=".">0.66</td>
<td align="char" valign="top" char="&#x00B1;"><italic>Fugu rubripes-Drosophila melanogaster</italic></td>
<td align="char" valign="top" char=".">0.91</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Homo sapiens-Caenorhabditis elegans</italic></td>
<td align="char" valign="top" char=".">0.52</td>
<td align="char" valign="top" char="&#x00B1;"><italic>Fugu rubripes-Caenorhabditis elegans</italic></td>
<td align="char" valign="top" char=".">0.68</td>
</tr>
<tr>
<td align="char" valign="top" char="."><italic>Mus musculus-Fugu rubripes</italic></td>
<td align="char" valign="top" char=".">0.10</td>
<td align="char" valign="top" char="&#x00B1;"><italic>Drosophila melanogaster-Caenorhabditis elegans</italic></td>
<td align="char" valign="top" char=".">0.64</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be seen from <xref rid="tab2" ref-type="table">Table 2</xref> that the <italic>p</italic>-values between any two species are &#x003E;0.05, it indicates that the differences between any two species were not statistically significant, means, in terms of the distributions on each site of the standardized first intron, the optimal matched segments exhibited high conservatism during species evolution. These results further confirmed the conservatism of the relative position distributions of optimal matched segments, which let us be sure that optimal matched segments are organized and functional sequences for each species.</p>
</sec>
</sec>
<sec id="sec11" sec-type="conclusions">
<title>Conclusion</title>
<p>We analyzed the matching features between the first intron and the reverse complementary sequences of other introns in each gene sequence, for the lengths of the optimal matched segments, we found the most probable lengths are distributed between 20 and 30 bp, it can be seen from <xref rid="fig1" ref-type="fig">Figure 1</xref>, and we calculated the ratios of the optimal matched segments whose lengths are between 21 and 30&#x2009;bp, they are 25%, 27%, 31%, 15%, and 26% in <italic>Homo sapiens</italic>, <italic>Mus musculus</italic>, <italic>Fugu rubripes</italic>, <italic>Drosophila melanogaster</italic> and <italic>Caenorhabditis elegans, respectively.</italic> Interestingly, we know that the siRNA, whose length is from 21 to 25&#x2009;bp, guiding mRNA to silent by perfect complementarity with target mRNA (<xref ref-type="bibr" rid="ref16">Volpe et al., 2002</xref>; <xref ref-type="bibr" rid="ref8">Lim et al., 2005</xref>), and the miRNA, whose length is from 18 to 25&#x2009;bp, restrains transcription and expression of target mRNA by different degree complementarity with target mRNA (<xref ref-type="bibr" rid="ref19">Zhang et al., 2016a</xref>), the results indicate that the probable lengths of the optimal matched segments are remarkably similar to the lengths of siRNA and miRNA. For the matching rates of the optimal matched segments, we found that most of the matching rates are distributed between 60% and 80%, they are very remarkably similar to the matching rate ranges with target mRNA of siRNA or miRNA. It means there is a high similarity between some optimal matched segments and siRNA or miRNA. Is this a coincidence? we do not think so. The basic interaction between introns is base complementary pairing, the matched sequences, especially between the first intron and corresponding complementary sequences of other introns in the same gene, must be related to some elements with special functions. Taking all the analyzes and conclusions above into account, we come to a conclusion that some optimal matched segments may be a kind of non-coding RNA with special biological functions, just like siRNA and miRNA, they are likely to participate in the process of gene expression and regulation. And we think, the optimal matched segments with special characteristics in the first introns may take part in regulating gene expression by RNA matching competition with other introns or exon.</p>
<p>In addition, we have got some interesting results by comparing the results of different species. In terms of the species selected in this work, <italic>Caenorhabditis elegans</italic>, <italic>Drosophila melanogaster</italic>, <italic>Fugu rubripes</italic>, <italic>Mus musculus</italic>, and <italic>Homo sapiens</italic> are listed from lower eukaryotes to higher eukaryotes. Based on this order and the calculations, we tried to analyze the variation law of matching features of introns along with the species evolution. For the lengths of the optimal matched segments, the average length of the optimal matched segments for the high eukaryotes are longer than that of the low eukaryotes. It suggests that the lengths of the optimal matched segments are increased in the evolution of eukaryotes. And the results showed that with the evolution of eukaryotes, the distributions of the length of the optimal matched segment become wider and wider. It means the lengths and the ranges of length distributions of the optimal matched segments are increased along with the evolution of eukaryotes. For the GC content of the optimal matched segments, the peak values of <italic>F</italic><sub>GCm</sub> are decreased along with the evolution of eukaryotes, it suggests the GC contents of the optimal matched segments are more widely distributed with the evolution of eukaryotes. If some functional elements are related to the optimal matched segments with special GC contents, the result means that higher organisms have more kinds of functional elements than lower organisms. But the corresponding GC content at the peak values are increased along with the evolution of eukaryotes. Both AT and GC basepairs form one set of hydrogen bonds, and it is a truth universally acknowledged that a GC base pair has three hydrogen bonds whereas AT has two, it means that DNA with high GC-content is more stable than DNA with low GC-content. Based on the above theories, it can be concluded that introns of higher organisms interacting with each other though weak bonds binding are more than that of lower organisms, we hypothesized that interactions through weak bonds can ensure the flexibility to take part in gene regulation. For the relative position of the optimal matched segments, the peaks of relative position distributions of optimal matched segments are increased with the evolution of eukaryotes, and the positions of the first two peaks exhibit significant conservatism. We think that some functional elements are related to the optimal matched segments at these proper positions, the results indicated that these elements of higher eukaryotes may have a more specific division of labor.</p>
<p>To conclude, in this work, we analyzed the possibility of interactions between the first introns and the other introns of mitochondrial ribosomal protein genes, then tried to interpret the modes of interactions between introns. We found some universal characteristics of the optimal matched segments between the first introns and the reverse complementary sequences of other introns, and we noticed there is a high similarity between some optimal matched segments and siRNA or miRNA, so, we believe that the characteristics of interactions among introns obtained in this work are the basic characteristics of the RNA&#x2013;RNA interactions. It means the optimal matched segments are probably functional non-coding RNAs including siRNA and miRNA. At the same time, we found some variation law of the optimal matched segments with the evolution of eukaryotes, which indicates that there may be a great difference in the complexity of the interactions of introns among species at different evolutionary levels, these results are of great significance in explaining the function of non-coding RNA. An increasing number of people are realizing the importance of non-coding RNA in gene expression regulating, but the functions and regulating mechanisms are not very clear, our results indicate that the base matching plays a key role in the interactions among introns. However, the related works have just started, the sample size in this study is relatively small, further large-sample studies are needed to obtain more detailed and clearer mechanism of introns interactions.</p>
</sec>
<sec id="sec12" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="sec13">
<title>Author contributions</title>
<p>RFL, XWS, SG, and SYP performed the work. RFL has developed the theoretical frame work, finished the data analysis, and wrote the manuscript. XWS, SG, and SYP finished the data collection and the calculations. All authors have read and approved this version of the article.</p>
</sec>
<sec id="sec14" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by the Natural Science Foundation of Inner Mongolia (2019MS03042) and the National Natural Science Foundation of China (31860304).</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abou</surname> <given-names>A. M.</given-names></name> <name><surname>Celli</surname> <given-names>L.</given-names></name> <name><surname>Belotti</surname> <given-names>G.</given-names></name> <name><surname>Lisa</surname> <given-names>A.</given-names></name> <name><surname>Bione</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>GC-AG introns features in long non-coding and protein-coding genes suggest their role in gene expression regulation</article-title>. <source>Front. Genet.</source> <volume>11</volume>:<fpage>488</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fgene.2020.00488</pub-id></citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bo</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Lu</surname> <given-names>Z.</given-names></name> <name><surname>Bao</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Potential relations between post-spliced introns and mature mRNAs in the Caenorhabditis elegans genome</article-title>. <source>J. Theor. Biol.</source> <volume>467</volume>, <fpage>7</fpage>&#x2013;<lpage>14</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jtbi.2019.01.031</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cui</surname> <given-names>J. G.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Sethi</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Y. Y.</given-names></name> <name><surname>Mahta</surname> <given-names>A.</given-names></name> <name><surname>Culicchia</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Micro-RNA-128 (miRNA-128) down-regulation in glioblastoma targets ARP5 (ANGPTL6), Bmi-1 and E2F-3a, key regulators of brain cell proliferation</article-title>. <source>J. Neuro-Oncol.</source> <volume>98</volume>, <fpage>297</fpage>&#x2013;<lpage>304</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11060-009-0077-0</pub-id></citation></ref>
<ref id="ref4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>L.</given-names></name> <name><surname>Crawford</surname> <given-names>L.</given-names></name> <name><surname>Tong</surname> <given-names>A.</given-names></name> <name><surname>Luu</surname> <given-names>N.</given-names></name> <name><surname>Tanizaki</surname> <given-names>Y.</given-names></name> <name><surname>Shi</surname> <given-names>Y. B.</given-names></name></person-group> (<year>2022</year>). <article-title>Sperm associated antigen 7 is activated by T3 during Xenopus tropicalis metamorphosis via a thyroid hormone response element within the first intron</article-title>. <source>Develop. Growth Differ.</source> <volume>64</volume>, <fpage>48</fpage>&#x2013;<lpage>58</lpage>. doi: <pub-id pub-id-type="doi">10.1111/dgd.12764</pub-id></citation></ref>
<ref id="ref5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>Z. P.</given-names></name> <name><surname>Chen</surname> <given-names>H. F.</given-names></name> <name><surname>Guo</surname> <given-names>Z. H.</given-names></name> <name><surname>Shen</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>W. F.</given-names></name> <name><surname>Xie</surname> <given-names>F. M.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Circular RNAs and their role in exosomes</article-title>. <source>Front. Oncol.</source> <volume>12</volume>:<fpage>848341</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fonc.2022.848341</pub-id></citation></ref>
<ref id="ref6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiao</surname> <given-names>S.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name> <name><surname>Huang</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name> <name><surname>Gao</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Advances in the identification of circular RNAs and research into circRNAs in human diseases</article-title>. <source>Front. Genet.</source> <volume>12</volume>:<fpage>665233</fpage>. doi: <pub-id pub-id-type="doi">10.3389/fgene.2021.665233</pub-id></citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>N. Q.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Cui</surname> <given-names>L.</given-names></name> <name><surname>Ma</surname> <given-names>N.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Hao</surname> <given-names>L. R.</given-names></name></person-group> (<year>2015</year>). <article-title>Expression of intronic miRNAs and their host gene Igf2 in a murine unilateral ureteral obstruction model</article-title>. <source>Braz. J. Med. Biol. Res.</source> <volume>48</volume>, <fpage>486</fpage>&#x2013;<lpage>492</lpage>. doi: <pub-id pub-id-type="doi">10.1590/1414-431X20143958</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lim</surname> <given-names>L. P.</given-names></name> <name><surname>Lau</surname> <given-names>N. C.</given-names></name> <name><surname>Garrett-Engele</surname> <given-names>P.</given-names></name> <name><surname>Grimson</surname> <given-names>A.</given-names></name> <name><surname>Schelter</surname> <given-names>J. M.</given-names></name> <name><surname>Castle</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2005</year>). <article-title>Microarray analysis shows that some microRNAs downregulate large numbers of target mRNAs</article-title>. <source>Nature</source> <volume>433</volume>, <fpage>769</fpage>&#x2013;<lpage>773</lpage>. doi: <pub-id pub-id-type="doi">10.1038/nature03315</pub-id></citation></ref>
<ref id="ref9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malekkou</surname> <given-names>A.</given-names></name> <name><surname>Sevastou</surname> <given-names>I.</given-names></name> <name><surname>Mavrikiou</surname> <given-names>G.</given-names></name> <name><surname>Georgiou</surname> <given-names>T.</given-names></name> <name><surname>Vilageliu</surname> <given-names>L.</given-names></name> <name><surname>Moraitou</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>A novel mutation deep within intron 7 of the GBA gene causes Gaucher disease</article-title>. <source>Mol. Genet. Genom. Med.</source> <volume>8</volume>:<fpage>e1090</fpage>. doi: <pub-id pub-id-type="doi">10.1002/mgg3.1090</pub-id></citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mattick</surname> <given-names>J. S.</given-names></name> <name><surname>Gagen</surname> <given-names>M. J.</given-names></name></person-group> (<year>2001</year>). <article-title>The evolution of controlled multitasked gene networks: the role of introns and other noncoding RNAs in the development of complex organisms</article-title>. <source>Mol. Biol. Evol.</source> <volume>18</volume>, <fpage>1611</fpage>&#x2013;<lpage>1630</lpage>. doi: <pub-id pub-id-type="doi">10.1093/oxfordjournals.molbev.a003951</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ong</surname> <given-names>C. T.</given-names></name> <name><surname>Adusumalli</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Increased intron retention is linked to Alzheimer's disease</article-title>. <source>Neural Regen. Res.</source> <volume>15</volume>, <fpage>259</fpage>&#x2013;<lpage>260</lpage>. doi: <pub-id pub-id-type="doi">10.4103/1673-5374.265549</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Palmiter</surname> <given-names>R. D.</given-names></name> <name><surname>Sandgren</surname> <given-names>E. P.</given-names></name> <name><surname>Avarbock</surname> <given-names>M. R.</given-names></name> <name><surname>Allen</surname> <given-names>D. D.</given-names></name> <name><surname>Brinster</surname> <given-names>R. L.</given-names></name></person-group> (<year>1991</year>). <article-title>Heterologous introns can enhance expression of transgenes in mice</article-title>. <source>Proc. Nati. Acad. Sci. U. S. A.</source> <volume>88</volume>, <fpage>478</fpage>&#x2013;<lpage>482</lpage>. doi: <pub-id pub-id-type="doi">10.1073/pnas.88.2.478</pub-id></citation></ref>
<ref id="ref13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>O. P.</given-names></name> <name><surname>Mishra</surname> <given-names>S.</given-names></name> <name><surname>Sharma</surname> <given-names>G.</given-names></name> <name><surname>Sindhania</surname> <given-names>A.</given-names></name> <name><surname>Kaur</surname> <given-names>T.</given-names></name> <name><surname>Sreehari</surname> <given-names>U.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Evaluation of intron-1 of odorant-binding protein-1 of Anopheles stephensi as a marker for the identification of biological forms or putative sibling species</article-title>. <source>PLoS One</source> <volume>17</volume>:<fpage>e0270760</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0270760</pub-id></citation></ref>
<ref id="ref14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sowalsky</surname> <given-names>A. G.</given-names></name> <name><surname>Xia</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <name><surname>Bubley</surname> <given-names>G. J.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Whole transcriptome sequencing reveals extensive Unspliced mRNA in metastatic castration-resistant prostate cancer</article-title>. <source>Mol. Cancer Res.</source> <volume>13</volume>, <fpage>98</fpage>&#x2013;<lpage>106</lpage>. doi: <pub-id pub-id-type="doi">10.1158/1541-7786.MCR-14-0273</pub-id></citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spijker</surname> <given-names>H. M. V.</given-names></name> <name><surname>Stackpole</surname> <given-names>E. E.</given-names></name> <name><surname>Almeida</surname> <given-names>S.</given-names></name> <name><surname>Katsara</surname> <given-names>O.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Shen</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Ribosome profiling reveals novel regulation of C9ORF72 GGGGCC repeat-containing RNA translation</article-title>. <source>RNA</source> <volume>28</volume>, <fpage>123</fpage>&#x2013;<lpage>138</lpage>. doi: <pub-id pub-id-type="doi">10.1261/rna.078963.121</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Volpe</surname> <given-names>T. A.</given-names></name> <name><surname>Kidner</surname> <given-names>C.</given-names></name> <name><surname>Hall</surname> <given-names>I. M.</given-names></name> <name><surname>Teng</surname> <given-names>G.</given-names></name> <name><surname>Grewal</surname> <given-names>S. I. S.</given-names></name> <name><surname>Martienssen</surname> <given-names>R. A.</given-names></name></person-group> (<year>2002</year>). <article-title>Regulation of heterochromatic silencing and histone H3 lysine-9 methylation by RNAi</article-title>. <source>Science</source> <volume>297</volume>, <fpage>1833</fpage>&#x2013;<lpage>1837</lpage>. doi: <pub-id pub-id-type="doi">10.1126/science.1074973</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vosseberg</surname> <given-names>J.</given-names></name> <name><surname>Schinkel</surname> <given-names>M.</given-names></name> <name><surname>Gremmen</surname> <given-names>S.</given-names></name> <name><surname>Snel</surname> <given-names>B.</given-names></name></person-group> (<year>2022</year>). <article-title>The spread of the first introns in proto-eukaryotic paralogs</article-title>. <source>Commun. Biol.</source> <volume>5</volume>:<fpage>476</fpage>. doi: <pub-id pub-id-type="doi">10.1038/s42003-022-03426-5</pub-id></citation></ref>
<ref id="ref18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoshihama</surname> <given-names>M.</given-names></name> <name><surname>Uechi</surname> <given-names>T.</given-names></name> <name><surname>Asakawa</surname> <given-names>S.</given-names></name> <name><surname>Kawasaki</surname> <given-names>K.</given-names></name> <name><surname>Kato</surname> <given-names>S.</given-names></name> <name><surname>Higa</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2002</year>). <article-title>The human ribosomal protein genes: sequencing and comparative analysis of 73 genes</article-title>. <source>Genome Res.</source> <volume>12</volume>, <fpage>379</fpage>&#x2013;<lpage>390</lpage>. doi: <pub-id pub-id-type="doi">10.1101/gr.214202</pub-id></citation></ref>
<ref id="ref20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>X. Q.</given-names></name> <name><surname>Xue</surname> <given-names>H.</given-names></name> <name><surname>Zheng</surname> <given-names>Y.</given-names></name> <name><surname>Meng</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2016b</year>). <article-title>The evolution mechanism of intron length</article-title>. <source>Genomics</source> <volume>108</volume>, <fpage>47</fpage>&#x2013;<lpage>55</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ygeno.2016.07.004</pub-id></citation></ref>
<ref id="ref19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>X. Q.</given-names></name> <name><surname>Zheng</surname> <given-names>Y.</given-names></name> <name><surname>Meng</surname> <given-names>H.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2016a</year>). <article-title>Analysis on the preference for sequence matching between mRNA sequences and the corresponding introns in ribosomal protein genes</article-title>. <source>J. Theor. Biol.</source> <volume>392</volume>, <fpage>113</fpage>&#x2013;<lpage>121</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jtbi.2015.12.003</pub-id></citation></ref>
<ref id="ref21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>X. O.</given-names></name> <name><surname>Wang</surname> <given-names>H. B.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Lu</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>L. L.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Complementary sequence-mediated exon circularization</article-title>. <source>Cells</source> <volume>159</volume>, <fpage>134</fpage>&#x2013;<lpage>147</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2014.09.001</pub-id>, PMID: <pub-id pub-id-type="pmid">25242744</pub-id></citation></ref>
<ref id="ref22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>X. O.</given-names></name> <name><surname>Chen</surname> <given-names>T.</given-names></name> <name><surname>Xiang</surname> <given-names>J. F.</given-names></name> <name><surname>Yin</surname> <given-names>Q. F.</given-names></name> <name><surname>Xing</surname> <given-names>Y. H.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Circular Intronic long noncoding RNAs</article-title>. <source>Mol. Cell</source> <volume>51</volume>, <fpage>792</fpage>&#x2013;<lpage>806</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.molcel.2013.08.017</pub-id></citation></ref></ref-list>
<fn-group><fn id="fn0004"><p><sup>1</sup><ext-link xlink:href="http://mobyle.pasteur.fr/cgi-bin" ext-link-type="uri">http://mobyle.pasteur.fr/cgi-bin</ext-link></p></fn></fn-group>
</back>
</article>