<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Mol. Biosci.</journal-id>
<journal-title>Frontiers in Molecular Biosciences</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Mol. Biosci.</abbrev-journal-title>
<issn pub-type="epub">2296-889X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">869601</article-id>
<article-id pub-id-type="doi">10.3389/fmolb.2022.869601</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Molecular Biosciences</subject>
<subj-group>
<subject>Mini Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning in RNA Structure Studies</article-title>
<alt-title alt-title-type="left-running-head">Yu et al.</alt-title>
<alt-title alt-title-type="right-running-head">Deep Learning in RNA Structure</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Yu</surname>
<given-names>Haopeng</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1665660/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Qi</surname>
<given-names>Yiman</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1739073/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Ding</surname>
<given-names>Yiliang</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/237915/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Department of Cell and Developmental Biology</institution>, <institution>John Innes Centre</institution>, <institution>Norwich Research Park</institution>, <addr-line>Norwich</addr-line>, <country>United Kingdom</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1286974/overview">Agnieszka Kiliszek</ext-link>, Institute of Bioorganic Chemistry (PAS), Poland</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1670370/overview">Maciej Antczak</ext-link>, Pozna&#x144; University of Technology, Poland</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/25594/overview">Francois Major</ext-link>, Universit&#xe9; de Montr&#xe9;al, Canada</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1675702/overview">Tomasz Zok</ext-link>, Pozna&#x144; University of Technology, Poland</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1166035/overview">Xiaojun Xu</ext-link>, Jiangsu University of Technology, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Haopeng Yu, <email>haopeng.yu@jic.ac.uk</email>; Yiliang Ding, <email>yiliang.ding@jic.ac.uk</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to RNA Networks and Biology, a section of the journal Frontiers in Molecular Biosciences</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>9</volume>
<elocation-id>869601</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Yu, Qi and Ding.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yu, Qi and Ding</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Deep learning, or artificial neural networks, is a type of machine learning algorithm that can decipher underlying relationships from large volumes of data and has been successfully applied to solve structural biology questions, such as RNA structure. RNA can fold into complex RNA structures by forming hydrogen bonds, thereby playing an essential role in biological processes. While experimental effort has enabled resolving RNA structure at the genome-wide scale, deep learning has been more recently introduced for studying RNA structure and its functionality. Here, we discuss successful applications of deep learning to solve RNA problems, including predictions of RNA structures, non-canonical G-quadruplex, RNA-protein interactions and RNA switches. Following these cases, we give a general guide to deep learning for solving RNA structure problems.</p>
</abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>RNA secondary structure</kwd>
<kwd>RNA tertiary structure</kwd>
<kwd>RNA structure prediction</kwd>
<kwd>RNA G-quadruplex</kwd>
<kwd>RNA-protein interaction</kwd>
</kwd-group>
<contract-sponsor id="cn001">Biotechnology and Biological Sciences Research Council<named-content content-type="fundref-id">10.13039/501100000268</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">European Research Council<named-content content-type="fundref-id">10.13039/501100000781</named-content>
</contract-sponsor>
<contract-sponsor id="cn003">Human Frontier Science Program<named-content content-type="fundref-id">10.13039/100004412</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>As a data-driven algorithm, deep learning has shown promise with successful applications in biology, healthcare, and drug discovery (<xref ref-type="bibr" rid="B44">Schmidhuber, 2015</xref>; <xref ref-type="bibr" rid="B3">Angermueller et al., 2016</xref>; <xref ref-type="bibr" rid="B19">Goh et al., 2017</xref>; <xref ref-type="bibr" rid="B11">Ching et al., 2018</xref>). One of the most recent deep learning breakthroughs has been to predict protein structure. In past decades, researchers needed to spend months or even years solving a complex protein structure using experimental methods like nuclear magnetic resonance (NMR) or cryo-electron microscopy (Cryo-EM). Based on this hard-earned data, deep learning models such as Alphafold2 and RoseTTAFold can predict protein structures from amino acid sequences that are remarkably close to those determined experimentally (<xref ref-type="bibr" rid="B4">Baek et al., 2021</xref>; <xref ref-type="bibr" rid="B25">Jumper et al., 2021</xref>). Currently, the protein structure database based on AlphaFold2 predictions has provided nearly one million protein structure models, far exceeding the experimentally determined structures in previous decades (<xref ref-type="bibr" rid="B51">Varadi et al., 2022</xref>). These advances in deep learning methods for predicting protein structures infer their applicability for studying RNA structure.</p>
<p>As a key in the central dogma, RNA is essential for gene expression. RNAs fold into RNA secondary structures by base pairing, which can fold further to form RNA tertiary structures. This RNA folding is extremely important for achieving RNA&#x2019;s diverse and complex biological functions (<xref ref-type="bibr" rid="B35">Mathews and Turner, 2006</xref>; <xref ref-type="bibr" rid="B37">Mortimer et al., 2014</xref>; <xref ref-type="bibr" rid="B58">Zhang and Ding, 2021</xref>). For instance, transfer RNAs usually have a cloverleaf secondary structure with an L-shaped tertiary structure which can fit into ribosomal P and A sites for the translation (<xref ref-type="bibr" rid="B24">Holley et al., 1965</xref>; <xref ref-type="bibr" rid="B26">Kim et al., 1974</xref>). While long noncoding RNAs (lncRNAs) regulate genomic functions through their specific RNA structures (<xref ref-type="bibr" rid="B40">Qian et al., 2019</xref>). Research into the exploits of individual RNA structures and their functional importance is ongoing, with more recent experimental efforts in probing RNA structures over tens of thousands of RNAs in one single experiment transforming the scale for such study. A dramatic increase in RNA structure data resources has laid the foundation for the application of deep learning algorithms in deciphering general features for predicting RNA structure and its functions.</p>
<p>In this review, we have compiled examples from previous research whereby deep learning methods were adopted to solve RNA structure-related problems. Firstly, we introduce the brief process of deep learning modelling through a G-quadruplex classification question. Secondly, we propose that the availability of high-throughput sequencing data has facilitated deep learning modelling of RNA structure-related problems. Next, the experience of deep learning architecture design is presented through the examples of RNA-protein binding prediction and toehold-switch prediction models. Subsequently, we introduce classification and regression in supervised learning through RNA secondary structure prediction and RNA tertiary structure scoring. Lastly, we present several solutions for the interpretation of deep learning models as mentioned in studies. Our review provides an overview of deep learning modelling approaches from the perspective of RNA structure-related research, before providing suggestions for future efforts to address more questions in RNA structure by deep learning.</p>
</sec>
<sec id="s2">
<title>The Basis of Deep Learning Modelling</title>
<p>Deep learning is a machine learning technique, capable of learning abstract features from high-dimensional data through multiple processing layers (<xref ref-type="bibr" rid="B31">LeCun et al., 2015</xref>). Imagine that we propose to build a model to determine whether a guanine-rich (G-rich) DNA or RNA sequence has the potential to fold into a G-quadruplex structure (GQS, a tertiary structure motif that is folded <italic>via</italic> Hoogsteen hydrogen-bonded guanines) (<xref ref-type="bibr" rid="B6">Bochman et al., 2012</xref>; <xref ref-type="bibr" rid="B29">Kwok et al., 2016a</xref>). In traditional modelling, the most important step is called &#x201c;feature extraction&#x201d;. For example, to build this GQS classification or scoring model, certain features need to be extracted based on previous knowledge, such as the number of adjacent guanines (&#x201c;GG&#x201d; or &#x201c;GGG&#x201d;), the length of the loops, the presence of bulge (like &#x201c;GGAG&#x201d;), whether the loop contains cytosine and the probability for the competition of adjacent canonical DNA or RNA secondary structure. However, if there are still features or non-linear combinations of features that are not considered, the model may struggle to achieve highly accurate predictions.</p>
<p>For deep learning, it is possible for modeling without feature extraction (<xref ref-type="fig" rid="F1">Figure 1A</xref>). This particular question can be considered a bi-classification problem, i.e., classifying G-rich sequences into GQS or non-GQS classes. Instead of feature extraction, we can simply input the entire G-rich sequences into the model. We first need to prepare a large number of GQS or non-GQS sequences with clear classification labels, for example, GQS as &#x201c;1&#x201d; and non-GQS as &#x201c;0&#x201d;. After designing a deep learning model, the training process begins. The GQS and non-GQS sequences are fed into the model as inputs and their model-estimated classifications as outputs. Ideally, the classification estimated by the model should be as close to the true class as possible, but usually not in the initial training. Therefore, we need to set an &#x201c;objective function&#x201d; (also known as the &#x201c;loss function&#x201d;) for evaluating the error of the estimated classification from the true classification (<xref ref-type="fig" rid="F1">Figure 1C</xref>). The model then updates its trainable parameters to reduce the error. Typically, deep learning models may have millions of trainable parameters, called weights. The model will calculate a gradient for each weight and determine the adjustment direction to reduce the error (known as &#x201c;gradient descent&#x201d;). Through continuous iterations corresponding to the constant updating of the weights, the classification predicted by the model progressively approaches the true classification (<xref ref-type="fig" rid="F1">Figure 1C</xref>). Ultimately, a powerful deep learning model is derived for predicting the foldability of the G-rich sequence.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Schematic overview of deep learning workflow. <bold>(A)</bold> Data processing. Supervised learning requires explicit labelling of the data, including class numbers in classification questions and values in regression questions. <bold>(B)</bold> Model design. Multilayer perceptron (MLP), convolutional neural network (CNN) and recurrent neural network (RNN) are the three main families of deep learning architecture. Typically, deep learning models assemble different architectures based on data structures. <bold>(C)</bold> Model training. The total training data is first divided into the training set, the validation set and the test set. Then the input data is passed into the model to obtain the predicted values. The loss function is applied to evaluate the difference between the predicted and the true values, whereby the model weights are updated. <bold>(D)</bold> Model interpretation. Features&#x2019; importance can be obtained by <italic>in silico</italic> mutations. For the CNN model, the features can also be evaluated by extracting the weight matrix of the filter.</p>
</caption>
<graphic xlink:href="fmolb-09-869601-g001.tif"/>
</fig>
<p>Several deep learning models are available for GQS classification prediction (<xref ref-type="table" rid="T1">Table 1</xref>). G4NN was trained using the MLP model on 149 experimentally identified RNA GQSs and 179 non-RNA GQSs from the G4RNA database, and the performance outperformed the scoring matrix-based RG4 prediction model (<xref ref-type="bibr" rid="B17">Garant et al., 2015</xref>, <xref ref-type="bibr" rid="B18">2017</xref>). &#x201c;PENGUINN&#x201d;, adopted a sequence as input and a prediction classification score as output (<xref ref-type="bibr" rid="B27">Klimentova et al., 2020</xref>). In addition, it has a higher area under the precision-recall curve value (AUC) than methods based on regular expressions and scoring matrices (<xref ref-type="bibr" rid="B27">Klimentova et al., 2020</xref>). The &#x201c;G4detector&#x201d; introduces RNA structure information to improve GQS prediction (<xref ref-type="bibr" rid="B5">Barshai et al., 2021</xref>) and the &#x201c;DeepG4&#x201d; was trained on <italic>in vivo</italic> G4 data (G4 ChIP-seq) and identified key DNA motifs associated with GQS region activity (<xref ref-type="bibr" rid="B41">Rocher et al., 2021</xref>). &#x201c;PENGUINN&#x201d;, &#x201c;G4detector&#x201d; and &#x201c;DeepG4&#x201d; have been applied to DNA GQS structure prediction at the genome-wide level, further deep learning models based on rG4-seq and SHALiPE-Seq datasets for RNA GQS prediction at the transcriptome-wide level can be expected to emerge in the future as well.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Deep learning-based models in RNA structure.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Function</th>
<th align="center">Name</th>
<th align="center">Model</th>
<th align="center">Method highlights</th>
<th align="center">Link</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="7" align="left">RNA secondary structure prediction</td>
<td align="left">SPOT-RNA <xref ref-type="bibr" rid="B45">Singh et al. (2019)</xref>
</td>
<td align="left">ResNet, LSTM</td>
<td align="left">The model was first trained with a large volume of RNA secondary structures, then trained again using a transfer learning strategy on a small number of validated RNA structures</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/jaswindersingh2/SPOT-RNA/">https://github.com/jaswindersingh2/SPOT-RNA/</ext-link>
</td>
</tr>
<tr>
<td align="left">CDPfold <xref ref-type="bibr" rid="B59">Zhang et al. (2019)</xref>
</td>
<td align="left">CNN, MLP</td>
<td align="left">Predicts the pairing probability matrix of RNA structures and applies dynamic programming methods to generate RNA structures</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/zhangch994/CDPfold">https://github.com/zhangch994/CDPfold</ext-link>
</td>
</tr>
<tr>
<td align="left">DMfold <xref ref-type="bibr" rid="B52">Wang et al. (2019)</xref>
</td>
<td align="left">Bi-LSTM</td>
<td align="left">Predicts the pairing probability matrix of RNA structures and applies IBPMP methods to generate RNA structures</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/linyuwangPHD/RNA-Secondary-Structure-Database">https://github.com/linyuwangPHD/RNA-Secondary-Structure-Database</ext-link>
</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B7">Calonaci et al. (2020)</xref>
</td>
<td align="left">CNN, MLP</td>
<td align="left">Integrates RNA thermodynamic method, chemical probing data and co-evolutionary information into the model</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/bussilab/shape-dca-data">https://github.com/bussilab/shape-dca-data</ext-link>
</td>
</tr>
<tr>
<td align="left">
<xref ref-type="bibr" rid="B53">Willmott et al. (2020)</xref>
</td>
<td align="left">Bi-LSTM</td>
<td align="left">Generates synthetic SHAPE data for RNA structure prediction</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/dwillmott/rna-state-inf">https://github.com/dwillmott/rna-state-inf</ext-link>
</td>
</tr>
<tr>
<td align="left">MXfold2 (<xref ref-type="bibr" rid="B43">Sato et al., 2021</xref>)</td>
<td align="left">CNN, Bi-LSTM</td>
<td align="left">Four types of the folding score were calculated for each nucleotide pair</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/keio-bioinformatics/mxfold2">https://github.com/keio-bioinformatics/mxfold2</ext-link>
</td>
</tr>
<tr>
<td align="left">Ufold (<xref ref-type="bibr" rid="B16">Fu et al., 2022</xref>)</td>
<td align="left">FCN</td>
<td align="left">The input is instead of RNA sequences but a matrix of 16 possible pairings and pairing features for each base pair</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/uci-cbcl/UFold">https://github.com/uci-cbcl/UFold</ext-link>
</td>
</tr>
<tr>
<td align="left">RNA tertiary structure scoring</td>
<td align="left">ARES <xref ref-type="bibr" rid="B50">Townshend et al. (2021)</xref>
</td>
<td align="left">MLP</td>
<td align="left">The model first generated many potential RNA structures by sampling and predicting their different score from the true structure, thus overcoming the problem of insufficient RNA tertiary structures</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="http://drorlab.stanford.edu/ares.html">http://drorlab.stanford.edu/ares.html</ext-link>
</td>
</tr>
<tr>
<td rowspan="4" align="left">G-quadruplexes structure prediction</td>
<td align="left">G4NN <xref ref-type="bibr" rid="B18">Garant et al. (2017)</xref>
</td>
<td align="left">MLP</td>
<td align="left">The model is trained on experimentally validated RNA GQSs and provides a stability score for RNA GQSs</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="http://scottgroup.med.usherbrooke.ca/G4RNA_screener/">http://scottgroup.med.usherbrooke.ca/G4RNA_screener/</ext-link>
</td>
</tr>
<tr>
<td align="left">PENGUINN <xref ref-type="bibr" rid="B27">Klimentova et al. (2020)</xref>
</td>
<td align="left">CNN</td>
<td align="left">Robustness to unbalanced data sets and easy-to-use web interface</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://ml-bioinfo-ceitec.github.io/penguinn/">https://ml-bioinfo-ceitec.github.io/penguinn/</ext-link>
</td>
</tr>
<tr>
<td align="left">G4detector <xref ref-type="bibr" rid="B5">Barshai et al. (2021)</xref>
</td>
<td align="left">CNN</td>
<td align="left">Introduces RNA secondary structure information into the model to improve G4 prediction</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/OrensteinLab/G4detector">https://github.com/OrensteinLab/G4detector</ext-link>
</td>
</tr>
<tr>
<td align="left">DeepG4 <xref ref-type="bibr" rid="B41">Rocher et al. (2021)</xref>
</td>
<td align="left">CNN, MLP</td>
<td align="left">The model is trained on <italic>in vivo</italic> G4 data (G4 ChIP-seq)</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/morphos30/DeepG4">https://github.com/morphos30/DeepG4</ext-link>
</td>
</tr>
<tr>
<td rowspan="2" align="left">RNA structure-mediated protein interactions prediction</td>
<td align="left">iDeepS <xref ref-type="bibr" rid="B38">Pan et al. (2018)</xref>
</td>
<td align="left">CNN, Bi-LSTM</td>
<td align="left">Combines RNA sequence and RNA structure as input during model training</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/xypan1232/iDeepS">https://github.com/xypan1232/iDeepS</ext-link>
</td>
</tr>
<tr>
<td align="left">PrismNet <xref ref-type="bibr" rid="B48">Sun et al. (2021)</xref>
</td>
<td align="left">CNN, ResNet, SE network</td>
<td align="left">Integrates experimental <italic>in vivo</italic> RNA structure data during model training</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/kuixu/PrismNet">https://github.com/kuixu/PrismNet</ext-link>
</td>
</tr>
<tr>
<td align="left">RNA structure-mediated regulatory elements prediction</td>
<td align="left">
<xref ref-type="bibr" rid="B2">Angenent-Mari et al. (2020)</xref>
</td>
<td align="left">MLP</td>
<td align="left">Comparably, this outperforming model was achieved by using RNA sequences directly as input data, rather than extracted features</td>
<td align="left">
<ext-link ext-link-type="uri" xlink:href="https://github.com/lrsoenksen/CL_RNA_SynthBio">https://github.com/lrsoenksen/CL_RNA_SynthBio</ext-link>
</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<title>Data First</title>
<p>In addition to improvements in computer power and high capacity models, the success of deep learning is largely attributable to the availability of large-scale annotated data (<xref ref-type="bibr" rid="B47">Sun et al., 2017</xref>). Fortunately, evolving technologies have provided researchers with a wealth of novel tools, especially high-throughput sequencing (HTS), allowing for an explosion of biological data (<xref ref-type="bibr" rid="B34">Mahmud et al., 2021</xref>). For example, several HTS methods were applied to detect GQS at both DNA and RNA levels (G4-seq and G4 ChIP-seq for DNA, rG4-seq and SHALiPE-seq for RNA), and thus induced the creation of deep learning models such as &#x201c;PENGUINN&#x201d;, &#x201c;G4detector&#x201d;, and &#x201c;DeepG4&#x201d; (<xref ref-type="bibr" rid="B9">Chambers et al., 2015</xref>; <xref ref-type="bibr" rid="B30">Kwok et al., 2016b</xref>; <xref ref-type="bibr" rid="B22">H&#xe4;nsel-Hertsch et al., 2016</xref>; <xref ref-type="bibr" rid="B27">Klimentova et al., 2020</xref>; <xref ref-type="bibr" rid="B54">Yang et al., 2020</xref>; <xref ref-type="bibr" rid="B5">Barshai et al., 2021</xref>; <xref ref-type="bibr" rid="B41">Rocher et al., 2021</xref>).</p>
<p>In RNA structure detection, recent high-throughput <italic>in vitro</italic> and <italic>in vivo</italic> RNA structure chemical probing methods can achieve nucleotide-resolution RNA structure information over tens of thousands of RNAs (RNA structure information of over 50 million nucleotides) in one single experiment, transforming the scale of RNA structure study to an unprecedented level (<xref ref-type="bibr" rid="B15">Ding et al., 2014</xref>; <xref ref-type="bibr" rid="B42">Rouskin et al., 2014</xref>; <xref ref-type="bibr" rid="B46">Spitale et al., 2015</xref>; <xref ref-type="bibr" rid="B56">Yu et al., 2020</xref>). These methods utilise chemicals such as dimethyl sulfate (DMS) and SHAPE (Selective 2&#x2032;-Hydroxyl Acylation analysed by Primer Extension) that determine the single-strandedness of RNA nucleotides. These large volumes of high-throughput sequencing data provide the potential for improving the accuracy of RNA structure prediction. Calonaci et al. established a compound deep learning model to combine multiple channels of RNA sequence information, chemical probing data (single-strandedness information) alongside direct coupling information (derived from co-evolutionary data) to build the thermodynamic prediction method (<xref ref-type="bibr" rid="B7">Calonaci et al., 2020</xref>). Further penalties derived from known RNA structures from the Protein Data Bank (PDB) database were applied as perturbations to the thermodynamic prediction (<xref ref-type="bibr" rid="B7">Calonaci et al., 2020</xref>).</p>
<p>RNA-protein interactions are integral to core biological processes, ranging from transcriptional and post-transcriptional regulation (<xref ref-type="bibr" rid="B8">Castello et al., 2012</xref>). With the increase of high throughput data on RNA binding protein binding sites, like CLIP-Seq, deep learning methods were developed as a consequence to better predict RNA-protein interactions (<xref ref-type="table" rid="T1">Table 1</xref>). Notably, RNA-binding proteins (RBP) recognise specific RNA sequences and specific RNA structure features (<xref ref-type="bibr" rid="B37">Mortimer et al., 2014</xref>; <xref ref-type="bibr" rid="B33">Lewis et al., 2017</xref>; <xref ref-type="bibr" rid="B58">Zhang and Ding, 2021</xref>). For example, PrismNet (Protein-RNA Interaction by Structure informed Modeling using deep neural NETwork) was constructed by integrating RNA sequence, RBP binding sites, and <italic>in vivo</italic> RNA structure information to predict the impact on RNA-protein interaction by one single-nucleotide variant (SNV) that disrupts RNA structures (<xref ref-type="bibr" rid="B48">Sun et al., 2021</xref>). This trained model can also predict the dynamics of the interaction between RNA structural mutations and RNA binding protein from a huge volume of disease-associated mutations; such large-scale assessments are impossible for experimental methods (<xref ref-type="bibr" rid="B48">Sun et al., 2021</xref>).</p>
</sec>
<sec id="s4">
<title>Design of Deep Learning Architectures</title>
<p>The deep learning model is not trained to fit existing data. Instead, it is required to predict independent, unknown data, i.e., generalisation. If the model only has a good fitting on the training set, it is called &#x201c;overfitting&#x201d;, that is, the model may have a large bias in predicting non-training set data. For this purpose, the input data is first normalised and thoroughly shuffled to ensure that the samples have the same distribution. Then, the data set is usually randomly divided into three parts: training set, validation set and test set (<xref ref-type="fig" rid="F1">Figure 1C</xref>) (<xref ref-type="bibr" rid="B20">Goodfellow et al., 2016</xref>). The training set is utilised to fit the model and the validation set for unbiased evaluation of an optimal model. Furthermore, a set of independent, unused samples is required for testing generalisability, and this is the test set. Typically, the ratio of the training set is maximised during model training. By way of illustration, for the ratio of training, validation and test sets in the prediction of RNA secondary structure, SPOT-RNA and E2Efold were established with the ratio of 8:1:1, while CDPfold adopted the ratio of 7:2:1 (<xref ref-type="bibr" rid="B45">Singh et al., 2019</xref>; <xref ref-type="bibr" rid="B59">Zhang et al., 2019</xref>; <xref ref-type="bibr" rid="B10">Chen et al., 2020</xref>). In addition, it is feasible to divide the data into k parts and use 1 of these parts as the test set and the remaining k-1 parts as the training set respectively to obtain the average performance of the model on this data set (also known as the &#x201c;k-fold cross-validation&#x201d;) (<xref ref-type="bibr" rid="B20">Goodfellow et al., 2016</xref>).</p>
<p>The next step is to consider the design of deep learning architectures. There are mainly three families of deep learning architectures: feed-forward neurons network, convolutional neurons network (CNN) and recurrent neurons network (RNN) (<xref ref-type="bibr" rid="B60">Zou et al., 2019</xref>) (<xref ref-type="fig" rid="F1">Figure 1B</xref>). The feed-forward network is the basic architecture and is also known as a multilayer perceptron (MLP) when each layer is a fully connected layer. CNN can receive input data in matrix form and scan the matrix by introducing &#x2018;filters&#x27; to calculate a sum of local weights so that local features can be captured (<xref ref-type="bibr" rid="B28">Krizhevsky et al., 2012</xref>). RNN was originally designed for sequential and time-series data and was enabled to &#x201c;remember&#x201d; the previous state of the series data to influence the current input and output (<xref ref-type="bibr" rid="B57">Zaremba et al., 2015</xref>).</p>
<p>Typically a deep learning model connects one or more architectures like &#x201c;building blocks&#x201d;. Then the entire model works like a pipeline, moving the input data &#x2018;through&#x27; the different architectures, layer by layer, to obtain the predicted values (<xref ref-type="fig" rid="F1">Figure 1C</xref>). For example, iDeepS, a deep learning-based method, combined a CNN and a bidirectional LSTM (Bi-LSTM, is a special kind of RNN) to predict RNA-protein binding preferences. The CNNs were first applied to determine the abstract features of both RNA sequence and <italic>in silico</italic> predicted RNA structure. The close relationship between RNA sequence and structure was then captured by Bi-LSTM for an estimate of possible long-range dependencies (<xref ref-type="bibr" rid="B38">Pan et al., 2018</xref>). The deep-learned weighted representations were then fed into a classification layer for predicting RNA binding protein (RBP) sites (<xref ref-type="bibr" rid="B38">Pan et al., 2018</xref>). Prediction values derived from iDeepS were verified by CLIP-seq by combining UV cross-linking with immunoprecipitation for analysing protein interactions with RNAs (<xref ref-type="bibr" rid="B38">Pan et al., 2018</xref>). This method outperforms the sequence-only prediction methods, indicating the importance of RNA structure features in RNA-protein binding.</p>
<p>In research on toehold switch prediction by deep learning, Angenent-Mari et al. adopted different deep learning models and performed a comparison. The toehold switch is the type of RNA switch that controls downstream translation by its hairpin structure and programmable trans-RNA sequence (<xref ref-type="bibr" rid="B21">Green et al., 2014</xref>). Much effort in RNA synthetic biology has attempted to improve the prediction of toehold switch functionality based on thermodynamic modelling and limited datasets. Angenent-Mari et al. expanded toehold switch datasets from &#x3c;1,000 to the 10<sup>5</sup> level by high-throughput DNA synthesis and sequencing pipeline and then presented different architectures for deep learning models to extract the desired sequence features (<xref ref-type="bibr" rid="B2">Angenent-Mari et al., 2020</xref>). The three-layer deep learning MLP model based on these datasets has a ten-fold improvement on the linear regression model. Then, the model with only RNA sequences as input and the model with 30 rational thermodynamic features as input were compared. Based on the results, the sequence-only model doubled the performance of the feature-extraction model, presumably as the 30 features were not fully inclusive of all the information hidden in the sequences (<xref ref-type="bibr" rid="B2">Angenent-Mari et al., 2020</xref>). Notably, the model with inputs of both thermodynamic features and RNA sequences did not significantly outperform the sequence-only model (<xref ref-type="bibr" rid="B2">Angenent-Mari et al., 2020</xref>). Interestingly, the more complex model architectures, CNN and LSTM, were also utilised for training the same toehold-switch datasets but did not outperform the MLP model (<xref ref-type="bibr" rid="B2">Angenent-Mari et al., 2020</xref>).</p>
</sec>
<sec id="s5">
<title>Supervised Learning in RNA Structure Prediction</title>
<p>The most common form of deep learning is the supervised learning (<xref ref-type="bibr" rid="B31">LeCun et al., 2015</xref>). In supervised learning, the goal of the model is to enable the predictions to be as close as possible to the labels, both discontinuous labels (classification) and continuous labels (regression). The model mentioned above for predicting whether a G-rich sequence can form a GQS is a typical classification question. Also, RNA secondary structure prediction can be achieved by classifying each base pair&#x2019;s status (pair or not) (<xref ref-type="table" rid="T1">Table 1</xref>).</p>
<p>For example, Singh et al. developed an RNA secondary structure prediction model, SPOT-RNA, with RNA sequence as input and the pairing classification status of each potential base pair as output (an L &#xd7; L matrix, L is the length of RNA sequence) (<xref ref-type="bibr" rid="B45">Singh et al., 2019</xref>). SPOT-RNA was developed to train an ensemble of ultra-deep hybrid networks of Residual Neural Network (ResNet) and LSTM with 13,419 RNA structures in the bpRNA database (<xref ref-type="bibr" rid="B13">Danaee et al., 2018</xref>; <xref ref-type="bibr" rid="B45">Singh et al., 2019</xref>). This large model was then trained on a small dataset of 217 validated high-resolution RNA structures. This transfer learning strategy was shown to improve prediction performance by 13% over the next-best model in direct RNA secondary structure prediction. Another software, E2Efold, adopts a deep learning approach to obtain the bi-classification scoring matrix of base pairs from the input RNA sequences, and then constrains the output space by an unrolled algorithm-based Post-Processing Network to achieve an end-to-end RNA structure prediction model (<xref ref-type="bibr" rid="B10">Chen et al., 2020</xref>). In addition to bi-classification, CDPfold adopted a CNN model to predict RNA pairing probability matrices of three labels ("(", ")" and ".") and further combined a dynamic programming algorithm to generate optimal RNA structures (<xref ref-type="bibr" rid="B59">Zhang et al., 2019</xref>). DMfold supports the prediction of seven RNA secondary structure dot-bracket symbols for each base, thus incorporating knowledge of the prediction of RNA pseudoknot structure (<xref ref-type="bibr" rid="B52">Wang et al., 2019</xref>). Non-classification deep learning algorithms have also been applied to RNA secondary structure prediction models. MXfold2 was trained with thermodynamic regularisation to ensure the predicted four types (helix stacking, helix opening, helix closing and unpaired region) of folding scores are close to the calculated free energy (<xref ref-type="bibr" rid="B43">Sato et al., 2021</xref>). Instead of inputting the RNA sequence directly, another model, Ufold, inputs an &#x2018;image&#x27; of the RNA sequence, a matrix of all possible base pairings (canonical and non-canonical base pairing) and pairing features (<xref ref-type="bibr" rid="B16">Fu et al., 2022</xref>). By employing the Fully Convolutional Networks (FCNs), Ufold transformed this RNA sequence &#x2018;image&#x27; into base-pairing probabilities for predicting RNA secondary structures (<xref ref-type="bibr" rid="B16">Fu et al., 2022</xref>).</p>
<p>In supervised learning, if the labels are continuous values, it becomes a regression question. Recently, the model for scoring RNA tertiary structures, ARES, is an example of a deep learning regression model developed with a small amount of training data (<xref ref-type="table" rid="T1">Table 1</xref>). In comparison with the &#x223c;100,000 unique protein structures, there are only 3,335 non-redundant RNA 3D structures (from &#x201c;the Representative Sets of RNA 3D Structures database&#x201d;, version 3.225), whereby most RNA tertiary structures are RNA fragments under 100bp (<xref ref-type="bibr" rid="B32">Leontis and Zirbel, 2012</xref>). This was mainly due to limitations of experimental methods to resolve RNA structures that are largely unstable, very dynamic, and have high plasticity. Unlike protein structure predictions using Alphafold2 and RoseTTAFold based on extensive data resources, only limited known RNA tertiary structures were available for RNA structure prediction. Townshend et al. trained a novel RNA tertiary structure scoring model, the Atomic Rotationally Equivariant Scorer (ARES), by 18 known RNA tertiary structures published between 1994 and 2006 (<xref ref-type="bibr" rid="B14">Das and Baker, 2007</xref>; <xref ref-type="bibr" rid="B50">Townshend et al., 2021</xref>). Unlike a direct prediction of RNA tertiary structure (sequence as input and tertiary structure as output), researchers first generated 1000 RNA tertiary structural models using the Rosetta FARFAR2 sampling method. Each derived RNA structural model was then assessed for the differences between each of its atom&#x2019;s positions and the corresponding atom of the known RNA structures, that is, the true root mean square deviation (RMSD) (<xref ref-type="bibr" rid="B50">Townshend et al., 2021</xref>). Next, the deep learning model was released, with the input being the atoms&#x27; features and the output being the RMSD for each generated RNA tertiary structure model. The ARES model is a sequential model containing an atomic embedding layer, a self-interaction layer, an equivariant convolution layer, and Multilayer Perceptron (MLP) with exponential linear units (<xref ref-type="bibr" rid="B23">Hinton and Salakhutdinov, 2006</xref>; <xref ref-type="bibr" rid="B12">Clevert et al., 2016</xref>; <xref ref-type="bibr" rid="B49">Thomas et al., 2018</xref>). As an RNA tertiary structure scoring model, ARES significantly outperforms the other scoring functions and models despite using a limited number of known RNA structures (<xref ref-type="bibr" rid="B50">Townshend et al., 2021</xref>).</p>
</sec>
<sec id="s6">
<title>Interpreting Deep Learning Models</title>
<p>A deep learning model is typically thought of as a &#x201c;black box&#x201d; containing millions of weights that predict input data as output values. But for researchers in biology, the biological features that the model learns and the biological questions it can explain are more important than just predictions. Contrary to standard statistical models and machine learning methods based on features extraction, deep learning models are challenging to interpret (<xref ref-type="bibr" rid="B60">Zou et al., 2019</xref>).</p>
<p>The most straightforward interpretation means is to perform an <italic>in silico</italic> mutations by algorithm (<xref ref-type="fig" rid="F1">Figure 1D</xref>) (<xref ref-type="bibr" rid="B60">Zou et al., 2019</xref>). This approach requires large-scale <italic>in silico</italic> mutations of the input data followed by re-prediction with the model to assess the impact of changes in the input on the output. For example, for a bi-classification model for GQS prediction, it is possible to simulate base-by-base mutations to alter the input sequence and predict its classification, thus evaluating which nucleotides affect GQS folding. During translation, ribosomes are known to actively unwind the RNA structure where a complex interaction between the ribosome and RNA structure occurs. DeepDRU is a deep learning model for predicting the unwinding state of RNA structures <italic>in vivo</italic> (<xref ref-type="bibr" rid="B55">Yu et al., 2019</xref>). This research demonstrated that ribosome occupancy has a greater impact on the unwinding degree of RNA structure <italic>in vivo</italic> than the sequence itself by simulating mutations of a feature while the rest of the features are fixed (<xref ref-type="bibr" rid="B55">Yu et al., 2019</xref>).</p>
<p>For interpreting CNN models, the convolutional filters in the model can be visualised as heat maps or position weight matrices to extract the high-level patterns learned (<xref ref-type="fig" rid="F1">Figure 1D</xref>). In the model for RBP prediction, DeepBind and iDeep adopted this approach to extract the parameter matrix of the filters from the first-layer convolutional network to identify the RBP binding motifs (<xref ref-type="bibr" rid="B1">Alipanahi et al., 2015</xref>; <xref ref-type="bibr" rid="B39">Pan and Shen, 2017</xref>). Another RPB prediction model, PrismNet, incorporates &#x201c;SmoothGrad&#x201d; to visualise enhanced saliency maps for identifying the high attention regions of RNA sequence leading to the extraction of RBP binding motifs (<xref ref-type="bibr" rid="B48">Sun et al., 2021</xref>). Notably, the interpretation of the model is a purely computational simulation based on a model with well-generalised properties, and the proof of the relevant conclusions may require subsequent experimental validation.</p>
</sec>
<sec sec-type="discussion" id="s7">
<title>Discussion</title>
<p>The emergence of data-driven deep learning approaches integrates technological innovation, &#x201c;Big Data&#x201d; exploitation, and huge computational power to significantly transform the scale for studying RNA structures and their functions. We introduce the basic concepts of deep learning, the importance of data volumn, supervised learning, design of deep learning architectures and model interpretation by reviewing recent deep learning applications in deciphering different aspects of RNA structure studies and highlighting those that demonstrate the best potential for future development.</p>
<p>Although deep learning has shown promise for application in the RNA structure field, there are still some issues that need to be addressed. Firstly overfitting is presently the major risk to deep learning models, especially when faced with limited data size. Advances in technology have led to the development of multi-layered, high-capacity models that can be applied to obtain features for more complex data structures. However, simultaneously, the risk of overfitting arises. In a recent study of RNA secondary structure prediction, it was suggested that E2Efold may suffer from overfitting and is therefore not suitable for predicting broader datasets (<xref ref-type="bibr" rid="B43">Sato et al., 2021</xref>; <xref ref-type="bibr" rid="B16">Fu et al., 2022</xref>). Hence it is far more important to develop highly available experimental datasets in the future than to adopt models with higher levels of capacity. Another challenge is to give a suitable biological interpretation to the purely computationally generated models and the relevant patterns learnt and how to apply deep learning models to complement human experience for functional RNA structure design. With sufficient data, more complex models always imply better performance, but at the same time become difficult to interpret. Typically, the complexity of a model is inversely proportional to the interpretability. In contrast to deep learning, &#x2018;non-deep&#x27; algorithms, such as decision tree algorithms, can have good interpretability by obtaining the weights of individual features. Therefore, we need to make a trade-off between model complexity and interpretability according to the specific objectives.</p>
<p>Encouragingly, with the dramatic increase of high throughput RNA structure data generated from different organisms under diverse conditions, deep learning will be increasingly appreciated by RNA structure researchers and be progressively used to deduce RNA structure information and associated functionality. As the rise of available deep learning models increases, it will become progressively easier for researchers to apply deep learning in their routine data analysis for studying RNA structures.</p>
</sec>
</body>
<back>
<sec id="s8">
<title>Author Contributions</title>
<p>HY and YD conceived the review study; HY and YQ collected research and design diagrams; HY, YQ, and YD wrote the manuscript and approved it for publication.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>YD is supported by the United Kingdom Biotechnology and Biological Sciences Research Council (BBSRC: BBS/E/J/000PR9788) and the European Research Council (ERC: 680324). HY is supported by the Human Frontier Science Program Fellowship (LT001077/2021-L).</p>
</sec>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We thank Ke Li (University of Exeter) for his suggestions. The figure was created with <ext-link ext-link-type="uri" xlink:href="http://BioRender.com">BioRender.com</ext-link>.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alipanahi</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Delong</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Weirauch</surname>
<given-names>M. T.</given-names>
</name>
<name>
<surname>Frey</surname>
<given-names>B. J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Predicting the Sequence Specificities of DNA- and RNA-Binding Proteins by Deep Learning</article-title>. <source>Nat. Biotechnol.</source> <volume>33</volume>, <fpage>831</fpage>&#x2013;<lpage>838</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3300</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Angenent-Mari</surname>
<given-names>N. M.</given-names>
</name>
<name>
<surname>Garruss</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Soenksen</surname>
<given-names>L. R.</given-names>
</name>
<name>
<surname>Church</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Collins</surname>
<given-names>J. J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Deep Learning Approach to Programmable RNA Switches</article-title>. <source>Nat. Commun.</source> <volume>11</volume>, <fpage>5057</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-18677-1</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Angermueller</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>P&#xe4;rnamaa</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Parts</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Stegle</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Deep Learning for Computational Biology</article-title>. <source>Mol. Syst. Biol.</source> <volume>12</volume>, <fpage>878</fpage>. <pub-id pub-id-type="doi">10.15252/msb.20156651</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baek</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>DiMaio</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Anishchenko</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Dauparas</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ovchinnikov</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>G. R.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Accurate Prediction of Protein Structures and Interactions Using a Three-Track Neural Network</article-title>. <source>Science</source> <volume>373</volume>, <fpage>871</fpage>&#x2013;<lpage>876</lpage>. <pub-id pub-id-type="doi">10.1126/science.abj8754</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barshai</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Aubert</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Orenstein</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>G4detector: Convolutional Neural Network to Predict DNA G-Quadruplexes</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinf.</source>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2021.3073595</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bochman</surname>
<given-names>M. L.</given-names>
</name>
<name>
<surname>Paeschke</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zakian</surname>
<given-names>V. A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>DNA Secondary Structures: Stability and Function of G-Quadruplex Structures</article-title>. <source>Nat. Rev. Genet.</source> <volume>13</volume>, <fpage>770</fpage>&#x2013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3296</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Calonaci</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cuturello</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sattler</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bussi</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Machine Learning a Model for RNA Structure Prediction</article-title>. <source>Nar. Genomics Bioinforma.</source> <volume>2</volume>, <fpage>lqaa090</fpage>. <pub-id pub-id-type="doi">10.1093/nargab/lqaa090</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Castello</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Eichelbaum</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Horos</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Beckmann</surname>
<given-names>B. M.</given-names>
</name>
<name>
<surname>Strein</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Insights into RNA Biology from an Atlas of Mammalian mRNA-Binding Proteins</article-title>. <source>Cell</source> <volume>149</volume>, <fpage>1393</fpage>&#x2013;<lpage>1406</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2012.04.031</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chambers</surname>
<given-names>V. S.</given-names>
</name>
<name>
<surname>Marsico</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Boutell</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Di Antonio</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>G. P.</given-names>
</name>
<name>
<surname>Balasubramanian</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>High-throughput Sequencing of DNA G-Quadruplex Structures in the Human Genome</article-title>. <source>Nat. Biotechnol.</source> <volume>33</volume>, <fpage>877</fpage>&#x2013;<lpage>881</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3295</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Umarov</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>RNA Secondary Structure Prediction by Learning Unrolled Algorithms</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2002.05810">http://arxiv.org/abs/2002.05810</ext-link> (Accessed February 17, 2022)</comment>. </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ching</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Himmelstein</surname>
<given-names>D. S.</given-names>
</name>
<name>
<surname>Beaulieu-Jones</surname>
<given-names>B. K.</given-names>
</name>
<name>
<surname>Kalinin</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Do</surname>
<given-names>B. T.</given-names>
</name>
<name>
<surname>Way</surname>
<given-names>G. P.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Opportunities and Obstacles for Deep Learning in Biology and Medicine</article-title>. <source>J. R. Soc. Interface.</source> <volume>15</volume>, <fpage>20170387</fpage>. <pub-id pub-id-type="doi">10.1098/rsif.2017.0387</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Clevert</surname>
<given-names>D.-A.</given-names>
</name>
<name>
<surname>Unterthiner</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hochreiter</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1511.07289">http://arxiv.org/abs/1511.07289</ext-link> (Accessed January 6, 2022)</comment>. </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Danaee</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Rouches</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wiley</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Hendrix</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>bpRNA: Large-Scale Automated Annotation and Analysis of RNA Secondary Structure</article-title>. <source>Nucleic Acids Res.</source> <volume>46</volume>, <fpage>5381</fpage>&#x2013;<lpage>5394</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky285</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Das</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Baker</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Automated De Novo Prediction of Native-like RNA Tertiary Structures</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>104</volume>, <fpage>14664</fpage>&#x2013;<lpage>14669</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0703836104</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kwok</surname>
<given-names>C. K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bevilacqua</surname>
<given-names>P. C.</given-names>
</name>
<name>
<surname>Assmann</surname>
<given-names>S. M.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>
<italic>In Vivo</italic> genome-wide Profiling of RNA Secondary Structure Reveals Novel Regulatory Features</article-title>. <source>Nature</source> <volume>505</volume>, <fpage>696</fpage>&#x2013;<lpage>700</lpage>. <pub-id pub-id-type="doi">10.1038/nature12756</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Nie</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>UFold: Fast and Accurate RNA Secondary Structure Prediction with Deep Learning</article-title>. <source>Nucleic Acids Res.</source> <volume>50</volume>, <fpage>e14</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkab1074</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Garant</surname>
<given-names>J.-M.</given-names>
</name>
<name>
<surname>Luce</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Scott</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Perreault</surname>
<given-names>J.-P.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>G4RNA: an RNA G-Quadruplex Database</article-title>, <source>Database</source>, <volume>2015</volume>, <fpage>bav059</fpage>. <pub-id pub-id-type="doi">10.1093/database/bav059</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Garant</surname>
<given-names>J.-M.</given-names>
</name>
<name>
<surname>Perreault</surname>
<given-names>J.-P.</given-names>
</name>
<name>
<surname>Scott</surname>
<given-names>M. S.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Motif Independent Identification of Potential RNA G-Quadruplexes by G4RNA Screener</article-title>. <source>Bioinformatics</source> <volume>33</volume>, <fpage>3532</fpage>&#x2013;<lpage>3537</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx498</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goh</surname>
<given-names>G. B.</given-names>
</name>
<name>
<surname>Hodas</surname>
<given-names>N. O.</given-names>
</name>
<name>
<surname>Vishnu</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deep Learning for Computational Chemistry</article-title>. <source>J. Comput. Chem.</source> <volume>38</volume>, <fpage>1291</fpage>&#x2013;<lpage>1307</lpage>. <pub-id pub-id-type="doi">10.1002/jcc.24764</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Courville</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Deep Learning</source>. <publisher-name>MIT Press</publisher-name>. </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Green</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Silver</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Collins</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Yin</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Toehold Switches: De-novo-designed Regulators of Gene Expression</article-title>. <source>Cell</source> <volume>159</volume>, <fpage>925</fpage>&#x2013;<lpage>939</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2014.10.002</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>H&#xe4;nsel-Hertsch</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Beraldi</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Lensing</surname>
<given-names>S. V.</given-names>
</name>
<name>
<surname>Marsico</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zyner</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Parry</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>G-quadruplex Structures Mark Human Regulatory Chromatin</article-title>. <source>Nat. Genet.</source> <volume>48</volume>, <fpage>1267</fpage>&#x2013;<lpage>1272</lpage>. <pub-id pub-id-type="doi">10.1038/ng.3662</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Salakhutdinov</surname>
<given-names>R. R.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Reducing the Dimensionality of Data with Neural Networks</article-title>. <source>Science</source> <volume>313</volume>, <fpage>504</fpage>&#x2013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127647</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Holley</surname>
<given-names>R. W.</given-names>
</name>
<name>
<surname>Apgar</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Everett</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Madison</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Marquisee</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Merrill</surname>
<given-names>S. H.</given-names>
</name>
<etal/>
</person-group> (<year>1965</year>). <article-title>Structure of a Ribonucleic Acid</article-title>. <source>Science</source> <volume>147</volume>, <fpage>1462</fpage>&#x2013;<lpage>1465</lpage>. <pub-id pub-id-type="doi">10.1126/science.147.3664.1462</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jumper</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Evans</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Pritzel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Figurnov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ronneberger</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Highly Accurate Protein Structure Prediction with AlphaFold</article-title>. <source>Nature</source> <volume>596</volume>, <fpage>583</fpage>&#x2013;<lpage>589</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Suddath</surname>
<given-names>F. L.</given-names>
</name>
<name>
<surname>Quigley</surname>
<given-names>G. J.</given-names>
</name>
<name>
<surname>McPherson</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sussman</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>A. H. J.</given-names>
</name>
<etal/>
</person-group> (<year>1974</year>). <article-title>Three-Dimensional Tertiary Structure of Yeast Phenylalanine Transfer RNA</article-title>. <source>Science</source> <volume>185</volume>, <fpage>435</fpage>&#x2013;<lpage>440</lpage>. <pub-id pub-id-type="doi">10.1126/science.185.4149.435</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Klimentova</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Polacek</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Simecek</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Alexiou</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>PENGUINN: Precise Exploration of Nuclear G-Quadruplexes Using Interpretable Neural Networks</article-title>. <source>Front. Genet.</source> <volume>11</volume>, <fpage>1287</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2020.568546</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>ImageNet Classification with Deep Convolutional Neural Networks</article-title>,&#x201d; in <source>Advances in Neural Information Processing Systems</source> (<publisher-name>Curran Associates, Inc.</publisher-name>). <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html">https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html</ext-link> (Accessed March 29, 2022)</comment>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kwok</surname>
<given-names>C. K.</given-names>
</name>
<name>
<surname>Marsico</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sahakyan</surname>
<given-names>A. B.</given-names>
</name>
<name>
<surname>Chambers</surname>
<given-names>V. S.</given-names>
</name>
<name>
<surname>Balasubramanian</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016a</year>). <article-title>rG4-seq Reveals Widespread Formation of G-Quadruplex Structures in the Human Transcriptome</article-title>. <source>Nat. Methods</source> <volume>13</volume>, <fpage>841</fpage>&#x2013;<lpage>844</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.3965</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kwok</surname>
<given-names>C. K.</given-names>
</name>
<name>
<surname>Sahakyan</surname>
<given-names>A. B.</given-names>
</name>
<name>
<surname>Balasubramanian</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016b</year>). <article-title>Structural Analysis Using SHALiPE to Reveal RNA G-Quadruplex Formation in Human Precursor MicroRNA</article-title>. <source>Angew. Chem.</source> <volume>128</volume>, <fpage>9104</fpage>&#x2013;<lpage>9107</lpage>. <pub-id pub-id-type="doi">10.1002/ange.201603562</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LeCun</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Deep Learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x2013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Leontis</surname>
<given-names>N. B.</given-names>
</name>
<name>
<surname>Zirbel</surname>
<given-names>C. L.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Nonredundant 3D Structure Datasets for RNA Knowledge Extraction and Benchmarking</article-title>,&#x201d; in <source>RNA 3D Structure Analysis and Prediction</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Leontis,</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Westhof</surname>
<given-names>E.</given-names>
</name>
</person-group> (<publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>281</fpage>&#x2013;<lpage>298</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-25740-7_13</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lewis</surname>
<given-names>C. J. T.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kalsotra</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>RNA Modifications and Structures Cooperate to Guide RNA-Protein Interactions</article-title>. <source>Nat. Rev. Mol. Cell Biol.</source> <volume>18</volume>, <fpage>202</fpage>&#x2013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1038/nrm.2016.163</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mahmud</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kaiser</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>McGinnity</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Hussain</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Learning in Mining Biological Data</article-title>. <source>Cogn. Comput.</source> <volume>13</volume>, <fpage>1</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1007/s12559-020-09773-x</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mathews</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Turner</surname>
<given-names>D. H.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Prediction of RNA Secondary Structure by Free Energy Minimization</article-title>. <source>Curr. Opin. Struct. Biol.</source> <volume>16</volume>, <fpage>270</fpage>&#x2013;<lpage>278</lpage>. <pub-id pub-id-type="doi">10.1016/j.sbi.2006.05.010</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mortimer</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Kidwell</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Doudna</surname>
<given-names>J. A.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Insights into RNA Structure and Function from Genome-wide Studies</article-title>. <source>Nat. Rev. Genet.</source> <volume>15</volume>, <fpage>469</fpage>&#x2013;<lpage>479</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3681</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Rijnbeek</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>H.-B.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Prediction of RNA-Protein Sequence and Structure Binding Preferences Using Deep Convolutional and Recurrent Neural Networks</article-title>. <source>BMC Genomics</source> <volume>19</volume>, <fpage>511</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-018-4889-1</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>H.-B.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>RNA-protein Binding Motifs Mining with a New Hybrid Deep Learning Based Cross-Domain Knowledge Integration Approach</article-title>. <source>BMC Bioinforma.</source> <volume>18</volume>, <fpage>136</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-017-1561-8</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qian</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yeung</surname>
<given-names>P. Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Q. C.</given-names>
</name>
<name>
<surname>Kwok</surname>
<given-names>C. K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Revealing lncRNA Structures and Interactions by Sequencing-Based Approaches</article-title>. <source>Trends Biochem. Sci.</source> <volume>44</volume>, <fpage>33</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/j.tibs.2018.09.012</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rocher</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Genais</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nassereddine</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Mourad</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>DeepG4: A Deep Learning Approach to Predict Cell-type Specific Active G-Quadruplex Regions</article-title>. <source>PLOS Comput. Biol.</source> <volume>17</volume>, <fpage>e1009308</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1009308</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rouskin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zubradt</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Washietl</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kellis</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Weissman</surname>
<given-names>J. S.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Genome-wide Probing of RNA Structure Reveals Active Unfolding of mRNA Structures <italic>In Vivo</italic>
</article-title>. <source>Nature</source> <volume>505</volume>, <fpage>701</fpage>&#x2013;<lpage>705</lpage>. <pub-id pub-id-type="doi">10.1038/nature12894</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sato</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Akiyama</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sakakibara</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>RNA Secondary Structure Prediction Using Deep Learning with Thermodynamic Integration</article-title>. <source>Nat. Commun.</source> <volume>12</volume>, <fpage>941</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-21194-4</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Deep Learning in Neural Networks: An Overview</article-title>. <source>Neural Netw.</source> <volume>61</volume>, <fpage>85</fpage>&#x2013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2014.09.003</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hanson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Paliwal</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>RNA Secondary Structure Prediction Using an Ensemble of Two-Dimensional Deep Neural Networks and Transfer Learning</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>5407</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-13395-9</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Spitale</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Flynn</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Q. C.</given-names>
</name>
<name>
<surname>Crisalli</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Jung</surname>
<given-names>J.-W.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Structural Imprints <italic>In Vivo</italic> Decode RNA Regulatory Mechanisms</article-title>. <source>Nature</source> <volume>519</volume>, <fpage>486</fpage>&#x2013;<lpage>490</lpage>. <pub-id pub-id-type="doi">10.1038/nature14263</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Shrivastava</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Revisiting Unreasonable Effectiveness of Data in Deep Learning Era</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>, <conf-loc>Venice, Italy</conf-loc>, <conf-date>October 2017</conf-date>, <fpage>843</fpage>&#x2013;<lpage>852</lpage>. <pub-id pub-id-type="doi">10.1109/iccv.2017.97</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y. T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Predicting Dynamic Cellular Protein-RNA Interactions by Deep Learning Using <italic>In Vivo</italic> RNA Structures</article-title>. <source>Cell Res.</source> <volume>31</volume>, <fpage>495</fpage>&#x2013;<lpage>516</lpage>. <pub-id pub-id-type="doi">10.1038/s41422-021-00476-y</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Thomas</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Smidt</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kearnes</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kohlhoff</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Tensor Field Networks: Rotation- and Translation-Equivariant Neural Networks for 3D Point Clouds</article-title>. <comment>ArXiv180208219 Cs. Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1802">http://arxiv.org/abs/1802</ext-link> (Accessed January 6, 2022)</comment>. </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Townshend</surname>
<given-names>R. J. L.</given-names>
</name>
<name>
<surname>Eismann</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Watkins</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Rangan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Karelina</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Das</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Geometric Deep Learning of RNA Structure</article-title>. <source>Science</source> <volume>373</volume>, <fpage>1047</fpage>&#x2013;<lpage>1051</lpage>. <pub-id pub-id-type="doi">10.1126/science.abe5650</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Varadi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Anyango</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Deshpande</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nair</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Natassia</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yordanova</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models</article-title>. <source>Nucleic Acids Res.</source> <volume>50</volume>, <fpage>D439</fpage>&#x2013;<lpage>D444</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkab1061</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>DMfold: A Novel Method to Predict RNA Secondary Structure with Pseudoknots Based on Deep Learning and Improved Base Pair Maximization Principle</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <fpage>143</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2019.00143</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Willmott</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Murrugarra</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Improving RNA Secondary Structure Prediction via State Inference with Deep Recurrent Neural Networks</article-title>. <source>Comput. Math. Biophys.</source> <volume>8</volume>, <fpage>36</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1515/cmb-2020-0002</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Cheema</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Duncan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Umar</surname>
<given-names>M. I.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>RNA G-Quadruplex Structures Exist and Function <italic>In Vivo</italic> in Plants</article-title>. <source>Genome Biol.</source> <volume>21</volume>, <fpage>226</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-02142-9</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Mao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Deciphering the Rules of mRNA Structure Differentiation in <italic>Saccharomyces cerevisiae In Vivo</italic> and <italic>In Vitro</italic> with Deep Neural Networks</article-title>. <source>RNA Biol.</source> <volume>16</volume>, <fpage>1044</fpage>&#x2013;<lpage>1054</lpage>. <pub-id pub-id-type="doi">10.1080/15476286.2019.1612692</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>RSVdb: a Comprehensive Database of Transcriptome RNA Structure</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume>. <pub-id pub-id-type="doi">10.1093/bib/bbaa071</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Zaremba</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Vinyals</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Recurrent Neural Network Regularization</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1409.2329">http://arxiv.org/abs/1409.2329</ext-link> (Accessed March 29, 2022)</comment>. </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Novel Insights into the Pervasive Role of RNA Structure in Post-transcriptional Regulation of Gene Expression in Plants</article-title>. <source>Biochem. Soc. Trans.</source> <volume>49</volume>, <fpage>1829</fpage>&#x2013;<lpage>1839</lpage>. <pub-id pub-id-type="doi">10.1042/BST20210318</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>A New Method of RNA Secondary Structure Prediction Based on Convolutional Neural Network and Dynamic Programming</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <fpage>467</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2019.00467</pub-id> </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Huss</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Abid</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mohammadi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Torkamani</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Telenti</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A Primer on Deep Learning in Genomics</article-title>. <source>Nat. Genet.</source> <volume>51</volume>, <fpage>12</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1038/s41588-018-0295-5</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>