<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2025.1626539</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Identification of DNA N6-methyladenine modifications in the rice genome with a fine-tuned large language model</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Yichi</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/3080159/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Hao</given-names>
</name>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Xiang</surname>
<given-names>Shicheng</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/3080283/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Lv</surname>
<given-names>Zhibin</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/778029/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/project-administration/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>College of Biomedical Engineering, Sichuan University</institution>, <addr-line>Chengdu</addr-line>,&#xa0;<country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Shanwen Sun, Northeast Forestry University, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Lixin Cheng, Jinan University, China</p>
<p>Fei Guo, Central South University, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Zhibin Lv, <email xlink:href="mailto:lvzhibin@pku.edu.cn">lvzhibin@pku.edu.cn</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>25</day>
<month>06</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>16</volume>
<elocation-id>1626539</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>05</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>06</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Zhang, Chen, Xiang and Lv</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Zhang, Chen, Xiang and Lv</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>DNA N6-methyladenine (6mA) plays a significant role in various biological processes. In the rice genome, 6mA is involved in important processes such as growth and development, influencing gene expression. Therefore, identifying the 6mA locus in rice is crucial for understanding its complex gene expression regulatory system. Although several useful prediction models have been proposed, there is still room for improvement. To address this, we propose an architecture named iRice6mA-LMXGB that integrates a fine-tuned large language model to identify the 6mA locus in rice. Specifically, our method consists of two main components: (1) a BERT model for feature extraction and (2) an XGBoost module for 6mA classification. We utilize a pre-trained DNABERT-2 model to initialize the parameters of the BERT component. Through transfer learning, we fine-tune the model on the rice 6mA recognition task, converting raw DNA sequences into high-dimensional feature vectors. These features are then processed by an XGBoost algorithm to generate predictions. To further validate the effectiveness of our fine-tuning strategy, we employ UMAP(Uniform Manifold Approximation and Projection) visualization. Our approach achieves a validation accuracy of 0.9903 in a five-fold cross-validation setting and produces a receiver operating characteristic (ROC) curve with an area under the curve (AUC) of 0.9994. Compared to existing predictors trained on the same dataset, our method demonstrates superior performance. This study provides a powerful tool for advancing research in rice 6mA epigenetics.</p>
</abstract>
<kwd-group>
<kwd>rice genome</kwd>
<kwd>N6-methyladenine</kwd>
<kwd>large language model</kwd>
<kwd>BERT</kwd>
<kwd>UMAP visualization</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="1"/>
<equation-count count="13"/>
<ref-count count="75"/>
<page-count count="10"/>
<word-count count="4869"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Functional and Applied Plant Genomics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>N6-methyladenine(6mA) is produced by methylation of the N6 position of adenine and has been found in bacteria, eukaryotes, and archaea (<xref ref-type="bibr" rid="B64">Zhang et&#xa0;al., 2015</xref>; <xref ref-type="bibr" rid="B37">O&#x2019;Brown and Greer, 2016</xref>). Rice is one of the most important cereal crops in the world. Within the rice genome, 6mA serves as a critical epigenetic modification, regulating gene expression through methylation at the N6 position of adenine (<xref ref-type="bibr" rid="B33">Lv et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B5">Chen et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B16">Jin et&#xa0;al., 2022</xref>). Studies have shown that 6mA in rice plays a vital role in many biological functions. For example, 6mA in rice is associated with stress response and helps rice to better adapt to adversity (<xref ref-type="bibr" rid="B65">Zhang et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B11">Ding et&#xa0;al., 2023</xref>). It is also associated with reproduction and regulates the growth and development of rice (<xref ref-type="bibr" rid="B67">Zhou et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B61">Yang et&#xa0;al., 2024</xref>). Zhou et&#xa0;al. discovered that 6mA is highly enriched in specific sequence motifs, conserved DNA sequence patterns that serve as recognition sites for epigenetic regulators. These motifs include AGG and GAGG, which are assumed to represent the binding elements of methyltransferase complexes or chromatin associated proteins. 6mA methylation preferentially occurred on these specific nucleotide motifs, indicating their functional significance in epigenetic regulation (<xref ref-type="bibr" rid="B22">Lee et&#xa0;al., 2018</xref>). And this methylation pattern is tightly linked to the drought stress response in rice (<xref ref-type="bibr" rid="B68">Zhou et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B61">Yang et&#xa0;al., 2024</xref>). In addition, 6mA can directly affect seed size and yield formation by regulating the expression of endosperm development-related genes (<xref ref-type="bibr" rid="B67">Zhou et&#xa0;al., 2021</xref>). In recent years, epigenetic breeding strategies based on CRISPR-6mA editing technology have provided new ideas to improve disease resistance and yield in rice by targeting modification of the 6mA locus (<xref ref-type="bibr" rid="B43">Romero and Gatica-Arias, 2019</xref>). However, traditional experimental methods such as SMRT-seq for detecting 6mA locus have the limitations of high cost and low throughput, and there is an urgent need to develop efficient computational prediction models to guide subsequent functional studies (<xref ref-type="bibr" rid="B70">Zhu et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B50">Wang L. et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B6">Chen et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B30">Liu et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B45">Shao et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B55">Xie H. et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B69">Zhou et&#xa0;al., 2024</xref>).</p>
<p>In recent years, machine and deep learning approaches have successfully addressed many challenges in identifying 6mA modifications in rice genomes (<xref ref-type="bibr" rid="B46">Sinha et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B51">Wang R. et&#xa0;al., 2023</xref>). In 2019, Chen et&#xa0;al. developed the first method for predicting DNA 6mA sites in rice, called i6mA-Pred, utilizing nucleotide chemical property (NCP) features and a support vector machine (SVM) as the classifier (<xref ref-type="bibr" rid="B7">Chen et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B74">Zou et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B35">Meher et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B53">Wang Y. et&#xa0;al., 2024</xref>). Subsequent research has seen the emergence of various single-classifier-based prediction methods, including MM-6mAPred (<xref ref-type="bibr" rid="B39">Pian et&#xa0;al., 2019</xref>), i6mA-DNCP (<xref ref-type="bibr" rid="B38">Park et&#xa0;al., 2020</xref>), iN6-methylat (<xref ref-type="bibr" rid="B21">Le, 2019</xref>), and iDNA6mA-rice (<xref ref-type="bibr" rid="B31">Lv et&#xa0;al., 2019</xref>). Moreover, ensemble learning models combining multiple classifiers, such as csDMA (<xref ref-type="bibr" rid="B29">Liu et&#xa0;al., 2019</xref>), SDM6A (<xref ref-type="bibr" rid="B3">Basith et&#xa0;al., 2019</xref>), 6mA-Finder (<xref ref-type="bibr" rid="B58">Xu et&#xa0;al., 2020</xref>), Meta-i6mA (<xref ref-type="bibr" rid="B13">Hasan et&#xa0;al., 2021</xref>), i6mA-VC (<xref ref-type="bibr" rid="B59">Xue et&#xa0;al., 2021</xref>), i6mA-Vote (<xref ref-type="bibr" rid="B48">Teng et&#xa0;al., 2022</xref>), and EpiSemble (<xref ref-type="bibr" rid="B46">Sinha et&#xa0;al., 2023</xref>), have been developed to enhance model performance and robustness. Deep learning techniques have evolved from traditional artificial neural network frameworks and have shown significant improvement in predictive power across multiple research domains. With the development of deep learning and its excellent performance, researchers began to apply it to the problem of DNA 6mA site prediction. In 2019, Yu et&#xa0;al. developed a prediction model called SNNRice6mA (<xref ref-type="bibr" rid="B62">Yu and Dai, 2019</xref>) based on convolutional neural networks (CNNs) through single-nucleotide one-hot coding, obtaining an accuracy of 0.920. Another group of researchers, Lv et&#xa0;al., proposed a convolutional neural network iRicem6A-CNN (<xref ref-type="bibr" rid="B32">Lv et&#xa0;al., 2021</xref>) based on a dinucleotide one-hot encoder in 2020, achieving an accuracy of 0.938 for 5-fold cross-validation. However, it is worth noting that CNNs are limited in focusing on only part of the information. Deep6mA (<xref ref-type="bibr" rid="B24">Li Z. et&#xa0;al., 2021</xref>), which consists of a convolutional neural network (CNN) and a bidirectional LSTM (BLSTM) module to solve the long-distance nucleotide association problem by learning contextual dependencies of the sequences, was proposed by Li et&#xa0;al. in 2021 and achieved a 5-fold cross-validation accuracy of 0.940.</p>
<p>Over the last few years, large-scale language modeling (LLM) has progressed tremendously (<xref ref-type="bibr" rid="B25">Li H. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B56">Xie X. et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B8">Chen et&#xa0;al., 2025</xref>). The well-known model, ChatGPT, is a fine-tuned version of the base GPT-3 model. By learning contextual text in a self-supervised manner, it can both understand and generate human language (<xref ref-type="bibr" rid="B10">Devlin et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B52">Wang G. et&#xa0;al., 2024</xref>). DNA sequences exhibit similarities to natural language. Nucleotides, the building blocks of nucleic acids, serve as &#x201c;words&#x201d; within biological systems&#x2019; &#x201c;languages&#x201d;. LLMs can be adapted for the analysis of biological sequence data by leveraging the structure of DNA and protein sequences as analogous to natural language texts (<xref ref-type="bibr" rid="B17">Jumper et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B42">Rives et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B54">Wei et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B26">Li T. et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B27">Li Y. et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B41">Qiao et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B19">Lai et&#xa0;al., 2025</xref>; <xref ref-type="bibr" rid="B57">Xie et&#xa0;al., 2025</xref>). There have been many breakthroughs in LLMs for applications in biology, such as AlphaFold2, a protein prediction model with very high accuracy (<xref ref-type="bibr" rid="B17">Jumper et&#xa0;al., 2021</xref>), the Geneformer model trained on data from about 10 million human single-cell RNA sequences (<xref ref-type="bibr" rid="B73">Zou et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B49">Theodoris et&#xa0;al., 2023</xref>), and DNABERT, a transformer-based DNA pre-training model (<xref ref-type="bibr" rid="B15">Ji et&#xa0;al., 2021</xref>). While LLMs demonstrate potential for identifying patterns and correlations in noisy biological datasets (<xref ref-type="bibr" rid="B20">Lam et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B47">Soylu and Sefer, 2024</xref>; <xref ref-type="bibr" rid="B56">Xie X. et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B28">Liu et&#xa0;al., 2025</xref>), they have yet to gain acceptance within plant science research. To date, LLMs have not been employed in the study of 6mA locus prediction in rice.</p>
<p>In this study, we develop a large language model-based transfer learning model called iRice6mA-LMXGB. it consists of a pre-trained DNABERT2 model and an XGBoost model. It contains a unique fine-tuning architecture that relies exclusively on DNA sequence data to distinguish 6mA sequences from non-6mA sequences. Experimental results demonstrate the model&#x2019;s outstanding performance, achieving a validation accuracy of 0.9903 through 5-fold cross-validation. Compared to all previous methods tested on standard datasets, iRice6mA-LMXGB significantly outperforms them, suggesting that this novel approach has the potential to transform biological sequence modeling.</p>
</sec>
<sec id="s2" sec-type="materials|methods">
<label>2</label>
<title>Materials and methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Benchmark dataset</title>
<p>In this study, we utilized the rice dataset constructed by Lv et&#xa0;al. (<xref ref-type="bibr" rid="B33">2020</xref>) for model training and evaluation using 5-fold cross-validation. To ensure the high quality of the data, sequences with greater than 80% similarity were removed via the CD-HIT program (<xref ref-type="bibr" rid="B23">Li and Godzik, 2006</xref>). The dataset is made of 154,000 sequences with 6mA sites and 154,000 sequences without 6mA sites. This is a widely adopted and balanced rice dataset. During model training, unbalanced datasets may lead to unreliable results. The majority class samples are dominant and the model will favor the majority class during training, thus ignoring the minority class. This may result in the model having high accuracy for the majority class but low recognition for the minority class during prediction. For ease of reference, we denote it as &#x201c;rice-Lv&#x201d; throughout this study. Both positive and negative sequences in the rice-Lv dataset are 41 base pairs in length. Positive sequences represent 6mA modifications at their centers, while negative sequences lack such modifications at theirs. By employing this well-established dataset, we enable a fair comparison between our method and those previously reported.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Architecture of iRice6mA-LMXGB</title>
<p>The architecture of iRice6A-LMXGB is presented in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>, comprising two main components: the pre-trained DNABERT-2 module and the XGBoost module. DNABERT-2 is a pre-trained BERT model specifically designed for encoding DNA sequences. It can efficiently identify complex long-range dependencies in these sequences (<xref ref-type="bibr" rid="B66">Zhou et&#xa0;al., 2023</xref>). And this module will undergo further fine-tuning in this study. XGBoost&#x2019;s superior performance, particularly in terms of speed and accuracy when processing large-scale datasets, enables its extensive use in solving classification problems (<xref ref-type="bibr" rid="B4">Chen and Guestrin, 2016</xref>; <xref ref-type="bibr" rid="B60">Yang et&#xa0;al., 2021</xref>). It utilizes the feature vectors output from the DNABERT-2 model to generate final prediction results. A detailed explanation of the model follows.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>The proposed modeling framework, iRice6mA-LMXGB.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1626539-g001.tif">
<alt-text content-type="machine-generated">Diagram illustrating a DNA sequence analysis model using DNABert2 and XGBoost. The input sequence undergoes token and positional embedding. Multi-head attention and feed-forward layers process embeddings. Outputs are classified via XGBoost, determining if sequences are 6mA or non-6mA.</alt-text>
</graphic>
</fig>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>DNABERT-2</title>
<p>DNABERT-2 is an iterative version of DNABERT. DNABERT is the first BERT-based DNA language model (<xref ref-type="bibr" rid="B15">Ji et&#xa0;al., 2021</xref>). Rigorously trained on a comprehensive genomic dataset encompassing the entire human genome, DNABERT offers a linguistic perspective for genomic analysis. While widely adopted, the initial version of DNABERT exhibited notable technical limitations. Specifically, DNABERT faced two critical challenges: first, its training data is limited to a single-species genome, which makes it difficult for the model to capture sequence-conserving patterns and diversity features across species; second, the k-mer sequence partitioning mechanism it employs not only triggers the hidden danger of data leakage during the training process, but also significantly increases the computational complexity (<xref ref-type="bibr" rid="B36">Moeckel et&#xa0;al., 2024</xref>). Such limitations underscore the pressing need for innovation and improvement in DNA-based language modeling research. To address these challenges, DNABERT-2 introduced significant improvements in both areas. First, it breaks through species boundaries and employs cross-species genomic datasets for pre-training, significantly enhancing the model&#x2019;s ability to recognize evolutionarily conserved regions and species specificity. Second, at the data processing stage, DNABERT-2 employs byte-pair encoding (BPE), a novel tokenization method that replaces traditional k-mer partitioning. This is a data compression algorithm widely used in large-scale language models (<xref ref-type="bibr" rid="B44">Sennrich et&#xa0;al., 2015</xref>), which effectively solves the risk of data leakage and improves computational efficiency, successfully overcoming the limitations of k-mer tokenization. As demonstrated by Zhihan et&#xa0;al.&#x2019;s comparative analysis, compared to conventional 6-mer tokenization methods, the byte-pair encoding (BPE) implementation exhibits superior sequence compression efficiency, reducing the tokenized sequence length by a factor of 5. The dramatic reduction in dimensionality directly improves the computational efficiency of processing genome sequences (<xref ref-type="bibr" rid="B66">Zhou et&#xa0;al., 2023</xref>).</p>
<p>The BERT model consists of two independent components: the module responsible for preprocessing BERT input and the pre-training BERT module. In the BERT input preprocessing module, DNABERT-2 utilizes BPE to tokenize DNA sequences. Byte Pair Encoding (BPE) is a subword tokenization algorithm commonly employed in NLP Natural Language Processing) tasks. Its key mechanism lies in iteratively merging character pairs of the highest frequency to construct a vocabulary of subwords. During tokenization, DNABERT-2 appends a [CLS] token at the sequence start and a [SEP] token at the end. Then, each token is put into an embedding module and converted into a vector. The DNABERT-2 model uses the ALiBi(Attention with Linear Biases) (<xref ref-type="bibr" rid="B40">Press et&#xa0;al., 2021</xref>) approach, which does not add positional embeddings to the input, but rather adds a non-learned embedding in every Attention computation to add a non-learning bias and a fixed set of statics to combine the location information with the Attention score. DNABERT-2 employs a transformer encoder architecture as the backbone of its pre-trained BERT module. The feature matrix is constructed by cascading encoders layer by layer across the network&#x2019;s layers (L). Each encoder comprises three components: multi-head self-attention units, position-wise feed-forward neural networks, and normalization layers. Within the i-th encoder stage, the multi-head self-attention mechanism operates as follows.</p>
<disp-formula>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mtext>Multihead</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mtext>Concat</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>O</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>For the i-th encoder, the input matrix <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mtext>&#xa0;</mml:mtext>
</mml:mrow>
</mml:math>
</inline-formula> is handled through n self-attentive heads for processing. The outputs of these heads are then transformed by the output transformation matrix <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, which is computed in detail for each <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:msup>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> as follows.</p>
<disp-formula>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mtext>head</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mtext>softmax</mml:mtext>
<mml:mo>(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo>)</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> serve as the transformation matrices for the query, key, and value components of each head, respectively. <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the dimension of the matrix.</p>
<p>Specifically, after computing <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>H</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> in the multi-head attention mechanism, this resultant output is added to the residual connection of the original input <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> for normalization. The computation proceeds according to the formula below.</p>
<disp-formula>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mtext>LayerNorm</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>MultiHead</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>After normalization, the processed data is passed through a feed-forward neural network using the following formula:</p>
<disp-formula>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi>max</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im13">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the trainable weight parameters within the feed-forward layer.</p>
<p>The output of the i-th encoder is achieved through normalization of the residual connection between <inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. Below is the corresponding formula.</p>
<disp-formula>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mtext>LayerNorm</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:mtext>FFN</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Finally, the output of the DNABERT-2 can be obtained by cascading the L encoders as follows.</p>
<disp-formula>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where d denotes the dimension of the word vector and N represents the total number of tokens.</p>
<p>DNABERT-2 follows the BERT model architecture, defined by three key parameters: L = 12, H = 768, and A = 12. The parameter L specifies the number of transformer layers (totaling 12). The parameter H determines the hidden layer size, with each token represented as a 768-dimensional vector. The parameter A specifies the number of attention heads (totaling 12). In this study, we use the full fine-tuning (FFT) (<xref ref-type="bibr" rid="B9">Church et&#xa0;al., 2021</xref>) method, which treats rice DNA sequences as &#x201c;sentences in natural language&#x201d; and inputs them into the DNABERT-2 module to adjust and update all the parameters. Finally, we use the BERT model to convert them into fixed-length feature vectors to obtain the original feature matrix before fine-tuning and the feature matrix after 200 cycles of updating.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>XGBOOST</title>
<p>The XGBoost classifier is a gradient boosting method that integrates regression trees (<xref ref-type="bibr" rid="B3">Basith et&#xa0;al., 2019</xref>). The objective function of the model is <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>j</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mtext>&#x3a9;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is the training loss function with the expression:</p>
<disp-formula>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>&#x3b8;</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mi>l</mml:mi>
</mml:mstyle>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> represents the training loss function for each sample. <inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the true value of the i-th sample. <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the estimated value of the i-th sample.</p>
<p>Then the estimated value of the i-th sample is expressed as:</p>
<disp-formula>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>K is the number of integrated trees, and F denotes the space of all possible decision trees. <inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is a specific categorical regression tree (CART).<inline-formula>
<mml:math display="inline" id="im22">
<mml:mrow>
<mml:mtext>&#x3a9;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>f</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is the tree structure complexity function, and its specific form is:</p>
<disp-formula>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mtext>&#x3a9;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>f</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mtext>&#x3b3;</mml:mtext>
<mml:mi>T</mml:mi>
<mml:mo>+</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mtext>&#x3bb;</mml:mtext>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msubsup>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The parameter <inline-formula>
<mml:math display="inline" id="im23">
<mml:mtext>&#x3b3;</mml:mtext>
</mml:math>
</inline-formula> restricts the number of leaf nodes <inline-formula>
<mml:math display="inline" id="im24">
<mml:mi>T</mml:mi>
</mml:math>
</inline-formula> of the tree to control the complexity of the model. And the parameter <inline-formula>
<mml:math display="inline" id="im25">
<mml:mtext>&#x3bb;</mml:mtext>
</mml:math>
</inline-formula> constrains the sum of the weights <inline-formula>
<mml:math display="inline" id="im26">
<mml:mrow>
<mml:msubsup>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> of each leaf node to suppress overfitting. The objective function is continuously optimized by adjusting the parameters for the optimal result. In this way, the XGBoost classifier finally outputs the prediction results of the rice sequence about 6mA by receiving the extracted feature vectors from DNABERT-2.</p>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Evaluation metrics and methods</title>
<p>In this study, we validate our approach using a traditional 5-fold cross-validation method and compare it to previous studies based on the benchmark dataset rice-Lv.</p>
<p>we will combine five metrics, including accuracy (ACC), sensitivity (Sn), specificity (Sp), Matthew&#x2019;s correlation coefficient (MCC), and area under the curve (AUC), to comprehensively evaluate the prediction performance of our model (<xref ref-type="bibr" rid="B72">Zou et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B75">Zulfiqar et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B12">Guo et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B14">Huang et&#xa0;al., 2024</xref>; <xref ref-type="bibr" rid="B71">Zhu et&#xa0;al., 2024</xref>).</p>
<p>ACC indicates the overall correctness of the model prediction and is a basic benchmark used to evaluate the model performance, which can be expressed as:</p>
<disp-formula>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The sensitivity Sn, also known as the true positive rate (TPR), is expressed as:</p>
<disp-formula>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The specificity Sp, also known as the true negative rate (TNR), is expressed as:</p>
<disp-formula>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>MCC is a composite metric that assesses the overall quality of classification model predictions by examining the performance of the classification model in each of the four quadrants of the confusion matrix. The superior score reflects the balanced excellence between true positives (TP), true negatives (TN), false negatives (FN) and false positives (FP). It can be defined as:</p>
<disp-formula>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The last performance metric we use is AUC, defined as the value of the area under the subject&#x2019;s operating characteristic curve. AUC is also an important measure of the performance of a dichotomous model. The larger the value of AUC, the better the model performs. AUC is a floating-point number between 0 and 1. 1 indicates that the model predicts perfectly, whereas 0.5 indicates that the model is similar to a random prediction (<xref ref-type="bibr" rid="B63">Zhang et&#xa0;al., 2025</xref>).</p>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results and discussion</title>
<sec id="s3_1">
<label>3.1</label>
<title>Model performance analysis</title>
<p>In this study, we developed three models. For the first model, we directly used the pre-trained DNABERT-2 to extract 768-dimensional features from rice DNA and fed them into an XGBoost classifier for prediction tasks. The XGBoost classifier shows unique advantages in genomics data classification tasks, mainly due to its ability to efficiently handle high-dimensional sparse data and its built-in regularization mechanism. Our dataset, with more than 300,000 samples, is characterized by high feature dimensionality, and XGBoost is able to efficiently capture nonlinear interaction effects through the gradient boosting framework combined with second-order derivative optimization. Its regularization term can in turn suppress overfitting and enhance model generalization (<xref ref-type="bibr" rid="B4">Chen and Guestrin, 2016</xref>). Cross-validation results showed ACC=0.6259, Sn=0.6207, Sp=0.6312, MCC=0.2519, and auROC=0.6728 for this configuration. For the second model, we loaded the rice-Lv dataset into the DNABERT-2 module and conducted 200 iteration loops to develop a fine-tuned version of the model. The 5-fold cross-validation scores were ACC=0.9903, Sn=0.9898, Sp=0.9907, MCC=0.9805, auROC=0.9994 which are 58.22%, 59.47%, 56.96%, 289.24%, and 48.54%, respectively higher than those of the non-fine-tuned model. For the third model, we utilized LightGBM&#x2019;s built-in function to assess and prioritize feature importance using features extracted from the fine-tuned DNABERT-2 model (<xref ref-type="bibr" rid="B18">Ke et&#xa0;al., 2017</xref>). The feature ranking principle of LightGBM is based on the Gradient Boosting Decision Tree (GBDT) framework, which evaluates feature importance by quantifying the contribution of features in the process of constructing the decision tree (<xref ref-type="bibr" rid="B18">Ke et&#xa0;al., 2017</xref>). Following this, we selected the top 300 features for modeling with XGBoost. The 5-fold cross-validation yielded ACC=0.9899, Sn=0.9890, Sp=0.9908, MCC=0.9799, and auROC=0.9994. As shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>, our cross-validation results indicate that: (1) Fine-tuned models outperformed non-fine-tuned counterparts significantly. (2) However, applying feature selection after fine-tuning caused minor performance degradation compared to models without feature selection, not much difference overall. These findings demonstrate the effectiveness of our fine-tuning strategy. The pre-training model is usually trained on multi species datasets, and may not be able to capture the 6mA distribution pattern unique to rice. Through the fine-tuning strategy, the model parameters are recalibrated, which can give priority to the local features in the rice genome, and the sensitivity of the model to the sequence context of rice 6mA is improved. Additionally, while XGBoost&#x2019;s tree-based architecture excels at managing high-dimensional data through regularization techniques, our results suggest that applying LightGBM-based feature selection after fine-tuning may slightly reduce model performance due to fewer feature interactions. We selected the second model with the best performance, performing fine-tuning for 200 iterations without feature selection, to name iRice6mA-LMXGB.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>
<bold>(A)</bold> Comparison of model performance with or without fine-tuning and with or without feature selection; <bold>(B)</bold> Average ROC curves for five-fold cross-validation of the three models. Where no fine-tuning_768 features denotes the model with no fine-tuning, 200 fine-tunings_768 features denotes the model with two hundred fine-tunings without feature selection, and 200 fine-tunings_300 features denotes the model that was fine-tuned 200 times and ranked for feature importance and the top 300 features are selected after the feature importance ranking.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1626539-g002.tif">
<alt-text content-type="machine-generated">Panel A shows a bar chart comparing performance metrics for models with different fine-tuning and feature combinations. Metrics include ACC, MCC, Sn, Sp, and auROC. Models with 200 fine-tunings perform better than those without. Panel B displays a ROC curve showing true positive rate against false positive rate for the same models, highlighting high AUC scores for models with fine-tuning. A dashed line represents random performance. An inset provides a zoomed-in view of the ROC curves near the top-left corner.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>UMAP dimensionality reduction visualization</title>
<p>In order to perform an in-depth analysis of the interpretability of the iRice6mA-LMXGB model after integrating DNABERT-2 with XGBoost, we used the UMAP (Uniform Manifold Approximation and Projection) technique. This is a nonlinear dimensionality reduction and visualization algorithm for large-scale datasets. Umap assumes that the data is distributed on a low dimensional manifold. Firstly, the probability weight is defined in the high dimensional space using the neighborhood graph to reflect the similarity between points. Then the cross entropy loss function is used to optimize the embedding in the low dimensional space to align the low dimensional similarity with the high dimensional structure. Based on graph theory and flow learning methods, it is assumed that the available data samples are uniformly distributed in the topological space and can be approximated and mapped from these finite data samples to a lower-dimensional space for visualization and analysis (<xref ref-type="bibr" rid="B34">McInnes and Healy, 2018</xref>).</p>
<p>To be more specific, we will visualize the distribution of 6mA and non-6mA by projecting each feature vector onto a 2D view using the UMAP technique. <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref> shows the arrangement of 6mA and non-6mA samples in 2D space before and after fine-tuning, and the decision boundary drawn in black by the XGBoost algorithm. Blue markers denote non-6mA samples, and orange markers denote 6mA samples. The first subplot represents the UMAP results of the original features without fine-tuning, which can be interpreted as all the sample points not showing any representative clustering. In <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3A</bold>
</xref>, poor separation indicates significant feature overlap between the 6mA sample points and the non-6mA sample points (<xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3A</bold>
</xref>), suggesting a high degree of overlap in their distributions. The second subfigure shows the results of projecting the high-dimensional feature space learned from the iRice6mA-LMXGB model into a 2D view, which shows much improved clustering, indicating a significant increase in separation and a decrease in overlap in the feature space (<xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3B</bold>
</xref>), resulting in improved performance. In summary, our approach allows for better learning of model decision boundaries. Through this visualization technique, we can more intuitively understand the impact of features on model predictions, further deepening our exploration of model interpretability.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>UMAP dimensionality reduction visualization. <bold>(A)</bold> UMAP results of the original features of the unfine-tuned model. <bold>(B)</bold> UMAP results of features learned by the iRice6mA-LMXGB model.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1626539-g003.tif">
<alt-text content-type="machine-generated">Two scatter plots labeled A and B with UMAP visualizations. Plot A shows mostly blue clusters centered with scattered points and black contour lines. Plot B features denser blue clusters near the center and dispersed orange clusters on the right, also outlined with black contours. Both axes are labeled UMAP1 and UMAP2.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Comparison of the proposed model with existing models</title>
<p>To better evaluate the performance of our model, we compare it with the following state-of-the-art methods, including MM-6mAPred (<xref ref-type="bibr" rid="B39">Pian et&#xa0;al., 2019</xref>), iDNA6mA-Rice (<xref ref-type="bibr" rid="B31">Lv et&#xa0;al., 2019</xref>), SNNRice6mA (<xref ref-type="bibr" rid="B62">Yu and Dai, 2019</xref>), iRicem6A-CNN (<xref ref-type="bibr" rid="B32">Lv et&#xa0;al., 2021</xref>), ENet-6mA (<xref ref-type="bibr" rid="B2">Abbas et&#xa0;al., 2022</xref>), Deep6mA (<xref ref-type="bibr" rid="B24">Li Z. et&#xa0;al., 2021</xref>) and SpineNet-6mA (<xref ref-type="bibr" rid="B1">Abbas et&#xa0;al., 2020</xref>). Our model is evaluated using the same five-fold cross-validation protocol on the same dataset as previous studies, employing the identical metrics: ACC, MCC, Sn, Sp, and AUC. As shown in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>, our iRice6mA-LMXGB model outperforms all previous predictors across all metrics and demonstrates more stable performance with less fluctuation in ACC, MCC, Sn, Sp, and AUC values. In ACC, MCC, Sn, and AUROC metrics, our model improves over the previous best predictor SpineNet-6mA by 5%, 11.42%, 3.42%, 6.62%, and 1.98%, respectively. Furthermore, it outperforms the previous best model, ENet-6mA, by 6.08% in Sp metric. To facilitate visualization of the comparison results, we created a box-and-whisker plot, as illustrated in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. To sum up, our iRice6mA-LMXGB model demonstrates superior performance compared to both machine learning-based and CNN/LSTM-based deep learning models for 6mA prediction in rice, showcasing its robustness as a predictive tool.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>5-fold cross-validation results of iRice6mA-LMXGB with several previous methods on the rice-Lv dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Method</th>
<th valign="top" align="left">ACC</th>
<th valign="top" align="left">MCC</th>
<th valign="top" align="left">Sn</th>
<th valign="top" align="left">Sp</th>
<th valign="top" align="left">AUROC</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">MM-6mAPred</td>
<td valign="top" align="left">0.9149</td>
<td valign="top" align="left">0.8300</td>
<td valign="top" align="left">0.9347</td>
<td valign="top" align="left">0.8951</td>
<td valign="top" align="left">0.9600</td>
</tr>
<tr>
<td valign="top" align="left">iDNA6mA-Rice</td>
<td valign="top" align="left">0.9170</td>
<td valign="top" align="left">0.8350</td>
<td valign="top" align="left">0.9300</td>
<td valign="top" align="left">0.9050</td>
<td valign="top" align="left">0.9640</td>
</tr>
<tr>
<td valign="top" align="left">SNNRice6mA</td>
<td valign="top" align="left">0.9204</td>
<td valign="top" align="left">0.8400</td>
<td valign="top" align="left">0.9433</td>
<td valign="top" align="left">0.8975</td>
<td valign="top" align="left">0.9700</td>
</tr>
<tr>
<td valign="top" align="left">iRicem6A-CNN</td>
<td valign="top" align="left">0.9382</td>
<td valign="top" align="left">0.8770</td>
<td valign="top" align="left">0.9434</td>
<td valign="top" align="left">0.9331</td>
<td valign="top" align="left">0.9790</td>
</tr>
<tr>
<td valign="top" align="left">ENet-6mA</td>
<td valign="top" align="left">0.9437</td>
<td valign="top" align="left">0.8700</td>
<td valign="top" align="left">0.9467</td>
<td valign="top" align="left">0.9339</td>
<td valign="top" align="left">0.9800</td>
</tr>
<tr>
<td valign="top" align="left">Deep6mA</td>
<td valign="top" align="left">0.9401</td>
<td valign="top" align="left">0.8800</td>
<td valign="top" align="left">0.9506</td>
<td valign="top" align="left">0.9296</td>
<td valign="top" align="left">0.9800</td>
</tr>
<tr>
<td valign="top" align="left">SpineNet-6mA</td>
<td valign="top" align="left">0.9431</td>
<td valign="top" align="left">0.8800</td>
<td valign="top" align="left">0.9571</td>
<td valign="top" align="left">0.9292</td>
<td valign="top" align="left">0.9800</td>
</tr>
<tr>
<td valign="top" align="left">iRice6mA-LMXGB (ours)</td>
<td valign="top" align="left">
<bold>0.9903</bold>
</td>
<td valign="top" align="left">
<bold>0.9805</bold>
</td>
<td valign="top" align="left">
<bold>0.9898</bold>
</td>
<td valign="top" align="left">
<bold>0.9907</bold>
</td>
<td valign="top" align="left">
<bold>0.9994</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values indicate that the model proposed in this study achieves optimal results in each of the assessment metrics.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Comparison of the proposed model with other existing models on the rice-Lv dataset.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1626539-g004.tif">
<alt-text content-type="machine-generated">Box plot comparing the performance of different models on a scale from 0.80 to 1.00. Models listed are MM-6mAPred, iDNA6mA-Rice, SNRRice6mA, iRicem6A-CNN, ENet-6mA, Deep6mA, SpineNet-6mA, and iRice6mA-LMXGB (Ours). Each box is color-coded, with iRice6mA-LMXGB showing the highest performance.</alt-text>
</graphic>
</fig>
</sec>
</sec>
<sec id="s4" sec-type="conclusions">
<label>4</label>
<title>Conclusions</title>
<p>In this article, we develop a novel computational model called iRice6mA-LMXGB that combines fine-tuned large language modeling to efficiently distinguish and identify 6mA and non-6mA loci in the rice genome. We utilized the large language model, DNABERT-2, to represent the DNA sequence as a continuous word vector, thus effectively capturing the DNA sequence features. Subsequently, we applied the robust machine learning method XGBoost to make accurate predictions based on the extracted features. We compare and analyze the performance of iRice6mA-LMXGB with other predictors, and the results show that iRice6mA-LMXGB obtains the best performance compared to previous models. Our model outperforms all existing models on ACC, SN, SP, MCC, and AUC (5-fold cross-validation: ACC = 0.9903, MCC = 0.9805, Sn = 0.9898, Sp = 0.9907, and auROC = 0.9994), suggesting that the iRice6mA-LMXGB is a powerful and robust predictor that can help researchers to identify and analyze the 6mA locus in the rice genome more effectively, thus providing a deeper understanding of the complex mechanisms of gene regulation and advancing the field of life sciences. It is demonstrated through UMAP visualization that the fine-tuning strategy for large language models significantly enhances the model&#x2019;s feature extraction ability. This raises the possibility that large language models can be fine-tuned for various purposes and deployed for plant-specific domains to solve biological problems. Moving ahead, we plan to expand our dataset and perform model optimization to enhance the generalizability of our model for broader applications.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The raw sequence data used in the study were obtained from the following URL: <uri xlink:href="http://lin-group.cn/server/iDNA6mA-Rice">http://lin-group.cn/server/iDNA6mA-Rice</uri>.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>YZ: Formal analysis, Investigation, Visualization, Writing &#x2013; original draft. HC: Investigation, Writing &#x2013; review &amp; editing. SX: Investigation, Writing &#x2013; review &amp; editing. ZL: Methodology, Project administration, Supervision, Writing &#x2013; review &amp; editing.</p>
</sec>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research and/or publication of this article. This work supports by the National Natural Science Foundation of China (No.62371318, No.62001090) and 2024 Foundation Cultivation Research Basic Research Cultivation Special Funding (No. 20826041H4211).</p>
</sec>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="ai-statement">
<title>Generative AI statement</title>
<p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abbas</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Tayara</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Chong</surname> <given-names>K. T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>SpineNet-6mA: A novel deep learning tool for predicting DNA N6-methyladenine sites in genomes</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>201450</fpage>&#x2013;<lpage>201457</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/Access.6287639</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abbas</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Tayara</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Chong</surname> <given-names>K. T.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>ENet-6mA: identification of 6mA modification sites in plant genomes using elasticNet and neural networks</article-title>. <source>Int. J. Mol. Sci.</source> <volume>23</volume>, <fpage>8314</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/ijms23158314</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Basith</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Manavalan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Shin</surname> <given-names>T. H.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>SDM6A: A web-based integrative machine-learning framework for predicting 6mA sites in the rice genome</article-title>. <source>Mol. Ther. Nucleic Acids</source> <volume>18</volume>, <fpage>131</fpage>&#x2013;<lpage>141</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.omtn.2019.08.011</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Guestrin</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>XGBoost: A scalable tree boosting system</article-title>,&#x201d; in <conf-name>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</conf-name>, <conf-loc>San Francisco, California, USA</conf-loc>. <fpage>785</fpage>&#x2013;<lpage>794</lpage> (<publisher-name>Association for Computing Machinery</publisher-name>).</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Climate-responsive DNA methylation is involved in the biosynthesis of lignin in birch</article-title>. <source>Front. Plant Sci.</source> <volume>13</volume>, <elocation-id>1090967</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2022.1090967</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Integrating machine learning and genome editing for crop improvement</article-title>. <source>aBIOTECH</source> <volume>5</volume>, <fpage>262</fpage>&#x2013;<lpage>277</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s42994-023-00133-5</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Lv</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Nie</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>i6mA-Pred: identifying DNA N6-methyladenine sites in the rice genome</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>2796</fpage>&#x2013;<lpage>2800</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btz015</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Protein language pragmatic analysis and progressive transfer learning for profiling peptide&#x2013;protein interactions</article-title>. <source>IEEE Trans. on Neural Networks and Learn. Syst</source> <volume>2025</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TNNLS.2025.3540291</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Church</surname> <given-names>K. W.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Emerging trends: A gentle introduction to fine-tuning</article-title>. <source>Nat. Lang. Eng.</source> <volume>27</volume>(<issue>6</issue>), <fpage>763</fpage>&#x2013;<lpage>778</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1017/S1351324921000322</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Devlin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>M.&#x2013;W.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Toutanova</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>. In: <conf-name>Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies</conf-name>, <volume>volume 1 (Long and Short Papers)</volume>: <fpage>4171</fpage>&#x2013;<lpage>4186</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.18653/v1/N19-1423</pub-id>
</citation>
</ref>   <ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Long</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhai</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhai</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>PlantCADB: A comprehensive plant chromatin accessibility database</article-title>. <source>Genom Proteomics Bioinf.</source> <volume>21</volume>, <fpage>311</fpage>&#x2013;<lpage>323</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.gpb.2022.10.005</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ju</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Highly accurate estimation of cell type abundance in bulk tissues based on single-cell reference and domain adaptive matching</article-title>. <source>Adv Sci.</source> <volume>11</volume>, <fpage>2306329</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/advs.202306329</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hasan</surname> <given-names>M. M.</given-names>
</name>
<name>
<surname>Basith</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Khatun</surname> <given-names>M. S.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Manavalan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Kurata</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Meta-i6mA: an interspecies predictor for identifying DNA N6-methyladenine sites of plant genomes by exploiting informative features in an integrative machine-learning framework</article-title>. <source>Brief Bioinform.</source> <volume>22</volume>, <elocation-id>bbaa202</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbaa202</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Qin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Ju</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Accurate RNA velocity estimation based on multibatch network reveals complex lineage in batch scRNA-seq data</article-title>. <source>BMC Biol.</source> <volume>22</volume>, <fpage>290</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12915-024-02085-8</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ji</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Davuluri</surname> <given-names>R. V.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome</article-title>. <source>Bioinformatics</source> <volume>37</volume>, <fpage>2112</fpage>&#x2013;<lpage>2120</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btab083</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Pang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>iDNA-ABF: multi-scale deep biological language learning model for the interpretable prediction of DNA methylations</article-title>. <source>Genome Biol.</source> <volume>23</volume>, <fpage>1</fpage>&#x2013;<lpage>23</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s13059-022-02780-1</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jumper</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Evans</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Pritzel</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Green</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Figurnov</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ronneberger</surname> <given-names>O.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Highly accurate protein structure prediction with AlphaFold</article-title>. <source>Nature</source> <volume>596</volume>, <fpage>583</fpage>&#x2013;<lpage>589</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ke</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Finley</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Ye</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>T.&#x2013;Y.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Lightgbm: A highly efficient gradient boosting decision tree</article-title>. <source>Adv. Neural Inf. Process Syst.</source> <volume>30</volume>, <fpage>3146</fpage>&#x2013;<lpage>3154</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.5555/3294996.3295074</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lai</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Deep generative models for therapeutic peptide discovery: A comprehensive review</article-title>. <source>ACM Comput. Surv</source> <volume>57</volume>, <fpage>1</fpage>&#x2013;<lpage>29</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1145/3714455</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lam</surname> <given-names>H. Y. I.</given-names>
</name>
<name>
<surname>Ong</surname> <given-names>X. E.</given-names>
</name>
<name>
<surname>Mutwil</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Large language models in plant biology</article-title>. <source>Trends Plant Sci.</source> <volume>29</volume>, <fpage>1145</fpage>&#x2013;<lpage>1155</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.tplants.2024.04.013</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Le</surname> <given-names>N. Q. K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>iN6-methylat (5-step): identifying DNA N(6)-methyladenine sites in rice genome using continuous bag of nucleobases via Chou&#x2019;s 5-step rule</article-title>. <source>Mol. Genet. Genomics</source> <volume>294</volume>, <fpage>1173</fpage>&#x2013;<lpage>1182</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00438-019-01570-y</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname> <given-names>N. K.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A comprehensive survey on genetic algorithms for DNA motif prediction</article-title>. <source>Inf. Sci.</source> <volume>466</volume>, <fpage>25</fpage>&#x2013;<lpage>43</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.ins.2018.07.004</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Godzik</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences</article-title>. <source>Bioinformatics</source> <volume>22</volume>, <fpage>1658</fpage>&#x2013;<lpage>1659</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btl158</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Kong</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Lang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Deep6mA: A deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species</article-title>. <source>PloS Comput. Biol.</source> <volume>17</volume>, <elocation-id>e1008767</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pcbi.1008767</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Pang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>BioSeq-BLM: a platform for analyzing DNA, RNA, and protein sequences based on biological language models</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <elocation-id>e129</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkab829</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>A foundation model identifies broad-spectrum antimicrobial peptides against drug-resistant bacterial infection</article-title>. <source>Nat. Commun.</source> <volume>15</volume>, <fpage>7538</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41467-024-51933-2</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Xiong</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zou</surname> <given-names>Q.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>msBERT-Promoter: a multi-scale ensemble predictor based on BERT pre-trained model for the two-stage prediction of DNA promoters and their strengths</article-title>. <source>BMC Biol.</source> <volume>22</volume>, <fpage>126</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12915-024-01923-z</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Bao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>PDLLMs: A group of tailored DNA large language models for analyzing plant genomes</article-title>. <source>Mol. Plant</source> <volume>18</volume>, <fpage>175</fpage>&#x2013;<lpage>178</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.molp.2024.12.006</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Dong</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>He</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>csDMA: an improved bioinformatics tool for identifying DNA 6&#x2009;mA modifications via Chou&#x2019;s 5-step rule</article-title>. <source>Sci. Rep.</source> <volume>9</volume>, <fpage>13109</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41598-019-49430-4</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Sequence Alignment/Map format: a comprehensive review of approaches and applications</article-title>. <source>Briefings Bioinf.</source> <volume>24</volume>, <elocation-id>bbad320</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbad320</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Dao</surname> <given-names>F. Y.</given-names>
</name>
<name>
<surname>Guan</surname> <given-names>Z. X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>J. X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>iDNA6mA-rice: A computational tool for detecting N6-methyladenine sites in rice</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <elocation-id>793</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2019.00793</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zou</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A convolutional neural network using dinucleotide one-hot encoder for identifying DNA N6-methyladenine sites in the rice genome</article-title>. <source>Neurocomputing</source> <volume>422</volume>, <fpage>214</fpage>&#x2013;<lpage>221</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.neucom.2020.09.056</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z. M.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>S. H.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>J. X.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Evaluation of different computational methods on 5-methylcytosine sites identification</article-title>. <source>Brief Bioinform.</source> <volume>21</volume>, <fpage>982</fpage>&#x2013;<lpage>995</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbz048</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McInnes</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Healy</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>UMAP: uniform manifold approximation and projection for dimension reduction</article-title>. <source>arXiv (USA)</source>, <elocation-id>abs/1802.03426</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1802.03426</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meher</surname> <given-names>P. K.</given-names>
</name>
<name>
<surname>Hati</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Sahu</surname> <given-names>T. K.</given-names>
</name>
<name>
<surname>Pradhan</surname> <given-names>U.</given-names>
</name>
<name>
<surname>Gupta</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Rath</surname> <given-names>S. N.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>SVM-root: identification of root-associated proteins in plants by employing the support vector machine with sequence-derived features</article-title>. <source>Curr. Bioinf.</source> <volume>19</volume>, <fpage>91</fpage>&#x2013;<lpage>102</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2174/1574893618666230417104543</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moeckel</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Mareboina</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Konnaris</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Chan</surname> <given-names>C. S. Y.</given-names>
</name>
<name>
<surname>Mouratidis</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Montgomery</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>A survey of k-mer methods and applications in bioinformatics</article-title>. <source>Comput. Struct. Biotechnol. J.</source> <volume>23</volume>, <fpage>2289</fpage>&#x2013;<lpage>2303</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.csbj.2024.05.025</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>O&#x2019;Brown</surname> <given-names>Z. K.</given-names>
</name>
<name>
<surname>Greer</surname> <given-names>E. L.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>N6-methyladenine: A conserved and dynamic DNA mark</article-title>. <source>Adv. Exp. Med. Biol.</source> <volume>945</volume>, <fpage>213</fpage>&#x2013;<lpage>246</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-319-43624-1_10</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Park</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wahab</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Nazari</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Ryu</surname> <given-names>J. H.</given-names>
</name>
<name>
<surname>Chong</surname> <given-names>K. T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>i6mA-DNC: Prediction of DNA N6-Methyladenosine sites in rice genome based on dinucleotide representation using deep learning</article-title>. <source>Chemom Intell Lab. Syst.</source> <volume>204</volume>, <fpage>104102</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.chemolab.2020.104102</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pian</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>MM-6mAPred: identifying DNA N6-methyladenine sites based on Markov model</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>388</fpage>&#x2013;<lpage>392</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btz556</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Press</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Smith</surname> <given-names>N. A.</given-names>
</name>
<name>
<surname>Lewis</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Train short, test long: attention with linear biases enables input length extrapolation</article-title>. <source>arXiv (USA)</source>, <elocation-id>abs/2108.12409</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2108.12409</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qiao</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Hou</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Identifying nucleotide-binding leucine-rich repeat receptor and pathogen effector pairing using transfer-learning and bilinear attention network</article-title>. <source>Bioinformatics</source> <volume>40</volume>, <elocation-id>btae581</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btae581</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rives</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Meier</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sercu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Goyal</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences</article-title>. <source>Proc. Natl. Acad. Sci. United States America</source> <volume>118</volume>, <elocation-id>e2016239118</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1073/pnas.2016239118</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Romero</surname> <given-names>F. M.</given-names>
</name>
<name>
<surname>Gatica-Arias</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>CRISPR/cas9: development and application in rice breeding</article-title>. <source>Rice Sci.</source> <volume>26</volume>, <fpage>265</fpage>&#x2013;<lpage>281</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.rsci.2019.08.001</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sennrich</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Haddow</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Birch</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Neural machine translation of rare words with subword units</article-title>. In <conf-name>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</conf-name>, <conf-loc>Berlin, Germany</conf-loc>. <publisher-name>Association for Computational Linguistics</publisher-name>. <volume>(Volume 1: Long Papers)</volume>: <fpage>1715</fpage>&#x2013;<lpage>1725</lpage>, doi:&#xa0;<pub-id pub-id-type="doi">10.18653/v1/P16-1162</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shao</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Leveraging random effects in cistrome-wide association studies for decoding the genetic determinants of prostate cancer</article-title>. <source>Adv Sci.</source> <volume>11</volume>, <fpage>2400815</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/advs.202400815</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sinha</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Dasmandal</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Yeasin</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Mishra</surname> <given-names>D. C.</given-names>
</name>
<name>
<surname>Rai</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Archak</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>EpiSemble: A novel ensemble-based machine-learning framework for prediction of DNA N6-methyladenine sites using hybrid features selection approach for crops</article-title>. <source>Curr. Bioinf.</source> <volume>18</volume>, <fpage>587</fpage>&#x2013;<lpage>597</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2174/1574893618666230316151648</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Soylu</surname> <given-names>N. N.</given-names>
</name>
<name>
<surname>Sefer</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>DeepPTM: protein post-translational modification prediction from protein sequences by combining deep protein language model with vision transformers</article-title>. <source>Curr. Bioinf.</source> <volume>19</volume>, <fpage>810</fpage>&#x2013;<lpage>824</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2174/0115748936283134240109054157</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Teng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Q.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>i6mA-vote: cross-species identification of DNA N6-methyladenine sites in plant genomes based on ensemble learning with voting</article-title>. <source>Front. Plant Sci.</source> <volume>13</volume>, <elocation-id>845835</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2022.845835</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Theodoris</surname> <given-names>C. V.</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Chopra</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Chaffin</surname> <given-names>M. D.</given-names>
</name>
<name>
<surname>Al Sayed</surname> <given-names>Z. R.</given-names>
</name>
<name>
<surname>Hill</surname> <given-names>M. C.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Transfer learning enables predictions in network biology</article-title>. <source>Nature</source> <volume>618</volume>, <fpage>616</fpage>&#x2013;<lpage>624</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41586-023-06139-9</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tiwari</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Muhammad</surname> <given-names>K.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>A deep multiple kernel learning-based higher-order fuzzy inference system for identifying DNA N4-methylcytosine sites</article-title>. <source>Inf. Sci.</source> <volume>630</volume>, <fpage>40</fpage>&#x2013;<lpage>52</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.ins.2023.01.149</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Jin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yin</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>F.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>DeepBIO: an automated and interpretable deep-learning platform for high-throughput biological sequence prediction, functional annotation and visualization analysis</article-title>. <source>Nucleic Acids Res.</source> <volume>51</volume>, <fpage>3017</fpage>&#x2013;<lpage>3029</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkad055</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Lou</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Kwok</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>EHR-HGCN: an enhanced hybrid approach for text classification using heterogeneous graph convolutional networks in electronic health records</article-title>. <source>IEEE J. Biomed. Health Inf.</source> <volume>28</volume>, <fpage>1668</fpage>&#x2013;<lpage>1679</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JBHI.2023.3346210</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhai</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zou</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>SBSM-Pro: support bio-sequence machine for proteins</article-title>. <source>Sci. China-Inf Sci.</source> <volume>67</volume>, <fpage>212106</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11432-024-4171-9</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname> <given-names>L.</given-names>
</name>
<name>
<surname>He</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Malik</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Manavalan</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Computational prediction and interpretation of cell-specific replication origin sites from multiple eukaryotes by exploiting stacking framework</article-title>. <source>Brief Bioinform.</source> <volume>22</volume>, <elocation-id>bbaa275</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbaa275</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tiwari</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Structured Sparse Regularization based Random Vector Functional Link Networks for DNA N4-methylcytosine sites prediction</article-title>. <source>Expert Syst. Appl.</source> <volume>235</volume>, <fpage>121157</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.eswa.2023.121157</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Gui</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Qiao</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Deep learning in template-free de novo biosynthetic pathway design of natural products</article-title>. <source>Brief Bioinform.</source> <volume>25</volume>, <elocation-id>bbae495</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbae495</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Methyl-GP: accurate generic DNA methylation prediction based on a language model and representation learning</article-title>. <source>Nucleic Acids Res.</source> <volume>53</volume>, <elocation-id>gkaf223</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkaf223</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>6mA-Finder: a novel online tool for predicting DNA N6-methyladenine sites in genomes</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>3257</fpage>&#x2013;<lpage>3259</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btaa113</pub-id>
</citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xue</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Qiao</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>i6mA-VC: A multi-classifier voting method for the computational identification of DNA N6-methyladenine sites</article-title>. <source>Interdiscip Sci.</source> <volume>13</volume>, <fpage>413</fpage>&#x2013;<lpage>425</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s12539-021-00429-4</pub-id>
</citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>He</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Risk Prediction of Diabetes: Big data mining with fusion of multifarious physical examination indicators</article-title>. <source>Inf. Fusion</source> <volume>75</volume>, <fpage>140</fpage>&#x2013;<lpage>149</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.inffus.2021.02.015</pub-id>
</citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>D.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Improving rice grain shape through upstream ORF editing-mediated translation regulation</article-title>. <source>Plant Physiol.</source> <volume>197</volume>, <elocation-id>kiae557</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/plphys/kiae557</pub-id>
</citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Dai</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>SNNRice6mA: A deep learning method for predicting DNA N6-methyladenine sites in rice genome</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <elocation-id>1071</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2019.01071</pub-id>
</citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>H. Q.</given-names>
</name>
<name>
<surname>Arif</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Thafar</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Albaradei</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Cai</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2025</year>). <article-title>PMPred-AE: a computational model for the detection and interpretation of pathological myopia based on artificial intelligence</article-title>. <source>Front. Med. (Lausanne)</source> <volume>12</volume>, <elocation-id>1529335</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fmed.2025.1529335</pub-id>
</citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>W.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>N6-methyladenine DNA modification in drosophila</article-title>. <source>Cell</source> <volume>161</volume>, <fpage>893</fpage>&#x2013;<lpage>906</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cell.2015.04.018</pub-id>
</citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>P.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). <article-title>N(6)-methyladenine DNA methylation in japonica and indica rice genomes and its association with gene expression, plant development, and stress responses</article-title>. <source>Mol. Plant</source> <volume>11</volume>, <fpage>1492</fpage>&#x2013;<lpage>1508</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.molp.2018.11.005</pub-id>
</citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Dutta</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Davuluri</surname> <given-names>R. V.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>DNABERT-2: efficient foundation model and benchmark for multi-species genome</article-title>. <source>arXiv (USA)</source>, <elocation-id>abs/2306.15006</elocation-id>. <uri xlink:href="https://ui.adsabs.harvard.edu/abs/2023arXiv230615006Z/abstract">https://ui.adsabs.harvard.edu/abs/2023arXiv230615006Z/abstract</uri>
</citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>DNA demethylases remodel DNA methylation in rice gametes and zygote and are required for reproduction</article-title>. <source>Mol. Plant</source> <volume>14</volume>, <fpage>1569</fpage>&#x2013;<lpage>1583</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.molp.2021.06.006</pub-id>
</citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). <article-title>Identification and analysis of adenine N6-methylation sites in the rice genome</article-title>. <source>Nat. Plants</source> <volume>4</volume>, <fpage>554</fpage>&#x2013;<lpage>563</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41477-018-0214-x</pub-id>
</citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yin</surname> <given-names>J.</given-names>
</name>
<name>
<surname>She</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>PSAC-6mA: 6mA site identifier using self-attention capsule network based on sequence-positioning</article-title>. <source>Comput. Biol. Med.</source> <volume>171</volume>, <fpage>108129</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compbiomed.2024.108129</pub-id>
</citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Beaulaurier</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Deikus</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>T. P.</given-names>
</name>
<name>
<surname>Strahl</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Hao</surname> <given-names>Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). <article-title>Mapping and characterizing N6-methyladenine in eukaryotic genomes using single-molecule real-time sequencing</article-title>. <source>Genome Res.</source> <volume>28</volume>, <fpage>1067</fpage>&#x2013;<lpage>1078</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1101/gr.231068.117</pub-id>
</citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Hao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Identification of microbe&#x2013;disease signed associations via multi-scale variational graph autoencoder based on signed message propagation</article-title>. <source>BMC Biol.</source> <volume>22</volume>, <fpage>172</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12915-024-01968-0</pub-id>
</citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>HAR_Locator: a novel protein subcellular location prediction model of immunohistochemistry images based on hybrid attention modules and residual units</article-title>. <source>Front. Mol. Biosci.</source> <volume>10</volume>, <elocation-id>1171429</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fmolb.2023.1171429</pub-id>
</citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Xing</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Gene2vec: gene subsequence embedding for prediction of mammalian N6-methyladenosine sites from mRNA</article-title>. <source>RNA</source> <volume>25</volume>, <fpage>205</fpage>&#x2013;<lpage>218</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1261/rna.069112.118</pub-id>
</citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname> <given-names>H. L.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Yin</surname> <given-names>Z. J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Integrating multiple sequence features for identifying anticancer peptides</article-title>. <source>Comput. Biol. Chem.</source> <volume>99</volume>, <fpage>7</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compbiolchem.2022.107711</pub-id>
</citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zulfiqar</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ahmad</surname> <given-names>R. M.</given-names>
</name>
<name>
<surname>Ahmed</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Cai</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Deep-STP: a deep learning-based approach to predict snake toxin proteins by using word embeddings</article-title>. <source>Front. Med. (Lausanne)</source> <volume>10</volume>, <fpage>1291352</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fmed.2023.1291352</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>