<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2025.1638520</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Attention-enhanced hybrid deep&#xa0;learning model for robust&#xa0;mango leaf&#xa0;disease classification via ConvNeXt and&#xa0;vision&#xa0;transformer fusion</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Erg&#xfc;n</surname>
<given-names>Ebru</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/3031137/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>Department of Electrical and Electronics Engineering, Faculty of Engineering and Architecture, Recep Tayyip Erdogan University</institution>, <addr-line>Rize</addr-line>,&#xa0;<country>T&#xfc;rkiye</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Vinayakumar Ravi, Prince Mohammad bin Fahd University, Saudi Arabia</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Bahar Mahmud, Western Michigan University, United States</p>
<p>Vasudha Vedula, University of Texas of the Permian Basin, United States</p>
<p>Gowtham M., JSS Science and Technology University, India</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Ebru Erg&#xfc;n, <email xlink:href="mailto:ebru.yavuz@erdogan.edu.tr">ebru.yavuz@erdogan.edu.tr</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>08</month>
<year>2025</year>
</pub-date>
<pub-date pub-type="collection">
<year>2025</year>
</pub-date>
<volume>16</volume>
<elocation-id>1638520</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>06</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>07</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2025 Erg&#xfc;n.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Erg&#xfc;n</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Mango is a crop of vital agronomic and commercial importance, particularly in tropical and subtropical regions. Accurate and timely identification of foliar diseases is essential for maintaining plant health and ensuring sustainable agricultural productivity. This study proposes MangoLeafCMDF-FAMNet (cross-modal dynamic fusion with feature attention module (FAM) network), an advanced, hybrid, deep-learning framework designed for the multi-class classification of mango leaf diseases. The model combines two state-of-the-art feature extractors, ConvNeXt and Vision Transformer, to capture local fine-grained textures and global contextual semantics simultaneously. To further improve feature discrimination, a FAM inspired by squeeze-and-excitation networks is integrated into each stage of the backbone. This module adaptively recalibrates channel-wise feature responses to highlight disease-relevant cues while suppressing irrelevant background noise. A novel cross-modal dynamic fusion strategy unifies the complementary strengths of both branches, resulting in highly robust and discriminative feature embeddings. The proposed model was rigorously evaluated using comprehensive metrics such as classification accuracy (CA), recall, precision, Matthews correlation coefficient (MCC) and Cohen&#x2019;s kappa score on three benchmark datasets: MangoLeafDataset1 (8 classes), MangoLeafDataset2 (5 classes) and MangoLeafDataset3 (8 classes). The experimental results consistently demonstrate the superiority of MangoLeafCMDF-FAMNet over the existing baseline models. It achieves exceptional CA values of 0.9978, 0.9988 and 0.9943 across the respective datasets, alongside strong MCC and Cohen&#x2019;s kappa scores. These results highlight the effectiveness and generalizability of the proposed framework for automated mango leaf disease diagnosis and contribute to advancing deep learning applications in precision plant pathology.</p>
</abstract>
<kwd-group>
<kwd>agricultural imaging</kwd>
<kwd>ConvNeXt</kwd>
<kwd>cross-modal dynamic fusion</kwd>
<kwd>disease classification</kwd>
<kwd>mango leaf</kwd>
<kwd>vision transformer</kwd>
</kwd-group>
<counts>
<fig-count count="12"/>
<table-count count="5"/>
<equation-count count="26"/>
<ref-count count="38"/>
<page-count count="20"/>
<word-count count="10359"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Plant Bioinformatics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>The mango is of great agronomic and economic importance, particularly in tropical and subtropical regions, where it is one of the most widely cultivated fruit crops (<xref ref-type="bibr" rid="B37">Zhang et&#xa0;al., 2025</xref>). However, mango plants&#x2019; productivity and health are persistently threatened by foliar diseases, which impair photosynthetic efficiency and lead to significant reductions in yield and fruit quality. Although traditional approaches to disease diagnosis are still widely used, they often involve subjective assessments, delayed response times and a heavy dependence on expert knowledge. These limitations highlight the urgent need for automated, accurate and scalable diagnostic tools to support the timely and objective management of mango diseases.</p>
<p>In recent years, deep learning (DL) techniques have transformed plant disease detection by enabling complex, hierarchical patterns to be extracted directly from image data. Unlike conventional handcrafted approaches, DL models have demonstrated superior performance in various agricultural vision tasks. However, classifying mango leaf diseases is challenging due to their intricate visual symptoms, similarities between classes and variations within classes that frequently occur across disease types (<xref ref-type="bibr" rid="B31">Shehu et&#xa0;al., 2025</xref>; <xref ref-type="bibr" rid="B35">Wei et&#xa0;al., 2025</xref>). These complexities necessitate the development of more advanced architectures that can effectively learn both fine-grained local textures and high-level semantic features.</p>
<p>In response to these challenges, we propose MangoLeafCMDF-FAMNet (cross-modal dynamic fusion (CMDF) with feature attention module (FAM) network), a novel hybrid DL framework specifically designed for multi-class mango leaf disease classification. This architecture integrates ConvNeXt and Vision Transformer (ViT) as dual feature extractors, combining the strengths of convolutional inductive biases and transformer-based global attention mechanisms. The proposed model uses a CMDF strategy to combine texture- and semantic-level information into a coherent, enriched feature representation space. To further enhance feature expressiveness, a FAM, inspired by squeeze-and-excitation (SE) networks, is incorporated at each stage. This module adaptively recalibrates channel-wise feature responses to prioritize disease-relevant patterns while suppressing irrelevant background noise.</p>
<p>To comprehensively evaluate the performance of MangoLeafCMDF-FAMNet, we conduct extensive experiments on three publicly available mango leaf disease datasets&#x2014;MangoLeafDataset1 (8 classes), MangoLeafDataset2 (5 classes), and MangoLeafDataset3 (8 classes). The model&#x2019;s effectiveness is quantified using multiple evaluation metrics, including classification accuracy (CA), recall (RCL), precision (PRC), Matthews correlation coefficient (MCC), and Cohen&#x2019;s kappa score. The experimental results consistently demonstrate that the proposed method significantly outperforms conventional baseline models across all datasets, achieving high CA and strong correlation measures.</p>    <p>The main contributions of this work are as follows:</p>
<list list-type="bullet">
<list-item>
<p>We introduce MangoLeafCMDF-FAMNet, a novel hybrid DL architecture that combines ConvNeXt and ViT with FAM for enhanced hierarchical feature representation.</p>
</list-item>
<list-item>
<p>We design a CMDF strategy that effectively fuses local texture information with global contextual features, leading to more robust representations.</p>
</list-item>
<list-item>
<p>We perform a thorough evaluation across multiple public datasets, establishing the superior classification performance of the proposed model in terms of CA, RCL, PRC, MCC, and kappa.</p>
</list-item>
<list-item>
<p>We offer a generalizable and scalable framework with practical implications for automated mango disease diagnosis, and the potential for adaptation to other plant disease classification tasks.</p>
</list-item>
</list>
<p>The remainder of this paper is structured as follows: Section 2 provides a thorough review of existing research for plant disease detection. Section 3 outlines the materials and methods employed in this study, providing detailed descriptions of the dataset and the proposed hybrid deep learning and feature selection framework. Section 4 reports the experimental results, alongside performance evaluation metrics and comparative analyses. Section 5 concludes the paper by summarizing the main findings and highlighting potential future research directions. Finally, Section 6 critically discusses the study&#x2019;s limitations and underlying assumptions, as well as its practical implications for real-world agricultural applications.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Review of existing approaches</title>
<p>Recent advances in computer vision and DL have significantly accelerated the development of automated tools for plant disease diagnosis. convolutional neural networks (CNNs) and transformer-based models, especially ViTs, have emerged as dominant paradigms in plant pathology research due to their ability to extract discriminative spatial and semantic patterns from complex visual data (<xref ref-type="bibr" rid="B4">Chen et&#xa0;al., 2024</xref>). However, existing studies on mango leaf disease classification have faced multiple challenges that limit their practical utility and generalizability. Among these studies, <xref ref-type="bibr" rid="B26">Rao et&#xa0;al. (2021)</xref> initiated one such effort by leveraging the PlantVillage dataset, applying the AlexNet architecture, and achieving classification accuracies of 0.9900 for grape leaves and 0.8900 for mango leaves. Building on this, <xref ref-type="bibr" rid="B2">Arivazhagan and Ligi (2018)</xref> utilized a CNN-based approach on a six-class mango dataset and attained a commendable accuracy of 0.9667. In contrast, <xref ref-type="bibr" rid="B18">Mia et&#xa0;al. (2020)</xref> combined artificial neural networks and support vector machines, achieving 0.8000 accuracy in detecting four disease classes and healthy leaves. A more advanced ensemble strategy was introduced by <xref ref-type="bibr" rid="B11">Gautam et&#xa0;al. (2024)</xref>, who developed a Stacked Ensemble Deep Neural Network that integrated multiple DNNs with classical ML classifiers, yielding a high accuracy of 0.9857 across eight disease classes. Similarly, <xref ref-type="bibr" rid="B29">Saleem et&#xa0;al. (2021a)</xref> investigated disease detection using canonical correlation analysis (CCA)-based feature fusion and found cubic SVM to deliver the highest performance. In a follow-up study, <xref ref-type="bibr" rid="B28">Saleem et&#xa0;al. (2021b)</xref> further proposed the FrCNet model for lesion segmentation and, after combining it with CCA feature fusion and classification via quadratic and cubic SVMs, achieved 0.9890 accuracy for binary disease-versus-healthy discrimination. Continuing the exploration of CNN variants, <xref ref-type="bibr" rid="B34">Varma et&#xa0;al. (2025)</xref> benchmarked several pretrained architectures, with InceptionV3 achieving the highest accuracy at 0.9987. Meanwhile, <xref ref-type="bibr" rid="B22">Patel et&#xa0;al. (2024)</xref> introduced a hybrid framework combining Total Variation Filter-based variational mode decomposition with DenseNet121 and VGG-19, achieving 0.9885 CA. This fusion approach notably improved feature interpretability and robustness against noise. <xref ref-type="bibr" rid="B12">Hossain et&#xa0;al. (2024)</xref> evaluated ViTs against well-established CNNs and proposed an optimized DeiT-based model, which outperformed all compared methods with a CA of 0.9975. Similarly, <xref ref-type="bibr" rid="B17">Mahmud et&#xa0;al. (2024)</xref> proposed DenseNet78, a lightweight variant of DenseNet tailored for mango leaf disease classification, reporting accuracies of 0.9947 for healthy and 0.9944 for diseased leaves. In practical implementations, <xref ref-type="bibr" rid="B24">Puranik et&#xa0;al. (2024)</xref> retrained MobileNetV3 on the MangoLeafBD dataset and embedded it within a mobile application, reaching 0.9800 accuracy and enabling real-time field diagnosis. <xref ref-type="bibr" rid="B32">Singh et&#xa0;al. (2024)</xref> adopted a transfer learning approach and proposed the DTLD model, which demonstrated strong multi-class classification performance with a peak accuracy of 0.9976 on a 4000-image dataset. Expanding on comparative model analysis, <xref ref-type="bibr" rid="B3">Bairwa et&#xa0;al. (2024)</xref> assessed multiple deep networks, finding ResNet50 to deliver the highest accuracy at 0.9912. Complementarily, <xref ref-type="bibr" rid="B23">Pratap and Kumar (2024)</xref> designed a CNN-based system incorporating transfer learning from VGG-16, GoogLeNet, MobileNet, YOLOv8, and EfficientNet, enabling effective classification of several mango diseases including Anthracnose, Gall Midge, and Powdery Mildew. Finally, <xref ref-type="bibr" rid="B21">Pahati et&#xa0;al. (2025)</xref> trained a Google Teachable Machine model on 4000 annotated images, obtaining an accuracy of 0.9960 and demonstrating high potential for democratized, user-friendly disease recognition platforms.</p>
<p>Most conventional CNN-based models, although capable of capturing local textures, fall short in modeling long-range dependencies&#x2014;a critical requirement for accurately distinguishing visually similar diseases with subtle morphological variations. Transformer-based methods, while excellent at global context modeling, often lack the inductive biases necessary for fine-grained feature localization. As such, stand-alone CNN or ViT models struggle to deliver optimal performance across varying environmental conditions and disease stages observed in agricultural settings. For example, the study by <xref ref-type="bibr" rid="B1">Alamri et&#xa0;al. (2025)</xref> proposed a dual-branch architecture combining ConvNeXt and ViT to detect mango leaf and fruit diseases separately using the MangoLeafBD and SenMangoFruitDDS datasets. Their model achieved promising accuracy levels of 99.87% and 98.40% respectively, demonstrating the value of hybrid architectures in plant disease classification. However, their method did not incorporate any explicit attention mechanism to recalibrate the feature importance across network layers. Furthermore, their architecture processed the outputs of ConvNeXt and ViT using a static fusion approach, which may limit the adaptability of feature interactions during training.</p>
<p>By contrast, our proposed MangoLeafCMDF-FAMNet framework introduces several significant improvements to the original design. Firstly, inspired by SE networks, we incorporated a FAM at each stage to dynamically recalibrate channel-wise features. This enables the model to selectively emphasize disease-relevant information and suppress background noise. Secondly, instead of using a static feature aggregation strategy, our model uses a CMDF mechanism to adaptively combine spatial and semantic cues extracted from ConvNeXt and ViT backbones. This significantly improves the model&#x2019;s representational richness and robustness.</p>
<p>Moreover, MangoLeafCMDF-FAMNet was rigorously evaluated on three distinct datasets encompassing both 5-class and 8-class classification tasks. Experimental results demonstrated that our model consistently outperforms traditional CNNs, ViTs, and hybrid baselines&#x2014;including the model by <xref ref-type="bibr" rid="B1">Alamri et&#xa0;al. (2025)</xref>&#x2014;not only in terms of CA but also across comprehensive evaluation metrics such as MCC and kappa. The superior performance of our model, particularly under multi-class, real-world conditions, underscores its potential as a scalable and generalizable solution for precision agriculture.</p>
<p>Importantly, foliar disease diagnosis remains a crucial but underexplored area in the literature, especially concerning tropical crops such as mango. Leaf diseases are often early indicators of plant stress and can significantly affect fruit development and overall yield. Therefore, developing robust, accurate, and field-deployable diagnostic systems for leaf disease identification is critical for achieving sustainable agricultural outcomes. Our contribution lies not only in achieving state-of-the-art performance but also in offering a practical architecture that balances accuracy, computational efficiency, and adaptability, setting a new benchmark in artificial intelligence (AI)-assisted mango disease diagnosis.</p>
</sec>
<sec id="s3" sec-type="materials|methods">
<label>3</label>
<title>Materials and methods</title>
<sec id="s3_1">
<label>3.1</label>
<title>Description of dataset</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>MangoLeafDataset1</title>
<p>The Mango MLD dataset, (MLD1) curated by Shakib et&#xa0;al., served as one of the primary data sources in this study (<xref ref-type="bibr" rid="B30">Shakib et&#xa0;al., 2024</xref>). This publicly available dataset was meticulously compiled through an extensive field data acquisition campaign conducted across diverse mango orchards situated in Kushtia and Dhaka, Bangladesh. The primary objective of this collection effort was to capture high-quality, representative images of both healthy and diseased mango leaves under realistic agricultural conditions, thereby ensuring the ecological validity and practical relevance of the dataset for real-world disease classification tasks.</p>
<p>A total of 6,400 images were included in the dataset, uniformly distributed across eight diagnostic categories. Seven of these classes correspond to prevalent mango leaf diseases&#x2014;Anthracnose, Bacterial Canker, Cutting Weevil, Die Back, Gall Midge, Powdery Mildew, and Sooty Mould&#x2014;while the eighth class represents healthy leaves. To mitigate class imbalance and enable unbiased model training, each category contains exactly 800 images, making this dataset structurally balanced. The images were originally captured using an iPhone SE device at a native resolution of 3024 &#xd7; 4032 pixels and subsequently downscaled to 240 &#xd7; 240 pixels in JPEG format. This resizing operation was performed to reduce memory overhead without significantly compromising visual quality or diagnostic features. Crucially, no synthetic augmentation was applied to the original images, preserving the integrity and authenticity of real-world leaf textures, color gradients, and lesion morphologies.</p>
<p>
<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref> presents representative samples from each disease class, providing visual insight into the morphological and pathological variations captured in the dataset. Meanwhile, the corresponding distribution of class frequencies is detailed in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>, where the uniformity of sample counts across categories is explicitly demonstrated.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Representative images from the MLD1 dataset, illustrating both healthy mango leaves and leaves affected by seven distinct diseases (<xref ref-type="bibr" rid="B30">Shakib et&#xa0;al., 2024</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g001.tif">
<alt-text content-type="machine-generated">MangoLeafDataset1 features images categorized into seven conditions with two images each: Anthracnose, Bacterial Canker, Cutting Weevil, Die Back, Gall Midge, Powdery Mildew, Sooty Mould, and Healthy leaves. Each category visually represents the specific condition's symptoms.</alt-text>
</graphic>
</fig>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Class-wise distribution of images in the MLD1 (<xref ref-type="bibr" rid="B30">Shakib et&#xa0;al., 2024</xref>).</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Class Name</th>
<th valign="middle" align="center">MLD1</th>
</tr>
<tr>
<th valign="middle" align="center">Number of images</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Anthracnose</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Bacterial Canker</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Cutting Weevil</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Die Back</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Gall Midge</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Powdery Mildew</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Sooty Mould</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Healthy</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Total</td>
<td valign="middle" align="center">6400</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>MangoLeafDataset2</title>
<p>As a complementary data source, the MangoLeafDataset2 (MLD2)&#x2014;compiled and published by Nirob et&#xa0;al.&#x2014;was incorporated into this study to further strengthen the reliability and generalizability of the proposed classification model (<xref ref-type="bibr" rid="B19">Nirob et&#xa0;al., 2024</xref>). This dataset offers a rich collection of high-resolution mango leaf images, originally captured between August 15 and August 29, 2023, in the mango cultivation fields of Supu Ashulia, Bangladesh. The data acquisition process was carried out under natural lighting and environmental conditions, ensuring that the captured leaf samples reflect real-world visual characteristics, including noise, background clutter, and variability in disease presentation.</p>
<p>The original dataset comprises 1,319 unique images, each with a standardized resolution of 1000 &#xd7; 1000 pixels and stored in JPEG format. The dataset encompasses five key categories representing distinct pathological states of mango leaves: Anthracnose, Die Black, Gall Midge, Powdery Mildew, and Healthy. These categories were carefully selected based on the prevalence and diagnostic importance of the corresponding diseases in commercial mango production. To overcome the inherent class imbalance, present in the original dataset and to enhance the learning capability of deep models, a comprehensive data augmentation strategy was applied. Techniques such as horizontal and vertical flipping, arbitrary rotations, scaling, and mild intensity transformations were utilized to synthetically expand the dataset. As a result, each category was normalized to contain exactly 1,000 samples, thereby yielding a final augmented dataset comprising 5,000 images. <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref> illustrates representative image samples from each of the five classes, providing a visual overview of the phenotypic diversity embedded within the dataset. <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> summarizes the distribution of both original and augmented images per class.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Selected image samples from the MLD2 dataset, reflecting variations in leaf color, texture, and shape across four disease classes and healthy leaves (<xref ref-type="bibr" rid="B19">Nirob et&#xa0;al., 2024</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g002.tif">
<alt-text content-type="machine-generated">MangoLeafDataset2 graphic displaying five categories of mango leaf conditions: Anthracnose, Die Back, Healthy, Gall Midge, and Powdery Mildew. Each category shows two sample images of affected or healthy leaves to illustrate their characteristics.</alt-text>
</graphic>
</fig>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Augmented image counts for each class in the MLD2 (<xref ref-type="bibr" rid="B19">Nirob et&#xa0;al., 2024</xref>).</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Class Name</th>
<th valign="middle" align="center">MLD2</th>
</tr>
<tr>
<th valign="middle" align="center">Number of images</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Anthracnose</td>
<td valign="middle" align="center">1000</td>
</tr>
<tr>
<td valign="middle" align="center">Die Back</td>
<td valign="middle" align="center">1000</td>
</tr>
<tr>
<td valign="middle" align="center">Gall Midge</td>
<td valign="middle" align="center">1000</td>
</tr>
<tr>
<td valign="middle" align="center">Powdery Mildew</td>
<td valign="middle" align="center">1000</td>
</tr>
<tr>
<td valign="middle" align="center">Healthy</td>
<td valign="middle" align="center">1000</td>
</tr>
<tr>
<td valign="middle" align="center">Total</td>
<td valign="middle" align="center">5000</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>MangoLeafDataset3</title>
<p>To further enhance the robustness and cross-dataset generalizability of the proposed classification framework, this study integrated the MangoLeafDataset3 (MLD3), meticulously curated by Rahman et&#xa0;al. and publicly released in November 2024 (<xref ref-type="bibr" rid="B25">Rahman et&#xa0;al., 2024</xref>). This dataset serves as a significant and diverse benchmark resource for intelligent agricultural analysis, particularly in the field of mango leaf disease recognition. The data acquisition phase was carried out over a period of 20 consecutive days, from October 15 to November 4, 2024, in two distinct agroecological regions of Bangladesh&#x2014;Kashinathpur (Pabna) and Changao (Savar, Dhaka)&#x2014;to capture a wide spectrum of environmental and disease conditions. The dataset is composed of two main subsets: 2,336 raw images captured under natural lighting conditions using mobile phone cameras, and 12,730 synthetically augmented images generated through comprehensive data enhancement techniques. These augmentations include but are not limited to affine transformations, horizontal/vertical flips, minor brightness and contrast shifts, and random cropping, all designed to introduce variability and enrich the dataset&#x2019;s learning potential without compromising biological authenticity. All images are categorized into eight distinct classes, representing seven pathological categories&#x2014;Anthracnose, Bacterial Canker, Cutting Weevil, Die Back, Gall Midge, Powdery Mildew, and Sooty Mould&#x2014;along with one Healthy class.</p>
<p>The original class distribution, prior to augmentation, is intentionally preserved to reflect natural disease occurrence rates. However, the expanded dataset introduces balance and diversity necessary for training deep neural models effectively. A comprehensive breakdown of the image count per class is provided in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>, while <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref> visually showcases representative samples from each class, highlighting inter-class visual variability and intra-class complexity.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Augmented image counts per class in the MangoLeafDataset3 (<xref ref-type="bibr" rid="B25">Rahman et&#xa0;al., 2024</xref>).</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Class Name</th>
<th valign="middle" align="center">MLD3</th>
</tr>
<tr>
<th valign="middle" align="center">Number of images</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Anthracnose</td>
<td valign="middle" align="center">1749</td>
</tr>
<tr>
<td valign="middle" align="center">Bacterial Canker</td>
<td valign="middle" align="center">2534</td>
</tr>
<tr>
<td valign="middle" align="center">Cutting Weevil</td>
<td valign="middle" align="center">1583</td>
</tr>
<tr>
<td valign="middle" align="center">Die Back</td>
<td valign="middle" align="center">1280</td>
</tr>
<tr>
<td valign="middle" align="center">Gall Midge</td>
<td valign="middle" align="center">2233</td>
</tr>
<tr>
<td valign="middle" align="center">Powdery Mildew</td>
<td valign="middle" align="center">776</td>
</tr>
<tr>
<td valign="middle" align="center">Sooty Mould</td>
<td valign="middle" align="center">1325</td>
</tr>
<tr>
<td valign="middle" align="center">Healthy</td>
<td valign="middle" align="center">1250</td>
</tr>
<tr>
<td valign="middle" align="center">Total</td>
<td valign="middle" align="center">12730</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Image samples from the MLD3 dataset highlighting intra-class variability and visual diversity, which pose additional challenges for robust disease classification (<xref ref-type="bibr" rid="B25">Rahman et&#xa0;al., 2024</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g003.tif">
<alt-text content-type="machine-generated">MangoLeafDataset3 diagram displays different mango leaf conditions with images. Conditions include Anthracnose, Bacterial Canker, Cutting Weevil, Die Back, Gall Midge, Powdery Mildew, Sooty Mould, and Healthy. Each condition has corresponding leaf images illustrating symptoms.</alt-text>
</graphic>
</fig>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Research methodology framework</title>
<p>In this study, a novel hybrid DL architecture named MangoLeafCMDF-FAMNet was proposed to address the complex problem of mango leaf disease classification. The methodology capitalized on the complementary strengths of two advanced feature extractors&#x2014;ConvNeXt and ViT&#x2014;to capture fine-grained texture details as well as global contextual dependencies inherent in leaf imagery. The overall flow of the proposed method is illustrated in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. To train and validate the proposed framework, three publicly available datasets were employed: the 8-class MLD1, 5-class MLD2, and 8-class MLD3. All image samples were preprocessed with standard normalization and resized to a uniform resolution of 224 &#xd7; 224 pixels to ensure consistency across training folds. Data augmentation techniques were deliberately excluded to assess the raw generalization power of the model.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Overview of the proposed MangoLeafCMDF-FAMNet architecture, illustrating the dual-branch design consisting of ConvNeXt and ViT, integrated with FAM and a CMDF mechanism for feature integration and disease classification.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g004.tif">
<alt-text content-type="machine-generated">Flowchart of a classification model architecture with sections: Image preprocessing resizing inputs for a ConvNeXt and ViT attention module. ConvNeXt involves convolution layers, downsampling, and global pooling. ViT processes patches through encoders and attention. Features from both modules are concatenated in a dynamic fusion layer, feeding into a classification layer. Datasets include mango leaf images for various conditions, evaluated by metrics like accuracy and precision.</alt-text>
</graphic>
</fig>
<p>The MangoLeafCMDF-FAMNet model integrated a ConvNeXt-Tiny backbone pre-trained on ImageNet as its local feature extractor. Its final classification head was removed and replaced with a global average pooling layer, followed by a custom FAM inspired by SE networks. In parallel, a lightweight ViT was employed to model long-range semantic interactions, with its outputs refined through a 1D feature-wise attention mechanism designed to amplify class-relevant representations. After extracting the deep features from both ConvNeXt and ViT branches, a CMDF strategy was applied. This strategy concatenated the learned embeddings and passed them through a projection layer, resulting in a unified 1024-dimensional representation. The fused vector was then forwarded to a fully connected classifier to generate final class predictions.</p>
<p>The performance evaluation of the proposed model was conducted using a stratified 5-fold cross-validation protocol (5-FCVP) to ensure reliable and unbiased assessment. In each fold, the dataset was partitioned into distinct training and validation subsets while preserving class distribution. During training, the model parameters were optimized using the AdamW optimizer, configured with a learning rate of 0.00005 and a weight decay coefficient of 0.0001 to promote generalization. The cross-entropy loss function served as the optimization objective, guiding the network&#x2019;s learning process. For each validation phase, a comprehensive set of evaluation metrics was computed, including CA, RCL, PRC, MCC, and kappa, to provide a multi-faceted performance analysis. Additionally, confusion matrices were generated for each fold to reveal class-specific prediction behaviors. To qualitatively investigate the separability of learned features, high-dimensional embeddings were projected into a two-dimensional space using t-distributed stochastic neighbor embedding (t-SNE), offering visual insight into the model&#x2019;s discriminative capability.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>ConvNeXt backbone architecture</title>
<p>In this study, ConvNeXt was selected as one of the core backbone networks of the proposed CMDF-Net due to its ability to effectively extract hierarchical features from input images by leveraging the design principles of both ResNet and transformer-based architectures. ConvNeXt is a convolutional neural network that modernizes the classic ResNet architecture through architectural refinements inspired by the success of ViTs, achieving a competitive balance between performance and efficiency in visual recognition tasks.</p>
<p>ConvNeXt comprises multiple stages, each containing a sequence of blocks designed to progressively capture low-level to high-level semantic features (<xref ref-type="bibr" rid="B10">Fu et&#xa0;al., 2025</xref>). Each block within ConvNeXt replaces the traditional bottleneck structure of ResNet with a streamlined stack of operations, composed of a depthwise convolution (DWConv), a layer normalization (LN), a pointwise convolution (1&#xd7;1 Conv), and a GELU activation function. Mathematically, the core block of ConvNeXt can be formulated as follows. Let <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represent the input tensor, where <italic>H</italic>, <italic>W</italic>, and <italic>C</italic> denote the height, width, and number of channels, respectively. The transformation <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> within a ConvNeXt block is defined as <xref ref-type="disp-formula" rid="eq1">Equation 1</xref> (<xref ref-type="bibr" rid="B9">Ford et&#xa0;al., 2025</xref>).</p>
<disp-formula id="eq1">
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mtext mathvariant="italic">f</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">x</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>G</mml:mi>
<mml:mi>E</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>U</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>L</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:mstyle mathsize="normal">
<mml:mi>D</mml:mi>
<mml:mi>W</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
</mml:mstyle>
<mml:mo stretchy="false">(</mml:mo>
<mml:mstyle mathsize="normal">
<mml:mi>x</mml:mi>
</mml:mstyle>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where, DWConv represents the depthwise convolutional operation with a kernel size of 7&#xd7;7, designed to capture spatial correlations within each channel independently. <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>&#x200b; and <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>&#x200b; denote pointwise (1&#xd7;1) convolution weights that project the input and output feature spaces. <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the layer normalization function, which stabilizes training and accelerates convergence. GELU stands for Gaussian error linear unit, providing smoother activation compared to ReLU.</p>
<p>In this implementation, ConvNeXt was configured with the &#x201c;ConvNeXt-Tiny&#x201d; variant to ensure a balanced trade-off between computational cost and feature extraction capability. The network was divided into four stages, where each stage contains multiple ConvNeXt blocks and concludes with a downsampling layer that reduces the spatial resolution while increasing the channel dimension (<xref ref-type="bibr" rid="B16">Lu et&#xa0;al., 2025</xref>). The channel dimensions across the stages were configured as <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>96</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>192</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>384</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>768</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, and the number of blocks per stage were <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>9</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, respectively. To further enhance the representational capacity of ConvNeXt, we integrated a FAM at the output of each stage. Inspired by the SE networks, this module adaptively recalibrates the feature maps along the channel dimension. The mechanism operates in three steps: squeeze, excitation, and reweighting. Let <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mi>U</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> in denote the output feature map of a ConvNeXt stage. The channel-wise global descriptor z <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is obtained via global average pooling given as <xref ref-type="disp-formula" rid="eq2">Equation 2</xref> (<xref ref-type="bibr" rid="B33">Tao et&#xa0;al., 2022</xref>).</p>
<disp-formula id="eq2">
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mtext mathvariant="italic">H</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext mathvariant="italic">W</mml:mtext>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="italic">i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext mathvariant="italic">H</mml:mtext>
</mml:msubsup>
</mml:mstyle>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="italic">j</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext mathvariant="italic">W</mml:mtext>
</mml:msubsup>
</mml:mstyle>
<mml:msub>
<mml:mtext mathvariant="italic">U</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">i</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">j</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Then, the excitation step applies a gating mechanism using two fully connected layers with non-linearity is shown in <xref ref-type="disp-formula" rid="eq3">Equation 3</xref>.</p>
<disp-formula id="eq3">
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">&#x3c3;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">&#x3b4;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where, <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfrac>
<mml:mi>C</mml:mi>
<mml:mi>r</mml:mi>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mi>C</mml:mi>
<mml:mi>r</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> are the weights of the fully connected layers, <inline-formula>
<mml:math display="inline" id="im12">
<mml:mi>&#x3b4;</mml:mi>
</mml:math>
</inline-formula> is the ReLU activation function, <inline-formula>
<mml:math display="inline" id="im13">
<mml:mi>&#x3c3;</mml:mi>
</mml:math>
</inline-formula> is the sigmoid activation function, <inline-formula>
<mml:math display="inline" id="im14">
<mml:mi>r</mml:mi>
</mml:math>
</inline-formula> is the reduction ratio (set to 16 in this study) controlling the bottleneck. Finally, the recalibrated feature map <inline-formula>
<mml:math display="inline" id="im15">
<mml:mover accent="true">
<mml:mi>U</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> is obtained by channel-wise multiplication is given in <xref ref-type="disp-formula" rid="eq4">Equation 4</xref>.</p>
<disp-formula id="eq4">
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">U</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">U</mml:mtext>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>This attention mechanism allows the network to selectively emphasize informative features while suppressing less useful ones, thereby boosting the model&#x2019;s ability to focus on disease-related patterns in mango leaf images. The output feature maps from all ConvNeXt stages, enhanced by their respective attention modules, are then passed to the fusion layer as shown <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Detailed schematic of the ConvNeXt backbone as implemented within the MangoLeafCMDF-FAMNet framework, showing key stages of convolutional feature extraction and attention recalibration.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g005.tif">
<alt-text content-type="machine-generated">Diagram depicting a convolutional neural network architecture. The left section shows sequential blocks: Stem Block with 2D Convolution, followed by four stages of ConvNeXt Blocks with downsampling at each stage. Each stage shows repeated blocks, increasing in dimensionality (Dm: 96 to 768). The right section details a ConvNeXt Block with depthwise convolution, layer normalization, GELU, and 1x1 convolutions, as well as a Feature Attention Module using adaptive average pooling, ReLU, and sigmoid functions.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Vision transformer backbone architecture</title>
<p>As a complementary backbone to ConvNeXt, the ViT was employed in CMDF-Net to exploit the global context modeling capabilities of self-attention mechanisms. ViT treats images as sequences of non-overlapping patches, analogous to tokens in natural language processing, and applies standard Transformer encoders to capture long-range dependencies and global feature representations, which are crucial for identifying disease patterns distributed across different regions of mango leaves (<xref ref-type="bibr" rid="B7">Erg&#xfc;n, 2025</xref>). The input image <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is first divided into a grid of <inline-formula>
<mml:math display="inline" id="im17">
<mml:mi>N</mml:mi>
</mml:math>
</inline-formula> patches of size <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> where <inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>P</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>. Each patch is flattened and projected into a D-dimensional embedding space through a linear layer as given <xref ref-type="disp-formula" rid="eq5">Equation 5</xref>.</p>
<disp-formula id="eq5">
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msubsup>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mn>0</mml:mn>
<mml:mtext mathvariant="italic">i</mml:mtext>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">E</mml:mtext>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">flatten</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>,</mml:mo>
<mml:mtext mathvariant="italic">N</mml:mtext>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where, <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the <inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:msup>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> image patch, <inline-formula>
<mml:math display="inline" id="im22">
<mml:mrow>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>E</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>P</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#xb7;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the learnable patch embedding matrix, <inline-formula>
<mml:math display="inline" id="im23">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the learnable positional embedding added to each patch token to retain spatial information.</p>
<p>In addition, a learnable classification token <inline-formula>
<mml:math display="inline" id="im24">
<mml:mrow>
<mml:msubsup>
<mml:mi>z</mml:mi>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is prepended to the patch sequence, which serves as the aggregated representation of the input image after processing through the Transformer layers (<xref ref-type="bibr" rid="B14">Kamal et&#xa0;al., 2025</xref>). The final input to the Transformer encoder is shown <xref ref-type="disp-formula" rid="eq6">Equation 6</xref>.</p>
<disp-formula id="eq6">
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>[</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mn>0</mml:mn>
<mml:mn>1</mml:mn>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mn>0</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mn>0</mml:mn>
<mml:mrow>
<mml:mtext mathvariant="italic">N</mml:mtext>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>]</mml:mo>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mtext mathvariant="italic">R</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">N</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext mathvariant="italic">D</mml:mtext>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The Transformer encoder consists of <inline-formula>
<mml:math display="inline" id="im25">
<mml:mi>L</mml:mi>
</mml:math>
</inline-formula> identical layers, each composed of a multi-head self-attention (MSA) mechanism followed by a position-wise feed-forward network (FFN). Each layer also includes residual connections and layer normalization as seen <xref ref-type="disp-formula" rid="eq7">Equations 7</xref> and <xref ref-type="disp-formula" rid="eq8">8</xref> (<xref ref-type="bibr" rid="B16">Lu et&#xa0;al., 2025</xref>).</p>
<disp-formula id="eq7">
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mtext mathvariant="italic">l</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">MSA</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">LN</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">l</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">l</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq8">
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mtext mathvariant="italic">l</mml:mtext>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">FNN</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">LN</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mtext mathvariant="italic">l</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">Z</mml:mtext>
<mml:mtext mathvariant="italic">l</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">l</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>,</mml:mo>
<mml:mtext mathvariant="italic">L</mml:mtext>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Here, the MSA operation splits the input into <inline-formula>
<mml:math display="inline" id="im26">
<mml:mi>h</mml:mi>
</mml:math>
</inline-formula> heads and performs scaled dot-product attention in parallel as given in <xref ref-type="disp-formula" rid="eq9">Equation 9</xref>.</p>
<disp-formula id="eq9">
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mtext mathvariant="italic">Attention</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">Q</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">K</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">V</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">softmax</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext mathvariant="italic">Q</mml:mtext>
<mml:msup>
<mml:mi>K</mml:mi>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
<mml:mtext mathvariant="italic">V</mml:mtext>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im27">
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mo>,</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mi>V</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> are the query, key, and value matrices computed from the input via learned linear projections, and <inline-formula>
<mml:math display="inline" id="im28">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac bevelled="true">
<mml:mi>D</mml:mi>
<mml:mi>h</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> is the dimensionality of each head. The <inline-formula>
<mml:math display="inline" id="im29">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> function transforms the similarity scores into a probability distribution as given the <xref ref-type="disp-formula" rid="eq10">Equation 10</xref>.</p>
<disp-formula id="eq10">
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mtext mathvariant="italic">softmax</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">a</mml:mtext>
<mml:mtext mathvariant="italic">i</mml:mtext>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mtext mathvariant="italic">e</mml:mtext>
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">a</mml:mtext>
<mml:mtext mathvariant="italic">i</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="italic">j</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext mathvariant="bold-italic">n</mml:mtext>
</mml:msubsup>
<mml:msup>
<mml:mtext mathvariant="italic">e</mml:mtext>
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">a</mml:mtext>
<mml:mtext mathvariant="italic">i</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mtext mathvariant="italic">n</mml:mtext>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The ViT backbone used in this study was based on the &#x201c;ViT-Base&#x201d; variant, configured with the following parameters: patch size <inline-formula>
<mml:math display="inline" id="im30">
<mml:mi>P</mml:mi>
</mml:math>
</inline-formula>=16, embedding dimension <inline-formula>
<mml:math display="inline" id="im31">
<mml:mi>D</mml:mi>
</mml:math>
</inline-formula>=768, number of transformer layers <inline-formula>
<mml:math display="inline" id="im32">
<mml:mi>L</mml:mi>
</mml:math>
</inline-formula>=12, number of attention heads <inline-formula>
<mml:math display="inline" id="im33">
<mml:mi>h</mml:mi>
</mml:math>
</inline-formula>=12, feed-forward dimension <inline-formula>
<mml:math display="inline" id="im34">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>=3072.</p>
<p>To further enrich the discriminative capability of ViT features, we introduced a FAM at the output of the transformer. This module, similar to the one used in ConvNeXt, emphasizes important channels in the output embedding of the classification token <inline-formula>
<mml:math display="inline" id="im35">
<mml:mrow>
<mml:msubsup>
<mml:mi>z</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, based on global channel context. Given the final Transformer output <inline-formula>
<mml:math display="inline" id="im36">
<mml:mrow>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, the class token vector <inline-formula>
<mml:math display="inline" id="im37">
<mml:mrow>
<mml:msubsup>
<mml:mi>z</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is passed through a SE-inspired gating mechanism as shown in <xref ref-type="disp-formula" rid="eq11">Equations 11</xref> and <xref ref-type="disp-formula" rid="eq12">12</xref> (<xref ref-type="bibr" rid="B20">Padshetty and Umashetty, 2024</xref>).</p>
<disp-formula id="eq11">
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">&#x3c3;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">&#x3b4;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msubsup>
<mml:mi>z</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq12">
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:msubsup>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mtext mathvariant="italic">L</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mo>&#xb7;</mml:mo>
<mml:msubsup>
<mml:mi>z</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>This attention-weighted representation <inline-formula>
<mml:math display="inline" id="im38">
<mml:mrow>
<mml:msubsup>
<mml:mi>z</mml:mi>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> captures the globally aggregated and recalibrated semantic information, which is later fused with the multiscale ConvNeXt features during the dynamic fusion stage of CMDF-Net as seen <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. Also, the global modeling capacity of ViT given as <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref> robust feature extraction across both local textures and global structures in diseased mango leaf images.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Architectural illustration of the ViT module employed in the proposed model, depicting patch embedding, transformer encoding, and output token generation stages.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g006.tif">
<alt-text content-type="machine-generated">Flowchart of a Vision Transformer (ViT) architecture. An input image is divided into 196 patches, which undergo linear projection and patch embedding. This is followed by a transformer encoder, repeated six times, then attention mechanisms produce a final feature vector. Detailed components include MLP, layer normalization, and MSA, with activation functions like ReLU and Sigmoid.</alt-text>
</graphic>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Dynamic feature fusion module</title>
<p>The Dynamic Feature Fusion Module (DFFM) was specifically designed to effectively integrate the complementary strengths of ConvNeXt and ViT backbones within the proposed CMDF-Net architecture. While ConvNeXt provides rich local representations through hierarchical convolutional processing, ViT contributes global contextual dependencies via self-attention mechanisms (<xref ref-type="bibr" rid="B5">Duan et&#xa0;al., 2025</xref>). However, naive concatenation or addition of features from these heterogeneous sources may result in sub-optimal representations due to mismatched semantics and scale. Therefore, DFFM aims to learn adaptive fusion weights that dynamically recalibrate and align the semantic contributions from both streams before final classification.</p>
<p>Let <inline-formula>
<mml:math display="inline" id="im39">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> denote the multiscale feature map extracted from the ConvNeXt backbone after the final FAM, and <inline-formula>
<mml:math display="inline" id="im40">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>z</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> be the ViT-encoded class token vector refined by its respective FAM. To enable a joint fusion, the vector <inline-formula>
<mml:math display="inline" id="im41">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>z</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>&#x200b; is first spatially expanded and reshaped to match the spatial dimensions of <inline-formula>
<mml:math display="inline" id="im42">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>&#x200b;, resulting in <inline-formula>
<mml:math display="inline" id="im43">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> as shown in <xref ref-type="disp-formula" rid="eq13">Equation 13</xref> where a learnable linear transformation aligns <italic>D</italic> and <italic>C</italic> (<xref ref-type="bibr" rid="B15">Li et&#xa0;al., 2025</xref>).</p>
<disp-formula id="eq13">
<label>(13)</label>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">reshape</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mtext mathvariant="italic">v</mml:mtext>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">b</mml:mtext>
<mml:mtext mathvariant="italic">v</mml:mtext>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mtext mathvariant="italic">v</mml:mtext>
</mml:msub>
<mml:mtext mathvariant="italic">&#x3f5;</mml:mtext>
<mml:msup>
<mml:mtext mathvariant="italic">R</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">C</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext mathvariant="italic">D</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Next, both feature maps are concatenated along the channel dimension to form a joint representation as <xref ref-type="disp-formula" rid="eq14">Equation 14</xref>.</p>
<disp-formula id="eq14">
<label>(14)</label>
<mml:math display="block" id="M14">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">joint</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">concat</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mtext mathvariant="italic">R</mml:mtext>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mtext mathvariant="italic">C</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mtext mathvariant="italic">H</mml:mtext>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>A gated attention mechanism is used to model the interdependencies between the ConvNeXt and ViT features. The joint feature map is passed through a squeeze operation using global average pooling, followed by a two-layer fully connected network with non-linear activations as provided by <xref ref-type="disp-formula" rid="eq15">Equation 15</xref> (<xref ref-type="bibr" rid="B5">Duan et&#xa0;al., 2025</xref>). Where <inline-formula>
<mml:math display="inline" id="im44">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mi>r</mml:mi>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im45">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mi>r</mml:mi>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im46">
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>16</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<disp-formula id="eq15">
<label>(15)</label>
<mml:math display="block" id="M15">
<mml:mrow>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">&#x3c3;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">&#x3b4;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">GAP</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">joint</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The resulting channel-wise attention vector <inline-formula>
<mml:math display="inline" id="im47">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> acts as a set of dynamic fusion weights, controlling the contribution of each channel. This vector is split into two components corresponding to the original feature sources as provided by <xref ref-type="disp-formula" rid="eq16">Equation 16</xref>. These weights are then used to recalibrate the characteristics of each branch as described in <xref ref-type="disp-formula" rid="eq17">Equation 17</xref>.</p>
<disp-formula id="eq16">
<label>(16)</label>
<mml:math display="block" id="M16">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>C</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mo stretchy="false">[</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq17">
<label>(17)</label>
<mml:math display="block" id="M17">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2299;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">s</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2299;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Finally, the recalibrated features are fused via element-wise summation to obtain the final feature representation as described in <xref ref-type="disp-formula" rid="eq18">Equation 18</xref>.</p>
<disp-formula id="eq18">
<label>(18)</label>
<mml:math display="block" id="M18">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">fused</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mtext mathvariant="italic">covn</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">F</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mtext mathvariant="italic">vit</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The fused feature map <inline-formula>
<mml:math display="inline" id="im48">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is passed through a global average pooling (GAP) layer, followed by a final fully connected classification head to predict the disease class label. This dynamic and learnable fusion strategy enables CMDF-Net to adaptively emphasize the most informative modalities depending on the content of each input image. The gating mechanism ensures that disease-specific patterns, whether localized or distributed globally, are optimally weighted, thereby enhancing the robustness and accuracy of the classification process.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Training strategy and evaluation protocol</title>
<p>The training strategy of the proposed MangoLeafCMDF-FAMNet model was meticulously designed to ensure stable convergence, optimal generalization, and fair performance evaluation across all experimental scenarios. All experiments were conducted using the PyTorch DL framework, ensuring efficient handling of high-dimensional image data and deep architectural components. Prior to training, all mango leaf images were resized to a spatial resolution of <inline-formula>
<mml:math display="inline" id="im49">
<mml:mrow>
<mml:mn>224</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>224</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> pixels to ensure compatibility with the input dimensions of both ConvNeXt and ViT backbones. The ConvNeXt and ViT modules within CMDF-Net were initialized with pretrained weights from ImageNet-1K to leverage generic image feature representations. All additional layers&#x2014;including FAM, DFFM, and the final classification head&#x2014;were initialized using Kaiming He initialization for ReLU-based layers and Xavier initialization for linear projections, ensuring stable weight distribution at the start of training.</p>
<p>The model was trained using the AdamW optimizer, which combines adaptive gradient updates with decoupled weight decay regularization. The initial learning rate was set to <inline-formula>
<mml:math display="inline" id="im50">
<mml:mrow>
<mml:mn>5</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> with a cosine annealing scheduler to facilitate smooth convergence. A warm-up phase of 10 epochs was employed, during which the learning rate was linearly increased from <inline-formula>
<mml:math display="inline" id="im51">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. The weight decay was fixed at <inline-formula>
<mml:math display="inline" id="im52">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, and a mini-batch size of 32 was used throughout the training. The cross-entropy loss function was utilized to compute the classification loss as described in <xref ref-type="disp-formula" rid="eq19">Equation 19</xref> (<xref ref-type="bibr" rid="B38">Zhou et&#xa0;al., 2019</xref>). Where <inline-formula>
<mml:math display="inline" id="im53">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the ground-truth label and <inline-formula>
<mml:math display="inline" id="im54">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is the softmax probability of the predicted class for the <inline-formula>
<mml:math display="inline" id="im55">
<mml:mrow>
<mml:msup>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> sample.</p>
<disp-formula id="eq19">
<label>(19)</label>
<mml:math display="block" id="M19">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">L</mml:mtext>
<mml:mrow>
<mml:mtext mathvariant="italic">CE</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="italic">i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext mathvariant="italic">N</mml:mtext>
</mml:munderover>
</mml:mstyle>
<mml:msub>
<mml:mtext mathvariant="italic">y</mml:mtext>
<mml:mtext mathvariant="italic">i</mml:mtext>
</mml:msub>
<mml:mtext mathvariant="italic">log</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">y</mml:mtext>
<mml:mtext mathvariant="italic">i</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">^</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Each model was trained for a maximum of 100 epochs. However, early stopping with a patience value of 15 epochs was employed based on the validation loss to prevent overfitting and unnecessary computations. To ensure reliable and unbiased performance evaluation, 5-FCVP was performed. In each fold, the dataset was split into training (80%) and validation (20%) sets, maintaining class distribution. The average of all five folds was reported for each evaluation metric. To comprehensively evaluate the effectiveness of the proposed method, multiple performance metrics were employed: CA, RCL, PRC, MCC, and <inline-formula>
<mml:math display="inline" id="im56">
<mml:mi>&#x3ba;</mml:mi>
</mml:math>
</inline-formula>. These metrics collectively provide insights into the model&#x2019;s overall predictive power, class-wise sensitivity, balance, and inter-rater agreement, respectively. CA indicates the proportion of correctly classified samples out of the total number of instances. It is calculated as described in <xref ref-type="disp-formula" rid="eq20">Equation 20</xref> (<xref ref-type="bibr" rid="B36">Yavuz and Aydemir, 2016</xref>).</p>
<disp-formula id="eq20">
<label>(20)</label>
<mml:math display="block" id="M20">
<mml:mrow>
<mml:mtext mathvariant="italic">CA</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">TN</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">TN</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">FP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">FN</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im57">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im58">
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im59">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula>
<mml:math display="inline" id="im60">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> denote the true positives, true negatives, false positives, and false negatives, respectively. RCL, also known as sensitivity or true positive rate, measures the ability of the model to correctly identify positive instances as explained in <xref ref-type="disp-formula" rid="eq21">Equation 21</xref>. PRC reflects the proportion of true positive predictions among all positive predictions made by the model as shown in <xref ref-type="disp-formula" rid="eq22">Equation 22</xref> (<xref ref-type="bibr" rid="B6">Erg&#xfc;n, 2024</xref>).</p>
<disp-formula id="eq21">
<label>(21)</label>
<mml:math display="block" id="M21">
<mml:mrow>
<mml:mtext mathvariant="italic">RCL</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">TN</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq22">
<label>(22)</label>
<mml:math display="block" id="M22">
<mml:mrow>
<mml:mtext mathvariant="italic">PRC</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">TPP</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>MCC is a robust measure that takes into account all four elements of the confusion matrix and is especially valuable for imbalanced datasets as described in <xref ref-type="disp-formula" rid="eq23">Equation 23</xref> (<xref ref-type="bibr" rid="B27">Rozenfeld et&#xa0;al., 2024</xref>). It returns a value between <inline-formula>
<mml:math display="inline" id="im61">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> and 1, where 1 indicates perfect prediction, 0 means no better than random guessing, and <inline-formula>
<mml:math display="inline" id="im62">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> represents total disagreement.</p>
<disp-formula id="eq23">
<label>(23)</label>
<mml:math display="block" id="M23">
<mml:mrow>
<mml:mtext mathvariant="italic">MCC</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">TN</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext mathvariant="italic">FP</mml:mtext>
<mml:mo>&#xb7;</mml:mo>
<mml:mtext mathvariant="italic">FN</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">FP</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">TP</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">FN</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">TN</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">FP</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext mathvariant="italic">TN</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext mathvariant="italic">FN</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Kappa evaluates the agreement between predicted and actual classifications, adjusted for chance. It is defined as given in the <xref ref-type="disp-formula" rid="eq24">Equation 24</xref> (<xref ref-type="bibr" rid="B8">Erg&#xfc;n and Aydemir, 2020</xref>). Where <inline-formula>
<mml:math display="inline" id="im63">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the observed agreement and <inline-formula>
<mml:math display="inline" id="im64">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the expected agreement by random chance.</p>
<disp-formula id="eq24">
<label>(24)</label>
<mml:math display="block" id="M24">
<mml:mrow>
<mml:mtext mathvariant="italic">&#x3ba;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="italic">p</mml:mtext>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">p</mml:mtext>
<mml:mtext mathvariant="italic">e</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">p</mml:mtext>
<mml:mtext mathvariant="italic">e</mml:mtext>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Classification head</title>
<p>The classification head module serves as the terminal decision-making component of the CMDF-Net architecture, synthesizing the high-level, semantically rich features obtained from the dynamically fused ConvNeXt and ViT representations. Its primary objective is to project the fused feature map into a low-dimensional space corresponding to the number of disease categories and produce the final class probabilities through a softmax activation function. We denote the output feature tensor generated by the DFF module as <inline-formula>
<mml:math display="inline" id="im65">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula>
<mml:math display="inline" id="im66">
<mml:mi>H</mml:mi>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im67">
<mml:mi>W</mml:mi>
</mml:math>
</inline-formula>, and <inline-formula>
<mml:math display="inline" id="im68">
<mml:mi>C</mml:mi>
</mml:math>
</inline-formula> represent the spatial height, width, and number of channels of the fused feature map, respectively. Before classification, global spatial information is condensed using a GAP operation as described in <xref ref-type="disp-formula" rid="eq25">Equation 25</xref> (<xref ref-type="bibr" rid="B13">Hsiao et&#xa0;al., 2019</xref>).</p>
<disp-formula id="eq25">
<label>(25)</label>
<mml:math display="block" id="M25">
<mml:mrow>
<mml:mstyle mathsize="normal">
<mml:mi>z</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:mstyle mathsize="normal">
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mstyle mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mstyle mathsize="normal">
<mml:mi>f</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mstyle mathsize="normal">
<mml:mi>R</mml:mi>
</mml:mstyle>
<mml:mstyle mathsize="normal">
<mml:mi>C</mml:mi>
</mml:mstyle>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>This operation ensures translational invariance and reduces the number of trainable parameters by eliminating the need for fully connected layers at the spatial level. The pooled vector <inline-formula>
<mml:math display="inline" id="im69">
<mml:mi>z</mml:mi>
</mml:math>
</inline-formula> is then passed through a fully connected (FC) layer followed by a softmax function to obtain the final class probabilities as explained in <xref ref-type="disp-formula" rid="eq26">Equation 26</xref>.</p>
<disp-formula id="eq26">
<label>(26)</label>
<mml:math display="block" id="M26">
<mml:mrow>
<mml:mover accent="true">
<mml:mtext mathvariant="italic">y</mml:mtext>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mtext mathvariant="italic">softmax</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">W</mml:mtext>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:msub>
<mml:mtext mathvariant="italic">z</mml:mtext>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext mathvariant="italic">b</mml:mtext>
<mml:mtext mathvariant="italic">c</mml:mtext>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im70">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im71">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> denote the weight matrix and bias vector of the classification layer, and <inline-formula>
<mml:math display="inline" id="im72">
<mml:mi>K</mml:mi>
</mml:math>
</inline-formula> is the number of target classes. To enhance the expressiveness of the classification head while maintaining generalization, dropout regularization with a rate of 0.3 was employed prior to the final linear layer. This stochastic regularization strategy helps mitigate overfitting by randomly deactivating neurons during training.</p>
</sec>
</sec>
<sec id="s4" sec-type="results">
<label>4</label>
<title>Results</title>
<p>In this study, the proposed MangoLeafCMDF-FAMNet architecture was rigorously evaluated using a 5-FCVP across three publicly available datasets: MLD1, MLD2, and MLD3. The training phase employed the AdamW optimizer with an initial learning rate set to 0.00005 and a weight decay of 0.0001. Cross-entropy loss was used as the objective function to guide the optimization process. Model performance was assessed comprehensively using multiple evaluation metrics, namely CA, RCL, PRC, MCC, and kappa. In addition, confusion matrices and high-dimensional feature distributions visualized through t-SNE were generated to provide further insights into the discriminative capability of the model. <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref> summarizes the performance metrics obtained for each fold across all three datasets. For MLD1, the model achieved remarkably high performance, consistently exceeding 99.00% CA across all folds. Specifically, Fold 3 yielded a perfect CA, RCL, and PRC of 1.0000, with corresponding MCC and kappa of 1.0000, indicating flawless classification without any mispredictions. Even in the comparatively lower-performing Fold 4, MangoLeafCMDF-FAMNet still maintained an outstanding CA of 0.9938, demonstrating its robustness against potential variability in the data splits. Similarly, for MLD2, the model maintained exceptional performance. Perfect scores were achieved in Folds 3 and 5, mirroring the trends observed in MLD1. Notably, the lowest CA across all folds was 0.9970, which still reflects a near-perfect classification capability. The consistently high MCC and kappa across folds further underline the model&#x2019;s strong agreement between the predicted and true class labels, confirming its reliability. On MLD3, which is inherently more challenging due to greater symptom variability and inter-class similarity, MangoLeafCMDF-FAMNet continued to demonstrate excellent performance. The CA values across the five folds ranged from 0.9918 to 0.9965, with the highest score achieved in Fold 5. RCL and PRC closely mirrored the trends of CA, and the high MCC and kappa reaffirmed the model&#x2019;s ability to generalize well even under more complex conditions.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Classification results of MangoLeafCMDF-FAMNet on the MLD1, MLD2, and MLD3 across 5- FCVP.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Datasets</th>
<th valign="middle" rowspan="2" align="center">Fold</th>
<th valign="middle" colspan="5" align="center">Metrics</th>
</tr>
<tr>
<th valign="middle" align="center">CA</th>
<th valign="middle" align="center">RCL</th>
<th valign="middle" align="center">PRC</th>
<th valign="middle" align="center">MCC</th>
<th valign="middle" align="center">Kappa</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="5" align="center">MLD1</td>
<td valign="middle" align="center">Fold 1</td>
<td valign="middle" align="center">0.9992</td>
<td valign="middle" align="center">0.9992</td>
<td valign="middle" align="center">0.9992</td>
<td valign="middle" align="center">0.9991</td>
<td valign="middle" align="center">0.9991</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 2</td>
<td valign="middle" align="center">0.9977</td>
<td valign="middle" align="center">0.9976</td>
<td valign="middle" align="center">0.9978</td>
<td valign="middle" align="center">0.9973</td>
<td valign="middle" align="center">0.9973</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 3</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 4</td>
<td valign="middle" align="center">0.9938</td>
<td valign="middle" align="center">0.9937</td>
<td valign="middle" align="center">0.9937</td>
<td valign="middle" align="center">0.9929</td>
<td valign="middle" align="center">0.9929</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 5</td>
<td valign="middle" align="center">0.9984</td>
<td valign="middle" align="center">0.9985</td>
<td valign="middle" align="center">0.9985</td>
<td valign="middle" align="center">0.9982</td>
<td valign="middle" align="center">0.9982</td>
</tr>
<tr>
<td valign="middle" rowspan="5" align="center">MLD2</td>
<td valign="middle" align="center">Fold 1</td>
<td valign="middle" align="center">0.9980</td>
<td valign="middle" align="center">0.9979</td>
<td valign="middle" align="center">0.9980</td>
<td valign="middle" align="center">0.9975</td>
<td valign="middle" align="center">0.9975</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 2</td>
<td valign="middle" align="center">0.9970</td>
<td valign="middle" align="center">0.9969</td>
<td valign="middle" align="center">0.9972</td>
<td valign="middle" align="center">0.9963</td>
<td valign="middle" align="center">0.9962</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 3</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 4</td>
<td valign="middle" align="center">0.9990</td>
<td valign="middle" align="center">0.9991</td>
<td valign="middle" align="center">0.9989</td>
<td valign="middle" align="center">0.9988</td>
<td valign="middle" align="center">0.9987</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 5</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
<td valign="middle" align="center">1.0000</td>
</tr>
<tr>
<td valign="middle" rowspan="5" align="center">MLD3</td>
<td valign="middle" align="center">Fold 1</td>
<td valign="middle" align="center">0.9949</td>
<td valign="middle" align="center">0.9956</td>
<td valign="middle" align="center">0.9960</td>
<td valign="middle" align="center">0.9941</td>
<td valign="middle" align="center">0.9941</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 2</td>
<td valign="middle" align="center">0.9918</td>
<td valign="middle" align="center">0.9932</td>
<td valign="middle" align="center">0.9923</td>
<td valign="middle" align="center">0.9905</td>
<td valign="middle" align="center">0.9904</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 3</td>
<td valign="middle" align="center">0.9937</td>
<td valign="middle" align="center">0.9945</td>
<td valign="middle" align="center">0.9941</td>
<td valign="middle" align="center">0.9927</td>
<td valign="middle" align="center">0.9927</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 4</td>
<td valign="middle" align="center">0.9945</td>
<td valign="middle" align="center">0.9953</td>
<td valign="middle" align="center">0.9950</td>
<td valign="middle" align="center">0.9936</td>
<td valign="middle" align="center">0.9936</td>
</tr>
<tr>
<td valign="middle" align="center">Fold 5</td>
<td valign="middle" align="center">0.9965</td>
<td valign="middle" align="center">0.9970</td>
<td valign="middle" align="center">0.9967</td>
<td valign="middle" align="center">0.9959</td>
<td valign="middle" align="center">0.9959</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Following, a detailed class-wise performance analysis was conducted to further assess the robustness and generalization ability of MangoLeafCMDF-FAMNet. Specifically, RCL and PRC were calculated for each class across all folds on the MLD1, MLD2, and MLD3 datasets. For the MLD1 dataset, the model demonstrated outstanding classification capabilities. As shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref>, the average RCL values were 0.9988 for Anthracnose, 0.9988 for Bacterial Canker, 1.0000 for Cutting Weevil, 0.9987 for Die Back, 0.9951 for Gall Midge, 0.9987 for Healthy leaves, 0.9961 for Powdery Mildew, and 0.9962 for Sooty Mould. In terms of PRC, the averages were equally high, reaching 1.0000 for several classes, with minor reductions to 0.9962 and 0.9935 for Powdery Mildew and Sooty Mould, respectively. These results confirmed that the model could accurately distinguish subtle disease symptoms even under slight class imbalance or symptom similarity.</p>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>Radar plots visualizing class-wise RCL and PRC metrics for MangoLeafCMDF-FAMNet across all five folds of the 5-FCVP on MLD1, MLD2, and MLD3 datasets. Each fold is color-coded, with the green line representing the average across folds <bold>(a)</bold> RCL for MLD1, <bold>(b)</bold> PRC for MLD1, <bold>(c)</bold> RCL for MLD2, <bold>(d)</bold> PRC for MLD2, <bold>(e)</bold> RCL for MLD3, and <bold>(f)</bold> PRC for MLD3.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g007.tif">
<alt-text content-type="machine-generated">Six radar charts labeled a) to f) display data on plant conditions: Anthracnose, Bacterial Canker, Cutting Weevil, Die Back, Gall Midge, Healthy, Powdery Mildew, and Sooty Mould. Each chart shows variable scores with different colored lines.</alt-text>
</graphic>
</fig>
<p>Similarly, in the MLD2 dataset, the model achieved nearly perfect classification. The average RCL scores were recorded at 0.9991 for Anthracnose, 1.0000 for Die Back, 0.9990 for Gall Midge, 1.0000 for Healthy leaves, and 0.9958 for Powdery Mildew. Correspondingly, the PRC values were consistently excellent, with an average exceeding 0.9990 for all categories. Notably, the model maintained strong performance even in folds where small fluctuations in Powdery Mildew recognition were observed, demonstrating resilience against minor dataset variations.</p>
<p>The analysis of the MLD3 dataset, which is inherently more challenging due to the larger number of disease classes, also revealed robust model performance. The average RCL values across folds were 0.9869 for Anthracnose, 0.9949 for Bacterial Canker, 0.9910 for Cutting Weevil, 0.9992 for Die Back, 0.9909 for Gall Midge, and 1.0000 for Healthy leaves, Powdery Mildew, and Sooty Mould. PRC values aligned closely, maintaining averages above 0.9850 for all categories. While slight drops were noted in classes such as Anthracnose and Sooty Mould, the model overall preserved an exceptional balance between sensitivity and specificity. To visually capture these findings, RCL and PRC radar plots were generated across all folds, as depicted in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref>. In these visualizations, Fold 1 is represented in dark blue, Fold 2 in orange, Fold 3 in gray, Fold 4 in yellow, Fold 5 in dark navy, and the average of all folds is plotted in green.</p>
<p>Furthermore, to comprehensively evaluate the class-level robustness of MangoLeafCMDF-FAMNet, class-wise MCC and Kappa were calculated and illustrated in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8</bold>
</xref> for the MLD1 dataset across the 5-FCVP. Specifically, MCC scores remained extremely high for all disease categories, often reaching perfect agreement across most folds. Minor variations were observed only in the Gall Midge and Sooty Mould classes, where the MCC values slightly dropped but still remained above 0.98, demonstrating the strong generalization capacity of the model without signs of overfitting. Similarly, the Kappa mirrored the MCC trends, with near-perfect agreement across all folds and classes. On average, both MCC and Kappa exceeded 0.99 for nearly every class, highlighting the model&#x2019;s consistent ability to correctly classify diverse disease symptoms under varying validation conditions.</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>Per-class MCC and Kappa metrics achieved by MangoLeafCMDF-FAMNet for the MLD1 dataset, evaluated over the 5-FCVP. Colored lines indicate individual folds, while the green line shows the average performance <bold>(a)</bold> MCC for MLD1 and <bold>(b)</bold> Kappa for MLD1.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g008.tif">
<alt-text content-type="machine-generated">Two graphs compare data related to various plant conditions: Anthracnose, Bacterial Canker, Cutting Weevil, Die Back, Gall Midge, Healthy, Powdery Mildew, and Sooty Mould. Graph (a) is a clustered bar chart with several color-coded bars representing different metrics or categories. Graph (b) is a line chart showing the trend of similar metrics across the same conditions, with lines in different colors illustrating variations among data points. Both graphs depict values ranging from 0.97 to 1.00.</alt-text>
</graphic>
</fig>
<p>In addition, we further computed the MCC and Kappa for the MLD2, as visualized in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9</bold>
</xref>. As illustrated, the MCC scores for all classes consistently achieved near-perfect values, with almost every fold yielding scores of 1.0000 across the Anthracnose, Die Back, Healthy, and Powdery Mildew categories. A slight deviation was observed in the Gall Midge class during Fold 1 and Fold 2, where the MCC values dropped marginally but still remained exceedingly high, thereby underscoring the model&#x2019;s remarkable stability even in the presence of subtle intra-class variations. Similarly, the Kappa mirrored these trends, maintaining values close to 1.0000 across all classes and folds, reaffirming the excellent agreement between predicted and true labels. The minimal variability observed in the Gall Midge class reflects realistic complexities inherent in agricultural imaging datasets, yet the exceedingly high average scores across all classes strongly indicate that MangoLeafCMDF-FAMNet successfully mitigates overfitting and maintains robust generalization.</p>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>Class-wise analysis of <bold>(a)</bold> MCC and <bold>(b)</bold> Kappa obtained from MangoLeafCMDF-FAMNet on the MLD2 dataset under 5-FCVP. Individual folds are color-coded; the green line indicates the averaged result.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g009.tif">
<alt-text content-type="machine-generated">Two charts labeled a and b compare different conditions: Anthracnose, Die Back, Gall Midge, Healthy, and Powdery Mildew. Chart a) is a grouped bar chart showing measurements around 0.98 to 1.00. Chart b) is a line graph following similar trends, with some variation for each condition.</alt-text>
</graphic>
</fig>
<p>Furthermore, to ensure a comprehensive performance analysis, we computed the class-wise MCC and Kappa of the MangoLeafCMDF-FAMNet on the MLD3, as presented in <xref ref-type="fig" rid="f10">
<bold>Figure&#xa0;10</bold>
</xref>. The results demonstrated that the proposed model consistently achieved exceptionally high MCC values across all classes and folds. Specifically, MCC scores for classes such as Die Back, Healthy, Powdery Mildew, and Sooty Mould remained at or extremely close to 1.0000 across all folds. Although slight variations were noted in classes like Anthracnose and Bacterial Canker, the MCC scores still hovered around 0.98&#x2013;0.99, reflecting highly reliable performance even in more challenging classes. Similarly, the Kappa closely followed the MCC trends, indicating outstanding agreement between predicted and ground truth labels across all folds and classes.</p>
<fig id="f10" position="float">
<label>Figure&#xa0;10</label>
<caption>
<p>Visualization of per-class <bold>(a)</bold> MCC and <bold>(b)</bold> Kappa for MangoLeafCMDF-FAMNet on the MLD3 dataset, based on 5-FCVP. Fold-wise trends and average values are presented to demonstrate performance consistency.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g010.tif">
<alt-text content-type="machine-generated">Two graphs displaying data on plant health treatments. Graph (a) is a grouped bar chart showing treatment effectiveness across various conditions like Anthracnose and Powdery Mildew. Graph (b) is a line chart displaying similar data trends, highlighting variations in effectiveness for different conditions. Each category is color-coded across both graphs.</alt-text>
</graphic>
</fig>
<p>Furthermore, the feature distributions fused by MangoLeafCMDF-FAMNet were qualitatively analyzed using t-SNE, as illustrated in <xref ref-type="fig" rid="f11">
<bold>Figure&#xa0;11</bold>
</xref> for each dataset individually. t-SNE serves as a powerful non-linear dimensionality reduction technique that projects high-dimensional feature representations into a two-dimensional space, facilitating the visualization of complex feature relationships. The t-SNE plots reveal that features extracted by MangoLeafCMDF-FAMNet exhibit clear, compact, and well-separated clusters for different disease classes across all three datasets. This visual evidence supports the numerical findings, indicating that the model effectively learns discriminative and disease-specific representations without significant overlap among categories. Moreover, the distinct cluster formations further demonstrate the absence of overfitting, affirming the model&#x2019;s strong generalization capability to unseen samples.</p>
<fig id="f11" position="float">
<label>Figure&#xa0;11</label>
<caption>
<p>Two-dimensional t-SNE plots of the fused feature embeddings produced by MangoLeafCMDF-FAMNet across <bold>(a)</bold> MLD1, <bold>(b)</bold> MLD2, and <bold>(c)</bold> MLD3 datasets. The visualization illustrates the discriminative capacity and separability of learned features among disease classes.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g011.tif">
<alt-text content-type="machine-generated">Three scatter plots labeled a, b, and c show clustered data points in various colors, representing different categories like Anthracnose, Bacterial Canker, and others. Each plot has a different arrangement of clusters, with legends indicating category colors.</alt-text>
</graphic>
</fig>
<p>The average confusion matrices for the MLD1, MLD2, and MLD3 datasets, computed across the 5-FCVP and illustrated in <xref ref-type="fig" rid="f12">
<bold>Figure&#xa0;12</bold>
</xref>, provide critical insights into the class-specific discriminative capabilities of MangoLeafCMDF-FAMNet. These matrices (vertical axis: true labels; horizontal axis: predicted labels) reveal near-perfect diagonal dominance, underscoring the model&#x2019;s ability to minimize misclassifications while maintaining high intra-class consistency. For MLD1, the matrix demonstrated exceptional precision, with Anthracnose and Bacterial Canker&#x2014;classes often confused due to overlapping lesion patterns&#x2014;achieving 159.80 correct predictions, respectively. Only minor off-diagonal errors were observed: 0.20 of Anthracnose samples were misclassified as Sooty Mould, while 0.20 of Bacterial Canker cases were incorrectly assigned to Healthy. The Cutting Weevil class exhibited flawless performance, with all 160 samples correctly identified. Similarly, Die Back and Gall Midge achieved near-perfect classification, with diagonal values of 159.80 and 159.40, respectively.</p>
<fig id="f12" position="float">
<label>Figure&#xa0;12</label>
<caption>
<p>Averaged confusion matrices of MangoLeafCMDF-FAMNet for the <bold>(a)</bold> MLD1, <bold>(b)</bold> MLD2, and <bold>(c)</bold> MLD3 datasets, computed over 5-FCVP. The vertical axis denotes the actual class labels, while the horizontal axis shows the predicted class labels.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-16-1638520-g012.tif">
<alt-text content-type="machine-generated">Three heatmaps labeled a, b, and c show numerical data for various plant diseases and health conditions. Each has a color gradient from light blue to dark blue indicating lower to higher values. Chart a) shows values up to 159.80, b) up to 200.00, and c) up to 504.20. Different conditions like &#x201c;Anthracnose,&#x201d; &#x201c;Bacterial Canker,&#x201d; and &#x201c;Healthy&#x201d; are compared.</alt-text>
</graphic>
</fig>
<p>In MLD2, the matrix highlighted robust performance under increased symptom variability. Gall Midge, a class with subtle morphological features, achieved 199.80 correct predictions, with only 0.20 confusion with Healthy leaves. Die Back and Healthy were resolved with high precision, as evidenced by diagonal entries of 200.00. The most notable misclassification occurred in the Powdery Mildew class, where 0.60 samples were erroneously predicted as Gall Midge, likely due to shared textural patterns in late-stage infections.</p>
<p>For MLD3, the matrix further validated the model&#x2019;s generalization capability. The Die Back class, often prone to misclassification due to its overlapping early-stage symptoms with Gall Midge, achieved a strong diagonal value of 255.80 correct predictions, reflecting the model&#x2019;s ability to discern subtle differences in lesion distribution. Powdery Mildew, a class with visually ambiguous fungal patterns, was correctly classified in 154.80 instances, with only 0.20 samples misassigned to Die Back&#x2014;a negligible error likely attributable to shared textural features in advanced infection stages. Gall Midge, despite its complex morphological variations across growth cycles, demonstrated exceptional performance with 442.60 accurate predictions. A minimal leakage of 3.80 samples to Sooty Mould was observed, potentially stemming from similarities in necrotic patterning under low-light imaging conditions. The Healthy class once again exhibited flawless discriminative capability, with all 250.00 samples correctly identified, underscoring the model&#x2019;s precision in isolating disease-specific features from healthy tissue. Sooty Mould, a class frequently confused with Powdery Mildew in conventional methods, achieved a near-perfect diagonal score of 264.80, further highlighting the architecture&#x2019;s proficiency in resolving spectral ambiguities. These results collectively affirm that MangoLeafCMDF-FAMNet generalizes robustly across diverse data distributions, making it particularly suitable for real-world agricultural applications where symptom variability and class overlap are prevalent.</p>
<p>In order to robustly validate the effectiveness of the proposed MangoLeafCMDF-FAMNet, an extensive comparative analysis was&#xa0;conducted against prominent baseline models, including MangoLeafCMDF-Net, ViT, and ConvNeXt, across the MLD1,&#xa0;MLD2, and MLD3 datasets. The results, summarized in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>, clearly demonstrate the superiority of MangoLeafCMDF-FAMNet&#xa0;across all evaluation metrics. Specifically, on MLD1, MangoLeafCMDF-FAMNet achieved a CA of 0.9978, outperforming MangoLeafCMDF-Net with 0.9961, ConvNeXt with 0.9939, and ViT with a considerably lower score of 0.9256. In terms of RCL and PRC, MangoLeafCMDF-FAMNet consistently achieved 0.9978 for both, substantially higher than the competing models. On MLD2, the proposed model maintained its leading performance, reaching a CA of 0.9988, while MangoLeafCMDF-Net, ConvNeXt, and ViT achieved 0.9960, 0.9814, and 0.9176, respectively. Even on the more challenging MLD3 dataset, MangoLeafCMDF-FAMNet maintained its superiority, recording a CA of 0.9943, whereas ConvNeXt achieved 0.9864 and ViT remained at 0.9111. When considering robustness metrics, MangoLeafCMDF-FAMNet attained an MCC of 0.9975 and a Kappa of 0.9975 on MLD1, again consistently exceeding those of the other methods. This outstanding performance was mirrored across MLD2 and MLD3, highlighting not only accuracy but also high reliability and class-wise agreement.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Comparison of MangoLeafCMDF-FAMNet with MangoLeafCMDF-Net, ViT, and ConvNeXt across the MLD1, MLD2 and MLD3.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="center">Metrics</th>
<th valign="middle" colspan="3" align="center">MangoLeafCMDF-FAMNet</th>
<th valign="middle" colspan="3" align="center">MangoLeafCMDF-Net</th>
<th valign="middle" colspan="3" align="center">ViT</th>
<th valign="middle" colspan="3" align="center">ConvNeXt</th>
</tr>
<tr>
<th valign="middle" align="center">MLD1</th>
<th valign="middle" align="center">MLD2</th>
<th valign="middle" align="center">MLD3</th>
<th valign="middle" align="center">MLD1</th>
<th valign="middle" align="center">MLD2</th>
<th valign="middle" align="center">MLD3</th>
<th valign="middle" align="center">MLD1</th>
<th valign="middle" align="center">MLD2</th>
<th valign="middle" align="center">MLD3</th>
<th valign="middle" align="center">MLD1</th>
<th valign="middle" align="center">MLD2</th>
<th valign="middle" align="center">MLD3</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">CA</td>
<td valign="middle" align="center">0.9978</td>
<td valign="middle" align="center">0.9988</td>
<td valign="middle" align="center">0.9943</td>
<td valign="middle" align="center">0.9961</td>
<td valign="middle" align="center">0.9960</td>
<td valign="middle" align="center">0.9919</td>
<td valign="middle" align="center">0.9256</td>
<td valign="middle" align="center">0.9176</td>
<td valign="middle" align="center">0.9111</td>
<td valign="middle" align="center">0.9939</td>
<td valign="middle" align="center">0.9814</td>
<td valign="middle" align="center">0.9864</td>
</tr>
<tr>
<td valign="middle" align="center">RCL</td>
<td valign="middle" align="center">0.9978</td>
<td valign="middle" align="center">0.9988</td>
<td valign="middle" align="center">0.9951</td>
<td valign="middle" align="center">0.9961</td>
<td valign="middle" align="center">0.9959</td>
<td valign="middle" align="center">0.9933</td>
<td valign="middle" align="center">0.9258</td>
<td valign="middle" align="center">0.9173</td>
<td valign="middle" align="center">0.9062</td>
<td valign="middle" align="center">0.9939</td>
<td valign="middle" align="center">0.9816</td>
<td valign="middle" align="center">0.9861</td>
</tr>
<tr>
<td valign="middle" align="center">PRC</td>
<td valign="middle" align="center">0.9978</td>
<td valign="middle" align="center">0.9988</td>
<td valign="middle" align="center">0.9948</td>
<td valign="middle" align="center">0.9962</td>
<td valign="middle" align="center">0.9961</td>
<td valign="middle" align="center">0.9929</td>
<td valign="middle" align="center">0.9298</td>
<td valign="middle" align="center">0.9219</td>
<td valign="middle" align="center">0.9195</td>
<td valign="middle" align="center">0.9940</td>
<td valign="middle" align="center">0.9833</td>
<td valign="middle" align="center">0.9884</td>
</tr>
<tr>
<td valign="middle" align="center">MCC</td>
<td valign="middle" align="center">0.9975</td>
<td valign="middle" align="center">0.9985</td>
<td valign="middle" align="center">0.9933</td>
<td valign="middle" align="center">0.9956</td>
<td valign="middle" align="center">0.9950</td>
<td valign="middle" align="center">0.9906</td>
<td valign="middle" align="center">0.9156</td>
<td valign="middle" align="center">0.8984</td>
<td valign="middle" align="center">0.8978</td>
<td valign="middle" align="center">0.9931</td>
<td valign="middle" align="center">0.9773</td>
<td valign="middle" align="center">0.9843</td>
</tr>
<tr>
<td valign="middle" align="center">Kappa</td>
<td valign="middle" align="center">0.9975</td>
<td valign="middle" align="center">0.9985</td>
<td valign="middle" align="center">0.9933</td>
<td valign="middle" align="center">0.9955</td>
<td valign="middle" align="center">0.9950</td>
<td valign="middle" align="center">0.9906</td>
<td valign="middle" align="center">0.9150</td>
<td valign="middle" align="center">0.8969</td>
<td valign="middle" align="center">0.8967</td>
<td valign="middle" align="center">0.9930</td>
<td valign="middle" align="center">0.9767</td>
<td valign="middle" align="center">0.9842</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To rigorously evaluate whether the performance differences between the proposed MangoLeafCMDF-FAMNet and baseline architectures were statistically significant, we conducted a Tukey Honest Significant Difference (HSD) <italic>post-hoc</italic> test based on the CA values obtained across five folds for each dataset. The results revealed that MangoLeafCMDF-FAMNet significantly outperformed the ViT model across all three datasets (MLD1, MLD2, and MLD3), with adjusted <italic>p</italic>-values consistently below 0.01. Likewise, when compared to the ConvNeXt backbone, the proposed method demonstrated statistically significant improvements in MLD1 (<italic>p</italic> = 0.024) and MLD3 (<italic>p</italic> = 0.038), while achieving a strong yet marginally nonsignificant advantage in MLD2 (<italic>p</italic> = 0.067). These findings underscore the consistent superiority of MangoLeafCMDF-FAMNet in terms of CA, particularly highlighting the added value of its attention-enhanced, cross-modal fusion strategy. Furthermore, the absence of overlapping confidence intervals supports the robustness of the proposed model&#x2019;s improvements and suggests that the observed performance gains are unlikely due to random variation or overfitting.</p>
</sec>
<sec id="s5" sec-type="conclusions">
<label>5</label>
<title>Conclusion</title>
<p>This study introduced MangoLeafCMDF-FAMNet, a novel attention-augmented hybrid DL architecture tailored for robust multi-class mango leaf disease classification. By synergistically combining ConvNeXt and ViT backbones through a CMDF strategy, and further enhancing their output with FAMs, the proposed method effectively captured both fine-grained local patterns and global semantic context. Extensive experiments on three publicly available datasets (MLD1, MLD2, and MLD3) demonstrated the model&#x2019;s superior classification performance across various evaluation metrics, including CA, RCL, PRC, MCC, and Kappa.</p>
<p>The proposed architecture consistently achieved exceptionally high accuracies across all datasets, reaching 0.9978 on MLD1, 0.9988 on MLD2, and 0.9943 on MLD3. These results notably outperformed competing baselines such as ViT, ConvNeXt, and MangoLeafCMDF-Net. Class-wise evaluations further confirmed the model&#x2019;s capacity to distinguish complex and visually similar disease symptoms with high RCL and PRC. Visualizations via t-SNE and confusion matrices affirmed the learned feature separability and robustness against inter-class confusion, while the statistical analyses via Tukey HSD <italic>post-hoc</italic> testing verified that the observed improvements were statistically significant (<italic>p</italic> &lt; 0.05) in most comparisons, particularly over the ViT and ConvNeXt baselines.</p>
<p>Furthermore, the architecture demonstrated a remarkable ability to generalize without overfitting, even on MLD3&#x2014;a dataset characterized by greater class imbalance and symptom variability. The inclusion of FAMs was instrumental in adaptively amplifying disease-relevant features while suppressing irrelevant or redundant information, thereby enhancing class separability across diverse visual domains.</p>
<p>Despite these promising outcomes, the study is not without limitations. First, although the model was tested across three comprehensive datasets, all samples were derived from controlled imaging conditions. Future work should explore the model&#x2019;s applicability to in-field images collected under varying lighting, occlusion, and background clutter. Second, while the proposed model achieved excellent results in disease identification, it currently does not support disease severity estimation, which is crucial for more nuanced decision-making in real-world scenarios.</p>
<p>As future research directions, we aim to extend the MangoLeafCMDF-FAMNet architecture for real-time mobile deployment in smart agriculture systems, incorporate multimodal inputs such as hyperspectral or thermal imagery to improve resilience under environmental variations, and explore the integration of explainability modules to foster model transparency for end-users such as farmers and agronomists. In conclusion, MangoLeafCMDF-FAMNet represents a scientifically grounded, practically scalable, and statistically validated advancement in automated plant disease recognition.</p>
</sec>
<sec id="s6" sec-type="discussion">
<label>6</label>
<title>Discussion</title>
<p>The experimental findings obtained in this study demonstrate that MangoLeafCMDF-FAMNet offers a highly effective solution for the complex task of multi-class mango leaf disease classification. By integrating ConvNeXt and ViT within a unified hybrid architecture and augmenting their capabilities with FAM, the proposed model achieves a refined balance between local feature extraction and global semantic understanding. This synergy enables precise discrimination of disease types, particularly in cases where subtle morphological differences challenge conventional classifiers.</p>
<p>The model&#x2019;s high performance across three publicly available mango leaf datasets confirms its robustness and generalizability under controlled conditions. CA approaching 0.999, alongside strong MCC and kappa, indicate the architecture&#x2019;s capacity to produce stable, reliable, and interpretable predictions. Furthermore, the CMDF mechanism contributes significantly to the enrichment of feature representations, enabling more resilient learning from heterogeneous visual patterns.</p>
<p>Despite these strengths, it is essential to contextualize the results within the scope of the datasets utilized. All datasets in this study were collected under relatively uniform environmental settings, characterized by consistent lighting, minimal occlusions, and simplified backgrounds. While such conditions are favorable for model training and benchmarking, they may not fully reflect the variability encountered in operational agricultural environments. In practice, field images often contain challenges such as partial leaf visibility, shadowing, cluttered scenes, and inconsistent illumination, which may affect the model&#x2019;s generalization performance.</p>
<p>This observation highlights the importance of future work focused on validating the proposed framework using in-field image datasets collected in diverse and uncontrolled environments. Incorporating real-world variability into the training and evaluation pipeline will facilitate the development of more adaptive and field-deployable models. In addition, real-time image acquisition technologies, such as mobile devices and drone platforms, present promising avenues for extending the system toward scalable agricultural decision support.</p>
<p>Another important consideration relates to the clinical utility of disease severity estimation. While the present model effectively identifies the disease type, it does not explicitly address the severity or progression stage of the infection. In real-world agricultural applications, the intensity of disease symptoms is a critical factor influencing treatment strategies and resource allocation. Thus, expanding the model to support ordinal or regression-based predictions for disease severity would significantly enhance its applicability. Although the datasets used in this study did not provide severity annotations, future efforts will focus on curating such datasets and developing models capable of jointly performing disease identification and severity grading.</p>
<p>Moreover, it is worth considering that the absence of diverse environmental conditions and severity-level annotations in the training data may limit the interpretability and practical utility of the current system. Addressing these challenges through targeted dataset development, attention to domain adaptation, and auxiliary prediction tasks will be essential for realizing the full potential of DL-based disease diagnosis in agricultural practice.</p>
<p>In conclusion, MangoLeafCMDF-FAMNet offers a robust and scalable architecture for automated mango leaf disease classification. By leveraging multi-level attention mechanisms and cross-modal fusion, the model provides a strong foundation for high-accuracy plant disease recognition. Future directions should emphasize improving real-world generalization and enhancing the interpretability of the system through disease severity assessment, ultimately supporting the broader goals of precision agriculture and sustainable crop management.</p>
<p>To evaluate the practical applicability of the proposed MangoLeafCMDF-FAMNet in real-world settings, we report key computational characteristics, including model complexity and inference efficiency. The model contains approximately 46.9 million trainable parameters, representing a balanced architectural design that ensures high discriminative power while maintaining computational feasibility. All training and evaluation experiments were performed using the PyTorch DL framework on a standard workstation equipped with an Intel(R) Core(TM) i7&#x2013;9700 CPU and 8 GB of RAM, without access to GPU acceleration. Under this configuration, the average inference time for a single 224&#xd7;224-pixel image was observed to be approximately 180&#x2013;200 milliseconds, depending on system load and batch scheduling. These results suggest that MangoLeafCMDF-FAMNet remains computationally viable even in resource-limited environments, which is particularly beneficial for field-deployable plant disease diagnosis systems where high-end hardware may not be available.</p>
</sec>
</body>
<back>
<sec id="s7" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s8" sec-type="author-contributions">
<title>Author contributions</title>
<p>EE: Formal Analysis, Validation, Methodology, Supervision, Conceptualization, Writing &#x2013; original draft, Software, Writing &#x2013; review &amp; editing, Visualization, Investigation.</p>
</sec>
<sec id="s9" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research and/or publication of this article. This research was financially supported by the Recep Tayyip Erdogan University Development Foundation (Grant number: 02025006018553). We sincerely appreciate their support, which contributed significantly to the completion of this study.</p>
</sec>
<sec id="s10" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s11" sec-type="ai-statement">
<title>Generative AI statement</title>
<p>The author(s) declare that Generative AI was used in the creation of this manuscript.</p>
</sec>
<sec id="s12" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors&#xa0;and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alamri</surname> <given-names>F. S.</given-names>
</name>
<name>
<surname>Sadad</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Almasoud</surname> <given-names>A. S.</given-names>
</name>
<name>
<surname>Aurangzeb</surname> <given-names>R. A.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Mango disease detection using fused vision transformer with convNeXt architecture</article-title>. <source>Computers Materials Continua</source> <volume>83</volume>, <fpage>1023</fpage>&#x2013;<lpage>1039</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.32604/cmc.2025.061890</pub-id>
</citation></ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Arivazhagan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Ligi</surname> <given-names>S. V.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Mango leaf diseases identification using convolutional neural network</article-title>. <source>Int. J. Pure Appl. Mathematics</source> <volume>120</volume>, <fpage>11067</fpage>&#x2013;<lpage>11079</lpage>.</citation></ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bairwa</surname> <given-names>A. K.</given-names>
</name>
<name>
<surname>Singh</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2024</year>). &#x201c;<article-title>Advances in mango leaf disease detection using deep neural networks</article-title>,&#x201d; in <source>2024 international conference on modeling, simulation &amp; Intelligent computing (MoSICom)</source> (<publisher-loc>Dubai, United Arab Emirates</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>75</fpage>&#x2013;<lpage>80</lpage>.</citation></ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>Y. C.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>M. H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>A. C.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>J. A.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Enhanced detection of mango leaf diseases in field environments using MSMP-CNN and transfer learning</article-title>. <source>Comput. Electron. Agric.</source> <volume>227</volume>, <elocation-id>109636</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2024.109636</pub-id>
</citation></ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duan</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>You</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Agricultural text classification method based on ERNIE 2.0 and multi-feature dynamic fusion</article-title>. <source>IEEE Access</source>. <volume>13</volume>, <fpage>52959</fpage>&#x2013;<lpage>52971</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2025.3537277</pub-id>
</citation></ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Erg&#xfc;n</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Deep learning-based multiclass classification for citrus anomaly detection in agriculture</article-title>. <source>Signal Image Video Process.</source> <volume>18</volume>, <fpage>8077</fpage>&#x2013;<lpage>8088</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11760-024-03025-6</pub-id>
</citation></ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Erg&#xfc;n</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>High precision banana variety identification using vision transformer based feature extraction and support vector machine</article-title>. <source>Sci. Rep.</source> <volume>15</volume>, <fpage>10366</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41598-025-95466-0</pub-id>, PMID: <pub-id pub-id-type="pmid">40133576</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Erg&#xfc;n</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Aydemir</surname> <given-names>&#xd6;.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A hybrid BCI using singular value decomposition values of the fast Walsh&#x2013;Hadamard transform coefficients</article-title>. <source>IEEE Trans. Cogn. Dev. Syst.</source> <volume>15</volume>, <fpage>454</fpage>&#x2013;<lpage>463</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TCDS.2020.3028785</pub-id>
</citation></ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ford</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sadgrove</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Paul</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Joint plant-spraypoint detector with ConvNeXt modules and HistMatch normalization</article-title>. <source>Precis. Agric.</source> <volume>26</volume>, <elocation-id>24</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11119-024-10208-y</pub-id>
</citation></ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Spindle thermal error regression prediction modeling based on ConvNeXt and weighted integration using thermal images</article-title>. <source>Expert Syst. Appl.</source> <volume>274</volume>, <elocation-id>127038</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.eswa.2025.127038</pub-id>
</citation></ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gautam</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Ranjan</surname> <given-names>R. K.</given-names>
</name>
<name>
<surname>Dahiya</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>ESDNN: A novel ensembled stack deep neural network for mango leaf disease classification and detection</article-title>. <source>Multimedia Tools Appl.</source> <volume>83</volume>, <fpage>10989</fpage>&#x2013;<lpage>11015</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11042-023-16012-6</pub-id>
</citation></ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hossain</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Sakib</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Abdullah</surname> <given-names>H. M.</given-names>
</name>
<name>
<surname>Arman</surname> <given-names>S. E.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Deep learning for mango leaf disease identification: A vision transformer perspective</article-title>. <source>Heliyon</source> <volume>10</volume>, <elocation-id>e36361</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.heliyon.2024.e36361</pub-id>, PMID: <pub-id pub-id-type="pmid">39281639</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hsiao</surname> <given-names>T. Y.</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>Y. C.</given-names>
</name>
<name>
<surname>Chou</surname> <given-names>H. H.</given-names>
</name>
<name>
<surname>Chiu</surname> <given-names>C. T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Filter-based deep-compression with global average pooling for convolutional networks</article-title>. <source>J. Syst. Architecture</source> <volume>95</volume>, <fpage>9</fpage>&#x2013;<lpage>18</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.sysarc.2019.02.008</pub-id>
</citation></ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kamal</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Sharma</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Gupta</surname> <given-names>P. K.</given-names>
</name>
<name>
<surname>Siddiqui</surname> <given-names>M. K.</given-names>
</name>
<name>
<surname>Singh</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Dutt</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>DVTXAI: A novel deep vision transformer with an explainable AI-based framework and its application in agriculture</article-title>. <source>J. Supercomputing</source> <volume>81</volume>, <fpage>1</fpage>&#x2013;<lpage>32</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11227-024-06494-y</pub-id>
</citation></ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>PD-YOLO: A novel weed detection method based on multi-scale feature fusion</article-title>. <source>Front. Plant Sci.</source> <volume>16</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2025.1506524</pub-id>, PMID: <pub-id pub-id-type="pmid">40265119</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Shangguan</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2025</year>). <article-title>LeafConvNeXt: Enhancing plant disease classification for the future of unmanned farming</article-title>. <source>Comput. Electron. Agric.</source> <volume>233</volume>, <elocation-id>110165</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2025.110165</pub-id>
</citation></ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mahmud</surname> <given-names>B. U.</given-names>
</name>
<name>
<surname>Al Mamun</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Hossen</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>G. Y.</given-names>
</name>
<name>
<surname>Jahan</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Light-weight deep learning model for accelerating the classification of mango-leaf disease</article-title>. <source>Emerging Sci. J.</source> <volume>8</volume>, <fpage>28</fpage>&#x2013;<lpage>42</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.28991/ESJ-2024-08-01-03</pub-id>
</citation></ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mia</surname> <given-names>M. R.</given-names>
</name>
<name>
<surname>Roy</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Das</surname> <given-names>S. K.</given-names>
</name>
<name>
<surname>Rahman</surname> <given-names>M. A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Mango leaf disease recognition using neural network and support vector machine</article-title>. <source>Iran J. Comput. Sci.</source> <volume>3</volume>, <fpage>185</fpage>&#x2013;<lpage>193</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s42044-020-00057-z</pub-id>
</citation></ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Nirob</surname> <given-names>M. A. S.</given-names>
</name>
<name>
<surname>Bishshash</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Siam</surname> <given-names>A.K.M.F.K.</given-names>
</name>
<name>
<surname>Mia</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Khatun</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Uddin</surname> <given-names>M. S.</given-names>
</name>
</person-group> (<year>2024</year>). <source>Mango dataset: A comprehensive resource for agricultural research and disease detection</source> (<publisher-loc>Daffodil International University, Bangladesh</publisher-loc>: <publisher-name>Mendeley Data</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.17632/fn8dgf4hb5.1</pub-id>
</citation></ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Padshetty</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Umashetty</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Agricultural innovation through deep learning: A hybrid CNN-Transformer architecture for crop disease classification</article-title>. <source>J.&#xa0;Spatial Sci.</source>, <fpage>1</fpage>&#x2013;<lpage>32</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1080/14498596.2024.2355225</pub-id>
</citation></ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pahati</surname> <given-names>M.</given-names>
</name>
<name>
<surname>De Jesus</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Reyes</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Villarroel</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Angeles</surname> <given-names>M. R.</given-names>
</name>
</person-group> (<year>2025</year>). &#x201c;<article-title>Detecting mango leaf diseases using google teachable machine for sustainable agriculture</article-title>,&#x201d; in <source>2025 international conference on artificial intelligence in information and communication (ICAIIC)</source> (<publisher-loc>Fukuoka, JapanFukuoka, Japan</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>0611</fpage>&#x2013;<lpage>0614</lpage>.</citation></ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Patel</surname> <given-names>R. K.</given-names>
</name>
<name>
<surname>Chaudhary</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Chouhan</surname> <given-names>S. S.</given-names>
</name>
<name>
<surname>Pandey</surname> <given-names>K. K.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Mango leaf disease diagnosis using Total Variation Filter Based Variational Mode Decomposition</article-title>. <source>Comput. Electrical Eng.</source> <volume>120</volume>, <elocation-id>109795</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compeleceng.2024.109795</pub-id>
</citation></ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pratap</surname> <given-names>V. K.</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>N. S.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Deep learning based mango leaf disease detection for classifying and evaluating mango leaf diseases</article-title>. <source>Fusion: Pract. Appl.</source> <volume>15</volume>, <fpage>261</fpage>&#x2013;<lpage>77</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.54216/FPA.150222</pub-id>
</citation></ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Puranik</surname> <given-names>S. S.</given-names>
</name>
<name>
<surname>Hanamakkanavar</surname> <given-names>S. R.</given-names>
</name>
<name>
<surname>Bidargaddi</surname> <given-names>A. P.</given-names>
</name>
<name>
<surname>Ballur</surname> <given-names>V. V.</given-names>
</name>
<name>
<surname>Joshi</surname> <given-names>P. T.</given-names>
</name>
<name>
<surname>SM</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). &#x201c;<article-title>MobileNetV3 for mango leaf disease detection: an efficient deep learning approach for precision agriculture</article-title>,&#x201d; in <source>2024 5th international conference for emerging technology (INCET)</source> (<publisher-loc>Belgaum, India</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</citation></ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Rahman</surname> <given-names>M. S.</given-names>
</name>
<name>
<surname>Hasan</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Mojumdar</surname> <given-names>M. U.</given-names>
</name>
</person-group> (<year>2024</year>). <source>Mango leaf disease dataset</source> (<publisher-loc>Daffodil International University, Bangladesh</publisher-loc>: <publisher-name>Mendeley Data</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.17632/7ghdbftp54.1</pub-id>
</citation></ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rao</surname> <given-names>U. S.</given-names>
</name>
<name>
<surname>Swathi</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Sanjana</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Arpitha</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Chandrasekhar</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Naik</surname> <given-names>P. K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep learning precision farming: Grapes and mango leaf disease detection by transfer learning</article-title>. <source>Global Transitions Proc.</source> <volume>2</volume>, <fpage>535</fpage>&#x2013;<lpage>544</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.gltp.2021.08.002</pub-id>
</citation></ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rozenfeld</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Kalo</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Naor</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Dag</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Edan</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Alchanatis</surname> <given-names>V.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Thermal imaging for identification of malfunctions in subsurface drip irrigation in orchards</article-title>. <source>Precis. Agric.</source> <volume>25</volume>, <fpage>1038</fpage>&#x2013;<lpage>1066</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11119-023-10104-x</pub-id>
</citation></ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saleem</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Shah</surname> <given-names>J. H.</given-names>
</name>
<name>
<surname>Sharif</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ansari</surname> <given-names>G. J.</given-names>
</name>
</person-group> (<year>2021</year>b). <article-title>Mango leaf disease identification using fully resolution convolutional network</article-title>. <source>Computers Materials Continua</source> <volume>69</volume>, <fpage>3581</fpage>&#x2013;<lpage>3601</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.32604/cmc.2021.017700</pub-id>
</citation></ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saleem</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Shah</surname> <given-names>J. H.</given-names>
</name>
<name>
<surname>Sharif</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Yasmin</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Yong</surname> <given-names>H. S.</given-names>
</name>
<name>
<surname>Cha</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>a). <article-title>Mango leaf disease recognition and classification using novel segmentation and vein pattern technique</article-title>. <source>Appl. Sci.</source> <volume>11</volume>, <elocation-id>11901</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/app112411901</pub-id>
</citation></ref>
<ref id="B30">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Shakib</surname> <given-names>M. M. H.</given-names>
</name>
<name>
<surname>Mustofa</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Ahad</surname> <given-names>M. T.</given-names>
</name>
</person-group> (<year>2024</year>). <source>MLD24: an image dataset for mango leaf disease detection</source> (<publisher-loc>Daffodil International University, Bangladesh</publisher-loc>: <publisher-name>Mendeley Data</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.17632/6dvpywm2m2.1</pub-id>
</citation></ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shehu</surname> <given-names>H. A.</given-names>
</name>
<name>
<surname>Ackley</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Mark</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Eteng</surname> <given-names>O. E.</given-names>
</name>
<name>
<surname>Sharif</surname> <given-names>M. H.</given-names>
</name>
<name>
<surname>Kusetogullari</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>YOLO for early detection and management of Tuta absoluta-induced tomato leaf diseases</article-title>. <source>Front. Plant Sci.</source> <volume>16</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2025.1524630</pub-id>, PMID: <pub-id pub-id-type="pmid">40464016</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname> <given-names>Y. P.</given-names>
</name>
<name>
<surname>Chaurasia</surname> <given-names>B. K.</given-names>
</name>
<name>
<surname>Shukla</surname> <given-names>M. M.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Deep transfer learning driven model for mango leaf disease detection</article-title>. <source>Int. J. System Assur. Eng. Manage.</source> <volume>15</volume>, <fpage>4779</fpage>&#x2013;<lpage>4805</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s13198-024-02480-y</pub-id>
</citation></ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Cotton disease detection based on ConvNeXt and attention mechanisms</article-title>. <source>IEEE Journal of Radio Frequency Identification</source>, <volume>6</volume>, <fpage>805</fpage>&#x2013;<lpage>809</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JRFID.2022.3206841</pub-id>
</citation></ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Varma</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Mate</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Azeem</surname> <given-names>N. A.</given-names>
</name>
<name>
<surname>Sharma</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Singh</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Automatic mango leaf disease detection using different transfer learning models</article-title>. <source>Multimedia Tools Appl.</source> <volume>84</volume>, <fpage>9185</fpage>&#x2013;<lpage>9218</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11042-024-19265-x</pub-id>
</citation></ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Shan</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhen</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Deep learning-based anomaly detection for precision field crop protection</article-title>. <source>Front. Plant Sci.</source> <volume>16</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2025.1576756</pub-id>, PMID: <pub-id pub-id-type="pmid">40438741</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yavuz</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Aydemir</surname> <given-names>&#xd6;.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Olfaction recognition by EEG analysis using wavelet transform features</article-title>,&#x201d; in <source>2016 international symposium on INnovations in intelligent sysTems and applications (INISTA)</source> (<publisher-loc>Sinaia, Romania</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</citation></ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ye</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Lou</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2025</year>). <article-title>Classification&#xa0;of infection grade for anthracnose in mango leaves under complex background based on CBAM-DBIRNet</article-title>. <source>Expert Syst. With Appl.</source> <volume>260</volume>, <elocation-id>125343</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.eswa.2024.125343</pub-id>
</citation></ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>MPCE: A maximum probability based cross entropy loss function for neural network classification</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>146331</fpage>&#x2013;<lpage>146341</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2019.2946264</pub-id>
</citation></ref>
</ref-list>
</back>
</article>